Capstone_Projects
The capstone project has been designed to apply knowledge into practice and to develop and improve critical skills such as problem-solving and collaboration skills. Students are matched with research labs within the UCY community and with industry partners to investigate pressing issues, applying data science areas. Capstone projects aim to give students some professional experience in a real work environment and help enhance their soft skills. These projects involve groups of roughly 3-4 students working in partnership.
The process is the following:
Final assessment is carried out by the company and the supervisor.

_Summer 2026
Spotter Analysis: Detection of Atypical Behaviour Across Supermarket Hierarchies
Background
Retail datasets generated from supermarket transactions are inherently complex, spanning multiple hierarchical levels such as product, brand, category, store, and region. Consumer behaviour across these levels is influenced by factors such as promotions, pricing, seasonality, availability, and external market conditions.
However, retail data often exhibits atypical patterns, including sudden spikes or drops in demand, inconsistencies across hierarchical levels, or unexpected deviations from historical trends. These may indicate data issues, stock disruptions, promotional effects, or genuine shifts in consumer behaviour.
Identifying such patterns across multiple levels of aggregation is essential for ensuring data reliability and generating accurate market insights. Traditional approaches rely heavily on manual inspection and are not scalable across large datasets.
This project introduces the concept of “Spotter Analysis”, a framework designed to automatically detect and highlight atypical behaviour across different levels of a supermarket chain.
Project Description
The objective of this project is to develop a multi-level analytical framework (“Spotter Analysis”) that detects atypical consumption behaviour across supermarket data.
The system will analyse transactional data across different hierarchical levels and identify deviations from expected patterns. These may include anomalies within a single level (e.g. product sales spike) or inconsistencies between levels (e.g. product increase not reflected at category level).
Students will design and implement methods to detect, compare, and interpret these atypical patterns, supporting both data validation and business insight generation.
Data
Students will work with transactional retail datasets that may include:
• product-level sales transactions
• product hierarchy (product → brand → category)
• store and regional structure
• price and promotion indicators
• time series of sales activity
Access to real transactional datasets will be granted under a Non-Disclosure Agreement (NDA). All results presented externally will be aggregated and anonymized.
Analytical Scope (Multi-Level Spotter Analysis)
The project will analyse behaviour across multiple levels:Product Level
• detect abnormal spikes/drops in SKU sales
• identify unusual patterns in individual product performance
Brand Level
• detect inconsistencies across products within the same brand
• identify brand-level shifts in demand
Category Level
• analyse category growth or decline
• detect abnormal category-level behaviour
Store / Region Level
• identify store-specific anomalies
• compare regional consumption patterns
• detect localised disruptions or trends
Cross-Level Consistency Checks
• compare behaviour across levels (e.g. SKU vs category)
• detect mismatches and structural inconsistencies
Tasks
Data Preparation and Exploration
• preprocess transactional datasets
• analyse patterns across hierarchical levels
Anomaly Detection
• develop statistical and machine learning models to detect atypical behaviour
• identify deviations at product, brand, category, and store levels
Cross-Level Analysis
• detect inconsistencies between hierarchical levels
• analyse how anomalies propagate across levels
Evaluation
• assess the effectiveness of detection methods
• validate results against known patterns or eventsDeliverables
• A multi-level “Spotter Analysis” framework for retail data
• Implementation of anomaly detection models across different hierarchical levels
• Documentation describing methodology and findings
• Final report and presentation with aggregated insights
Extra-Mile Tasks
Students may extend the project by:
• developing a dashboard to visualise anomalies across levels
• building an alert system for real-time anomaly detection
• linking detected anomalies to potential drivers (promotions, pricing, availability)
• creating a scoring system to rank the severity of anomalies
Data-Driven Recommendations and Search Intelligence for the Jinius Marketplace
Background:
Jinius Marketplace, developed by the Bank of Cyprus, is a digital e-commerce platform that connects consumers with local retailers across a wide product range. It offers features like centralised checkout, wishlists, promotions, and loyalty rewards. Businesses benefit from simplified logistics and exposure, while shoppers enjoy a more personalised and convenient shopping experience.
The platform generates diverse data: basket contents, order history, wishlists, product descriptions, images, catalogue metadata, and potentially search logs. These data sources present a rich opportunity to apply machine learning and data science techniques to improve recommendations, product discovery, search relevance, and commercial insights.
Previous capstone projects have explored product matching and product attribute extraction. This project will focus on new directions using similar data sources, with particular emphasis on recommendations and search.
Objective:
The goal of this project is to develop intelligent, data-driven features for the Jinius platform using available customer interaction, catalogue, and search-related data. The team can focus on one or both of the following tasks:
Tasks:
1. Product Association, Bundle, and Recommendation Intelligence
Analyse basket, order, wishlist, and catalogue data to identify relationships between products, categories, brands, or retailers, and use these insights to generate recommendations or product bundles.
Examples include “frequently bought together” products, “complete your basket” suggestions, wishlist-based recommendations, seasonal bundles, and complementary product recommendations.
Suggested Techniques: Association rule mining, Apriori, FP-growth, collaborative filtering, matrix factorisation, product embeddings, graph analysis, co-purchase networks, sequence-aware recommendation.
Impact: Improved personalisation, higher basket value, better product discovery, and actionable insights for promotions and merchandising.
2. Search Relevance and AI-Powered Product Discovery
Improve how customers search for and discover products on the marketplace, especially when searches are vague, informal, incomplete, or intent-based. The aim is to help users find relevant products even when their query does not exactly match product titles or descriptions.Examples include semantic search, natural-language product discovery, query expansion, ranking improvement, identification of poor or no-result searches, and alternative product suggestions.
Suggested Techniques: NLP, semantic search, text embeddings, multimodal embeddings, vector databases, learning-to-rank, transformer models, CLIP, retrieval-augmented generation, large language models.
Impact: Improved search relevance, reduced customer frustration, better product visibility, and increased search-to-purchase conversion.
Data:
Basket and order history, anonymised
Product catalogue, text, images, and metadata
Wishlists, anonymised
Search logs
Enhancing Entity Benchmarking Through Public Data Integration in Power BI
Objective:
This project focuses on enriching existing Microsoft Power BI dashboards with publicly available external data. The goal is to identify, connect to, and integrate relevant public data sources into Power BI — using APIs, web URLs, and other data connectivity methods — in order to enhance the benchmarking capabilities of individual entity dashboards. By overlaying external reference points onto internal data (e.g., comparing an entity’s average salary against national averages, or contextualising performance within broader economic and industry trends), entities will be able to assess their position relative to the market, discover insights, and make more informed decisions.
Scope of Work
Participants will:
• Research and identify relevant public data sources (e.g., government statistics portals, open data platforms, industry bodies, international organisations)
• Connect to external data using APIs, web connectors, or other available ingestion methods/tools (python/R/SQL)
• Transform and model the acquired data so that it integrates seamlessly with existing dashboard structures
• Design and build visualisations that enable clear, intuitive comparison between an entity’s internal metrics and external benchmarks
• Document data sources, refresh schedules, transformation logic, and any limitations
Expected Outcomes:
• A set of Power BI data connections to curated public sources, ready for ongoing use
• Enhanced Power BI reports and dashboards that provide entities with external context and benchmarking insights
• Documentation enabling the team to maintain and extend the solution beyond the project period
Bonus 1: Implementation of AgenticAI tools/add-ons to enhance the quality and extend the scope of data injected into Power BI.
Bonus 2: Develop a robust data processing layer using python/R/SQL, to prepare analysis-ready datasets suitable for both statistical exploration and visualisation.
What Participants Will Gain:
• Hands-on experience with Microsoft Power BI (data connectivity, Power Query, data modelling, DAX, and visualisation), python/R and SQL.
• Practical exposure to working with APIs and public data sources• Understanding of how external data adds analytical value in a professional environment
• Experience working within a corporate team on a deliverable with real business impact
Title: Road Assistance Delay Time Prediction
Description: The project aims to address the challenge of accurately predicting the delay time for Road Assistance services. In motor insurance operations, response time is a critical service quality indicator directly affecting customer satisfaction and operational efficiency. However, delays can vary significantly depending on factors such as location, time of day, traffic conditions, weather, vehicle type, and service provider availability.
This project offers an opportunity to apply advanced data science and machine learning techniques to develop a predictive model capable of estimating delay times based on operational and contextual data. The objective is to provide actionable insights that can help improve dispatch decisions, optimize resource allocation, and enhance customer experience.
Deliverables:
• A predictive machine learning model capable of estimating the delay time from the initial service request to the arrival of the service vehicle.
• A well-documented analysis pipeline including data preprocessing, exploratory data analysis (EDA) with visual insights, feature engineering, model development, and validation.
• A short final presentation outlining the project and summarizing key findings, along with recommendations on how the model can be integrated into the company’s dispatch system to improve customer communication.
Available Data:
Students will work with a comprehensive dataset from historical road assistance logs. The dataset includes features such as request timestamps, pickup and drop-off coordinates, service vehicle types, and incident categories.
Instructions:
Students are encouraged to follow the full data science lifecycle and apply the tools and techniques acquired during their studies. They are also expected to maintain regular communication with the company’s Data Team to validate findings and stay aligned with project goals. The project emphasizes team collaboration and offers students the opportunity to work with real-world, time-sensitive data in a practical business context.
Main Challenges:
1. Handling missing, inconsistent, or erroneous records, which may require cross-referencing with external data sources such as traffic or weather APIs to be identified by the students.
2. Accounting for high variability in delay times driven by rare or extreme events, such as peak demand periods or large-scale incidents.
3. Applying innovative modeling approaches to accurately predict delay times across diverse geographic areas and incident types with limited historical coverage.
4. Ensuring the model architecture is optimized for quick inference to provide near-instant updates to customers.
5. Deriving accurate delay times from provider event timestamps, which may be delayed, recorded in batches, or missing, requiring careful data cleaning and robust target definition.
Predictive Analytics and Clinical Decision Support in a Hospital Setting
Project Description
This project involves the development of a Machine Learning model to support clinical and operational decision-making at the German Medical Institute (GMI), a private hospital based in Limassol, Cyprus. Depending on the selected focus area — agreed upon with the student at the outset — the model may address challenges such as patient readmission prediction, length-of-stay estimation, or appointment no-show forecasting. More specifically, with regards to the length-of-stay estimation, students will be provided with retrospective patient records related to their admission that includes primary and secondary diagnoses, procedure codes (i.e., what surgery or other procedure brought them to the hospital), their age, gender, nationality and other relevant demographic information to develop a model that accurately predicts how long a patient is expected to stay in the hospital based on the above information. The data required for the development of the above mentioned models will be anonymised and provided to the students, ensuring compliance with GDPR and Data Sharing Agreements that will have already been put in place.
By analysing anonymised patient and operational data, the project aims to generate actionable insights that can support clinical staff, enhance hospital resource planning and ultimately improve patient outcomes. Students will gain hands-on experience in healthcare data preprocessing, feature engineering, and the application of traditional and advanced Machine Learning techniques, while navigating the unique ethical, regulatory, and data privacy considerations that define real-world healthcare environments.
Expected Deliverables
At the conclusion of the project, the student is expected to provide the following:
• A cleaned and well-documented dataset, including all preprocessing and feature engineering steps
• An exploratory data analysis (EDA) report describing key patterns and clinical/operational insights
• A trained and evaluated ML model with documented methodology, performance metrics, and limitations
• A final written report suitable for presentation to both technical and non-technical stakeholders
• A short presentation of findings to the GMI AI team
Extra-Mile Tasks
For students who demonstrate exceptional performance and are willing to go above and beyond, the following additional tasks are encouraged:
• Development of an interactive dashboard to visualise model outputs for clinical or operational end-users
Enhanced Fraud Modelling with LLM based Fraud Analyst assistant
Project Goal
Develop a machine learning model to detect fraudulent card transactions and extend it with an LLM-based chatbot that can generate structured fraud analysis for individual transactions using model outputs, entity history, and data insights.
Problem Statement
Fraud detection requires balancing accurate classification with clear, actionable insights. This project focuses on building a reliable detection model and enhancing it with an AI agent that can interpret results and provide human-readable explanations for transaction-level decisions.
Data Sources
Card transaction dataset with fraud indicators
Tasks
• Perform exploratory data analysis to understand patterns and relationships
• Develop and evaluate a fraud detection model
• Incorporate entity-level historical context
• Build an LLM-based agent to act as a fraud analyst assistant
Deliverables
1. Trained Model
A fraud detection model with performance evaluation
2. LLM Fraud Analysis Agent
An AI agent that can generate insights for individual transactions using model outputs and contextual data
3. Technical Documentation
Summary of approach, findings, model performance, and agent design
4. Presentation
Overview of the project, key insights, and outcomes
Grant Thornton has put forward two proposals with a similar scope. The final choice will depend on data availability and student interest.
Proposal 1
Project Title: Building a Custom Legal Chatbot Using LLMs and Cypriot Legislation
This project focuses on transforming how users interact with Cypriot legal texts by developing a custom chatbot powered by large language models (LLMs). The goal is to make legislation and court decisions, available through www.cylaw.org, more searchable and understandable.
Students will design a prototype that combines semantic search with an LLM-powered question-answering system, allowing users to ask legal questions in plain language and receive accurate, well-sourced responses.
Key deliverables include a semantic retrieval system that understands the meaning of user queries and a chatbot interface that generates summaries or explanations of legal content. The project will use publicly available legal documents from CyLaw, with the option to focus on a specific legal domain such as labour or environmental law to ensure a manageable and meaningful scope.
Proposal 2
Project Title: “RAG-Based Policy Assistant (Prototype)”
Project Description
This project focuses on the design and implementation of a prototype Retrieval-Augmented Generation (RAG) agent using Python, aimed at improving access to internal organisational knowledge. Students will develop a streamlined framework that combines document retrieval techniques with large language models to enable users to query a limited set of company policies— such as HR, compliance, or cybersecurity documents—through a conversational interface. The emphasis will be on building a working end-to-end pipeline, ensuring that responses are grounded in source documents and include references to improve transparency and trust.
Following the development of the core framework, the solution will be applied to a selected subset of real company policies to demonstrate practical usability. The final output will include a functional prototype with a simple chat interface and a proof-of-concept integration into Microsoft Teams or a comparable environment. The project will prioritise correctness, usability, and architectural clarity over full production deployment, while also introducing key concepts such as responsible AI, data privacy considerations, and system scalability at a conceptual level.
Expected Deliverables
1. Solution Design
· High-level architecture of the RAG system (ingestion, retrieval, generation)
· Technology stack justification (Python, vector DB, LLM choice)
2. Core RAG Prototype (Python)
· Document ingestion and preprocessing pipeline (limited dataset)· Embedding generation and storage
· Retrieval and LLM-based response generation
3. Policy Dataset Preparation
· Structured processing of a small number (1-2) of company policy documents
· Chunking and metadata tagging approach
4. Chat Interface (Prototype)
· Simple UI (e.g., Streamlit or lightweight web app)
· Ability to query documents and receive contextual responses
5. Response Quality Enhancements
· Basic prompt engineering
· Inclusion of source citations in responses
6. Evaluation Approach
· Simple testing framework (manual test cases)
· Basic metrics: relevance, accuracy, hallucination observations
7. Microsoft Teams Integration (PoC)
· Basic bot or webhook-based integration (non-production)
· Demonstration of how the assistant can be accessed in a workplace tool
8. Governance & Risk Considerations (Conceptual)
· Short paper covering:
o Data privacy considerations
o Access control concepts
o Risks of LLM usage (e.g., hallucination, bias)
9. Documentation
· Technical setup guide
· User guide for interacting with the assistant
10. Final Presentation & Demo
· Demonstration of the working prototype
· Key findings and limitations
11. Final Report
· Methodology, implementation steps, challenges
· Recommendations for scaling to production
Next-Generation Customer Segmentation Using Transformer-Based Encoders
Project Description
Customer segmentation is a long-standing problem in banking analytics. Traditionally, banks segment customers using manually engineered tabular features and classical clustering techniques (e.g. k-means, hierarchical clustering). While these methods are simple and interpretable, they may fail to capture complex relationships in modern banking data.
Recent advances in encoder-based and transformer-inspired models for tabular data allow customer representations to be learned rather than hand-crafted. These models generate dense embeddings that may encode latent financial and behavioural patterns.
This capstone explores a core empirical question, not a predefined conclusion:
Do encoder-based representations actually improve customer segmentation in banking, compared to traditional approaches?
Students will design experiments to test this claim rigorously, using tabular banking data.
What You Will Work On
You will solve the same segmentation problem twice, using two fundamentally different paradigms.
1. Traditional Segmentation (Baseline)
You will:
• Work with tabular banking data (customer attributes, balances, product holdings, transaction aggregates)
• Engineer and select features
• Apply classical clustering techniques
• Interpret and profile resulting customer segments
This represents how segmentation is typically done today in many banking environments.
2. Next Generation Encoder-Based Segmentation
You will:
• Use the same tabular banking data
• Fine-tune or train an encoder-based model adapted for tabular inputs
• Generate customer embeddings
• Perform clustering in the learned latent space
• Analyse whether and how the resulting segments differ from the baseline
The encoder replaces manual feature engineering, not the clustering step.What You Are Expected to Prove (or Disprove)
This is a research-style project, not a demo.
You must:
• Compare the two approaches fairly (same data, same number of clusters where applicable)
• Use objective metrics, stability tests, and qualitative analysis
• Critically assess:
o Segment coherence
o Interpretability
o Stability
o Complexity vs. benefit
Negative or neutral results are acceptable if they are well-argued.
Data and Tools
• Domain: Banking
• Data Modality: Fully tabular
• Models:
o Classical clustering methods
o Encoder-based models (including fine-tuned large or pre-trained models adapted to
tabular data)
• Programming: Python-based data science/ AI stack
The focus is methodology and reasoning, not production deployment.
What You Will Deliver
1. Final Technical Report
o Clear problem definition
o Description of both segmentation pipelines
o Experimental setup and evaluation
o Critical comparison and conclusions
o Explicit discussion of limitations
2. Code Repository
o Reproducible experiments
o Clear separation between baseline and encoder-based approaches
o Documented assumptions and parameters
3. Final Presentation
o Clear answer to the core research questiono Evidence-based conclusions
o Honest discussion of trade-offs
High Level Project Milestones
1. Problem Definition & Data Understanding
2. Baseline (Traditional) Segmentation
3. Encoder Selection & Fine-Tuning
4. Embedding-Based Segmentation
5. Comparative Evaluation
6. Final Synthesis & Conclusions
Background:
RiskMatrix is an international Credit Risk Advisory and Software Services firm that has been appointed by Moody’s Analytics as its exclusive representative relating to credit risk software and services offering across the number of territories including Cyprus, Greece, the Middle East, Egypt and East Africa. We serve more than 130 banking groups.
RiskMatrix’ Credit Analytics team provides advisory services to many of RiskMatrix’ customers. These services centre around the measurement of credit risk.
A key class of credit risk models comprise models to determine probabilities of default (PDs).
RiskMatrix uses its own proprietary approach to build such models. The approach is structured and designed to avoid over- and under-fitting of the models. This is important as the banks supported by RiskMatrix usually cannot provide extensive historical information for their business customers and therefore data limitations pose a significant challenge to building robust models.
The Capstone Project:
Recent developments in machine learning and data science provide alternative approaches to developing PD models. The Capstone will explore how such approaches perform in a real-world domain and to understand the reasons why such approaches may or may not achieve better outcomes than more traditional approaches.
Objective:
1) Use the XG-Boost approach to develop a PD model using financial statement data for a “typical” dataset comprising business customers of a bank. Assess the performance of the model, the strengths and limitations of the approach and the impact of different choices in parameterisation and structure of the model. Compare XG-Boost to a benchmark modelling approach (to be agreed as part of the Capstone Project).
2) Perform an assessment of the resulting model structure and behaviour. This will involve using an approach such as SHAP to provide an explanation and understanding of the model produced and identify weakness of the model. Compare SHAP (or another selected approaches) to alternate “standard” approaches to assess the model behaviour (e.g. sensitivity analysis). Assess the appropriateness, suitability and strengths and limitations of SHAP for this purpose.
3) Optional (if time permits): Assess how dataset size affects the quality of the model produced
4) Optional (if time permits): Compare XG-Boost to an alternative contemporary ML approach, for example, CatBoost.
Dataset:
The data will comprise a subset of “typical” data used in a modelling project by RiskMatrix. This will include candidate model input values such as financial ratios together with outcome values (whether a company defaulted within twelve months of the observation point).
Approach:
A senior member of RiskMatrix’ Credit Analytics team will partner with the students for this project.
Using this approach, RiskMatrix can assist the students to gain a deeper understanding of the data and its meaning, provide expertise in “best practices” in PD model development, and provide guidance in the interpretation of the results and the appropriateness and effectiveness of the models produced.
The focus will be on understanding how the selected techniques work on a real-world modelling task, with real world problems and complications. The emphasis will be on gaining deeper understanding as opposed to building a high performance model.
_Summer 2025
Data-Driven Insights for the Jinius Marketplace
Background:
Jinius Marketplace, developed by the Bank of Cyprus, is a digital e-commerce
platform that connects consumers with local retailers across a wide product range.
It offers features like centralised checkout, wishlists, promotions, and loyalty
rewards. Businesses benefit from simplified logistics and exposure, while shoppers
enjoy a more personalised experience.
The platform generates diverse data: basket contents, order history, wishlists,
product descriptions, images, and catalogue metadata. These data sources present
a rich opportunity to apply machine learning and data science techniques to
improve recommendations, product search, and commercial insights.
Objective:
The goal of this project is to develop intelligent, data-driven features for the Jinius
platform using available customer interaction and catalogue data. The team can
focus on one or more of the following tasks:
Tasks:
Data:
LLM-Powered Chatbot for Regulatory and Policy Question Answering
Background:
Professionals operating in financial, legal, and compliance domains often rely on dense PDF documents containing regulatory and policy information (unstructured datasets). These documents are not easily searchable or accessible through natural language, leading
to inefficiencies, misinterpretation, and time-consuming manual analysis. With recent advancements in Large Language Models (LLMs) and vector search technologies, it is now possible to build intelligent systems capable of understanding, retrieving, and contextualizing
information from such complex documents.
This presents an opportunity to work on a project with real-world impact, utilizing real datasets to address challenges faced by professionals daily. It allows exploration of an emerging area of AI that is in high demand, promising significant advancements and efficiencies in information handling and interpretation.
Objective:
Data:
Data will consist of PDF reports, specifically focusing on regulatory and policy information with regards to the financial, legal, and compliance domains. These reports are typically dense and lengthy, containing valuable insights that need to be efficiently parsed and indexed for effective retrieval by the chatbot.
Desired Outcome:
Bonus 1: Custom UI Interface: Develop a user-friendly, custom UI interface for intuitive interaction with the chatbot. This interface will enhance user experience, allowing seamless navigation through queries and responses.
Bonus 2: Agentic RAG – Implement an Agentic RAG approach to handle more complex and nuanced questions. This advanced capability will allow the chatbot to navigate not just simple queries, but also multi-faceted questions that require deeper analysis
and understanding.
Used Car Price Prediction using Classified Ads Data
Project Description:
The proposed project aims to address the challenge of accurately estimating the selling price of used cars. The second-hand automotive market is highly dynamic, with vehicle value influenced by a wide range of factors including technical specifications, age, mileage, and seller profile. This project offers an opportunity to leverage advanced data science techniques to develop a robust, data-driven car valuation model.
Deliverables:
Available Data:
Students will work with a large-scale dataset collected from various classified ads platforms.
The dataset includes detailed features such as brand, model, engine specifications, mileage,
and selling price, among others. The data also supports analysis over time and by region.
Instructions:
Students are encouraged to follow the full data science lifecycle and apply the tools and
techniques they have acquired during their studies. They are also expected to maintain
regular communication with the company’s team to validate findings and to stay aligned with
project’s goals. The project emphasizes team collaboration and gives students the
opportunity to work with real-world data in a practical business context.
Main Challenges:
Customer Experience Analytics using Machine Learning Techniques
Project Description:
This project involves developing a Machine Learning model to support and optimize call center operations for a utility provider. Depending on the selected topic, the model may focus on forecasting call volumes, predicting first-call resolution, classifying call topics, or clustering agent performance. By analyzing historical call data, the project aims to uncover actionable insights that can improve efficiency, resource planning, and customer service. This will provide students with practical experience in data preprocessing, feature engineering, and traditional Machine Learning techniques, while addressing real-world challenges in the context of customer support operations.
Resources: Our company will provide
Deliverables:
Extra-Mile Tasks: Development of a simple user interface (UI) that allows users to interact with the model and view live outputs (e.g., predictions, classifications, or insights) in a user-friendly way.
Fraud Detection Model
Project Goal:
Develop a machine learning model to detect fraudulent card transactions. You will explore the dataset, identify key fraud indicators, assess sources of bias, and create and calibrate a fraud detection model using appropriate performance metrics.
Problem Statement:
Card fraud detection is critical for protecting businesses and customers from monetary loss. However, effective fraud detection models must balance sensitivity and specificity in a highly imbalanced dataset. This project aims to extract meaningful features, analyse fraud patterns, identify relationships and biases, and develop a robust model for detecting fraudulent transactions.
Data Sources:
Card transaction dataset with fraud indicators.
Tasks:
Deliverables:
Building a Custom Legal Chatbot Using LLMs and Cypriot Legislation
Project Description:
This project focuses on transforming how users interact with Cypriot legal texts by developing a custom chatbot powered by large language models (LLMs). The goal is to make legislation and court decisions, available through www.cylaw.org, more searchable and understandable. Students will design a prototype that combines semantic search with an LLM-powered question-answering system, allowing users to ask legal questions in plain language and receive accurate, well-sourced responses. Key deliverables include a semantic retrieval system that understands the meaning of user queries and a chatbot interface that generates summaries or explanations of legal content. The project will use publicly available legal documents from CyLaw, with the option to focus on a specific legal domain such as labour or environmental law to ensure a manageable and meaningful scope.
Understanding Cyprus’s real estate market
Project Description:
As part of its Real Estate offering, Deloitte would like to build an end-to-end real estate asset monitoring and analysis infrastructure through capturing data on a recurring basis from various publicly and non-publicly available sources. These datasets will be subsequently used to populate Deloitte’s internally developed PowerBI tool, which the business is solely going to use for internal purposes.
Objective:
Available Data:
Deliverables:
Assessment of actual healthcare needs of National Healthcare System (GeSY) beneficiaries and identification of areas of over provision and/or under provision based on the current utilization of healthcare services
Background:
The introduction of the General Healthcare System (GeSY) in Cyprus in 2019 has improved significantly the access to healthcare services for the whole population and it has led to the financial protection of all patients. As a result, we have seen a notable increase in demand for healthcare services and, at the same time, a drastic reduction of out-of-pocket payments and a minimisation of unmet healthcare needs. Even though the System has now reached a steady state, the growth of the utilisation of healthcare services cannot be accurately assessed vis-à-vis the actual healthcare needs of the beneficiaries, leading potentially to some areas being overserved and/or others remaining underserved. There is therefore a growing need to assess whether the current distribution and utilisation of healthcare services, align with the actual needs of the beneficiaries.
Factors such as demographic changes and epidemiological factors influence healthcare needs. Mismatching of actual needs against service provision leads to oversupply or undersupply of services. On one hand, oversupply of services leads to the provision of unnecessary treatments, inefficient allocation of resources and increased healthcare expenditure, while on the other hand, undersupply of services leads to unmet needs, poorer health outcomes and increased healthcare expenditure in the long term. In order to maximise the healthcare system’s effectiveness, it is crucial to understand and quantify the actual healthcare needs of the beneficiaries.
Project Description:
This project aims to estimate the actual healthcare needs of beneficiaries for specific services/ procedures/products and compare them against current utilisation of healthcare services/products in order to identify areas of over provision or under provision. More specifically areas that should be addressed are:
The findings of this project will help identify the areas of over provision and/or under provision of healthcare services, in order to support evidence-based strategic decision-making to align healthcare provision to actual patient needs, thus maximising the efficiency of the System.
Data:
The Health Insurance Organisation (HIO) will provide data available in its IT system, both for beneficiaries’ demographics (e.g. analysis by age, gender, district etc.), as well as for utilisation of services/products (e.g. number of outpatient visits, procedures, diagnostic and lab tests, inpatient incidents etc.) per category/subcategory and/or specific services/products.
Other demographic and socioeconomic data can be obtained from the Statistical Service.
Benchmark data for the utilisation of services in other countries can be obtained from the database of Eurostat, OECD etc.
Customer Churn Prediction and Proactive Retention Strategy for Telco Provider
Project Description:
This project will use machine learning to predict which customers are likely to churn based on usage, billing, and interaction data. It will also generate actionable recommendations, such as targeted offers or service improvements to help the telco provider proactively retain high-risk customers and reduce churn rates.
The outcome will empower the telco provider with data-driven insights to optimize customer retention initiatives and enhance overall customer satisfaction.
Resources:
Our company will provide:
Desired Deliverables:
_Summer 2024
Internal Fraud Detection System
Description
The goal of this project is to develop an Internal Fraud Detection System. The system aims to identify and flag suspicious activities suggestive of internal fraud by detecting abnormal patterns in employee behaviour and actions. Using unsupervised learning techniques, students will develop models to identify anomalies. This project includes data collection and analysis, model development and model serving.
Milestones
Possible set of data sources
Curating the Emotional Journey: A Museum Music Recommendation System with Emotion Recognition
Project Goal:
This capstone project challenges you to design and develop a data science application that personalizes the museum visitor experience through music. The application will leverage emotion recognition and stress level technologies to analyze visitor emotional state from video footage and heart rate signal to generate appropriate background music to fit on their visiting experience on the thematic era of the exhibit.
Problem Statement:
Museums strive to create engaging and impactful experiences for visitors. Music plays a significant role in setting the atmosphere and influencing visitor emotions. However, traditional, static background music may not effectively cater to the diverse emotional states and interests of visitors navigating various exhibits. This project aims to bridge this gap by developing a music generation system tailored to human individual emotional state.
Data Sources:
Technical Stack:
Project Deliverables:
Success Criteria:
A successful project will demonstrate the following:
Description:
As part of its Real Estate offering, Deloitte would like to build an end-to-end real estate asset monitoring and analysis infrastructure through capturing data on a recurring basis from various publicly and non-publicly available sources. These datasets will be subsequently used to populate Deloitte’s internally developed PowerBI tool, which the business is solely going to use for internal purposes.
Objective:
Data:
Deliverables:
Title: Telco Network Failure Prediction
Duration: 2 months
Project Description:
This project aims to leverage advanced data science techniques to anticipate and mitigate network failures in telecommunication systems. Predicting network failures in telecommunication systems is crucial for maintaining service quality, minimizing downtime, and optimizing resource allocation. By analyzing historical data such as network performance metrics, environmental conditions, equipment status, and maintenance logs, the data science team will develop predictive models to anticipate potential failures before they occur. This will allow the Telco provider to proactively address issues, improve network reliability, and enhance customer satisfaction.
Resources:
Our company will provide:
Desired Deliverables:
Measuring the Impact of Wildfire Risk on Property Market Prices
Objective:
Develop a predictive model to measure the impact of wildfire risk on property market prices.
Data:
The project will utilize a combination of the following datasets:
Deliverables:
Resources:
Expected Outcome:
The project aims to provide a robust predictive model that accurately measures the impact of wildfire risk on property market prices. Interactive visualization presentation of results is expected in the form of a WebApp (through Shiny, Flask etc.). The findings will help users to make informed decisions regarding property investments and risk management.
Machine Learning for Fraud Detection in Insurance Claims
Background
Insurance claim fraud is a global issue that drains billions of dollars from the industry each year, escalating costs for both insurers and policyholders. In Greece, for instance, fraudulent claims tally up to €200 million annually, accounting for about 10% of all claim payouts. These challenges are not unique to Greece but are mirrored worldwide, necessitating a unified and innovative approach across the sector to tackle this pervasive problem.
Insurance claim fraud is defined as any deliberate and misleading act or omission by any individual or legal entity aimed at gaining an unlawful financial benefit for themselves or facilitating such gain for others within the context of a valid or invalid insurance contract.
The need for advanced data analytics and machine learning in fraud detection is becoming increasingly critical as fraudsters employ more sophisticated methods to circumvent traditional detection systems. Globally, insurers are turning to technology to gain an edge against these tactics, utilising big data and predictive analytics to spot irregularities and inconsistencies in claim submissions. This technological shift represents a move from reactive to proactive fraud management, where potential frauds are flagged and investigated before they result in financial loss.
Moreover, the ongoing digital transformation in the insurance industry highlights the importance of continuous innovation in fraud detection systems. The integration of AI and machine learning technologies enhances fraud detection accuracy and speeds up the process, allowing insurers to handle claims more efficiently while reducing the chances of fraudulent claims slipping through the net.
Engaging with this project provides an opportunity to contribute to significant advancements in improving this procedure and applying cutting-edge data science techniques to real-world problems. This initiative is an excellent initiative for aspiring data scientists to make a tangible difference, helping to shape the future of an industry on which everyone relies.
Objectives
Briefly, the core objectives of this project are defined as:
Methodology
The principal methodology pipeline will consist of:
Resources
All the necessary resources for the completion of the project will be delivered, such as:
Title: Customer Behavior prediction
Project Description:
The project focuses on analyzing customer data of a Telecommunication company, specifically related to their mobile plans. The data contains information related to customers such as monthly payments, mobile usage, contract information and competitive data from other telecom vendors. The objective is to develop a methodology to process the data, extract meaningful insights through analysis and create new features that will be used to build predictive models. The models will be trained to predict customer behavior that will lead to switch to a different mobile plan. This project will provide an opportunity to gain practical experience in Data Analysis, Feature Engineering, and Machine Learning, while also gaining valuable insights into solving real-world problems in the telecommunications industry.
Resources:
Our company will provide:
Desired Deliverables:
Health Data Analysis and/or Pricing model development
Project Description:
This dynamic and interactive project will provide students with hands-on experience applying their knowledge in data science to the actuarial field, particularly in data analysis and optimization as well as Health Insurance Pricing model development. It features a unique collaboration between Milliman’s local and international teams, ETHNIKI, THE, HELLENIC GENERAL INSURANCE CO. S.A., one of Greece’s largest insurance companies, and the University of Cyprus that successfully promotes the project.
To get a flavor of what the project will include, have a look into the below. (Note that this is not an exhaustive list of tasks):
*The intention will be to anonymize all data utilized for this Capstone project.
Resources:
UCY students will have the opportunity to work with real life Insurance market data that will include information regarding historical claims with a wealth of parameters around each claim. Milliman will provide guidance and support throughout the duration of the project and will ensure that materials provided to the students are adequate to perform the tasks.
Desired Deliverables:
Scope
Today, accurate flight delay prediction is among the most critical and challenging problems when scheduling and managing flights for all involved aviation data value chain stakeholders. As part of the capstone project, the flight delay prediction problem will be addressed from the perspective of an airport and for a day-ahead time horizon in terms of:
The flight delay forecasts need to be accompanied by appropriate explanations in order to shed light into the provided forecasts, so Explainable AI (XAI) techniques (i.e. for feature relevance and counterfactual explanations) need to be leveraged.
The aviation data that will be used as part of the capstone project are confidential flight data from an airport in Europe.
Deliverables
D1. XAI Models for the Flight Delay Forecasting (Classification) & Problematic Flight Forecast Index – end of June 2024
D2. XAI Models for the Flight Delay Forecasting (Regression) & XAI Models Updates for the Flight Delay Forecasting (Classification) & Problematic Flight Forecast Index – end of July 2024
_Summer 2023
As part of its risk assessment, the bank is required to predict the returns from the sales of its real estate collaterals. One of the parameters that determines those returns is the ‘recovery rate’, i.e., the % of the property market value that the Bank will recover by selling the property.
Objective: develop a predictive model for recovery rate parameter of collaterals, using the Bank’s most recent internal database.
Data: contains information on collateral sales transactions (e.g., open market value, sale price, sale date) and property characteristics (e.g., location, property type, size), for historic collaterals onboarded from 2016 onwards. Publicly available data from CY Statistical Service related to property price indexes can be used to supplement the analysis.
Deliverables:
Short description: The project aims to develop and evaluate a solution to automate the processes of model validation and identify variables that canchallenge and enhance the performance of the Bank’s behavioural credit models.
Objective: The objective of the project is to improve the efficiency and effectiveness of the model validation function in the Bank, and to ensure compliance with validation unit’s internal procedures and methods.
Data: The project will use the following data sources:
Deliverables: The expected outcomes and deliverables of the project are:
Project Title: Cancer Prevention AI tools
Project Description: The Cancer Prevention AI tools project is a cross-functional capstone project that aims to develop a mobile application that monitors the daily living of people and suggests new food, activity, and lifestyle routines to reduce the percentage of cancer disease. The application will be built by a team of two or three students with expertise in data analytics (visual, textual, multisensory signals analysis) and one in business development (optional).
The data analytics student(s) will be responsible for developing AI models that will analyse user data, including daily food intake, physical activity, mood estimation and lifestyle habits, to provide personalized recommendations for reducing the risk of cancer. The AI models will use machine learning algorithms to identify patterns and correlations in the data and suggest changes to the user’s routine accordingly. The Data Analytics student(s) should have expertise in programming languages such as Python, and be familiar with relevant libraries and tools such as NumPy, Pandas, Scikit-learn, TensorFlow, Keras, and Matplotlib. They should also have experience in machine learning algorithms, deep learning models, and data visualization techniques.
Overall, the Cancer Prevention AI tools project will provide a valuable tools for people to monitor their daily habits and receive personalized recommendations for reducing their risk of cancer. The project will also provide valuable experience for the student team members in data analytics, and business development, as well as a tangible product to showcase their skills to potential employers (optional).
Deliverables:
AI Models: The project should deliver appropriate AI models to analyse user data, including daily food intake, physical activity, and lifestyle habits, to provide personalized recommendations for reducing the risk of cancer. The model should use machine learning algorithms to identify patterns and correlations in the data and suggest changes to the user’s routine accordingly.
Business Plan: The project should include a business plan that outlines the target market for the AI tools, the competitive landscape, and the marketing and sales strategy. The plan should also include financial projections and revenue models.
Technical Documentation: The project should include technical documentation that describes the AI tools design, and functionality, as well as instructions for installing and running the AI tools.
Presentation: The project team should prepare a final presentation that showcases the AI models, and the business plan. The presentation should include a demonstration of the AI tools’ functionality, and the business plan’s revenue projections.
|
Project title: |
Use of ML Techniques to Enable Intrusion Detection in IoT Networks |
|
Background: |
Networks are vulnerable to costly attacks. Thus, the ability to detect these intrusions early on and minimize their impact is imperative to the financial security and reputation of an institution. There are two mainstream systems of intrusion detection (IDS), signature-based and anomaly-based IDS. Signature-based IDS identify intrusions by referencing a database of known identity, or signature, for each of the previous intrusion events. Anomaly-based IDS attempt to identify intrusions by referencing a baseline or learned patterns of normal behavior. Under this approach, deviations from the baseline are considered intrusions. |
|
Project aims:
|
This MSc project will utilize existing work on Lightweight Intrusion Detection for Wireless Sensor Networks (using BLR, SVM, SOM, Isol. Trees) and extend it, considering the characteristics of IoT networks, new attacks, new topologies, and especially new classification algorithms |
|
Reading / Datasets |
Works by V. Vassiliou and C. Ioannou (University of Cyprus) Find suitable datasets from https://www.kaggle.com/ |
|
Contact Details
|
Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy |
|
Project title: |
Predicting YouTube Trending Video |
|
Background:
|
A new trend in 5G and 6G networks is using the network edge (base station) for caching and processing and for supporting the quality of service and the experience of end users consuming visual content and interactive media. |
|
Project aims:
|
Within this project, will provide the ability to group and predict users’ and/or content’s needs. One way is to predict videos trending on social media. |
|
Readings / Datasets |
Youtube Video Info https://www.kaggle.com/datasets/datasnaek/youtube-new |
|
Contact Details
|
Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy |
|
Project title: |
Predicting Motor-vehicle Accidents |
|
Background:
|
Motor vehicle crashes cause loss of life, property and finances. Vehicle accidents are a focus of traffic safety research, uncovering useful information that can be directly applied to reduce these losses. Traditionally, modeling crash events has been done using machine learning techniques, considering crash level variables, such as roadway characteristics, lighting conditions, weather conditions and the prevalence of drugs or alcohol.
|
|
Project aims:
|
Incorporate crash report data, road data, and demographic data to better understand crash locations and the associated cost. Use a mixed linear modeling technique that enables data fusion in a principled way to build a better predictive model. Analyze the natural clustering of events in space by different geographic levels. |
|
Readings / Datasets |
Need to get relevant information from the Cyprus Police and the Association of Insurance Companies. |
|
Contact Details
|
Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy |
|
Project title: |
Preventive to Predictive Maintenance |
|
Background:
|
Maintenance is an integral component of operating manufacturing equipment. Preventive maintenance occurs on the same schedule every cycle — whether or not the upkeep is actually needed. Preventive maintenance is designed to keep parts in good repair but does not take the state of a component or process into account. Predictive maintenance occurs as needed, drawing on real-time collection and analysis of machine operation data to identify issues at the nascent stage before they can interrupt production. With predictive maintenance, repairs happen during machine operation and address an actual problem. If a shutdown is required, it will be shorter and more targeted. While the planned downtime in preventive maintenance may be inconvenient and represents a decrease in overall capacity availability, it’s highly preferable to the unplanned downtime of reactive maintenance, where costs and duration may be unknown until the problem is diagnosed and addressed. Preventive to Predicitve Maintenance is about the transition of a preventive maintenance strategy to a predictive maintenance strategy of a replaceable part |
|
Project aims:
|
The objective of the project is to use the associated detailed dataset is to precisely predict the remaining useful life (RUL) of the element in question, so a transition to predictive maintenance is made possible. |
|
Readings / Datasets |
https://www.kaggle.com/datasets/prognosticshse/preventive-to-predicitve-maintenance |
|
Contact Details
|
Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy |
|
Project title: |
Analysis and Visualization of Mobile Cellular Telephony Network Coverage and Electromagnetic Measurements |
|
Background:
|
An interesting challenge in 5G and 6G Mobile Cellular Telephony is the need to have a large(r) number of small(er) base stations to achieve the data rates, delays and number of users expected. The Mobile Network Operators report to the Telecommunication Regulators of each country a number of parameters, including coverage, location of stations and radiation levels.
|
|
Project aims:
|
The objective of the project is to collate information available in different systems/platforms and generate an up-to-date map of coverage and mobile telephony stations, alongside the information gathered through periodic and ad-hoc measurements. These will be related with measurements of mobile telephony quality of service and shown on a common reference map with the ability to explore the changes during the years. |
|
Readings / Datasets
|
Open data from the Department of Electronic Communications. Information from the ICT market observatory of OCECPR and other public data. http://www.emf.mcw.gov.cy/emf/?page=emfmeasurements https://www.data.gov.cy/search/field_topic/%CE%B5%CF%80%CE%B9%CF%83%CF%84%CE%AE%CE%BC%CE%B7-%CE%BA%CE%B1%CE%B9-%CF%84%CE%B5%CF%87%CE%BD%CE%BF%CE%BB%CE%BF%CE%B3%CE%AF%CE%B1-40
|
|
Contact Details
|
Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy |
Title: Measuring consumer price inflation in Cyprus
We calculate Cyprus inflation in terms of the Consumer Price Index (CPI), which is compiled using the online prices (big data) for a pre-selected/fixed and representative basket of goods and services. In doing so, we employ ‘web scraping’ algorithms to visit the websites/pages of large retailers in Cyprus and collect and store the prices of goods and services available online. We then implement standard techniques and proprietary methodologies to calculate price statistics and indices. This work is inspired by the Billion Prices Project, an academic initiative at MIT and Harvard that used prices collected from hundreds of online large retailers around the world on a daily basis to conduct research in macro and international economics.
Title: Studying the crude oil price pass-through into fuel and/or consumer prices in Cyprus
We leverage a high-frequency (weekly) online price series dataset in an econometric framework for pass-through estimation, forecasting and policy making. We aim to quantify how (in what degree and time) the fuel and/or consumer prices respond to an oil price shock and provide policy makers with important implications and useful information regarding consumer welfare.
Title: Create an end-to-end MLops pipeline for a specific model (“anti-scrapping model”)
Duration: Hard deadline the end of August
Team: 3 Data Science MSc Students with backgrounds: 2 in CS, 1 in Statistics, plus one supervisor on the UCY side. The supervisor can offer weekly one-hour coaching and review sessions to the students.
Project Description:
Scope of this project is to create an end-to-end MLops pipeline for a specific model (detecting competitors trying to scrape/crawl prices from our website). The students will work with the strong guidance of our Data science team to design/implement/deploy an ML model to production.
Resources:
Our company will provide
Desired Deliverables:
Title: Loyalty Scheme – Client Churn Prediction
Duration: 2 months (June 2023 – July 2023)
Team Members:
Data Scientist with a background in Computer Science: Responsible for data preprocessing, feature engineering, and model development.
Statistician: Responsible for statistical analysis, model evaluation, and validation.
Business Analyst: Responsible for understanding the business requirements, interpreting the results, and providing insights for decision-making.
Supervisor: A subject matter expert who will provide guidance and support throughout the project.
Project Description:
This project aims to predict the churn of COMPANY’s loyalty customers by analyzing customer data and building predictive models. The team will work closely with COMPANY’s internal stakeholders to gather relevant data related to customer behavior, transactions, and engagement metrics. The collected data will be processed, cleaned, and transformed to create meaningful features. The team will then explore various machine learning algorithms to develop predictive models that can identify potential churners accurately.
The project will involve conducting an in-depth exploratory data analysis to uncover insights and patterns within the data. The team will perform feature engineering to derive additional relevant features from the existing dataset. These features will be used to train and evaluate predictive models using techniques such as logistic regression, decision trees, random forests, or neural networks. The models will be assessed based on metrics like accuracy, precision, and F1-score.
Resources:
Our company will provide:
A designated contact person who will offer guidance, support, and domain expertise.
Access to relevant customer data, subject to a non-disclosure agreement (NDA).
Desired Deliverables:
Churn prediction models and evaluation report: The report should detail the methodology, model performance metrics, and provide insights into key features driving churn.
Final Project Report: A comprehensive document summarizing the entire project, including data preprocessing, model development, insights, and recommendations for COMPANY.
Presentation: A final presentation summarizing the project, highlighting the key findings, and providing actionable recommendations for COMPANY to reduce churn among loyalty customers.
Title: Quality control, cleaning, and analyses of wearable data in the context of clinical trials.
Wearables have enabled the non-intrusive monitoring of subjects in medical research studies during their normal daily lives. At Stremble we have already expanded our in-house analytics platform to automatically collect data from several of our clinical trials and studies in flat JSON files. However, cleaning, organizing, and analyzing these data is a challenge. Therefore, standardizing the quality control storage and access to the data would benefit many ongoing cancer research projects.
Project Aims:
Skills:
Students in Computer Science, Mathematics or Statistics would be eligible for this position.
Programming/Scripting knowledge preferably Python
R and ideally Shinny for the analyses.
Who we are:
Stremble Ventures is a company based in Limassol, Cyprus established in 2011. We offer contract research to companies and research institutions around the world with a focus on Bioinformatics and Computational Biology. We have several ongoing EU, RIF, and commercially funded projects.
Tickmill is a retail FX broker operating globally and employing a dedicated team of over 250 professionals. With an average monthly trading volume surpassing $150 billion, Tickmill stands as a major player in the financial industry. One of its key assets is its own quantitative research team that has developed proprietary trading systems. These systems leverage extensive tabular data stored across diverse databases. The team’s primary objective is to make informed investment decisions by analyzing unconventional data. For instance, is it possible to predict Starbucks’ stock price if we know the number of people visiting their stores? If so, would combining this information with the average amount spent by customers in Starbucks stores yield improved results?
During the capstone project, students who join Tickmill’s Quantitative Research team will actively participate in tasks such as data mining, data storage utilizing various database types based on their specific requirements, and, notably, data analysis. In the financial industry, data analysis poses significant challenges for two primary reasons. Firstly, unlike voice, image, or video data, financial data cannot be generated or created. Secondly, the signal-to-noise ratio is typically low, which significantly contributes to the complexity of this particular task in machine learning and data science.
In this project, participating students will undergo a two-week training program focused on the financial industry and trading. Prior knowledge in this field is not required. Following the training, students will delve into various types of data and explore the corresponding data analysis within that domain. They will have the flexibility to use their preferred data analysis tools (although we utilize Jupyter notebook, they are free to choose the tool they are most comfortable with) to conduct their analyses. By the conclusion of the project, students will be equipped with the ability to determine which types of data are valuable for predicting the price change of financial assets and which data contain excessive noise. This understanding encompasses the concepts of correlation and causation, although it is not limited to these factors alone. In addition to this phase, which we refer to as the first stage of data analysis, students will have the opportunity to create their own features and explore the freedom to combine different types of data in order to derive meaningful insights and conclusions.
Due to confidentiality agreements, students will not be allowed to disclose data vendors names in their capstone reports.
