>>>Master

... in Data Science

Capstone_Projects

>>Capstone Project

The capstone project has been designed to apply knowledge into practice and to develop and improve critical skills such as problem-solving and collaboration skills. Students are matched with research labs within the UCY community and with industry partners to investigate pressing issues, applying data science areas. Capstone projects aim to give students some professional experience in a real work environment and help enhance their soft skills.  These projects involve groups of roughly 3-4 students working in partnership.

The process is the following:

  • A short description of projects are announced to students.
  • Students bid up to three projects taking into account the fields of their interest or research.
  • The data science directors make the final assignment of projects to students. The projects are under the supervision of a member of the Programme’s academic staff.
  • Specific learning outcomes are stipulated in a learning agreement between the student, the supervisor and the company.
  • The student keeps a log file of his/her work and at the end writes a progress report (6000 words).
  • The company is obliged to monitor the progress of the students and to provide relevant mentorship.

Final assessment is carried out by the company and the supervisor.

Available Capstone Projects 

_Summer 2026

Retail Zoom

Spotter Analysis: Detection of Atypical Behaviour Across Supermarket Hierarchies

Background

Retail datasets generated from supermarket transactions are inherently complex, spanning multiple hierarchical levels such as product, brand, category, store, and region. Consumer behaviour across these levels is influenced by factors such as promotions, pricing, seasonality, availability, and external market conditions.

However, retail data often exhibits atypical patterns, including sudden spikes or drops in demand, inconsistencies across hierarchical levels, or unexpected deviations from historical trends. These may indicate data issues, stock disruptions, promotional effects, or genuine shifts in consumer behaviour.

Identifying such patterns across multiple levels of aggregation is essential for ensuring data reliability and generating accurate market insights. Traditional approaches rely heavily on manual inspection and are not scalable across large datasets.

This project introduces the concept of “Spotter Analysis”, a framework designed to automatically detect and highlight atypical behaviour across different levels of a supermarket chain.

Project Description

The objective of this project is to develop a multi-level analytical framework (“Spotter Analysis”) that detects atypical consumption behaviour across supermarket data.

The system will analyse transactional data across different hierarchical levels and identify deviations from expected patterns. These may include anomalies within a single level (e.g. product sales spike) or inconsistencies between levels (e.g. product increase not reflected at category level).

Students will design and implement methods to detect, compare, and interpret these atypical patterns, supporting both data validation and business insight generation.

Data

Students will work with transactional retail datasets that may include:

• product-level sales transactions

• product hierarchy (product → brand → category)

• store and regional structure

• price and promotion indicators

• time series of sales activity

Access to real transactional datasets will be granted under a Non-Disclosure Agreement (NDA). All results presented externally will be aggregated and anonymized.

Analytical Scope (Multi-Level Spotter Analysis)

The project will analyse behaviour across multiple levels:Product Level

• detect abnormal spikes/drops in SKU sales

• identify unusual patterns in individual product performance

Brand Level

• detect inconsistencies across products within the same brand

• identify brand-level shifts in demand

Category Level

• analyse category growth or decline

• detect abnormal category-level behaviour

Store / Region Level

• identify store-specific anomalies

• compare regional consumption patterns

• detect localised disruptions or trends

Cross-Level Consistency Checks

• compare behaviour across levels (e.g. SKU vs category)

• detect mismatches and structural inconsistencies

Tasks

Data Preparation and Exploration

• preprocess transactional datasets

• analyse patterns across hierarchical levels

Anomaly Detection

• develop statistical and machine learning models to detect atypical behaviour

• identify deviations at product, brand, category, and store levels

Cross-Level Analysis

• detect inconsistencies between hierarchical levels

• analyse how anomalies propagate across levels

Evaluation

• assess the effectiveness of detection methods

• validate results against known patterns or eventsDeliverables

• A multi-level “Spotter Analysis” framework for retail data

• Implementation of anomaly detection models across different hierarchical levels

• Documentation describing methodology and findings

• Final report and presentation with aggregated insights

Extra-Mile Tasks

Students may extend the project by:

• developing a dashboard to visualise anomalies across levels

• building an alert system for real-time anomaly detection

• linking detected anomalies to potential drivers (promotions, pricing, availability)

• creating a scoring system to rank the severity of anomalies

Jinius Marketplace

Data-Driven Recommendations and Search Intelligence for the Jinius Marketplace

Background:

Jinius Marketplace, developed by the Bank of Cyprus, is a digital e-commerce platform that connects consumers with local retailers across a wide product range. It offers features like centralised checkout, wishlists, promotions, and loyalty rewards. Businesses benefit from simplified logistics and exposure, while shoppers enjoy a more personalised and convenient shopping experience.

The platform generates diverse data: basket contents, order history, wishlists, product descriptions, images, catalogue metadata, and potentially search logs. These data sources present a rich opportunity to apply machine learning and data science techniques to improve recommendations, product discovery, search relevance, and commercial insights.

Previous capstone projects have explored product matching and product attribute extraction. This project will focus on new directions using similar data sources, with particular emphasis on recommendations and search.

Objective:

The goal of this project is to develop intelligent, data-driven features for the Jinius platform using available customer interaction, catalogue, and search-related data. The team can focus on one or both of the following tasks:

Tasks:

1. Product Association, Bundle, and Recommendation Intelligence

Analyse basket, order, wishlist, and catalogue data to identify relationships between products, categories, brands, or retailers, and use these insights to generate recommendations or product bundles.

Examples include “frequently bought together” products, “complete your basket” suggestions, wishlist-based recommendations, seasonal bundles, and complementary product recommendations.

Suggested Techniques: Association rule mining, Apriori, FP-growth, collaborative filtering, matrix factorisation, product embeddings, graph analysis, co-purchase networks, sequence-aware recommendation.

Impact: Improved personalisation, higher basket value, better product discovery, and actionable insights for promotions and merchandising.

2. Search Relevance and AI-Powered Product Discovery

Improve how customers search for and discover products on the marketplace, especially when searches are vague, informal, incomplete, or intent-based. The aim is to help users find relevant products even when their query does not exactly match product titles or descriptions.Examples include semantic search, natural-language product discovery, query expansion, ranking improvement, identification of poor or no-result searches, and alternative product suggestions.

Suggested Techniques: NLP, semantic search, text embeddings, multimodal embeddings, vector databases, learning-to-rank, transformer models, CLIP, retrieval-augmented generation, large language models.

Impact: Improved search relevance, reduced customer frustration, better product visibility, and increased search-to-purchase conversion.

Data:

Basket and order history, anonymised

Product catalogue, text, images, and metadata

Wishlists, anonymised

Search logs

PwC

Enhancing Entity Benchmarking Through Public Data Integration in Power BI

Objective:

This project focuses on enriching existing Microsoft Power BI dashboards with publicly available external data. The goal is to identify, connect to, and integrate relevant public data sources into Power BI — using APIs, web URLs, and other data connectivity methods — in order to enhance the benchmarking capabilities of individual entity dashboards. By overlaying external reference points onto internal data (e.g., comparing an entity’s average salary against national averages, or contextualising performance within broader economic and industry trends), entities will be able to assess their position relative to the market, discover insights, and make more informed decisions.

Scope of Work

Participants will:

• Research and identify relevant public data sources (e.g., government statistics portals, open data platforms, industry bodies, international organisations)

• Connect to external data using APIs, web connectors, or other available ingestion methods/tools (python/R/SQL)

• Transform and model the acquired data so that it integrates seamlessly with existing dashboard structures

• Design and build visualisations that enable clear, intuitive comparison between an entity’s internal metrics and external benchmarks

• Document data sources, refresh schedules, transformation logic, and any limitations

Expected Outcomes:

• A set of Power BI data connections to curated public sources, ready for ongoing use

• Enhanced Power BI reports and dashboards that provide entities with external context and benchmarking insights

• Documentation enabling the team to maintain and extend the solution beyond the project period

Bonus 1: Implementation of AgenticAI tools/add-ons to enhance the quality and extend the scope of data injected into Power BI.

Bonus 2: Develop a robust data processing layer using python/R/SQL, to prepare analysis-ready datasets suitable for both statistical exploration and visualisation.

What Participants Will Gain:

• Hands-on experience with Microsoft Power BI (data connectivity, Power Query, data modelling, DAX, and visualisation), python/R and SQL.

• Practical exposure to working with APIs and public data sources• Understanding of how external data adds analytical value in a professional environment

• Experience working within a corporate team on a deliverable with real business impact

Hellas Direct

Title: Road Assistance Delay Time Prediction

Description: The project aims to address the challenge of accurately predicting the delay time for Road Assistance services. In motor insurance operations, response time is a critical service quality indicator directly affecting customer satisfaction and operational efficiency. However, delays can vary significantly depending on factors such as location, time of day, traffic conditions, weather, vehicle type, and service provider availability.

This project offers an opportunity to apply advanced data science and machine learning techniques to develop a predictive model capable of estimating delay times based on operational and contextual data. The objective is to provide actionable insights that can help improve dispatch decisions, optimize resource allocation, and enhance customer experience.

Deliverables:

• A predictive machine learning model capable of estimating the delay time from the initial service request to the arrival of the service vehicle.

• A well-documented analysis pipeline including data preprocessing, exploratory data analysis (EDA) with visual insights, feature engineering, model development, and validation.

• A short final presentation outlining the project and summarizing key findings, along with recommendations on how the model can be integrated into the company’s dispatch system to improve customer communication.

Available Data:

Students will work with a comprehensive dataset from historical road assistance logs. The dataset includes features such as request timestamps, pickup and drop-off coordinates, service vehicle types, and incident categories.

Instructions:

Students are encouraged to follow the full data science lifecycle and apply the tools and techniques acquired during their studies. They are also expected to maintain regular communication with the company’s Data Team to validate findings and stay aligned with project goals. The project emphasizes team collaboration and offers students the opportunity to work with real-world, time-sensitive data in a practical business context.

Main Challenges:

1. Handling missing, inconsistent, or erroneous records, which may require cross-referencing with external data sources such as traffic or weather APIs to be identified by the students.

2. Accounting for high variability in delay times driven by rare or extreme events, such as peak demand periods or large-scale incidents.

3. Applying innovative modeling approaches to accurately predict delay times across diverse geographic areas and incident types with limited historical coverage.

4. Ensuring the model architecture is optimized for quick inference to provide near-instant updates to customers.

5. Deriving accurate delay times from provider event timestamps, which may be delayed, recorded in batches, or missing, requiring careful data cleaning and robust target definition.

KPMG

Predictive Analytics and Clinical Decision Support in a Hospital Setting

Project Description

This project involves the development of a Machine Learning model to support clinical and operational decision-making at the German Medical Institute (GMI), a private hospital based in Limassol, Cyprus. Depending on the selected focus area — agreed upon with the student at the outset — the model may address challenges such as patient readmission prediction, length-of-stay estimation, or appointment no-show forecasting. More specifically, with regards to the length-of-stay estimation, students will be provided with retrospective patient records related to their admission that includes primary and secondary diagnoses, procedure codes (i.e., what surgery or other procedure brought them to the hospital), their age, gender, nationality and other relevant demographic information to develop a model that accurately predicts how long a patient is expected to stay in the hospital based on the above information. The data required for the development of the above mentioned models will be anonymised and provided to the students, ensuring compliance with GDPR and Data Sharing Agreements that will have already been put in place.

By analysing anonymised patient and operational data, the project aims to generate actionable insights that can support clinical staff, enhance hospital resource planning and ultimately improve patient outcomes. Students will gain hands-on experience in healthcare data preprocessing, feature engineering, and the application of traditional and advanced Machine Learning techniques, while navigating the unique ethical, regulatory, and data privacy considerations that define real-world healthcare environments.

Expected Deliverables

At the conclusion of the project, the student is expected to provide the following:

• A cleaned and well-documented dataset, including all preprocessing and feature engineering steps

• An exploratory data analysis (EDA) report describing key patterns and clinical/operational insights

• A trained and evaluated ML model with documented methodology, performance metrics, and limitations

• A final written report suitable for presentation to both technical and non-technical stakeholders

• A short presentation of findings to the GMI AI team

Extra-Mile Tasks

For students who demonstrate exceptional performance and are willing to go above and beyond, the following additional tasks are encouraged:

• Development of an interactive dashboard to visualise model outputs for clinical or operational end-users

Synectics

Enhanced Fraud Modelling with LLM based Fraud Analyst assistant

Project Goal

Develop a machine learning model to detect fraudulent card transactions and extend it with an LLM-based chatbot that can generate structured fraud analysis for individual transactions using model outputs, entity history, and data insights.

Problem Statement

Fraud detection requires balancing accurate classification with clear, actionable insights. This project focuses on building a reliable detection model and enhancing it with an AI agent that can interpret results and provide human-readable explanations for transaction-level decisions.

Data Sources

Card transaction dataset with fraud indicators

Tasks

• Perform exploratory data analysis to understand patterns and relationships

• Develop and evaluate a fraud detection model

• Incorporate entity-level historical context

• Build an LLM-based agent to act as a fraud analyst assistant

Deliverables

1. Trained Model

A fraud detection model with performance evaluation

2. LLM Fraud Analysis Agent

An AI agent that can generate insights for individual transactions using model outputs and contextual data

3. Technical Documentation

Summary of approach, findings, model performance, and agent design

4. Presentation

Overview of the project, key insights, and outcomes

Grant Thornton

Grant Thornton has put forward two proposals with a similar scope. The final choice will depend on data availability and student interest.

Proposal 1

Project Title: Building a Custom Legal Chatbot Using LLMs and Cypriot Legislation

This project focuses on transforming how users interact with Cypriot legal texts by developing a custom chatbot powered by large language models (LLMs). The goal is to make legislation and court decisions, available through www.cylaw.org, more searchable and understandable.

Students will design a prototype that combines semantic search with an LLM-powered question-answering system, allowing users to ask legal questions in plain language and receive accurate, well-sourced responses.

Key deliverables include a semantic retrieval system that understands the meaning of user queries and a chatbot interface that generates summaries or explanations of legal content. The project will use publicly available legal documents from CyLaw, with the option to focus on a specific legal domain such as labour or environmental law to ensure a manageable and meaningful scope.

Proposal 2

Project Title: “RAG-Based Policy Assistant (Prototype)”

Project Description

This project focuses on the design and implementation of a prototype Retrieval-Augmented Generation (RAG) agent using Python, aimed at improving access to internal organisational knowledge. Students will develop a streamlined framework that combines document retrieval techniques with large language models to enable users to query a limited set of company policies— such as HR, compliance, or cybersecurity documents—through a conversational interface. The emphasis will be on building a working end-to-end pipeline, ensuring that responses are grounded in source documents and include references to improve transparency and trust.

Following the development of the core framework, the solution will be applied to a selected subset of real company policies to demonstrate practical usability. The final output will include a functional prototype with a simple chat interface and a proof-of-concept integration into Microsoft Teams or a comparable environment. The project will prioritise correctness, usability, and architectural clarity over full production deployment, while also introducing key concepts such as responsible AI, data privacy considerations, and system scalability at a conceptual level.

Expected Deliverables

1. Solution Design

· High-level architecture of the RAG system (ingestion, retrieval, generation)

· Technology stack justification (Python, vector DB, LLM choice)

2. Core RAG Prototype (Python)

· Document ingestion and preprocessing pipeline (limited dataset)· Embedding generation and storage

· Retrieval and LLM-based response generation

3. Policy Dataset Preparation

· Structured processing of a small number (1-2) of company policy documents

· Chunking and metadata tagging approach

4. Chat Interface (Prototype)

· Simple UI (e.g., Streamlit or lightweight web app)

· Ability to query documents and receive contextual responses

5. Response Quality Enhancements

· Basic prompt engineering

· Inclusion of source citations in responses

6. Evaluation Approach

· Simple testing framework (manual test cases)

· Basic metrics: relevance, accuracy, hallucination observations

7. Microsoft Teams Integration (PoC)

· Basic bot or webhook-based integration (non-production)

· Demonstration of how the assistant can be accessed in a workplace tool

8. Governance & Risk Considerations (Conceptual)

· Short paper covering:

o Data privacy considerations

o Access control concepts

o Risks of LLM usage (e.g., hallucination, bias)

9. Documentation

· Technical setup guide

· User guide for interacting with the assistant

10. Final Presentation & Demo

· Demonstration of the working prototype

· Key findings and limitations

11. Final Report

· Methodology, implementation steps, challenges

· Recommendations for scaling to production

Bank of Cyprus

Next-Generation Customer Segmentation Using Transformer-Based Encoders

Project Description

Customer segmentation is a long-standing problem in banking analytics. Traditionally, banks segment customers using manually engineered tabular features and classical clustering techniques (e.g. k-means, hierarchical clustering). While these methods are simple and interpretable, they may fail to capture complex relationships in modern banking data.

Recent advances in encoder-based and transformer-inspired models for tabular data allow customer representations to be learned rather than hand-crafted. These models generate dense embeddings that may encode latent financial and behavioural patterns.

This capstone explores a core empirical question, not a predefined conclusion:

Do encoder-based representations actually improve customer segmentation in banking, compared to traditional approaches?

Students will design experiments to test this claim rigorously, using tabular banking data.

What You Will Work On

You will solve the same segmentation problem twice, using two fundamentally different paradigms.

1. Traditional Segmentation (Baseline)

You will:

• Work with tabular banking data (customer attributes, balances, product holdings, transaction aggregates)

• Engineer and select features

• Apply classical clustering techniques

• Interpret and profile resulting customer segments

This represents how segmentation is typically done today in many banking environments.

2. Next Generation Encoder-Based Segmentation

You will:

• Use the same tabular banking data

• Fine-tune or train an encoder-based model adapted for tabular inputs

• Generate customer embeddings

• Perform clustering in the learned latent space

• Analyse whether and how the resulting segments differ from the baseline

The encoder replaces manual feature engineering, not the clustering step.What You Are Expected to Prove (or Disprove)

This is a research-style project, not a demo.

You must:

• Compare the two approaches fairly (same data, same number of clusters where applicable)

• Use objective metrics, stability tests, and qualitative analysis

• Critically assess:

o Segment coherence

o Interpretability

o Stability

o Complexity vs. benefit

Negative or neutral results are acceptable if they are well-argued.

Data and Tools

• Domain: Banking

• Data Modality: Fully tabular

• Models:

o Classical clustering methods

o Encoder-based models (including fine-tuned large or pre-trained models adapted to

tabular data)

• Programming: Python-based data science/ AI stack

The focus is methodology and reasoning, not production deployment.

What You Will Deliver

1. Final Technical Report

o Clear problem definition

o Description of both segmentation pipelines

o Experimental setup and evaluation

o Critical comparison and conclusions

o Explicit discussion of limitations

2. Code Repository

o Reproducible experiments

o Clear separation between baseline and encoder-based approaches

o Documented assumptions and parameters

3. Final Presentation

o Clear answer to the core research questiono Evidence-based conclusions

o Honest discussion of trade-offs

High Level Project Milestones

1. Problem Definition & Data Understanding

2. Baseline (Traditional) Segmentation

3. Encoder Selection & Fine-Tuning

4. Embedding-Based Segmentation

5. Comparative Evaluation

6. Final Synthesis & Conclusions

Risk Matrix

Background:

RiskMatrix is an international Credit Risk Advisory and Software Services firm that has been appointed by Moody’s Analytics as its exclusive representative relating to credit risk software and services offering across the number of territories including Cyprus, Greece, the Middle East, Egypt and East Africa. We serve more than 130 banking groups.

RiskMatrix’ Credit Analytics team provides advisory services to many of RiskMatrix’ customers. These services centre around the measurement of credit risk.

A key class of credit risk models comprise models to determine probabilities of default (PDs).

RiskMatrix uses its own proprietary approach to build such models. The approach is structured and designed to avoid over- and under-fitting of the models. This is important as the banks supported by RiskMatrix usually cannot provide extensive historical information for their business customers and therefore data limitations pose a significant challenge to building robust models.

The Capstone Project:

Recent developments in machine learning and data science provide alternative approaches to developing PD models. The Capstone will explore how such approaches perform in a real-world domain and to understand the reasons why such approaches may or may not achieve better outcomes than more traditional approaches.

Objective:

1) Use the XG-Boost approach to develop a PD model using financial statement data for a “typical” dataset comprising business customers of a bank. Assess the performance of the model, the strengths and limitations of the approach and the impact of different choices in parameterisation and structure of the model. Compare XG-Boost to a benchmark modelling approach (to be agreed as part of the Capstone Project).

2) Perform an assessment of the resulting model structure and behaviour. This will involve using an approach such as SHAP to provide an explanation and understanding of the model produced and identify weakness of the model. Compare SHAP (or another selected approaches) to alternate “standard” approaches to assess the model behaviour (e.g. sensitivity analysis). Assess the appropriateness, suitability and strengths and limitations of SHAP for this purpose.

3) Optional (if time permits): Assess how dataset size affects the quality of the model produced

4) Optional (if time permits): Compare XG-Boost to an alternative contemporary ML approach, for example, CatBoost.

Dataset:

The data will comprise a subset of “typical” data used in a modelling project by RiskMatrix. This will include candidate model input values such as financial ratios together with outcome values (whether a company defaulted within twelve months of the observation point).

Approach:

A senior member of RiskMatrix’ Credit Analytics team will partner with the students for this project.

Using this approach, RiskMatrix can assist the students to gain a deeper understanding of the data and its meaning, provide expertise in “best practices” in PD model development, and provide guidance in the interpretation of the results and the appropriateness and effectiveness of the models produced.

The focus will be on understanding how the selected techniques work on a real-world modelling task, with real world problems and complications. The emphasis will be on gaining deeper understanding as opposed to building a high performance model.

 

_Summer 2025

Jinius Marketplace

Data-Driven Insights for the Jinius Marketplace

Background:

Jinius Marketplace, developed by the Bank of Cyprus, is a digital e-commerce
platform that connects consumers with local retailers across a wide product range.
It offers features like centralised checkout, wishlists, promotions, and loyalty
rewards. Businesses benefit from simplified logistics and exposure, while shoppers
enjoy a more personalised experience.
The platform generates diverse data: basket contents, order history, wishlists,
product descriptions, images, and catalogue metadata. These data sources present
a rich opportunity to apply machine learning and data science techniques to
improve recommendations, product search, and commercial insights.

Objective:

The goal of this project is to develop intelligent, data-driven features for the Jinius
platform using available customer interaction and catalogue data. The team can
focus on one or more of the following tasks:

Tasks:

  • Product Attribute Extraction (Text & Image)
    Extract structured attributes (e.g. colour, material, style) from unstructured
    product descriptions and images to enrich product metadata.
    • Suggested Techniques: NLP (e.g. BERT, spaCy), computer vision (e.g. CLIP, ResNet)
    • Impact: Improved product filtering, tagging, and search
  • Outfit or Bundle Recommendation Engine
    Build a model to suggest complementary products or full outfits based on
    basket history, wishlists, or seasonal patterns.
    • Suggested Techniques: Collaborative filtering, product embeddings,
      matrix factorisation
    • Impact: Personalised suggestions and bundles
  • Product Association & Trend Analysis
    Discover patterns in customer purchases and wishlist behaviours.
    • Suggested Techniques: Association rule mining (e.g. Apriori), co-
      purchase graph analysis (e.g. NetworkX)
    • Impact: Insights for marketing and merchandising (e.g. “frequently
      bought together”)

Data:

  • Basket and order history (anonymised)
  • Product catalogue (text, images, metadata)
  • Wishlists (anonymised)

 

 

 

PwC

LLM-Powered Chatbot for Regulatory and Policy Question Answering

Background:

Professionals operating in financial, legal, and compliance domains often rely on dense PDF documents containing regulatory and policy information (unstructured datasets). These documents are not easily searchable or accessible through natural language, leading
to inefficiencies, misinterpretation, and time-consuming manual analysis. With recent advancements in Large Language Models (LLMs) and vector search technologies, it is now possible to build intelligent systems capable of understanding, retrieving, and contextualizing
information from such complex documents.

This presents an opportunity to work on a project with real-world impact, utilizing real datasets to address challenges faced by professionals daily. It allows exploration of an emerging area of AI that is in high demand, promising significant advancements and efficiencies in information handling and interpretation.

Objective:

  • Develop a chatbot powered by state-of-the-art LLMs capable of answering regulatory and policy-related questions in natural language.
  • Design a robust pipeline for parsing, processing, and semantically indexing PDF documents for high-accuracy retrieval.
  • Implement a Retrieval-Augmented Generation (RAG) architecture that retrieves relevant document segments and uses them to generate grounded responses.
  • Ensure traceability by linking each response to the specific document sections used as context. 

Data:

Data will consist of PDF reports, specifically focusing on regulatory and policy information with regards to the financial, legal, and compliance domains. These reports are typically dense and lengthy, containing valuable insights that need to be efficiently parsed and indexed for effective retrieval by the chatbot.

Desired Outcome:

  • A chatbot that can accurately respond to policy and regulatory questions by referencing content from official documentation (grounded and explainable answers).
  • Supporting Documentation (technical & user manual)

 

Bonus 1: Custom UI Interface: Develop a user-friendly, custom UI interface for intuitive interaction with the chatbot. This interface will enhance user experience, allowing seamless navigation through queries and responses.

Bonus 2: Agentic RAG – Implement an Agentic RAG approach to handle more complex and nuanced questions. This advanced capability will allow the chatbot to navigate not just simple queries, but also multi-faceted questions that require deeper analysis
and understanding.

 

Hellas Direct

Used Car Price Prediction using Classified Ads Data

 

Project Description:

The proposed project aims to address the challenge of accurately estimating the selling price of used cars. The second-hand automotive market is highly dynamic, with vehicle value influenced by a wide range of factors including technical specifications, age, mileage, and seller profile. This project offers an opportunity to leverage advanced data science techniques to develop a robust, data-driven car valuation model.

 

Deliverables:

  • A predictive machine learning model capable of estimating the selling price of used
    cars based on their attributes.
  • A well-documented analysis pipeline including data preprocessing, exploratory data
    analysis (EDA) with visual insights, feature engineering and selection, model
    development, and validation.
  • A short final presentation outlining the project and summarizing its key findings, along
    with recommendations on how the developed model can be applied or further
    improved by the company.

Available Data:

Students will work with a large-scale dataset collected from various classified ads platforms.
The dataset includes detailed features such as brand, model, engine specifications, mileage,
and selling price, among others. The data also supports analysis over time and by region.

 

Instructions:

Students are encouraged to follow the full data science lifecycle and apply the tools and
techniques they have acquired during their studies. They are also expected to maintain
regular communication with the company’s team to validate findings and to stay aligned with
project’s goals. The project emphasizes team collaboration and gives students the
opportunity to work with real-world data in a practical business context.

 

Main Challenges:

  • Managing and processing a large-scale dataset efficiently.
  • Handling noisy data and erroneous values. This may require cross-referencing with
    external data sources to be identified by the students.
  • Identifying and treating a high number of outliers.
  • Applying innovative modeling approaches to effectively predict prices for car models
    with limited data coverage.

KPMG

Customer Experience Analytics using Machine Learning Techniques

Project Description:

This project involves developing a Machine Learning model to support and optimize call center operations for a utility provider. Depending on the selected topic, the model may focus on forecasting call volumes, predicting first-call resolution, classifying call topics, or clustering agent performance. By analyzing historical call data, the project aims to uncover actionable insights that can improve efficiency, resource planning, and customer service. This will provide students with practical experience in data preprocessing, feature engineering, and traditional Machine Learning techniques, while addressing real-world challenges in the context of customer support operations.

 

Resources: Our company will provide

  • contact person that will provide guidance and support/mentoring
  • access to all necessary data (all covered by NDA)

Deliverables:

  • Exploratory data analysis report that showcases meaningful insights (e.g., summary statistics, correlation analysis, customer/product segmentation)
  • Implementation of forecasting and predictive models
  • Final Project report/dashboard

Extra-Mile Tasks: Development of a simple user interface (UI) that allows users to interact with the model and view live outputs (e.g., predictions, classifications, or insights) in a user-friendly way.

 

 

 

Synectics

Fraud Detection Model

 

Project Goal:

Develop a machine learning model to detect fraudulent card transactions. You will explore the dataset, identify key fraud indicators, assess sources of bias, and create and calibrate a fraud detection model using appropriate performance metrics.

Problem Statement:

Card fraud detection is critical for protecting businesses and customers from monetary loss. However, effective fraud detection models must balance sensitivity and specificity in a highly imbalanced dataset. This project aims to extract meaningful features, analyse fraud patterns, identify relationships and biases, and develop a robust model for detecting fraudulent transactions.

 

Data Sources:

Card transaction dataset with fraud indicators.

Tasks:

  • Exploratory Data Analysis: Feature extraction, visualization, and correlation analysis, etc.
  • Model Creation: Create a model of your choosing for detection of frauds.

Deliverables:

  • Trained Model: Fraud detection model with performance evaluation.
  • Technical Documentation: Explanation of methodology, EDA findings, model training and
    calibration, identified biases, and results.
  • Presentation: A summary of findings, challenges, and key takeaways.

Grant Thornton

Building a Custom Legal Chatbot Using LLMs and Cypriot Legislation

 

Project Description:

This project focuses on transforming how users interact with Cypriot legal texts by developing a custom chatbot powered by large language models (LLMs). The goal is to make legislation and court decisions, available through www.cylaw.org, more searchable and understandable. Students will design a prototype that combines semantic search with an LLM-powered question-answering system, allowing users to ask legal questions in plain language and receive accurate, well-sourced responses. Key deliverables include a semantic retrieval system that understands the meaning of user queries and a chatbot interface that generates summaries or explanations of legal content. The project will use publicly available legal documents from CyLaw, with the option to focus on a specific legal domain such as labour or environmental law to ensure a manageable and meaningful scope.

Bank of Cyprus

Deloitte

Understanding Cyprus’s real estate market

 

Project Description:

As part of its Real Estate offering, Deloitte would like to build an end-to-end real estate asset monitoring and analysis infrastructure through capturing data on a recurring basis from various publicly and non-publicly available sources. These datasets will be subsequently used to populate Deloitte’s internally developed PowerBI tool, which the business is solely going to use for internal purposes.

 

Objective:

  • Develop automated web-scraping scripts in Python to download sale and rental-related information from publicly available sources of information.
  • Combine the downloaded datasets from the publicly available sources to Deloitte’s purchased datasets and create analyses on Cyprus’s real estate market.

Available Data:

  • Downloaded information from publicly available domains that include pricing information on sale and rental asking prices, descriptive characteristics of the underlying real estate assets etc.
  • Deloitte-purchased datasets that list all officially recorded purchases of real estate assets since 2010.

Deliverables:

  • Automated web-scraping scripts in Python that will be downloading information on a recurring basis.
  • PowerBI dashboards that will include pre-determined analyses describing Cyprus’s real estate market. These analyses will be geared towards understanding asset performance (value development, rental yields, prices per sqm etc.), and identify overarching market trends.

 

Health Insurance Organization

Assessment of actual healthcare needs of National Healthcare System (GeSY) beneficiaries and identification of areas of over provision and/or under provision based on the current utilization of healthcare services

Background:

The introduction of the General Healthcare System (GeSY) in Cyprus in 2019 has improved significantly the access to healthcare services for the whole population and it has led to the financial protection of all patients. As a result, we have seen a notable increase in demand for healthcare services and, at the same time, a drastic reduction of out-of-pocket payments and a minimisation of unmet healthcare needs. Even though the System has now reached a steady state, the growth of the utilisation of healthcare services cannot be accurately assessed vis-à-vis the actual healthcare needs of the beneficiaries, leading potentially to some areas being overserved and/or others remaining underserved. There is therefore a growing need to assess whether the current distribution and utilisation of healthcare services, align with the actual needs of the beneficiaries.

Factors such as demographic changes and epidemiological factors influence healthcare needs. Mismatching of actual needs against service provision leads to oversupply or undersupply of services. On one hand, oversupply of services leads to the provision of unnecessary treatments, inefficient allocation of resources and increased healthcare expenditure, while on the other hand, undersupply of services leads to unmet needs, poorer health outcomes and increased healthcare expenditure in the long term. In order to maximise the healthcare system’s effectiveness, it is crucial to understand and quantify the actual healthcare needs of the beneficiaries.

Project Description:

This project aims to estimate the actual healthcare needs of beneficiaries for specific services/ procedures/products and compare them against current utilisation of healthcare services/products in order to identify areas of over provision or under provision. More specifically areas that should be addressed are:

  • Assessment of actual healthcare needs: This will involve analysing demographic, epidemiological and socioeconomic data as well as benchmark data from European or other countries, in order to estimate the actual healthcare needs of the beneficiaries. Healthcare needs could be analysed by category/subcategory of services/ drugs / consumables and/or per specific key procedures.
  • Identification of areas of over provision and/or under provision of healthcare services: Mismatches between actual beneficiaries’ needs and current utilisation of healthcare services should be identified. Healthcare services utilisation could be analysed by category/subcategory of services/ drugs / consumables and/or per specific key procedures.

The findings of this project will help identify the areas of over provision and/or under provision of healthcare services, in order to support evidence-based strategic decision-making to align healthcare provision to actual patient needs, thus maximising the efficiency of the System.

Data:

The Health Insurance Organisation (HIO) will provide data available in its IT system, both for beneficiaries’ demographics (e.g. analysis by age, gender, district etc.), as well as for utilisation of services/products (e.g. number of outpatient visits, procedures, diagnostic and lab tests, inpatient incidents etc.) per category/subcategory and/or specific services/products.

Other demographic and socioeconomic data can be obtained from the Statistical Service.

Benchmark data for the utilisation of services in other countries can be obtained from the database of Eurostat, OECD etc.

Drastyc

Customer Churn Prediction and Proactive Retention Strategy for Telco Provider

 

Project Description:

This project will use machine learning to predict which customers are likely to churn based on usage, billing, and interaction data. It will also generate actionable recommendations, such as targeted offers or service improvements to help the telco provider proactively retain high-risk customers and reduce churn rates.

The outcome will empower the telco provider with data-driven insights to optimize customer retention initiatives and enhance overall customer satisfaction.

Resources:

Our company will provide:

  • A designated contact person who will offer guidance and support
  • Access to relevant customer data, subject to a non-disclosure agreement (NDA)
  • Experience with major enterprise customer
  •  Startup and innovative environment

Desired Deliverables:

  • Data cleansing, exploratory analysis and feature construction
  • Model development and evaluation
  • Documentation or Presentation

_Summer 2024

Bank of Cyprus

Internal Fraud Detection System

Description

The goal of this project is to develop an Internal Fraud Detection System. The system aims to identify and flag suspicious activities suggestive of internal fraud by detecting abnormal patterns in employee behaviour and actions. Using unsupervised learning techniques, students will develop models to identify anomalies. This project includes data collection and analysis, model development and model serving.

Milestones

  1. Data Collection and Preprocessing:
  • Objective: Gather and prepare data for analysis.
  • Tasks:
    • Collect data on logs, relations, transactions, demographics.
    • Clean, normalize and preprocess the data and perform feature engineering to ensure the data are prepared for analysis.
  1. Model Development and Evaluation:
  • Objective: Develop and evaluate fraud detection models.
  • Tasks:
    • Conduct exploratory data analysis (EDA) to understand the distributions and relationships between different variables.
    • Apply unsupervised machine learning algorithms (e.g. clustering algorithms, anomaly detection methods etc) to develop fraud detection models that identify unusual patterns and potentially fraudulent activities.
    • Since this is an unsupervised challenge use metrics whenever applicable like:
      • Isolation Forest Anomaly Score
      • Reconstruction Error
      • Z-score
      • One-class SVM
      • Local Outlier Factor
  1. Implementation and Deployment:
  • Objective: Implement and test the model in a simulated environment.
  • Tasks:
    • Deploy the selected model in a simulated environment where it will be able to accept requests.
    • Use Event Driven Architecture (Kafka, RabbitMQ)
      • Producers: Responsible for sending events containing the necessary information for inference.
      • Consumers: Responsible for model inference and notify for possible fraudulent activity.
    • Or, REST API where the model will be behind a GET/POST request and according to the payload it will infer if it is suspicious or not.

Possible set of data sources 

  • Employee Viewing Activity Logs
  • Transactional Data
  • Relations
  • Customer Demographics

 



Catalink Ltd

Curating the Emotional Journey: A Museum Music Recommendation System with Emotion Recognition

Project Goal:

This capstone project challenges you to design and develop a data science application that personalizes the museum visitor experience through music. The application will leverage emotion recognition and stress level technologies to analyze visitor emotional state from video footage and heart rate signal to generate appropriate background music to fit on their visiting experience on the thematic era of the exhibit.

Problem Statement:

Museums strive to create engaging and impactful experiences for visitors. Music plays a significant role in setting the atmosphere and influencing visitor emotions. However, traditional, static background music may not effectively cater to the diverse emotional states and interests of visitors navigating various exhibits. This project aims to bridge this gap by developing a music generation system tailored to human individual emotional state.

Data Sources:

  • Museum Layout: This data will provide a map of the museum layout, including information on exhibit locations and themes (e.g., Ancient Egypt, Renaissance Art, Modernism).
  • Visitor Video Footage: Anonymized video footage captured ethically and with informed consent from visitors will be used to train the emotion recognition model.
  • Heart rate signal: Anonymized heart rate signal captured from visitors to train stress estimation models.
  • Music Library: A curated music library categorized by genre, mood, and historical period will provide the music recommendations.

Technical Stack:

  • Computer Vision: Techniques like facial expression recognition will be used to analyze visitor emotions from video data.
  • Machine Learning: A machine learning model will be trained to classify visitor emotions (e.g., joy, sadness, interest) based on facial expressions and stress level based on heart-rate signals.
  • Recommendation Systems: An algorithm will be developed to generate music that complements both the visitor’s emotional and stress state and the thematic era of the exhibit they are visiting.

Project Deliverables:

  • Data Preprocessing and Annotation (optional): Develop a system to anonymize video footage and extract relevant data for emotion recognition model training.
  • Emotion Recognition Model Development: Train a machine learning model to accurately classify visitor emotions from video data.
  • Stress estimation model development: Train a machine learning model to identify visitors’ stress level during their museum visiting experience.
  • Music generation System Design: Develop an algorithm that generate music based on visitor emotions, stress and exhibit themes.
  • Mobile Application Development (Optional): Create a mobile application that allows visitors to receive personalized music recommendations based on their location within the museum.

Success Criteria:

A successful project will demonstrate the following:

  • Ethical and responsible use of video footage for data collection. (optional)
  • Development of a high-performing emotion recognition model (happy, sad, neutral).
  • Development of a high-performing stress estimation model.
  • Implementation of a music generation system that considers both visitor emotions, stress and exhibit themes.
  • A functional prototype or mobile application showcasing the system’s capabilities. (optional)
  • Clear and concise documentation of the project methodology and results.

Deloitte

Description:

As part of its Real Estate offering, Deloitte would like to build an end-to-end real estate asset monitoring and analysis infrastructure through capturing data on a recurring basis from various publicly and non-publicly available sources. These datasets will be subsequently used to populate Deloitte’s internally developed PowerBI tool, which the business is solely going to use for internal purposes.

Objective:

  1. Develop automated web-scraping scripts in Python to download sale and rental-related information from publicly available sources of information. 
  2. Combine the downloaded datasets from the publicly available sources to Deloitte’s purchased datasets and create analyses on Cyprus’s real estate market.

Data:

  1. Downloaded information from publicly available domains that include pricing information on sale and rental asking prices, descriptive characteristics of the underlying real estate assets etc. 
  2. Deloitte-purchased datasets that list all officially recorded purchases of real estate assets since 2010. 

Deliverables:

  1. Automated web-scraping scripts in Python that will be downloading information on a recurring basis. 
  2. PowerBI dashboards that will include pre-determined analyses describing Cyprus’s real estate market. These analyses will be geared towards understanding asset performance (value development, rental yields, prices per sqm etc.), and identify overarching market trends. 
  3. Relevant documentation.

Drastyc

Title: Telco Network Failure Prediction

Duration: 2 months 

Project Description:

This project aims to leverage advanced data science techniques to anticipate and mitigate network failures in telecommunication systems. Predicting network failures in telecommunication systems is crucial for maintaining service quality, minimizing downtime, and optimizing resource allocation. By analyzing historical data such as network performance metrics, environmental conditions, equipment status, and maintenance logs, the data science team will develop predictive models to anticipate potential failures before they occur. This will allow the Telco provider to proactively address issues, improve network reliability, and enhance customer satisfaction. 

Resources:

Our company will provide:

  • A designated contact person who will offer guidance and support
  • Access to relevant customer data, subject to a non-disclosure agreement (NDA)
  • Experience with major enterprise customer
  • Startup and innovative environment

Desired Deliverables:

  • Data cleansing, exploratory analysis and feature construction
  • Model development and evaluation
  • Documentation

Grant Thornton

Measuring the Impact of Wildfire Risk on Property Market Prices

Objective:

Develop a predictive model to measure the impact of wildfire risk on property market prices.

Data:

The project will utilize a combination of the following datasets:

  1. Historical property market data: Includes information on property sales transactions (e.g., market value, sale price, sale date) and property characteristics (e.g., location, property type, size) from the past decade.
  2. Wildfire risk data: Historical data on wildfire occurrences, intensity, and affected areas, sourced from national and regional fire management agencies.
  3. Environmental and climatic data: Data on weather patterns, temperature, and other relevant environmental factors from meteorological services.
  4. Socioeconomic data: Information on population density, economic activity, and other relevant factors from public databases.
  5. Supplementary data: Publicly available data from real estate and statistical services related to property price indexes.

Deliverables:

  1. Data Collection & Integration:
  • Gather and compile datasets from various sources.
  • Ensure proper integration and alignment of the data for analysis.
  1. Data Cleansing, Exploratory Analysis & Feature Construction:
  • Clean and preprocess the data to handle missing values, outliers, and inconsistencies.
  • Conduct exploratory data analysis (EDA) to understand the relationships and patterns within the data.
  • Construct relevant features that capture the influence of wildfire risk and other factors on property prices.
  1. Model Development & Evaluation:
  • Develop predictive models using appropriate machine learning algorithms to estimate the impact of wildfire risk on property market prices.
  • Evaluate the models using standard metrics and validate their performance on test datasets.
  • Compare different models to identify the most effective approach for prediction.
  1. Documentation:
  • Provide comprehensive documentation of the entire process, including data sources, methodologies, model development, and evaluation results.
  • Include detailed explanations of the assumptions, limitations, and potential implications of the findings.
  1. Presentation:
  • Prepare a presentation summarizing the project’s objectives, methodologies, key findings, and recommendations.
  • Present the results to the stakeholders, highlighting the practical applications of the model and its potential impact on property market assessments.

Resources:

  • Access to relevant databases and data sources
  • Computational resources for data processing and model training
  • Software tools for data analysis, visualization, and model development (e.g., R, Python, Shiny, Flask)

Expected Outcome:

The project aims to provide a robust predictive model that accurately measures the impact of wildfire risk on property market prices. Interactive visualization presentation of results is expected in the form of a WebApp (through Shiny, Flask etc.). The findings will help users to make informed decisions regarding property investments and risk management.

Hellas Direct

Machine Learning for Fraud Detection in Insurance Claims

Background

Insurance claim fraud is a global issue that drains billions of dollars from the industry each year, escalating costs for both insurers and policyholders. In Greece, for instance, fraudulent claims tally up to €200 million annually, accounting for about 10% of all claim payouts. These challenges are not unique to Greece but are mirrored worldwide, necessitating a unified and innovative approach across the sector to tackle this pervasive problem.

Insurance claim fraud is defined as any deliberate and misleading act or omission by any individual or legal entity aimed at gaining an unlawful financial benefit for themselves or facilitating such gain for others within the context of a valid or invalid insurance contract. 

The need for advanced data analytics and machine learning in fraud detection is becoming increasingly critical as fraudsters employ more sophisticated methods to circumvent traditional detection systems. Globally, insurers are turning to technology to gain an edge against these tactics, utilising big data and predictive analytics to spot irregularities and inconsistencies in claim submissions. This technological shift represents a move from reactive to proactive fraud management, where potential frauds are flagged and investigated before they result in financial loss.

Moreover, the ongoing digital transformation in the insurance industry highlights the importance of continuous innovation in fraud detection systems. The integration of AI and machine learning technologies enhances fraud detection accuracy and speeds up the process, allowing insurers to handle claims more efficiently while reducing the chances of fraudulent claims slipping through the net.

Engaging with this project provides an opportunity to contribute to significant advancements in improving this procedure and applying cutting-edge data science techniques to real-world problems. This initiative is an excellent initiative for aspiring data scientists to make a tangible difference, helping to shape the future of an industry on which everyone relies.

Objectives

Briefly, the core objectives of this project are defined as:

  1. Development of a Predictive Statistical Model: Conceptual Design and Implementation of a robust Machine Learning Model towards the accurate identification of irregular patterns and potential fraud in claim submissions.
  2. Data Profiling, Analysis and Processing: Processing and Analysis of historical claims to pinpoint key features and patterns associated with fraudulent activities by encapsulating and adhering to business knowledge from experts.
  3. Imbalanced Model Optimization: Design, Implementation and Optimization of predictive models in a real-world case scenario with imbalanced datasets by experimenting, employing, and benchmarking diverse techniques.
  4. Stakeholder Management, with the view of engaging the prototype solution with the direct Subject Matter Experts (SMEs) on a high-level commercial discussion.
  5. Model Validation, establishing functionalities and methodologies for actual measurable outcomes.

Methodology

The principal methodology pipeline will consist of:

  1. Data Preprocessing: The preprocessing phase will involve cleaning data, standardising formats, and engineering new features that highlight patterns indicative of fraudulent behaviour.
  2. Exploratory Data Analysis (EDA): This phase will involve statistical analysis to understand distributions and identify outliers, correlation analysis to explore relationships between variables, and the use of visualisation tools to present these insights clearly. Effectively communicating findings helps refine the data for modelling and gain stakeholder buy-in.
  3. Model Development: Several models will be evaluated, focusing on handling the corresponding feature distributions and the dataset’s imbalance.
  4.   Model Validation and Testing: The model’s performance will be validated using several metrics, always driving the stakeholders’ strategic business decisions. This entails fine-tuning the trade-off between different metrics that map directly into associated business metrics. 
  5. Stakeholder Engagement and Feedback: Feedback loops with end-users in the respective departments will be conducted, with the results officially presented. Their insights will guide iterative improvements to the model and its interfaces, ensuring the tool remains user-friendly and aligned with operational needs.

Resources

All the necessary resources for the completion of the project will be delivered, such as:

  • Respective anonymized claims data of the last months, on top of additionally demanded datasets
  • Interaction with the team with regular meetings, according to the demanded application
  • Technical leadership

KPMG

Title: Customer Behavior prediction

Project Description:

The project focuses on analyzing customer data of a Telecommunication company, specifically related to their mobile plans. The data contains information related to customers such as monthly payments, mobile usage, contract information and competitive data from other telecom vendors. The objective is to develop a methodology to process the data, extract meaningful insights through analysis and create new features that will be used to build predictive models. The models will be trained to predict customer behavior that will lead to switch to a different mobile plan. This project will provide an opportunity to gain practical experience in Data Analysis, Feature Engineering, and Machine Learning, while also gaining valuable insights into solving real-world problems in the telecommunications industry.

Resources:

Our company will provide:

  • contact person that will provide guidance and support
  • access to all necessary data (all covered by NDA)

Desired Deliverables:

  • Exploratory data analysis report that showcases meaningful insights (e.g., summary statistics, correlation analysis, customer segmentation)
  • Predictions on customer switch plan for different time intervals
  • Model evaluation report
  • Final Project report


Milliman

Health Data Analysis and/or Pricing model development

Project Description:

This dynamic and interactive project will provide students with hands-on experience applying their knowledge in data science to the actuarial field, particularly in data analysis and optimization as well as Health Insurance Pricing model development. It features a unique collaboration between Milliman’s local and international teams, ETHNIKI, THE, HELLENIC GENERAL INSURANCE CO. S.A., one of Greece’s largest insurance companies, and the University of Cyprus that successfully promotes the project.

To get a flavor of what the project will include, have a look into the below. (Note that this is not an exhaustive list of tasks):

  • Data preparation: UCY students will need to collect health related historical claims data* from different sources (provided by Milliman), organize the data in a way that will enable easy access and analysis, clean the data by removing any inconsistencies, duplicates or other that could diminish their reliability.
  • Data enhancement: To improve, if possible, the quality of the data by introducing additional relevant information.
  •  Inflation analysis & Assumption building: For this task students will extract the trends and aim to model the inflation inherent in the historical claims of the company.
  • Model design: Create a robust pricing model that would utilize the cleaned and enhanced data incorporating assumptions for inflation (and any other trends identified). Use of GLM or other approaches that the students will find useful can be made.

*The intention will be to anonymize all data utilized for this Capstone project.

Resources:

UCY students will have the opportunity to work with real life Insurance market data that will include information regarding historical claims with a wealth of parameters around each claim. Milliman will provide guidance and support throughout the duration of the project and will ensure that materials provided to the students are adequate to perform the tasks. 

Desired Deliverables:

  • The deliverable will include:
    • a short report and/or presentation that would summarize the work performed (including conclusions)
    • working files used (including pricing model)
  • To be further discussed with the UCY students and supervisor upon project kick-off


Suite5

Scope

Today, accurate flight delay prediction is among the most critical and challenging problems when scheduling and managing flights for all involved aviation data value chain stakeholders. As part of the capstone project, the flight delay prediction problem will be addressed from the perspective of an airport and for a day-ahead time horizon in terms of:

  • The exact (positive/negative) forecasted delays (in terms of minutes) for both the departure and arrival flights, with minimal errors/high accuracy (to the extent possible).
  • The forecasted flight delays (for both the departure and arrival flights) in the following classes: (a) early arrival/departure (before the scheduled time), (b) delay equal to or less than 15’ (15:59), (c) delay between 16’ (16:00) and 30’ (30:59), (d) between 31’ (31:00) and 45’ (45:59), (e) delay over 46’ (46:00) along with their probability, e.g. flight X has 58% probability of delay over 46’, 20% delay between 16’ and 30’, 12% delay between 30’ and 45’, 10% delay less than 15’. 
  • A problematic flight forecast index including the problematic (cancelled) flights (e.g. flight X with 45% risk of cancellation due to weather conditions based on historical data)

The flight delay forecasts need to be accompanied by appropriate explanations in order to shed light into the provided forecasts, so Explainable AI (XAI) techniques (i.e. for feature relevance and counterfactual explanations) need to be leveraged. 

The aviation data that will be used as part of the capstone project are confidential flight data from an airport in Europe. 

Deliverables

D1. XAI Models for the Flight Delay Forecasting (Classification) & Problematic Flight Forecast Index – end of June 2024

D2. XAI Models for the Flight Delay Forecasting (Regression) & XAI Models Updates for the Flight Delay Forecasting (Classification) & Problematic Flight Forecast Index – end of July 2024



_Summer 2023

Bank of Cyprus - Project 1

As part of its risk assessment, the bank is required to predict the returns from the sales of its real estate collaterals. One of the parameters that determines those returns is the ‘recovery rate’, i.e., the % of the property market value that the Bank will recover by selling the property.

Objective: develop a predictive model for recovery rate parameter of collaterals, using the Bank’s most recent internal database.

Data: contains information on collateral sales transactions (e.g., open market value, sale price, sale date) and property characteristics (e.g., location, property type, size), for historic collaterals onboarded from 2016 onwards. Publicly available data from CY Statistical Service related to property price indexes can be used to supplement the analysis.

Deliverables:

  • Data cleansing, exploratory analysis of & feature construction
  • Model development & evaluation
  • Documentation

Bank of Cyprus - Project 2

Short description: The project aims to develop and evaluate a solution to automate the processes of model validation and identify variables that canchallenge and enhance the performance of the Bank’s behavioural credit models.

Objective: The objective of the project is to improve the efficiency and effectiveness of the model validation function in the Bank, and to ensure compliance with validation unit’s internal procedures and methods.

Data: The project will use the following data sources:

  • The Bank’s internal data on its behavioural credit models, such as model specifications, parameters, inputs, outputs, assumptions, limitations, and performance metrics.
  • The Bank’s internal data on its customers, such as credit scores, loan amounts, repayment history, default rates, and other relevant variables.
  • External data sources, such as market data, macroeconomic indicators, industry benchmarks, and peer comparisons.

Deliverables: The expected outcomes and deliverables of the project are:

  • Automation of the processes of model validation, such as data processing, model testing, back-testing, benchmarking, sensitivity analysis, and reporting.
  • Development of a machine learning algorithm that can identify variables that can enhance the performance of the Bank’s behavioural credit models, such as new features, interactions, transformations, or selection methods.
  • A report that documents the data analysis, machine learning pipeline development, machine learning pipeline evaluation, and machine learning pipeline deployment processes and results.
  • A presentation that showcases the project objectives, methods, findings, applications, and implications for the Bank’s model validation unit.

Catalink Ltd

Project Title: Cancer Prevention AI tools

Project Description: The Cancer Prevention AI tools project is a cross-functional capstone project that aims to develop a mobile application that monitors the daily living of people and suggests new food, activity, and lifestyle routines to reduce the percentage of cancer disease. The application will be built by a team of two or three students with expertise in data analytics (visual, textual, multisensory signals analysis) and one in business development (optional).

The data analytics student(s) will be responsible for developing AI models that will analyse user data, including daily food intake, physical activity, mood estimation and lifestyle habits, to provide personalized recommendations for reducing the risk of cancer. The AI models will use machine learning algorithms to identify patterns and correlations in the data and suggest changes to the user’s routine accordingly. The Data Analytics student(s) should have expertise in programming languages such as Python, and be familiar with relevant libraries and tools such as NumPy, Pandas, Scikit-learn, TensorFlow, Keras, and Matplotlib. They should also have experience in machine learning algorithms, deep learning models, and data visualization techniques.

Overall, the Cancer Prevention AI tools project will provide a valuable tools for people to monitor their daily habits and receive personalized recommendations for reducing their risk of cancer. The project will also provide valuable experience for the student team members in data analytics, and business development, as well as a tangible product to showcase their skills to potential employers (optional).

Deliverables:

AI Models: The project should deliver appropriate AI models to analyse user data, including daily food intake, physical activity, and lifestyle habits, to provide personalized recommendations for reducing the risk of cancer. The model should use machine learning algorithms to identify patterns and correlations in the data and suggest changes to the user’s routine accordingly.

Business Plan: The project should include a business plan that outlines the target market for the AI tools, the competitive landscape, and the marketing and sales strategy. The plan should also include financial projections and revenue models.

Technical Documentation: The project should include technical documentation that describes the AI tools design, and functionality, as well as instructions for installing and running the AI tools.

Presentation: The project team should prepare a final presentation that showcases the AI models, and the business plan. The presentation should include a demonstration of the AI tools’ functionality,  and the business plan’s revenue projections.

 

CYENS Center of Excellence - Project 1

Project title:

Use of ML Techniques to Enable Intrusion Detection in IoT Networks

Background:

Networks are vulnerable to costly attacks. Thus, the ability to detect these intrusions early on and minimize their impact is imperative to the financial security and reputation of an institution.

There are two mainstream systems of intrusion detection (IDS), signature-based and anomaly-based IDS. Signature-based IDS identify intrusions by referencing a database of known identity, or signature, for each of the previous intrusion events. Anomaly-based IDS attempt to identify intrusions by referencing a baseline or learned patterns of normal behavior. Under this approach, deviations from the baseline are considered intrusions.

Project aims:

 

This MSc project will utilize existing work on Lightweight Intrusion Detection for Wireless Sensor Networks  (using BLR, SVM, SOM, Isol. Trees) and extend it, considering the characteristics of IoT networks, new attacks, new topologies, and especially new classification algorithms

Reading / Datasets

Works by V. Vassiliou and C. Ioannou (University of Cyprus)

Find suitable datasets from https://www.kaggle.com/

Contact Details

 

Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy

CYENS Center of Excellence - Project 2

Project title:

Predicting YouTube Trending Video

Background:

 

A new trend in 5G and 6G networks is using the network edge (base station) for caching and processing and for supporting the quality of service and the experience of end users consuming visual content and interactive media.

Project aims:

 

Within this project, will provide the ability to group and predict users’ and/or content’s needs.  One way is to predict videos trending on social media.

Readings / Datasets

Youtube Video Info  https://www.kaggle.com/datasets/datasnaek/youtube-new

Contact Details

 

Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy

 

CYENS Center of Excellence - Project 3

Project title:

Predicting Motor-vehicle Accidents

Background:

 

Motor vehicle crashes cause loss of life, property and finances. Vehicle accidents are a focus of traffic safety research, uncovering useful information that can be directly applied to reduce these losses.

Traditionally, modeling crash events has been done using machine learning techniques, considering crash level variables, such as roadway characteristics, lighting conditions, weather conditions and the prevalence of drugs or alcohol.

 

Project aims:

 

Incorporate crash report data, road data, and demographic data to better understand crash locations and the associated cost.  Use a mixed linear modeling technique that enables data fusion in a principled way to build a better predictive model. Analyze the natural clustering of events in space by different geographic levels.

Readings / Datasets

Need to get relevant information from the Cyprus Police and the Association of Insurance Companies. 

Contact Details

 

Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy

CYENS Center of Excellence - Project 4

Project title:

Preventive to Predictive Maintenance

Background:

 

Maintenance is an integral component of operating manufacturing equipment.

Preventive maintenance occurs on the same schedule every cycle — whether or not the upkeep is actually needed. Preventive maintenance is designed to keep parts in good repair but does not take the state of a component or process into account.

Predictive maintenance occurs as needed, drawing on real-time collection and analysis of machine operation data to identify issues at the nascent stage before they can interrupt production. With predictive maintenance, repairs happen during machine operation and address an actual problem. If a shutdown is required, it will be shorter and more targeted.

While the planned downtime in preventive maintenance may be inconvenient and represents a decrease in overall capacity availability, it’s highly preferable to the unplanned downtime of reactive maintenance, where costs and duration may be unknown until the problem is diagnosed and addressed.

Preventive to Predicitve Maintenance is about the transition of a preventive maintenance strategy to a predictive maintenance strategy of a replaceable part

Project aims:

 

The objective of the project is to use the associated detailed dataset is to precisely predict the remaining useful life (RUL) of the element in question, so a transition to predictive maintenance is made possible. 

Readings / Datasets

https://www.kaggle.com/datasets/prognosticshse/preventive-to-predicitve-maintenance

Contact Details

 

Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy

 

CYENS Center of Excellence - Project 5

Project title:

Analysis and Visualization of Mobile Cellular Telephony Network Coverage and Electromagnetic Measurements

Background:

 

An interesting challenge in 5G and 6G Mobile Cellular Telephony is the need to have a large(r) number of small(er) base stations to achieve the data rates, delays and number of users expected.  The Mobile Network Operators report to the Telecommunication Regulators of each country a number of parameters, including coverage, location of stations and radiation levels.

 

Project aims:

 

The objective of the project is to collate information available in different systems/platforms and generate an up-to-date map of coverage and mobile telephony stations, alongside the information gathered through periodic and ad-hoc measurements. These will be related with measurements of mobile telephony quality of service and shown on a common reference map with the ability to explore the changes during the years.

Readings / Datasets

 

Open data from the Department of Electronic Communications. Information from the ICT market observatory of OCECPR and other public data.

http://www.emf.mcw.gov.cy/emf/?page=emfmeasurements

https://www.data.gov.cy/search/field_topic/%CE%B5%CF%80%CE%B9%CF%83%CF%84%CE%AE%CE%BC%CE%B7-%CE%BA%CE%B1%CE%B9-%CF%84%CE%B5%CF%87%CE%BD%CE%BF%CE%BB%CE%BF%CE%B3%CE%AF%CE%B1-40

 

Contact Details

 

Dr. Vasos Vassiliou, Smart Networked Systems Research Group Leader, Associate Professor, Computer Science, University of Cyprus vasosv@cyens.org.cy

Economics Research Center (CypERC) - Project 1

Title: Measuring consumer price inflation in Cyprus

We calculate Cyprus inflation in terms of the Consumer Price Index (CPI), which is compiled using the online prices (big data) for a pre-selected/fixed and representative basket of goods and services. In doing so, we employ ‘web scraping’ algorithms to visit the websites/pages of large retailers in Cyprus and collect and store the prices of goods and services available online. We then implement standard techniques and proprietary methodologies to calculate price statistics and indices. This work is inspired by the Billion Prices Project, an academic initiative at MIT and Harvard that used prices collected from hundreds of online large retailers around the world on a daily basis to conduct research in macro and international economics.

Economics Research Center (CypERC) - Project 2

Title: Studying the crude oil price pass-through into fuel and/or consumer prices in Cyprus

We leverage a high-frequency (weekly) online price series dataset in an econometric framework for pass-through estimation, forecasting and policy making. We aim to quantify how (in what degree and time) the fuel and/or consumer prices respond to an oil price shock and provide policy makers with important implications and useful information regarding consumer welfare.

Hellas Direct

Title:  Create an end-to-end MLops pipeline for a specific model (“anti-scrapping model”)

Duration: Hard deadline the end of August

Team: 3 Data Science MSc Students with backgrounds: 2 in CS, 1 in Statistics, plus one supervisor on the UCY side. The supervisor can offer weekly one-hour coaching and review sessions to the students.

Project Description:

Scope of this project is to create an end-to-end MLops pipeline for a specific model (detecting competitors trying to scrape/crawl prices from our website). The students will work with the strong guidance of our Data science team to design/implement/deploy an ML model to production.

Resources:

Our company will provide

  • quotations and sales data of the last 6 months (all covered by NDA).
  • Strong interaction with the our team.
  • A lot of positive energy.

 

Desired Deliverables:

  • A end-to-end MLops working solution (in python)
  • A final report of 3 slides

KPMG

Title: Loyalty Scheme – Client Churn Prediction

Duration: 2 months (June 2023 – July 2023)

Team Members:

Data Scientist with a background in Computer Science: Responsible for data preprocessing, feature engineering, and model development.

Statistician: Responsible for statistical analysis, model evaluation, and validation.

Business Analyst: Responsible for understanding the business requirements, interpreting the results, and providing insights for decision-making.

Supervisor: A subject matter expert who will provide guidance and support throughout the project.

Project Description:

This project aims to predict the churn of COMPANY’s loyalty customers by analyzing customer data and building predictive models. The team will work closely with COMPANY’s internal stakeholders to gather relevant data related to customer behavior, transactions, and engagement metrics. The collected data will be processed, cleaned, and transformed to create meaningful features. The team will then explore various machine learning algorithms to develop predictive models that can identify potential churners accurately.

The project will involve conducting an in-depth exploratory data analysis to uncover insights and patterns within the data. The team will perform feature engineering to derive additional relevant features from the existing dataset. These features will be used to train and evaluate predictive models using techniques such as logistic regression, decision trees, random forests, or neural networks. The models will be assessed based on metrics like accuracy, precision, and F1-score.

Resources:

Our company will provide:

A designated contact person who will offer guidance, support, and domain expertise.

Access to relevant customer data, subject to a non-disclosure agreement (NDA).

Desired Deliverables:

Churn prediction models and evaluation report: The report should detail the methodology, model performance metrics, and provide insights into key features driving churn.

Final Project Report: A comprehensive document summarizing the entire project, including data preprocessing, model development, insights, and recommendations for COMPANY.

Presentation: A final presentation summarizing the project, highlighting the key findings, and providing actionable recommendations for COMPANY to reduce churn among loyalty customers.

Stremble Ventures

Title: Quality control, cleaning, and analyses of wearable data in the context of clinical trials.

Wearables have enabled the non-intrusive monitoring of subjects in medical research studies during their normal daily lives. At Stremble we have already expanded our in-house analytics platform to automatically collect data from several of our clinical trials and studies in flat JSON files. However, cleaning, organizing, and analyzing these data is a challenge. Therefore, standardizing the quality control storage and access to the data would benefit many ongoing cancer research projects.

Project Aims:

  1. Establish a cleaning and quality control workflow that accumulates data ready for analyses in a database.
  2. Develop some analytical routines that utilize mathematics, statistics, and computer science to test hypothesis, explore the data utilizing AI approaches and produce visualizations.
  3. Integrate the analyses workflow.

Skills:

            Students in Computer Science, Mathematics or Statistics would be eligible for this position.

            Programming/Scripting knowledge preferably Python

            R and ideally Shinny for the analyses.

Who we are:

Stremble Ventures is a company based in Limassol, Cyprus established in 2011. We offer contract research to companies and research institutions around the world with a focus on Bioinformatics and Computational Biology. We have several ongoing EU, RIF, and commercially funded projects.

Tickmill

Tickmill is a retail FX broker operating globally and employing a dedicated team of over 250 professionals. With an average monthly trading volume surpassing $150 billion, Tickmill stands as a major player in the financial industry. One of its key assets is its own quantitative research team that has developed proprietary trading systems. These systems leverage extensive tabular data stored across diverse databases. The team’s primary objective is to make informed investment decisions by analyzing unconventional data. For instance, is it possible to predict Starbucks’ stock price if we know the number of people visiting their stores? If so, would combining this information with the average amount spent by customers in Starbucks stores yield improved results?  

During the capstone project, students who join Tickmill’s Quantitative Research team will actively participate in tasks such as data mining, data storage utilizing various database types based on their specific requirements, and, notably, data analysis. In the financial industry, data analysis poses significant challenges for two primary reasons. Firstly, unlike voice, image, or video data, financial data cannot be generated or created. Secondly, the signal-to-noise ratio is typically low, which significantly contributes to the complexity of this particular task in machine learning and data science.

In this project, participating students will undergo a two-week training program focused on the financial industry and trading. Prior knowledge in this field is not required. Following the training, students will delve into various types of data and explore the corresponding data analysis within that domain. They will have the flexibility to use their preferred data analysis tools (although we utilize Jupyter notebook, they are free to choose the tool they are most comfortable with) to conduct their analyses. By the conclusion of the project, students will be equipped with the ability to determine which types of data are valuable for predicting the price change of financial assets and which data contain excessive noise. This understanding encompasses the concepts of correlation and causation, although it is not limited to these factors alone. In addition to this phase, which we refer to as the first stage of data analysis, students will have the opportunity to create their own features and explore the freedom to combine different types of data in order to derive meaningful insights and conclusions.

Due to confidentiality agreements, students will not be allowed to disclose data vendors names in their capstone reports.