CAMBRIDGEGENOMIC MEDICINE

GM5 · MPHIL & MRES · 2026–27

Bioinformatics, interpretation and data quality assurance in genomic analysis

Previously: Advanced Bioinformatics from genomes to systems. About the module name changes

A new name for a familiar module

This is the module previously called Advanced Bioinformatics — from genomes to systems. The name has changed; the module retains its content.

About this module

The aim of this module is to equip students with the skills to analyse and interpret sequencing data, including variant identification, annotation and filtering using appropriate computational and statistical methods. Students will also critically assess the challenges of managing large genomic datasets and the role of databases in variant annotation. This module covers the fundamental principles of bioinformatics with a focus on its impact in both clinical genomics and in research. Students will gain hands-on experience in analysing sequencing data, including quality assessment, alignment to reference genomes, variant calling, and annotation. They will also learn to apply computational skills and statistical methods to handle large genomic datasets and critically evaluate the use of different databases for variant annotation. Through a combination of theoretical sessions and practical exercises, students will develop the essential skills to analyse and interpret genomic data, preparing them to address current and future challenges in managing and utilizing large-scale genomic information in clinical and research settings.

15 credits · Module Leads: Dr Tim Hearn (University of Cambridge) and Dr Kenneth Langlands (University of Cambridge)

Before the module

Suggested preparation

Module-specific materials and instructions are provided through your course Moodle/VLE. These refreshers are optional support, not additional assessment requirements.

What you will study

  • bioinformatics principles and their application in clinical genomics
  • techniques for assessing the quality of raw sequencing data
  • common quality metrics (e.g., base quality scores, read length distribution, coverage) and the use of appropriate tools and software
  • genome alignment and variant calling:
    • aligning sequencing data to reference genomes using tools like BWA and Bowtie.
    • variant calling methods: single nucleotide polymorphisms (SNPs), insertions and deletions (indels)
    • best practice in variant calling and filtering strategies to identify clinically- relevant variants
  • variant annotation and interpretation:
    • annotating variants using databases like dbSNP , ClinVar, and Ensembl.
    • tools for variant annotation
    • interpretation of annotated variants and their clinical significance.
  • computational and statistical methods in genomic data analysis:
    • basic computational skills for handling large datasets (e.g., command-line tools, scripting).
    • statistical methods for variant analysis and interpretation.
    • data normalization, variant frequency calculations and significance testing
  • databases and resources for genomic data:
    • overview of major genomic databases and their use in variant annotation (e.g., 1000 Genomes, ExAC, gnomAD)
    • evaluating the reliability and relevance of different databases
    • integrating data from multiple sources for comprehensive variant interpretation.
  • challenges in managing large genomic datasets:
    • current and future challenges in genomic data management (e.g., storage, processing power, data security).
    • techniques for efficient data handling and processing (e.g., cloud computing, parallel processing).
    • strategies for maintaining data integrity and privacy in large-scale genomic research.
    • ethical, legal, and social issues in the use of genomic data.
  • future trends in genomic data analysis:
    • emerging technologies and methodologies in genomic data analysis.
    • the role of artificial intelligence and machine learning in variant interpretation.

Learning objectives

By the end of this module you will be able to:

  • analyse the quality of sequencing data, align to a reference genome, call and annotate variants, and apply filtering strategies to identify variants in sequencing data
  • appraise the use of different databases in variant annotation
  • apply relevant basic computational skills and statistical methods to handle and analyse genomic data.
  • examine the current and future challenges of managing and manipulating large genomic data sets

Teaching

30 November–4 December 2026

School of Clinical Medicine, Hills Road, Cambridge CB2 0QQ

Dates follow the 2026–27 handbook. Check your course communications for timetable updates.

Learning objectives follow the 2026–27 curriculum, including programme-team updates.