Modern Statistics for Interdisciplinary Omics and Big Data
The end of the IMforFUTURE project was celebrated with a final workshop focused on mathematical and statistical methodologies applied to omics research. This workshop was held online, from the 28th to the 30th of June 2021.
The workshop was a joint event of the IMforFUTURE network meeting and the Leeds Annual Statistical Research (LASR) Workshop. There was also be a celebration of Kanti Mardia's 85th birthday.
We were pleased to have the following keynote speakers: Cornelia van Duijn, Hongzhe Li, and Janice Scealy.
The workshop also included a short course on omics integration in R and a mindfulness session on Zoom Fatigue.
Congratulations to Ruheyan Nuermaimaiti for winning the poster prize!
Speakers
The following keynote and invited speakers gave presentations at the meeting:
Keynote speakers
Hongzhe Li is a Professor in biostatistics, epidemiology, and informatics at the Perelman School of Medicine, University of Pennsylvania. He is the head of the Statistical Genetics and Genomics Laboratory. His research is mostly motivated by problems in statistical genetics and genomics, including methods for family-based genetic linkage and association analysis, methods for admixture mapping, methods for genome-wide association analysis, methods for analysis of microarray time course gene expression data, high dimensional regression analysis for genomic data, methods for copy number variation analysis, and methods for analysis of next generation sequence data. He has extensively published in both statistical methodological research in top statistics/biostatistics journals (JASA, AOS, AOAS, Biometrika, Biometrics, Biostatistics etc) and in top genetics journals (AJHG, Plos Genetics, etc) and collaborative research in top scientific journals (Science, NEJM, Nature, Nature Genetics, PNAS, Developmental Cell etc).
Janice Scealy is a senior lecturer of statistics at the Australian National University and a 2018 Australian Research Council Discovery Early Career Researcher Award Fellow. She was awarded a prominent research recognition - the Moran Medal in 2021 by the Australian Academy of Science for her substantial contributions to statistical science. Her research interests include: compositional data analysis; directional statistics, shape analysis and statistics for manifold valued data; robust statistics; model selection in linear mixed models; analysis of sample survey data; geostatistics; and analysis of spatial data. Her work has been funded by multiple Australian Research Council programs. She has published in leading academic journals including Journal of the American Statistical Association, Journal of the Royal Statistical Society Series B, Australian and New Zealand Journal of Statistics, Statistics and Computing, and Statistical Science. She is currently a chief investigator on the Australian Research Council Discovery Project entitled A New Generation of Palaeomagnetic Statistics.
Cornelia van Duijn is a Professor of epidemiology at Nuffield Department of Population Health and a Fellow of St Cross College, Oxford. Her research within the Oxford Big Data Institute focuses on large-scale –omics studies of neurodegenerative disorders. She also studies systemic vascular, endocrine and gastrointestinal pathology that is relevant for brain and ocular function. Her current research portfolio includes cross-omics research integrating (epi)genetic, transcriptomic, proteomic, metabolomic and microbiome data of epidemiological cohorts with state of the art brain imaging and cellular model systems. Over the years, she has been a leading figure in several international consortia including ENGAGE, CHARGE , IGAP, ADSP and IGGC . At present, she is the leader of two major consortia: the Horizon2020 CoSTREAM consortium aiming to understand the link between stroke and Alzheimer disease and the MEMORABEL Gut-Brain consortium aiming to unravel the role of the gut microbiome in Alzheimer disease and brain pathology. She further leads the human proteomics and metabolomics discovery research in the Innovative Medicine Initiative (IMI) ADAPTED program, which aims to identify new Alzheimer medicines through understanding the function of the APOE gene.
Invited speakers
Linda Chaba, Strathmore University, Kenya
Said el Bouhaddani, UMC Utrecht, The Netherlands
Ian Dryden, University of Nottingham, UK
Thomas Hamelryck, University of Copenhagen, Denmark
Peter Jupp, University of St Andrews, UK
John Kent, University of Leeds, UK
Lucija Klarić, University of Edinburgh, UK
Renee Ruhaak, Leiden University Medical Center, The Netherlands
Programme
Modern Statistics for Interdisciplinary
Omics and Big Data
[All times are UK BST]
Monday 28 June: Omics + Epidemiology
[ Session chairs: Arianna Landini and Azra Frkatović ]
| 09.00 - 09.15 | Introduction, Jeanine Houwing-Duistermaat |
| 09.15 - 09.30 | Official opening: Kurt Langfeld, Head of School of Mathematics |
| 09.30 - 10.15 | Cornelia van Duijn (keynote speaker) "Multi-omics studies in complex diseases" |
| 10.15 - 10.45 | Renee Ruhaak (invited speaker) "The challenging journey of biomarker translation targeting unmet clinical needs" |
| 10.45 - 11.00 | Break |
| 11.00 - 11.15 | Introduction to posters, John Kent |
| 11.15 - 12.00 | Poster session 1
|
| 12.00 - 12.30 | Lunch |
| 12.30 - 13.15 | Poster session 2
|
| 13.15 - 13.45 | Lucija Klarić (invited speaker) "Molecular Mechanisms from Omics Data: Applications in Glycomics" |
| 13.45 - 15.00 | Contributed talks - Session 1 [Moderator: Frances Williams]
|
| 15.00 - 15.15 | Break |
| 15.15 - 17.15 | IMforFUTURE meeting (closed) PhD students Leeds (closed) |
Tuesday 29 June: Omics + Methods + Bioinformatics
[ Session chairs: Morning - Luisa Cutillo and Seppo Virtanen;
Afternoon - Annah Muli and Zhujie Gu ]
| 09.00 - 09.15 | Introduction, Jeanine Houwing-Duistermaat |
| 09.15 - 10.00 | Janice Scealy (keynote speaker) "Score matching for microbiome compositional data" |
| 10.00 - 10.30 | Peter Jupp (invited speaker) "Statistics of crystallographic orientations" [ NO SOCIAL MEDIA ] |
| 10.30 - 10.45 | Break |
| 10.45 - 11.15 | John Kent (invited speaker) "Flexible and tractable models for directional statistics" |
| 11.15 - 11.45 | Thomas Hamelryck (invited speaker) "Deep probabilistic programming for protein structure prediction" |
| 11.45 - 12.15 | Ian Dryden (invited speaker) "Manifold valued data analysis of samples of networks" |
| 12.15 - 12.45 | Kanti Mardia "LASR Workshops and my Journey to Statistics on Manifolds" |
| 12.45 - 13.30 | Lunch |
| 13.30 - 14.15 | Hongzhe Li (keynote speaker) "Interrogating the Gut Microbiome: Estimation of Bacterial Growth Rates and Prediction of Biosynthetic Gene Clusters" |
| 14.15 - 14.30 | Break |
| 14.30 - 15.45 | Contributed talks - Session 2 [Moderator: Charles Taylor]
|
| 15.45 - 16.30 | Poster session 3 + break
|
| 16.30 - 17.30 | Mindfulness session, Kitty Wheater Dr Kitty Wheater is a mindfulness practitioner, social and medical anthropologist and writer. Her research focuses on embodiment and ethics in mindfulness-based interventions, and she is broadly interested in the anthropology of medicine and religion. Dr Wheater will discuss why online meetings can be more tiring and draining than in-person meetings, a phenomenon known as 'Zoom fatigue'. She then will guide participants through mindfulness exercises to cope with the new online meeting routine established by the COVID-19 pandemic. |
Wednesday 30 June: Big Data + the Future
[ Session chairs: Ruheyan Nuermaimaiti and Minzhen Xie ]
| 09.40 | Welcome |
| 09.45 - 10.15 | Said el Bouhaddani (invited speaker) "Statistical integration of multiple omics datasets with probabilistic PLS methods" |
| 10.15 - 10.45 | Break |
| 10.45 - 11.15 | Linda Chaba (invited speaker) "A machine learning based approach to cancer classification using RNA-Seq data" |
| 11.15 - 12.30 | Contributed talks - Session 3 [Moderator: Hae-Won Uh]
|
| 12.30 - 12.45 | Closing remarks |
| 12.45 - 13.30 | Lunch |
| 13.30 - 16.30 | Short course on omics, Said el Bouhaddani and Zhujie Gu |
R course on omics integration
This three-hour short course is aimed at researchers who want to learn how to jointly analyse two omics datasets and interpret the results. We will be using the “OmicsPLS” R package with real data examples.
The course consists of two parts. The first two-hour session is a mix of theory and practice. You will be introduced to the basics of data integration approaches and their sparse variants. We will also discuss how to incorporate external biological knowledge. You will practice with these approaches using OmicsPLS and try out different visualisation tools to interpret the results.
The second one-hour breakout session will be hands-on where you apply what you’ve learned to two omics datasets.

There will be a brief rejoinder at the end of the session.
By the end of this course, you should be able to:
- Have a deeper understanding of data integration methods (PLS, O2PLS, GO2PLS)
- Understand how feature selection works
- Know how to incorporate external biological information
- Implement O2PLS/GO2PLS on your own data with “OmicsPLS” R package
- Interpret and visualise O2PLS/GO2PLS results
Meeting committees
Contact email: [email protected]
Organising committee:
Jessica Brennan, University of Leeds
Helen Copeland, University of Leeds
Arief Gusnanto, University of Leeds
Jeanine Houwing-Duistermaat, University of Leeds (chair)
Arianna Landini, University of Edinburgh
Annah Muli, University of Leeds
Ruheyan Nuermaimaiti, University of Leeds
Hae-Won Uh, UMC Utrecht
Minzhen Xie, University of Leeds
Scientific committee:
Arief Gusnanto, University of Leeds
Jeanine Houwing-Duistermaat, University of Leeds
Hae-Won Uh, UMC Utrecht
Frances Williams, King's College London
Poster committee:
René Eijkemans, UMC Utrecht (chair)
Haiyan Liu, University of Leeds
Claudia Sala, University of Bologna

