Bioinformatics Scientist with over a decade of experience designing, building, and deploying production-grade NGS workflows and clinical genomics platforms for multi-omics research.
Specialized in analysing short-read (Illumina) and long-read (PacBio, Nanopore) sequencing data (WGS, WES, RNA-seq, germline/somatic variant calling).
Proficient in Nextflow/WDL, Python, Bash, Docker, and cloud/HPC, alongside skills in full-stack web development (Django, React) for research and clinical applications.
Carl von Ossietzky University of Oldenburg · Oldenburg, Germany
Lead bioinformatics operations at the Core Facility Genomics and Bioinformatics, supporting multi-omics research across WGS, WES, RNA-seq, single-cell, Illumina, and Oxford Nanopore datasets.
Build and maintain in-house Nextflow workflows for short-read (Illumina) and long-read (Oxford Nanopore) such as allele-specific expression, ONT methylation profiling.
Support research groups with bioinformatics input on study design, data QC, and downstream analysis planning.
Perform germline variant calling using DRAGEN and DeepVariant, and analyze bulk and single-cell RNA-seq datasets.
Manage sequencing data storage and access control across projects on HPC infrastructure.
Maintain project-level documentation, including code documentation and methodology.
Collaborate with clinical researchers on the analysis of genomic and other clinical research datasets, and support reproducible germline variant-analysis workflows within the hospital computing environment.
Deploy and maintain the in-house clinical genomics platform, enabling clinicians to perform variant analysis and prioritisation.
Ensure compliance with data-protection (GDPR), access-control, and documentation requirements across the hospital's clinical governance framework and the university's research environment.
Jan 2023 - Aug 2025Work
Bioinformatician
University of Oxford · Oxford, United Kingdom
Designed and implemented workflows for the research group's Terra.bio platform, a cloud-based trusted research environment (TRE) on Google Cloud, including a WDL-based workflow for variant calling and annotation using PacBio long-read sequencing data, and a machine learning workflow for proteomic data.
Led the processing and analysis of diverse 'omics datasets for research projects at the Oxford-GSK Institute for Molecular and Computational Medicine (IMCM).
Provided bioinformatics support to IMCM research teams, devising new analysis strategies and working closely with the bioinformatics core group.
Maintained extensive documentation of all analyses and code within the IMCM data platform.
Oct 2018 - Dec 2022Work
Research Scholar
Adam Mickiewicz University · Poznan, Poland
Developed a bioinformatics pipeline for automated exploration of NCBI SRA datasets, enabling efficient sequence-based searches and analysis of large-scale NGS data.
Analyzed publicly available RNA-seq data to understand gene expression patterns of snoRNAs, tRNAs, and tRNA-like genes.
Performed comparative genomic analyses of draft and complete genomes to identify and characterize novel tRNA-Cys gene clusters in Arabidopsis thaliana.
Developed a snoRNA expression atlas web application, including gene annotations, expression levels, and predicted target genes.
Developed a tRNA expression database providing insights into tRNA and tRNA-like gene expression, transcript-level coverage, and structural details.
Apr 2016 - Sept 2018Work
Bioinformatics Analyst
Medgenome Labs Ltd. · Banglore, India
Delivered clinical WES, WGS, RNA-seq, and neo-epitope prediction projects to clinicians and clients as part of a clinical research team operating within a CAP-accredited, NABL-certified diagnostic environment.
Developed an in-house variant annotation database to support the interpretation and reporting of clinically relevant genomic variants.
Automated the in-house neoepitope prioritisation pipeline for somatic variants, integrating somatic variant calling, HLA typing, and ML-based immunogenicity scoring from tumor-normal DNA-seq and RNA-seq data.
Built a web application on top of the pipeline for automated neoepitope prioritisation, generating reports with ranked peptide candidates and their corresponding HLA allele predictions.
Built an internal QC dashboard for real-time monitoring of NGS pipeline runs across WES, WGS, and RNA-seq projects on on-premise servers.
Oct 2015 - Apr 2016Work
Bioinformatics Trainee
Genotypic Technology Pvt. Ltd. · Banglore, India
Enhanced the in-house variant annotation pipeline by adding newer annotation sources.
Analyzed clinical whole-exome sequencing (WES) data to aid clinicians in diagnosing genetic diseases.
Conducted benchmarking of variant calling tools, including Illumina's BaseSpace and Agilent SureCall.
2013 - 2015Education
Master of Technology (M.Tech) in Bioinformatics
Sam Higginbottom Institute of Agriculture, Technology & Sciences · Prayagraj, India
2008 - 2012Education
Bachelor of Technology (B.Tech) in Bioinformatics
Bharath University · Chennai, India
Projects
SnapVar
A web application for visualization of human genetic variants in protein domain and transcript context.
proteomics-ML-workflow
A cloud-based proteomics machine learning workflow using deep learning and classical ML models for biomarker discovery.
ARA (Automatic Record Analysis)
A pipeline for automated exploration of NCBI SRA datasets using nucleotide sequences as queries.
nf-rna-wasp-allele-count
A genotype-aware RNA-seq pipeline using STAR+WASP for allele-specific read counting and variant allele fraction estimation.
wgs-varcall
A cloud-based WGS variant calling pipeline for short-read (Illumina) data with an additional feature for targeted GBA variant detection using the Illumina Gauchian tool.
nf-ont-methpro
A haplotype-resolved DNA methylation profiling pipeline for Oxford Nanopore long-read sequencing.
pb-variant-call
A WDL-based workflow for variant calling and annotation using PacBio HiFi long-read sequencing data, optimized for execution on the Terra.bio cloud platform.
Gauchian-enrich
A variant annotator for GBA variants called by the Illumina Gauchian tool.
getBamDepth
A tool to generate depth of coverage from a BAM/SAM/CRAM file or parse the output generated by samtools depth.
variant-liftover
A command-line tool to lift over SNVs/indels from hg19 to hg38.
gene-to-protein-domains
A command-line utility to fetch protein domain and transcript information via UniProt/Ensembl APIs.
SRA-annotator
A command-line tool for retrieving annotations from the NCBI SRA database.
Skills
Bioinformatics
Short-read (Illumina) & long-read (PacBio, ONT) data analysis (WGS, WES, RNA-seq)Pipeline development and automationGermline/Somatic Variant Calling