Ricardo R. Pavan
  • Home
  • About
  • Projects
  • Tutorials & Data
    • Bioinformatics Tutorials
    • Data & Statistics
  • Publications
  • Software
  • CV
  • Contact

On this page

  • What I do
  • Featured projects
  • Selected outputs
  • Latest from the blog
  • Let’s talk

Home

Bioinformatics scientist developing reproducible computational methods for microbiome, virome, and biomedical research.

Ricardo R. Pavan

Bioinformatics Scientist  |  Microbiome, Viromics & Reproducible Data Science

I develop reproducible bioinformatics pipelines and statistical/machine learning methods for microbiome, virome, and biomedical data — from raw sequencing reads to publication-ready results.

View Projects Download CV

GitHub  ·  LinkedIn  ·  Google Scholar  ·  ORCID  ·  Email

Portrait of Ricardo R. Pavan

What I do

Microbiome & Viromics

I design and run analyses that turn raw sequencing reads — 16S amplicon, shotgun metagenomic, and viral metagenomic data — into reliable, interpretable pictures of microbial and viral community structure, dynamics, and host interactions.

Bioinformatics Software & Pipelines

I build and maintain open-source tools (e.g. CRESSENT) and workflow pipelines that make specialized analyses reusable, tested, documented, and installable by other researchers — not one-off scripts.

Statistical Analysis & Machine Learning

I apply multivariate statistics, regression, classification, dimensionality reduction, and clustering to high-dimensional biological and environmental data, with attention to study design and validation, not just model fitting.

Reproducible Pipelines & HPC

I implement workflows in Nextflow and Snakemake, containerized with Docker/Apptainer and run on Slurm-managed HPC and research cyberinfrastructure (CyVerse, ACCESS, Tapis), so analyses can be rerun and verified by others.

Featured projects

CRESSENT

Problem: ssDNA viruses (CRESS viruses) lack consistent, reproducible annotation tools. Role: Lead developer. Methods: Python/R, motif & recombination analysis, CLI packaging. Impact: Published open-source toolkit distributed via PyPI/Bioconda (Microbial Genomics).

Python R PyPI Bioconda CLI

Project details · Repository · Publication

VirION3

Problem: Short-read-only sequencing struggles to recover complete, accurate viral genomes. Role: Developer, pipeline design. Methods: Long-read (Nanopore/PacBio) + Illumina hybrid assembly, benchmarking. Impact: Extends the VirION/VirION2 line of methods for viral genome recovery.

Nanopore PacBio Illumina Nextflow

Project details · Repository: link pending

BATS Diel Virome

Problem: Viral population dynamics at sub-daily timescales are rarely resolved in marine systems. Role: Co-lead, statistical/bioinformatic analysis. Methods: Time-series viral metagenomics, host prediction, depth-stratified sampling. Impact: Published in PLOS Biology.

Metagenomics Time-series Host prediction

Project details · Publication

Pig Burn Wound Microbiome

Problem: Understanding microbial community succession in a wound-healing model. Role: Bioinformatics analysis (assembly, binning, taxonomy). Methods: Nextflow, metagenomic assembly, genome binning, CheckM2, GTDB-Tk. Impact: Ongoing — generalized methods description (no unpublished results shown).

Nextflow Slurm Apptainer Metagenomics

Project details

Spinal Cord Injury Microbiome & Virome

Problem: How microbiome, virome, and phage–host dynamics shift after spinal cord injury. Role: Bioinformatics/statistical analysis. Methods: Longitudinal microbiome/virome profiling, mobile-element detection. Impact: Ongoing — generalized methods description.

Longitudinal Microbiome Virome

Project details

Macquarie Harbour Microbial Ecology

Problem: How environmental gradients and aquaculture shape microbial communities in a stratified estuary. Role: PhD research — lead. Methods: 16S/18S sequencing, network analysis, machine learning, db-RDA/PERMANOVA. Impact: Multiple peer-reviewed publications (2021–2022).

R Machine Learning Ecological Statistics

Project details

See all projects →

Selected outputs

  • Publication — Pavan, Sullivan & Tisza (2026). CRESSENT: a bioinformatics toolkit to explore and improve ssDNA virus annotation. Microbial Genomics. DOI
  • Publication — Carrillo et al. (2026). Sub-daily virus sampling at the Bermuda Atlantic Time Series reveals diel and depth-structured population dynamics. PLOS Biology. DOI
  • Software — CRESSENT: open-source ssDNA virus annotation toolkit (PyPI, Bioconda).
  • Talk — CRESSENT: A Toolkit to Enhance ssDNA Virus Annotation. Midwest Microbiome Symposium, Columbus, OH (2025).

See all publications → · See all software →

Latest from the blog

Diagram illustrating a Nextflow metagenomics pipeline

Building a Reproducible Nextflow Pipeline for Metagenomics
Nextflow
Metagenomics
HPC
A minimal, worked example of structuring a metagenomics workflow in Nextflow so it is portable and reproducible.
Jul 24, 2026

Illustration comparing clustering methods

Comparing Clustering Methods on the Same Dataset
Statistics
Machine Learning
K-means, hierarchical clustering, and DBSCAN applied to the same simulated ecological dataset — and why they disagree.
Jul 24, 2026
No matching items

All tutorials → · All data & statistics posts →

Let’s talk

I’m glad to hear from people exploring research collaborations, bioinformatics or data-science projects, scientific software, biomedical/microbiome data analysis, employment opportunities, or workshop/talk invitations.

Get in touch

© 2026 Ricardo R. Pavan. Content licensed CC BY-NC-ND 4.0.

Built with Quarto.

  • ORCID