atavide lite

Introduction

  • Overview
    • What problem does atavide lite solve?
    • Supported analyses
    • What the project provides
    • Scope and responsibilities
  • Design and rationale
    • Explicit stages instead of one opaque run
    • Failure containment and restartability
    • Separate profiles instead of universal scheduler abstraction
    • A small analysis contract
    • Similar processing across sequencing technologies
    • Read-based and assembly-based branches
    • Resource-aware data movement
    • Conservative use of automation and AI

User guide

  • Installation
    • Prerequisites
    • Clone the repository
    • Build the bundled utilities
    • Create software environments
    • Databases and references
    • Verify before using real data
    • Updating
  • Choosing a profile
    • Read type
    • Supplied cluster profiles
      • Pawsey Setonix
      • Flinders Deepthought
      • NCI Gadi
      • Generic and component workflows
    • Selection checklist
  • Quick start
    • 1. Install and choose a profile
    • 2. Create an analysis directory
    • 3. Configure the dataset
    • 4. Build the sample list
    • 5. Create log directories
    • 6. Prepare databases
    • 7. Submit quality control
    • 8. Submit optional host removal
    • 9. Start the two analysis branches
    • 10. Monitor before scaling up
  • Configuration and input preparation
    • Analysis directory layout
    • DEFINITIONS.sh
    • Paired-read naming
    • Single/long-read naming
    • Host-removal choices
    • Database provenance
    • Resource settings
  • Pipeline stages
    • Dependency model
    • 1. Quality control with fastp
    • 2. Optional host-read separation
    • 3. FASTQ-to-FASTA conversion
    • 4. Read-based taxonomy with MMseqs2
    • 5. Taxonomy completion and summaries
    • 6. Functional annotation with BV-BRC Subsystems
    • 7. Optional PHROG annotation
    • 8. Assembly with MEGAHIT
    • 9. Preparing contigs for VAMB
    • 10. Binning with VAMB
    • 11. MAG quality with CheckM
    • 12. Optional grouped VAMB
    • 13. Read-fate and Sankey summaries
  • Outputs and interpretation
    • Scheduler logs
    • Quality-controlled reads
    • Host and non-host reads
    • MMseqs2 results
    • Taxonomy tables
    • Functional and Subsystems tables
    • Assemblies
    • VAMB products
    • CheckM results
    • Sankey and read-fate outputs
    • Reproducibility record
  • Troubleshooting
    • A practical diagnosis order
    • Unbound variables
    • No samples or wrong array indices
    • Missing paired reads
    • Conda activation fails in a batch job
    • Tool not found or wrong version
    • Job exceeded wall time
    • Out of memory
    • Files disappeared from scratch
    • MMseqs2 database not found
    • TaxonKit errors or incomplete taxonomy
    • Empty output after fastp
    • Dependency will never run
    • Asking for help

Technical guide

  • Scientific methods
    • Quality control
    • Host-read removal
    • Read-based taxonomic profiling
    • Functional profiling
    • Assembly
    • MAG recovery
    • Optional analyses
    • Reproducible reporting
  • Components and utilities
    • bin/
    • summarise_taxonomy/
    • minion_taxonomy_and_function/
    • adapters/
    • pawsey_lib/
    • bash/
    • Custom taxon exclusion
  • Adding support for another cluster
    • Begin with a cluster-support issue
    • Choose the closest source profile
    • Porting checklist
    • Resource mapping by stage
    • Worked example: Pawsey Setonix
    • Validate the new profile
  • Contributing
    • Start with an issue
    • Create a focused branch
    • Make changes that remain inspectable
    • Using agentic AI
    • Validate locally and on the cluster
    • Commit and open a pull request
    • Documentation contributions
atavide lite
  • Search


© Copyright 2026, The atavide lite contributors.

Built with Sphinx using a theme provided by Read the Docs.