Hapo-G, haplotype-aware polishing of genome assemblies with accurate reads

Archive ouverte

Aury, Jean-Marc | Istace, Benjamin

Edité par CCSD ; Oxford University Press -

International audience. Single-molecule sequencing technologies have recently been commercialized by Pacific Biosciences and Oxford Nanopore with the promise of sequencing long DNA fragments (kilobases to megabases order) and then, using efficient algorithms, provide high quality assemblies in terms of contiguity and completeness of repetitive regions. However, the error rate of long-read technologies is higher than that of short-read technologies. This has a direct consequence on the base quality of genome assemblies, particularly in coding regions where sequencing errors can disrupt the coding frame of genes. In the case of diploid genomes, the consensus of a given gene can be a mixture between the two haplotypes and can lead to premature stop codons. Several methods have been developed to polish genome assemblies using short reads and generally, they inspect the nucleotide one by one, and provide a correction for each nucleotide of the input assembly. As a result, these algorithms are not able to properly process diploid genomes and they typically switch from one haplotype to another. Herein we proposed Hapo-G (Haplotype-Aware Polishing Of Genomes), a new algorithm capable of incorporating phasing information from high-quality reads (short or long-reads) to polish genome assemblies and in particular assemblies of diploid and heterozygous genomes.

Suggestions

Du même auteur

Chromosome-Scale Assemblies of Plant Genomes using Nanopore Long Reads and Optical Maps

Archive ouverte | Aury, Jean-Marc | CCSD

International audience. Plant genomes are often characterized by a high level of repetitiveness and polyploid nature. Consequently, creating genome assemblies for plant genomes is challenging. The introduction of sh...

De novo sequencing and assembly of complex genomes

Archive ouverte | Aury, Jean-Marc | CCSD

International audience. Reconstructing plant genomes is a difficult task due to their often large sizes, unusual ploidy and large numbers of repeated elements. However, the field of sequencing is changing very rapid...

De novo sequencing and assembly of genomes using long reads

Archive ouverte | Aury, Jean-Marc | CCSD

International audience. The introduction of short-reads technologies ten years ago, significantly increased the number of available genomes. Generally, these assemblies are incomplete and fragmented, and only a few ...

Chargement des enrichissements...