Whole-genome sequences of 240 indigenous African cattle from Egypt, Uganda, and South Africa

Abstract

Indigenous cattle are central to livestock production in Africa, valued for their adaptability to harsh tropical environments despite lower productivity than commercial breeds. Genome analyses offer critical insights into the genetic potential for enhancing both resilience and productive traits, supporting the advancement of worldwide cattle farming systems. Here, we generated whole-genome sequence data for 240 indigenous cattle representing breeds from distinct agro-climatic regions in Egypt, Uganda, and South Africa. The dataset comprises over ten terabytes of paired-end reads generated using the Illumina NovaSeq. 6000 platform, with an average genome coverage of approximately 10×. Post-filtering reads were mapped to the ARS-UCD1.2 reference genome with a mean mapping rate of 99.2% (range: 64.5–99.9%). Variant calling identified ~43 million SNPs and 6 million indels (≤50 bp) unevenly distributed across the genome. Functional annotation indicated that many variants were located within or near known genes. This comprehensive genomic resource provides a foundation for future studies of genetic diversity, breed identity, population structure, local adaptation, breed-specific traits, or strategies for global cattle conservation.

Description

DATA AVAILABILITY : Raw sequencing data are available in the ENA under Project accession PRJEB90914. VCFs, including both the primary dataset and the filtered dataset used for downstream analyses, are available in EVA under the following project accessions: PRJEB110435 (ARS-UCD1.2; raw data), PRJEB94226 (ARS-UCD1.2; filtered data), and PRJEB93868 (ARS-UCD2.0 Y chromosome; raw data). Publicly available reference variant datasets were obtained from the 1000 Bull Genomes Project (PRJNA391427) and the African Genomic Reference Resource (PRJEB74565). CODE AVAILABILITY : The data analyses were conducted using standard bioinformatics tools on a Linux system. Representative command-line scripts and parameter settings for the main sequencing quality control, read preprocessing, mapping, variant calling, variant filtering, annotation, and population structure analyses are publicly available on GitHub (https://github.com/junxingao888/BovWGS-Pipeline) and Zenodo (https://doi.org/10.5281/zenodo.19189075).

Keywords

Next-generation sequencing (NGS), Whole genome sequencing (WGS), Indigenous cattle, Livestock production, Africa, Agricultural genetics, Deoxyribonucleic acid (DNA), DNA sequencing

Sustainable Development Goals

SDG-02: Zero hunger

Citation

Dlamini, N., Gao, J., Ginja, C. et al. Whole-genome sequences of 240 indigenous African cattle from Egypt, Uganda, and South Africa. Scientific Data 13, 906: 1-13 (2026). https://doi.org/10.1038/s41597-026-07458-y.