> For the complete documentation index, see [llms.txt](https://kobic.gitbook.io/kna-documentation/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://kobic.gitbook.io/kna-documentation/kna-user-guide/1.-overview.md).

# 1. Overview

## Introduction

**Korean Nucleotide sequence Archive (KNA)** is a repository that collects, stores, and manages nucleotide sequence data across a wide range of biological sources, including nucleotide sequences, sequence assemblies, sequence annotations, PCR primer records, or other nucleotide sequence-derived records.

KNA supports the submission and management of the following data categories:

* **Nucleotide sequences:** general DNA or RNA sequence records, including genes, mRNAs, non-coding RNAs, viral sequences, phage sequences, plasmids, and organelle genomes such as mitochondria and chloroplasts.
* **Genome assemblies:** prokaryotic and eukaryotic genome assemblies, including chromosomes, scaffolds, and contigs produced from assembled sequence data.
* **MAGs (Metagenome-Assembled Genomes):** reconstructed genomes derived from environmental metagenomic samples through binning and assembly, accompanied by quality metrics such as completeness and contamination.
* **TSA (Transcriptome Shotgun Assemblies):** computationally assembled transcript sequences generated from RNA sequencing or related transcriptome data.
* **TLS (Targeted Locus Sequences):** targeted marker or locus sequences, such as 16S rRNA, 18S rRNA, ITS, COI, or other conserved loci.
* **Synthetic constructs:** artificially designed or engineered nucleotide sequences, including expression constructs, vectors, and other designed DNA sequences.
* **PCR primers:** forward and reverse primer sequence information submitted as primer records.

KNA functions as a component of the **Korean BioData Station (K-BDS)**, providing a standardized platform for the submission, management, and dissemination of nucleotide sequence data generated through biological and genomic research.\
Through harmonized metadata and accession systems, KNA aims to ensure national-scale data archiving and interoperability with international repositories (e.g., GenBank, ENA, DDBJ)

## Key features of KNA

* **Integrated Nucleotide data management**\
  KNA provides a centralized platform to collect, store, and manage diverse nucleotide sequence data, including genomic sequences, transcriptomes, marker sequences, synthetic constructs, and PCR primers.
* **Compliance and Alignment with International Standards**\
  KNA is aligning its data standards and submission framework with INSDC formats (GenBank, ENA, DDBJ) to support global data sharing and interoperability. This ensures that submitted data can be efficiently integrated with international repositories once the archive joins INSDC.
* **Support for Non-Human and Environmental Samples**\
  KNA accepts sequences from humans, non-human organisms, and environmental samples, including metagenomes and organelle genomes, expanding research utility across biodiversity and ecology studies.
* **Secure and Scalable Data Storage**\
  KNA maintains robust infrastructure for secure storage, backup, and scalable management of large-scale nucleotide datasets, ensuring long-term preservation and accessibility.
* **Facilitating Transparent and Reproducible Research**\
  Through standardized accession systems and harmonized metadata, KNA promotes reproducibility and transparency in genomic research, while currently preparing for integration with international standards and repositories.
* **Data Validation and Quality Assurance**\
  Submitted sequences undergo automated and manual validation to ensure integrity, completeness, and adherence to submission guidelines. Metadata curation ensures high-quality, standardized information.

## Cite

**To describe data deposition in your manuscript, use the following sentence:** The nucleotide sequence data reported in this paper have been deposited in the Korea Nucleotide Sequence Archive (KNA) in Korea Bioinformation Center, Korea Research Institute of Bioscience and Biotechnology (KAPxxxxxxx) that are publicly accessible at <https://kbds.re.kr/KNA>.

> Contact KNA staff for assistance at [kna@kribb.re.kr](https://app.gitbook.com/u/17Tn8ORjrdemlt3xrt7zevXbbSw2)
