> For the complete documentation index, see [llms.txt](https://waqasnayab.gitbook.io/pipeline/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://waqasnayab.gitbook.io/pipeline/master/implementation-of-cz-id-mini-wdl-based-sars-cov-2-consensus-genome-workflow-pipeline-at-aku/sars-cov-2-consensus-genome-qc-mixed-sites-correction.md).

# SARS CoV-2 Consensus Genome QC- Mixed sites Correction

This page will guide you how to remove mixed sites (MS) from your FASTA consensus genome files before submitting them to GISAID.

> *Khan W, Kanwar S*
>
> *AKU CITRIC Center for Bioinformatics and Computational Biology, Depratment of Pediatrics and Child Health, Faculty of Health Sciences, Medical College, The Aga Khan Universitry, Karachi-74800, Pakistan.*

Mixed sites are mixtures of two or more different bases at a given necleotide position within the sequence.

![IUPAC nucleotide code for single and mixed bases ](https://3625414896-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MdH8WvV6vWCdYeCPAEf%2Fuploads%2FwIZXiPCCyYbhRPXoFXZ3%2Fimage.png?alt=media\&token=43861b11-205b-46fd-8068-f503def39817)

On [NextClade](https://clades.nextstrain.org/), mixed sites are represented by "M". Color coding represents the mixed sites numbers in the sequence.

* Red indicates the number of mixed sites >10.
* Yellow indicates the number of mixed sites >2.
* Green indicates the number of mixed sites <=2 (acceptable).

!["M" in the QC column represents the mixed sites in FASTA file.](https://3625414896-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MdH8WvV6vWCdYeCPAEf%2Fuploads%2Fvr1qzaug9QNp7sorshey%2Fimage.png?alt=media\&token=bcb14edd-6344-46e1-9e53-9bbf4ac21aeb)

These can be corrected by following below steps:

**1)** Use “grep” command to find ambiguous code in consensus.fa file.

```
grep 'Y\|R\|W\|S\|K\|M\|D\|V\|H\|B\|X' consensus.fa
```

<div align="center"><img src="https://3625414896-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MdH8WvV6vWCdYeCPAEf%2Fuploads%2FnEP3p3jiWMYzvK4EOxzu%2Fimage.png?alt=media&amp;token=87f6801c-768e-4148-8a23-21c397cebb9d" alt="Identification of ambiguous nucleotides in genome."></div>

**2)** Identify the position of ambiguous code by opening the alignment file, for example, muscle.out.fasta (generated as intermediate file in [CZ ID pipeline](/pipeline/master/implementation-of-cz-id-mini-wdl-based-sars-cov-2-consensus-genome-workflow-pipeline-at-aku.md)) using [MegaX](https://www.megasoftware.net/). Copy three nucleotides upstream and dowstream around MS position, click on "Find Motif" option in MegaX dropdown menu, paste the sequence and search.

***Hint:*** To find the exact MS position, copy adjacent/flanking nucleotides (+/-3) around that specific MS and search in consensus genome FASTA file.

As shown in the image below, the MS `"R"` is located at `6,459` position (displayed at the bottom left corner).

![Visualization of nucleotide position on MegaX](https://3625414896-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MdH8WvV6vWCdYeCPAEf%2Fuploads%2FppCCQLlfNYX7IEWk297o%2Fimage.png?alt=media\&token=773c7888-eb7a-47f7-a309-a4e15a2d68e6)

**3)** Find the frequency of nucleotides at identified position using [Integrative Genomics Viewer](https://software.broadinstitute.org/software/igv/) (IGV). You can do this by opening IGV and firdt uploading the reference genome FASTA file (in our case [MN908947.3](https://www.ebi.ac.uk/ena/browser/api/fasta/MN908947.3?lineLimit=1000)), then upload the BAM file of sample containing MS and search the position (that was identified through MegaX) by pasting it in IGV's search box.

***Hint:** In IGV's search box, paste the position after  `:` . Do not remove the entire line written in search box, e.g.,* `MN908947.3`**`:`**`position`

![IGV's search box](https://3625414896-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MdH8WvV6vWCdYeCPAEf%2Fuploads%2FA40v4tLDkmpQqAXH5tMq%2Fimage.png?alt=media\&token=bfa1a34b-f23e-4146-8ab6-7f30f53aea81)

Click on coverage bar to view the frequency of nucleotides at specified position. If the frequency of a nucleotide is >=75% then replace the ambiguous code in consensus.fa file with most occurring/frequent nucleotide. As shown in the below figure, frequency of "A" is 90% so R can be replaced by "A" in consensus.fa file.&#x20;

![IGV is used to identify the nucleotide frequency](https://3625414896-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MdH8WvV6vWCdYeCPAEf%2Fuploads%2FH7BGtadGOuapXGLlBpQh%2Fimage.png?alt=media\&token=49611ba1-b71d-4244-9271-1f93364f924e)

One thing to keep in mind is the threshold frequency value. Do not modify the nucleotide position if the nucleotide frequency is less than 75 percent. Previously, the threshold criteria was 90% in CZ ID pipeline which then changed to 75% few months ago (see the exact context on [GitHub](https://github.com/chanzuckerberg/czid-workflows/blob/b498b35a56139421813ef0046621dafc812f8a4b/consensus-genome/run.wdl#L58)). It is always better to stay updated regarding the standards accepted by scientific community *(**Hint:*** maintain your CZ ID pipeline by continously updating it!!!).

![NextClade QC results after MS manual correction](https://3625414896-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MdH8WvV6vWCdYeCPAEf%2Fuploads%2FNiA3ZGc7JkoFP0Xzvlkd%2Fimage.png?alt=media\&token=e12f8f47-80bb-4f72-a608-df63dcfb3845)

Hope this will help you to improve the QC of consensus genome.
