Self-correction approach (SCA) was proposed in the ref (Reducing assembly complexity of microbial genomes with single-molecule sequencin, Genome Biology 2013).
We used all SMRT cells and randomly selected four and six SMRT cells three times for each, and access the correctness by Quast.
Statistics without reference | All Data | 4 SMRT cells : 1st Set | 4 SMRT cells : 2nd Set | 4 SMRT cells : 3rd Set | 6 SMRT cells : 1st Set | 6 SMRT cells : 2nd Set | 6 SMRT cells : 3rd Set |
# contigs | 2 | 8 | 10 | 14 | 1 | 1 | 4 |
Largest contig | 4 278 957 | 2 277 010 | 1 213 670 | 984 459 | 4 641 350 | 4 640 250 | 3 162 440 |
Total length | 4 650 771 | 4 648 304 | 4 644 602 | 4 656 274 | 4 641 350 | 4 640 250 | 4 653 394 |
N50 | 4 278 957 | 2 043 590 | 2 044 147 | 2 135 225 | 3 162 440 | 4 640 250 | 4 641 350 |
Misassemblies | |||||||
# misassemblies | 8 | 0 | 0 | 0 | 7 | 7 | 8 |
Misassembled contigs length | 4 278 957 | 3 530 352 | 2 949 761 | 3 653 461 | 4 641 350 | 4 640 250 | 3 209 090 |
Mismatches | |||||||
# mismatches per 100kbp | 0.8 | 0.43 | 0.58 | 1.36 | 0.15 | 0.95 | 0.58 |
# indels per 100kbp | 5.71 | 2.98 | 4.45 | 9.56 | 1.77 | 8.02 | 6.88 |
# N's per 100kbp | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
Genome Statistics | |||||||
Genome fraction(%) | 100 | 100 | 99.815 | 99.87 | 100 | 99.995 | 99.979 |
Duplication ratio | 1.037 | 1.016 | 1.017 | 1.025 | 1.022 | 1.038 | 1.025 |
# genes | 4494+3 part | 4494+3 part | 4480+7 part | 4485+9 part | 4494+3 part | 4493+4 part | 4492+5 part |
NGA50 | 615 234 | 1 205 052 | 572 342 | 875 953 | 844 482 | 633 220 | 1 267 242 |
Running Time | 19hr 06m | 13hr 34m | 13hr 21m | 12hr 38m | 21hr 28m | 22hr 56m | 22hr 07m |