2-bit packing for the four canonical DNA nucleotide symbols (A, C, G, T), four symbols per byte, giving 4:1 on pure nucleotide data. Bytes outside that alphabet are recorded in an exception list (position plus original value) and packed as a placeholder code, so arbitrary byte streams still round-trip exactly. Byte-for-byte identical to CompressionWorkbench’s BB_Dna reference block.
| Property | Value |
|---|---|
| Category | Compression Algorithms |
| Sub-category | Bioinformatics |
| Security status | 🎓 Educational Only |
| Complexity | Intermediate |
| Inventor | W. James Kent (UCSC 2bit format) |
| Year | 2002 |
| Origin | 🇺🇸 United States |
| Source | algorithms/compression/dna-compression.js |
Status: 🎓 Educational Only
No vulnerabilities are recorded for this implementation.
4 vectors ship with this algorithm and run in the test suite. Byte values are hexadecimal.
Vector 1 — Empty DNA sequence
| Field | Value |
|---|---|
input |
(empty) |
expected |
0000000000000000 |
Vector 2 — Basic nucleotides - 2-bit encoding
| Field | Value |
|---|---|
input |
41434754 |
expected |
04000000000000001b |
Vector 3 — Simple nucleotide sequence
| Field | Value |
|---|---|
input |
414347544743 |
expected |
06000000000000001b90 |
Vector 4 — Ambiguity code N escaped through the exception list
| Field | Value |
|---|---|
input |
414347544e414347544e |
expected |
0a00000002000000040000004e090000004e1b06c0 |