Challenges and envisaged approach Despite its potential, making DNA a practical and efficient storage medium requires overcoming several key challenges:
(i) Data transformation: converting digital data into a quaternary alphabet (A, C, G, T).
(ii) DNA synthesis: writing data through the physical synthesis of DNA strands.
(iii) DNA sequencing: reading the stored data by sequencing DNA.
(iv) Data retrieval: reconstructing the original digital data from the sequenced symbols.
This PhD project focuses on the first and fourth challenges, by developing joint compression and error-correction algorithms that are robust to sequencing errors arising during step (iii).
Efficient DNA storage critically depends on fast sequencing technologies, which often come at the cost of increased error rates. For example, nanopore sequencing, developed by Oxford Nanopore Technologies (ONT), enables real-time analysis but introduces relatively high error rates [4,5]. Unlike traditional sequencing technologies, nanopore sequencing produces not only substitution errors but also insertion and deletion errors. Deletions are particularly challenging, as they differ from erasure errors where the position of missing data is known
e.g., packet losses in digital communications). In the case of deletions, neither the existence nor the location of the missing symbols is known, significantly complicating error correction.
Two main approaches have been proposed in the literature to address these errors. The first consists of avoiding error-prone patterns by designing constrained DNA sequences, such as limiting homopolymers or enforcing balanced GC content [6]. The second approach embraces sequencing errors and focuses on correcting them using coding techniques [7]. These strategies are typically considered contradictory, as one seeks to prevent errors while the other assumes their presence [8].
In this project, we propose to explore an alternative route that combines both strategies: avoiding the majority of sequencing errors while correcting the remaining ones. This will be achieved by jointly structuring the compressed DNA stream and designing error-correction mechanisms tailored to nanopore sequencing. In particular, we will exploit the properties of the quaternary alphabet and the interaction between compression and error-correction algorithms. The work will build upon concepts from non-binary channel coding [10,11] and DNA-specific transcoding techniques [6], while also integrating efficient post-sequencing data retrieval methods such as those proposed in [9].
Bibliography
[1] David Reinsel-John Gantz-John Rydning, John Reinsel, and John Gantz. “The digitization of the world from edge to core.Framingham: International Data Corporation”, 16:1–28, 2018.
[2] Luis Ceze, Jeff Nivala, and Karin Strauss. “Molecular digital data storage using DNA”. NatureReviews Genetics, 20(8):456–466, 2019.
[3] Victor Zhirnov, Reza M Zadegan, Gurtej S Sandhu, George M Church, and William L Hughes. “Nucleic acid memory”. Nature materials, 15(4):366–370, 2016.
[4] Delahaye, Clara, and Jacques Nicolas. “Nanopore MinION Long Read Sequencer: An Overview of Its Error Landscape,” November 23, 2020. https://hal.inria.fr/hal-03123133.
[5] ———. “Sequencing DNA with Nanopores: Troubles and Biases.” PLoS ONE, October 1, 202
[6] S. Al Sayyed and A. Roumy and T. Maugey. ``Efficient constraining of transcoding for DNA-based data storage'', IEEE International Conference on Image Processing (ICIP), 2025.
[7] R. Khabbaz, M. Antonini, S. Kas Hanna, Marker Guess & Check Plus (MGC+): An Efficient Short Blocklength Code for Random Edit Errors, International Symposium on Topics in Coding (ISTC), 2025.
[8] F. Weindel, A. L. Gimpel, R. N. Grass and R. Heckel, ”Embracing errors is more effective than avoiding them through constrained coding for DNA data storage,” Allerton Conference on Communication, Control, and Computing 2023.
[9] F. De Moor, O. Boullé, D. Lavenier, De Bruijn Graph Partitioning for Scalable and Accurate DNA Storage Processing, BioRxiv, 2025.
[10] M. C. Davey and D. J. C. Mackay, “Low density parity check codes over GF(q)”, Information Theory Workshop, 1998, pp. 70-71.
[11] D. Declercq and M. Fossorier, “Decoding algorithms for nonbinary LDPC codes over GF(q)”, IEEE Trans. on Commun., vol. 55, pp. 633- 643, April 2007