Encoding genetic sequence information and its uses

JP2025515360A5Pending Publication Date: 2026-05-08INST PASTEUR DE DAKAR +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INST PASTEUR DE DAKAR
Filing Date
2023-04-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The prior art is difficult to conduct rigorous mathematical analysis and use symmetry to process information on biological sequences.

Method used

Sequence tagging and modification detection are performed by defining the tag using symmetry properties, sequence tagging and modification detection.

Benefits of technology

Mathematical analysis of biological sequences and labeling and detection using symmetry characteristics is realized, improving the processing efficiency and accuracy of sequence information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method for representing a selected sequence of RNA and / or DNA nucleotides and / or nucleotide analogs, comprising the steps of: a) defining a cube using nucleotide bases, each of the eight vertices of the cube being assigned to a nitrogenous base from a subset [A (adenine), C (cytosine), G (guanine), T (thymine)], the nitrogenous bases T (thymine) and U (uracil) being understood herein to be equivalent and therefore interchangeable, each of the vertices of the cube being assigned to a nucleotide base from the subset such that, for each vertex, the assigned base is directly connected via edges to all other base types of the subset; and b) defining a first nucleotide base of the selected sequence of nucleotides using the nucleotide bases defined in the cube. and assigning a nitrogenous base type of each nucleotide of the selected sequence of nucleotides to an assigned vertex of the cube, and sequentially assigning each subsequent nucleotide of the selected sequence of nucleotides to a corresponding vertex of the cube or tetrahedron such that the assigned vertex is directly connected to the vertex of the previous nucleotide base through an edge of the cube, wherein if the nucleotide base is equal to the previous nucleotide base, the nucleotide base is assigned to the same vertex of the cube as the vertex to which the previous nucleotide base was assigned, and wherein the selected sequence of nucleotides and / or nucleotide analogs comprises a nitrogenous base selected from A (adenine), C (cytosine), G (guanine), T (thymine) and / or U (uracil).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the field of bioinformatics, in particular to the field of genome information processing.

[0002] The present disclosure relates to a novel method for representing nucleotide sequences and the use of these representations.More specifically, the present disclosure relates to a method for representing nucleotide sequences in the form of a cube or tetrahedron or in the form of a matrix that represents such a cube or tetrahedron, a method for tagging sequences using such representations, a method for determining modifications to sequences, and a method for defining one or more markers that are characterized by symmetric invariance within nucleotide sequences. [Background technology]

[0003] Many fundamental laws of nature, such as the conservation of energy, the conservation of momentum, or the structure of space-time in general relativity, emerge from the concept of symmetry (https: / / www.pnas.org / doi / pdf / 10.1073 / pnas.93.25.14256). Although symmetric patterns were known to the Egyptians, the mathematical structures underlying the study of symmetries, a branch of abstract algebra known as group theory, were first discovered in the 19th century by Abel, Galois, and Lie. Since then, group theory has revolutionized our understanding of many scientific disciplines, including physics, chemistry, cryptography, and information theory.

[0004] The use of symbolic representations (character strings) to describe biological sequences has aided in the understanding of genes and genomes. However, the lack of a "grammar" for these symbols prevents rigorous mathematical analysis of genetic information.

[0005] Thus, there is a need for better representations of biological sequences that can enable mathematical analysis of genetic information, including the exploitation of symmetries. Summary of the Invention

[0006] A first aspect of the present invention relates to a computer-implemented method of representing a sequence of nucleotides, the method comprising: a) defining a cube or tetrahedron using nucleotide bases, where each vertex of the cube or tetrahedron is assigned a nucleotide base, and for each vertex, the assigned base is directly connected to every other base type via an edge; b) from a selected sequence of nucleotides comprising nucleotides from one of the nucleotide subsets [A,C,G,T] or [A,C,G,U], assigning a first nucleotide base of the selected sequence of nucleotides to a vertex to which the base type of the nucleotide was assigned and sequentially assigning each subsequent nucleotide of the selected sequence of nucleotides to a corresponding vertex of a cube or tetrahedron, such that either the assigned vertex is directly connected to the vertex of the previous nucleotide base through an edge of the cube or tetrahedron, respectively, or if the nucleotide base is equal to the previous nucleotide base, the nucleotide base is assigned to the same vertex as the previous nucleotide base was assigned; c) for each vertex of the cube or tetrahedron, determining the number of nucleotides assigned to each vertex of the cube or tetrahedron; Includes.

[0007] In a preferred embodiment, the first nucleotide is assigned to its corresponding vertex within a given initial face of a cube or tetrahedron.

[0008] In a further preferred embodiment, in which the nucleotide sequence is represented in the form of a cube, the predetermined initial face is a face containing the nucleotide bases in the order GATC or GAUC in a clockwise sense.

[0009] In yet a further preferred embodiment, the sequence of nucleotides is further represented in the form of a matrix. d) assigning each vertex of the cube or tetrahedron to an element of a matrix having at least as many elements as the cube or tetrahedron has vertices; e) assigning the total number of nucleotide bases assigned to each vertex of the cube or tetrahedron to the value of the assigned element of the matrix; Includes.

[0010] In a preferred embodiment of the matrix representation, the matrix is ​​a square matrix, and preferably the vertices of a cube or tetrahedron are represented on the diagonal elements of the matrix. More preferably, when the nucleotide sequence is represented in the form of a cube, the square matrix is ​​preferably a 4x4 matrix.

[0011] In a preferred embodiment of the matrix representation, the sequence of nucleotides is represented in the form of a cube, and the vertices of the cube are assigned to the matrix by a 3D projection of the cube onto a 2D Euclidean plane defining an inner and outer set of nucleotide bases each representing an opposite face of the cube, preferably with the projection axis perpendicular to the given initial face.

[0012] In a preferred embodiment of any of the embodiments of the first aspect of the invention, the received sequence of nucleotides further comprises information indicating the 5' and 3' ends of the sequence, the nucleotides being assigned sequentially to the bases of the cube or tetrahedron in a 5'→3' sense.

[0013] A second aspect of the invention relates to a computer-implemented method of tagging a nucleotide sequence by encoding the nucleotide sequence in the form of a matrix according to any of the embodiments of the matrix representation method of the first aspect of the invention.

[0014] A third aspect of the present invention relates to a computer-implemented method for determining modifications of a nucleotide sequence, the method comprising: a) calculating from at least one selected nucleotide sequence representing a template sequence and at least one selected sequence representing at least one nucleotide sequence for which a modification is to be determined a matrix representation of the sequences according to any of the embodiments of the matrix representation method of the first aspect of the invention; b) individually comparing the equivalent matrix elements of at least one template sequence and at least one sequence for which an alteration in the matrix representation is determined; c) determining that a modification between the at least one template sequence and the at least one sequence for which a modification is to be determined has occurred if any of the equivalent matrix elements are not equal; Includes.

[0015] Alternatively, the method comprises: a) calculating, from at least one selected nucleotide sequence representing a template sequence and at least one selected sequence representing at least one nucleotide sequence for which a modification is to be determined, a matrix representation of the first and second sequences according to any of the embodiments of the matrix representation method of the first aspect of the invention, wherein the matrix is ​​a square matrix; b) calculating the determinant of a matrix representation of the array of step a); c) comparing the determinant of at least one template sequence and at least one sequence for which modifications of the matrix representation are determined; d) determining that a modification between the at least one template sequence and the at least one sequence in which a modification is determined has occurred if the determinants of step c) are not equal; Includes.

[0016] In a preferred embodiment, the modification is one or more of a SNP, a nucleotide insertion, a nucleotide deletion or a nucleotide transition.

[0017] In a further preferred embodiment of any of the embodiments of the third aspect of the invention, the method further comprises the step of determining that the integrity of the biological sequence has been compromised if an alteration of the sequence is determined.

[0018] A fourth aspect of the invention relates to a computer-implemented method for defining one or more markers in a nucleotide sequence, the markers being characterized by symmetry invariance, the method comprising: a) receiving a sequence of nucleotides; b) The length of the base is an integer n 2selecting a portion of the sequence to be analyzed that is a power of c) Base length n 2 into n fragments of length n; d) assembling the array into an n×n matrix; e) computing several transformations of the matrix using one or more of the selected symmetries until all feasible transformations for each of the selected symmetries are exhausted; f) determining for each of the selected symmetries the elements of the matrix that are invariant under all transformations; g) assigning invariant elements of the matrix as invariant markers; h) repeating steps a) to g) using another portion of the sequence, preferably by sliding the start of the sequence to be analyzed by one or more nucleotides, until the entire sequence has been analyzed; Includes.

[0019] In a preferred embodiment, the one or more symmetries include rotational symmetry and / or mirror symmetry.

[0020] In another preferred embodiment, n is an odd number, preferably a prime number, more preferably 5 or greater.

[0021] A fifth aspect of the present invention relates to a computer program product comprising instructions which, when executed by a processor, cause a computer to perform any of the steps of the methods according to the first, second, third and / or fourth aspects of the present invention. [Brief description of the drawings]

[0022] To enable a better understanding of the present disclosure and to show how it may be put into effect, reference will now be made, by way of example only, to the accompanying schematic drawings. [Figure 1] 1 shows list, triangle, square and matrix representations of a nucleotide sequence, as well as the strand orientation of each representation, according to one or more embodiments. [Figure 2A] 1 shows the non-redundant tetranucleotide permutations of the GATC alphabet. [Figure 2B] 1 shows the non-redundant trinucleotide permutations of the GATC alphabet. [Diagram 3] 1 illustrates how a cube of nucleotides can be formed from six copies of the square representation, according to one or more embodiments. [Figure 4] 1 illustrates how a tetrahedron of a nucleotide can be formed from four copies of a triangular representation, according to one or more embodiments. [Figure 5A] 1 shows a cube of a nucleotide base according to one or more embodiments. [Figure 5B] FIG. 5B is a close-up of a square representation of the cube of FIG. 5A according to one or more embodiments. [Figure 5C] 5B shows the formation of a rhombic cuboctahedron of nucleotides from a close-up of the square representation of FIG. 5B, according to one or more embodiments. [Figure 6] 1 illustrates the relationship between a cube and an octahedron via a rhombic cuboctahedron, according to one or more embodiments. [Figure 7A] 1 illustrates a perspective view of a GenoCube, according to one or more embodiments. [Figure 7B] 1 illustrates a floor plan view of a GenoCube, according to one or more embodiments. [Figure 7C] 1 shows another plan view of the GenoCube, according to one or more embodiments, which distinguishes between inner and outer sets of nucleotides. [Figure 8A] 1 shows an example of a nucleotide sequence according to one or more embodiments. [Figure 8B] 8B illustrates an example of a cube definition using nucleotide bases and the assignment of the sequence of nucleotides of FIG. 8A to the vertices of a cube, according to one or more embodiments. [Figure 8C] 8C illustrates a conformal 3D projection of a cube and the assignment of an array to the cube of FIG. 8B in accordance with one or more embodiments. [Figure 8D]A matrix representation of the array in FIG. 8A and its cubic representation in FIG. 8B and FIG. 8C are shown. [Figure 9] 1 shows an example of a matrix encoding a large nucleotide sequence, according to one or more embodiments. [Figure 10] 1 illustrates an example of the sensitivity of matrix encoding to small changes, according to one or more embodiments. [Figure 11] FIG. 1 shows a diagram of a process for determining invariant markers in a nucleotide sequence, according to one or more embodiments. [Figure 12] 1 shows an example of a process for determining invariant markers in a nucleotide sequence, according to one or more embodiments. [Figure 13] 1 illustrates a computer-implemented product according to one or more embodiments. [Figure 14] 1 illustrates a diagram of the use of a TISA deconvolution matrix (TDM) in accordance with one or more embodiments. [Figure 15] FIG. 1 illustrates an example of a method for constructing a TISA deconvolution matrix (TDM) in accordance with one or more embodiments. [Figure 16] 1 illustrates a diagram of an example method for decoding a TISA deconvolution matrix (TDM) in accordance with one or more embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] Description of the Invention definition The term "symmetric group" refers to a group whose elements are bijections from a set onto itself and whose group operations are composition of functions. In the specific case of the present invention, a finite symmetric group S4 defined over a finite set of four symbols consists of permutations that can be performed on the four symbols. There are n! such permutation operations, so the order (number of elements) of the symmetric group S4 is 4!=24.

[0024] The term "dihedral group" refers to the group formed by the different rotational and reflection symmetries of regular polygons with n sides. There are n rotational symmetries and n reflection symmetries. The dihedral group D4 is the group of symmetries of a square.

[0025] The term "alternating group" refers to the group of even permutations of a finite set. An alternating group on a set of n elements is called an alternating group of order n, or an alternating group on n letters, and is denoted An or Alt(n).

[0026] The term "genocube" refers to a cube in which each nucleotide is represented twice, on diagonally opposite vertices of the cube.

[0027] explanation This work examines whether biological sequences encoded by genes and genomes contain symmetries and, as a result, can be addressed using the mathematical power of abstract algebra and group theory. It was hypothesized that the apparent conflict between symmetric and evolutionary transformations may in fact represent two different manifestations of the same phenomenon. It was first recognized that changes in nucleotide sequence (i.e., mutations) could be mimicked by mathematical rearrangements of ordered sets of nucleotides (i.e., permutations).

[0028] Although the study focuses on the four-letter nucleotide alphabets used in DNA, the findings and methods described herein may be readily extendable to any four-letter nucleotide alphabet, such as one or other modified nucleotide four-letter alphabets of RNA. As reported in Figure 1, the four-letter DNA can be represented in several different formats, such as linear (A), circular (B and C), and tabular (D). As any polynucleotide can be represented in all three models, these representations are informationally equivalent but differ by their symmetry properties. However, when examined under the prism of symmetry, representations B, C, and D exhibited several unique properties. The "list representation" contains no a priori symmetry (except for repetitive, monotonic and / or palindromic sequences). The "equilateral triangle representation" exhibits dihedral symmetry of order 6. The "square representation" and "matrix representation" both exhibit dihedral symmetry of order 8, but with different strand orientations (antiparallel vs. parallel).

[0029] Since biological sequences, including genetic sequences, are linear polymers of consecutive nucleotides, we studied permutations of GATC. By performing a systematic analysis of n permutations (n ​​varies from 2 to 12), as shown in Figure 2, we surprisingly found that 3- and 4-permutations of GATC result in permutation families containing the same number of elements (24). This observation is consistent with the symmetric group S n This can be explained by the existence of a universal mathematical structure, called GATC, which describes the number of possible permutations associated with a set of n elements as n!. In our case, the number of non-redundant 4-permutations of GATC is given by calculating 4! = 4 × 3 × 2 × 1 = 24. Similarly, the number of non-redundant 3-permutations of GATC is given by the probability formula 4 × 3 × 2 = 24. These two sets are bijectively related to the symmetry group S4 and are consequently isomorphic to this group. Since the group S4 is isomorphic to the group of symmetries of a cube, this result establishes a one-to-one correspondence between the groups of 3- and 4-nucleotide permutations and the group symmetries of the cube. If the vertices of a cube are numbered from 1 to 4 and the opposite vertices are given the same numbers, a permutation corresponding to the symmetry can be read off one of the faces.

[0030] Also of great interest is the fact that the S4 group can be decomposed into three finite subgroups: (i) the dihedral group D3, which is isomorphic to the group of triangular symmetry, (ii) the dihedral group D4, which is isomorphic to the group of square symmetry, and the alternating group (A4), which is isomorphic to the group of tetrahedral symmetry. From a biological point of view, the existence of these group isomorphisms means that any nucleotide sequence composed of 3- and 4-permutations of the {G,A,T,C} alphabet can be represented as a unique path through the vertices of a tetrahedron or a cube.

[0031] With reference to FIG. 3, when depicted using a "square representation", the 24 non-redundant tetranucleotide permutations of the GATC alphabet can be seen to fall into three subsets of eight different permutations. Each subset contains eight dihedral symmetries corresponding to the symmetry group D4. This group has eight members, namely one identity (ID), three rotational symmetries (R1-R3), and three mirror symmetries (M1-M4). More interestingly, when represented twice, these 48 permutations give rise to a cube (called a genocube) in which each nucleotide is represented twice on diagonally opposite vertices of the genocube. This representation is advantageous since any nucleotide sequence can be represented on such a genocube. The resulting groups (two copies of S4) are then divided into three subgroups with the composition (S4 ゜ The octahedral symmetry group of order 48 (O h ) is the same type as

[0032] The same kind of inference was performed with a set of three permutations. As shown in Figure 4, the 24 non-redundant triplet permutations of GATC can be represented by placing a single nucleotide at each vertex of an equilateral triangle, i.e., a "triangular representation". The resulting operation classifies all three permutations into four subsets of six different permutations (6 × 4 = 24), each of which corresponds to a symmetry of the dihedral group D3, which represents the group of symmetries of equilateral triangles. This group has six members, namely, one identity (ID), two rotational symmetries (R1-R2), and three reflection mirror symmetries (M1-M3). This discovery allows the results obtained for equilateral triangles to be assembled into a regular tetrahedron, which has the property of representing any nucleotide sequence (DNA, RNA or sequences containing modified nucleotides) at its vertices, as observed for the genocube.

[0033] The 24 non-redundant permutations of GATC are redundant 3-letter codes (24 out of 4 3 = 64 possible words) or 4-letter code (4 4This covers only a small subset of the genetic information that can be encoded in a cubic matrix (=24 of 256 possible words). Note that physically expanding by separating the faces overcomes this hurdle because the geometric shape of the cube (see Figures 5A and 5B) generates a large number of original nucleotides at each vertex (see the example of three A's and T's shown in Figure 5C). Thus, the 24 non-redundant 3- and 4-permutations of the GATC alphabet can be expanded to cover any redundant permutation (e.g., AAA, GAA, TTT, AAAA, GAAA, ATTT, etc.) by simply expanding the faces of the genocube. This produces three copies of the letters present at each vertex of each square. When connected, this transforms the cube into a novel shape intermediate between a cube and an octahedron, which is called a rhombiccuboctahedron. This rhombiccuboctahedron can represent any of the 64 triplets that make up the genetic code. This observation is supported by a mathematical property called duality, which transforms a cube topologically into a regular octahedron by expanding the triangles at each vertex of the cube. As shown in Figure 6, the rhombic cube octahedron can then be transformed into a regular octahedron by vanishing the squares located at each vertex of the triangle. Conversely, the octahedron can be transformed into a rhombic cube octahedron by blowing a square into each of its vertices. Note that the regular octahedron has 24 rotational (or orientation-preserving) symmetries, for a total of 48 symmetries. These include transformations that combine reflections and rotations. The cube is a polyhedron that is dual to the octahedron, and thus has the same set of symmetries.

[0034] This dual model can be used to represent the non-redundant permutations on the cube and / or tetrahedron, and the geometry can be locally / temporally expanded to represent the remaining redundant permutations. Overall, we hypothesize that this new representation has profound implications for understanding the nature, structure, and evolution of the genetic code and genetic information in general.

[0035] Therefore, we describe herein a new method to compactly represent a nucleotide sequence as a unique virtual path along the vertices of a cube or tetrahedron.

[0036] A first aspect of the invention relates to a computer-implemented method of representing a sequence of nucleotides, the method comprising the steps of: a) defining a cube or tetrahedron using nucleotide bases such that each vertex of the cube or tetrahedron is assigned a nucleotide base. Each vertex is assigned a nucleotide base, and for each vertex, the assigned base is directly connected to every other base type via an edge. Thus, when determining which nucleotide base is assigned to a vertex, the base of the vertex directly connected to that vertex via an edge must be considered, and is assigned a different nucleotide base than the base of the directly connected vertex. Since each vertex is directly connected to three other vertices via edges, once the definition of the cube or tetrahedron is complete, any vertex of interest is directly connected to three vertices via edges, and the three vertices have different bases assigned between them and the base assigned to the vertex of interest. It should be noted that when a nucleotide sequence is represented in the form of a cube, each nucleotide base type is represented twice, at diagonally opposite vertices of the cube. Advantageously, this allows any nucleotide sequence to be represented on such a cube, since in all cases each subsequent nucleotide base can be assigned to the same vertex as the previous nucleotide base was assigned to, or to a vertex directly connected to the previous nucleotide base, with each vertex being directly connected to every other base type via an edge. b) from a selected sequence of nucleotides comprising nucleotides from one of the nucleotide subsets [A,C,G,T] or [A,C,G,U], assigning a first nucleotide base of the selected sequence of nucleotides to a vertex to which the base type of the nucleotide was assigned and sequentially assigning each subsequent nucleotide of the selected sequence of nucleotides to a corresponding vertex of a cube or tetrahedron, such that the assigned vertex is either directly connected to the vertex of the previous nucleotide base through an edge of the cube or tetrahedron, respectively, or, if the nucleotide base is equal to the previous nucleotide base, the nucleotide base is assigned to the same vertex as the vertex to which the previous nucleotide base was assigned. It should be noted that any sequence of DNA, RNA or a set of four modified nucleotides can therefore be represented. It should also be noted that since a tetrahedron has only four vertices, each vertex is assigned a different nucleotide base. Thus, the first nucleotide is assigned to the only vertex that is assigned a nucleotide base type. However, in the case of a cube, there are eight vertices, so for each nucleotide base, there are two vertices that are assigned a nucleotide base type. Thus, the first nucleotide base of a sequence of nucleotides can be assigned to either of the two vertices that are assigned a nucleotide base type. c) determining for each vertex of the cube or tetrahedron the number of nucleotides assigned to that vertex of the cube or tetrahedron, respectively, thus obtaining a cubic or tetrahedral representation having a sequence of nucleotides, a value being determined for each vertex of the cube or tetrahedron.

[0037] According to a preferred embodiment in which the nucleotide sequence is represented in the form of a cube or tetrahedron, the first nucleotide is assigned to its corresponding vertex in a predetermined initial face of the cube or tetrahedron, advantageously, by having a predetermined initial face in the cube, the problem of redundancy in the number of vertices to which the base type of the nucleotide is assigned is solved. By establishing a predetermined face, a consistent representation of any sequence in the form of a cube can be achieved. It should be noted that this does not occur in tetrahedrons, since in tetrahedrons there is only one vertex to which the first nucleotide base can be assigned, and therefore there is no redundancy.

[0038] In another preferred embodiment, in which the nucleotide sequence is represented in the form of a cube, the predetermined initial face is a face that contains the nucleotide bases in the order GATC or GAUC in a clockwise sense. Advantageously, this provides for a consistent selection of the predetermined initial face.

[0039] To further exploit the properties of such encoding, a matrix representation of the GenoCube can be designed. In a preferred embodiment of any of the above aspects, the sequence of nucleotides is further represented in the form of a matrix. Advantageously, representing the nucleotide sequence in the form of a matrix allows the use of different matrix tools for the representation of the nucleotide sequence. Thus, preferably, the method further comprises the following steps: d) assigning each vertex of the cube or tetrahedron to an element of a matrix having at least as many elements as the cube or tetrahedron has, respectively. If the nucleotide sequence is represented in the form of a tetrahedron, each vertex of the tetrahedron is assigned to an element of a matrix having at least four elements. If the nucleotide sequence is represented in the form of a cube, each vertex of the tetrahedron is assigned to an element of a matrix having at least eight elements. e) assigning the total number of nucleotide bases assigned to each vertex of the cube or tetrahedron to the value of the assigned element of the matrix, thus obtaining a matrix containing at least the number of nucleotides for each vertex of the cube or tetrahedron in the representation of the sequence of nucleotides.

[0040] Hereinafter, the matrix representing the vertices of the GenoCube may be referred to as a Topologically Invariant Sequence Array (TISA). Note that topologically invariant expression in the context of the present invention refers to the fact that sequence identity remains invariant through various transformations and / or representations, i.e., the matrix is ​​simply another representation of the sequence of amino acids.

[0041] It should be noted that several designs can be envisaged when assigning the vertices of a cube or tetrahedron to a matrix, since multiple choices can be made when determining how to assign vertices to each element in the matrix representing the vertices. For example, if a nucleotide sequence is represented in the form of a tetrahedron, a matrix containing at least four elements can be shaped as at least (column by column) 4×1, 2×2 or 1×4 matrices. If a nucleotide sequence is represented in the form of a cube, the number of sizes of the matrix can be expanded. Furthermore, multiple matrix designs can satisfy the requirement of having at least as many elements as a cube or tetrahedron has vertices. Thus, if a nucleotide sequence is represented in the form of a cube, the sequence can be represented using a 4×4 matrix, with half of the elements of the matrix (8) having assigned cube vertices and the other half (8) having no assigned cube vertices.

[0042] In a more preferred embodiment, the matrix is ​​a square matrix. Advantageously, a square matrix has further properties, such as the ability to calculate the determinant.

[0043] In a further preferred embodiment, when the nucleotide sequence is represented in the form of a cube, the sequence is represented by a 4x4 matrix. In a further preferred embodiment, the vertices of the cube or tetrahedron are represented on the elements in the diagonal of the matrix. Advantageously, this ensures that the determinant of the matrix is ​​not zero every time the sequence is wrapped at least once through all the vertices of the cube. In another further preferred embodiment, when the nucleotide sequence is represented in the form of a cube, the cube vertices are assigned to a square 4x4 matrix, and the vertex values ​​are assigned to its diagonal elements.

[0044] The cube can be represented using a conformal 3D projection, as shown in Figure 7A, or projected onto a 2D Euclidean plane, as shown in Figure 7B. Interestingly, the 2D projection transforms the cube into a planar graph that defines an exterior and interior set of nucleotide bases, as shown in Figure 7C, with the exterior set of nucleotides highlighted. This property is interesting because it mimics the inherent polarity of polynucleotide sequences.

[0045] In a more preferred embodiment of any of the above preferred embodiments, in which the sequence of nucleotides is represented in the form of a cube, the vertices of the cube are assigned to the matrix by a 3D projection of the cube onto a 2D Euclidean plane that defines an inner and outer set of nucleotide bases, each of which represents an opposite face of the cube. Advantageously, this allows for a consistent representation of the cube in a square 4x4 matrix, with the vertex values ​​assigned to its diagonal elements, as shown, for example, in Figure 8D. More preferably, the projection axis is a projection axis perpendicular to the given initial face.

[0046] In each sequence of nucleotides, 5' and 3' ends can be defined. In another preferred embodiment of any of the above embodiments, the received sequence of nucleotides further comprises information indicating the 5' and 3' ends of the sequence, and the nucleotides are assigned sequentially to the bases of a cube or tetrahedron in a 5'→3' sense. Advantageously, this further prevents that two different representations of the same sequence can be obtained when the starting end is not taken into account. It should be noted that alternatively, the nucleotides can be assigned sequentially to the bases of a cube or tetrahedron in a 3'→5' sense.

[0047] Given a particular nucleotide sequence, the nucleotide sequence can be tagged so that it can be easily distinguished among other sequences. The matrix representation method described herein can serve as a method for tagging a sequence. Thus, the second aspect of the present invention relates to a computer-implemented method for tagging a sequence using a matrix representation by encoding the nucleotide sequence in the form of a matrix according to any of the methods described in the first aspect of the present invention.

[0048] A third aspect of the present invention relates to a computer-implemented method for determining modifications of nucleotide sequences. Deep sequencing, human genome polymorphisms, or pan-genome studies are activities that require dealing with large amounts of nucleotide sequences. During the process of processing, exchanging, or storing this information, errors can occur that threaten the quality of the conclusions extracted from these data. As shown in Figure 10 and in Example 2, the TISA matrix is ​​very sensitive to sequence variations. Indeed, any modification, even a small modification (SNP), changes the winding path of the considered sequence, and therefore the structure of the TISA matrix and the results of its determinant. Similarly, TISA can be used to represent and score genetic mutations between related or unrelated sequences.

[0049] The computer-implemented method for determining modifications in a nucleotide sequence comprises the following steps. a) calculating from at least one selected nucleotide sequence representing a template sequence and at least one selected sequence representing at least one nucleotide sequence for which a modification is to be determined a matrix representation of the sequence according to any of the methods of the first aspect of the invention, in which the sequence of nucleotides is represented in the form of a matrix, preferably the template sequence being the expected sequence of the sequence for which a modification is to be determined. b) individually comparing equivalent matrix elements of at least one template sequence and at least one sequence for which an alteration in the matrix representation is determined, the equivalent matrix elements of each matrix being those elements located in the same position (i.e., column and row) of the represented matrix. c) determining that a modification has occurred between the at least one template sequence and the at least one sequence in which the modification is to be determined when any of the equivalent matrix elements are not equal.

[0050] Note that perfect palindrome sequences and sequences with the same compositions 5→3' and 3'-5' have the same TISA signature, but whenever any of the equivalent matrix elements are not equal, there is certainly an alteration between the first and second sequences.

[0051] Alternatively, a computer-implemented method for determining modifications of a nucleotide sequence comprises the following steps. a) from at least one selected nucleotide sequence representing a template sequence and at least one selected sequence representing at least one nucleotide sequence in which a modification is to be determined. Calculating a matrix representation of the sequence according to any of the methods of the first aspect of the invention, wherein the sequence of nucleotides is represented in the form of a square matrix, preferably the template sequence is the predicted sequence of the sequence for which sequence modifications are to be determined. b) computing the determinant of the matrix representation of the array. c) Comparing the determinant of at least one template sequence and at least one sequence for which modifications of the matrix representation are determined. d) determining that a modification has occurred between the at least one template sequence and the at least one sequence in which a modification is to be determined when the determinants of step c) are not equal.

[0052] Similarly, while perfect palindrome sequences and sequences with the same compositions 5→3' and 3'-5' will have the same TISA signature, it should be noted that the chances of two completely different matrices having the same determinant are minimal, and whenever the determinants are not equal there is a definite alteration between the first and second sequence.

[0053] Advantageously, using determinants instead of the equivalent elements of the TISA matrix further simplifies the comparison process, further compressing the relevant information by simply assigning a determinant to each array.

[0054] In a preferred embodiment of any of the embodiments of the third aspect of the invention, the modification is one or more of a SNP, a nucleotide insertion, a nucleotide deletion or a nucleotide transition. As shown in Example 2, a nucleotide insertion, a nucleotide deletion or a nucleotide transition during storage / retrieval or exchange of a sequence of nucleotides through a network generates a different TISA matrix.

[0055] Therefore, in another preferred embodiment of any of the embodiments of the third aspect of the present invention, the method further comprises determining that the integrity of the biological sequence is compromised if sequence modification is determined.As also shown in Example 2, the TISA matrix representation is a good checksum system, because any changes due to information loss, noise, errors will generate a different TISA matrix, so when the biological sequence is compromised, it can be determined by comparing the current TISA matrix with the previous TISA matrix, or by comparing the current determinant with the previous determinant.

[0056] After discovering the laws of symmetry embedded within the DNA alphabet, the presence of nucleotides or sequences that are invariant to transformations due to symmetry was analyzed. For that purpose, the RIP, MIP and RMIP methodologies were devised, standing for Rotation (R), Mirror (M) or Rotation + Mirror (RM) Invariant Patterns (IP), respectively. The principles of all these methodologies are exemplified for RIP in Figure 12, but are similar for MIP and RMIP (the only change is in the nature of the symmetry applied).

[0057] A fourth aspect of the invention relates to a computer-implemented method for defining one or more markers in a nucleotide sequence, the markers being characterized by symmetric invariance. The method comprises the steps of: a) receiving a sequence of nucleotides, the sequence of nucleotides comprising nucleotides from one of the nucleotide subsets [A, C, G, T] or [A, C, G, U]. Note that this can therefore represent any sequence of DNA, RNA or the set of four modified nucleotides. b) The length of the base is an integer n 2 Selecting a window of the sequence that is a power of 2 for analysis. The sequence of nucleotide bases in length is at least n 2 Note that, it must be the case that,. This can be seen in Figure 11, step 1. c) Base length n 2 into n fragments of length n. Note that the sequence is divided into n fragments such that the first nucleotide of each fragment is contiguous with the last nucleotide of the previous fragment. d) Assembling the sequences into an n×n matrix. The matrix is ​​preferably assembled by placing each sequence fragment as a row in the matrix, with consecutive fragments in consecutive rows. Note, however, that other obvious assembly strategies may be used, such as placing the fragments as columns, or using other distributions of fragments. This can be seen in step 2 of FIG. 11. e) Computing several transformations of the matrix using one or more of the selected symmetries until all feasible transformations for each of the selected symmetries have been exhausted. This can be seen in step 3 of FIG. f) determining the elements of the matrix that are invariant under all transformations for each of the selected symmetries. To do this, the equivalent matrix elements in each of the matrix symmetries are compared and those elements that are invariant under all transformations are so determined. This can be seen in steps 4 and 5 of FIG. 11. g) assigning invariant elements of the matrix as invariant markers. h) Repeating steps a) to g) using another portion of the sequence until all the sequence has been analyzed, preferably by sliding the start of the sequence to be analyzed by one or more nucleotides. This can be seen in step 6 of FIG. 11.

[0058] The RIP methodology, as well as its derived variants (MIP and RMIP), is advantageous because it can identify and distinguish markers and genetic landmarks within a single nucleotide sequence, even in the absence of polymorphisms. Thus, RIP, MIP and RMIP methodologies are suitable tools for studying the structure, function, and evolution of genes and genomes.

[0059] In another preferred embodiment, the one or more symmetries are either rotational symmetry and / or mirror symmetry.As shown in Example 3, the square matrix has several rotational and mirror symmetries.It should be noted that the matrix has three rotational symmetries (see FIG. 11) and four mirror symmetries (across each diagonal and through the middle row and middle column).Any combination of rotational symmetry and / or mirror symmetry can be applied.

[0060] Interestingly, the number of invariant markers is a multiple of the number of applied transformations, and optionally, there is one more invariant corresponding to the center of rotation. Note that for the RIP in Figure 12, there are 3 rotations + identity (center of rotation), so the number of invariant markers is a multiple of 4.

[0061] In a preferred embodiment of any of the embodiments of the fourth aspect of the present invention, n is preferably an odd number, more preferably a prime number, more preferably 5 or more. Note that when n is an odd number, the matrix is ​​centered on the nucleotide, and therefore the invariant element is always found in the central element (n / 2, n / 2) of the matrix. This element should not be considered as a marker.

[0062] RIPs, MIPs and RMIPs are the tools of choice for studying genotypes, phenotypes and biological traits, since they can address any biological sequence (nucleotides, proteins, epigenetic markers, sugars, etc.). Indeed, like previously described markers such as SNPs, RFLPs, mass spectrometry signatures, they can be used to establish associations between the presence or absence of a given invariant marker and a given biological characteristic relevant for health prevention, disease detection, breeding, population studies, etc.

[0063] The presence or absence of invariant marker and invariant marker pattern can be related to physiological state.Furthermore, there can be a correlation between the presence or absence of one or more potential biological markers on one or more n×n matrices and the expression of a given genotype, phenotype or biological trait.Therefore, when a strong correlation is found, the invariant marker can be used to indicate physiological state and / or the expression of a given genotype, phenotype or biological trait.

[0064] The computer-implemented methods of any of the foregoing embodiments of the first, second, third and / or fourth aspects of the invention may be executed on any suitable computing system including one or more processing units such as a microprocessor, GPU, CPU, multi-core processor, etc., a server, or a distributed computing system such as a cloud.

[0065] A fifth aspect of the invention relates to a computer program product comprising instructions which, when executed by a processor, cause a computer to perform any of the steps of the methods according to the first, second, third and / or fourth aspects of the invention. The computer program product may be implemented in hardware and / or in software.

[0066] 13 illustrates a computer program product 100 for implementing the methods disclosed herein, which may be implemented in software, hardware, or a combination thereof. The modules may be stored in the memory of the system, or may be stored remotely, for example, on a remote server or distributed computing system communicatively connected to the system. The memory may include one or more volatile or non-volatile memory devices, such as DRAM, SRAM, flash memory, read-only memory, ferroelectric RAM, hard disk drives, floppy disks, magnetic tapes, optical disks, and the like.

[0067] In some embodiments, the modules may include a representation module 110 configured to represent a sequence of nucleotides as a cube or tetrahedron, and optionally as a matrix, according to any of the embodiments of the first aspect of the invention.

[0068] In some embodiments, the module may include a tagging module 120 for receiving a matrix representing the sequence nucleotides from the representation module 110 and assigning it to sequence nucleotides according to the second aspect of the invention, thereby assigning the matrix to the sequence of nucleotides.

[0069] In some embodiments, the module may include a modification determination module 130 configured to determine a modification of a sequence of nucleotides by comparing the sequence of nucleotides to a template sequence according to any of the embodiments of the third aspect of the present invention.

[0070] In some embodiments, the modules may include a transformation module 140 configured to obtain a matrix resulting from applying a symmetric transformation to a square matrix.

[0071] In some embodiments, the modules may include an invariance determination module 150 configured to overlay the original matrices and the transformed matrices from the transformation module and calculate matrix elements that are equal (same rows and columns) across all matrices.

[0072] On the other hand, as shown in the examples, as shown in FIG. 14, TISA matrices are particularly advantageous for supporting the exchange and analysis of large amounts of genomic information between two remote locations. They provide a fast, robust, agile, and lightweight data structure that can be easily exchanged over telecommunication networks such as the Internet, wide area networks (WANs), local area networks (LANs), satellite networks, and cellular networks. TISA matrices are compact convolutional representations of genomic information that carry only a small fraction of the weight of the transported information (measured in bits, typically less than 1% of the original weight). When considered together with a secondary matrix called the TISA deconvolution matrix (TDM), the TISA matrix can be easily decoded to extract the original genomic information in a lossless manner. This communication scheme does not violate the Shannon information theory principle, since the amount of information contained in the sum of the TISA+TDM matrices is equal to or greater than the calculated Shannon entropy.

[0073] The construction method of the TDM matrix is ​​illustrated in FIG.

[0074] Meanwhile, Figure 15D represents a typical convoluted TDM matrix. In this configuration, the position and location of each nucleotide in the original sequence can be obtained by dividing each and every TDM individual number (Gf, Af, Tf, Cf, Gb, Ab, Tb, Cb) by its corresponding prime number. Only one of the eight possible numbers (Gf, Af, Tf, Cf, Gb, Ab, Tb, Cb) will be divisible by this prime number and thus indicate its position on the GenoCube.

[0075] Thus, a sixth aspect of the present invention relates to a computer implemented method for representing a selected sequence of RNA and / or DNA nucleotides and / or nucleotide analogues, the method comprising: a) defining a cube using nucleotide bases, each of the eight vertices of the cube being assigned to a nitrogenous base from a subset [A (adenine), C (cytosine), G (guanine), T (thymine)], the nitrogenous bases T (thymine) and U (uracil) being understood herein to be equivalent and therefore interchangeable, each of the vertices of the cube being assigned to a nucleotide base from the subset such that for each vertex, the assigned base is directly connected via edges to all other base types of the subset; b) assigning the first nucleotide base of the selected sequence of nucleotides to a vertex of a cube to which the nitrogenous base type of the nucleotide has been assigned, and assigning each subsequent nucleotide of the selected sequence of nucleotides sequentially to a corresponding vertex of a cube or tetrahedron such that the assigned vertex is directly connected to the vertex of the previous nucleotide base through an edge of the cube; If the nucleotide base is equal to the previous nucleotide base, then the nucleotide base is assigned to the same vertex of the cube as the previous nucleotide base was assigned to; the selected sequence of nucleotides and / or nucleotide analogues comprises nitrogenous bases selected from A (adenine), C (cytosine), G (guanine), T (thymine) and / or U (uracil); Includes.

[0076] In a preferred embodiment of the sixth aspect of the invention, the method comprises the steps of: c) assigning a prime number to each consecutive nucleotide position throughout the length of the selected sequence, where the prime number assigned is different for each position in the selected sequence unless a pattern of one type of nitrogenous base (A (adenine), C (cytosine), G (guanine), T (thymine), or U (uracil)) is repeated and two or more repeats of the same type of nitrogenous base are immediately adjacent to each other, in which case each position in the pattern is assigned the same prime number; d) identifying all nucleotide positions assigned to each vertex of the cube and calculating the product of all prime numbers assigned to each of said nucleotide positions, including those repeated, if any, for each vertex of the cube; Further includes:

[0077] In a preferred embodiment of the sixth aspect of the present invention, the prime number is ascending order from the 5' end of the sequence of selected nucleotides to the 3' end of the sequence of selected nucleotides; in descending order from the 5' end of the sequence of selected nucleotides to the 3' end of the sequence of selected nucleotides; ·random, in ascending order from 5' to 3' of the sequence of selected nucleotides, beginning with the prime number 2 being assigned to the first nucleotide of the sequence of selected nucleotides located at the 5' end of the sequence of selected nucleotides; or in ascending order from 3' to 5' of the sequence of selected nucleotides, beginning with the prime number 2 being assigned to the first nucleotide of the sequence of selected nucleotides located at the 3' end of the sequence of selected nucleotides; is assigned to each consecutive nucleotide position over the entire length of the selected sequence according to one of the following:

[0078] In the present invention, prime numbers are preferably assigned to each successive nucleotide position throughout the entire length of the selected sequence of nucleotides in ascending order from the 5' to 3' direction, starting with the prime number 2 being assigned to the first nucleotide of the selected sequence of nucleotides located at the 5' end of the selected sequence of nucleotides.

[0079] In another preferred embodiment of the sixth aspect of the invention, the sequence of the selected nucleotides is further represented in the form of a matrix, in particular in the form of a TISA deconvolution matrix (TDM), and the method comprises the steps of: e) assigning each vertex of the cube to an element or component of a matrix having at least as many elements or components as there are vertices of the cube; f) for each element or component of the matrix representing the vertices of the cube, assigning the corresponding product to said vertices calculated in step d) of the sixth aspect of the invention; Further includes:

[0080] In yet another preferred embodiment of the sixth aspect of the present invention, the sequence of the selected nucleotides is represented as a TISA deconvolution matrix (TDM) in the form of a cube, preferably as a 4x4 matrix.

[0081] In yet another preferred embodiment of the sixth aspect of the invention, the sequence of selected nucleotides is represented in the form of a cube, the vertices of the cube being assigned to the matrix by a 3D projection of the cube onto a 2D Euclidean plane defining an inner and outer set of nucleotide bases each representing an opposite face of the cube, preferably with the projection axis perpendicular to the given initial face.

[0082] In yet another preferred embodiment of the sixth aspect of the invention, the sequence of selected nucleotides further comprises information indicating the 5' and 3' ends of the sequence of selected nucleotides, the nucleotides being assigned sequentially to the bases in the cube in the 5'→3' direction.

[0083] A seventh aspect of the invention relates to a computer implemented method of tagging a nucleotide sequence by encoding the nucleotide sequence in the form of a matrix according to any of the methods of the sixth aspect of the invention.

[0084] An eighth aspect of the present invention provides a computer implemented method for decoding a nucleotide sequence in the form of a TISA deconvolution matrix (TDM) as defined in the sixth aspect of the present invention, the method comprising: Integer factorization of each of the elements or components of the matrix representing the positions of the different vertices of the cube (corresponding to Gf, Af, Tf, Cf, Gb, Ab, Tb, Cb) into their prime factors; - assigning all the obtained prime factors, and therefore their corresponding nitrogenous bases, to positions in the sequence of selected nucleotides, each of the prime factors being at a predetermined position in the selected sequence; Includes.

[0085] In preferred embodiments of the eight aspects of the invention, prime numbers are assigned to each consecutive nucleotide position throughout the entire length of the selected nucleotide sequence in ascending order from 5' to 3', beginning with prime number 2 being assigned to the first nucleotide of the sequence of selected nucleotides located at the 5' end of the selected nucleotide sequence, and each resulting prime factor, and therefore its corresponding nitrogenous base, is assigned to a position in the sequence of selected nucleotides according to the ascending position of each prime number in the selected sequence.

[0086] In another preferred embodiment of the eight aspects of the invention, prime numbers are randomly assigned to each consecutive nucleotide position throughout the entire length of the selected sequence, and each resulting prime number, and therefore its corresponding nitrogenous base, is assigned to a position in the sequence of selected nucleotides according to the predetermined position of each of said prime numbers in the selected sequence.

[0087] The following examples are merely illustrative of the present invention and are not intended to be limiting thereof. EXAMPLES

[0088] Example 1: Array matrix encoding example As a first example, the encoding of a nucleotide sequence into a matrix is ​​done in a stepwise process, as shown in Figures 8A-8D.

[0089] The initial nucleotide sequence in the 5'→3' sense is GATTGCCTACCT as shown in FIG. 8A. Then, in FIG. 8B, a cube is defined, with each base connected to all base types through its edges. Then, the front of the cube is defined as the given initial face, and the sequence is wound along the vertices, such that the assigned base is connected to the base of the previous nucleotide through the edge of the cube. Then, in FIG. 8C, a 3D projection of the cube on a 2D Euclidean plane that defines the inner and outer squares is shown, as well as how the winding of the sequence looks from that perspective. Finally, in FIG. 8D, the number of nucleotides assigned to each base of the cube is assigned to each such vertex of the cube, respectively, and the nucleotide sequence is encoded into a square matrix, with the vertices represented on the diagonal elements of the matrix. The square matrix is ​​a Topology Invariant Sequence Array (TISA).

[0090] Note that the TISA matrix preserves the topology of the genocube. If the sequence to be encoded contains a stretch of repeated nucleotides (e.g., TT or CC in Example A), the redundancy is captured by calculating 2 units at the corresponding vertices. The sum of all positions in the TISA matrix is ​​identical to the length of the encoded nucleotide sequence (here 12 nucleotides).

[0091] Additionally, Figure 9 shows the TISA matrix of a more complex nucleotide sequence, namely the partial sequence of E. coli strain U 5 / 41 16S ribosomal RNA (accession number NR_024570.1). The final TISA matrix is ​​a 4x4 square matrix with the vertices of the cube represented on the diagonal elements of the matrix. The TISA wrapping of an actual DNA sequence around a genomic cube illustrates the level of compaction and compactness obtained by this method.

[0092] This shows that it is possible to encode nucleotide sequences via matrices while preserving the topological information of the wrapped nucleotide segments in the cubic form. Note that the same applies to winding a wire on a tetrahedron, resulting in a simple matrix.

[0093] Example 2: Example of sensitivity of matrix encoding to changes The second example is directed to illustrating the high sensitivity of the proposed matrix encoding to small sequence changes, in accordance with one or more embodiments.

[0094] As shown in Figure 10, the elements of the TISA matrix from the sequence introduced in Figure 9 (a partial sequence of E. coli strain U 5 / 41 16S ribosomal RNA (accession number NR_024570.1)) change significantly when even small modifications are made.

[0095] First, two compensatory changes were made by substituting two base positions so that each base was present in the sequence in equal proportions. All elements of the matrix representing the vertices are changed.

[0096] Second, even smaller changes were tested: single nucleotide polymorphisms (i.e. changing one nucleotide to another using a different base), in which case all the elements of the matrix representing the vertices are changed.

[0097] The changes in the TISA matrices are significant, but can be further noticed by calculating the determinants of each of the TISA matrices: for the compensatory changes, the determinant decreased by more than 11 million, whereas for the single nucleotide polymorphisms, it increased by more than 8 million.

[0098] Thus, it can be seen that the obtained unique ID (either the TISA matrix or its determinant) can be used to check the integrity of the nucleotide sequence during storage / retrieval or exchange over a network. Variations due to information loss, noise, and errors will generate different TISA matrices. The TISA matrix can be treated as a compact representation of the encoded sequence / gene / genome since it preserves the topological information of the wrapped nucleotide segments. Thus, TISA encoding can serve as a method of tagging sequences. Moreover, this further indicates that TISA encoding can be used to study the effect of mutations on nucleotide sequences by representing and scoring genetic mutations among related or unrelated sequences.

[0099] Example 3: Example of the determination of invariant markers in a sequence of nucleotides by rotational symmetry (RIP) Finally, we present a fourth example showing different rotationally invariant patterns (RIPs) in two randomly generated nucleotide sequences.

[0100] For a given sequence, a window length n of 17 is preselected according to one or more embodiments. Thus, a 17×17 matrix is ​​first given in FIG. 13 (step 2). After applying some rotational symmetry, symmetry-invariant elements are identified from which positional information, i.e., the position in the matrix of the symmetry-invariant elements and a qualitative value, i.e., the nature of the element (base type), can be obtained.

[0101] Note that since n is odd, the center of the matrix is ​​always invariant (the center of rotation). This element is not informative in itself and must be discarded. If a smaller n is chosen, an alternative matrix can be constructed and a different symmetric invariant element placed. Similarly, for longer sequences, further RIPs can be calculated by shifting the first nucleotide. Note that although this example refers to RIPs, the same considerations can be applied to other variants such as mirror invariant patterns (MIPs) and rotated mirror invariant patterns (RMIPs).

[0102] RIP, MIP and RMIP methodologies can be suitable tools to study the structure, function, and evolution of genes and genomes.

[0103] Example 4 - Convolution of the TDM matrix As shown in Figure 14, TISA matrices are particularly advantageous for supporting the exchange and analysis of large amounts of genomic information between two distant locations. They provide a fast, robust, agile, and lightweight data structure that can be easily exchanged over telecommunication networks such as the Internet, wide area networks (WANs), local area networks (LANs), satellite networks, and cellular networks. TISA matrices are compact convolutional representations of genomic information that carry only a small fraction of the weight of the transported information (measured in bits, typically less than 1% of the original weight). When considered together with a second-order matrix called the TISA deconvolution matrix (TDM), the TISA matrix can be easily decoded to extract the original genomic information in a lossless manner. This communication scheme does not violate Shannon information theory principles, since the amount of information contained in the sum of the TISA+TDM matrices is equal to or greater than the calculated Shannon entropy.

[0104] The construction method of the TDM matrix is ​​illustrated in Figure 15. Using the following scheme, the first 40 DNA nucleotides constituting the M. tuberculosis esaX gene encoding the ESAT-6 immunogenic protein (Figure 15A) were subjected to both TISA encoding (Figure 15B) and TDM encoding (Figure 15D). TISA encoding was performed as shown in the previous example. The TDM was encoded as shown in Figure 15C using the following three successive steps: · Step #1- The gene sequence corresponding to the esaX gene (40 nucleotides long) was tabulated. ○ The first line contained the nucleotides of the sequence being considered (in the 4 to ''4 direction). The second row contained the prime numbers attributed to each nucleotide, starting with 2. For consecutive redundant nucleotides such as GG, AA, TTT, GG and CC in this example, the same prime number was assigned multiple times. · Step #2 - The third row of the table noted the location of each nucleotide / prime pair once represented on the GenoCube. Here, the GenoCube was chosen to display the GATC sequence on the front of the cube arranged in a clockwise orientation. The corresponding sequence of nucleotides and associated primes were assigned specific symbols depending on whether they were in a corner located in the front (Gf, Af, Tf, Cf) or back (Gb, Ab, Tb, Cb) of the GenoCube. · Step #3- The prime numbers corresponding to each of the eight corners of the cube, labeled as Gf, Af, Tf, Cf, Gb, Ab, Tb, and Cb respectively, were multiplied together and the corresponding products were reported in the TDM matrix (Figure 15D).

[0105] Example 5 - Deconvolution of the TDM matrix Figure 15D shows a typical convoluted TDM matrix. In this configuration, the position and location of each nucleotide in the original sequence can be obtained by dividing each and every TDM individual number (Gf, Af, Tf, Cf, Gb, Ab, Tb, Cb) by its corresponding prime number. Only one of the eight possible numbers (Gf, Af, Tf, Cf, Gb, Ab, Tb, Cb) will be divisible by this prime number and thus indicate its position on the GenoCube.

[0106] For example, the start of the esaX sequence (the A of ATG) can be identified at the Af position, since that is the only position divisible by 2. In general, the position of the first nucleotide of the sequence is found at a position of even parity, since 2 is the only even prime number.

[0107] To reconstruct the original nucleotide sequences encoded in the TISA and TDM matrices, a two-step method can be followed: Step #1 - Proceed with integer factorization into prime factors at each of the positions (Gf, Af, Tf, Cf, Gb, Ab, Tb, Cb). Step #2- Collect all the obtained prime factors in a list and arrange them in ascending order. Assign each prime factor its corresponding nucleotide value as deduced from its position in the TDM matrix.

[0108] An example of steps 1 and 2 is shown in Figures 16A and 16B, respectively. Note that the value of every position in the TISA matrix corresponds to the number of factors of the corresponding number in the TDM matrix. As an example, the number 14 found in position Gb of the TISA matrix counts the number of prime factors (14) of the number 2389558903510844901830255 (5, 17, 23, 47, 59, 59, 73, 83, 83, 83, 107, 113, 113, 131).

Claims

1. A computer implementation method for representing selected sequences of RNA and / or DNA nucleotides and / or nucleotide analogs, g) A step of defining a cube using nucleotide bases, wherein each of the eight vertices of the cube is assigned a nitrogen-containing base from a subset [A (adenine), C (cytosine), G (guanine), T (thymine)], where the nitrogen-containing bases T (thymine) and U (uracil) are understood herein to be equivalent and therefore interchangeable, and each of the vertices of the cube is assigned, for each vertex, to a nucleotide base from the subset such that the assigned base is directly connected to all other base types of the subset via the edges, h) The steps of assigning the first nucleotide base of the selected nucleotide sequence to the vertex of the cube to which the nitrogen-containing base type of the nucleotide is assigned, and sequentially assigning each subsequent nucleotide of the selected nucleotide sequence to the corresponding vertex of the cube or tetrahedron, such that the assigned vertices are directly connected to the vertex of the previous nucleotide base via the edge of the cube, If the nucleotide base is equal to the previous nucleotide base, the nucleotide base is assigned to the same vertex of the cube to which the previous nucleotide base was assigned. The selected sequence of the nucleotide and / or nucleotide analog comprises a nitrogen-containing base selected from A (adenine), C (cytosine), G (guanine), T (thymine), and / or U (uracil), Computer implementation methods, including those mentioned above.

2. i) A step of assigning a prime number to each consecutive nucleotide position throughout the entire length of the selected sequence, wherein a pattern of one type of nitrogen-containing base (A (adenine), C (cytosine), G (guanine), T (thymine), or U (uracil)) is repeated, and two or more repetitions of the same type of nitrogen-containing base are directly adjacent to each other, in which case the assigned prime numbers are different for each position in the selected sequence unless each position in the pattern is assigned the same prime number; j) Identifying all the nucleotide positions assigned to each vertex of the cube and calculating the product of all the prime numbers assigned to each of the nucleotide positions, including any repetitions for each vertex of the cube, The method according to claim 1, further comprising:

3. The aforementioned prime numbers are a. In ascending order from the 5' end of the selected nucleotide sequence to the 3' end of the selected nucleotide sequence, b. Descending order from the 5' end of the selected nucleotide sequence to the 3' end of the selected nucleotide sequence, c. Random, d. Starting in ascending order from 5' to 3' in the selected nucleotide sequence, the prime number 2 is assigned to the first nucleotide of the selected nucleotide sequence located at the 5' end of the sequence, or e. Starting from the 3' to 5' direction of the selected nucleotide sequence, the prime number 2 is assigned to the first nucleotide of the selected nucleotide sequence located at the 3' end of the sequence, in ascending order. The method according to claim 2, wherein each consecutive nucleotide position is assigned over the entire length of the selected sequence according to one of the following:

4. The sequence of the selected nucleotides is further represented in matrix form, and the method is k) The step of assigning each vertex of the cube to an element or component of a matrix having at least the same number of elements or components as the vertices of the cube, l) For each of the elements or components of the matrix representing the vertices of the cube, assign the calculated corresponding product to the vertex calculated in step d) of claim 2, The method according to any one of claims 2 or 3, further comprising:

5. The method according to claim 4, wherein the matrix is ​​a square matrix.

6. The method according to claim 5, wherein the sequence of selected nucleotides is represented in the form of a cube.

7. The method according to claim 6, wherein the sequence of selected nucleotides is represented in the form of a cube, and the vertices of the cube are assigned to the matrix by a 3D projection of the cube onto a 2D Euclidean plane that defines an inner and outer set of nucleotide bases, each representing an opposing face of the cube.

8. The method according to any one of claims 1 to 3, wherein the sequence of the selected nucleotides further includes information indicating the 5' and 3' ends of the sequence of the selected nucleotides, and the nucleotides are sequentially assigned to the bases in the cube in the 5'→3' direction.

9. A computer implementation method for tagging a nucleotide sequence by encoding the nucleotide sequence in matrix form according to the method of claim 4.

10. A computer implementation method for decoding the nucleotide sequence in the matrix form described in claim 9, a. A step of factoring each of the elements or components of the matrix representing different vertex positions of the cube into their prime factors, b. A step of assigning all obtained prime factors, and therefore their corresponding nitrogen-containing bases, to positions in the sequence of the selected nucleotides, such that each of the primes is at a predetermined position in the sequence of the selected nucleotides; Computer implementation methods, including those mentioned above.

11. The computer implementation method according to claim 10, wherein the prime numbers are assigned to each consecutive nucleotide position along the entire length of the selected nucleotide sequence in ascending order from the 5' to 3' direction of the selected nucleotide sequence, with prime number 2 being assigned to the first nucleotide of the selected nucleotide sequence located at the 5' end of the selected nucleotide sequence, and each of the obtained prime factors, and therefore its corresponding nitrogen-containing base, is assigned to the position in the selected nucleotide sequence according to the ascending position of each prime number in the selected sequence.

12. The computer implementation method according to claim 10, wherein the prime numbers are randomly assigned to each consecutive nucleotide position over the entire length of the selected sequence, and each of the obtained prime factors, and therefore its corresponding nitrogen-containing base, is assigned to a position in the sequence of the selected nucleotide according to the predetermined position of each of the prime numbers in the selected sequence.