Sequencing positioning and matching method and apparatus, spatial omics sequencing device, and medium
By building an index information database and a matching information database, combining comparison priority, editing distance and spatial distance, the problem of low matching rate of position barcodes in spatial numerology sequencing is solved, which improves comparison efficiency and efficiency, and provides more data for downstream analysis.
Patent Information
- Application Number
- PCT/CN2024/140208
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-24
AI Technical Summary
In the prior art, the position barcode matching rate is low during spatial acoustics sequencing, resulting in the position barcode matching rate obtained by secondary sequencing, especially when processing position barcodes of larger lengths, the relative alignment rate is lower and the effective efficiency is reduced.
By splitting the first position barcode into multiple reference short sequences to build an index information database, splitting the second position barcode into multiple short sequences to be selected for comparison, and determining the successful matching position barcode based on the comparison priority, editing distance and spatial distance, improving the comparison efficiency.
Improve the efficiency and effectiveness of position barcode comparison, reduce mismatch, and provide more downstream analysis data.
Smart Images

Figure CN2024140208_24072025_PF_FP_ABST
Abstract
Description
Sequencing positioning matching method and device, spatial omics sequencing equipment and medium
[0001] The present invention claims priority to Chinese patent application number 202410076175.4, filed with the Patent Office of China on January 18, 2024, entitled “Sequencing Positioning Matching Method and Device, Spatial Omics Sequencing Equipment and Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of spatial omics information technology, and in particular to a spatial omics sequencing positioning matching method and apparatus, a spatial omics sequencing device, and a computer-readable storage medium. Background Art
[0003] Transcription is the process by which genetic information flows from DNA to RNA. It involves the synthesis of RNA using one strand of double-stranded DNA (the template strand used for transcription, not the coding strand) as a template and the four ribonucleotides A, U, C, and G as raw materials, catalyzed by RNA polymerase. As the first step in protein biosynthesis, transcription involves the reading and copying of a gene into mRNA. Specifically, a specific DNA fragment serves as a template for genetic information, and DNA-dependent RNA polymerase acts as a catalyst to synthesize precursor mRNA based on the principle of base complementarity. During mRNA transcription, the double-stranded DNA molecule opens, and under the action of RNA polymerase, the four free ribonucleotides bind to the single strand of DNA according to the principle of base pairing. RNA polymerase then reacts to form a single-stranded mRNA molecule, completing transcription.
[0004] Traditional gene expression analysis, such as RNA-Seq, can provide rich gene expression information, but it often requires grinding sample tissue into a mixture of single cells or RNA molecules, resulting in a loss of spatial information in the gene expression obtained through subsequent sequencing. Spatial transcriptomics, collectively referred to herein as spatial omics, is an interdisciplinary field that combines histology and gene expression analysis. It measures mRNA in intact tissue sections, combines the spatial information of mRNA with morphological content, and maps the locations of all gene expression to obtain a complete gene expression map of the organism, thereby aiming to understand the spatial heterogeneity of gene expression at the cellular level.
[0005] Spatial omics is the process of retaining and analyzing spatial information by applying microarray technology with positional barcodes on fixed tissue sections, allowing researchers to map the gene expression data obtained by sequencing back to its original spatial position based on this spatial information, providing a new perspective for understanding tissue structure and function, disease occurrence and progression. In the prior art, two-step sequencing is usually used to determine the spatial position of gene expression. First, the tissue section fixed on the slide is contacted with a microarray with positional barcodes, allowing the positional barcodes to mark the RNA molecules in the tissue. Then, the sequence (barcode) of each positional barcode and their corresponding coordinate information (spatial coordinates X, Y) are obtained through the first sequencing, and the RNA is reverse transcribed and amplified while retaining the positional barcode information; finally, a library, amplification and other processes are established for a second sequencing to obtain the RNA sequence and its corresponding positional barcode. By comparing the positional barcodes of the two sequencings, the RNA sequence is mapped back to its original spatial position.
[0006] However, errors may occur during both sequencing processes, resulting in a low matching rate of the position barcodes obtained by the secondary sequencing. Errors may arise from basic errors in the sequencing process, sequence duplications, or barcode design defects. Currently, the most popular technology in the industry for spatial omics is to use position barcodes to retain spatial information. For example, 10X Genomics has introduced a whitelist of pre-defined position barcodes, and then uses the correspondence between the position barcodes and the whitelist to determine the spatial position. These whitelists do not require re-sequencing for each experiment, which can avoid certain sequencing errors and improve the effective matching rate of the corresponding position barcodes. Commonly used gene quantitative alignment software that uses positional barcodes to compare with whitelists, such as STARsolo, is a tool for aligning RNA-Seq data based on STAR (Spliced Transcripts Alignment to a Reference). Although their performance is already very good, when processing spatial omics data, positional barcodes are often aligned directly and the matching standards are very strict to reduce the possibility of false matches. Even if a whitelist is introduced for alignment, it means that a certain proportion of actually valid positional barcodes may be mistakenly excluded; especially for alignment scenarios with positional barcodes of longer lengths, the alignment rate is low, which greatly reduces the effectiveness of the positional barcodes. Summary of the Invention
[0007] The present application provides a spatial omics barcode positioning method and apparatus, a computer device, and a computer-readable storage medium with a higher matching rate and the ability to improve the efficiency of position barcodes.
[0008] In a first aspect of the present application, a method for sequencing and positioning matching of spatial omics is provided, comprising:
[0009] Obtaining a first position barcode with spatial position information;
[0010] For each first position barcode, multiple reference short sequences are obtained, index information is established according to the sequence number and position of the first position barcode where each reference short sequence is located, and an index information library of the first position barcode is constructed;
[0011] Obtaining a second position barcode obtained by secondary sequencing;
[0012] For each second position barcode, a plurality of candidate short sequences are obtained, the candidate short sequences are compared with the index information library to determine matching short sequences, and a matching information library of the matching short sequences and the first position barcode is constructed;
[0013] Determining the comparison priority of the first position barcode according to the matching information database;
[0014] For each second position barcode, compare it with the first position barcode according to the comparison priority, determine the candidate first position barcode whose editing distance with the second position barcode meets the target requirements, calculate the spatial distance based on the spatial position information of the candidate first position barcode, and determine the second position barcode that has been successfully compared based on the spatial distance meeting the target conditions.
[0015] In a second aspect, a spatial omics sequencing positioning matching device is also provided, comprising:
[0016] An acquisition module, configured to acquire a first position barcode with spatial position information;
[0017] An index establishment module, configured to obtain a plurality of reference short sequences for each first position barcode, establish index information according to the sequence number and position of the first position barcode where each reference short sequence is located, and construct an index information library for the first position barcode;
[0018] The acquisition module is further used to obtain the second position barcode obtained by secondary sequencing;
[0019] a matching module configured to obtain, for each second position barcode, a plurality of candidate short sequences, compare the candidate short sequences with the index information library to determine matching short sequences, and construct a matching information library of the matching short sequences and the first position barcode;
[0020] a priority module, configured to determine a comparison priority of the first position barcode according to the matching information library;
[0021] A comparison module is used to compare each second position barcode with the first position barcode according to the comparison priority, determine a candidate first position barcode whose editing distance with the second position barcode meets the target requirements, calculate the spatial distance based on the spatial position information of the candidate first position barcode, and determine the second position barcode that has been successfully compared based on the spatial distance meeting the target conditions.
[0022] In a third aspect, a spatial omics sequencing device is provided, comprising a processor and a memory connected to the processor, wherein the memory stores a computer program that can be executed by the processor, and when the computer program is executed by the processor, it implements the spatial omics sequencing positioning matching method as described in any embodiment of the present application.
[0023] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the spatial omics sequencing positioning matching method as described in any embodiment of the present application is implemented.
[0024] In the above embodiment, an index information library is constructed by splitting the first position barcode into multiple reference short sequences, and the second position barcode is split into multiple short sequences to be selected. The selected short sequences are compared with the reference short sequences to screen and determine the matching short sequences, and a matching information library of the matching short sequences and the first position barcode is constructed to obtain information on potential first position barcodes that have a higher probability of having a matching relationship with the second position barcode, and determine the comparison priority of the first position barcode. In the process of comparing the second position barcode with the first position barcode, the comparison can be performed according to the comparison priority of the first position barcode, which can improve the comparison efficiency; secondly, the first position barcode whose edit distance meets the target requirements can be locked more accurately and quickly as a candidate, and the relative spatial distance is calculated based on the spatial position information of the candidate first position barcode. The second position barcode that has been successfully matched is determined based on the spatial distance meeting the target conditions. In this way, the comparison priority, edit distance and the spatial distance between multiple candidate first position barcodes are combined to jointly assist in judging whether the second position barcode can be retained. The matching efficiency of the position barcode can be improved on the basis of reducing mismatches. Compared with the method of directly comparing the position barcodes in the prior art, it can be found that a large number of position barcodes abandoned by the existing matching standards can be recovered, which greatly improves the efficiency of the position barcode and provides more data for downstream analysis.
[0025] In the above embodiments, the spatial omics sequencing positioning matching device, spatial omics sequencing equipment and computer-readable storage medium belong to the same concept as the corresponding spatial omics sequencing positioning matching method embodiments, and thus have the same technical effects as the corresponding spatial omics sequencing positioning matching method embodiments, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG1 is a schematic diagram of an application scenario of a spatial omics sequencing positioning matching method in one embodiment.
[0027] FIG2 is a flow chart of a method for sequencing and positioning matching of spatial omics in one embodiment.
[0028] FIG3 is a flowchart of a method for sequencing location matching of spatial omics in an optional specific example.
[0029] FIG4 is a schematic diagram of the structure of a spatial omics sequencing positioning matching device in one embodiment.
[0030] FIG5 is a schematic diagram of the structure of a spatial omics sequencing device in one embodiment. DETAILED DESCRIPTION
[0031] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0033] In the following description, the expression "some embodiments" is involved, which describes a subset of all possible embodiments. It should be noted that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0034] In the following description, the terms "first, second, and third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first, second, and third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0035] Before further explaining the embodiments of the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations.
[0036] 1. Positional barcodes: A unique nucleotide sequence (DNA or RNA) used to label and identify individual molecules or cells in biological samples. Positional barcodes are typically character strings of a certain length. These barcodes play a crucial role in spatial transcriptomics, enabling RNA molecules to be mapped to their original locations in tissue samples.
[0037] 2. Position barcode whitelist refers to a list of all known barcode sequences included in the detection kit (designed in the spatial omics sequencing platform) and is available during library preparation.
[0038] 3. RNA, or ribonucleic acid, is a linear macromolecule composed of ribonucleotides polymerized through 3',5'-phosphodiester bonds. RNA in nature is usually single-stranded, and the four most basic bases in RNA are adenine (A), uracil (U), guanine (G), and cytosine (C). In contrast, DNA, which is also a nucleic acid like RNA, is usually a double-stranded molecule, and its nitrogenous bases replace RNA's uracil with thymine (T).
[0039] 4. Edit distance: This is a quantitative measurement parameter for the degree of difference between two character strings (e.g., the first position barcode and the second position barcode in this application). The measurement method is to see how many times the processing is required to transform one character string into another. Here, the number of processing times usually corresponds to the edit distance.
[0040] Spatial omics is a technique that spatially analyzes RNA-Seq data, analyzing all mRNAs within a single tissue section to locate and differentiate the active expression of functional genes within specific tissue regions. See Figure 1 for an example application scenario for sequencing-based mapping in spatial omics. The sequencing process of spatial omics mainly includes: 1. Sample preparation and quality inspection; 2. Tissue sectioning, staining and imaging, etc.; 3. Tissue permeabilization and cDNA synthesis, which mainly includes placing the tissue section on a slide containing capture probes bound to RNA (referred to as spatial transcriptome chips in the embodiments of this application, and the capture probes have unique position barcodes, thereby retaining the spatial position information of RNA), and fixing and permeabilizing (so that the mRNA in the cell is released and binds to the corresponding capture probes to obtain gene expression information). This process is a one-time sequencing; 4. Library construction, after cDNA synthesis using the captured RNA as a template, a sequencing library is constructed for the cDNA sample; 5. Sequencing the prepared library, such as using a high-throughput sequencing platform NGS (Next Generation Sequencing, second-generation sequencing technology) sequencing method for sequencing, this process is a secondary sequencing; 6. Data analysis and visualization (quality control and analysis of the data, and obtaining information on genes expressed in spatial positions, etc.).
[0041] Among them, the spatial omics sequencing positioning matching method provided in the embodiment of the present application is usually used in data analysis and visualization. The comparison of the first position barcode and the second position barcode is a key link in finally obtaining the gene expression at the spatial position in the tissue.
[0042] Please refer to FIG2 , which shows a method for spatial omics sequencing and positioning matching provided in one embodiment of the present application, including the following steps:
[0043] S101, obtaining a first position barcode with spatial position information.
[0044] The first position barcode (first barcode) is the barcode information obtained from a single sequencing run of spatial omics, recording the spatial position of the RNA on the spatial transcriptome chip. In one optional example, the processed first barcode file includes three columns of data: one column for the position barcode sequence consisting of the characters A, T, C, and G, one column for the spatial position coordinates X, and one column for the spatial position coordinates Y.
[0045] Optionally, step S101 includes:
[0046] Obtaining a location barcode whitelist, and obtaining a first location barcode with spatial location information from the location barcode whitelist; or,
[0047] A first position barcode with spatial position information is obtained by performing one sequencing based on the spatial group chip.
[0048] Here, the method for obtaining the first barcode with spatial location information is not limited to the whitelist or real-time sequencing. It can refer to reading the first barcode file obtained by a previous sequencing run, such as reading the position barcode whitelist to obtain the first barcode data; it can also refer to reading the first barcode data in real time during the sequencing run.
[0049] S102 : For each first position barcode, obtain multiple reference short sequences, create index information according to the sequence number and position of the first position barcode where each reference short sequence is located, and construct an index information library of the first position barcode.
[0050] Each reference short sequence refers to a string of characters within a first barcode. The position of the reference short sequence within the first barcode corresponds to the segment of the first barcode to which the reference short sequence belongs. This is typically represented by the relative distance of the reference short sequence from a reference position in the first barcode (e.g., the starting point of the first barcode). For example, if the position of a reference short sequence is 10, it means that the first position of the reference short sequence is located at the 10th position of the corresponding first barcode. Each first barcode is split into multiple reference short sequences. For each reference short sequence, its index information is established based on the serial number of the first barcode from which it originated and its position within the first barcode. An index information library for the first barcode is constructed based on the index information of the reference short sequences obtained from all first barcodes.
[0051] It should be noted that the index information library represents a data set containing the index information of all reference short sequences, in the form of, but not limited to, tables, documents, key values, etc. In the embodiment of the present application, the index information of the reference short sequence is a hash value, and the index information library refers to a hash table containing the index information corresponding to all reference short sequences.
[0052] S103, obtaining a second position barcode obtained by secondary sequencing.
[0053] The second barcode refers to the barcode information of the positional barcode sequence obtained by spatial omics secondary sequencing. According to the sequencing principle of spatial omics, the second barcode sequencing library has been separated from the spatial transcriptome chip during the construction process, and the spatial position information has been lost.
[0054] S104 , for each second position barcode, obtain multiple short sequences to be selected, compare the short sequences to be selected with the index information library to determine matching short sequences, and construct a matching information library of the matching short sequences and the first position barcode.
[0055] Each candidate short sequence refers to a string of characters within a second barcode. The method for obtaining a candidate short sequence from a second barcode is generally the same as the method for obtaining a reference short sequence from a first barcode, and the length of the candidate short sequence is the same as that of the reference short sequence. The candidate short sequence is compared with the index information library to determine a matching short sequence. For a candidate short sequence determined to be a matching short sequence, the matching relationship between the matching short sequence and the first barcode can be determined based on the sequence number of the first barcode containing the reference short sequence matched by the candidate short sequence, thereby constructing a matching information library for the matching short sequence and the first barcode. It is understood that the matching information library contains candidate short sequences that can be successfully matched with at least one reference short sequence in the index information library, and candidate short sequences that fail to match are not stored.
[0056] S105: Determine the comparison priority of the first position barcode according to the matching information database.
[0057] The first barcode's comparison priority refers to the order in which the first barcode is subsequently compared with the second barcode when it is used as a matching object. The higher the first barcode's comparison priority, the higher its comparison priority as a matching object, and the more preferentially it is compared with the second barcode during the comparison process. Based on the matching relationship between the short sequences and the first barcode in the matching information library, the matching rate between the short sequences to be selected in the second barcode and each first barcode can be calculated. The comparison priority of each first barcode is correspondingly associated with the matching rate of each first barcode determined based on the matching information library. In one optional example, the first barcode with the highest number of matches (highest matching rate) with the short sequences to be selected in the second barcode has the highest matching priority.
[0058] S106, for each second position barcode, compare it with the first position barcode according to the comparison priority, determine the candidate first position barcode whose editing distance with the second position barcode meets the target requirement, calculate the spatial distance based on the spatial position information of the candidate first position barcode, and determine the second position barcode that has been successfully compared based on the spatial distance meeting the target condition.
[0059] In the implementation process for matching the second barcode with the first barcode, each second barcode is sequence-aligned with the current second barcode based on the first barcode's alignment priority. In one optional example, the matching information for the short sequence and the first barcode is: {8:5,117:3,3890:1}. The left side of the colon is the first barcode's sequence number, and the right side is the frequency of occurrence of the matching short sequence in the first barcode. Based on the matching information between the short sequence and the first barcode, three first barcodes (sequence numbers 8, 117, and 3890) have the same short sequence as the second barcode. However, since the first barcode with sequence number 8 has the highest short sequence matching frequency, the first barcode with sequence number 8 has the highest alignment priority, and a complete sequence alignment with the second barcode will be performed starting from it.
[0060] Meeting the target edit distance requirement may mean that the edit distance is less than a preset value E. Meeting the target spatial distance requirement may mean that the spatial distance is less than a preset value D. The specific values of D and E can be adjusted in practice and are not limited in this application. In one optional example, E is 2 and D is 1. For each second barcode, the first barcode whose edit distance meets the target requirement is determined as a candidate first barcode. The relative spatial distance is then calculated based on the spatial position information of each candidate first barcode. If the spatial distance between multiple candidate first barcodes meets the target requirement, the second barcode is considered to be successfully matched.
[0061] In the above embodiment, the first barcode is split into multiple reference short sequences to construct an index information library, and the second barcode is split into multiple candidate short sequences. The candidate short sequences are compared with the reference short sequences to screen and determine matching short sequences, and a matching information library of the matching short sequences and the first barcode is constructed. In this way, information about potential first barcodes with a higher probability of matching with the second barcode is obtained, and the comparison priority of the first barcode is determined. When comparing the second barcode with the first barcode, the comparison can be performed according to the comparison priority of the first barcode, which can improve the comparison efficiency. Secondly, the first barcode whose edit distance meets the target requirements can be more accurately and quickly locked as a candidate, and the relative spatial distance is calculated based on the spatial position information of the candidate first barcode. The second barcode that has been successfully matched is determined based on whether the spatial distance meets the target conditions. In this way, the combination of the comparison priority, the edit distance, and the spatial distance between multiple candidate first barcodes can jointly assist in determining whether the second barcode can be retained. It can improve the matching efficiency of the position barcode while reducing mismatches. Compared with the method of directly comparing the position barcodes in the prior art, it can be found that a large number of second barcodes discarded by the existing matching standards can be recovered, greatly improving the efficiency of the position barcode and providing more data for downstream analysis.
[0062] In some embodiments, step S106 includes:
[0063] Determine a candidate first position barcode whose edit distance with the second position barcode meets a target requirement;
[0064] Taking the spatial position information of the candidate first position barcode corresponding to the minimum edit distance as the standard point, calculate the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the standard point, and determine the second position barcode that has been successfully matched based on whether the spatial distance meets the target condition.
[0065] In this embodiment, the spatial distance is calculated based on the spatial position information of the candidate first barcode. The corresponding candidate first barcode with the smallest edit distance to the second barcode is used as a reference point, and the spatial distances of the spatial position information of other candidate first barcodes relative to the reference point are calculated.
[0066] Take the following Table 1 as an example:
[0067] For a particular second barcode, the candidate first barcodes whose edit distances meet the target requirements include candidate first barcode1, candidate first barcode2, and candidate first barcode3 corresponding to edit distance 1, edit distance 2, and edit distance 3, respectively. Taking candidate first barcode1, corresponding to edit distance 1, which has the smallest edit distance, as the reference point, the spatial distance D1 between the spatial position information of candidate first barcode2 and the reference point and the spatial distance D2 between the spatial position information of candidate first barcode3 and the reference point are calculated. If spatial distances D1 and D2 respectively meet the target conditions, the second barcode is determined to be a successfully aligned second position barcode.
[0068] In the above embodiment, the second barcode is further screened for retention based on the spatial distance between multiple candidate first barcodes whose edit distances meet the target requirement, thereby reducing mismatches and improving the matching efficiency of the second barcode position barcode.
[0069] In some embodiments, the method of using the spatial position information of the candidate first position barcode corresponding to the minimum edit distance as a reference point, calculating the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the reference point, and determining the second position barcode that has been successfully matched based on the spatial distance satisfying the target condition includes:
[0070] Determine whether the candidate first position barcode corresponding to the minimum edit distance is unique;
[0071] If so, taking the spatial position information of the candidate first position barcode corresponding to the minimum edit distance as the reference point, calculate the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the reference point. When the spatial distance is less than the target threshold, it is considered that the comparison is successful, and the second position barcode is determined as the second position barcode that has been successfully compared;
[0072] If not, the spatial position information of the multiple candidate first position barcodes corresponding to the smallest edit distance is used as the standard point for iterative calculation. In any iterative calculation, the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the standard point is calculated. When the spatial distance is less than the target threshold, it is considered that the comparison is successful, and the second position barcode is determined to be the second position barcode that has been successfully compared.
[0073] In this embodiment, when the spatial distance is calculated using the candidate first barcode having the smallest edit distance to the second barcode as the standard point, an implementation scheme is further provided for possibly including multiple candidate first barcodes at the same edit distance.
[0074] The candidate first barcode with the minimum edit distance is not unique, as shown in Table 2 below:
[0075] For a certain second barcode, the candidate first barcode corresponding to the minimum edit distance includes the candidate first barcode1 and the first barcode4 corresponding to the edit distance 1. Then, the candidate first barcode1 and the first barcode4 are respectively used as the standard points for iterative calculation. In one iterative calculation with the candidate first barcode1 as the standard point, the spatial distance D1 between the spatial position information of the candidate first barcode2 and the standard point (candidate first barcode1) and the spatial distance D2 between the spatial position information of the candidate first barcode3 and the standard point (candidate first barcode1) are respectively calculated. When the spatial distances D1 and D2 respectively meet the target conditions, the second barcode is determined to be the second position barcode with successful matching. In one iterative calculation with the candidate first barcode4 as the standard point, the spatial distance D3 between the spatial position information of the candidate first barcode2 and the standard point (candidate first barcode4) and the spatial distance D4 between the spatial position information of the candidate first barcode3 and the standard point (candidate first barcode4) are respectively calculated. When the spatial distances D3 and D4 respectively meet the target conditions, the second barcode is determined to be the second position barcode with successful matching.
[0076] It should be noted that the second barcode alignment is successful if either the spatial distances D1 and D2 calculated using the candidate first barcode 1 as the reference point meet the target condition or the spatial distances D3 and D4 calculated using the candidate first barcode 4 as the reference point meet the target condition.
[0077] In the above embodiment, candidate first barcodes whose edit distances to the second barcode meet the target requirements may be initially located at different positions in the tissue. By selecting different candidate first barcodes as reference points for iteration, the error caused by initial positioning mismatches can be reduced, thereby improving the matching efficiency of the second barcode position barcode while reducing mismatches.
[0078] In some embodiments, determining whether the candidate first position barcode corresponding to the minimum edit distance is unique includes:
[0079] Based on the comparison between the second position barcode and the first position barcode, the candidate first position barcodes whose edit distance meets the target requirement are retained in the same vector, and whether the corresponding candidate first position barcodes at the same edit distance have repeated marker bits in a sequence of a preset length are recorded;
[0080] It is determined according to the mark bit whether the candidate first position barcode corresponding to the minimum edit distance is unique.
[0081] When the second barcode is compared with the first barcode, if the second barcode and the first barcode are completely matched (the edit distance is considered to be 0), the matching process can be stopped immediately and the second barcode is considered to be matched successfully. If the second barcode and the first barcode are not completely matched, all candidate first barcodes whose edit distance with the second barcode meets the target condition can be calculated and retained in the same vector. For example:
[0082] If the edit distance threshold is set to 3, the edit distance between the candidate first barcode 1 and the matching short sequence is 2, which means that the edit distance meets the target requirement.
[0083] Optionally, for each second barcode comparison result, for the candidate first barcode whose edit distance meets the target requirement, the following can be performed:
[0084] Edit distance 1(b2,b1_1)=1 b1_1(X,Y)=100.5,234.2
[0085] Edit distance 2(b2,b1_2)=2 b1_2(X,Y)=100.1,234.9
[0086] Edit distance 2(b2,b1_3)=2 b1_3(X,Y)=32,1198
[0087] Where b2 refers to the second barcode currently participating in the alignment, the left side is the edit distance value determined by the alignment result, and b1_1, b1_2, and b1_3 respectively represent the serial numbers of the first barcodes corresponding to the corresponding edit distance values.
[0088] Optionally, an integer can be used to represent the addition of candidate first barcodes with the same edit distance. The marker bit in the integer, such as the last bit, can be used to mark whether multiple candidates with the same edit distance have been added. For example, in 00001010, the second and fourth bits, counting from right to left, are marked as 1, so that it can be determined that alignments with edit distances of 2 and 4 appear in the alignment results. If the last bit, counting from right to left, is 0, it means that there are no two or more results with an edit distance of 2 or 4. If more alignment results are added, when there are other alignment results with an edit distance of 2 or 4, the last bit is marked as 1. In this way, the value of the marker bit in the integer can be used to determine whether the candidate first barcode corresponding to the minimum edit distance is unique.
[0089] In the above embodiment, a vector is used to record the candidates for comparison between the second barcode and the first barcode, and a flag bit is set in an integer to record whether there are duplicate first barcode candidates at the same edit distance, which is conducive to improving comparison efficiency.
[0090] In some embodiments, calculating the spatial distance between the spatial position information of the candidate first position barcode corresponding to the other edit distance and the reference point, and determining the second position barcode that has been successfully matched based on the spatial distance satisfying the target condition, includes:
[0091] According to the spatial position information of the candidate first position barcodes corresponding to other edit distances and the spatial position information of the standard point, the spatial distance between each candidate first position barcode and the standard point in different coordinate axis directions is calculated. When the spatial distances in different coordinate axis directions are all less than the target threshold, the comparison is deemed successful, and the second position barcode is determined as the second position barcode that has been successfully compared.
[0092] For the spatial distance to meet the target condition, it means that the coordinates in different coordinate axis directions (such as X and Y axis directions) in the spatial position information of the candidate first barcode are used to calculate the relative spatial distance with the reference point in the corresponding coordinate axis directions. When the spatial distances in different coordinate axis directions are all less than the target threshold, the comparison is considered successful.
[0093] In the above embodiment, the spatial distance between any candidate first barcode and the reference point is calculated based on whether the spatial distance in any axis direction is less than the target threshold, so as to better ensure that the error caused by mismatch is reduced and the comparison accuracy is improved.
[0094] In some embodiments, step S104 includes:
[0095] For each second position barcode, a sliding window is used to sequentially slide and interval extract a short sequence of length K to be selected on the second position barcode; wherein the interval between two adjacent short sequences to be selected is W;
[0096] Comparing each of the candidate short sequences with the index information library to determine whether there is a hit reference short sequence that is identical to the current candidate short sequence;
[0097] If so, based on the index information of the hit reference short sequence, determine the hit first position barcode corresponding to the current short sequence to be selected, and the position deviation of the current short sequence to be selected and the hit reference short sequence in their respective position barcodes; based on the position deviation satisfying a preset range, determine that the current short sequence to be selected is a matching short sequence; based on the serial number of the hit first position barcode corresponding to the matching short sequence and the number of repetitions of the hit reference short sequence, convert them into a hash value; establish directory information corresponding to the matching short sequence; and form a matching information library of the matching short sequence and the first position barcode based on the directory information.
[0098] Among them, splitting the second barcode into short sequences can improve the efficiency of short sequence splitting by using a sliding window. In this embodiment, a sliding window with a width of W is used to slide sequentially on the second barcode. After each time a candidate short sequence of length K is extracted, another candidate short sequence of length K is extracted again at a distance of the width of the sliding window W. For each second barcode, each time a candidate short sequence of length K is extracted, the candidate short sequence is compared with the reference short sequence in the index information library. If the same reference short sequence is found, the same reference short sequence is considered a hit reference short sequence. The position deviation is determined based on the position of the current candidate short sequence in the second barcode and the position of the hit reference short sequence in the corresponding first barcode. If the position deviation meets a preset range (such as M), the candidate short sequence is regarded as a matching short sequence, and the directory information corresponding to the matching short sequence is established accordingly. Finally, a matching information library is formed based on the directory information of all the short sequences determined to be matching short sequences.
[0099] It should be noted that the matching information library represents a data set containing directory information of all matching short sequences, in the form of, but not limited to, tables, documents, key values, etc. In the embodiment of the present application, the directory information of the matching short sequence is a hash value, and the matching information library refers to a hash table containing the directory information corresponding to all matching short sequences.
[0100] In the above embodiment, splitting the second barcode into short sequences and comparing them with the index information library built based on the first barcode can help improve comparison efficiency. While tolerating a certain comparison error, the first barcode with the highest matching probability can be selected as a candidate. This can better ensure that the error caused by mismatches is reduced while improving comparison efficiency and accuracy.
[0101] Optionally, the sequencing positioning and matching method further includes:
[0102] If not, discard the current short sequence to be selected; or,
[0103] If so, the current short sequence to be selected is determined to be a false match based on the position deviation exceeding the preset range, and the current short sequence to be selected is discarded.
[0104] For each second barcode, each time a candidate short sequence of length K is extracted, the candidate short sequence is compared with the reference short sequence in the index information library. If the same reference short sequence cannot be found, the current candidate short sequence is not saved and the next candidate short sequence is extracted and compared.
[0105] Alternatively, if the same reference short sequence is found, the same reference short sequence is used as the hit reference short sequence, and the position deviation is determined based on the position of the current candidate short sequence in the second barcode and the position of the hit reference short sequence in the corresponding first barcode. If the position deviation exceeds the preset range M, the current match is considered a false match, and the current candidate short sequence is not saved, and the next candidate short sequence is extracted and compared. For example,
[0106] second barcode:GCGGTCTGGATGTGCGAAACTACA, position: 13
[0107] first barcodeX: CGCGAAACTCGGGGCGAAACTCTA, position: 1
[0108] Assume M is 2. The candidate short sequence "GCGAAACT" in the second barcode is compared with the index information library and hits the reference short sequence "GCGAAACT" in the first barcode X. Based on the position "13" of the current candidate short sequence in the second barcode and the position "1" of the hit reference short sequence in the corresponding first barcode, it can be seen that the position deviation is 12. The position deviation exceeds the preset range M. Therefore, the current candidate short sequence and the hit reference short sequence are a false match and are not saved.
[0109] In the above embodiment, by splitting the second barcode into short sequences and comparing them with the index information library built based on the first barcode, the matching short sequences that successfully match are screened by combining the alignment identity and position deviation. Only the matching short sequence information is stored. This allows the first barcode with the highest matching probability to be screened as a candidate while tolerating a certain comparison error. This facilitates the rapid and accurate screening of valid second barcodes based on the candidate first barcodes.
[0110] In some embodiments, step S102 includes:
[0111] For each first position barcode, a sliding window is used to sequentially slide and intermittently extract a reference short sequence of length K on the first position barcode; wherein the number of bits between two adjacent reference short sequences is W;
[0112] The serial number and position of the first position barcode of each reference short sequence are converted into a hash value, index information corresponding to each reference short sequence is established, and an index information library of the first position barcode is formed based on the index information.
[0113] The length of the reference short sequence is the same as the length of the short sequence to be selected. In this embodiment, the first barcode is split into reference short sequences in the same manner as the second barcode is split into short sequences to be selected. A sliding window of width W is sequentially slid over the first barcode. After each reference short sequence of length K is extracted, another reference short sequence of length K is extracted after a distance equal to the width of the sliding window W. The index information of each reference short sequence is calculated based on the sequence number of the first barcode corresponding to the reference short sequence and its position in the corresponding first barcode. The index information is represented by a hash value to construct an index information library for the first barcode.
[0114] In an optional example, referring to the serial number of the short sequence on the first barcode (such as 1, 2, 3) and its position on the first barcode (such as a sequence of length K starting from the 10th position), it is converted into a 64-bit integer representation through bit operations. An example of binary representation is:
[0115] 0010110001011010101010000110000000000000000000000000000000000000000010. The first 32 bits represent the first barcode sequence number, and the last 32 bits represent the starting position of the first barcode. After conversion, they are stored in the hash table as the hash value of the corresponding reference short sequence.
[0116] In the above embodiment, an index is established for the first barcode in the form of a reference short sequence. The reference short sequence is extracted from the corresponding first barcode at intervals. This can reduce the overall data volume for subsequent comparison and facilitate subsequent comparison based on the index information library. All first barcodes with potential matching rates can be screened out while tolerating a certain comparison error, thereby improving efficiency and avoiding errors and misjudgments.
[0117] In some embodiments, for each first position barcode, extracting a reference short sequence of length K by sequentially sliding intervals on the first position barcode using a sliding window includes:
[0118] For each first position barcode, a sliding window is used to sequentially slide on the first position barcode to extract a reference short sequence of length K at intervals. During the sliding process of the sliding window on the first position barcode, the endpoint digits of length O are skipped according to the endpoint position of the first position barcode.
[0119] Wherein, based on the principle of sequencing library construction, the junction (i.e., endpoint) where the position barcode is connected to other sequences (such as connector adapter) is usually relatively low in sequencing quality value. For example, due to the inconsistency of biochemical reactions in the synthesis process of the probe on the spatial transcriptome chip, the length of the synthetic probe may also be incorrect; due to sequencing errors in the sequencing process, the boundary between the barcode and adapter may be incorrect, etc. By setting the sliding window in the sliding process on the first barcode, the endpoint position of the first barcode is skipped with a length of 0 respectively to avoid extracting the sequence at the endpoint position as a reference short sequence. Wherein, the value of length 0 can be adjusted in actual applications, and this application is not limited to this. The endpoint position of the first barcode usually includes two left and right endpoints. In actual applications, a certain position or an interval in sequencing can be preset as an endpoint, or sequencing quality judgment can be performed in the sequencing process, and the position where the sequencing quality is significantly reduced is regarded as the endpoint position.
[0120] In the above embodiment, by setting the endpoint to be skipped when extracting the reference short sequence from the first barcode, the accuracy of the extracted reference short sequence in representing the corresponding first barcode can be improved.
[0121] In some embodiments, converting the serial number and position of the barcode at the first position of each reference short sequence into a hash value and establishing index information corresponding to each reference short sequence further includes:
[0122] Hash values corresponding to multiple identical reference short sequences are stored in the same vector to form common index information of the reference short sequences.
[0123] In this embodiment, when creating a first barcode index information library, if duplicate reference short sequences are encountered, their corresponding index information is stored in a vector, such as [117, 235, 449]. This creates shared index information that corresponds to multiple first barcodes for the same reference short sequence. By designing this shared index information vector, index information for identical reference short sequences can be merged, simplifying the index information library and improving subsequent comparison efficiency.
[0124] In order to have a more comprehensive understanding of the spatial omics sequencing positioning matching method provided in the embodiments of the present application, please refer to FIG3 . The spatial omics sequencing positioning matching method is described below using a specific example. The spatial omics sequencing positioning matching method includes:
[0125] S11, read the first barcode, extract the reference short sequence and establish the index information library.
[0126] Extract a reference short sequence of length K from the first barcode at intervals W. Use bitwise operations to convert the first barcode's serial number and position corresponding to the reference short sequence into an integer and store it in a hash table. During the extraction of the reference short sequence K, adjust the endpoint position O of the first barcode.
[0127] S12, read in the second barcode, extract the short sequence to be selected and compare it with the index information library, and establish a matching information library based on the comparison results.
[0128] The candidate short sequence that is successfully aligned with the reference short sequence and whose position deviation on the corresponding position barcode is less than M is determined as the matching short sequence. According to the matching situation between the matching short sequence and the reference short sequence, the directory information of the matching short sequence (the serial number of the corresponding first barcode + the number of occurrences in the corresponding first barcode) is formed to build a matching information library.
[0129] S13, counting the first barcodes containing the most short sequence matches, and determining the alignment priority of each first barcode based on the statistics.
[0130] S14: Compare the second barcode according to the comparison priority of the first barcode.
[0131] S15, if the current second barcode completely matches the first barcode, stop comparing the current second barcode and save the successfully matched second barcode.
[0132] S16: If the current second barcode does not completely match the first barcode, the first barcode within the edit distance E is counted as the candidate first barcode.
[0133] S161 , taking the candidate first barcode with the smallest edit distance as the reference point, calculate whether the spatial distances between the candidate first barcodes corresponding to other edit distances and the reference point are within a range D.
[0134] S162: If yes, stop comparing and save the second barcode that has been successfully compared.
[0135] If not, S163 stops the comparison and discards the current second barcode.
[0136] It should be noted that the values of parameters K, W, O, M, E, and D can all be adjusted in actual applications to implement dynamic planning to complete the comparison of position barcodes.
[0137] The spatial omics sequencing positioning matching method provided in the embodiments of the present application has at least the following characteristics:
[0138] First, splitting the first barcode and second barcode into short sequences and using a dynamic programming method to locate and match them is beneficial to improving the alignment efficiency.
[0139] Second, for each second barcode, the matching priority is determined by identifying the first barcode with a higher matching probability based on short sequence positioning. The edit distance is introduced to first screen candidate first barcodes, and then the spatial distance is introduced to obtain the spatial position relationship between multiple candidate first barcodes. Therefore, the matching priority, edit distance, and the spatial distance between multiple candidate first barcodes are combined to assist in determining whether the second barcode is matched successfully and to stop matching in a timely manner. This can improve matching efficiency and increase the matching efficiency of position barcodes while reducing mismatches. Compared with the existing method of directly matching position barcodes, it can be found that a large number of second barcodes discarded by existing matching standards can be recovered, greatly improving the efficiency of position barcodes and providing more data for downstream analysis.
[0140] Third, in the prior art method of directly comparing position barcodes, it is assumed that there are no mismatches in whitelist comparisons, and strict matching criteria are introduced for whether the position barcodes are identical. However, this assumption does not hold true in actual applications, resulting in a large number of wasted position barcodes. The spatial omics sequencing positioning matching method provided in the embodiments of the present application combines alignment priority, edit distance, and the spatial distance between multiple candidate first barcodes to assist in determining whether the second barcode is successfully matched. It is found that a large number of position barcodes discarded by existing algorithms can be recovered by increasing mismatches and utilizing spatial information, greatly improving the efficiency of position barcodes and providing more data for downstream analysis.
[0141] On the other hand, the present application, referring to FIG4 , provides a sequencing positioning matching device for spatial omics, comprising: an acquisition module 11 for acquiring a first position barcode with spatial position information; an index establishment module 12 for acquiring a plurality of reference short sequences for each first position barcode, establishing index information according to the serial number and position of the first position barcode where each reference short sequence is located, and constructing an index information library for the first position barcode; the acquisition module 11 is also used to acquire a second position barcode obtained by secondary sequencing; a matching module 13 is used to acquire a plurality of short sequences to be selected for each second position barcode, and compare the short sequences to be selected with the selected short sequences. The index information library is compared to determine a matching short sequence, and a matching information library of the matching short sequence and the first position barcode is constructed; a priority module 14 is used to determine the comparison priority of the first position barcode according to the matching information library; a comparison module 15 is used to compare each second position barcode with the first position barcode according to the comparison priority, determine a candidate first position barcode whose editing distance with the second position barcode meets the target requirement, calculate the spatial distance based on the spatial position information of the candidate first position barcode, and determine the second position barcode that has been successfully compared based on the spatial distance meeting the target condition.
[0142] Optionally, the comparison module 15 is used to determine whether the candidate first position barcode corresponding to the minimum edit distance is unique; if so, the spatial position information of the candidate first position barcode corresponding to the minimum edit distance is used as the standard point, and the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the standard point is calculated. When the spatial distance is less than the target threshold, it is deemed that the comparison is successful, and the second position barcode is determined as the second position barcode that has been successfully compared; if not, the spatial position information of multiple candidate first position barcodes corresponding to the minimum edit distance is used as the standard point for iterative calculation. In any iterative calculation, the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the standard point is calculated. When the spatial distance is less than the target threshold, it is deemed that the comparison is successful, and the second position barcode is determined as the second position barcode that has been successfully compared.
[0143] Optionally, the comparison module 15 is also used to retain the candidate first position barcodes whose edit distance meets the target requirements in the same vector based on the comparison between the second position barcode and the first position barcode, and record whether the corresponding candidate first position barcodes at the same edit distance have a mark bit that is repeated in a sequence of a preset length; and determine whether the candidate first position barcode corresponding to the minimum edit distance is unique based on the mark bit.
[0144] Optionally, the comparison module 15 is also used to calculate the spatial distance between each candidate first position barcode and the standard point in different coordinate axis directions based on the spatial position information of the candidate first position barcode corresponding to other editing distances and the spatial position information of the standard point. When the spatial distances in different coordinate axis directions are all less than the target threshold, the comparison is deemed to be successful, and the second position barcode is determined to be the second position barcode that has been successfully compared.
[0145] Optionally, the matching module 13 is further configured to extract, for each second position barcode, a candidate short sequence of length K by sliding the second position barcode in sequence using a sliding window; wherein the interval between two adjacent candidate short sequences is W; compare each candidate short sequence with the index information library to determine whether there is a hit reference short sequence that is the same as the current candidate short sequence; if so, determine, based on the index information of the hit reference short sequence, the hit first position barcode corresponding to the current candidate short sequence, and the position deviation between the current candidate short sequence and the hit reference short sequence in their respective position barcodes; determine, based on the position deviation satisfying a preset range, that the current candidate short sequence is a matching short sequence; convert the serial number of the hit first position barcode corresponding to the matching short sequence and the number of repetitions of the hit reference short sequence into a hash value, establish directory information corresponding to the matching short sequence, and form a matching information library of the matching short sequence and the first position barcode based on the directory information.
[0146] Optionally, the matching module 13 is further configured to, if not, discard the current short sequence to be selected; or, if so, determine that the current short sequence to be selected is a false match based on the position deviation exceeding the preset range, and discard the current short sequence to be selected.
[0147] Optionally, the index establishment module 12 is further used to extract a reference short sequence of length K for each first position barcode by sliding a window in sequence on the first position barcode; wherein the number of bits between two adjacent reference short sequences is W; convert the serial number and position of the first position barcode where each reference short sequence is located into a hash value, establish index information corresponding to each reference short sequence, and form an index information library of the first position barcode based on the index information.
[0148] Optionally, the index establishment module 12 is also used to extract a reference short sequence of length K at intervals by sliding a sliding window on each first position barcode. During the sliding process of the sliding window on the first position barcode, the endpoint digits of length O are skipped according to the endpoint position of the first position barcode.
[0149] Optionally, the index establishing module 12 is further configured to store hash values corresponding to a plurality of identical reference short sequences in the same vector to form common index information of the reference short sequences.
[0150] Optionally, the acquisition module 11 is specifically used to obtain a position barcode whitelist, and obtain a first position barcode with spatial position information from the position barcode whitelist; or, obtain a first position barcode with spatial position information by sequencing a spatial group chip.
[0151] It should be noted that the spatial omics sequencing positioning matching device provided in the above embodiment only uses the division of the above-mentioned program modules as an example to illustrate the process of performing position barcode comparison. In actual applications, the above-mentioned processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the method steps described above. In addition, the spatial omics sequencing positioning matching device provided in the above embodiment and the spatial omics sequencing positioning matching method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0152] On the other hand, the present application further provides a spatial omics sequencing device. Refer to Figure 5, which is an optional hardware structure diagram of a spatial omics sequencing device, wherein the spatial omics sequencing device includes a processor 212 and a memory 211 connected to the processor 212, and the memory 211 stores a computer program for implementing the sequencing positioning matching method of the spatial omics provided in any embodiment of the present application, so that when the corresponding computer program is executed by the processor, the step of the sequencing positioning matching method of the spatial omics provided in any embodiment of the present application is implemented. The spatial omics sequencing device loaded with the corresponding computer program has the same technical effect as the corresponding method embodiment, and will not be repeated here to avoid repetition.
[0153] On the other hand, the embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, each process of the above-mentioned sequencing positioning matching method embodiment based on spatial omics is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. Wherein, the computer-readable storage medium is such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0154] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0155] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for a terminal (which can be a mobile phone, computer, server, spatial genomics sequencing platform, gene sequencer, or network equipment, etc.) to execute the methods described in each embodiment of the present invention.
[0156] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A sequencing and localization matching method for spatial omics, characterized in that, Including: Obtaining a first position barcode with spatial position information; For each of the first position barcodes, obtaining a plurality of reference short sequences, establishing index information according to the serial number and position of the first position barcode where each reference short sequence is located, and constructing an index information library of the first position barcode; Obtaining a second position barcode obtained by secondary sequencing; For each of the second position barcodes, obtaining a plurality of candidate short sequences, comparing the candidate short sequences with the index information library to determine matching short sequences, and constructing a matching information library of the matching short sequences and the first position barcode; According to the matching information library, determining the comparison priority of the first position barcode; For each of the second position barcodes, comparing it with the first position barcode according to the comparison priority, determining candidate first position barcodes whose edit distance from the second position barcode meets the target requirements, calculating a spatial distance based on the spatial position information of the candidate first position barcodes, and determining a second position barcode with successful comparison based on the spatial distance meeting the target conditions.
2. The sequencing and positioning matching method according to claim 1, wherein The determining candidate first position barcodes whose edit distance from the second position barcode meets the target requirements, calculating a spatial distance based on the spatial position information of the candidate first position barcodes, and determining a second position barcode with successful comparison based on the spatial distance meeting the target conditions includes: Determining candidate first position barcodes whose edit distance from the second position barcode meets the target requirements; Taking the spatial position information of the candidate first position barcode corresponding to the minimum edit distance as a reference point, calculating the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the reference point, and determining a second position barcode with successful comparison based on the spatial distance meeting the target conditions.
3. The sequencing and positioning matching method according to claim 2, wherein The taking the spatial position information of the candidate first position barcode corresponding to the minimum edit distance as a reference point, calculating the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the reference point, and determining a second position barcode with successful comparison based on the spatial distance meeting the target conditions includes: Judging whether the candidate first position barcode corresponding to the minimum edit distance is unique; If so, taking the spatial position information of the candidate first position barcode corresponding to the minimum edit distance as a reference point, calculating the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the reference point, regarding it as a successful comparison when the spatial distance is less than the target threshold, and determining the second position barcode as the second position barcode with successful comparison; If not, taking the spatial position information of the plurality of candidate first position barcodes corresponding to the minimum edit distance as reference points respectively for iterative calculation. In any iterative calculation, calculating the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the reference point, regarding it as a successful comparison when the spatial distance is less than the target threshold, and determining the second position barcode as the second position barcode with successful comparison.
4. The sequencing and positioning matching method according to claim 3, wherein The judging whether the candidate first position barcode corresponding to the minimum edit distance is unique includes: Based on the comparison between the second position barcode and the first position barcode, the candidate first position barcodes with an edit distance meeting the target requirements are retained in the same vector, and whether there are duplicates among the candidate first position barcodes corresponding to the same edit distance is recorded through the flag bits in a sequence of a preset length; Based on the flag bits, it is determined whether the candidate first position barcode corresponding to the minimum edit distance is unique.
5. The sequencing and positioning matching method according to claim 2, characterized in that Calculating the spatial distance between the spatial position information of the candidate first position barcodes corresponding to other edit distances and the standard point, and determining the successfully compared second position barcodes based on the spatial distance meeting the target conditions includes: According to the spatial position information of the candidate first position barcodes corresponding to other edit distances and the spatial position information of the standard point, calculate the spatial distance between each candidate first position barcode and the standard point in different coordinate axis directions. When the spatial distances in different coordinate axis directions are all less than the target threshold, it is regarded as a successful comparison, and the second position barcode is determined as the successfully compared second position barcode.
6. The sequencing and positioning matching method according to claim 1, wherein, For each of the second position barcodes, obtaining a plurality of candidate short sequences, comparing the candidate short sequences with the index information library to determine the matching short sequences, and constructing a matching information library of the matching short sequences and the first position barcodes, including: For each of the second position barcodes, use a sliding window to sequentially slide on the second position barcode at intervals to extract candidate short sequences of length K; wherein, the interval between two adjacent candidate short sequences is W; Compare each of the candidate short sequences with the index information library to determine whether there is a hit reference short sequence identical to the current candidate short sequence; If so, according to the index information of the hit reference short sequence, determine the hit first position barcode corresponding to the current candidate short sequence, and the position deviation between the current candidate short sequence and the hit reference short sequence in their respective position barcodes. Based on the position deviation meeting the preset range, determine the current candidate short sequence as a matching short sequence. Convert the serial number of the hit first position barcode corresponding to the matching short sequence and the number of repeated occurrences of the hit reference short sequence into a hash value, establish the directory information corresponding to the matching short sequence, and form the matching information library of the matching short sequence and the first position barcode based on the directory information.
7. The sequencing and positioning matching method according to claim 6, wherein It further includes: If not, discard the current candidate short sequence; Or, If so, based on the position deviation exceeding the preset range, determine the current candidate short sequence as a false match and discard the current candidate short sequence.
8. The sequencing and positioning matching method according to claim 1, wherein For each of the first position barcodes, obtaining a plurality of reference short sequences, establishing index information according to the serial number and position of each reference short sequence in the first position barcode, and constructing an index information library of the first position barcodes, including: For each of the first position barcodes, use a sliding window to sequentially slide on the first position barcode at intervals to extract reference short sequences of length K; wherein, the number of interval bits between two adjacent reference short sequences is W; Convert the serial number and position of each reference short sequence in the first position barcode into a hash value, establish index information corresponding to each reference short sequence, and form an index information library of the first position barcode based on the index information.
9. The sequencing and positioning matching method according to claim 8, wherein For each of the first position barcodes, using a sliding window to sequentially and intermittently extract reference short sequences of length K on the first position barcode, includes: For each of the first position barcodes, using a sliding window to sequentially and intermittently extract reference short sequences of length K on the first position barcode. During the sliding process of the sliding window on the first position barcode, skip the end bits of length O according to the end point positions of the first position barcode.
10. The sequencing positioning and matching method according to claim 8, wherein The step of converting the serial number and position of each reference short sequence in the first position barcode into a hash value and establishing index information corresponding to each reference short sequence, further includes: Store the hash values corresponding to multiple identical reference short sequences in the same vector to form shared index information of the reference short sequences.
11. The sequencing and positioning matching method according to claim 1, characterized in that, The step of obtaining the first position barcode with spatial position information, includes: Obtain a whitelist of position barcodes, and obtain the first position barcode with spatial position information from the whitelist of position barcodes; or, Obtain the first position barcode with spatial position information obtained by performing a single sequencing based on a spatial group chip.
12. The sequencing and positioning matching method according to claim 1, wherein The length of the candidate short sequence is equal to the length of the reference short sequence.
13. A sequencing and positioning matching device for spatial omics, characterized in that, Includes: An acquisition module, configured to acquire the first position barcode with spatial position information; An index establishment module, configured to, for each of the first position barcodes, acquire a plurality of reference short sequences, establish index information according to the serial number and position of each reference short sequence in the first position barcode, and construct an index information library of the first position barcode; The acquisition module is further configured to acquire the second position barcode obtained by secondary sequencing; A matching module, configured to, for each of the second position barcodes, acquire a plurality of candidate short sequences, compare the candidate short sequences with the index information library to determine matching short sequences, and construct a matching information library of the matching short sequences and the first position barcode; A priority module, configured to determine the comparison priority of the first position barcode according to the matching information library; A comparison module, configured to, for each of the second position barcodes, compare it with the first position barcode according to the comparison priority, determine candidate first position barcodes whose edit distance from the second position barcode meets the target requirements, calculate the spatial distance based on the spatial position information of the candidate first position barcodes, and determine the second position barcodes with successful comparison based on the spatial distance meeting the target conditions.
14. A spatial omics sequencing device, characterized in that, Includes a processor and a memory connected to the processor. A computer program executable by the processor is stored on the memory. When the computer program is executed by the processor, it implements the spatial omics sequencing positioning and matching method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the spatial omics sequencing positioning and matching method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Method and device for analyzing sequencing data of space transcriptome chip
CN115331733A
Spatial barcoding
CN115461469A
Simplified sequencing method of space transcriptome and application thereof
CN116287160A
Method and system for barcode error correction
CN116529827A
Sequencing positioning matching method and device, space omics sequencing equipment and medium
CN117854594A