Method and device for querying mRNA (messenger ribonucleic acid) sequence based on distributed SQL (structured query language) query engine
By using specified query statements and Spark SQL in the distributed SQL query engine, the problem of inefficient mRNA sequence query efficiency in the existing technology is solved, and efficient and accurate gene sequencing data mutation analysis is achieved.
Patent Information
- Application Number
- CN202311714523.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing distributed SQL query engine is inefficient in mRNA sequence query and cannot meet the explosive growth of gene sequencing data.
The specified distributed SQL query statement and Spark SQL query statement are used to query the RefGene database through the specified distributed SQL query engine statement to obtain the gene ID, and the mRNA sequence is queried using Spark SQL, and efficient query is carried out in combination with bicombination conditions.
It realizes efficient and accurate query of mRNA sequences, and improves the efficiency of variation analysis of gene sequencing data.
Smart Images

Figure CN120148657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed SQL query engines, and particularly to a method and device for querying mRNA sequences based on a distributed SQL query engine. Background Art
[0002] Gene sequencing refers to the analysis of blood, body fluids or cells by a sequencing instrument to obtain the base sequence of deoxyribonucleic acid (i.e., DNA). The mRNA (messenger RNA) sequence is a single-stranded ribonucleic acid that is transcribed from one strand of DNA as a template and carries genetic information to guide protein synthesis.
[0003] With the rapid decline in cost, gene sequencing is gradually moving towards clinical applications, the sequencing data shows an explosive growth, and the data that needs to be analyzed for mutations has also increased sharply. However, for the existing gene data analysis based on databases such as RefGene, the distributed SQL query engine is limited by the algorithm efficiency of interval queries in these two databases, resulting in a very low query efficiency for mRNA sequences. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method and device for querying mRNA sequences based on a distributed SQL query engine, which can perform mutation analysis efficiently and accurately.
[0005] Based on the above purpose, a method for querying mRNA sequences based on a distributed SQL query engine provided by the present invention includes:
[0006] Query the RefGene database using the specified distributed SQL query engine statement and return the ID that uniquely identifies the gene. The specified Spark SQL statement means using the query statement "select * from s rgjoin r on goverlap((s.txStart, s.txEnd, s.exonCount, s.exonStarts, s.exonEnds, s.chr, s.strand), (r.start, r.end, r.chr))". In this query statement, s represents the RefGene database in table form, and r represents the variant to be annotated in table form. Use a tuple as the condition for "on". Each parameter in the tuple is represented as follows: s.txStart represents the start field of the variant in table s, s.txEnd represents the end field of the variant in table s, s.exonCount represents the exon count field in table s, s.exonStarts represents the set of start points of each exon in table s, s.exonEnds represents the set of end points of each exon in table s, s.chr represents the chromosome number field in table s, s.strand represents the gene direction (i.e., positive strand and negative strand) field in table s; r.start represents the start field of the variant in table r, r.end represents the end field of the variant in table r, and r.chr represents the chromosome number field in table r.
[0007] According to the returned gene ID, query the mRNA sequence and return the query result. For the query of the mRNA sequence, use the standard query statement of SparkSQL.
[0008] An embodiment of the present invention also provides a query device for mRNA sequences based on a distributed SQL query engine. The annotation device includes: a central processing unit, which can perform various appropriate actions and processes according to the data and programs stored in the memory. Through the bus, the central processing unit, the memory, the input / output part, the external storage part, and the network part are interconnected. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 FLOWCHART
[0010] Figure 2 SCHEMATIC DIAGRAM OF THE DEVICE DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0012] REFER TO Figure 1 , which is the flowchart of the embodiment of the present invention.
[0013] The gene analysis and annotation method includes the following steps:
[0014] Step 101: Query the RefGene database using a specified distributed SQL query engine statement and return the ID that uniquely identifies a gene. The specified distributed SQL query engine statement means using the query statement "select * from s rgjoin r on goverlap((s.txStart, s.txEnd, s.exonCount, s.exonStarts, s.exonEnds, s.chr, s.strand), (r.start, r.end, r.chr))". In this query statement, s represents the RefGene database in table form, and r represents the variant to be annotated in table form. Use a binary tuple as the condition for on, and the parameters in the binary tuple are represented as follows: s.txStart represents the start field of the variant in table s, s.txEnd represents the end field of the variant in table s, s.exonCount represents the exon count field in table s, s.exonStarts represents the set of starts of each exon in table s, s.exonEnds represents the set of ends of each exon in table s, s.chr represents the chromosome number field in table s, and s.strand represents the gene direction (i.e., positive strand and negative strand) field in table s; r.start represents the start field of the variant in table r, r.end represents the end field of the variant in table r, and r.chr represents the chromosome number field in table r.
[0015] Step 102: Query the mRNA sequence based on the returned gene ID and return the query result. For the query of the mRNA sequence, use the standard query statement of Spark SQL, such as "select * from mrnaseq where ID = 'NM_002714'". The conditional parameter includes the ID that uniquely identifies a gene in RefGene for the SNP.
[0016] Reference Figure 2 , which shows a schematic structural diagram of a computer system of a terminal device / server suitable for implementing the embodiments of the present application. Figure 2 The shown terminal device / server is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0017] As Figure 2 shown, the device includes a central processing unit 201, which can perform various appropriate actions and processes according to the data and programs stored in the memory 202. Through the bus 203, interconnections are achieved among the central processing unit 201, the memory 202, the input / output section 204, the external storage section 205, and the network section 206.
[0018] The device of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.
[0019] Those of ordinary skill in the art should understand that: the discussion of any above embodiment is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity.
[0020] Embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An mRNA sequence query method based on a distributed SQL query engine, characterized in that it includes: Query the RefGene database using a specified distributed SQL query engine statement and return the ID that uniquely identifies the gene; The specified distributed SQL query engine statement refers to using the query statement select*from s rgjoin r ongoverlap((s.txStart, s.txEnd, s.exonCount, s.exonStarts, s.exonEnds, s.chr, s.strand), (r.start, r.end, r.chr)). In this query statement, s represents the RefGene database in table form, and r represents the variant to be annotated in table form. Use a binary tuple as the condition for on, and the parameters in the binary tuple are represented as follows: s.txStart represents the start field of the variant in table s, s.txEnd represents the end field of the variant in table s, s.exonCount represents the exon count field in table s, s.exonStarts represents the set of start points of each exon in table s, s.exonEnds represents the set of end points of each exon in table s, s.chr represents the chromosome number field in table s, s.strand represents the gene direction (i.e., positive and negative strands) field in table s; r.start represents the start field of the variant in table r, r.end represents the end field of the variant in table r, and r.chr represents the chromosome number field in table r; According to the returned gene ID, query the mRNA sequence and return the query result; For the query of the mRNA sequence, use the standard query statement of Spark SQL.
2. An annotation device for mRNA sequences based on a distributed SQL query engine, characterized in that it includes: Through the bus, the central processing unit, the memory, the input / output part, the external storage part, and the network part are interconnected. The central processing unit can perform various appropriate actions and processes according to the data and programs stored in the memory.