Query method and device of MiRBase gene pool based on distributed SQL (Structured Query Language) query engine

By querying the RefGene database using the specified Spark SQL statement in the distributed SQL query engine to obtain the gene ID, and using this ID to query the MiRBase gene library, the problem of low efficiency of MiRBase query in the existing technology is solved, and efficient and accurate mutation analysis is achieved.

CN120148656APending Publication Date: 2025-06-13XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311714329.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing genetic data analysis based on RefGene, ResSeq and other databases is limited by the algorithm efficiency of the distributed SQL query engine for interval query, resulting in the low query efficiency of MiRBase.

Method used

A query method of MiRBase gene library based on distributed SQL query engine is used to query the RefGene database through the specified Spark SQL statement to obtain the gene ID, and the MiRBase gene library is used to query the standard query statement of the distributed SQL query engine.

Benefits of technology

It realizes efficient and accurate query of the MiRBase gene library, and improves the efficiency of variant analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148656A_ABST
    Figure CN120148656A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a query method of a MiRBase gene pool based on a distributed SQL query engine. The method comprises the following steps: querying a RefGene database by using a specified distributed SQL query engine statement, and returning an ID (Identity) of a unique identification gene; and querying the MiRBase according to the returned gene ID, and returning a query result. In addition, the embodiment of the invention provides a query device of the MiRBase gene pool based on the distributed SQL query engine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed SQL query engines, and particularly to a query method and device for the MiRBase gene library based on a distributed SQL query engine. Background Art

[0002] Gene sequencing refers to the analysis of blood, body fluids or cells by a sequencing instrument to measure the base sequence of deoxyribonucleic acid (i.e., DNA). The miRBase gene library provides a comprehensive database including miRNA sequence data, annotations, predicted gene targets, etc., and is one of the most important public databases for storing miRNA information.

[0003] With the rapid decline in cost, gene sequencing is gradually moving towards clinical applications, the sequencing data shows an explosive growth, and the data that needs to be analyzed for mutations has also increased sharply. However, the existing gene data analysis based on databases such as RefGene and ResSeq is limited by the algorithm efficiency of the interval query of these two databases by the distributed SQL query engine, resulting in a very low query efficiency of MiRBase. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a query method and device for the MiRBase gene library based on a distributed SQL query engine, which can perform mutation analysis efficiently and accurately.

[0005] Based on the above purpose, a query method for the MiRBase gene library based on a distributed SQL query engine provided by the present invention includes:

[0006] Query the RefGene database using the specified distributed SQL query statement and return the IDs that uniquely identify genes. The specified distributed SQL query statement means using the query statement "select * from s rgjoin r on goverlap((s.txStart, s.txEnd, s.exonCount, s.exonStarts, s.exonEnds, s.chr, s.strand), (r.start, r.end, r.chr))". In this query statement, s represents the RefGene database in table form, and r represents the variant to be annotated in table form. Use a tuple as the condition for "on", and the parameters in the tuple are represented as follows: s.txStart represents the start field of the variant in table s, s.txEnd represents the end field of the variant in table s, s.exonCount represents the exon count field in table s, s.exonStarts represents the set of start points of each exon in table s, s.exonEnds represents the set of end points of each exon in table s, s.chr represents the chromosome number field in table s, and s.strand represents the gene direction (i.e., positive strand and negative strand) field in table s; r.start represents the start field of the variant in table r, r.end represents the end field of the variant in table r, and r.chr represents the chromosome number field in table r.

[0007] According to the returned gene IDs, query the miRBase gene library and return the query results. The query of the miRBase gene library uses the standard query statement of the distributed SQL query engine.

[0008] An embodiment of the present invention also provides a query device for the miRBase gene library based on a distributed SQL query engine. The annotation device includes: a central processing unit, which can perform various appropriate actions and processes according to the data and programs stored in the memory. Through the bus, the central processing unit, the memory, the input / output part, the external storage part, and the network part are interconnected. Description of the Drawings

[0009] Figure 1 Flowchart

[0010] Figure 2 Schematic Diagram of the Device Detailed Embodiments

[0011] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0012] Refer to Figure 1 , which is the flowchart of an embodiment of the present invention.

[0013] The gene analysis and annotation method includes the following steps:

[0014] Step 101: Query the RefGene database using a specified Spark SQL statement and return the ID that uniquely identifies a gene. The specified Spark SQL statement refers to using the query statement "select * from s rgjoin r on goverlap((s.txStart, s.txEnd, s.exonCount, s.exonStarts, s.exonEnds, s.chr, s.strand), (r.start, r.end, r.chr))". In this query statement, s represents the RefGene database in table form, and r represents the variant to be annotated in table form. A binary tuple is used as the condition for "on", and the parameters in the binary tuple are as follows: s.txStart represents the start field of the variant in table s, s.txEnd represents the end field of the variant in table s, s.exonCount represents the exon count field in table s, s.exonStarts represents the set of start points of each exon in table s, s.exonEnds represents the set of end points of each exon in table s, s.chr represents the chromosome number field in table s, and s.strand represents the gene direction (i.e., positive strand and negative strand) field in table s; r.start represents the start field of the variant in table r, r.end represents the end field of the variant in table r, and r.chr represents the chromosome number field in table r.

[0015] Step 102: Query the MiRBase gene library based on the returned gene ID and return the query result. For the query of the MiRBase gene library, use the standard query statement of Spark SQL, such as "select * from mirbase where ID = 'NM_147191'". The conditional parameter includes the ID that uniquely identifies the gene of the SNP in RefGene.

[0016] Reference Figure 2 , which shows a schematic structural diagram of a computer system of a terminal device / server suitable for implementing the embodiments of the present application. Figure 2 The shown terminal device / server is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.

[0017] As Figure 2 shown, the device includes a central processing unit 201, which can perform various appropriate actions and processes according to the data and programs stored in the memory 202. Through the bus 203, connections are achieved among the central processing unit 201, the memory 202, the input / output part 204, the external storage part 205, and the network part 206.

[0018] The device of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated herein.

[0019] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity.

[0020] Embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A query method for the miRBase gene bank based on a distributed SQL query engine, characterized in that it includes: Query the RefGene database using a specified distributed SQL query statement and return the ID that uniquely identifies the gene; The specified distributed SQL query statement refers to using the query statement select * from s rgjoin r ongoverlap((s.txStart, s.txEnd, s.exonCount, s.exonStarts, s.exonEnds, s.chr, s.strand), (r.start, r.end, r.chr)). In this query statement, s represents the RefGene database in table form, and r represents the variant to be annotated in table form. A binary tuple is used as the condition for on, and each parameter in the binary tuple is represented as follows: s.txStart represents the start field of the variant in table s, s.txEnd represents the end field of the variant in table s, s.exonCount represents the exon count field in table s, s.exonStarts represents the set of start points of each exon in table s, s.exonEnds represents the set of end points of each exon in table s, s.chr represents the chromosome number field in table s, s.strand represents the gene direction (i.e., positive strand and negative strand) field in table s; r.start represents the start field of the variant in table r, r.end represents the end field of the variant in table r, r.chr represents the chromosome number field in table r; According to the returned gene ID, query the miRBase gene bank and return the query result; For the query of the miRBase gene bank, use the standard query statement of the distributed SQL query engine.

2. An annotation device for miRBase based on a distributed SQL query engine, characterized in that it includes: Through the bus, the central processing unit, the memory, the input / output part, the external storage part, and the network part are interconnected. The central processing unit can perform various appropriate actions and processes according to the data and programs stored in the memory.