Method, system and equipment for quickly searching matched gene mutation sites
By creating indexes in the gnomAD database and using parallel processing technology, the problem of inefficient retrieval of genomic data is solved, and the rapid search of matching gene mutation sites is achieved, which significantly improves the search speed and efficiency and meets clinical needs.
Patent Information
- Application Number
- CN202510078342.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-27
AI Technical Summary
When processing huge and complex genomic data, the search efficiency is inefficient and data integration is difficult, and it cannot meet the clinical needs of rapid diagnosis and timely treatment.
By obtaining the vcf file of the gnomAD database, creating an index, and using a thread pool to perform tasks in parallel, extracting the information of specified gene loci, collecting results, merging data, and outputting a result list, it is possible to quickly find matching gene mutation sites.
It significantly improves the speed and efficiency of gene mutation site retrieval, and can find records corresponding to specific chromosomes and location ranges in milliseconds, meeting the clinical needs of rapid diagnosis and timely treatment.
Smart Images

Figure CN120220801A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics and genomics technology, and in particular to a method, system and device for quickly searching for matching gene mutation sites. Background Art
[0002] The Genome Aggregation Database (gnomAD) is a valuable resource carefully created by an international consortium of researchers. It brings together exome and genome sequencing data from many large-scale sequencing projects. The main function of the GnomAD database is to record the mutation frequencies of different ethnic groups at all detected mutation sites, providing researchers with a comprehensive and authoritative reference for gene mutation frequencies. In practical applications, the analysis process of many in vitro diagnostic genetic test results is inseparable from the filtering and screening of sites with mutation frequencies in the GnomAD database. Through this process, those mutation sites that appear frequently in the general population and may be harmless can be effectively excluded, thereby focusing attention on those relatively rare mutation sites that may be pathogenic. This not only helps to improve the accuracy of disease diagnosis, but also provides a more accurate basis for subsequent treatment decisions.
[0003] However, regarding the retrieval of the GnomAD database, the existing technology still has some obvious shortcomings, which are mainly reflected in the following aspects:
[0004] 1. The GnomAD database has a huge amount of data, resulting in low retrieval efficiency
[0005] The GnomAD database contains extremely rich gene mutation sites and data information, and the file size has basically reached an astonishing 30T. This has directly led to a significant reduction in the efficiency of the genome mutation site screening process and is unable to meet the clinical needs of rapid diagnosis and timely treatment.
[0006] 2. High data retrieval complexity
[0007] Due to the diversity and complexity of gene mutation sites and the differences in data formats of different sequencing projects, searching for specific gene mutation sites in the GnomAD database becomes extremely complicated, further increasing the difficulty and workload of data retrieval.
[0008] 3. Data integration and analysis are difficult
[0009] In the actual process of screening gene mutation sites, it is often necessary to integrate and analyze the data in the GnomAD database with other bioinformatics databases or clinical data. However, differences in data formats, standards, and interfaces between different databases make data integration extremely difficult.
[0010] Therefore, the current related technologies and application research on finding matching gene mutation sites still need to be further improved. Summary of the Invention
[0011] In view of this, the present invention proposes a method, system and device for quickly finding matching gene mutation sites, which solves the technical problems of low retrieval efficiency and difficult data integration when facing huge and complex genomic data in the prior art.
[0012] The technical solution of the present invention is implemented as follows:
[0013] The present invention provides a method for quickly finding matching gene mutation sites, including the following steps:
[0014] S1. Obtain the vcf file of the gnomAD database and create an index for the vcf file;
[0015] S2. Create a thread pool, execute multiple tasks in parallel, extract the specified gene locus information in the gnomAD database, collect the execution results of each task, add records of matching information to the result list, and output the debugging information of the unmatched records;
[0016] S3. Add entries, convert the result list into a data frame form, merge it with the vcf file in step S1, save the merged data to a specified output file and output it in a fixed format, and output a prompt message indicating that the search is completed.
[0017] On the basis of this technical solution, further preferably, creating an index for the vcf file further includes creating an index stored in tbi format and storing it in the same folder as the vcf file.
[0018] On the basis of this technical solution, further preferably, the specified gene locus information includes chromosome, chromosome start, end position and RS ID.
[0019] On the basis of this technical solution, further preferably, the range of the index is the region of the chromosome start, end position and their upstream and downstream ±50bp.
[0020] On the basis of this technical solution, further preferably, the specific steps for extracting the specified gene locus information in the gnomAD database are:
[0021] According to the extracted chromosome and RS ID, construct the corresponding gnomAD file path and query parameters, open the indexed gnomAD file, and retrieve the matching gene mutation site information by the index positioning method.
[0022] Based on this technical solution, further preferably, the added entry is the entry with RS ID = -1.
[0023] Based on this technical solution, further preferably, the record of the matching information is added to the result list, and the debugging information of the unmatched records output includes:
[0024] Perform RS ID mutation matching. If a gene mutation is matched, extract the gnomAD record of the gene mutation information;
[0025] If no gene mutation is matched, record it as -1, then merge the matched information with the vcf file in step S1, record the unmatched variant information, and output the debugging information of the unmatched records.
[0026] Based on this technical solution, further preferably, the fixed format is the csv format.
[0027] In a second aspect, the present invention provides a system for quickly searching and matching gene mutation sites, including:
[0028] A preprocessing and indexing module, configured to obtain the vcf file of the gnomAD database and create an index for the vcf file;
[0029] A parallel retrieval and matching module, configured to create a thread pool, execute multiple tasks in parallel, extract the specified gene locus information in gnomAD, collect the execution results of each task, add the record of the matching information to the result list, and output the debugging information of the unmatched records;
[0030] A data integration and output module, adding entries, converting the result list into the form of a data frame, merging it with the vcf file in step S1, saving the merged data to a specified output file and outputting it in a fixed format, and outputting a prompt message indicating that the search is completed.
[0031] Based on this technical solution, further preferably, the parallel retrieval and matching module includes a parallel retrieval module and a result matching module.
[0032] In a third aspect, the present invention further provides an electronic device, including a memory and a processor, where the memory has a computer program, and when the processor executes the program, it implements the method for quickly searching and matching gene mutation sites according to any one of the first aspect.
[0033] The method, system and device for quickly searching and matching gene mutation sites according to the present invention have the following beneficial effects compared with the prior art:
[0034] A method for quickly finding matching gene mutation sites, which is based on an efficient data indexing mechanism and parallel processing technology. First, parallel retrieval is achieved by introducing a thread pool, breaking through the limitations of traditional single-threaded or serial processing. The retrieval tasks are distributed to multiple threads for simultaneous execution, making full use of the computing resources of multi-core CPUs and significantly improving the speed and efficiency of gene mutation site retrieval.
[0035] The present invention creates an index on the dataset of input data based on chromosome numbers and start and end positions (Start, End) and their upstream and downstream ±50bp regions. Through this index optimization, rapid positioning and access to gene mutation site data are realized, avoiding traversing the entire dataset. In practical applications, even when facing ClinVar data containing millions of records, records corresponding to specific chromosomes and position ranges can be found within milliseconds, significantly improving the efficiency of gene mutation site screening.
[0036] Finally, the present invention designs a flexible data integration method, which can effectively integrate the extracted gnomAD information with the original input data, and realizes data merging through the pandas.merge function. At the same time, it supports saving the results to a CSV file. It provides comprehensive and accurate data support for subsequent bioinformatics analysis and clinical applications, expands the application scope of gene mutation site screening results, improves the readability and usability of data, and enables researchers to more conveniently obtain and use relevant information on gene mutation sites. Brief Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a flowchart of a method for quickly finding matching gene mutation sites according to Embodiment 1 of the present invention;
[0039] Figure 2 It is a schematic diagram of a method for quickly finding matching gene mutation sites according to Embodiment 2 of the present invention;
[0040] Figure 3 It is a system framework diagram of a method for quickly finding matching gene mutation sites according to Embodiment 3 of the present invention. Detailed Embodiments
[0041] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] In the prior art, after a person's gene sequencing is completed, millions of mutations may be detected, and during the filtering process, it is necessary to know the mutation frequencies of these variant sites. GnomAD records the mutation frequencies of genomes of different ethnic groups. Therefore, it is necessary to find the mutation frequencies of millions of sites according to the chromosome number and genomic position in a short time. However, the sites recorded in the gnomAD database are extremely large. Although this database is stored in units of chromosomes, the mutation information recorded on each chromosome is also very large. Therefore, the process of opening each chromosome file and searching for mutations is very time-consuming, especially when searching for mutations in the hundreds, the time and computer resources consumed are extremely large.
[0043] The present invention provides a method, system and device for quickly searching and matching gene mutation sites. First, the genomic data is preprocessed to construct a fine index structure to achieve rapid positioning of gene mutation sites; then, a multi-threaded parallel retrieval strategy is used to simultaneously search for matching gene mutation site information from multiple data sources, greatly improving the speed and accuracy of data retrieval.
[0044] Example 1
[0045] As Figure 1 shown, a method for quickly searching and matching gene mutation sites includes the following steps:
[0046] S1. Obtain the vcf file of the gnomAD database and create an index for the vcf file;
[0047] S2. Create a thread pool, execute multiple tasks in parallel, extract the specified gene site information in the gnomAD database, collect the execution results of each task, add records of matching information to the result list, and output the debugging information of the unmatched records;
[0048] S3. Add entries, convert the result list into the form of a data frame, merge it with the vcf file in step S1, save the merged data to a specified output file and output it in a fixed format, and output a prompt message indicating that the search is completed.
[0049] In this embodiment, specifically, creating an index for the vcf file further includes creating an index stored in tbi format and storing it in the same folder as the vcf file.
[0050] In this embodiment, specifically, the specified gene locus information includes chromosome, chromosome start, end position, and RS ID.
[0051] In this embodiment, specifically, the range of the index is the region of ±50bp upstream and downstream of the chromosome start, end position.
[0052] In this embodiment, specifically, the steps for extracting the specified gene locus information in the gnomAD database are as follows:
[0053] According to the extracted chromosome and RS ID, construct the corresponding file path and query parameters of gnomAD, open the indexed gnomAD file, and retrieve the matching gene mutation site information through the index positioning method.
[0054] In this embodiment, specifically, the added entry is the entry with RS ID = -1.
[0055] In this embodiment, specifically, record the matching information in the result list, and the debugging information output for the unmatched records includes:
[0056] Perform RS ID mutation matching. If a gene mutation is matched, extract the gnomAD record of the gene mutation information;
[0057] If no gene mutation is matched, record it as -1, then merge the matched information with the vcf file in step S1, record the unmatched variant information, and output the debugging information of the unmatched records.
[0058] In this embodiment, specifically, the fixed format is csv format.
[0059] For example, when processing a large-scale ClinVar dataset, it can achieve an effect close to linear acceleration, greatly reducing the total data retrieval time and meeting the clinical needs of rapid diagnosis and timely treatment.
[0060] Example 2
[0061] As Figure 2 shown, a method for quickly finding and matching gene mutation sites includes the following steps:
[0062] S1. Obtain the vcf file of the gnomAD database and create an index for the vcf file;
[0063] Specifically, use the read_csv function in the pandas library of Python to read files, such as ClinVar, and read the files into a data frame. To optimize memory usage and improve data loading efficiency, set the low_memory parameter to False during reading. After completion, the data frame contains all the information of ClinVar, facilitating subsequent analysis and processing.
[0064] Specifically, create an index. Create an index on the input data, such as the DataFrame of ClinVar. The index is based on the chromosome and the start and end positions (Start, End) and the regions ±50bp upstream and downstream of them. By creating this index based on the position range, specific gene mutation site data can be quickly located and accessed, laying a foundation for subsequent efficient retrieval and matching. This indexing method can effectively cover the potential impact regions of gene mutation sites and improve the accuracy and efficiency of retrieval.
[0065] S2. Create a thread pool to execute multiple tasks in parallel, extract the specified gene locus information from the gnomAD database, collect the execution results of each task, add records of matching information to the result list, and output the debugging information of the unmatched records;
[0066] Specifically, extract the position information of the input gene mutation site, the path of the gnomAD folder, and the required number of cores. Among them, the gnomAD folder includes the vcf file and tbi file described in step S1, and the number of cores is the number of CPUs to be called, with a default value of 4.
[0067] Extract the four main pieces of information from the file recording the position information of the mutations to be searched: chromosome number, RSID, mutation start and end positions, and check whether this necessary information exists in the input file. Create the corresponding number of worker threads according to the input number of cores, and use the thread pool to execute the task of extracting gnomAD information; for each mutation site record, call the built-in function extract_gnomad_info to extract mutation data from the corresponding gnomAD file using the position index and perform mutation matching for the RS ID.
[0068] If a matching mutation is found, extract all the information related to the gnomAD record mutation. If there is no matching mutation in the gnomAD database, it will be recorded as -1.
[0069] Merge the matching information with the original input file data; collect and process the results returned by each thread, record the variant information that cannot be matched, and output the debugging information for subsequent tracking.
[0070] S3. Add entries, convert the result list into a data frame, merge it with the vcf file in step S1, save the merged data to a specified output file in a fixed format, and output a prompt message indicating that the search is completed.
[0071] Specifically, add entries with RS=-1: Add records with an RS ID of -1 in the input data to the result list. These records have no gnomAD information. Through this step, ensure that all original input data can be reflected in the final result.
[0072] Then, data merging: Convert the result list into a data frame and merge it with the original input data. The merging is based on the chromosome and RS# (dbSNP) columns. Through merging, the extracted gnomAD information can be associated with the input data to form a complete data set.
[0073] Save to CSV: Save the merged data to a specified output file, output it in CSV format using the to_csv method, and set the index=False parameter to not save the row index. Output a prompt message indicating that the extraction is completed.
[0074] Example 3
[0075] As Figure 3 shown, a system for quickly searching for matching gene mutation sites includes:
[0076] A preprocessing and indexing module for obtaining the vcf file of the gnomAD database and creating an index for the vcf file.
[0077] Specifically, after creating an index based on the position range, this module can quickly locate specific gene mutation sites and their potential impact regions in the input data (such as ClinVar), without traversing the entire data set. For example, when processing ClinVar data containing millions of records, the records corresponding to a specific chromosome and position range can be found within milliseconds through the index, significantly improving the speed of data retrieval.
[0078] A parallel retrieval and matching module for creating a thread pool, executing multiple tasks in parallel, extracting the specified gene locus information in gnomAD, collecting the execution results of each task, adding records with matching information to the result list, and outputting the debugging information of the unmatched records.
[0079] Through parallel processing technology, the module distributes the task of retrieving gene mutation sites that originally needed to be executed sequentially to multiple threads for simultaneous execution, significantly reducing the total time for data retrieval. For example, when processing large-scale ClinVar data, compared with the traditional single-threaded retrieval method, detecting 10 sites basically takes 40 minutes. The present invention can increase the retrieval efficiency by hundreds of times, and basically can retrieve the information of 10 sites within 1 second. And when using ClinVar data as the test input data, more than 3 million variant sites can be locked in 17 hours, thus meeting the clinical needs of rapid diagnosis and timely treatment.
[0080] The data integration and output module adds entries, converts the result list into a data frame form, merges it with the vcf file in step S1, saves the merged data to a specified output file for output in a fixed format, and outputs a prompt message indicating that the search is completed.
[0081] In a preferred embodiment, specifically, the parallel retrieval and matching module includes a parallel retrieval module and a result matching module.
[0082] In a preferred embodiment, an electronic device includes a memory and a processor. The memory has a computer program, and when the processor executes the program, it implements the method for quickly searching and matching gene mutation sites described in Embodiment 1.
[0083] In summary, the present invention provides a method, system and device for quickly searching and matching gene mutation sites. First, the genomic data is preprocessed to construct a fine index structure to achieve rapid positioning of gene mutation sites; then, a multi-threaded parallel retrieval strategy is used to simultaneously search for matching gene mutation site information from multiple data sources, greatly improving the speed and accuracy of data retrieval.
[0084] In addition, the present invention also includes a system for quickly searching and matching gene mutation sites. Through a flexible data integration module, it can effectively integrate gene mutation data in different formats and sources, providing a solid data foundation for subsequent bioinformatics analysis and clinical applications. Compared with the prior art, the present invention has achieved a qualitative leap in the efficiency of gene mutation site screening, significantly reducing the time cost of data processing, while improving the flexibility and accuracy of data integration, providing strong technical support for the development of genomics research and precision medicine.
[0085] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for quickly finding matching gene mutation sites, characterized in that: The steps include: S1. Obtain a vcf file of the gnomAD database, and create an index for the vcf file; S2, create a thread pool, execute multiple tasks in parallel, extract the gene locus information specified in the gnomAD database, collect the execution results of each task, add the records of matching information to the result list, and output the debugging information of the unmatched records; S3, add entries, convert the result list into a data frame, merge it with the vcf file in step S1, save the merged data to the specified output file and output it in a fixed format, and output a prompt message indicating that the search is complete.
2. The method for quickly searching for matching gene mutation sites according to claim 1, characterized in that: Creating an index for the vcf file also includes storing the created index in a tbi format and storing it in the same folder as the vcf file.
3. The method for quickly searching for matching gene mutation sites according to claim 1, characterized in that: The specified locus information includes chromosome, chromosome start, end position and RS ID.
4. The method for quickly searching for matching gene mutation sites according to claim 3, characterized in that: The index ranges from the chromosome start and end positions and the region ±50 bp upstream and downstream thereof.
5. The method for quickly searching for matching gene mutation sites according to claim 4, characterized in that: Extract the gene locus information specified in the gnomAD database. The specific steps are: According to the extracted chromosome and RS ID, the corresponding gnomAD file path and query parameters are constructed, the indexed gnomAD file is opened, and the matching gene mutation site information is retrieved through the index positioning method.
6. The method for quickly searching for matching gene mutation sites according to claim 4, characterized in that: Adding an entry is adding an entry with RS ID=-1.
7. The method for quickly searching for matching gene mutation sites according to claim 6, characterized in that: Add matching records to the result list, and output debugging information for unmatched records including: Perform RS ID mutation matching. If a gene mutation is matched, extract the gene mutation information recorded in gnomAD. If no gene mutation is matched, it is recorded as -1, and then the matched information is merged with the vcf file of step S1, the unmatched mutation information is recorded, and the debugging information of the unmatched record is output.
8. A system for quickly finding matching gene mutation sites, characterized in that: include: A preprocessing and indexing module, used to obtain the vcf file of the gnomAD database and create an index for the vcf file; The parallel search and matching module is used to create a thread pool, execute multiple tasks in parallel, extract the gene locus information specified in the gnomAD database, collect the execution results of each task, add the matching information records to the result list, and output the debugging information of the unmatched records; The data integration output module adds entries, converts the result list into a data frame, merges it with the vcf file of step S1, saves the merged data to the specified output file and outputs it in a fixed format, and outputs a prompt message indicating that the search is complete.
9. The system for quickly searching for matching gene mutation sites according to claim 8, characterized in that: The parallel search and matching module includes a parallel search module and a result matching module.
10. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory has a computer program, and when the processor executes the program, the method for quickly searching for matching gene mutation sites according to any one of claims 1 to 7 is implemented.