A method and system for storage and visual retrieval based on family genomic data
By constructing a difference database and embedding model, generating a joint index and using cosine similarity calculation, combined with a sliding window caching strategy, the problem of low storage and retrieval efficiency of family genome data is solved, and efficient multi-dimensional data display and analysis is achieved.
Patent Information
- Application Number
- CN202411842257.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-12-13
Smart Images

Figure CN119763659B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of molecular biology, in particular to a storage and visual retrieval method and system based on family genomic data. BACKGROUND
[0002] The analysis of family genomic data has important value in modern biomedical, genetics and precision medicine fields, especially in studying the genetic factors of complex diseases. Family data not only contains rich genetic information, but also contains multi-modal information such as family relationships, intergenerational transmission of genetic variations, and environmental influences. By systematically storing and effectively retrieving these data, researchers can identify patterns of genetic mutations between generations, predict the risk of genetic diseases, and analyze hidden associations in genomic data. However, in the face of the diversity, scale and high dimensionality of data, traditional storage and retrieval techniques face problems such as poor retrieval effect and unsmooth data visualization.
[0003] In existing genomic data storage, file storage is usually used to save data. These methods have good support for single genomic data, but their efficiency and flexibility are greatly limited when facing family data. On the one hand, file storage is difficult to achieve ideal results in the efficient storage and complex association of multi-modal and unstructured data; on the other hand, file storage lacks a fast retrieval mechanism and is difficult to meet real-time query requirements. Therefore, the existing storage and retrieval methods for family genomic data have the problem of low storage and retrieval efficiency. SUMMARY
[0004] The purpose of the present application is to provide a storage and visual retrieval method and system based on family genomic data to solve the problem of low storage and retrieval efficiency of existing family genomic data storage and retrieval methods.
[0005] The technical scheme adopted by the present application to solve the above technical problems is:
[0006] A storage and visual retrieval method based on family genomic data, comprising the following steps:
[0007] Step 1: Obtain family genomic data, and construct the chromosome information of each family member M i including gene sequence and base sequence, and mark the variation site information in the gene sequence and base sequence;
[0008] Step 2: For each generation i of family members in the family genomic data, calculate the difference D i,i-1 between each generation i and the parent generation, obtain the difference information, and store the difference information in the difference database;
[0009] Step three: encode the gene sequence, base sequence and variant site information in the family genomic data through the embedding model to generate gene sequence vector v gene , base sequence vector v base and variant site information vector v variant . At the same time, compress the original data, and after establishing a joint index of the compressed data and embedding vectors v gene , v base and v variant , obtain the index vector and store it in the embedding database;
[0010] Step four: for a given query sequence q, encode the gene sequence, base sequence and variant site information in the query sequence q using the embedding model, and if there is data missing, replace it with a zero vector. Then, perform weighted summation on the encoded vectors to obtain the query vector v query .
[0011] Step five: perform cosine similarity calculation on the query vector v query and the index vector in the embedding database, and select the k results with the highest cosine similarity, i.e. k gene sequences or base sequences, and obtain the variant site information according to the k gene sequences or base sequences;
[0012] Step six: based on the k gene sequences or base sequences, obtain the difference information in the difference database according to the range of the start site and the end site of each gene sequence or base sequence;
[0013] Step seven: visualize the variant site information obtained in step five and the difference information obtained in step six.
[0014] Further, the difference D i,i-1 is represented as:
[0015] D i,i-1 =f diff (G i ,G i-1 )
[0016] f diff (G i ,G i-1 )={(j,h j )|h j ≠g j}
[0017] where G i =[g1,g2,…,g n ], G i-1 =[h1,h2,…,h n ] and hj and g j denote the base of the parent and child at the jth position, j = 1, 2,... n, f diff denote the difference function.
[0018] Further, the compression in step three is BWT compression.
[0019] Further, the cosine similarity is represented as:
[0020]
[0021] where, v db denote the index vector stored in the embedding database.
[0022] Further, the visualization is specifically:
[0023] Step one: define the size W of the cache window and the moving step length of the window;
[0024] Step two: initialize a cache structure for storing data segments within the window and the start and end positions of the window;
[0025] Step three: determine whether the cache hits, if the cache hits, i.e. the request position is within the cache window, then directly return the data in the cache, if it does not hit, i.e. the request position is not within the cache window, then update the data in the cache;
[0026] The specific steps for updating the data in the cache are:
[0027] determine the request position p request , and let the start position of the new window p start = p request + W / 2, then according to p start and W, get the new window range [p start , p end ];
[0028] Call the data source interface to extract the data in the range [p start , p end ] from the search results and store it in the cache.
[0029] A storage and visual retrieval system based on family genomic data, the system comprises a data storage module, a data retrieval module and a visualization module;
[0030] The data storage module specifically performs the following steps:
[0031] Step one: obtain family genomic data, for family genomic data, construct each family member M iChromosome information including gene sequence and base sequence, and annotating mutation site information in the gene sequence and the base sequence;
[0032] Step two: for each generation i of family members in the family genomic data, calculate the difference D i,i-1 between each generation i of family members and their parents respectively, obtain difference information, and store the difference information in the difference database;
[0033] Step three: encode the gene sequence, base sequence and mutation site information in the family genomic data through the embedding model to generate gene sequence vector v gene , base sequence vector v base and mutation site information vector v variant , and at the same time, compress the original data, and after establishing a joint index of the compressed data and embedding vectors v gene , v base and v variant , obtain an index vector and store it in the embedding database;
[0034] The data retrieval module specifically performs the following steps:
[0035] Step four: for a given query sequence q, encode the gene sequence, base sequence and mutation site information in the query sequence q using the embedding model respectively, and if there is data missing, replace it with a zero vector, then perform weighted summation on the encoded vectors to obtain a query vector v quert ;
[0036] Step five: calculate the cosine similarity of the query vector v query and the index vector in the embedding database respectively, and select the k results with the highest cosine similarity, that is, the k gene sequences or base sequences, and obtain the mutation site information according to the k gene sequences or base sequences;
[0037] Step six: based on the k gene sequences or base sequences, according to the range of the start site and the end site of each gene sequence or base sequence, obtain the difference information in the difference database;
[0038] The visualization module is used to visualize the mutation site information obtained in step five and the difference information obtained in step six.
[0039] Further, the difference D i,i-1 is expressed as:
[0040] D i,i-1 =f diff (G i ,G i-1 )
[0041] fdiff (G i ,G i-1 )={(j,h j )|h j ≠g j}
[0042] where G i =[g1,g2,…,g n ], G i-1 =[h1,h2,…,h n ], h j and g j denote the bases of the parent and child at the jth position, j=1,2,...n, f diff denotes the difference function.
[0043] Further, the compression in the step three is BWT compression.
[0044] Further, the cosine similarity is represented as:
[0045]
[0046] where v db denotes the index vector stored in the embedding database.
[0047] Further, the visualization is specifically:
[0048] Step one: define the size W of the cache window and the moving step length of the window;
[0049] Step two: initialize a cache structure for storing the data segments within the window and the start and end positions of the window;
[0050] Step three: determine whether the cache hits, if the cache hits, i.e. the request position is within the cache window, then directly return the data in the cache, if it does not hit, i.e. the request position is not within the cache window, then update the data in the cache;
[0051] The specific steps of updating the data in the cache are:
[0052] Determine the request position p request , and let the start position of the new window p start =p request +W / 2, then according to p start and W, get the new window range [p start , p end ];
[0053] Call the data source interface to extract the data in the range [p start , p end ] from the search results and store it in the cache.
[0054] The beneficial effects of the present application are:
[0055] The present application organically integrates gene sequences, base variations, variation site information and family relationships in family genomic data, and stores the encoded vectors and the original data by establishing a joint index. When retrieving, the cosine similarity is used to determine the gene sequence or base sequence, and then the variation site information is obtained. The technical solution of the present application can improve the storage efficiency and retrieval efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 It is a schematic diagram of the overall structure of the present application;
[0057] Figure 2 It is a visual specific flowchart of the present application. DETAILED DESCRIPTION
[0058] It should be particularly noted that the various embodiments disclosed in the present application can be combined with each other without conflict.
[0059] Embodiment one: the storage and visualization retrieval method based on family genomic data according to the present embodiment comprises the following steps:
[0060] Step one: obtain the family genomic data, and construct the chromosome information of each family member M i including the gene sequence and the base sequence, and mark the variation site information in the gene sequence and the base sequence;
[0061] Step two: calculate the difference D i,i-1 of each family member of each generation i in the family genomic data with the parent, obtain the difference information, and store the difference information in the difference database;
[0062] Step three: encode the gene sequence, base sequence and variation site information in the family genomic data through the embedding model, respectively generate the gene sequence vector v gene , the base sequence vector v base and the variation site information vector v variant , and at the same time, compress the original data, and store the compressed data and the embedding vectors v gene , v base and v variant in the embedding database after establishing a joint index;
[0063] Step four: for a given query sequence q, the gene sequence, base sequence and variant site information in the query sequence q are respectively encoded using an embedding model, if there is data missing, it is replaced by a zero vector, then the encoded vectors are weighted and summed to obtain the query vector v query ;
[0064] Step five: the query vector v query is respectively calculated with the index vector in the embedding database Cosine similarity, and the k results with the highest cosine similarity, i.e. k gene sequences or base sequences, are selected, and the variant site information is obtained according to the k gene sequences or base sequences;
[0065] Step six: based on the k gene sequences or base sequences, the difference information is obtained in the difference database according to the range of the start site and the end site of each gene sequence or base sequence;
[0066] Step seven: the variant site information obtained in step five and the difference information obtained in step six are visualized.
[0067] The application also constructs a multi-level, dynamic visual display system, so that users can intuitively understand the genetic variation pattern and intergenerational transmission of family members in multi-dimensional data.
[0068] Specific implementation method two: this implementation method is a further description of the specific implementation method one, the difference between this implementation method and the specific implementation method one is that the difference D i,i-1 is expressed as:
[0069] D i,i-1 =f diff (G i ,G i-1 )
[0070] f diff (G i ,G i-1 )={(j,h j )|h j ≠g j}
[0071] Wherein, G i =[g1,g2,…,g n ], G i-1 =[h1,h2,…,h n ], h j and g j represent the base at the jth position of the parent and child, j=1,2,...n, f diff represents the difference function.
[0072] Specific implementation three: this implementation is a further description of the specific implementation two, the difference between this implementation and the specific implementation two is that the compression in step three is BWT compression.
[0073] Specific implementation four: this implementation is a further description of the specific implementation three, the difference between this implementation and the specific implementation three is that the cosine similarity is expressed as:
[0074]
[0075] Wherein, v db represents the index vector stored in the embedding database.
[0076] Specific implementation five: this implementation is a further description of the specific implementation four, the difference between this implementation and the specific implementation four is that the visualization is specifically:
[0077] Step one: define the size W of the cache window and the moving step of the window;
[0078] Step two: initialize a cache structure for storing data segments within the window and the start and end positions of the window;
[0079] Step three: determine whether the cache hits, if the cache hits, that is, the request position is within the cache window, then directly return the data in the cache, if not, that is, the request position is not within the cache window, then update the data in the cache;
[0080] The specific steps of updating the data in the cache are:
[0081] Determine the request position p request , and let the start position of the new window p start =p request +W / 2, then according to p start and W, get the new window range [p start , p end ];
[0082] Call the data source interface to extract the data in the range [p start , p end ] from the search results and store it in the cache.
[0083] Specific implementation six: this implementation is a storage and visualization retrieval system based on family genomic data, the system includes a data storage module, a data retrieval module and a visualization module;
[0084] The data storage module specifically performs the following steps:
[0085] Step one: obtain the family genomic data, and construct the chromosome information of each family member M i , including gene sequence and base sequence, and mark the variation site information in the gene sequence and base sequence;
[0086] Step two: calculate the difference D i,i-1 of each family member of each generation i in the family genomic data with the parent, obtain the difference information, and store the difference information in the difference database;
[0087] Step three: encode the gene sequence, base sequence and variation site information in the family genomic data through the embedding model to generate gene sequence vector v gene , base sequence vector v base and variation site information vector v variant , and compress the original data, and store the compressed data and embedding vectors v gene , v base and v variant in the embedding database after establishing a joint index;
[0088] The data retrieval module specifically performs the following steps:
[0089] Step four: for a given query sequence q, encode the gene sequence, base sequence and variation site information in the query sequence q using the embedding model, and replace the missing data with a zero vector, then perform weighted summation on the encoded vectors to obtain the query vector v query ;
[0090] Step five: calculate the cosine similarity of the query vector v query and the index vector in the embedding database, and select the k results with the highest cosine similarity, i.e. k gene sequences or base sequences, and obtain the variation site information according to the k gene sequences or base sequences;
[0091] Step six: based on the k gene sequences or base sequences, obtain the difference information in the difference database according to the range of the start site and the end site of each gene sequence or base sequence;
[0092] The visualization module is used to visualize the variation site information obtained in step five and the difference information obtained in step six.
[0093] Specific implementation method seven: this implementation method is a further description of the specific implementation method six, and the difference between this implementation method and the specific implementation method six is that the difference D i,i-1 is represented as:
[0094] D i,i-1= f diff (G i , G i-1 )
[0095] f diff (G i , G i-1 ) = {(j, h j ) | h j ≠ g j}
[0096] where G i = [g1, g2,..., g n ], G i-1 = [h1, h2,..., h n ], h j and g j represent the bases of the parent and child at the ith position, j = 1, 2,..., n, f diff represents the difference function.
[0097] Eighth Embodiment: This embodiment is a further illustration of the sixth embodiment, and the difference between this embodiment and the sixth embodiment is that the compression in the step three is BWT compression.
[0098] Ninth Embodiment: This embodiment is a further illustration of the sixth embodiment, and the difference between this embodiment and the sixth embodiment is that the cosine similarity is expressed as:
[0099]
[0100] where v db represents the index vector stored in the embedding database.
[0101] Tenth Embodiment: This embodiment is a further illustration of the sixth embodiment, and the difference between this embodiment and the sixth embodiment is that the visualization is specifically:
[0102] Step one: define the size W of the cache window and the moving step length of the window;
[0103] Step two: initialize a cache structure for storing the data segments within the window and the start and end positions of the window;
[0104] Step three: determine whether the cache hits, if the cache hits, i.e. the request position is within the cache window, then directly return the data in the cache, if it does not hit, i.e. the request position is not within the cache window, then update the data in the cache;
[0105] The specific steps for updating the data in the cache are:
[0106] determine the request position p requestLet p start be the start position of the new window request , and W / 2 the window size. Then, according to p start and W, the new window range [p start , p end ] is obtained.
[0107] The data source interface is called to extract the data in the range [p start , p end ] from the search results and store it in the cache.
[0108] 1. Family data initialization: For each family member M i , its genetic sequence G i , base sequence B i , and variant site information P i (including multiple chromosome numbers and corresponding variant base sites s) (a total of 24 chromosomes) are established.
[0109] 2. Differential storage: For each generation i of family members, the difference D i,i-1 with the parent is calculated to obtain the difference information and store:
[0110] D i,i-1 = f diff (G i , G i-1 )
[0111] Let G i = [g1, g2, …, g n ] and G i-1 = [h1, h2, …, h n ], where h j and g j represent the bases at the jth position of the parent and child. The difference function f diff is defined as:
[0112] f diff (G i , G i-1 ) = {(j, h j ) | h j ≠ g j}
[0113] That is, record all the different positions j and their corresponding bases h j in the child and parent sequences. In this way, only the variant points of the child and the parent are stored, greatly reducing the storage requirement.
[0114] 3. Modal hierarchical compression: BWT compression algorithm is used for the genetic sequence layer, the base sequence layer, and the variant site information layer, respectively:
[0115] SG = Compress(G i ), SB = Compress(B i ), SP = Compress(P i ),
[0116] where S G , S B , S P represent the compressed gene, base and variant site information storage structures respectively.
[0117] 4. Multi-modal joint index construction:
[0118] Construct a multi-modal joint index Index(M i ) for the gene sequence, base sequence and variant site information in the database, so that the gene sequence, base sequence and variant site information of each family member are associated together:
[0119] Index(M i ) = {(G i , B i , P i ) | i e {1,..., n}}
[0120] Multi-modal joint retrieval based on gene, base sequence and variant site information
[0121] In the retrieval process, the present application uses three data types of gene, base sequence and variant site information, which require different processing methods and similarity measures, and at the same time supports multi-information mixed retrieval. The specific implementation process is as follows:
[0122] Receive one or a combination of site information, base sequence and gene sequence to be retrieved;
[0123] 1. Data preprocessing: In order to realize joint retrieval, we design a mechanism to enable gene sequence, base sequence and site sequence to be associated in similarity calculation. By defining an embedding model, the three different types of data are mapped to a unified embedding space, thereby realizing cross-modal data representation and similarity calculation. The specific process is as follows:
[0124] (1) Embedding model construction
[0125] To analyze the correlation degree of gene sequence, base sequence and variant site information, the model first converts the gene sequence, base sequence and variant site information into low-dimensional vector representation through an independent embedding layer, then extracts features through a convolutional layer, and finally combines these feature vectors in a shared space. Through contrastive loss (Contrastive Loss), the model is trained to ensure that similar input data (such as similar genes, bases and variant sites) are closer in the embedding space, and dissimilar data are farther apart. The embedding model construction flowchart is shown in Figure 1 .
[0126] (2) Data conversion
[0127] Convert the search input gene sequence, base sequence and variant site sequence into a format acceptable to the embedding model, input it into the model, obtain the output result of each data type, and generate the corresponding embedding vector through the model.
[0128] Gene sequence: After inputting the gene sequence, use the trained embedding model to convert it into a vector v gene , representing the semantic information of the gene sequence.
[0129] Base sequence: After inputting the base sequence, use the trained embedding model to convert it into a vector v base .
[0130] Variant site: After inputting the variant site information, use the trained embedding model to convert it into a vector v variant .
[0131] (3) Vector combination
[0132] The corresponding vectors of the gene sequence, base sequence and variant site sequence are input into the shared embedding space as joint input. If a type of input is missing, replace it with a zero vector. This embedding space aims to map all data types (genes, bases, variants) to a unified semantic space. Specifically, the three vectors are integrated into a comprehensive vector v query through weighted summation, and this vector is used for retrieval of embedding vectors in the database.
[0133] 2. Index data embedding vector: In order to be able to match data from the database, we need to pre-convert all data (gene sequence, base sequence, variant site, etc.) in the database into embedding vectors through the same embedding model, and store these embedding vectors in an index database. Each data entry will have a corresponding embedding vector v db , representing the semantic information of the data.
[0134] 3. Similarity measurement method: During the retrieval process, the query vector (v query) in the database. To measure their similarity, cosine similarity measure is used: db
[0135]
[0136] The closer the value of cosine similarity is to 1, the more similar the two vectors are.
[0137] 4. Retrieval process: In the retrieval process, for a given query sequence q, it is first converted into a format acceptable by the embedding model, and the corresponding gene, base, and variant embedding vectors are generated respectively. If some type of data is missing (e.g., only gene sequence or variant sequence), a zero vector is used to replace the input of that type. Then, according to the above method, the three embedding vectors (v gene , v base , and v variant ) are weighted and summed to obtain a comprehensive query vector v gene . The query vector v query will be compared with the embedding vectors stored in the database for similarity calculation. To measure the similarity, cosine similarity is used to calculate the similarity between the query vector and each index field in the database. Cosine similarity judges the correlation between vectors by calculating the angle similarity between them, the closer the value is to 1, the more similar, the closer the value is to 0, the less similar.
[0138] 5. Ranking and selection: All data items are ranked according to the comprehensive similarity measure, ensuring that the data items with the highest similarity are ranked first. Then, from the ranking results, the top k data items with the highest similarity value are selected as the final retrieval results.
[0139] 6. Output results: The top k most similar data items after sorting are output as the final retrieval results returned to the user. The specific content and structure of the returned results are as follows:
[0140] Matching gene sequence: gene sequence information similar to the query gene sequence, from the start site to the end site.
[0141] Matching base sequence: gene fragment information similar to the query base sequence, from the start site to the end site.
[0142] Matching variant: variant information similar to the query variant, such as SNP site, variant type, etc.
[0143] Family data difference information: according to the range of matching sites, obtain the difference site information.
[0144] Similarity score: The matching score calculated based on similarity (e.g. cosine similarity) can help users evaluate the relevance of the returned results.
[0145] Sliding window cache strategy visual retrieval results
[0146] As shown in Figure 2 , for the results of user retrieval, the sliding window cache strategy can be used to dynamically present the results on the visualization interface. The specific implementation process is as follows:
[0147] 1. Initialize the cache
[0148] Cache window initialization: When the system starts or is first loaded, define the size W (e.g. 1,000 bases) of the cache window and the moving step size S of the window.
[0149] Cache structure creation: Initialize a cache structure to store the data segments within the window and the start and end positions of the window.
[0150] 2. Determine cache hit
[0151] Check if result is within the range of the cache window. If the cache is hit (i.e. the request position is within the cache range (cache window)), return the data in the cache directly. If the cache is not hit (i.e. the request position is outside the current cache range), the cache needs to be updated.
[0152] 3. Update mechanism of the cache window
[0153] Determine the new window position:
[0154] Calculate the new window range [p request , p start ] containing the request position p end , and let the start position of the new window p start = p request +W / 2, so that the window center is close to the request position, thereby increasing the hit rate of subsequent retrieval.
[0155] Data interface call to load new window data:
[0156] Call the data source interface to extract the data in the range [p start , p end ] from the retrieval results and store it in the cache. Update the boundary positions start_position and end_position of the cache window.
[0157] The present application protects an efficient storage and visual retrieval method for family genomic data. The core content includes multi-modal data fusion storage and differential compression technology, dynamic caching strategy based on sliding window, multi-dimensional flexible query mechanism, and multi-level, interactive visualization display method. These innovative solutions effectively improve the storage efficiency, retrieval speed and display effect of family genomic data, providing strong technical support for genetics and genome research.
[0158] When considering the overall alternatives or modifications of the present application, competitors may explore different technologies to achieve similar effects. For example, instead of sliding window caching technology, competitors may use distributed databases or cloud computing resources to improve the storage and access speed of genomic data. In addition, instead of the method of integrating family genetic relationship display, competitors may develop tools based on different visualization technologies, such as using virtual reality (VR) or augmented reality (AR) to provide a more immersive user experience.
[0159] From the perspective of data processing, competitors may use machine learning algorithms to optimize genetic variation tracking analysis, and use advanced data analysis tools to replace traditional query and analysis methods, improving the accuracy of prediction and the convenience of operation. In general, although the present application proposes an effective solution, there may be or will be other technical routes in the market that achieve similar purposes through different combinations of technologies or emerging technologies, which may serve as potential alternatives or modifications of the present application.
[0160] It should be noted that the specific embodiments are only an explanation and description of the technical solutions of the present application, and cannot limit the scope of protection. Any partial change made according to the claims and description of the present application shall still fall within the scope of protection of the present application.
Claims
1. A method for storing and visualizing retrieval based on family genomic data, characterized in that The method comprises the following steps: Step one: obtain family genomic data, for the family genomic data, build the chromosome information of each family member, including gene sequence and base sequence, and label the variation site information in the gene sequence and base sequence; Step two: obtain the family member's phenotype data, and build the phenotype data of each family member; Step two: for each generation i of the family members in the family genomic data, calculate the difference between each generation i and the parent generation respectively , obtain the difference information, and store the difference information in the difference database; Step three: encode the gene sequence, base sequence and mutation site information in the family genome data through the embedding model to generate gene sequence vectors , base sequence vectors and mutation site information vectors , and at the same time, compress the original data, and establish a joint index of the compressed data and the embedding vectors , and to obtain an index vector, which is stored in the embedding database; Step four: for a given query sequence q, the gene sequence, base sequence and variant site information in the query sequence q are respectively encoded using an embedding model, if there is data missing, a zero vector is used instead, then the encoded vectors are weighted and summed to obtain a query vector ; Step 5: Query vector Calculate the cosine similarity with the index vectors embedded in the database and select the one with the highest cosine similarity The result is gene sequence or base sequence, and according to Gene sequence or base sequence to obtain the mutation site information; Step 6: Based on Gene sequences or base sequences are obtained from a difference database according to the range of the starting site and the ending site of each gene sequence or base sequence; Step seven: visualizing the variation site information obtained in step five and the difference information obtained in step six; The difference is represented as: wherein , , and represent the bases of the parent and child at position j = 1, 2,... n, denotes the difference function. 2. The method of claim 1, wherein the method further comprises: The compression in step three is BWT compression.
3. The method of claim 1, wherein the method further comprises: determining a plurality of genomic regions of interest; and determining a plurality of genomic regions of interest in the genomic data of the plurality of individuals in the family. The cosine similarity is expressed as: wherein, represents an index vector stored in an embedded database.
4. The method of claim 1, wherein the method further comprises: determining a plurality of genomic regions of interest; and determining a plurality of genomic regions of interest in the genomic data of the plurality of individuals. The visualization is specifically: Step one: define the size of the cache window and the moving step of the window; Step two: initializing a cache structure for storing data segments within a window and start and end positions of the window; Step three: judging whether the cache hits, if the cache hits, i.e. the request position is within the cache window, directly returning the data in the cache, if not, i.e. the request position is not within the cache window, updating the data in the cache; The specific steps of updating the data in the cache are: Determine the request position , and let the start position of the new window = , then according to and , get the new window range ; The data source interface is invoked to extract the data from the search results and store it in the cache.
5. A storage and visual retrieval system based on family genomic data, characterized in that The system comprises a data storage module, a data retrieval module and a visualization module; The data storage module specifically performs the following steps: Step one: obtain the family genomic data, for the family genomic data, build the chromosome information of each family member, including gene sequence and base sequence, and label the variation site information in the gene sequence and base sequence; Step two: obtain the family phenotype data, and build the phenotype information of each family member; Step two: for each generation i of the family members in the family genomic data, calculate the difference between each generation i and the parent generation respectively , obtain the difference information, and store the difference information in the difference database; Step three: encode the gene sequence, base sequence and mutation site information in the family genome data through the embedding model to generate gene sequence vectors , base sequence vectors and mutation site information vectors Meanwhile, the original data is compressed, and after establishing a joint index of the compressed data and embedding vectors , and , an index vector is obtained and stored in the embedding database; The data retrieval module specifically performs the following steps: Step four: for a given query sequence q, the gene sequence, base sequence and variant site information in the query sequence q are respectively encoded using an embedding model, if there is data missing, a zero vector is used instead, then the encoded vectors are weighted and summed to obtain a query vector ; Step 5: Query vector Calculate the cosine similarity with the index vectors embedded in the database and select the one with the highest cosine similarity The result is gene sequence or base sequence, and according to Gene sequence or base sequence to obtain the mutation site information; Step 6: Based on Gene sequences or base sequences are obtained from a difference database according to the range of the starting site and the ending site of each gene sequence or base sequence; The visualization module is used for visualizing the variation site information obtained in step five and the difference information obtained in step six; The difference is represented as: wherein , , and represent the bases of the parent and child at position j = 1, 2,... n, , denotes the difference function.
6. The system according to claim 5, wherein The compression in step three is BWT compression.
7. The system according to claim 5, wherein The cosine similarity is expressed as: wherein, represents an index vector stored in the embedded database.
8. The system according to claim 5, wherein The visualization is specifically: Step one: define the size of the cache window and the step size of the window movement; Step two: initializing a cache structure for storing data segments within a window and start and end positions of the window; Step three: judging whether the cache hits, if the cache hits, i.e. the request position is within the cache window, directly returning the data in the cache, if not, i.e. the request position is not within the cache window, updating the data in the cache; The specific steps of updating the data in the cache are: Determine the request position , and let the start position of the new window = (request position - window size) mod window size , then according to and , get the new window range ; The data source interface is invoked to extract the data in the result set and store it in the cache.
Citation Information
Patent Citations
Large-scale gene sequencing data storage and query system
CN113901006A
Method, device and equipment for constructing multi-modal knowledge retrieval system fused with large model
CN118394978A