A genomic data analysis method

Through adaptive multidimensional spatial compression algorithm and gene variant-phenotype association analysis method, genomic data is efficiently compressed and accurately analyzed, solving the problems of inaccurate and inefficient handling in traditional technologies, and achieving the effect of efficient storage and accurate screening of disease markers.

CN120199337BActive Publication Date: 2025-08-26XIDIAN GRP HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510677106.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-26
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Traditional genomic data storage and processing techniques have problems such as inaccurate processing, low computational efficiency and low analysis accuracy in the genomic data analysis process.

Method used

The genomic data is intelligently compressed by adaptive multidimensional spatial compression algorithm, and combined with the gene variant-phenotype correlation analysis method, the genomic data is compressed locally and globally through the adaptive multidimensional spatial compression algorithm, and the variant information is further compressed using the variant differential coding mechanism, and the disease markers are screened in combination with nonlinear weighted recursion.

Benefits of technology

It effectively reduces the storage needs of genomic data, retains key information, improves the accuracy of data compression efficiency and correlation analysis, and can efficiently screen out potential markers related to diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199337B_ABST
    Figure CN120199337B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing technology, and in particular to a genomic data analysis method. The method comprises the following steps: preprocessing genomic data, performing intelligent compression processing on the preprocessed genomic data using an adaptive multidimensional space compression algorithm to obtain compressed genomic data; performing gene variation detection and filtering on the compressed genomic data to obtain variant genomic data, and performing functional annotation to obtain genomic data with annotation results; performing gene expression analysis on the genomic data with annotation results to obtain gene expression data; and performing association analysis based on the genomic data with annotation results and gene expression data using a gene variation-phenotype association analysis method to screen out potential disease markers. The method solves the technical problems of inaccurate genomic data processing, low computational efficiency, and low analysis accuracy in the genomic data analysis process of traditional genomic data storage and processing technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a genome data analysis method. Background Art

[0002] With the rapid development of genomic research, the scale of genomic data continues to expand, especially in the fields of large-scale genome sequencing and personalized medicine, where the amount of data generated is growing exponentially. Traditional data storage and processing technologies are no longer able to meet this growing data demand. This is especially true in the application of high-throughput sequencing technologies, which generate large amounts of raw genomic data that require rigorous variant detection, data compression, and storage, and which must also be efficiently queried and analyzed.

[0003] However, traditional genomic data storage and processing technologies have the following technical problems: inaccurate genomic data processing, low computational efficiency, and low analysis accuracy during genomic data analysis. Summary of the Invention

[0004] The present invention provides a genome data analysis method to solve the technical problems of inaccurate genome data processing, low computational efficiency and low analysis accuracy in the genome data analysis process in traditional genome data storage and processing technologies.

[0005] A genome data analysis method of the present invention specifically includes the following technical solutions:

[0006] A genomic data analysis method comprises the following steps:

[0007] S1. Preprocessing the genomic data to obtain preprocessed genomic data; performing intelligent compression processing on the preprocessed genomic data using an adaptive multidimensional space compression algorithm to obtain compressed genomic data; performing gene variation detection and filtering on the compressed genomic data to obtain variant genomic data, and performing functional annotation to obtain genomic data with annotation results;

[0008] S2. Perform gene expression analysis on the annotated genomic data to obtain gene expression data; based on the annotated genomic data and gene expression data, perform association analysis using gene variation-phenotype association analysis to screen for potential disease markers.

[0009] Preferably, the S1 specifically includes:

[0010] In the implementation process of the adaptive multidimensional space compression algorithm, the preprocessed genomic data is regarded as a point cloud in a high-dimensional space, and high-dimensional space mapping is applied to map the preprocessed genomic data to a low-dimensional space to obtain the mapped genomic data.

[0011] Preferably, the S1 specifically includes:

[0012] The mapped genomic data is divided into data blocks, and an adaptive local compression strategy is adopted for the data blocks. The compression method is dynamically selected based on the local characteristics of the data blocks to perform local compression.

[0013] Preferably, the S1 specifically includes:

[0014] In the process of local compression, a local compression function is defined. Through weighted summation and logarithmic transformation, the compression strategy is adjusted to obtain the local compression cost of the data block. The specific formula of the local compression function is as follows:

[0015] ;

[0016] in, It is data blocks The storage space required after local compression represents the local compression cost of the data block; is the total number of feature dimensions considered during compression; is the weighting factor, which indicates the The weight of each feature; is the scaling factor used to control each variation feature The degree of contribution to the compression effect; It is in data blocks The mapped genomic data in The variation function on each feature is calculated from the specific features in the genome sequence;

[0017] Compression processing is performed based on the local compression cost of the data block to obtain a locally compressed data block.

[0018] Preferably, the S1 specifically includes:

[0019] Based on the local compression cost of the data block and the similarity measurement between the locally compressed data blocks, the mapped genome data is globally compressed to obtain the globally compressed genome data. The specific formula is:

[0020] ;

[0021] in, It is the globally compressed genomic data; is the total number of mapped genomic data blocks; It is The weight of a locally compressed data block; It is data blocks The local compression cost of It is Partially compressed data blocks Hedi Partially compressed data blocks Similarity measure between ; is a tuning parameter.

[0022] Preferably, the S1 specifically includes:

[0023] A mutation difference coding mechanism is introduced to calculate the difference coding of the mutation point based on the mutation point and adjacent mutation points in the globally compressed genomic data; a global index is introduced to locate the data position; and the compressed genomic data is obtained based on the local compression cost of the data block, the globally compressed genomic data, the difference coding of the mutation point and the global index.

[0024] Preferably, the S2 specifically includes:

[0025] The variation data in the genomic data with annotation results is recorded as gene variation data; the gene variation-phenotype association analysis method is based on gene variation data and gene expression data, combined with disease phenotype data, and dynamically adjusts the contribution of various types of data to disease marker screening through a nonlinear weighted recursive method.

[0026] Preferably, the S2 specifically includes:

[0027] In the implementation of the gene variation-phenotype association analysis method, the gene variation data is preprocessed, and the gene expression data and disease phenotype data are standardized; by calculating the correlation between the gene variation data, gene expression data and disease phenotype data, the preliminary weighting factors of the gene variation data and the preliminary weighting factors of the gene expression data are obtained, and the weighting factors are dynamically updated in a recursive manner.

[0028] Preferably, the S2 specifically includes:

[0029] Based on the weighting factor, the relationship between gene variation data and disease phenotype data is modeled through a nonlinear regression model, and a regularization term is introduced to optimize the nonlinear regression model to obtain an optimized nonlinear regression model; based on the optimized nonlinear regression model, the gene variation-phenotype correlation is calculated and ranked, and the genes most relevant to the disease phenotype are selected as potential disease markers.

[0030] The beneficial effects of the technical solution of the present invention are:

[0031] 1. Through an adaptive multidimensional spatial compression algorithm and a variation differential encoding mechanism, genomic data is compressed into a very small storage space, not only reducing storage requirements but also effectively preserving key information, such as gene variation and gene expression changes. An adaptive compression strategy dynamically adjusts the compression method based on the local characteristics and global redundancy of data blocks, making data compression more efficient. This is particularly effective in compressing variation information and redundant portions of gene sequences, improving the overall compression ratio.

[0032] 2. Perform association analysis through gene variation-phenotype association analysis. Based on gene variation data, gene expression data, and disease phenotype data, the contribution of various types of data to the final disease marker screening is dynamically adjusted through nonlinear weighted recursion to improve the accuracy of association analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a flow chart of a genome data analysis method according to the present invention. DETAILED DESCRIPTION

[0034] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0035] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0036] The following describes in detail a specific scheme of a genome data analysis method provided by the present invention with reference to the accompanying drawings.

[0037] Refer to the attached Figure 1 , which shows a flow chart of a genome data analysis method provided by one embodiment of the present invention, the method comprising the following steps:

[0038] S1. Preprocessing the genomic data to obtain preprocessed genomic data; performing intelligent compression processing on the preprocessed genomic data using an adaptive multidimensional space compression algorithm to obtain compressed genomic data; performing gene variation detection and filtering on the compressed genomic data to obtain variant genomic data, and performing functional annotation to obtain genomic data with annotation results;

[0039] The genomic data is preprocessed (such as removing adapter sequences, trimming poor quality sequences, and removing contaminating sequences) to obtain preprocessed genomic data; the technical means used in the preprocessing process are all well known to those skilled in the art and are not described in detail here.

[0040] The preprocessed genomic data is intelligently compressed using an adaptive multidimensional space compression algorithm to obtain compressed genomic data. This algorithm utilizes an adaptive compression strategy to maximize data compression. Furthermore, considering the specific characteristics of genomic data (such as the presence of variant information), specialized differential encoding is performed on the variants in the preprocessed genomic data to further compress the variant data. Furthermore, the adaptive multidimensional space compression algorithm supports efficient random access to the compressed data. The specific implementation of the adaptive multidimensional space compression algorithm is as follows:

[0041] First, the preprocessed genomic data is regarded as a point cloud in a high-dimensional space, represented as ,in, Represents a point cloud; It is a point cloud Middle data points, represented as a dimensional vector, which contains the actual information of the gene sequence, the mutation location, the mutation type, the sequence repetition degree and other metadata; is the total number of data points in the point cloud. To obtain a compact compressed representation, a high-dimensional space mapping is applied to map the pre-processed genomic data into a new low-dimensional space using existing spatial embedding methods (such as PCA or t-SNE). This mapping must not only consider local structure (such as the similarity between similar sequences) but also maintain global structure (such as the distribution of variant points in the genome). The mapping of each data point is represented as:

[0042] ;

[0043] in, It is data points Representation in low-dimensional space, i.e., mapped genomic data; is the dimension of the low-dimensional space; It is a spatial embedding function obtained by spatial embedding methods such as PCA or t-SNE; It is data points In low-dimensional space The value of the dimension;

[0044] Furthermore, after high-dimensional space mapping, there is a lot of repeated information and structural redundancy in the local area, so it is necessary to locally compress the mapped genomic data. The mapped genomic data is divided into For each data block, an adaptive local compression strategy is adopted, and a compression method is dynamically selected according to the local characteristics of the data block (such as repetition, variation density, etc.), and local compression is performed to obtain the locally compressed data block;

[0045] The goal of local compression is to make the genomic data representation more compact through quantization and encoding. Therefore, we define a local compression function based on the various feature dimensions of the data block. We adjust the compression strategy through weighted summation and logarithmic transformation to obtain the local compression cost of the data block. The specific formula of the local compression function is as follows:

[0046] ;

[0047] in, It is data blocks The storage space required after local compression, i.e. the local compression cost of a data block, is an indicator of compression efficiency. A smaller value indicates a better compression effect. is the total number of feature dimensions considered during compression; is the weighting factor, which indicates the The weight of each feature is used to adjust the influence of each feature on the compression effect, which is determined based on expert experience; is the scaling factor used to control each variation feature The contribution to the compression effect determines the variation characteristics The magnitude of the effect during compression; It is in data blocks The mapped genomic data in The variation function on a feature is calculated based on specific features in the genome sequence (such as gene mutations, insertion / deletion (InDel) variations, changes in gene expression, or other values ​​related to biological characteristics), such as ,in, Indicates the data blocks Some statistics related to the variation characteristics in is the number of types of statistics related to the variation characteristics, It is with The weight coefficients related to the statistics, It is data blocks Middle and First For example, if it is the frequency of mutation, the frequency of each mutation can be used as ; Indicates the data blocks Changes in gene expression, is the coefficient associated with changes in gene expression, Is the gene expression difference, which indicates the difference in gene expression levels under different conditions (such as comparing the expression differences under two conditions). For example, if the expression values ​​of a gene under different conditions are and ,So The purpose of the above local compression function is to adjust the compression strength according to the density and importance of each feature. By introducing the logarithmic transformation, the eigenvalues ​​with large numerical variations can be effectively compressed while retaining important structural information.

[0048] Based on the local compression cost of the data block, an existing compression technology such as Huffman coding, LZ77, arithmetic coding, etc. is selected according to expert experience to perform compression processing to obtain a locally compressed data block;

[0049] Global compression is performed based on the locally compressed data blocks. The goal of global compression is to further reduce the storage capacity of genomic data by analyzing the global redundancy relationship between the mapped genomic data. Taking into account the global repetitiveness and variant data in the mapped genomic data, combined with the similarity measurement between the locally compressed data blocks, the local compression costs of all data blocks are weighted summed to obtain the globally compressed genomic data. The specific formula is:

[0050] ;

[0051] in, It is the globally compressed genomic data; is the total number of mapped genomic data blocks; It is The weight of a locally compressed data block indicates the importance of the locally compressed data block in global compression. , It is The variability of the locally compressed data block reflects the number of variations; It is The size of a partially compressed data block; 、 、 is the weighting coefficient, which is used to control the influence of different factors on the weight calculation and is determined according to the expert experience method; It is The entropy value of a locally compressed data block indicates the amount of information in the locally compressed data block; It is Partially compressed data blocks Hedi Partially compressed data blocks The similarity measure between them indicates the similarity between the two data blocks in structure or content; It is a tuning parameter used to control the impact of similarity measurement on the compression result.

[0052] For variations in genomic data (such as single nucleotide polymorphisms (SNPs) and insertions and deletions (InDels)), a variation difference encoding mechanism is introduced to further compress the data based on the globally compressed genomic data, especially at continuous variation positions, which can effectively reduce redundant information. The specific content of the variation difference encoding mechanism is: for variation points in the globally compressed genomic data, and adjacent mutation points , the difference coding formula is used to calculate the difference coding of the variation point. The difference coding formula is as follows:

[0053] ;

[0054] in, is the first Differential coding of mutation points; and are the positions of the current and previous mutation points in the globally compressed genomic data; is the first The type code of each variant point (such as SNP is 1, InDel is 2, etc.);

[0055] Through differential encoding, the variant data is compressed into a more concise representation, thereby further improving the compression efficiency. Especially for those variants with small changes, differential encoding can significantly reduce the storage requirements;

[0056] Furthermore, in order to support efficient random access, the existing global index structure is used to quickly locate any location of the data, support efficient query and decompression operations, and the global index is recorded as ,Through global indexing, users can directly jump to a specific location in the compressed ,data for random access, and the decompression process will only be performed on the ,required data block without decompressing the entire data;

[0057] Based on the local compression cost of the data block, the globally compressed genome data, the difference encoding of the mutation point and the global index, the compressed genome data is obtained. , ;

[0058] Genetic variant detection and filtering are performed on the compressed genomic data. Efficient variant detection tools, such as GATK or SAMtools, are used to identify SNPs (single nucleotide polymorphisms) and Indels (insertion-deletion variants) in the compressed genomic data. Pre-set filtering criteria (such as depth and quality) based on expert experience are then used to screen for possible false positives and low-quality variants, thereby obtaining variant genomic data. Furthermore, functional annotation of the variant genomic data is performed, i.e., by connecting it to known gene libraries (such as RefSeq and Ensembl) to understand the biological significance of each variant. Functional annotation not only helps determine the potential impact of variants on gene function but also helps discover disease-related variants, ultimately obtaining annotated genomic data.

[0059] S2. Perform gene expression analysis on the annotated genomic data to obtain gene expression data; based on the annotated genomic data and gene expression data, perform association analysis using gene variation-phenotype association analysis to screen for potential disease markers.

[0060] Gene expression analysis is performed on genomic data with annotation results, where the genomic data with annotation results contain variation information (such as SNPs, Indels, etc.) and have the location and function information of the genes marked. The genomic data with annotation results are analyzed using differential expression analysis tools (such as DESeq2, EdgeR, etc.), and the expression differences of each gene under different conditions are calculated and the statistical significance of the expression differences is evaluated to obtain gene expression data, including the expression profiles and differentially expressed genes of the genomic data under different conditions.

[0061] Furthermore, association analysis was performed using the gene variation-phenotype association analysis method. This method is based on variant data (gene variation data) in annotated genomic data, gene expression data, and disease phenotype data from existing databases. Through a nonlinear weighted recursive approach, it dynamically adjusts the contribution of each type of data to the final disease marker screening. The specific implementation process of the gene variation-phenotype association analysis method is as follows:

[0062] First, the gene variation data Preprocessing is performed using techniques known to those skilled in the art to remove low-frequency variants and low-quality variants, and the gene variation data are used for subsequent analysis, expressed as the genotype of each variant site; the gene expression data are used as the Standardize existing disease phenotypic data to ensure that all disease phenotypic data are in a unified framework to facilitate subsequent association analysis. For disease phenotypic data, binary classification (0 or 1) is used to represent disease status, for example, "1" represents the disease group and "0" represents the control group;

[0063] Furthermore, a weighting factor is introduced to dynamically weight the contribution of gene variation data and gene expression data to disease phenotype data, calculate the correlation between gene variation data and disease phenotype data, and obtain the preliminary weighting factor of gene variation data. The calculation formula of the preliminary weighting factor of gene variation data is as follows:

[0064] ;

[0065] in, It is the preliminary weighting factor of gene variation data, reflecting the correlation between gene variation data and disease phenotype data; is the total number of gene variation data; The first The amount of variation in a gene; The disease phenotype data Disease status of gene variation data; similarly, preliminary weighting factors of gene expression data are calculated ;

[0066] In order to make the weighting factor gradually tend to be stable, it is necessary to dynamically update the weighting factor through recursion. The update formula of the weighting factor is as follows:

[0067] ;

[0068] ;

[0069] in, and For the current iteration (i.e. iterations); and is a smoothing factor, which is used to control the influence of the results of the previous iteration and is obtained through experiments; and For the last iteration (i.e. Through the recursive update process, the weighting factors can be gradually optimized to make the weighted contributions of gene variation data and gene expression data more accurate;

[0070] Subsequently, a nonlinear regression model is used to model the relationship between gene variation data and disease phenotype data to capture the possible nonlinear association between the two and obtain the gene variation-phenotype correlation. The specific expression of the nonlinear regression model is:

[0071] ;

[0072] in, It is The variation-phenotype association of each gene; is the regression coefficient, determined experimentally; The gene expression data By adding sine, logarithmic, and square root functions, the nonlinear regression model can more flexibly fit complex nonlinear relationships, thereby more accurately capturing the implicit connection between gene variation data and disease phenotype data, making the nonlinear regression model more robust and adaptable when processing data; in order to ensure stable optimization of the nonlinear regression model in multiple rounds of iterations, regularization terms are further introduced to prevent overfitting, such as adding the L2 norm regularization term when minimizing the objective function;

[0073] Finally, based on the optimized nonlinear regression model, the gene variation-phenotype correlation was recalculated and sorted according to its value to obtain the most relevant marker gene set after sorting. This set of genes was used as candidates for disease marker screening, and the genes most relevant to the disease phenotype were selected as potential disease markers for subsequent verification.

[0074] In summary, a genomic data analysis method was completed.

[0075] The order in which the embodiments of the invention are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0076] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0077] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A genome data analysis method, characterized in that: The following steps are involved: S1. Preprocessing genomic data to obtain preprocessed genomic data; performing intelligent compression processing on the preprocessed genomic data using an adaptive multidimensional space compression algorithm to obtain compressed genomic data; in the implementation process of the adaptive multidimensional space compression algorithm, treating the preprocessed genomic data as a point cloud in a high-dimensional space, and applying high-dimensional space mapping to map the preprocessed genomic data to a low-dimensional space to obtain mapped genomic data; dividing the mapped genomic data into data blocks, adopting an adaptive local compression strategy for the data blocks, and dynamically selecting a compression method based on the local features of the data blocks to perform local compression; in the local compression process, defining a local compression function, adjusting the compression strategy through weighted summation and logarithmic transformation to obtain a local compression cost of the data block; performing compression processing based on the local compression cost of the data block to obtain a locally compressed data block; Perform gene variation detection and filtering on the compressed genomic data to obtain mutated genomic data, and perform functional annotation to obtain genomic data with annotation results; S2. Perform gene expression analysis on the annotated genomic data to obtain gene expression data; perform association analysis using gene variation-phenotype association analysis methods based on the annotated genomic data and gene expression data to screen for potential disease markers; The variation data in the genomic data with annotation results is recorded as gene variation data; the gene variation-phenotype association analysis method is based on gene variation data and gene expression data, combined with disease phenotype data, and dynamically adjusts the contribution of various types of data to disease marker screening through a nonlinear weighted recursive method.

2. A genome data analysis method according to claim 1, characterized in that: Said S1 specifically includes: The specific formula of the local compression function is as follows: , in, It is data blocks The storage space required after local compression represents the local compression cost of the data block; The total number of feature dimensions considered during compression; Weighting factor, which indicates the The weight of each feature; is the scaling factor used to control each variation feature The degree of contribution to the compression effect; It is in data blocks The mapped genomic data in The variation function on a feature is calculated from specific features in the genome sequence.

3. A genome data analysis method according to claim 1, characterized in that: Said S1 specifically includes: Based on the local compression cost of the data block and the similarity measurement between the locally compressed data blocks, the mapped genome data is globally compressed to obtain the globally compressed genome data. The specific formula is: , Among them, is the globally compressed genome data; is the total number of mapped genomic data blocks; It is The weight of a locally compressed data block; It is data blocks The local compression cost of It is Partially compressed data blocks Hedi Partially compressed data blocks Similarity measure between ; is a tuning parameter.

4. A genome data analysis method according to claim 3, characterized in that: Said S1 specifically includes: A mutation difference coding mechanism is introduced. Based on the difference between the mutation point and the adjacent mutation points in the globally compressed genomic data, as well as the type coding of the mutation point in the globally compressed genomic data, the difference coding of the mutation point is calculated. A global index is introduced to locate the data position. Based on the local compression cost of the data block, the globally compressed genomic data, the difference coding of the mutation point and the global index, the compressed genomic data is obtained.

5. A genome data analysis method according to claim 1, characterized in that: Said S2 specifically includes: In the implementation of the gene variation-phenotype association analysis method, the gene variation data is preprocessed, and the gene expression data and disease phenotype data are standardized; by calculating the correlation between the gene variation data, gene expression data and disease phenotype data, the preliminary weighting factors of the gene variation data and the preliminary weighting factors of the gene expression data are obtained, and the weighting factors are dynamically updated in a recursive manner.

6. A genome data analysis method according to claim 5, characterized in that: Said S2 specifically includes: Based on the weighting factor, the relationship between gene variation data and disease phenotype data is modeled through a nonlinear regression model, and a regularization term is introduced to optimize the nonlinear regression model to obtain an optimized nonlinear regression model; based on the optimized nonlinear regression model, the gene variation-phenotype correlation is calculated and ranked, and the genes most relevant to the disease phenotype are selected as potential disease markers.

Citation Information

Patent Citations

  • Solid tumor mutant gene detection and analysis method and system

    CN115132276A

  • Methods and systems for storing genomic data in a file structure comprising an information metadata structure

    EP4226382A1

  • Disease profiling information providing system based on multiple database information and method therefor

    KR102483880B1