Method for collecting and converging phenotype-haplotype associated annotation data based on online analysis platform

By using data mining and cross-validation on an online analytics platform, a reliable haplotype-phenotype association labeled dataset was generated, solving the problems of high resource consumption and insufficient quality in biological breeding data collection, and achieving efficient data aggregation and model training.

CN121171372APending Publication Date: 2025-12-19INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511311062.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

In existing technologies, the acquisition of biological breeding data, especially the acquisition of genotype-phenotype relationship annotation data, suffers from problems such as high resource consumption, lack of multi-party verification, and insufficient data quality, which affect the training efficiency and accuracy of intelligent design breeding models.

Method used

By utilizing the analytical functions of an online analysis platform, significant haplotype-phenotypic relationship data are obtained by scanning the temporary folder in the background. Combined with analysis of variance and t-test information, comprehensive analysis and cross-validation are performed to generate a labeled dataset that can be used to train large language models, thus achieving efficient data aggregation.

Benefits of technology

It significantly improves data collection efficiency, reduces time and cost, and enhances the reliability and applicability of datasets, making it suitable for intelligent design breeding data aggregation for a variety of crops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent design breeding, and particularly discloses a phenotype-haplotype associated annotation data acquisition and convergence method based on an online analysis platform. Comprising the following steps: 1) obtaining process information generated based on variance analysis and t test by scanning a background temporary folder of an analysis platform; 2) obtaining process information according to haplotype association analysis by scanning a background temporary folder of an analysis platform; and 3) carrying out comprehensive analysis and cross validation on the two types of information by using background codes to obtain reliable specific interval / gene haplotype-phenotype associated information, carrying out conversion processing on the information by using the codes, and storing the information as a format data set required by large language model training to complete data convergence. According to the invention, the method collects and gathers annotation data, can achieve the real-time and long-term accumulation of data through the continuous online use of an analysis function of an online analysis platform, and is of great benefit for the iterative upgrade of a trained intelligent design model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent design breeding technology, specifically relating to a method for collecting and aggregating phenotypic-haplotype association annotation data based on an online analysis platform. Background Technology

[0002] With the increasing popularity of artificial intelligence, especially generative artificial intelligence (GPT, DeepSeek, etc.), intelligent design has been put on the breeding agenda. One of the important aspects of intelligent design breeding is model prediction, and the accuracy of model prediction largely depends on the quality of the dataset used (LeCun, Y., Bengio, Y. & Hinton, G. Deeplearning. Nature, 2015 (521), 436–444). A common challenge in training robust models is the lack of large-scale multi-omics data on crops with sufficient data points and sample variability; the aggregation of data related to biological breeding, especially the collection of genotype-phenotype relationship labeled data, has become a bottleneck. Whether it's machine learning or deep learning, model training requires a massive amount of data. Statistics show that 80% of data scientists spend over 40% of their time collecting, preparing, and annotating data, rather than building models (How do Data Professionals Spend their Time on Data Science Projects? By Bob Hayes on February 19, 2019 in Analytics, Data Science, https: / / businessoverbroadway.com / 2019 / 02 / 19 / how-do-data-professionals-spend-their-time-on-data-science-projects / ). Genome selection often stems from genotype-phenotype relationship data; while artificial intelligence based on large language models requires massive amounts of labeled data, especially biologically significant or statistically substantial genotype-phenotype relationship annotations. Currently, acquiring such data primarily relies on offline computation methods, which consume significant additional computing resources and lacks multi-source validation.

[0003] SNP information is the most polymorphic molecular marker in the genome, widely distributed throughout the genome, appearing in both intragenic and intergenic regions. SNP loci on chromosomes exhibit varying degrees of linkage. Some tightly linked SNP loci often represent chromosome segments that are transferred as a whole during natural selection (evolution) and artificial selection (breeding). A small number of genotypes representing these SNP loci can effectively represent the genotype of the entire chromosome segment. Since chromosome segment transfer occurs through gametes, which are haploids, their genotype is called a haplotype (HT). HT variations often have a more direct impact on crop phenotypes, and genomic selection targeting HT is more effective (Front. Plant Sci., 05 September 2023, Volume 14 - 2023 | https: / / doi.org / 10.3389 / fpls.2023.1217589 , Haplotypeblocks for genomic prediction: a comparative evaluation in multiple cropdatasets, Sven E. Weber, Matthias Frisch, Rod J. Snowdon, Kai P. Voss-Fels).

[0004] In 2020, the Institute of Crop Science, Chinese Academy of Agricultural Sciences, and other institutions jointly released the second version (RFGB v2.0) of the Functional Genomics and Breeding (FGB) data platform for rice (https: / / rfgb.rmbreeding.cn). This platform integrates 18M SNPs, 2.3M InDels, 40,000 gene haplotypes, and 12 phenotypic data from 3,024 rice germplasm resources. It provides functions such as gene retrieval, visualization of genomic variation, BLAST, and reconstructed sequence download. In particular, the joint analysis of haplotype and phenotypic data can support the discovery of favorable haplotypes. The release of this platform provides a convenient and practical online tool for rice researchers worldwide to conduct molecular design breeding. In 2022, the FGB paradigm was expanded with the addition of the SoyFGBv2.0 soybean data platform (https: / / sfgb.rmbreeding.cn). Its main features include: First, providing discrete-value phenotypic data to help users identify "useful" germplasm resources for breeding or genetic research, enabling non-download sharing of 33 quantitative traits and 9 qualitative traits from 2214 soybean sequencing resources (2K-SG). Second, utilizing FGB or user-owned phenotypic data to achieve online analysis of the correlation between phenotypic and haplotype variations at different genomic resolutions. Third, through "search" and "browse" modules, users can obtain genomic variations for experimental validation. According to user experience, compared to traditional spreadsheet-assisted haplotype analysis, using FGB for specific gene haplotype analysis can improve efficiency by nearly 60 times. As of March 31, 2025, compared to 6278 domestic and international biological data sharing platforms, FGB is one of the few platforms providing online haplotype analysis, with its rice sub-platform ranking in the top 14% and its soybean sub-platform ranking in the top 60%. From December 2023 to March 31, 2025 alone, the FGB platform provided more than 23,000 online services to researchers in biological breeding, with online analysis accounting for more than 50% of these services. Leveraging the FGB analysis platform, it is expected to aggregate analytical validation from multiple users worldwide, providing important references for haplotype-phenotypic data aggregation. Summary of the Invention

[0005] In response to the aforementioned research background, this invention utilizes the existing analytical functions of online analysis platforms, including the FGB platform, to acquire significant haplotype-phenotypic relationship data pairs by obtaining temporary analysis files, thereby enhancing data collection and aggregation capabilities and providing labeled data for parameter training of large language models. It is primarily applied to the aggregation of biological breeding training data required for the intelligent design of crops such as rice.

[0006] This invention provides a method for collecting and aggregating phenotypic-haplotype association annotation data based on an online analysis platform, which proceeds according to the following steps:

[0007] 1) By scanning the temporary folder in the background of the analysis platform, obtain the following information based on the process information generated by ANOVA and t-test: gene name, gene structural region, chromosome, physical location of SNP variation, haplotype classification, list of samples involved, trait name, significance of ANOVA of haplotype-phenotypic value association, significance of pairwise comparison t-test of phenotypic values ​​corresponding to haplotypes, etc.

[0008] 2) By scanning the temporary folder in the background of the analysis platform, obtain the following information based on the process information of haplotype association analysis: phenotypic code, chromosome, physical interval, haplotype classification, and statistical significance of interval-phenotype association;

[0009] 3) Use the backend code to perform comprehensive analysis and cross-validation on the above two types of information to obtain reliable haplotype-phenotype association information for specific intervals / genes. Then use the code to convert and process the above information and save it as a phenotype-haplotype association labeled dataset in a format required for training large language models, thus completing the data collection and aggregation.

[0010] This invention provides a method for collecting and aggregating phenotypic-haplotype association annotation data based on an online analysis platform, which proceeds according to the following steps:

[0011] 1) By scanning and analyzing the temporary folder in the background of the online analysis platform, the following information is obtained based on the process information generated by ANOVA and t-test: gene name, gene structural region, chromosome, physical location of SNP variation, haplotype classification, list of involved samples, trait name, significance of ANOVA in haplotype-phenotypic value association, and significance of pairwise comparison t-test of phenotypic values ​​corresponding to haplotypes. The obtained data set is called dataset H.

[0012] 2) By scanning the temporary folder in the background of the analysis platform, the following information is obtained based on the process information of haplotype association analysis: phenotypic code, chromosome, physical interval, haplotype classification, and statistical significance of interval-phenotype association. The obtained data set is called dataset G.

[0013] 3) Use the backend code to perform comprehensive analysis and cross-validation on the above two types of information to obtain reliable haplotype-phenotype association information for specific intervals / genes. Then use the code to convert and process the above information and save it as a phenotype-haplotype association labeled dataset in a format required for training large language models, thus completing the data collection and aggregation.

[0014] Specifically, it targets plant varieties, such as crops, including soybeans, rice, corn, wheat, millet, and potatoes.

[0015] More specifically, the construction method of the online analytics platform is as follows:

[0016] (1) DNA extraction and high-throughput genome sequencing were performed on the collected germplasm resources with genetic diversity. (If a read has more than 50% of the bases with a quality value less than 5 or has adapter contamination, it will be filtered out.)

[0017] (2) Align the reads obtained from each sample with the reference genome, generate a BAM format file from the alignment results, and extract SNP information (specifically, the quality control parameters are set as follows: the mapping quality value of each site is greater than 20, the variation quality value is greater than 50, and each base is supported by at least 2 reads, with a MAF value > 0.001). Randomly select the high-quality SNP data subset with the fewest missing data from the extracted SNP dataset, with a total number not exceeding 200K. Calculate the genetic distance matrix of the sequencing germplasm (for example, use a free tool such as TreeBeST to construct a clustering tree to show the kinship between sequencing germplasm, with the boot straps parameter set to 1000).

[0018] (3) Haplotype information extraction and database platform construction.

[0019] In a specific implementation, a database module is built to complete the above analysis. For example, a web server is built based on tools such as Apache or Nginx, phenotypic data and gene annotation information are imported into databases such as MySQL, and the website interface is implemented using PHP or JavaScript.

[0020] More specifically, for soybeans, the Soy_Haplotype module of the soybean FGB data platform is used directly, with the URL https: / / sfgb.rmbreeding.cn / analysis / haplotype.

[0021] Specifically, in step 1), the following information is obtained based on the process information generated by analysis of variance and t-test: gene name, gene structural region, chromosome, physical location of SNP variation, haplotype classification, sample list, trait name, significance of analysis of variance for haplotype-phenotypic value association (ANOVA_Sig), and significance of pairwise comparison t-test of phenotypic values ​​corresponding to haplotypes (Sig_ttest). The obtained data set is called dataset H.

[0022] Specifically, in step 2), the haplotype association analysis process obtains the following information: trait code, chromosome, physical interval (Start, End), and statistical significance of interval-phenotype association (corrected p-value, p.value.adj). The obtained data set is called dataset G.

[0023] Furthermore, step 3) data aggregation involves taking the intersection of the above datasets H and G through gene IDs using code, that is, obtaining the gene IDs and related information detected under both analysis methods, forming the data intersection C;

[0024] The data in the intersection C will be a significant subset D in both types of analysis, which will be used to generate labeled data E, including traits, haplotypes, chromosomes, start physical location, end physical location, and whether cross-validation is passed, which can then be used for model training.

[0025] The method of this invention can be applied to the aggregation of biological breeding data required for intelligent crop design, specifically, the crop being rice, corn, or wheat.

[0026] Compared with existing technologies, this invention has the following advantages and effects: By utilizing information mining from the online analysis process based on an online analysis platform, it reuses data from automatically generated process files (temporary files); it achieves the aggregation of crop haplotype-phenotype association data without requiring additional de novo association analysis. Compared with methods that involve complete de novo association processing, it significantly improves the efficiency of acquiring labeled data, saving time and costs. For different analysis modules of the online analysis platform, including various analysis methods such as ANOVA / t-tests and association analysis, analyzing the same gene / haplotype and trait, it is expected to achieve cross-validation between different analysis methods and data sources, improving the reliability of the dataset; this method is applicable to database platforms providing online analysis and has relatively broad universality.

[0027] In summary, by collecting and aggregating labeled data through this invention, and through the continuous online use of the analysis function of the online analysis platform, the real-time long-term accumulation of data can be achieved, which is of great benefit to the iterative upgrading of the trained intelligent design model. Attached Figure Description

[0028] Figure 1 A comparison diagram of the data acquisition process for intelligent crop design annotation based on the FGB platform and traditional methods. Detailed Implementation

[0029] The invention is further illustrated below with specific implementation examples. Unless otherwise specified, all methods used are conventional methods. The following examples are not intended to limit the invention in any way.

[0030] (I) Genomic Information Acquisition

[0031] 1. Test materials

[0032] Leaf samples of soybean germplasm resources (including varieties and breeding parents) with genetic diversity were collected for genome sequencing, hereinafter referred to as sequencing germplasm.

[0033] 2. DNA extraction and high-throughput genome sequencing

[0034] Following the DNA extraction method described in Temnykh, S., Park, W., Ayres, N. et al. Mapping and genome organization of microsatellite sequences in rice (Oryza sativa L.). Theor Appl Genet 100, 697–712 (2000), or using a kit to further improve purity, genomic DNA was extracted from homozygous lines for sequencing. Considering cost, shot-gun sequencing technology can be used for genome sequencing, with library preparation and sequencing methods following standard procedures. A coverage of 10X or higher is recommended for obtaining high-quality data. To ensure the quality of sequencing data, reads with more than 50% of their bases having a quality value less than 5 or containing adapter contamination should be filtered out and discarded.

[0035] (II) SNP Information Extraction and Breeding Parent Cluster Analysis

[0036] Given that it is necessary to maintain the diversity of sequencing germplasm as much as possible, we need to have a basic understanding of the phylogenetic relationships of sequencing germplasm.

[0037] Based on the soybean genomic DNA sequencing data described above, we compared the reads obtained from each sample with the soybean reference genome (e.g., GmaxW82.a2) using free analysis tools such as BWA, generating BAM format files from the alignment results. Based on the BAM files, we extracted SNP information using free analysis tools such as the Genome Analysis Toolkit (GATK). To improve the reliability of SNP information extraction, the quality control parameters were set as follows: mapping quality value greater than 20, variation quality value greater than 50 for each site, and each base supported by at least two reads, with a MAF value > 0.001. From the extracted SNP dataset, a high-quality SNP subset with the fewest missing data was randomly selected, totaling no more than 200K, for the next step of sequencing germplasm clustering analysis.

[0038] Based on the aforementioned high-quality SNP data subsets, the genetic distance matrix of the sequencing germplasms was calculated. Free tools such as TreeBeST were used to construct cluster trees to show the phylogenetic relationships between the sequencing germplasms, with the boot straps parameter set to 1000.

[0039] (III) Haplotype Information Extraction and Database Platform Construction

[0040] Based on the GFF3 annotation file for the genome provided by GmaxW82.a2 (which contains information on the physical location of genes on chromosomes and gene structure), SNPs in specific gene regions (coding regions, promoter regions, or untranslated regions can be selected as appropriate) are extracted to obtain SNP data with high polymorphism (MAF value > 0.05). The non-synonymous variants formed by combining all SNPs within a specific gene region are defined as haplotype (HT) variants within that region. Samples are grouped by type according to HT variants, thereby obtaining SNP information for each HT variant type, a list of sequencing samples carrying that HT type, and phenotypic statistics of the sequencing samples.

[0041] To facilitate application, a database module can be built to complete the above analysis. A web server can be set up using tools such as Apache or Nginx, and phenotypic data and gene annotation information can be imported into databases such as MySQL. The website interface can be implemented using PHP or JavaScript. This module supports three input formats: gene ID, chromosomal region, and SNP site. Users can select different gene regions, set MAF filters, and submit candidate sequencing germplasm lists to achieve personalized online analysis based on their research results.

[0042] (iv) Haplotype Analysis Data Acquisition

[0043] The Soy_Haplotype module (https: / / sfgb.rmbreeding.cn / analysis / haplotype) of the soybean FGB data platform created according to the above steps already provides online analysis. The following steps can be implemented directly through the Soy_Haplotype module.

[0044] Strengthen the management of temporary files generated by the analysis module so that the following information can be obtained by scanning the temporary folder of the FGB analysis platform module and based on the process information generated by Haplotype analysis: Gene name, gene structural region, chromosome, physical location of SNP variation (Position) (Table 1); Haplotype classification, sample list (Table 2); Trait name, significance of ANOVA-Sig association between haplotype and phenotypic value (Table 3); significance information of pairwise comparison t-test of phenotypic values ​​corresponding to haplotypes (Sig_ttest) (Table 4).

[0045] Table 1. Examples of Haplotype analysis genotype (.geno) data formats obtained through the FGB platform

[0046]

[0047] Table 2. Examples of Haplotype (.haplo) data format obtained through the FGB platform for Haplotype analysis.

[0048]

[0049] Table 3. Example of ANOVA significance data format obtained from the FGB platform for haplotype-phenotypic value association analysis using the FGB platform.

[0050]

[0051] The data set obtained through these steps is called dataset H.

[0052] Table 4. Example of data format for pairwise comparison t-test significance data of haplotype-corresponding phenotypic values ​​obtained through the FGB platform for Haplotype analysis.

[0053]

[0054] (v) Acquisition of haplotype-based genome association analysis (Hap-GWAS analysis) data

[0055] The following can be accomplished using the haplotype genome association analysis module of an online platform such as Soybean FGB (https: / / sfgb.rmbreeding.cn / analysis / hapGwas), created according to the steps described above. By scanning the temporary folder of the FGB platform module, information obtained based on the Hap-GWAS analysis process information includes: trait code, chromosome (Chr), physical interval (Start, End), and statistical significance of interval-phenotype association (p.value.adj). The collection of data obtained through these steps is called dataset G.

[0056] Table 5. Examples of data formats for haplotype-phenotypic association statistical analysis obtained through the FGB platform in Hap-GWAS analysis.

[0057]

[0058] (vi) Generation of labeled data

[0059] By comprehensively analyzing and cross-validating the above two types of information, relatively reliable haplotype-phenotype association information for specific intervals / gene haplotypes can be obtained. This information is then transformed using code and saved as a dataset in a format suitable for training large language models. Specifically, the intersection of datasets H and G can be obtained through code, using gene IDs to capture the gene IDs and related information detected by both analysis methods, forming the data intersection C. The data in intersection C will be a significant subset D in both analyses, used to generate labeled data E, including traits, haplotypes, chromosomes, start physical location, end physical location, and whether cross-validation was successful. This data can be used as specific interval / gene haplotype-phenotype association annotation data, which can be provided to crop intelligent design based on technologies such as large language models for model parameter training.

[0060] Table 6. Examples of data formats that can be used for training crop intelligent design models

[0061]

[0062] Traditional methods for acquiring training data require collecting multi-year, multi-location phenotypic data from sequencing germplasm from scratch, obtaining genotype data through sequencing, and performing massive analysis (for soybean, the computational complexity is 40,000 genes * N haplotypes * M phenotypes). In contrast, this invention only requires starting with several temporary files generated by online analysis platforms such as FGB (e.g., two sets of temporary files for Haplotype analysis and Hap_GWAS analysis) for data scanning. Since users typically perform haplotype analysis on genes with well-defined functions, this significantly reduces the number of genes required for haplotype analysis, making the target more specific and significantly reducing computational complexity. Figure 1 For online analysis platforms like FGB, only routine user calculations need to be maintained. The main computational load is in the subsequent data integration and labeling, which is also essential for traditional methods.

[0063] In summary, the method described in this invention can fully utilize the temporary file information of online analysis platforms such as FGB to improve the efficiency of phenotypic-haplotype association data collection. This method can be applied to the aggregation of bio-breeding data required for intelligent design of rice and other crops.

Claims

1. A method for collecting and aggregating phenotypic-haplotype association labeled data based on an online analysis platform, characterized in that, Follow these steps: 1) By scanning and analyzing the temporary folder in the background of the online analysis platform, the following information is obtained based on the process information generated by ANOVA and t-test: gene name, gene structural region, chromosome, physical location of SNP variation, haplotype classification, list of involved samples, trait name, significance of ANOVA in haplotype-phenotypic value association, and significance of pairwise comparison t-test of phenotypic values ​​corresponding to haplotypes. The obtained data set is called dataset H. 2) By scanning the temporary folder in the background of the analysis platform, the following information is obtained based on the process information of haplotype association analysis: phenotypic code, chromosome, physical interval, haplotype classification, and statistical significance of interval-phenotype association. The obtained data set is called dataset G. 3) Use the backend code to perform comprehensive analysis and cross-validation on the above two types of information to obtain reliable haplotype-phenotype association information for specific intervals / genes. Then use the code to convert and process the above information and save it as a phenotype-haplotype association labeled dataset in a format required for training large language models, thus completing the data collection and aggregation.

2. The data acquisition and aggregation method as described in claim 1, characterized in that, It targets plant varieties, such as crops, specifically soybeans, rice, corn, wheat, millet, and potatoes.

3. The data acquisition and aggregation method as described in claim 1, characterized in that, The construction method of the online analysis platform is as follows: (1) DNA extraction and high-throughput genome sequencing were performed on the collected germplasm resources with genetic diversity; (Reads with more than 50% base quality values ​​less than 5 or adapter contamination in the raw data were filtered out and discarded.) (2) Align the reads obtained from each sample with the reference genome, generate a BAM format file from the alignment results, and extract SNP information (specifically, the quality control parameters are set as follows: the mapping quality value of each site is greater than 20, the variation quality value is greater than 50, and each base is supported by at least 2 reads, with a MAF value > 0.001). Randomly select the high-quality SNP data subset with the fewest missing data from the extracted SNP dataset, with a total number not exceeding 200K. Calculate the genetic distance matrix of the sequencing germplasm (for example, use a free tool such as TreeBeST to construct a clustering tree to show the kinship between sequencing germplasm, with the boot straps parameter set to 1000). (3) Haplotype information extraction and database platform construction.

4. The data acquisition and aggregation method as described in claim 1, characterized in that, To complete the above analysis, a database module is built. For example, a web server is built based on tools such as Apache or Nginx, phenotypic data and gene annotation information are imported into databases such as MySQL, and the website interface is implemented using PHP or JavaScript. More specifically, for soybeans, the Soy_Haplotype module of the soybean FGB data platform is used directly, with the URL https: / / sfgb.rmbreeding.cn / analysis / haplotype.

5. The data acquisition and aggregation method as described in claim 1, characterized in that, Step 1) Based on the process information generated by ANOVA and t-test, obtain the following information: gene name, gene structural region, chromosome, physical location of SNP variation, haplotype classification, sample list, trait name, significance of ANOVA for haplotype-phenotypic value association (ANOVA_Sig), and significance of pairwise comparison t-test of phenotypic values ​​corresponding to haplotypes (Sig_ttest). The obtained data set is called dataset H.

6. The data acquisition and aggregation method as described in claim 1, characterized in that, In step 2), the following information is obtained during the haplotype association analysis process: trait code, chromosome (Chr), physical interval (Start, End), and statistical significance of interval-phenotype association (corrected p-value, p.value.adj). The obtained data set is called dataset G.

7. The data acquisition and aggregation method as described in claim 1, characterized in that, Step 3) Data aggregation involves taking the intersection of the above datasets H and G through gene IDs using code. This means obtaining the gene IDs and related information detected by both analysis methods and forming the data intersection C. The data in the intersection C will be a significant subset D in both types of analysis, which will be used to generate labeled data E, including traits, haplotypes, chromosomes, start physical location, end physical location, and whether cross-validation is passed, which can then be used for model training.

8. The method of any one of claims 1 to 7 is applied in the aggregation of bio-breeding data required for intelligent crop design, wherein the crop is rice, corn, or wheat.

Citation Information

Patent Citations

  • Method for merging linkage disequilibrium SNPs in association analysis results in batches

    CN117292746A

  • Analysis method, system and device based on single cell whole transcriptome sequencing data

    CN118866127A

  • Genome selection method and system based on natural language processing

    CN120412727A

  • Method for detecting single nucleotide polymorphism susceptibility that is target of drug resistance in ankylosing spondylitis

    TW202523849A

  • System and method for analyzing genotype using genetic variation information on individual's genome

    US20190087540A1