A method and system for genetic testing management

CN120089195BActive Publication Date: 2026-08-21CHERRY VALLEY BREEDING TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510179321.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2026-08-21
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

[0004]本发明解决的技术问题是:现有技术不能做到多层次特征提取,不能同时包括基因组DNA、RNA、蛋白质、miRNA和lncRNA多种生物信息的提取和分析,无法提升基因检测的深度,无法将基因-疾病数据库与优化特征选择实现现实结合,识别出与疾病关联性最强的基因变异特征,无法通过连续的特征选择、回归模型训练和验证过程使基因-疾病数据库持续更新和精确化,提升基因与疾病之间关联的预测水平,无法通过监测数据库和验证集的差异性,生成概率分布图,不能确保随着时间的推移令基因-疾病数据库稳定收敛到一个准确且可信的水平

Benefits of technology

[0015]本发明的有益效果:多层次特征提取,同时包括基因组DNA、RNA、蛋白质、miRNA和lncRNA多种生物信息的提取和分析,提升基因检测的深度,将基因-疾病数据库与优化特征选择实现现实结合,识别出与疾病关联性最强的基因变异特征,通过连续的特征选择、回归模型训练和验证过程使基因-疾病数据库持续更新和精确化,提升基因与疾病之间关联的预测水平,通过监测数据库和验证集的差异性,生成概率分布图,确保随着时间的推移令基因-疾病数据库稳定收敛到一个准确且可信的水平,不仅能够对种鸭的基因信息进行详细分析,还能基于大数据和机器学习对潜在疾病风险进行精准预测,提高早期诊断的准确性,还能够根据分析结果,自动生成个性化的位点变异分析报告,为育种专家提供准确的选种依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089195B_ABST
    Figure CN120089195B_ABST
Patent Text Reader

Abstract

The application discloses a kind of gene detection management method and system, including the complete process from biological sample collection, biological sample processing, gene detection, data analysis to result report, realize multilevel feature extraction, simultaneously including the extraction and analysis of multiple biological information of genomic DNA, RNA, protein, miRNA and lncRNA, improve the depth of gene detection, realize real combination of gene-disease database and optimized feature selection, identify the strongest gene variation feature associated with disease, continuously update and refine gene-disease database through continuous feature selection, regression model training and verification process, improve the prediction level of the association between genes and diseases, generate probability distribution map by monitoring the difference between database and validation set, ensure that gene-disease database converges to an accurate and reliable level over time, automatically generate personalized site variation analysis report according to the analysis result, and provide accurate selection basis for breeding experts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of gene feature extraction, and in particular to a gene detection and management method and system. Background Technology

[0002] Genetic testing technology has made significant progress in fields such as medicine, agriculture, and the environment, showing great potential, especially in disease prevention, precision medicine, and personalized treatment. However, with the widespread development of genomics research and clinical applications, how to efficiently, safely, and reliably manage genetic testing data, ensuring data accuracy, privacy protection, and the effectiveness of clinical applications, has become an urgent problem to be solved. Therefore, genetic testing management methods and systems have emerged, with the main purpose of building a complete solution for data management, analysis, storage, transmission, and security protection that meets clinical testing needs while also complying with ethical, legal, and technical standards.

[0003] Currently, Chinese invention patent application number CN202310956410.2 discloses a method and system for managing information in a gene testing laboratory. This method collects relevant information from gene testing; associates and encrypts the relevant information before uploading it to a blockchain for traceability and querying; and constructs a network security situation model to evaluate the blockchain network. This application not only provides complete data support for traceability and querying but also ensures the confidentiality, privacy, and security of gene testing information. However, existing technologies cannot achieve multi-level feature extraction, cannot simultaneously extract and analyze multiple biological information such as genomic DNA, RNA, protein, miRNA, and lncRNA, cannot improve the depth of gene testing, cannot realistically combine gene-disease databases with optimized feature selection to identify the gene variants most strongly associated with diseases, cannot continuously update and refine the gene-disease database through continuous feature selection, regression model training, and validation processes to improve the predictive level of the association between genes and diseases, cannot generate probability distribution maps by monitoring the differences between the database and the validation set, and cannot ensure that the gene-disease database stably converges to an accurate and reliable level over time. Summary of the Invention

[0004] The technical problem solved by this invention is that existing technologies cannot achieve multi-level feature extraction, cannot simultaneously extract and analyze multiple biological information such as genomic DNA, RNA, protein, miRNA, and lncRNA, cannot improve the depth of gene detection, cannot realistically combine gene-disease databases with optimized feature selection to identify gene variant features with the strongest association with diseases, cannot continuously update and refine the gene-disease database through continuous feature selection, regression model training, and validation processes to improve the predictive level of the association between genes and diseases, cannot generate probability distribution maps by monitoring the differences between the database and the validation set, and cannot ensure that the gene-disease database stably converges to an accurate and reliable level over time.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a gene detection management method, comprising the following steps: Step S1: Collect biological samples using intelligent collection equipment and record detailed information about the biological samples; Step S2: Automated preprocessing of the extracted biological sample data; Step S3: Using sequencing technology and algorithms, the genes in the biological sample are sequenced, the quality of the sequencing data is monitored in real time, and the data quality is controlled. Step S4: Utilize big data analytics and machine learning techniques to deeply mine and analyze sequencing data, identifying disease-related gene variations and patterns; Step S5: Automatically generate personalized test reports based on disease-related gene variations and patterns.

[0006] Preferably, step S1 includes: The intelligent collection device includes an intelligent sampling instrument, which includes a mobile blood drawer; The biological samples include blood and feather tissue samples. During the collection process, detailed information about the biological samples is recorded, including sampling time, sampling location, breeding duck identification information, and biological sample type. The collected biological samples are recorded using a unique barcode, which is automatically synchronized to the hospital or laboratory's information management system.

[0007] Preferably, step S2 includes: The collected biological samples undergo preprocessing steps involving extraction and purification using automated equipment, and key features are extracted from the preprocessed data. The extracted key features are ligated into sequencing adapters to construct sequencing libraries.

[0008] Preferably, extracting key features includes: Genomic DNA is extracted from cells using chemical reagents, including CTAB and a salt solution. Gene regions are amplified using PCR, and the sequence information of the genome is obtained. The sequence information includes the genetic information of the organism, including genes encoding proteins, promoters of regulatory genes, and terminator regions. Single nucleotide polymorphisms (SNPs) were detected using high-throughput sequencing technology. These SNPs characterize gene variations among individuals, identify mutated gene regions, and extract histone modification and transcription factor binding regions. DNA methylation status was detected by sulfite sequencing; Total RNA was extracted from the sample using TRIzol, mRNA was extracted using magnetic beads, rRNA and tRNA were removed, the expression levels of several genes were detected by real-time quantitative PCR, miRNA was extracted from cells or tissues using a specialized miRNA extraction kit, miRNA expression and differential expression were extracted using miRNA microarray technology, and lncRNA expression patterns were obtained using transcriptome sequencing technology. Total protein was extracted from cells or tissues using lysis buffer, protein expression levels were detected by Western blotting, proteins were separated and identified by liquid chromatography-mass spectrometry, protein functions, modifications and interactions were annotated using proteomics databases, and post-translational modifications were analyzed by mass spectrometry to obtain protein functions and regulatory mechanisms. Detection of cell surface markers using antibody labeling; The extracted genome sequence information, variant gene regions, histone modification and transcription factor binding regions, DNA methylation status, expression levels of several genes, miRNA expression and differences, lncRNA expression patterns, protein expression levels, protein function and regulatory mechanisms, and cell markers are used as key data quantities and input into a neural network to extract key features of biological samples in a unified data format. These key features are then set into a key feature set of biological samples.

[0009] Preferably, step S3 includes: DNA sequences are read using the Illumina high-throughput sequencing platform, and a data quality report is automatically generated. The data quality report includes the quality Q score based on each reading and the distribution of the quality Q score, GC content, and statistical information on repetitive sequences and short gene sequences. FastQC was used to remove duplicates, trim and filter DNA sequences, removing short gene sequences with Q scores below a preset threshold. Picard was used to remove duplicate short gene sequences. The alignment tool BWA was used to align the sequencing results with a pre-acquired reference genome. Based on the alignment results, GATK was used to call up variants and identify SNPs and InDels. Multiple sequencing was performed in the regions of variant genes and regions of histone modification and transcription factor binding.

[0010] Preferably, big data is used to collect the association between gene mutations or hereditary traits corresponding to diseases to obtain a gene-disease database. The gene-disease database includes a first disease, a second disease, ... and an nth disease, where n is a constant. The pointers of each disease address point to the associated genes included in the disease. Each gene address stores the association between the gene and the corresponding disease in the form of a dictionary. Feature selection is optimized using ant colony optimization algorithm. Combined with the gene-disease database, the gene variant features with the strongest association with the disease and the corresponding feature correlation are obtained. A regression model is used to establish a quantitative relationship between the strongest gene variant features and the corresponding feature correlation and the disease risk. The gene-disease database is updated according to the quantitative relationship.

[0011] Preferably, the process of selecting the most representative features of the disease from the set of key features of the biological samples using the ant colony algorithm, and reducing the data dimensionality, includes: Based on the gene-disease database, a genotype feature matrix is ​​constructed according to the genotype of the disease. The row vectors of the genotype feature matrix represent the disease type, and the column vectors represent the gene types associated with the disease type. The node data of the genotype feature matrix is ​​represented by one-hot encoding. The node data corresponding to the genes included in the mutated gene region are uniformly overlapped and marked as mutated encoding data that does not repeat with all one-hot encodings. The initialization logic includes constructing an initial genotype feature matrix based on a gene-disease database crawled from big data, setting the node data in the initial genotype feature matrix as ant colony parameters, using non-zero node data as initial gene feature values, setting the disease correlation of the initial gene feature values ​​in the gene-disease database as heuristic information, and finding the most correlated gene feature value based on the heuristic information to obtain the most correlated gene type corresponding to the disease. All node data marked with variant coding data are extracted as variant gene feature values. The disease association of the variant gene feature values ​​in the gene-disease database is set as heuristic variant information. The most associated gene feature value is found according to the heuristic variant information, and the gene type corresponding to the variant gene feature value with the greatest pathogenic potential corresponding to the disease is obtained. The process of finding the most relevant gene trait values ​​includes: Each gene feature value is assigned a feature vector. A regression model is trained using selected features to calculate the performance of the feature vectors. The performance is represented by the mean squared error evaluation index. The probability of feature selection is automatically updated based on the performance. The probability is repeatedly updated until it can no longer be updated, and the optimal feature vector is obtained and output.

[0012] Preferably, the most relevant gene type corresponding to the disease selected by the ant colony algorithm and the gene type corresponding to the most pathogenic variant gene feature value of the disease are used as the first training set and the second validation set. A linear regression model is trained based on the first training set to obtain a gene-disease prediction model. The gene-disease prediction model is used to predict the second validation set to obtain the first variant gene prediction of the disease. The first variant gene prediction is added to the gene-disease database to update the gene-disease database. An updated second validation set is obtained based on the updated gene-disease database. The gene-disease database and the second validation set are iteratively updated. The difference between the gene-disease database and the second validation set during the update process is calculated according to the time series, and the difference is plotted as a probability distribution map. The probability distribution map should converge over time. When the probability distribution at time t converges to below a preset convergence threshold, the updated gene-disease database is obtained. The risk of recessive gene inheritance and variation in breeding ducks due to inbreeding was assessed by predicting the association between genes and diseases in a gene-disease database.

[0013] Preferably, step S5 includes: When a breeding duck biological sample is input, the system searches the gene-disease database for the corresponding genotype-related disease type extracted from the biological sample. It then uses a linear regression model to predict the probability of the disease and automatically generates a personalized test report. This provides doctors with accurate diagnostic information. The test report includes a detailed description of the gene mutation, the potential health risks of the mutation, the susceptibility probability to related diseases, and drug response.

[0014] A gene detection management system includes an acquisition module, a processing module, a sequencing module, a data analysis module, and a reporting module. The acquisition module is used to collect biological samples through intelligent acquisition devices and record detailed information about the biological samples. The processing module is used to automatically perform preprocessing steps on the extracted biological sample data; The sequencing module is used to sequence genes in biological samples using sequencing technology and algorithms, monitor the quality of sequencing data in real time, and perform data quality control. The data analysis module is used to perform in-depth mining and analysis of sequencing data using big data analysis and machine learning technologies to identify disease-related gene variations and patterns. The reporting module is used to automatically generate personalized test reports based on disease-related gene variations and patterns.

[0015] The beneficial effects of this invention are as follows: multi-level feature extraction, including the extraction and analysis of multiple biological information such as genomic DNA, RNA, protein, miRNA, and lncRNA, enhances the depth of gene detection, realistically combines gene-disease database with optimized feature selection, identifies gene variation features with the strongest association with diseases, and continuously updates and refines the gene-disease database through continuous feature selection, regression model training, and validation processes, improving the predictive level of the association between genes and diseases. By monitoring the differences between the database and the validation set, a probability distribution map is generated to ensure that the gene-disease database stably converges to an accurate and reliable level over time. It can not only perform detailed analysis of the genetic information of breeding ducks, but also accurately predict potential disease risks based on big data and machine learning, improving the accuracy of early diagnosis. Furthermore, it can automatically generate personalized site variation analysis reports based on the analysis results, providing breeding experts with accurate selection criteria. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the basic process of a gene detection management method provided in one embodiment of the present invention. Detailed Implementation

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] Reference Figure 1 As an embodiment of the present invention, a gene detection management method is provided, comprising the following steps: Step S1: Collect biological samples using intelligent collection equipment and record detailed information about the biological samples; Step S2: Automated preprocessing of the extracted biological sample data; Step S3: Using advanced sequencing technologies and algorithms, the genes in the biological sample are sequenced, the quality of the sequencing data is monitored in real time, and the data quality is controlled to ensure the reliability of the sequencing results; Step S4: Utilize big data analytics and machine learning techniques to deeply mine and analyze sequencing data, identifying disease-related gene variations and patterns; Step S5: Automatically generate personalized test reports based on disease-related gene variations and patterns.

[0019] Step S1 includes: The intelligent collection device includes an intelligent sampling instrument, which includes a moving blood drawer to ensure the efficiency and accuracy of sampling. The biological samples include blood and feather tissue samples. During the collection process, detailed information about the biological samples is recorded, including sampling time, sampling location, breeding duck identification information, and biological sample type. The collected biological samples are recorded with unique barcodes to ensure that each biological sample corresponds one-to-one with the breeding duck information, avoiding incorrect biological sample labeling. The barcodes are automatically synchronized to the hospital or laboratory's information management system to ensure data integrity and traceability.

[0020] The purpose of genetic testing is to understand the genetic background, kinship, and disease resistance of breeding ducks through scientific means.

[0021] Step S2 includes: The preprocessing steps of collecting biological samples are carried out through automated equipment for extraction and purification, and key features are extracted from the preprocessed data. Automated processing improves processing efficiency and accuracy. According to the requirements of the sequencing platform, the extracted key features are ligated with sequencing adapters to construct a sequencing library.

[0022] Key features extracted include: Genomic DNA is extracted from cells using chemical reagents, including CTAB and a salt solution. Gene regions are amplified using PCR, and the sequence information of the genome is obtained. The sequence information includes the genetic information of the organism, including genes encoding proteins, promoters of regulatory genes, and terminator regions. Single nucleotide polymorphisms (SNPs) were detected using high-throughput sequencing technology. These SNPs characterize gene variations among individuals, identify mutated gene regions, and extract histone modification and transcription factor binding regions. DNA methylation status was detected by sulfite sequencing; Total RNA was extracted from the sample using TRIzol, mRNA was extracted using magnetic beads, rRNA and tRNA were removed, and the expression levels of several genes, including user-specified gene types, were detected by real-time quantitative PCR. miRNA was extracted from cells or tissues using a specialized miRNA extraction kit, and the expression and differential expression of miRNA were extracted using miRNA microarray technology. The expression pattern of lncRNA was obtained using transcriptome sequencing technology. Total protein was extracted from cells or tissues using lysis buffer, protein expression levels were detected by Western blotting, proteins were separated and identified by liquid chromatography-mass spectrometry, protein functions, modifications and interactions were annotated using proteomics databases, and post-translational modifications were analyzed by mass spectrometry to obtain protein functions and regulatory mechanisms. Detection of cell surface markers using antibody labeling; The extracted genome sequence information, variant gene regions, histone modification and transcription factor binding regions, DNA methylation status, expression levels of several genes, miRNA expression and differences, lncRNA expression patterns, protein expression levels, protein function and regulatory mechanisms, and cell markers are used as key data quantities and input into a neural network to extract key features of biological samples in a unified data format. These key features are then set into a key feature set of biological samples.

[0023] Step S3 includes: DNA sequences are read using the Illumina high-throughput sequencing platform, and a data quality report is automatically generated. The data quality report includes the quality Q score based on each reading and the distribution of the quality Q score, GC content, and statistical information on repetitive sequences and short gene sequences. FastQC was used to remove duplicates, trim, and filter DNA sequences, eliminating short gene sequences with Q scores below a preset threshold. Due to PCR amplification bias, sequencing results may contain repetitive sequences. Using the Picard tool to remove repetitive short gene sequences helps improve data quality. The BWA alignment tool was used to align the sequencing results with a pre-acquired reference genome to ensure data accuracy. Based on the alignment results, GATK was used to call variants, identifying SNPs and InDels. Multiple sequencing was performed in the regions of variant genes and regions of histone modifications and transcription factor binding to reduce incidental errors.

[0024] Employing advanced sequencing technologies and algorithms to sequence genes in biological samples, and monitoring the quality of sequencing data in real time, is a crucial step in ensuring the reliability of sequencing results. By using high-throughput sequencing platforms, coupled with real-time quality monitoring tools and stringent quality control strategies, sequencing errors can be minimized, ensuring that data quality meets the standards for research or clinical analysis.

[0025] By using big data to collect the associations of gene mutations or hereditary traits corresponding to diseases, a gene-disease database is obtained. The gene-disease database includes a first disease, a second disease, ... and an nth disease, where n is a constant. The pointers of each disease address point to the associated genes included in the disease. Each gene address stores the association between the gene and the corresponding disease in the form of a dictionary. Feature selection is optimized using ant colony optimization algorithm. Combined with the gene-disease database, the gene variant features with the strongest association with the disease and the corresponding feature correlation are obtained. A regression model is used to establish a quantitative relationship between the strongest gene variant features and the corresponding feature correlation and the disease risk. The gene-disease database is updated according to the quantitative relationship.

[0026] The process of selecting the most representative features of the disease from the key feature set of the biological samples using the ant colony algorithm, and reducing data dimensionality, includes: Gene data is usually high-dimensional and requires data preprocessing first. Based on the gene-disease database, a genotype feature matrix is ​​constructed based on the genotype of the disease. The row vectors of the genotype feature matrix represent the disease type, and the column vectors represent the gene types associated with the disease type. The node data of the genotype feature matrix is ​​represented by one-hot encoding. The node data corresponding to the genes included in the mutated gene region are uniformly overlapped and marked as mutated encoding data that does not repeat all one-hot encodings. The initialization logic includes constructing an initial genotype feature matrix based on a gene-disease database crawled from big data, setting the node data in the initial genotype feature matrix as ant colony parameters, using non-zero node data as initial gene feature values, setting the disease correlation of the initial gene feature values ​​in the gene-disease database as heuristic information, and finding the most correlated gene feature value based on the heuristic information to obtain the most correlated gene type corresponding to the disease. All node data marked with variant coding data are extracted as variant gene feature values. The disease association of the variant gene feature values ​​in the gene-disease database is set as heuristic variant information. The most associated gene feature value is found according to the heuristic variant information, and the gene type corresponding to the variant gene feature value with the greatest pathogenic potential corresponding to the disease is obtained. The process of finding the most relevant gene trait values ​​includes: Each gene feature value is assigned a feature vector. A regression model is trained using selected features to calculate the performance of the feature vectors. The performance is represented by the mean squared error evaluation index. The probability of feature selection is automatically updated based on the performance. The probability is a weight automatically calculated by the computer. The probability is repeatedly updated until it can no longer be updated, and the optimal feature vector is obtained and output.

[0027] The most relevant gene type for the disease selected by the ant colony algorithm and the gene type corresponding to the most pathogenic variant gene feature value for the disease are used as the first training set and the second validation set. A linear regression model is trained based on the first training set to obtain a gene-disease prediction model. The gene-disease prediction model is used to predict the second validation set to obtain the first variant gene prediction for the disease. The first variant gene prediction is added to the gene-disease database to update the gene-disease database. An updated second validation set is obtained based on the updated gene-disease database. The gene-disease database and the second validation set are iteratively updated. The differences between the gene-disease database and the second validation set during the update process are calculated according to the time series, and the differences are plotted as a probability distribution map. The probability distribution map should converge over time. When the probability distribution at time t converges to below a preset convergence threshold, the updated gene-disease database is obtained. The association data between genes and diseases is added, and the association level between genes and diseases is refined, especially increasing the disease prediction level of variant genes. The risk of recessive gene inheritance and variation in breeding ducks due to inbreeding was assessed by predicting the association between genes and diseases in gene-disease databases. For example, the association between BRCA1 and BRCA2 gene mutations and breast cancer, or the relationship between the APOE gene and Alzheimer's disease.

[0028] Step S5 includes: When a breeding duck biological sample is input, the system searches the gene-disease database for the corresponding genotype-related disease type extracted from the biological sample. It then uses a linear regression model to predict the probability of the disease and automatically generates a personalized test report. This provides doctors with accurate diagnostic information. The test report includes a detailed description of the gene mutation, the potential health risks of the mutation, the susceptibility probability to related diseases, and drug response.

[0029] The generation of reports needs to be understood by both doctors and breeding ducks. It typically includes concise charts, risk warnings, and big data searches to obtain the necessary medical interpretations to help doctors make decisions.

[0030] A gene detection management system includes an acquisition module, a processing module, a sequencing module, a data analysis module, and a reporting module. The acquisition module is used to collect biological samples through intelligent acquisition devices and record detailed information about the biological samples. The processing module is used to automatically perform preprocessing steps on the extracted biological sample data; The sequencing module is used to sequence genes in biological samples using sequencing technology and algorithms, monitor the quality of sequencing data in real time, and perform data quality control. The data analysis module is used to perform in-depth mining and analysis of sequencing data using big data analysis and machine learning technologies to identify disease-related gene variations and patterns. The reporting module is used to automatically generate personalized test reports based on disease-related gene variations and patterns.

[0031] The gene detection and management system is used to determine whether there is excessive inbreeding in the population. The judgment logic includes: calculating the total probability mean of recessive hereditary traits and variations caused by inbreeding; when the preset genetic defect probability threshold is reached, artificial intervention in mating is carried out to prevent genetic defect problems caused by inbreeding. The gene detection management system is used to identify preferred disease-resistant genes through gene testing, thereby selecting breeding ducks with strong disease resistance to improve the disease resistance of the population. It can also help predict the production performance of ducks and improve economic benefits.

[0032] This invention achieves multi-level feature extraction, including the extraction and analysis of various biological information such as genomic DNA, RNA, protein, miRNA, and lncRNA, enhancing the depth of gene detection. It realistically combines gene-disease databases with optimized feature selection, identifying gene variants with the strongest disease association. Through continuous feature selection, regression model training, and validation processes, the gene-disease database is continuously updated and refined, improving the predictive level of the association between genes and diseases. By monitoring the differences between the database and the validation set, a probability distribution map is generated, ensuring that the gene-disease database stably converges to an accurate and reliable level over time. It can not only perform detailed analysis of the genetic information of breeding ducks, but also accurately predict potential disease risks based on big data and machine learning, improving the accuracy of early diagnosis. By combining the gene variants of breeding ducks, it can also automatically generate personalized site variant analysis reports based on the analysis results, providing breeding experts with accurate selection criteria.

[0033] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0034] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A gene detection management method, characterized in that, Includes the following steps: Step S1: Collect biological samples using intelligent collection equipment and record detailed information about the biological samples; Step S2: Automated preprocessing of the extracted biological sample data; Step S3: Using sequencing technology and algorithms, the genes in the biological sample are sequenced, the quality of the sequencing data is monitored in real time, and the data quality is controlled. Step S4: Utilize big data analytics and machine learning techniques to deeply mine and analyze sequencing data, identifying disease-related gene variations and patterns; Step S5: Automatically generate personalized test reports based on disease-related gene variations and patterns; By using big data to collect the associations of gene mutations or hereditary traits corresponding to diseases, a gene-disease database is obtained. The gene-disease database includes a first disease, a second disease, ... and an nth disease, where n is a constant. The pointers of each disease address point to the associated genes included in the disease. Each gene address stores the association between the gene and the corresponding disease in the form of a dictionary. Feature selection is optimized using ant colony optimization algorithm. Combined with the gene-disease database, the gene variant features with the strongest association with the disease and the corresponding feature association are obtained. A regression model is used to establish a quantitative relationship between the gene variant features with the strongest association and the corresponding feature association and the disease risk. The gene-disease database is updated according to the quantitative relationship. The process of selecting the most representative features of a disease from a set of key features in biological samples using the ant colony algorithm, and reducing data dimensionality, includes: Based on the gene-disease database, a genotype feature matrix is ​​constructed according to the genotype of the disease. The row vectors of the genotype feature matrix represent the disease type, and the column vectors represent the gene types associated with the disease type. The node data of the genotype feature matrix is ​​represented by one-hot encoding. The node data corresponding to the genes included in the mutated gene region are uniformly overlapped and marked as mutated encoding data that does not repeat with all one-hot encodings. The initialization logic includes constructing an initial genotype feature matrix based on a gene-disease database crawled from big data, setting the node data in the initial genotype feature matrix as ant colony parameters, using non-zero node data as initial gene feature values, setting the disease correlation of the initial gene feature values ​​in the gene-disease database as heuristic information, and finding the most correlated gene feature value based on the heuristic information to obtain the most correlated gene type corresponding to the disease. All node data marked with variant coding data are extracted as variant gene feature values. The disease association of the variant gene feature values ​​in the gene-disease database is set as heuristic variant information. The most associated gene feature value is found according to the heuristic variant information, and the gene type corresponding to the variant gene feature value with the greatest pathogenic potential corresponding to the disease is obtained. The process of finding the most relevant gene trait values ​​includes: Each gene feature value is assigned a feature vector. A regression model is trained using the selected features to calculate the performance of the feature vector. The performance is represented by the mean squared error evaluation index. The probability of feature selection is automatically updated based on the performance. The probability is repeatedly updated until it can no longer be updated, and the optimal feature vector is obtained and output. The most relevant gene type for the disease selected by the ant colony algorithm and the gene type corresponding to the most pathogenic variant gene feature value for the disease are used as the first training set and the second validation set. A linear regression model is trained based on the first training set to obtain a gene-disease prediction model. The gene-disease prediction model is used to predict the second validation set to obtain the first variant gene prediction for the disease. The first variant gene prediction is added to the gene-disease database to update the gene-disease database. An updated second validation set is obtained based on the updated gene-disease database. The gene-disease database and the second validation set are iteratively updated. The differences between the gene-disease database and the second validation set during the update process are calculated according to the time series, and the differences are plotted as a probability distribution map. The probability distribution map should converge over time. When the probability distribution at time t converges to below a preset convergence threshold, the updated gene-disease database is obtained. The risk of recessive gene inheritance and variation in breeding ducks due to inbreeding was assessed by predicting the association between genes and diseases in a gene-disease database.

2. The gene detection management method as described in claim 1, characterized in that, Step S1 includes: The intelligent collection device includes an intelligent sampling instrument, which includes a dynamic blood drawer and a saliva collector; The biological samples include blood, saliva, and feather tissue samples. During the collection process, detailed information about the biological samples is recorded, including sampling time, sampling location, breeding duck identification information, and biological sample type. The collected biological samples are recorded using a unique barcode, which is automatically synchronized to the hospital or laboratory's information management system.

3. The gene detection management method as described in claim 2, characterized in that, Step S2 includes: The collected biological samples undergo preprocessing steps involving extraction and purification using automated equipment, and key features are extracted from the preprocessed data. The extracted key features are ligated into sequencing adapters to construct sequencing libraries.

4. The gene detection management method as described in claim 3, characterized in that, Key features extracted include: Genomic DNA is extracted from cells using chemical reagents, including CTAB and a salt solution. Gene regions are amplified using PCR, and the sequence information of the genome is obtained. The sequence information includes the genetic information of the organism, including genes encoding proteins, promoters of regulatory genes, and terminator regions. Single nucleotide polymorphisms (SNPs) were detected using high-throughput sequencing technology. These SNPs characterize gene variations among individuals, identify mutated gene regions, and extract histone modification and transcription factor binding regions. DNA methylation status was detected by sulfite sequencing; Total RNA was extracted from the sample using TRIzol, mRNA was extracted using magnetic beads, rRNA and tRNA were removed, the expression levels of several genes were detected by real-time quantitative PCR, miRNA was extracted from cells or tissues using a specialized miRNA extraction kit, miRNA expression and differential expression were extracted using miRNA microarray technology, and lncRNA expression patterns were obtained using transcriptome sequencing technology. Total protein was extracted from cells or tissues using lysis buffer, protein expression levels were detected by Western blotting, proteins were separated and identified by liquid chromatography-mass spectrometry, protein functions, modifications and interactions were annotated using proteomics databases, and post-translational modifications were analyzed by mass spectrometry to obtain protein functions and regulatory mechanisms. Detection of cell surface markers using antibody labeling; The extracted genome sequence information, variant gene regions, histone modification and transcription factor binding regions, DNA methylation status, expression levels of several genes, miRNA expression and differences, lncRNA expression patterns, protein expression levels, protein function and regulatory mechanisms, and cell markers are used as key data quantities and input into a neural network to extract key features of biological samples in a unified data format. These key features are then set into a key feature set of biological samples.

5. The gene detection management method as described in claim 4, characterized in that, Step S3 includes: DNA sequences are read using the Illumina high-throughput sequencing platform, and a data quality report is automatically generated. The data quality report includes the quality Q score based on each reading and the distribution of the quality Q score, GC content, and statistical information on repetitive sequences and short gene sequences. FastQC was used to remove duplicates, trim and filter DNA sequences, removing short gene sequences with Q scores below a preset threshold. Picard was used to remove duplicate short gene sequences. The alignment tool BWA was used to align the sequencing results with a pre-acquired reference genome. Based on the alignment results, GATK was used to call up variants and identify SNPs and InDels. Multiple sequencing was performed in the regions of variant genes and regions of histone modification and transcription factor binding.

6. The gene detection management method as described in claim 5, characterized in that, Step S5 includes: When a breeding duck biological sample is input, the system searches the gene-disease database for the corresponding genotype-related disease type extracted from the biological sample. It then uses a linear regression model to predict the probability of the disease and automatically generates a personalized test report. This provides doctors with accurate diagnostic information. The test report includes a detailed description of the gene mutation, the potential health risks of the mutation, the susceptibility probability to related diseases, and drug response.

7. A gene detection management system for running the gene detection management method according to any one of claims 1-6, characterized in that, It includes an acquisition module, a processing module, a sequencing module, a data analysis module, and a reporting module: The acquisition module is used to collect biological samples through intelligent acquisition devices and record detailed information about the biological samples. The processing module is used to automatically perform preprocessing steps on the extracted biological sample data; The sequencing module is used to sequence genes in biological samples using sequencing technology and algorithms, monitor the quality of sequencing data in real time, and perform data quality control. The data analysis module is used to perform in-depth mining and analysis of sequencing data using big data analysis and machine learning technologies to identify disease-related gene variations and patterns. The reporting module is used to automatically generate personalized test reports based on disease-related gene variations and patterns.

Citation Information

Patent Citations

  • Gene detection and inspection laboratory information management method and system

    CN117034361A

  • Genetic locus excavation method based on multi-target ant colony optimization algorithm

    CN105205344A

  • Extraction and detection integrated gene detection system

    CN114005488A

  • Method for predicting disease risk based on analysis of complex genetic information

    US20190385696A1