Gene detection management method and system
Through intelligent collection equipment and big data analysis technology, multi-level feature extraction and in-depth analysis of biological samples is solved, and the problem that the existing technology cannot extract multiple biological information at the same time and improve the depth of gene detection is achieved, the continuous update and accuracy of the gene-disease database is achieved, the level of disease prediction is improved, and the probability distribution map is generated, ensuring the stable convergence of the database.
Patent Information
- Application Number
- CN202510179321.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The existing technology cannot perform multi-level feature extraction, and cannot include the extraction and analysis of multiple biological information of genomic DNA, RNA, protein, miRNA and lncRNA at the same time. It cannot improve the depth of gene detection, cannot combine the gene-disease database with optimized feature selection, cannot identify the gene variant characteristics with the strongest correlation with disease, cannot continuously update and be accurate through continuous feature selection, regression model training and verification processes, cannot improve the prediction level of the association between genes and diseases, cannot generate probability distribution maps through monitoring the differences between databases and verification sets, and cannot ensure that the gene-disease database steadily converges to an accurate and credible level over time.
Biological samples are collected through intelligent collection equipment, and automated preprocessing is performed. Genes are sequenced using sequencing technology and algorithms. Big data analysis and machine learning technology are used for in-depth mining and analysis, identifying gene mutations and patterns related to diseases, and automatically generating personalized detection reports.
Multi-level feature extraction is realized, the depth of gene detection is improved, the gene-disease database is combined with optimized feature selection is realized, and the gene mutation characteristics with the strongest correlation with disease are identified. By continuously updating and preciselying the gene-disease database, the prediction level of the association between genes and diseases is improved, and the probability distribution map is generated, ensuring the stable convergence of the gene-disease database is provided, and detailed analysis of gene information and prediction of potential disease risks is provided, which improves the accuracy of early diagnosis.
Smart Images

Figure CN120089195A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene feature extraction, and particularly to a gene detection management method and system. Background Art
[0002] Gene detection technology has made remarkable progress in the fields of medicine, agriculture, environment, etc., and has shown great potential especially in disease prevention, precision medicine, personalized treatment, etc. However, with the extensive development of genomics research and clinical applications, how to efficiently, safely and reliably manage gene detection data, ensure the accuracy of data, privacy protection and the effectiveness of clinical applications has become an urgent problem to be solved. Therefore, gene detection management methods and systems have emerged. Their main purpose is to construct a complete set of solutions for data management, analysis, storage, transmission and security protection, which not only meet the clinical detection needs but also comply with ethical, legal and technical standards.
[0003] Currently, the Chinese invention patent with the application number CN202310956410.2 discloses a gene detection and inspection laboratory information management method and system, which collects various relevant information on gene detection and inspection; correlates the relevant information, encrypts and uploads it to the blockchain, and provides traceability query; constructs a network security situation model to evaluate the blockchain network. This application can not only provide complete data support for traceability query, but also ensure the confidentiality, privacy and security of gene detection and inspection information; however, the prior art cannot perform multi-level feature extraction, cannot simultaneously include the extraction and analysis of multiple biological information such as genomic DNA, RNA, protein, miRNA and lncRNA, cannot improve the depth of gene detection, cannot realize the practical combination of gene-disease database and optimized feature selection to identify the gene mutation features with the strongest correlation with diseases, cannot continuously update and refine the gene-disease database through continuous feature selection, regression model training and verification processes to improve the prediction level of the association between genes and diseases, cannot generate a probability distribution map by monitoring the differences between the database and the validation set, and cannot ensure that the gene-disease database converges stably to an accurate and credible level over time. Summary of the Invention
[0004] The technical problem solved by the present invention is that the prior art cannot perform multi-level feature extraction, cannot simultaneously include the extraction and analysis of various biological information such as genomic DNA, RNA, protein, miRNA, and lncRNA, cannot improve the depth of gene detection, cannot realize the practical combination of gene-disease databases and optimized feature selection to identify the gene mutation features with the strongest correlation with diseases, cannot continuously update and refine the gene-disease database through continuous feature selection, regression model training, and verification processes to improve the prediction level of the association between genes and diseases, cannot generate a probability distribution map by monitoring the differences between the database and the validation set, and cannot ensure that the gene-disease database converges stably to an accurate and reliable level over time.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: A gene detection management method, comprising the following steps: Step S1: Collect biological samples through intelligent collection devices and record the detailed information of the biological samples; Step S2: Automatically perform preprocessing steps on the extracted biological sample data; Step S3: Use sequencing technology and algorithms to sequence the genes in the biological samples, monitor the quality of the sequencing data in real time, and perform quality control on the data; Step S4: Use big data analysis and machine learning technologies to deeply mine and analyze the sequencing data to identify gene mutations and patterns related to diseases; Step S5: Automatically generate a personalized detection report based on the gene mutations and patterns related to diseases.
[0006] Preferably, the step S1 includes: The intelligent collection device includes an intelligent sampling instrument, and the intelligent sampling instrument includes a dynamic blood extractor; The biological samples include blood and feather tissue samples, and the detailed information of the biological samples is recorded during the collection process. The detailed information includes the sampling time, sampling location, duck identity information, and biological sample type; The collected biological samples are recorded through a unique barcode, and the barcode is automatically synchronized to the information management system of the hospital or laboratory.
[0007] Preferably, the step S2 includes: Perform preprocessing steps of extraction and purification on the collected biological samples through automated equipment, and extract key features from the preprocessed data; Connect the extracted key features to sequencing adapters to construct a sequencing library.
[0008] Preferably, the extraction of key features includes: Genomic DNA is extracted from cells using chemical reagents, which include CTAB and a salt solution. The gene region is amplified using PCR, and the sequence information of the genome is obtained. The sequence information includes the genetic information of an organism, and the genetic information includes genes encoding proteins, promoter and terminator regions regulating genes; Single nucleotide polymorphisms are detected by high-throughput sequencing technology. The single nucleotide polymorphisms characterize genetic variations between individuals, the variant gene regions are identified, and regions of histone modification and transcription factor binding are extracted; The DNA methylation status is detected by bisulfite sequencing; Total RNA is extracted from samples using TRIzol, mRNA is extracted by magnetic bead method, rRNA and tRNA are removed, the expression levels of several genes are detected by real-time quantitative PCR, miRNA is extracted from cells or tissues using a specialized miRNA extraction kit, the expression and differences of miRNA are extracted by miRNA chip technology, and the expression pattern of lncRNA is obtained by transcriptome sequencing technology; Total protein is extracted from cells or tissues using lysis buffer, the protein expression level is detected by Western Blot technology, proteins are separated and identified by liquid chromatography-mass spectrometry technology, the functions, modifications and interactions between proteins are annotated using a proteomics database, and post-translational modifications are analyzed by mass spectrometry to obtain the functions and regulatory mechanisms of proteins; Cell surface markers are detected by antibody labeling; The sequence information of the extracted genome, the variant gene regions, the regions of histone modification and transcription factor binding, the DNA methylation status, the expression levels of several genes, the expression and differences of miRNA, the expression pattern of lncRNA, the protein expression level, the functions and regulatory mechanisms of proteins, and the cell markers are respectively input into a neural network as key data volumes to extract the key feature quantities of biological samples in a unified data format, and the key feature quantities are aggregated into a key feature set of biological samples.
[0009] Preferably, step S3 includes: The DNA sequence is read based on the Illumina high-throughput sequencing platform, and a data quality report is automatically generated. The data quality report includes statistical information on the quality Q score based on each read and the distribution of the quality Q score, GC content, repetitive sequences, and gene short sequences; Use FastQC to de-duplicate, trim, and filter DNA sequences, remove short gene sequences with Q scores below a preset score threshold, use the tool Picard to remove duplicate short gene sequences, use the alignment tool BWA to align the sequencing results with a pre-obtained reference genome, and according to the alignment results, use GATK for variant calling to identify SNPs and InDels, and perform multiple sequencing in the regions of mutated genes and histone modifications and transcription factor bindings.
[0010] Preferably, use big data to collect the associations between gene mutations or genetic traits corresponding to diseases to obtain a gene-disease database. The gene-disease database includes the first disease, the second disease,..., and the nth disease, where n is a constant. The pointers of each disease address respectively point to the associated genes included in the disease, and the associations between genes and the corresponding diseases are stored in the form of a dictionary in each gene address. Optimize feature selection through the ant colony algorithm, combine the gene-disease database, obtain the gene mutation features with the strongest association with the disease and the corresponding feature associations, use a regression model to establish the quantitative relationship between the gene mutation features with the strongest association and the corresponding feature associations and the disease risk, and update the gene-disease database according to the quantitative relationship.
[0011] Preferably, use the ant colony algorithm to select the features that can best represent the disease from the key feature set of biological samples and reduce the data dimension. The process of selecting the features that can best represent the disease includes: Construct a genotype feature matrix based on the genotypes of diseases according to the gene-disease database. The row vectors of the genotype feature matrix represent disease types, and the column vectors represent the gene types associated with the corresponding disease types. The node data of the genotype feature matrix is represented by one-hot encoding, and the node data corresponding to the genes included in the mutated gene region is uniformly overlapped and marked with a mutated encoding data that is not repeated with all one-hot encodings. Initialize the gene-disease feature space. The initialization logic includes constructing an initial genotype feature matrix according to the gene-disease database captured by big data, setting the node data in the initial genotype feature matrix as ant colony parameters, setting the non-zero node data as the initial gene feature values, setting the corresponding disease associations of the initial gene feature values in the gene-disease database as heuristic information, and finding the most relevant gene feature values according to the heuristic information to obtain the most associated gene types corresponding to the disease. Extract all node data marked with mutant coding data as mutant gene eigenvalues, set the corresponding disease associations of the mutant gene eigenvalues in the gene-disease database as heuristic mutant information, find the most relevant gene eigenvalues according to the heuristic mutant information, and obtain the gene type corresponding to the mutant gene eigenvalue with the most pathogenic potential for the disease; The process of finding the most relevant gene eigenvalues includes: Arrange a column of feature vectors for each gene eigenvalue respectively, calculate the performance of the feature vectors using a regression model trained based on selected features, where the performance is represented by the evaluation index of mean squared error, automatically update the probability of feature selection according to the performance, repeat updating the probability until it cannot be updated, obtain the optimal feature vector, and output the optimal feature vector.
[0012] Preferably, use the most relevant gene type corresponding to the disease selected by the ant colony algorithm and the gene type corresponding to the mutant gene eigenvalue with the most pathogenic potential for the disease as the first training set and the second validation set, train a linear regression model based on the first training set to obtain a gene-disease prediction model, use the gene-disease prediction model to predict the second validation set to obtain the first mutant gene prediction for the disease, add the first mutant gene prediction to the gene-disease database to update the gene-disease database, obtain an updated second validation set according to the updated gene-disease database, iteratively update the gene-disease database and the second validation set, calculate the difference between the gene-disease database and the second validation set during the update process in time series, and plot the difference as a probability distribution graph, and the probability distribution graph should converge over time. When the probability distribution at time t converges to below a preset convergence threshold, obtain the updated gene-disease database; Predict and evaluate the risk of recessive gene inheritance and mutation in inbred breeding of breeding ducks through the association between genes and diseases in the gene-disease database.
[0013] Preferably, step S5 includes: When a biological sample of a breeding duck is input, find the disease types associated with the corresponding genotype extracted from the biological sample in the gene-disease database, predict the probability of suffering from the disease through a linear regression model, automatically generate a personalized detection report, and provide accurate diagnostic basis for doctors. The detection report includes a detailed description of the gene mutation, the potential health risks brought by the mutation, the susceptibility probability of related diseases, and drug reactions.
[0014] A gene detection management system includes a collection module, a processing module, a sequencing module, a data analysis module, and a report module: The collection module is used to collect biological samples through intelligent collection devices and record the detailed information of the biological samples; The processing module is used to automatically perform preprocessing steps on the extracted biological sample data; The sequencing module is used to sequence the genes in the biological sample by using sequencing technologies and algorithms, monitor the quality of the sequencing data in real time, and perform quality control on the data; The data analysis module is used to deeply mine and analyze the sequencing data by using big data analysis and machine learning technologies, and identify gene variations and patterns related to diseases; The reporting module is used to automatically generate a personalized detection report according to the gene variations and patterns related to diseases.
[0015] Advantages of the present invention: Multi-level feature extraction, including the extraction and analysis of various biological information such as genomic DNA, RNA, protein, miRNA, and lncRNA, improving the depth of gene detection, realizing the practical combination of the gene-disease database and optimized feature selection, identifying the gene variation features with the strongest correlation with diseases, continuously updating and refining the gene-disease database through continuous feature selection, regression model training, and verification processes, improving the prediction level of the association between genes and diseases, generating a probability distribution map by monitoring the differences between the database and the validation set, ensuring that the gene-disease database converges stably to an accurate and reliable level over time, not only being able to analyze the gene information of breeding ducks in detail, but also accurately predicting potential disease risks based on big data and machine learning, improving the accuracy of early diagnosis, and also being able to automatically generate a personalized site variation analysis report according to the analysis results, providing an accurate basis for breeding experts to select seeds. Brief Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the basic process of a gene detection management method provided by an embodiment of the present invention. Detailed Embodiments
[0017] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is made in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.
[0018] Refer to Figure 1 , an embodiment of the present invention, provides a gene detection management method, including the following steps: Step S1: Collect biological samples through intelligent collection devices and record the detailed information of the biological samples; Step S2: Automatically perform preprocessing steps on the extracted biological sample data; Step S3: Use advanced sequencing technologies and algorithms to sequence the genes in the biological sample, monitor the quality of the sequencing data in real time, and perform quality control on the data to ensure the reliability of the sequencing results; Step S4: Utilize big data analysis and machine learning technologies to deeply mine and analyze the sequencing data, and identify gene variations and patterns related to diseases; Step S5: Automatically generate a personalized detection report based on the gene variations and patterns related to diseases.
[0019] The said step S1 includes: The intelligent collection device includes an intelligent sampling instrument, and the intelligent sampling instrument includes a mobile blood extractor to ensure the efficiency and accuracy of sampling; The biological sample includes a blood sample and a feather tissue sample, and detailed information of the biological sample is recorded during the collection process. The detailed information includes the sampling time, sampling location, breeding duck identity information, and biological sample type; The collected biological sample is recorded by a unique barcode to ensure the one-to-one correspondence between the biological sample and the breeding duck information, avoid incorrect biological sample labeling, and the barcode is automatically synchronized to the information management system of the hospital or laboratory to ensure data integrity and traceability.
[0020] The purpose of gene detection is to understand the genetic background, genetic relationship, and disease resistance information of breeding ducks through scientific means.
[0021] The said step S2 includes: Perform a preprocessing step of extraction and purification on the collected biological sample through an automated device, and extract key features from the preprocessed data. The automated processing improves the processing efficiency and accuracy; According to the requirements of the sequencing platform, connect the extracted key features with sequencing adapters to construct a sequencing library.
[0022] The extraction of key features includes: Extract genomic DNA from cells using chemical reagents. The chemical reagents include CTAB and salt solution, and use PCR to amplify gene regions and the sequence information of the genome. The sequence information includes the genetic information of the organism, and the genetic information includes genes encoding proteins, promoter and terminator regions regulating genes; Detect single nucleotide polymorphisms through high-throughput sequencing technology. The single nucleotide polymorphisms characterize gene variations between individuals, identify the variant gene regions, and extract regions of histone modification and transcription factor binding; Detect the DNA methylation status through bisulfite sequencing; Total RNA was extracted from samples by TRIzol, mRNA was extracted by magnetic bead method, rRNA and tRNA were removed, and the expression levels of several genes, including the gene types specified by the user, were detected by real-time quantitative PCR. miRNA was extracted from cells or tissues by a specialized miRNA extraction kit, and the expression and differences of miRNA were extracted by miRNA chip technology. The expression pattern of lncRNA was obtained by transcriptome sequencing technology; Total protein was extracted from cells or tissues using lysis buffer, and the protein expression level was detected by Western Blot technology. Proteins were separated and identified by liquid chromatography-mass spectrometry. The functions, modifications, and interactions between proteins were annotated using a proteomics database. Post-translational modifications were analyzed by mass spectrometry to obtain the functions and regulatory mechanisms of proteins; Cell surface markers were detected by antibody labeling; The sequence information of the extracted genome, the mutated gene regions, the regions of histone modification and transcription factor binding, the DNA methylation status, the expression levels of several genes, the expression and differences of miRNA, the expression pattern of lncRNA, the protein expression level, the functions and regulatory mechanisms of proteins, and the cell markers were used as key data volumes and respectively input into a neural network to extract the key feature quantities of biological samples in a unified data format. The key feature quantities were aggregated into a key feature set of biological samples.
[0023] Step S3 includes: DNA sequences were read based on the Illumina high-throughput sequencing platform, and a data quality report was automatically generated. The data quality report included the quality Q score based on each read and the distribution of the quality Q score, GC content, statistical information on repetitive sequences, and gene short sequences; FastQC was used to de-duplicate, trim, and filter the DNA sequences, and gene short sequences with a Q score lower than a preset score threshold were removed. Due to PCR amplification bias, the sequencing results may contain repetitive sequences. The tool Picard was used to remove the repetitive gene short sequences, which helped improve the data quality. The alignment tool BWA was used to align the sequencing results with a pre-obtained reference genome to ensure the accuracy of the data. According to the alignment results, GATK was used for variant calling to identify SNPs and InDels. Multiple sequencing was performed in the mutated gene regions and the regions of histone modification and transcription factor binding to reduce accidental errors.
[0024] Sequencing the genes in biological samples using advanced sequencing technologies and algorithms and monitoring the quality of sequencing data in real time are important steps to ensure the reliability of sequencing results. By using high-throughput sequencing platforms, combined with real-time quality monitoring tools and strict quality control strategies, sequencing errors can be minimized to ensure that the data quality meets the standards of research or clinical analysis.
[0025] Collect the associations between gene mutations or genetic traits corresponding to diseases using big data to obtain a gene-disease database. The gene-disease database includes the first disease, the second disease,..., and the nth disease, where n is a constant. The pointers of each disease address respectively point to the associated genes included in the disease, and the associations between genes and the corresponding diseases are stored in the form of a dictionary in each gene address. Optimize feature selection through the ant colony algorithm, combine with the gene-disease database to obtain the gene mutation features with the strongest association with the disease and the corresponding feature associations. Use a regression model to establish the quantitative relationship between the gene mutation features with the strongest association and the corresponding feature associations and the disease risk, and update the gene-disease database according to the quantitative relationship.
[0026] Use the ant colony algorithm to select the features that best represent the disease from the key feature set of biological samples and reduce the data dimension. The process of selecting the features that best represent the disease includes: Gene data is usually high-dimensional and needs to be preprocessed first. Based on the gene-disease database, construct a genotype feature matrix based on the genotypes of diseases. The row vectors of the genotype feature matrix represent disease types, and the column vectors represent the gene types associated with the corresponding disease types. The node data of the genotype feature matrix is represented by one-hot encoding, and the node data corresponding to the genes included in the mutated gene region is uniformly overlapped and marked as mutated coding data that is not repeated with all one-hot encodings. Initialize the gene-disease feature space. The initialization logic includes constructing an initial genotype feature matrix based on the gene-disease database captured by big data, setting the node data in the initial genotype feature matrix as ant colony parameters, setting the non-zero node data as the initial gene feature values, setting the corresponding disease associations of the initial gene feature values in the gene-disease database as heuristic information, and finding the most relevant gene feature values according to the heuristic information to obtain the most associated gene types corresponding to the disease. Extract all the node data marked with mutated coding data as mutated gene feature values, set the corresponding disease associations of the mutated gene feature values in the gene-disease database as heuristic mutation information, find the most relevant gene feature values according to the heuristic mutation information, and obtain the gene types corresponding to the mutated gene feature values with the most pathogenic potential corresponding to the disease. The process of finding the most relevant gene eigenvalues includes: Arrange a column of eigenvectors for each gene eigenvalue respectively, calculate the performance of the eigenvectors by using a regression model trained based on selected features, where the performance is represented by the evaluation index of mean square error, automatically update the probability of feature selection according to the performance, and the probability is the weight automatically calculated by the computer. Repeat updating the probability until it cannot be updated to obtain the optimal eigenvector, and output the optimal eigenvector.
[0027] Take the most relevant gene types corresponding to the disease selected by the ant colony algorithm and the gene types corresponding to the most pathogenic potential mutant gene eigenvalues of the disease as the first training set and the second validation set. Train a linear regression model based on the first training set to obtain a gene-disease prediction model. Use the gene-disease prediction model to predict the second validation set to obtain the first mutant gene prediction of the disease. Add the first mutant gene prediction to the gene-disease database to update the gene-disease database. Obtain an updated second validation set according to the updated gene-disease database. Iteratively update the gene-disease database and the second validation set. Calculate the difference between the gene-disease database and the second validation set during the update process in time series, and plot the difference as a probability distribution graph. The probability distribution graph should converge over time. When the probability distribution at time t converges to be lower than a preset convergence threshold, obtain the updated gene-disease database, add the associated data between genes and diseases, and refine the association level between genes and diseases, especially increasing the disease prediction level of mutant genes. Predict and evaluate the risk of recessive gene inheritance and mutation in inbred breeding of breeding ducks through the association between genes and diseases in the gene-disease database. For example, the correlation between BRCA1 and BRCA2 gene mutations and breast cancer, or the relationship between the APOE gene and Alzheimer's disease.
[0028] The step S5 includes: When inputting a biological sample of a breeding duck, search for the disease types associated with the genotype extracted from the biological sample in the gene-disease database, and predict the probability of suffering from the disease through a linear regression model, and automatically generate a personalized detection report to provide an accurate diagnosis basis for doctors. The detection report includes a detailed description of the gene mutation, the possible health risks brought by the mutation, the susceptibility probability of related diseases, and drug reactions.
[0029] The generation of the report needs to take into account the understanding of doctors and breeding ducks, and usually includes concise charts, risk warnings, and necessary medical explanations obtained through big data search to help doctors make decisions.
[0030] A gene detection management system includes a collection module, a processing module, a sequencing module, a data analysis module, and a reporting module: The acquisition module is used to collect biological samples through intelligent acquisition devices and record the detailed information of the biological samples; The processing module is used to automatically perform preprocessing steps on the extracted biological sample data; The sequencing module is used to sequence the genes in biological samples by using sequencing technologies and algorithms, monitor the quality of sequencing data in real time, and perform quality control on the data; The data analysis module is used to deeply mine and analyze the sequencing data by using big data analysis and machine learning technologies, and identify gene variations and patterns related to diseases; The reporting module is used to automatically generate personalized detection reports according to gene variations and patterns related to diseases.
[0031] The gene detection management system is used to judge whether there is excessive inbreeding within a population. The judgment logic includes: calculating the total probability mean of recessive genetic traits and variations caused by inbreeding. When the preset genetic defect probability threshold is reached, artificial intervention in mating is carried out to prevent genetic defect problems caused by inbreeding; It is used to identify preferred disease-resistant genes through gene detection, so as to select breeding ducks with strong disease resistance to improve the disease resistance of the population. The gene detection management system can help predict the production performance of ducks and improve economic benefits.
[0032] The present invention realizes multi-level feature extraction, including the extraction and analysis of various biological information such as genomic DNA, RNA, protein, miRNA, and lncRNA at the same time, improves the depth of gene detection, combines the gene-disease database with optimized feature selection to achieve a practical combination, identifies the gene variation features with the strongest correlation with diseases, and continuously updates and refines the gene-disease database through continuous feature selection, regression model training, and verification processes, improves the prediction level of the association between genes and diseases, generates a probability distribution map by monitoring the differences between the database and the validation set, ensures that the gene-disease database converges stably to an accurate and reliable level over time, can not only analyze the gene information of breeding ducks in detail, but also accurately predict potential disease risks based on big data and machine learning, improve the accuracy of early diagnosis, and can also automatically generate personalized locus variation analysis reports according to the analysis results by combining the gene variations of breeding ducks, providing accurate seed selection basis for breeding experts.
[0033] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media that contain computer-usable program code. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk, or optical disk. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 or the functions specified in multiple boxes.
[0034] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A gene detection management method, characterized in that: The following steps are involved: Step S1: Collect biological samples through intelligent collection equipment and record detailed information of the biological samples; Step S2: automatically preprocessing the extracted biological sample data; Step S3: Sequencing the genes in the biological sample using sequencing technology and algorithms, monitoring the quality of the sequencing data in real time, and performing quality control on the data; Step S4: Using big data analysis and machine learning techniques to deeply mine and analyze sequencing data to identify disease-related gene variants and patterns; Step S5: Automatically generate a personalized test report based on disease-related gene variations and patterns.
2. The gene detection management method according to claim 1, characterized in that: The step S1 comprises: The intelligent collection equipment includes an intelligent sampling instrument, and the intelligent sampling instrument includes a blood drawer; The biological samples include blood and feather tissue samples, and detailed information of the biological samples is recorded during the collection process, wherein the detailed information includes sampling time, sampling location, identification information of the breeder ducks and type of biological samples; The collected biological samples are recorded by a unique barcode, which is automatically synchronized to the information management system of the hospital or laboratory.
3. The gene detection management method according to claim 2, characterized in that: The step S2 comprises: The collected biological samples are subjected to the pre-processing steps of extraction and purification by automated equipment, and key features are extracted from the pre-processed data; The extracted key features are connected to the sequencing adapters to construct the sequencing library.
4. The gene detection management method according to claim 3, characterized in that: The key features extracted include: Extracting genomic DNA from cells using chemical reagents, the chemical reagents including CTAB and salt solution, using PCR to amplify gene regions, genomic sequence information, the sequence information including genetic information of the organism, the genetic information including protein encoding genes, promoter and terminator regions of regulatory genes; Detecting single nucleotide polymorphisms (SNPs) by high-throughput sequencing technology, wherein the SNPs characterize gene variation between individuals, identifying the mutated gene regions, and extracting the regions of histone modification and transcription factor binding; DNA methylation status was detected by bisulfite sequencing; Total RNA was extracted from the samples by TRIzol, mRNA was extracted by magnetic bead method, rRNA and tRNA were removed, the expression of several genes was detected by real-time quantitative PCR, miRNA was extracted from cells or tissues by a special miRNA extraction kit, the expression and difference of miRNA was extracted by miRNA chip technology, and the expression pattern of lncRNA was obtained by transcriptome sequencing technology; Total protein is extracted from cells or tissues using lysis buffer, protein expression is detected by Western Blot, proteins are separated and identified by liquid chromatography-mass spectrometry, functions, modifications and interactions between proteins are annotated using proteomics databases, and post-translational modifications are analyzed by mass spectrometry to obtain protein functions and regulatory mechanisms; Detection of cell surface markers by antibody labeling; The extracted genome sequence information, variant gene regions, histone modification and transcription factor binding regions, DNA methylation status, expression levels of several genes, expression and differences of miRNA, expression pattern of lncRNA, protein expression level, protein function and regulatory mechanism and cell markers are respectively input into the neural network as key data quantities to extract key feature quantities of biological samples in a unified data format, and the key feature quantities are aggregated into a key feature set of biological samples.
5. The gene detection management method according to claim 4, characterized in that: The step S3 comprises: Reading DNA sequences based on the Illumina high-throughput sequencing platform, and automatically generating a data quality report, wherein the data quality report includes a quality Q score based on each read and the distribution of the quality Q score, GC content, repeated sequences, and statistical information of short gene sequences; FastQC was used to deduplicate, trim, and filter the DNA sequences, and short gene sequences with Q scores below the preset score threshold were removed. The tool Picard was used to remove repeated short gene sequences, and the alignment tool BWA was used to align the sequencing results with the pre-acquired reference genome. Based on the alignment results, GATK was used to call variants, identify SNPs and InDels, and perform multiple sequencing in the variant gene regions and the regions of histone modification and transcription factor binding.
6. The gene detection management method according to claim 5, characterized in that: Using big data to collect the association of gene mutations or genetic traits corresponding to the disease, a gene-disease database is obtained, wherein the gene-disease database includes a first disease, a second disease, ... and an nth disease, wherein n is a constant, and pointers of each disease address point to genes associated with the disease, and each gene address stores genes and the associations between the genes and the corresponding diseases in the form of a dictionary; The feature selection is optimized by the ant colony algorithm, and the gene-disease database is combined to obtain the gene variation features with the strongest correlation with the disease and the corresponding feature correlation. The regression model is used to establish a quantitative relationship between the gene variation features with the strongest correlation and the corresponding feature correlation and the disease risk, and the gene-disease database is updated according to the quantitative relationship.
7. The gene detection management method according to claim 6, characterized in that: The ant colony algorithm is used to select the feature that best represents the disease from the key feature set of the biological sample and reduce the data dimension. The process of selecting the feature that best represents the disease includes: Constructing a genotype feature matrix based on the genotype of the disease according to the gene-disease database, wherein the row vector of the genotype feature matrix represents the disease type, the column vector represents the gene type associated with the disease type, the node data of the genotype feature matrix is represented by one-hot coding, and the node data corresponding to the genes included in the mutated gene region are uniformly overlapped and marked as variant coding data that is not repeated with all one-hot codings; Initializing the gene-disease feature space, the initialization logic includes constructing an initial genotype feature matrix according to a gene-disease database captured by big data, setting node data in the initial genotype feature matrix as ant colony parameters, using node data that is not 0 as initial gene feature values, setting the disease association corresponding to the initial gene feature value in the gene-disease database as heuristic information, searching for the most correlated gene feature value according to the heuristic information, and obtaining the most correlated gene type corresponding to the disease; Extract all node data marked with variant coding data as variant gene feature values, set the corresponding disease association of the variant gene feature values in the gene-disease database as heuristic variation information, search for the most relevant gene feature value according to the heuristic variation information, and obtain the gene type corresponding to the variant gene feature value with the most pathogenic potential corresponding to the disease; The process of finding the most relevant gene feature values includes: Arrange a column of feature vectors for each gene feature value respectively, use the selected feature training regression model to calculate the performance of the feature vector, the performance is represented by an evaluation index of mean square error, automatically update the probability of feature selection according to the performance, repeatedly update the probability until it cannot be updated, obtain the optimal feature vector, and output the optimal feature vector.
8. The gene detection management method according to claim 7, characterized in that: The most relevant gene type corresponding to the disease selected by the ant colony algorithm and the gene type corresponding to the variant gene feature value with the most pathogenic potential corresponding to the disease are used as the first training set and the second validation set, a linear regression model is trained based on the first training set to obtain a gene-disease prediction model, the second validation set is predicted using the gene-disease prediction model to obtain a first variant gene prediction of the disease, the first variant gene prediction is added to the gene-disease database, the gene-disease database is updated, an updated second validation set is obtained according to the updated gene-disease database, the gene-disease database and the second validation set are iteratively updated, the difference between the gene-disease database and the second validation set during the updating process is calculated according to the time series, and the difference is plotted into a probability distribution graph, the probability distribution graph should converge over time, and when the probability distribution at time t converges to below a preset convergence threshold, an updated gene-disease database is obtained; The risk of recessive gene inheritance and mutation in inbreeding of breeder ducks was assessed by predicting the association between genes and diseases in the gene-disease database.
9. The gene detection management method according to claim 8, characterized in that: The step S5 comprises: When a biological sample of a breeding duck is input, the disease type associated with the corresponding genotype extracted from the biological sample is searched in the gene-disease database, and the probability of suffering from the disease is predicted through a linear regression model. A personalized test report is automatically generated to provide doctors with accurate diagnostic basis. The test report includes a detailed description of the gene mutation, the health risks that the mutation may bring, the susceptibility probability of related diseases and drug response.
10. A gene detection management system, characterized in that: Including acquisition module, processing module, sequencing module, data analysis module and reporting module: The collection module is used to collect biological samples through intelligent collection equipment and record detailed information of the biological samples; The processing module is used to automatically perform pre-processing steps on the extracted biological sample data; The sequencing module is used to sequence genes in biological samples using sequencing technology and algorithms, monitor the quality of sequencing data in real time, and perform quality control on the data; The data analysis module is used to use big data analysis and machine learning technology to deeply mine and analyze sequencing data and identify gene variations and patterns associated with diseases; The reporting module is used to automatically generate personalized test reports based on disease-related gene variations and patterns.
Citation Information
Patent Citations
Gene detection and inspection laboratory information management method and system
CN117034361A
Genetic locus excavation method based on multi-target ant colony optimization algorithm
CN105205344A
Gene detection management method and system
CN107066836A
Extraction and detection integrated gene detection system
CN114005488A
Intelligent gene detection database management system
CN119132395A