Method and system for processing genetic variation sites, and model training method
Through the site processing model, the genetic variant site characteristics and the alignment order of the site sequence are solved, and the problem of low accuracy of genetic variant site processing results in the prior art is solved, and higher processing accuracy is achieved, supporting the smooth development of biological genetic research.
Patent Information
- Application Number
- CN202411148099.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-20
AI Technical Summary
When using neural network models to analyze and process genetic variant sites, the prior art ignores the connection between different genetic variant sites, resulting in low accuracy of processing results, affecting the smooth development of biological genetic research.
By determining the initial site sequence of biological samples to be processed, it is input to the site processing model, using this model to process the characteristics of genetic variant sites, determine the association relationship between sites, and adjust the arrangement order of site sequences according to the association relationship, and obtain the target site sequence for processing.
The accuracy of genetic variation site processing results is improved, the correlation relationship between sites is taken into account, and the problem of low accuracy of neural network model output is avoided, ensuring the smooth progress of biological genetic research.
Smart Images

Figure CN119274650B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of biological genetic technology, and in particular, to a method for processing genetic variation sites; one or more embodiments of this specification also relate to a model training method, another method for processing genetic variation sites, a phenotype prediction method, a system for processing genetic variation sites, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In the field of biological genetics, analyzing and processing genetic variation sites in gene sequences is an important task in biological genetics research. With the continuous development of artificial intelligence technology, neural network models have also been introduced in biological genetics research to improve the efficiency of biological genetics research.
[0003] At present, in the process of using neural network models to analyze and process genetic variation sites, the connection between different genetic variation sites is often ignored, resulting in the problem of low accuracy of the processing results output by the neural network model. Therefore, it seriously affects the smooth progress of biological genetic research. Summary of the invention
[0004] In view of this, an embodiment of this specification provides a method for processing genetic variation sites. One or more embodiments of this specification also involve a model training method, another method for processing genetic variation sites, a phenotype prediction method, a genetic variation site processing system, a genetic variation site processing device, a model training device, another genetic variation site processing device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a method for processing a genetic variation site is provided, comprising:
[0006] Determining an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0007] The initial site sequence is input into a site processing model to obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result.
[0008] According to a second aspect of the embodiments of this specification, a device for processing a genetic variation site is provided, comprising:
[0009] A sequence determination module is configured to determine an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0010] A sequence processing module is configured to input the initial site sequence into a site processing model to obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result.
[0011] According to a third aspect of an embodiment of this specification, a model training method is provided, which is applied to a cloud-side device, including:
[0012] Determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence;
[0013] Using the site processing model to be trained, the genetic variation site characteristics of each sample genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of the genetic variation sites of each sample in the sample site sequence is adjusted according to the association relationship to obtain an adjusted sample site sequence;
[0014] Processing the adjusted sample site sequence to obtain a sample processing result;
[0015] The sample processing results and the sample labels are used to adjust the model parameters of the site processing model to be trained to obtain a trained site processing model.
[0016] According to a fourth aspect of an embodiment of this specification, a model training device is provided, which is applied to a cloud-side device, including:
[0017] A data determination module is configured to determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence;
[0018] a feature processing module, configured to process the genetic variation site features of each sample genetic variation site using the site processing model to be trained, determine the association relationship between the genetic variation sites, and adjust the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the association relationship to obtain an adjusted sample site sequence;
[0019] A result determination module is configured to process the adjusted sample site sequence to obtain a sample processing result;
[0020] The parameter adjustment module is configured to use the sample processing results and the sample labels to adjust the model parameters of the site processing model to be trained to obtain a trained site processing model.
[0021] According to a fifth aspect of the embodiments of this specification, a method for processing genetic variation sites is provided, which is applied to a cloud-side device, and includes:
[0022] A processing request for a genetic variation site sent by a receiving end-side device, wherein the processing request carries an initial site sequence of a biological sample to be processed, and the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0023] In response to the processing request, the initial site sequence is input into a site processing model to obtain processing results corresponding to the plurality of genetic variation sites, wherein the site processing model is used to process genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result;
[0024] The processing result is sent to the terminal side device.
[0025] According to a sixth aspect of the embodiments of this specification, a genetic variation site processing device is provided, which is applied to a cloud-side device, including:
[0026] A request receiving module is configured to receive a processing request for a genetic variation site sent by a terminal device, wherein the processing request carries an initial site sequence of a biological sample to be processed, and the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0027] A result acquisition module is configured to, in response to the processing request, input the initial site sequence into a site processing model to obtain processing results corresponding to the plurality of genetic variation sites, wherein the site processing model is used to process genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result;
[0028] The result sending module is configured to send the processing result to the terminal side device.
[0029] According to a sixth aspect of the embodiments of this specification, a genetic variation site processing system is provided, including a client and a server, wherein:
[0030] The client is configured to send an initial site sequence of a biological sample to be processed to the server, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0031] The server is configured to receive the initial site sequence of the biological sample to be processed sent by the client, input the initial site sequence into a site processing model, and obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result.
[0032] According to a seventh aspect of the embodiments of this specification, a phenotype prediction method is provided, comprising:
[0033] Determining an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0034] The initial site sequence is input into a phenotype prediction processing model to obtain phenotype prediction results corresponding to the multiple genetic variation sites, wherein the phenotype prediction processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the phenotype prediction result.
[0035] According to an eighth aspect of the embodiments of this specification, a computing device is provided, including:
[0036] Memory and processor;
[0037] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of any of the above-mentioned methods are implemented.
[0038] According to a ninth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of any of the methods described above are implemented.
[0039] According to a tenth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which implements the steps of any of the methods described above when executed by a processor.
[0040] One or more embodiments of the present specification provide a method for processing genetic variation sites. In the process of processing the genetic variation sites, an initial site sequence including a plurality of genetic variation sites arranged in sequence can be input into a site processing model for processing. The site processing model is used to process the genetic variation site characteristics of each genetic variation site, thereby determining the correlation between each genetic variation site, adjusting the arrangement order of each genetic variation site in the initial site sequence according to the correlation, and obtaining a target site sequence, thereby achieving the goal of considering the correlation between each genetic variation site in the process of processing the genetic variation sites, and obtaining a processing result with higher accuracy by processing the target site sequence determined based on the correlation, thereby avoiding the problem of low accuracy of the processing result output by the neural network model, and avoiding the serious impact on biological genetic research work caused by the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a schematic diagram of an application of a method for processing a genetic variation site provided by an embodiment of this specification;
[0042] Figure 2 It is a flow chart of a method for processing a genetic variation site provided by one embodiment of this specification;
[0043] Figure 3 is a processing flow chart of a method for processing a genetic variation site provided by an embodiment of this specification;
[0044] Figure 4 is a flow chart of a model training method provided by an embodiment of this specification;
[0045] Figure 5 is a flow chart of another method for processing genetic variation sites provided by one embodiment of this specification;
[0046] Figure 6 It is a schematic diagram of the structure of a genetic variation site processing system provided by one embodiment of this specification;
[0047] Figure 7 It is a schematic diagram of the structure of a genetic variation site processing device provided by one embodiment of this specification;
[0048] Figure 8 It is a structural schematic diagram of a model training device provided by an embodiment of this specification;
[0049] Fig. 9 It is a schematic diagram of the structure of another genetic variation site processing device provided by an embodiment of this specification;
[0050] Fig.10 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0051] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0052] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0053] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0054] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0055] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a foundation model / foundation model (Foundation Model 1). The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (Large Language Model 1, LLM), a multi-modal pre-training model (multi-modal pre-training model 1), etc.
[0056] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image description (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0057] First, the terms involved in one or more embodiments of this specification are explained.
[0058] Genotype: refers to the general term for all gene combinations of an individual organism, which reflects the genetic makeup of the organism, that is, the sum of all genes obtained from its parents; the genotype used specifically in genetics often refers to the genotype of a certain trait; if two organisms have only one different genotype, then their genotypes are different, so genotype refers to all combinations of all alleles at all loci of an individual.
[0059] Phenotype: refers to the phenotype, abbreviated as phenotype. Phenotype refers to the sum of characteristics exhibited by an individual with a specific genotype under certain environmental conditions; the so-called characteristics refer to the morphology, structure, physiology, biochemistry and other characteristics of an organism.
[0060] Deep learning: Deep learning is the process of learning the inherent laws and representation levels of sample data. The information obtained in the learning process is of great help in interpreting data such as text, images, and sounds. Its ultimate goal is to enable machines to have analytical learning capabilities like humans and to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm that has achieved results in speech and image recognition that far exceed previous related technologies.
[0061] Convolutional neural network: It is a type of feedforward neural network with a deep structure that includes convolution calculations. It is one of the representative algorithms of deep learning. Convolutional neural network has representational learning capabilities and can classify input information in a translation-invariant manner according to its hierarchical structure. Therefore, it is also called "translation-invariant artificial neural network."
[0062] Reinforcement Learning (RL): Also known as reinforcement learning, evaluation learning or enhanced learning, it is one of the paradigms and methodologies of machine learning. It is used to describe and solve the problem of how an agent can maximize rewards or achieve specific goals through learning strategies during its interaction with the environment.
[0063] SNP: refers to single nucleotide polymorphism, which is a DNA sequence polymorphism caused by the variation of a single nucleotide at the genome level.
[0064] AI (Artificial Intelligence): refers to artificial intelligence.
[0065] GBLUP: A model for estimating breeding values (BLUP) of individuals.
[0066] SVM (Support Vector Machine): refers to support vector machine, which is a generalized linear classifier that performs binary classification of data using supervised learning.
[0067] SVs (Structure Variations, abbreviated as SVs): refers to structural variations. SVs refer to long-length sequence changes and positional relationship changes in the genome.
[0068] Reward: It is a feedback signal given by the environment to the agent to evaluate the quality of the action taken by the agent at a certain moment. The design of reward value is one of the core parts of the reinforcement learning framework, which directly affects the learning direction and final strategy of the agent.
[0069] In the field of biological genetics, analyzing and processing genetic variation sites in gene sequences is an important task in biological genetics research. With the continuous development of artificial intelligence technology, neural network models have also been introduced in biological genetics research to improve the efficiency of biological genetics research. For example, AI-based smart breeding has become a research hotspot in the current field of biological genetics, and predicting phenotypes (traits) through genetic variation SNPs is the core of AI breeding. However, genetic variation sites are often of high dimensionality and are not tightly connected to each other.
[0070] In response to the above problems, this specification provides a solution that uses bioinformatics knowledge or feature engineering to reduce the dimension of the original full genetic variation data, and then builds a prediction model on the reduced-dimensional data. The prediction models are mainly divided into two categories:
[0071] The first type is: machine learning methods.
[0072] Among them, the representative of this method is the GBLUP model, and many variant solutions have evolved with the GBLUP model as the core. Machine learning methods such as random forests and SVM can handle scenarios with higher dimensions and small sample sizes. However, since the models are often based on artificially designed features and have limited model representation capabilities, they do not perform well for complex tasks.
[0073] The second type is: deep learning methods.
[0074] This method automatically extracts useful features from the input data through a deep model. The most commonly used one is the convolutional neural network, which can be used to extract local features. However, the model in this method has major defects: first, due to the large number of learning parameters, the model is often unable to handle particularly long site inputs; second, the model cannot aggregate the associations between distant sites, which limits the further development of this type of method.
[0075] Based on the defects of the above scheme, it can be seen that in the process of predicting phenotypes through genetic variation SNPs, how to effectively extract the characteristics between associated sites becomes a problem that needs to be solved.
[0076] Based on this, in this specification, a method for processing genetic variation sites is provided. One or more embodiments of this specification also involve a model training method, another method for processing genetic variation sites, a phenotype prediction method, a genetic variation site processing system, a genetic variation site processing device, a model training device, another genetic variation site processing device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0077] See also Figure 1 , Figure 1 A schematic diagram of an application of a method for processing a genetic variation site according to an embodiment of the present specification is shown. Figure 1 It can be seen that the user can send the SNPs site sequence to the server 104 through the terminal 102 for phenotype prediction. After the server 104 receives the SNPs site sequence (i.e., input data) for phenotype prediction, it inputs the input data into the representation prediction framework (i.e., site processing model), and uses the agent AG in the representation prediction framework to adjust the combination order of the genetic variation sites to obtain the genetic variation sites arranged in the new order; then, the phenotype prediction model in the representation prediction framework is used to extract features of the genetic variation sites arranged in the new order, thereby effectively extracting the site features between the associated sites, and performing phenotype prediction based on the site features to obtain the predicted value of the phenotype; the server 104 sends the predicted value of the phenotype output by the representation prediction framework to the terminal 102, thereby meeting the user's phenotype prediction needs.
[0078] See also Figure 2 , Figure 2 A flow chart of a method for processing a genetic variation site provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0079] Step 202: Determine an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence.
[0080] Among them, the biological sample to be processed can be understood as a biological sample whose genetic variation sites need to be processed. When the method for processing genetic variation sites provided in this specification is applied to different scenarios, the biological sample to be processed is also different; when the method for processing genetic variation sites is applied to agricultural breeding scenarios, the biological sample to be processed can be a crop sample to be processed; when the method for processing genetic variation sites is applied to animal husbandry scenarios, the biological sample to be processed can be an animal sample to be processed; when the method for processing genetic variation sites is applied to microbial scenarios, the biological sample to be processed can be a microbial sample to be processed; that is, the biological sample to be processed can be an economic crop sample to be processed, an animal sample to be processed, a plant sample to be processed and / or a microbial sample to be processed, etc.
[0081] The initial site sequence can be understood as a sequence composed of multiple genetic variation sites arranged in sequence; the sequential arrangement can be performed in a random arrangement manner, sorted according to the acquisition time of the genetic variation sites, or sorted according to the site position of the genetic variation sites in the gene sequence.
[0082] The genetic variation site can be understood as a difference site occurring on a DNA sequence with a fixed position in the genome, and the genetic variation site includes but is not limited to SNPs and SVs.
[0083] In one or more embodiments provided in this specification, in order to meet the user's processing operations on genetic variation sites, the method provided in this specification can receive an initial site sequence sent by the user and process the initial site sequence. The specific implementation method is as follows.
[0084] Determine the initial site sequence of the biological sample to be processed, including:
[0085] The initial site sequence of the biological sample to be processed sent by a client is received, wherein the initial site sequence is sent by the client when a user performs an initial site sequence upload operation.
[0086] Specifically, the method for processing genetic variation sites provided in this specification can be applied to a server, which can receive the initial site sequence of the biological sample to be processed sent by a client; the specific method is:
[0087] The user performs an initial site sequence upload operation on the client, thereby providing the initial site sequence of the biological sample to be processed to the client, and the client uploads the initial site sequence of the biological sample to be processed to the server;
[0088] The server can receive the initial site sequence of the biological sample to be processed sent by the client, and subsequently determine the processing results corresponding to multiple genetic variation sites based on the initial site sequence; thereby, through the interaction between the client and the server, the user's processing needs for genetic variation sites are met.
[0089] Step 204: Input the initial site sequence into a site processing model to obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result.
[0090] The site processing model can be understood as a model that can process genetic variation sites; the processing result can be understood as the processing result of the site processing model for the genetic variation site; when the genetic variation site processing method provided in this specification is applied to different scenarios, the site processing model is different, and the processing result is also different;
[0091] In the case where the method for processing genetic variation sites is applied to a phenotype prediction scenario, the site processing model may be a phenotype prediction processing model for predicting phenotypes through genetic variation SNPs, or the site processing model may be a phenotype prediction framework for predicting phenotypes through genetic variation SNPs, and correspondingly, the processing result may be a phenotype prediction result; in the case where the method for processing genetic variation sites is applied to a genetic disease analysis scenario, the site processing model may be a genetic disease prediction model for predicting genetic diseases through genetic variation SNPs, and correspondingly, the processing result may be a predicted genetic disease; in the case where the method for processing genetic variation sites is applied to an identity recognition scenario, In some embodiments, the site processing model may be an identity recognition model for performing identity recognition through genetic variation SNPs, and correspondingly, the processing result may be an identity recognition result; it should be noted that the site processing model may be a large model, and the site processing model (such as a phenotype prediction model) may include a feature extraction unit and a feature processing unit; the processing result may be a phenotype prediction result, and the phenotype prediction result may be understood as a result of characterizing the phenotypic situation corresponding to the genetic variation site; the phenotype prediction result may be a phenotype, a predicted value of a phenotype, etc., and the predicted value of the phenotype may be represented by a preset numerical value, for example, the predicted value may be represented by a floating-point numerical value.
[0092] The genetic variation site feature can be understood as the genetic variation feature corresponding to the genetic variation site, and the association relationship can be understood as the relationship that characterizes the association between each genetic variation site in multiple genetic variation sites.
[0093] The target site sequence can be understood as a site sequence obtained after adjusting the arrangement order of each genetic variation site according to the association relationship. It should be noted that the adjustment of the arrangement order of each genetic variation site can be understood as adjusting at least two associated genetic variation sites to similar arrangement positions, and / or adjusting at least two unrelated genetic variation sites to distant arrangement positions.
[0094] In one or more embodiments provided in this specification, in the process of processing the initial site sequence using the site processing model to obtain processing results corresponding to multiple genetic variation sites, the order adjustment unit and the sequence processing unit in the site processing model can be used, and the specific implementation method is as follows.
[0095] Inputting the initial site sequence into a site processing model to obtain processing results corresponding to the multiple genetic variation sites includes:
[0096] Inputting the initial site sequence into the site processing model, wherein the site processing model includes a sequence adjustment unit and a sequence processing unit;
[0097] Using the sequence adjustment unit, the genetic variation site characteristics of each genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of each genetic variation site in the initial site sequence is adjusted according to the association relationship to obtain a target site sequence;
[0098] The target site sequence is processed using the sequence processing unit to obtain the processing result.
[0099] Among them, the order adjustment unit can be understood as a unit used to determine the association relationship between each genetic variation site, and adjust the arrangement order of each genetic variation site in the initial site sequence according to the association relationship; the order adjustment unit can be one or more network layers in the site processing model; the order adjustment unit can be a sub-model or reinforcement learning model in the site processing model; or, the order adjustment unit can be an intelligent agent in the site processing model; it should be noted that the order adjustment unit can be obtained by training through reinforcement learning; for example, the order adjustment unit can be an intelligent agent used to adjust the combination order of genetic variation sites.
[0100] The sequence processing unit can be understood as a unit used to process the target site sequence; the sequence processing unit can be one or more network layers in the site processing model; the sequence processing unit can be a sub-model in the site processing model; the sequence processing unit can be a convolutional neural network in the site processing model; or, the sequence processing unit can be a phenotype prediction model in the site processing model, and the phenotype prediction model can be composed of a one-dimensional convolutional neural network; for example, the sequence processing unit can be a phenotype prediction model for phenotype prediction based on genetic variation SNPs.
[0101] Taking the application of the method for processing genetic variation sites provided in this specification in the phenotype prediction scenario as an example, the method for processing the genetic variation sites is explained, wherein the biological sample to be processed can be a crop sample to be processed, multiple genetic variation sites can be SNPs, the site processing model can be a phenotype prediction framework, the sequence adjustment unit can be an intelligent agent, the sequence processing unit can be a phenotype prediction model, and the processing result can be a predicted phenotype. Based on this, in order to allow the phenotype prediction model to effectively extract the characteristics between the associated sites in SNPs, an agent AG (i.e., a sequence adjustment unit) is used to determine the associated sites in SNPs, and adjust the SNPs site sequence of each batch of input data (i.e., the initial site sequence) based on the associated sites; after completing the sequence adjustment, the new genetic variation site sequence is input into the phenotype prediction model for processing, thereby obtaining a predicted phenotype; wherein the learning method of the Agent (i.e., the agent AG) is achieved through reinforcement learning.
[0102] Based on the above embodiments, it can be seen that in the process of using the site processing model to process the initial site sequence to obtain processing results corresponding to multiple genetic variation sites, a sequence adjustment unit can be used to adjust the arrangement order of each genetic variation site in the initial site sequence according to the association relationship, so that the sequence processing unit can effectively extract the features between the associated sites and improve the accuracy of the processing results.
[0103] In one or more embodiments provided in this specification, the sequence adjustment unit is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, and adjust the arrangement order of each genetic variation site in the initial site sequence according to the association relationship to obtain the target site sequence, including:
[0104] Determine the plurality of genetic variation sites in the initial site sequence by using the sequence adjustment unit, and extract features of each genetic variation site to obtain the genetic variation site features corresponding to each genetic variation site;
[0105] Perform feature analysis on multiple genetic variation site features to obtain correlation information between the features of each genetic variation site;
[0106] Based on the association information, the association relationship between the genetic variation sites is determined.
[0107] Among them, the association information can be understood as information that characterizes the degree of association between the characteristics of a genetic variation site and the characteristics of other genetic variation sites. The association information can be represented by a matrix or a vector; or, the association information can be site type information, or, the association information can be similarity.
[0108] Among them, feature analysis of multiple genetic variation site features can be understood as clustering analysis of the multiple genetic variation site features, so as to determine the association information between each genetic variation site feature and other genetic variation site features; or, it can be understood as similarity calculation of the multiple genetic variation site features, and determining the association information between each genetic variation site feature and other genetic variation site features through the similarity.
[0109] Adjusting the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship can be understood as arranging associated genetic variation sites at closer positions and arranging unassociated genetic variation sites at farther positions; it should be noted that the higher the degree of association between the genetic variation sites, the closer the arrangement positions between the genetic variation sites, and the lower the degree of association between the genetic variation sites, the farther the arrangement positions between the genetic variation sites.
[0110] Continuing with the above example, in the process of adjusting the order of SNPs sites, the agent AG after reinforcement learning can be used to achieve position adjustment; the current agent AG of reinforcement learning can be replaced by a deep learning model, called deep reinforcement learning DRL, so AG can be directly used to achieve position adjustment of sites. The specific method is: first, the SNPs site data is input into a neural network model (i.e., agent AG), and k categories are set, where k is less than or equal to the SNPs site ns, and is a hyperparameter; secondly, the model extracts features from the SNPs site, determines the site features corresponding to each SNPs site, and performs cluster analysis on the site features corresponding to each SNPs site, and outputs k categories (i.e., association information); finally, based on the k categories, sites with the same category are placed in similar positions, otherwise, no position adjustment is performed.
[0111] Based on the above embodiments, the method for processing genetic variation sites provided in this specification uses a sequence adjustment unit to determine the association information between the characteristics of each genetic variation site, and uses the association information as the association relationship between each genetic variation site, so as to facilitate the subsequent effective extraction of features between associated sites.
[0112] In one or more embodiments provided in this specification, feature analysis is performed on multiple genetic variation site features to obtain association information between the features of each genetic variation site, including:
[0113] Performing feature cluster analysis on the multiple genetic variation site features to obtain at least two feature sets, determining the genetic variation site features of the same feature set as associated genetic variation site features, determining the genetic variation site features of different feature sets as unassociated genetic variation site features, and determining the association information based on the associated genetic variation site features and the unassociated genetic variation site features; or
[0114] A similarity calculation is performed based on the plurality of genetic variation site features to determine the similarity between the genetic variation site features, and the similarity is determined as the association information between the genetic variation site features.
[0115] Continuing with the above example, this method can be implemented in two ways in the process of obtaining the association information between the characteristics of each genetic variation site.
[0116] The first method is to perform feature clustering analysis on the features of multiple genetic variation sites to determine the method of association information. The specific execution steps are:
[0117] 1. Input each batch of input data into the agent AG to adjust the order of SNPs sites and set k categories, where k is less than or equal to the number of SNPs sites ns and is a hyperparameter.
[0118] 2. The agent AG extracts features from the SNPs sites, determines the site features corresponding to each SNPs site, and performs cluster analysis on the site features corresponding to each SNPs site, and outputs k categories (i.e., site sets).
[0119] 3. Determine the SNP sites of the same category as associated SNPs sites, and determine the SNP sites of different categories as unassociated SNPs sites, so as to determine the association information between the SNP sites.
[0120] The second method is to calculate the similarity of multiple genetic variation site features to determine the associated information. The specific execution steps are:
[0121] 1. Agent AG extracts features from the sampled SNPs and determines the site features corresponding to each SNPs site.
[0122] 2. Agent AG calculates the similarity of the site features corresponding to each SNPs site, thereby determining similar SNPs sites and dissimilar SNPs sites for each SNPs site.
[0123] The specific implementation method is to cluster the sampled SNPs sites, and determine the SNPs sites similar to each SNPs site based on the clustering analysis results; the similar SNPs sites are used as associated SNPs sites associated with each SNPs site; and the dissimilar SNPs are used as unassociated SNPs sites.
[0124] 3. The agent AG determines the association information between each SNP site and other SNP sites based on the associated SNPs site and the unassociated SNPs site.
[0125] Subsequently, the association information is used as an association relationship to adjust the order of SNPs sites, so that the associated SNPs sites are arranged in similar positions. Otherwise, no position adjustment is performed.
[0126] Based on the above embodiments, the method for processing genetic variation sites provided in this specification utilizes a sequence adjustment unit to mine the association relationship between genetic variation sites through cluster analysis, so as to facilitate the subsequent effective extraction of features between associated sites.
[0127] In one or more embodiments provided in this specification, the processing of the target site sequence to obtain the processing result includes:
[0128] The target site sequence is input into the sequence processing unit, and the sequence processing unit is used to determine the multiple genetic variation sites in the target site sequence.
[0129] Extracting features from each of the genetic variation sites to obtain associated genetic variation site features between associated genetic variation sites;
[0130] Result prediction is performed based on the characteristics of the associated genetic variation sites to obtain processing results corresponding to the multiple genetic variation sites.
[0131] Among them, the characteristics of associated genetic variation sites can be understood as the characteristics between each associated site. The sequence processing unit provided in this specification can extract local features by arranging the associated genetic variation sites in similar positions. Therefore, in the process of feature extraction, the sequence processing unit can effectively extract the characteristics between the associated sites and make predictions based on the characteristics, thereby obtaining accurate processing results.
[0132] Continuing with the above example, after the agent AG adjusts the SNPs site order of the input data, the new site data can be input into the phenotype prediction model, which can effectively extract the features between the associated sites and perform phenotype prediction based on the features, thereby accurately obtaining the predicted value of the phenotype, wherein the phenotype prediction model can be composed of a one-dimensional convolutional neural network.
[0133] In one or more embodiments provided in this specification, the site processing model is a phenotype prediction processing model, and the processing result is a phenotype prediction result;
[0134] The step of inputting the initial site sequence into a site processing model to obtain processing results corresponding to the plurality of genetic variation sites includes:
[0135] The initial site sequence is input into the phenotype prediction processing model to obtain the phenotype prediction results corresponding to the multiple genetic variation sites.
[0136] The phenotype prediction processing model includes a sequence adjustment unit and a sequence processing unit.
[0137] Continuing with the above example, the new site data is input into the phenotype prediction framework, which can effectively extract the features between the associated sites and perform phenotype prediction based on the features, thereby accurately obtaining the predicted value of the phenotype.
[0138] In one or more embodiments provided in this specification, after inputting the initial site sequence into the site processing model to obtain the processing results corresponding to the multiple genetic variation sites, the process further includes:
[0139] The processing result is sent to the client for display.
[0140] Specifically, after obtaining the processing result, the processing can be sent to the client, so that the processing result can be displayed to the user through the client, thereby meeting the user's processing needs for genetic variation sites.
[0141] In one or more embodiments provided in this specification, before inputting the initial site sequence into the site processing model to obtain the processing results corresponding to the multiple genetic variation sites, the method further includes:
[0142] Determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence;
[0143] Using the to-be-trained site processing model to process the genetic variation site characteristics of the genetic variation sites of each sample, and determining the association relationship between the genetic variation sites;
[0144] Adjusting the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the association relationship to obtain an adjusted sample site sequence;
[0145] Processing the adjusted sample site sequence to obtain a sample processing result;
[0146] The sample processing results and sample labels are used to adjust the model parameters of the site processing model to be trained to obtain the site processing model.
[0147] For the training steps of the site processing model in this embodiment, please refer to the corresponding or corresponding description in the following model training method, which will not be elaborated here.
[0148] One or more embodiments of the present specification provide a method for processing genetic variation sites. In the process of processing the genetic variation sites, an initial site sequence including a plurality of genetic variation sites arranged in sequence can be input into a site processing model for processing. The site processing model is used to process the genetic variation site characteristics of each genetic variation site, thereby determining the correlation between each genetic variation site, adjusting the arrangement order of each genetic variation site in the initial site sequence according to the correlation, and obtaining a target site sequence, thereby achieving the goal of considering the correlation between each genetic variation site in the process of processing the genetic variation sites, and obtaining a processing result with higher accuracy by processing the target site sequence determined based on the correlation, thereby avoiding the problem of low accuracy of the processing result output by the neural network model, and avoiding the serious impact on biological genetic research work caused by the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work.
[0149] The following combination Figure 3 Taking the application of the method for processing genetic variation sites provided in this specification in a phenotype prediction scenario as an example, the method for processing genetic variation sites is further described. Figure 3 A processing flow chart of a method for processing a genetic variation site provided in an embodiment of the present specification is shown, which specifically includes the following steps.
[0150] Step 302: Use the agent (AG) to adjust the combination order of genetic variation sites.
[0151] Specifically, in order to enable the phenotype prediction model to effectively extract the features between associated sites, this method introduces an agent AG (i.e., reinforcement learning model) to adjust the order of SNPs sites in each batch of input data, that is, Figure 3 The specific execution method is:
[0152] 1. Sample the site data to obtain the sampled site data.
[0153] Considering that global adjustment of site arrangement will cause great disturbance to the training process and is not conducive to the convergence of the model, it is implemented by updating the positions of small batches of sites each time; and for small batches of sites, sampling can be used to determine them.
[0154] The sampling method is as follows: for each batch of training data (i.e. Figure 3 The shown is the dimension of (s,n). The site data of the second dimension (n) is randomly sampled to obtain data of (s,ns) dimension (that is, the sampled site data). In other words, ns sites are randomly selected from the n sites for this action.
[0155] It should be noted that for the agent service (AG), the processing object is the site, so the (s, ns) data is converted to (ns, s), that is, the feature of each site is the genotype vector of s samples.
[0156] 2. Adjust the order of the sampled site data.
[0157] This solution can be implemented in two ways during the process of sequentially adjusting the sampled site data.
[0158] The first method is to calculate the similarity between ns sampling sites, and adjust the order according to the similarity. Similar sites are close in position, otherwise no operation is performed. The specific implementation steps are as follows:
[0159] (1) Agent AG extracts features from the sampled SNPs and determines the site features corresponding to each SNP.
[0160] (2) The agent AG performs cluster analysis on the site characteristics corresponding to each SNPs site, and determines the SNPs sites similar to each SNPs site based on the cluster analysis results; the similar SNPs sites are regarded as associated SNPs sites associated with each SNPs site; and the dissimilar SNPs are regarded as unassociated SNPs sites.
[0161] (3) Agent AG, based on the associated SNPs and unassociated SNPs, determines the association between each SNP and other SNPs, and adjusts the order of SNPs according to the association, so as to arrange the associated SNPs in similar positions. Otherwise, no position adjustment is performed.
[0162] Among them, the second method is: use the agent AG after reinforcement learning to achieve position adjustment; the current reinforcement learning agent AG can be replaced by a deep learning model, called deep reinforcement learning DRL, so the position adjustment of the site can be directly achieved by using AG. The specific implementation steps are:
[0163] (1) Input the input data (ns, s) into a neural network model (i.e., agent AG) and set k categories, where k is less than or equal to ns and is a hyperparameter.
[0164] (2) The model clusters the input data and outputs k categories.
[0165] (3) Sites of the same category are placed in close positions; otherwise, no position adjustment is performed.
[0166] In the above process of sorting SNPs sites, the agent AG uses the original input data as the initial state State0 (S0) and the initial environment. After the execution of Act ion is completed, it returns to the new state S1 and updates the environment to obtain the site data arranged in a new order.
[0167] In the above process of sorting SNPs sites, the agent AG uses the original input data as the initial state State0 (S0) and the initial environment. After the execution of Act ion is completed, it returns to the new state S1 and updates the environment to obtain the site data arranged in a new order.
[0168] Step 304: Perform phenotype prediction using the phenotype prediction model to obtain a predicted value of the phenotype.
[0169] In the process of phenotype prediction, this method uses the agent AG and combines it with the phenotype prediction model based on convolutional neural network. The specific steps of phenotype prediction are as follows:
[0170] 1. Input the site data arranged in the new order into the phenotype prediction model.
[0171] 2. Phenotype prediction model, which performs feature extraction on the input site data, thereby effectively extracting site features between associated sites.
[0172] 3. The phenotype prediction model performs phenotype prediction based on the characteristics of the site to obtain the predicted value of the phenotype.
[0173] It should be noted that the phenotypic prediction model of the present method is capable of extracting local features by arranging associated genetic variation sites in similar positions. Thus, during the feature extraction process, the phenotypic prediction model can effectively extract features between associated sites and make predictions based on the features, thereby obtaining accurate phenotypic prediction values.
[0174] Step 306: Perform model training based on the predicted values and observed values.
[0175] During the training process, the mean square error loss function is calculated through the phenotypic prediction values and the phenotypic observation values, and the mean square error loss function is used to adjust the model parameters of the phenotypic prediction model, thereby optimizing the phenotypic prediction model.
[0176] The agent AG is continuously improved through reinforcement learning. The specific learning method is as follows:
[0177] First, the difference (eg, similarity) between the predicted phenotype and the observed phenotype is determined.
[0178] Secondly, the reward value (Reward) is set according to the reduction or enlargement of the difference between the phenotypic prediction value and the phenotypic observation value.
[0179] Finally, the agent AG is reinforced based on the reward value.
[0180] The above model training steps are repeated until the model training stop condition is reached, and the trained phenotype prediction model and agent AG are obtained.
[0181] It should be noted that after completing the above model training operation, the trained representation prediction framework (including the phenotype prediction model and the agent AG framework) can be applied to the phenotype prediction scenario for use. The specific steps of applying the representation prediction framework are:
[0182] 1. Determine the input data required for phenotype prediction (i.e., SNPs site sequence) and input the input data into the phenotype prediction framework.
[0183] 2. Use the proxy AG in the representation prediction framework to adjust the combination order of genetic variation sites.
[0184] Specifically, the agent AG after reinforcement learning can analyze the correlation between each genetic variation site, and adjust the order between each genetic variation site smoothly based on the correlation, so as to obtain the genetic variation sites arranged in a new order, so that the associated SNPs sites are arranged in close positions, otherwise, they are arranged in farther positions.
[0185] 3. Using the phenotype prediction model in the representation prediction framework, feature extraction is performed on the genetic variation sites arranged in the new order, thereby effectively extracting the site characteristics between the associated sites, and performing phenotype prediction based on the site characteristics to obtain the predicted value of the phenotype.
[0186] 4. The representation prediction framework outputs the predicted value of the phenotype.
[0187] Based on the above steps, it can be seen that the method for processing genetic variation sites in one or more embodiments of the present specification provides a method for predicting phenotypes from genotype data based on reinforcement learning. The method uses a reinforcement learning model to adjust the combination order of genetic variation sites, and combines a phenotype prediction model based on a convolutional neural network to perform phenotype prediction. The reinforcement learning model and the phenotype prediction model are optimized respectively by setting rewards and losses based on the difference between the predicted value and the true value, thereby obtaining a reinforcement learning model and a phenotype prediction model with better performance.
[0188] It should be noted that the method of predicting phenotypes using genotype data based on reinforcement learning has advantages including but not limited to: avoiding the problem that the input data dimension is large, which makes it difficult for the model to fully extract the features between associated sites; because in this scheme, by introducing the reinforcement learning model into the phenotype prediction task, the input data is adjusted step by step, so that the final associated sites are continuously close, so that the convolutional network can fully extract the features of these site clusters, thereby solving the above problem and obtaining a more accurate prediction value of the phenotype.
[0189] See also Figure 4 , Figure 4 A flowchart of a model training method provided according to an embodiment of the present specification is shown. The model training method is applied to a cloud-side device and specifically includes the following steps.
[0190] Step 402: Determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence.
[0191] Step 404: Using the site processing model to be trained, the genetic variation site characteristics of each sample genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of the genetic variation sites of each sample in the sample site sequence is adjusted according to the association relationship to obtain an adjusted sample site sequence.
[0192] Step 406: Process the adjusted sample site sequence to obtain a sample processing result.
[0193] Step 408: Using the sample processing results and the sample labels, adjust the model parameters of the site processing model to be trained to obtain a trained site processing model.
[0194] The training biological sample can be understood as a biological sample used as training data, the sample genetic variation site can be understood as a genetic variation site used as a training sample, and the sample processing result can be understood as a processing result corresponding to the sample genetic variation site.
[0195] The sample label can be understood as a real processing result. For example, the sample label can be a phenotypic prediction value of the sample label, or a processing result such as a predicted genetic disease.
[0196] In one or more embodiments provided in this specification, the cloud-side device may be a central cloud device of a distributed architecture or an edge cloud device of a distributed architecture. The cloud-side device may be a cloud-side device on which a cloud desktop system or cloud desktop software is installed and deployed, such as a cloud server, a cloud host, etc.
[0197] It should be noted that the way the training data is processed by the site processing model to be trained in the above steps is consistent with the way the site processing model processes the initial site sequence in the embodiment of the above-mentioned method for processing genetic variation sites, and will not be elaborated on here.
[0198] In one or more embodiments provided in this specification, the method of using the site processing model to be trained to process the genetic variation site characteristics of each sample genetic variation site, determining the association relationship between the genetic variation sites, and adjusting the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the association relationship to obtain the adjusted sample site sequence includes:
[0199] Using the sequence adjustment unit in the training site processing model, the genetic variation site characteristics of each sample genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of the genetic variation sites of each sample in the sample site sequence is adjusted according to the association relationship to obtain an adjusted sample site sequence;
[0200] The step of processing the adjusted sample site sequence to obtain a sample processing result includes:
[0201] The adjusted sample site sequence is processed using a sequence processing unit in the site processing model to be trained to obtain a sample processing result.
[0202] Among them, for the explanation of the order adjustment unit and the sequence processing unit, reference may be made to the corresponding or corresponding contents in the embodiment of the above-mentioned method for processing genetic variation sites, and no further details will be given here.
[0203] In one or more embodiments provided in this specification, the step of adjusting the model parameters of the site processing model to be trained by using the sample processing result and the sample label to obtain the trained site processing model includes:
[0204] Calculating a loss function for the sequence processing unit using the sample processing result and the sample label;
[0205] Determine a reward value for the sequence adjustment unit by using the sample processing result and the similarity between the sample labels;
[0206] The loss function is used to optimize the model parameters of the sequence processing unit, and the reward value is used to perform reinforcement learning processing on the sequence adjustment unit, and the site processing model to be trained is continued to be trained until the model training stop condition is reached, thereby obtaining the trained site processing model.
[0207] Among them, the loss function can be a mean square error loss function, a root mean square error loss function, etc., and this specification does not make any specific restrictions on this.
[0208] Taking the application of the model training method provided in this specification in the phenotype prediction scenario as an example, the model training method is explained, the multiple genetic variation sites can be SNPs, the site processing model can be a phenotype prediction model framework, the sequence adjustment unit can be an intelligent agent, the sequence processing unit can be a phenotype prediction model, the sample processing result can be a phenotype prediction value corresponding to the training sample, and the sample label can be a real phenotype observation value; based on this, during the training process, the mean square error loss function is calculated by the phenotype prediction value and the phenotype observation value, and the mean square error loss function is used to adjust the model parameters of the phenotype prediction model, so as to optimize the phenotype prediction model; and the agent AG is continuously improved through reinforcement learning, and the specific learning method is: first, determine the difference between the phenotype prediction value and the phenotype observation value, and secondly, set the reward value (Reward) according to the reduction and enlargement of the difference between the phenotype prediction value and the phenotype observation value; finally, the agent AG is subjected to reinforcement learning based on the reward value.
[0209] The above model training steps are repeated until the model training stop condition is reached, and the trained phenotype prediction model and agent AG are obtained.
[0210] Based on the above embodiments, it can be seen that the reward value and the loss function are set by the gap between the predicted value and the true value to optimize the reinforcement learning model (i.e., the agent AG) and the phenotype prediction model respectively, so as to obtain a phenotype prediction model and a reinforcement learning model with better performance.
[0211] One or more embodiments of the present specification provide a model training method. In the process of training a site processing model, the site processing model to be trained can be used to process the genetic variation site characteristics of each sample genetic variation site, determine the correlation between the genetic variation sites, and adjust the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the correlation to obtain the adjusted sample site sequence; thereby, in the process of training the site processing model, the correlation between the genetic variation sites can be considered, and by processing the sample site sequence adjusted based on the correlation, a sample processing result with high accuracy can be obtained; the model parameters of the site processing model to be trained are adjusted using the sample processing results and the sample labels, so as to obtain a site processing model that can output a processing result with high accuracy, so that the site processing model can consider the correlation between the genetic variation sites in the process of processing the genetic variation sites, so as to obtain a processing result with high accuracy, avoid the problem of low accuracy of the processing result output by the neural network model, and avoid the serious impact on biological genetic research work caused by the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work.
[0212] The above is a schematic scheme of a model training method of this embodiment. It should be noted that the technical scheme of this model training method and the technical scheme of the above-mentioned method for processing a genetic variation site belong to the same concept, and the details of the technical scheme of a model training method that are not described in detail can be referred to the description of the technical scheme of the above-mentioned method for processing a genetic variation site.
[0213] See also Figure 5 , Figure 5 A flowchart of another method for processing genetic variation sites provided according to an embodiment of the present specification is shown. The method for processing genetic variation sites is applied to a cloud-side device and specifically includes the following steps.
[0214] Step 502: A processing request for a genetic variation site sent by a receiving end-side device, wherein the processing request carries an initial site sequence of a biological sample to be processed, and the initial site sequence includes a plurality of genetic variation sites arranged in sequence.
[0215] Step 504: In response to the processing request, the initial site sequence is input into the site processing model to obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain the target site sequence, and process the target site sequence to obtain the processing result.
[0216] Step 506: Send the processing result to the terminal side device.
[0217] In one or more embodiments provided in this specification, the cloud-side device may be a central cloud device of a distributed architecture or an edge cloud device of a distributed architecture. The cloud-side device may be a cloud-side device with a cloud desktop system or cloud desktop software installed and deployed, such as a cloud server, cloud host, etc. The terminal-side device may be understood as any terminal that interacts with the cloud-side device for data exchange. The terminal may be a laptop, desktop computer, tablet computer, smart device, server, etc.
[0218] One or more embodiments of the present specification provide another method for processing genetic variation sites that can be applied to cloud-side devices. In the process of processing genetic variation sites, the method can receive a processing request for genetic variation sites sent by a terminal-side device, and the processing request carries an initial site sequence of a biological sample to be processed; based on this, the cloud-side device can input an initial site sequence containing a plurality of genetic variation sites arranged in sequence into a site processing model for processing, and the site processing model is used to process the genetic variation site characteristics of each genetic variation site, thereby determining the association relationship between each genetic variation site, adjusting the arrangement order of each genetic variation site in the initial site sequence according to the association relationship, and obtaining a target site sequence, thereby achieving the process of processing genetic variation sites. The association relationship between each genetic variation site can be considered, and by processing the target site sequence determined based on the association relationship, a processing result with higher accuracy can be obtained, thereby avoiding the problem of low accuracy of the processing result output by the neural network model, and avoiding the serious impact on biological genetic research work due to the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work.
[0219] The above is a schematic scheme of another method for processing genetic variation sites of this embodiment. It should be noted that the technical scheme of the another method for processing genetic variation sites and the technical scheme of the above method for processing one genetic variation site belong to the same concept, and the details of the technical scheme of the another method for processing genetic variation sites that are not described in detail can all be referred to the description of the technical scheme of the above method for processing one genetic variation site.
[0220] One or more embodiments of this specification provide a phenotype prediction method, comprising:
[0221] Determining an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0222] The initial site sequence is input into a phenotype prediction processing model to obtain phenotype prediction results corresponding to the multiple genetic variation sites, wherein the phenotype prediction processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the phenotype prediction result.
[0223] One or more embodiments of the present specification provide a phenotype prediction method. In the process of phenotype prediction of genetic variation sites, an initial site sequence including a plurality of genetic variation sites arranged in sequence can be input into a phenotype prediction processing model for processing. The phenotype prediction processing model is used to process the genetic variation site characteristics of each genetic variation site, thereby determining the correlation between each genetic variation site, adjusting the arrangement order of each genetic variation site in the initial site sequence according to the correlation, and obtaining a target site sequence, thereby achieving the goal of considering the correlation between each genetic variation site in the process of phenotype prediction of genetic variation sites, and obtaining a phenotype prediction result with high accuracy by performing phenotype prediction on the target site sequence determined based on the correlation, thereby avoiding the problem of low accuracy of the phenotype prediction result output by the neural network model, and avoiding the serious impact on biological genetic research work due to the low accuracy of the phenotype prediction result, thereby ensuring the smooth progress of biological genetic research work.
[0224] The above is a schematic scheme of a phenotype prediction method of this embodiment. It should be noted that the technical scheme of this phenotype prediction method and the technical scheme of the above-mentioned method for processing a genetic variation site belong to the same concept, and the details of the technical scheme of the phenotype prediction method that are not described in detail can all be referred to the description of the technical scheme of the above-mentioned method for processing a genetic variation site.
[0225] See also Figure 6 , Figure 6 FIG. 1 is a schematic diagram showing a structure of a genetic variation site processing system provided by an embodiment of the present specification. Figure 6 As shown, the device includes a client 602 and a server 604, wherein:
[0226] The client 602 is configured to send an initial site sequence of a biological sample to be processed to the server 604, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0227] The server 604 is configured to receive the initial site sequence of the biological sample to be processed sent by the client 602, input the initial site sequence into a site processing model, and obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain the target site sequence, and process the target site sequence to obtain the processing result.
[0228] One or more embodiments of the present specification provide a genetic variation site processing system including a client and a server. In the process of processing the genetic variation sites, the client can send the initial site sequence of the biological sample to be processed to the server. The server can input the initial site sequence including multiple genetic variation sites arranged in sequence into a site processing model for processing. The site processing model is used to process the genetic variation site characteristics of each genetic variation site, thereby determining the association relationship between each genetic variation site, adjusting the arrangement order of each genetic variation site in the initial site sequence according to the association relationship, and obtaining a target site sequence, thereby achieving that in the process of processing the genetic variation sites, the association relationship between each genetic variation site can be considered, and by processing the target site sequence determined based on the association relationship, a processing result with higher accuracy can be obtained, thereby avoiding the problem of low accuracy of the processing result output by the neural network model, and avoiding the serious impact on biological genetic research work due to the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work.
[0229] The above is a schematic scheme of a genetic variation site processing system of this embodiment. It should be noted that the technical scheme of the genetic variation site processing system and the technical scheme of the genetic variation site processing method described above belong to the same concept, and the details of the technical scheme of the genetic variation site processing system that are not described in detail can all be referred to the description of the technical scheme of the genetic variation site processing method described above.
[0230] Corresponding to the above method embodiment, this specification also provides an embodiment of a device for processing genetic variation sites, Figure 7 FIG. 2 shows a schematic diagram of a genetic variation site processing device provided by an embodiment of the present specification. Figure 7 As shown, the device comprises:
[0231] A sequence determination module 702 is configured to determine an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0232] The sequence processing module 704 is configured to input the initial site sequence into a site processing model to obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result.
[0233] Optionally, the sequence processing module 704 is further configured to:
[0234] Inputting the initial site sequence into the site processing model, wherein the site processing model includes a sequence adjustment unit and a sequence processing unit;
[0235] Using the sequence adjustment unit, the genetic variation site characteristics of each genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of each genetic variation site in the initial site sequence is adjusted according to the association relationship to obtain a target site sequence;
[0236] The target site sequence is processed using the sequence processing unit to obtain the processing result.
[0237] Optionally, the sequence processing module 704 is further configured to:
[0238] Determine the plurality of genetic variation sites in the initial site sequence by using the sequence adjustment unit, and extract features of each genetic variation site to obtain the genetic variation site features corresponding to each genetic variation site;
[0239] Perform feature analysis on multiple genetic variation site features to obtain correlation information between the features of each genetic variation site;
[0240] Based on the association information, the association relationship between the genetic variation sites is determined.
[0241] Optionally, the sequence processing module 704 is further configured to:
[0242] Performing feature cluster analysis on the multiple genetic variation site features to obtain at least two feature sets, determining the genetic variation site features of the same feature set as associated genetic variation site features, determining the genetic variation site features of different feature sets as unassociated genetic variation site features, and determining the association information based on the associated genetic variation site features and the unassociated genetic variation site features; or
[0243] A similarity calculation is performed based on the plurality of genetic variation site features to determine the similarity between the genetic variation site features, and the similarity is determined as the association information between the genetic variation site features.
[0244] Optionally, the sequence processing module 704 is further configured to:
[0245] Inputting the target site sequence into the sequence processing unit, and using the sequence processing unit to determine the multiple genetic variation sites in the target site sequence;
[0246] Extracting features from each of the genetic variation sites to obtain associated genetic variation site features between associated genetic variation sites;
[0247] Result prediction is performed based on the characteristics of the associated genetic variation sites to obtain processing results corresponding to the multiple genetic variation sites.
[0248] Optionally, the sequence determination module 702 is further configured to:
[0249] receiving the initial site sequence of the biological sample to be processed sent by a client, wherein the initial site sequence is sent by the client when the user performs an initial site sequence upload operation;
[0250] After inputting the initial site sequence into the site processing model to obtain the processing results corresponding to the plurality of genetic variation sites, the method further includes:
[0251] The processing result is sent to the client for display.
[0252] Optionally, the site processing model is a phenotype prediction processing model, and the processing result is a phenotype prediction result;
[0253] The sequence processing module 704 is further configured to:
[0254] The initial site sequence is input into the phenotype prediction processing model to obtain the phenotype prediction results corresponding to the multiple genetic variation sites.
[0255] Optionally, the genetic variation site processing device further includes a model training module configured to:
[0256] Determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence;
[0257] Using the to-be-trained site processing model to process the genetic variation site characteristics of the genetic variation sites of each sample, and determining the association relationship between the genetic variation sites;
[0258] Adjusting the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the association relationship to obtain an adjusted sample site sequence;
[0259] Processing the adjusted sample site sequence to obtain a sample processing result;
[0260] The sample processing results and sample labels are used to adjust the model parameters of the site processing model to be trained to obtain the site processing model.
[0261] One or more embodiments of the present specification provide a device for processing genetic variation sites. In the process of processing the genetic variation sites, an initial site sequence including a plurality of genetic variation sites arranged in sequence can be input into a site processing model for processing. The site processing model is used to process the genetic variation site characteristics of each genetic variation site, thereby determining the correlation between each genetic variation site, adjusting the arrangement order of each genetic variation site in the initial site sequence according to the correlation, and obtaining a target site sequence, thereby achieving the goal of considering the correlation between each genetic variation site in the process of processing the genetic variation sites, and obtaining a processing result with higher accuracy by processing the target site sequence determined based on the correlation, thereby avoiding the problem of low accuracy of the processing result output by the neural network model, and avoiding the serious impact on biological genetic research work caused by the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work.
[0262] The above is a schematic scheme of a processing device for a genetic variation site of this embodiment. It should be noted that the technical scheme of the processing device for a genetic variation site and the technical scheme of the processing method for a genetic variation site described above belong to the same concept, and the details of the technical scheme of the processing device for a genetic variation site that are not described in detail can all be referred to the description of the technical scheme of the processing method for a genetic variation site described above.
[0263] Corresponding to the above method embodiment, this specification also provides a model training device embodiment, Figure 8 FIG. 1 shows a schematic diagram of a model training device provided by an embodiment of the present specification. Figure 8 As shown, the device comprises:
[0264] The data determination module 802 is configured to determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence;
[0265] The feature processing module 804 is configured to process the genetic variation site features of each sample genetic variation site using the site processing model to be trained, determine the association relationship between the genetic variation sites, and adjust the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the association relationship to obtain an adjusted sample site sequence;
[0266] A result determination module 806 is configured to process the adjusted sample site sequence to obtain a sample processing result;
[0267] The parameter adjustment module 808 is configured to use the sample processing results and the sample labels to adjust the model parameters of the site processing model to be trained to obtain a trained site processing model.
[0268] Optionally, the feature processing module 804 is further configured to:
[0269] Using the sequence adjustment unit in the training site processing model, the genetic variation site characteristics of each sample genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of the genetic variation sites of each sample in the sample site sequence is adjusted according to the association relationship to obtain an adjusted sample site sequence;
[0270] The step of processing the adjusted sample site sequence to obtain a sample processing result includes:
[0271] The adjusted sample site sequence is processed using a sequence processing unit in the site processing model to be trained to obtain a sample processing result.
[0272] Optionally, the parameter adjustment module 808 is further configured to:
[0273] Calculating a loss function for the sequence processing unit using the sample processing result and the sample label;
[0274] Determine a reward value for the sequence adjustment unit by using the sample processing result and the similarity between the sample labels;
[0275] The loss function is used to optimize the model parameters of the sequence processing unit, and the reward value is used to perform reinforcement learning processing on the sequence adjustment unit, and the site processing model to be trained is continued to be trained until the model training stop condition is reached, thereby obtaining the trained site processing model.
[0276] One or more embodiments of the present specification provide a model training device. In the process of training a site processing model, the site processing model to be trained can be used to process the genetic variation site characteristics of each sample genetic variation site, determine the correlation between the genetic variation sites, and adjust the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the correlation to obtain the adjusted sample site sequence; thereby, in the process of training the site processing model, the correlation between the genetic variation sites can be considered, and by processing the sample site sequence adjusted based on the correlation, a sample processing result with high accuracy can be obtained; the model parameters of the site processing model to be trained are adjusted using the sample processing results and the sample labels, so as to obtain a site processing model that can output a processing result with high accuracy, so that the site processing model can consider the correlation between the genetic variation sites in the process of processing the genetic variation sites, so as to obtain a processing result with high accuracy, avoid the problem of low accuracy of the processing result output by the neural network model, and avoid the serious impact on biological genetic research work caused by the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work.
[0277] The above is a schematic scheme of a model training device of this embodiment. It should be noted that the technical scheme of the model training device and the technical scheme of the above-mentioned model training method belong to the same concept, and the details not described in detail in the technical scheme of the model training device can be referred to the description of the technical scheme of the above-mentioned model training method.
[0278] Corresponding to the above method embodiment, this specification also provides another embodiment of a device for processing genetic variation sites, Fig. 9 FIG. 2 shows a schematic diagram of the structure of another genetic variation site processing device provided by an embodiment of the present specification. Fig. 9 As shown, the device is applied to a cloud-side device, including:
[0279] The request receiving module 902 is configured to receive a processing request for a genetic variation site sent by a terminal device, wherein the processing request carries an initial site sequence of a biological sample to be processed, and the initial site sequence includes a plurality of genetic variation sites arranged in sequence;
[0280] The result acquisition module 904 is configured to, in response to the processing request, input the initial site sequence into a site processing model to obtain processing results corresponding to the plurality of genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result;
[0281] The result sending module 906 is configured to send the processing result to the terminal side device.
[0282] One or more embodiments of the present specification provide another genetic variation site processing device that can be applied to a cloud-side device. In the process of processing the genetic variation site, the method can receive a processing request for the genetic variation site sent by the terminal-side device, and the processing request carries an initial site sequence of the biological sample to be processed; based on this, the cloud-side device can input the initial site sequence containing multiple genetic variation sites arranged in sequence into a site processing model for processing, and the site processing model is used to process the genetic variation site characteristics of each genetic variation site, so as to determine the association relationship between each genetic variation site, and adjust the arrangement order of each genetic variation site in the initial site sequence according to the association relationship to obtain a target site sequence, thereby achieving that in the process of processing the genetic variation site, the association relationship between each genetic variation site can be considered, and by processing the target site sequence determined based on the association relationship, a processing result with higher accuracy can be obtained, thereby avoiding the problem of low accuracy of the processing result output by the neural network model, and avoiding the serious impact on biological genetic research work caused by the problem of low accuracy of the processing result, thereby ensuring the smooth progress of biological genetic research work.
[0283] The above is a schematic scheme of another genetic variation site processing device of this embodiment. It should be noted that the technical scheme of the another genetic variation site processing device and the technical scheme of the another genetic variation site processing method described above belong to the same concept, and the details of the technical scheme of the another genetic variation site processing device that are not described in detail can all be referred to the description of the technical scheme of the another genetic variation site processing method described above.
[0284] Fig.10 The block diagram of a computing device 1000 according to an embodiment of the present specification is shown. The components of the computing device 1000 include but are not limited to a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 via a bus 1030, and the database 1050 is used to store data.
[0285] The computing device 1000 also includes an access device 1040 that enables the computing device 1000 to communicate via one or more networks 1060. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1040 may include one or more of any type of network interface, wired or wireless (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Wide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0286] In one embodiment of the present specification, the above components of the computing device 1000 and Fig.10 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.10 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0287] The computing device 1000 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1000 may also be a mobile or stationary server.
[0288] The processor 1020 is used to execute the following computer program / instruction, which implements the steps of any one of the above methods when executed by the processor.
[0289] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to any of the above method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of any of the above method embodiments.
[0290] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0291] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to any of the above method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of any of the above method embodiments.
[0292] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0293] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of any of the above methods belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of any of the above methods.
[0294] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0295] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0296] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0297] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0298] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for processing genetic variation sites, comprising: Determining an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence; The initial site sequence is input into a site processing model to obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result, the association relationship is determined according to the association information between the characteristics of each genetic variation site, the association information is determined by clustering analysis of the multiple genetic variation site characteristics, or is determined according to the similarity obtained by similarity calculation of the multiple genetic variation site characteristics, and the method for obtaining the processing result includes: inputting the target site sequence into a sequence processing unit in the site processing model, using the sequence processing unit to determine multiple genetic variation sites in the target site sequence, extracting features of each genetic variation site, obtaining associated genetic variation site features between associated genetic variation sites, predicting results based on the associated genetic variation site features, and obtaining the processing results corresponding to the multiple genetic variation sites.
2. The method according to claim 1, inputting the initial site sequence into a site processing model to obtain processing results corresponding to the plurality of genetic variation sites, comprising: Inputting the initial site sequence into the site processing model, wherein the site processing model includes a sequence adjustment unit and a sequence processing unit; Using the sequence adjustment unit, the genetic variation site characteristics of each genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of each genetic variation site in the initial site sequence is adjusted according to the association relationship to obtain a target site sequence; The target site sequence is processed using the sequence processing unit to obtain the processing result.
3. The method according to claim 2, wherein the sequence adjustment unit is used to process the genetic variation site characteristics of each genetic variation site to determine the association relationship between the genetic variation sites, comprising: Determine the plurality of genetic variation sites in the initial site sequence by using the sequence adjustment unit, and extract features of each genetic variation site to obtain the genetic variation site features corresponding to each genetic variation site; Perform feature analysis on multiple genetic variation site features to obtain correlation information between the features of each genetic variation site; Based on the association information, the association relationship between the genetic variation sites is determined.
4. The method according to claim 3, wherein the feature analysis is performed on a plurality of genetic variation site features to obtain association information between the features of each genetic variation site, comprising: Performing feature cluster analysis on the multiple genetic variation site features to obtain at least two feature sets, determining the genetic variation site features of the same feature set as associated genetic variation site features, determining the genetic variation site features of different feature sets as unassociated genetic variation site features, and determining the association information based on the associated genetic variation site features and the unassociated genetic variation site features; or A similarity calculation is performed based on the plurality of genetic variation site features to determine the similarity between the genetic variation site features, and the similarity is determined as the association information between the genetic variation site features.
5. The method of claim 1, wherein determining the initial site sequence of the biological sample to be processed comprises: receiving the initial site sequence of the biological sample to be processed sent by a client, wherein the initial site sequence is sent by the client when the user performs an initial site sequence upload operation; After inputting the initial site sequence into the site processing model to obtain the processing results corresponding to the plurality of genetic variation sites, the method further includes: The processing result is sent to the client for display.
6. The method according to any one of claims 2 to 5, wherein the site processing model is a phenotype prediction processing model, and the processing result is a phenotype prediction result; The step of inputting the initial site sequence into a site processing model to obtain processing results corresponding to the plurality of genetic variation sites includes: The initial site sequence is input into the phenotype prediction processing model to obtain the phenotype prediction results corresponding to the multiple genetic variation sites.
7. The method according to any one of claims 1 to 5, before inputting the initial site sequence into the site processing model to obtain the processing results corresponding to the plurality of genetic variation sites, further comprising: Determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence; Using the to-be-trained site processing model to process the genetic variation site characteristics of each sample genetic variation site, and determining the association relationship between the genetic variation sites; Adjusting the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the association relationship to obtain an adjusted sample site sequence; Processing the adjusted sample site sequence to obtain a sample processing result; The sample processing results and sample labels are used to adjust the model parameters of the site processing model to be trained to obtain the site processing model.
8. A model training method, applied to a cloud-side device, comprising: Determine a site processing model to be trained and training data, wherein the training data includes a sample site sequence and a sample label of a training biological sample, and the sample site sequence includes a plurality of sample genetic variation sites arranged in sequence; Using the site processing model to be trained, the genetic variation site features of each sample genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of the genetic variation sites of each sample in the sample site sequence is adjusted according to the association relationship to obtain an adjusted sample site sequence, wherein the association relationship is determined according to the association information between the features of each genetic variation site, and the association information is determined by performing a cluster analysis on the features of multiple genetic variation sites, or is determined according to the similarity obtained by performing a similarity calculation on the features of the multiple genetic variation sites; Processing the adjusted sample site sequence to obtain a sample processing result, wherein the method for obtaining the sample processing result includes: inputting the adjusted sample site sequence into a sequence processing unit in the site processing model to be trained, using the sequence processing unit to determine multiple sample genetic variation sites in the adjusted sample site sequence, performing feature extraction on each sample genetic variation site, obtaining associated genetic variation site features between associated sample genetic variation sites, performing result prediction based on the associated genetic variation site features, and obtaining the sample processing results corresponding to the multiple sample genetic variation sites; The sample processing results and the sample labels are used to adjust the model parameters of the site processing model to be trained to obtain a trained site processing model.
9. The method according to claim 8, using the site processing model to be trained to process the genetic variation site characteristics of each sample genetic variation site, determine the association relationship between the genetic variation sites, and adjust the arrangement order of the genetic variation sites of each sample in the sample site sequence according to the association relationship to obtain the adjusted sample site sequence, comprising: Using the sequence adjustment unit in the training site processing model, the genetic variation site characteristics of each sample genetic variation site are processed to determine the association relationship between the genetic variation sites, and the arrangement order of the genetic variation sites of each sample in the sample site sequence is adjusted according to the association relationship to obtain an adjusted sample site sequence; The step of processing the adjusted sample site sequence to obtain a sample processing result includes: The adjusted sample site sequence is processed using a sequence processing unit in the site processing model to be trained to obtain a sample processing result.
10. The method according to claim 9, using the sample processing result and the sample label to adjust the model parameters of the site processing model to be trained to obtain a trained site processing model, comprising: Calculating a loss function for the sequence processing unit using the sample processing result and the sample label; Determine a reward value for the sequence adjustment unit by using the sample processing result and the similarity between the sample labels; The loss function is used to optimize the model parameters of the sequence processing unit, and the reward value is used to perform reinforcement learning processing on the sequence adjustment unit, and the site processing model to be trained is continued to be trained until the model training stop condition is reached, thereby obtaining the trained site processing model.
11. A method for processing genetic variation sites, applied to a cloud-side device, comprising: A processing request for a genetic variation site sent by a receiving end-side device, wherein the processing request carries an initial site sequence of a biological sample to be processed, and the initial site sequence includes a plurality of genetic variation sites arranged in sequence; In response to the processing request, the initial site sequence is input into a site processing model to obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site features of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result, wherein the association relationship is determined according to association information between the features of each genetic variation site, the association information is determined by clustering analysis of the features of multiple genetic variation sites, or is determined according to similarity obtained by similarity calculation of the features of the multiple genetic variation sites, and the processing result is obtained in a manner including: inputting the target site sequence into a sequence processing unit in the site processing model, determining multiple genetic variation sites in the target site sequence by using the sequence processing unit, extracting features of each genetic variation site, obtaining associated genetic variation site features between associated genetic variation sites, predicting results based on the associated genetic variation site features, and obtaining the processing results corresponding to the multiple genetic variation sites; The processing result is sent to the terminal side device.
12. A phenotype prediction method, comprising: Determining an initial site sequence of a biological sample to be processed, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence; The initial site sequence is input into a phenotype prediction processing model to obtain phenotype prediction results corresponding to the multiple genetic variation sites, wherein the phenotype prediction processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the phenotype prediction result, the association relationship is determined according to the association information between the characteristics of each genetic variation site, the association information is determined by clustering analysis of the multiple genetic variation site characteristics, or is determined according to the similarity obtained by similarity calculation of the multiple genetic variation site characteristics, and the method for obtaining the phenotype prediction result includes: inputting the target site sequence into a sequence processing unit in the phenotype prediction processing model, using the sequence processing unit to determine multiple genetic variation sites in the target site sequence, extracting features of each genetic variation site, obtaining associated genetic variation site features between associated genetic variation sites, performing result prediction based on the associated genetic variation site features, and obtaining the phenotype prediction results corresponding to the multiple genetic variation sites.
13. A genetic variation site processing system, comprising a client and a server, wherein: The client is configured to send an initial site sequence of a biological sample to be processed to the server, wherein the initial site sequence includes a plurality of genetic variation sites arranged in sequence; The server is configured to receive the initial site sequence of the biological sample to be processed sent by the client, input the initial site sequence into a site processing model, and obtain processing results corresponding to the multiple genetic variation sites, wherein the site processing model is used to process the genetic variation site characteristics of each genetic variation site, determine the association relationship between the genetic variation sites, adjust the arrangement order of the genetic variation sites in the initial site sequence according to the association relationship, obtain a target site sequence, and process the target site sequence to obtain the processing result, the association relationship is determined according to the association information between the characteristics of each genetic variation site, the association information is determined by clustering analysis of the multiple genetic variation site characteristics, or determined according to the similarity obtained by similarity calculation of the multiple genetic variation site characteristics, and the method for obtaining the processing result includes: inputting the target site sequence into a sequence processing unit in the site processing model, using the sequence processing unit to determine multiple genetic variation sites in the target site sequence, extracting features of each genetic variation site, obtaining associated genetic variation site features between associated genetic variation sites, predicting results based on the associated genetic variation site features, and obtaining the processing results corresponding to the multiple genetic variation sites.
14. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
16. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Neural network model training method and device and heating and ventilation system energy efficiency optimization method
CN111310905A
Target data determination method and device
CN113157760A