Gene prediction method and apparatus, computer device, and computer readable storage medium

Through a multi-layer neural network model, combined with the template gene sequence and the genetic and free gene sequences of the object to be tested, the target gene site is determined and the characteristic data is extracted, which solves the problem of low gene prediction accuracy of the traditional Bayesian probability model and achieves more efficient gene sequence recognition and prediction.

WO2025208288A1PCT designated stage Publication Date: 2025-10-09SHENZHEN HUADA GENE INST
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/085303
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-01
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Traditional Bayesian probability models have low accuracy in predicting the genomes of genetic offspring and have difficulty accurately identifying and distinguishing the presence and location of specific genes.

Method used

A multi-layer neural network model is used, including the first network, the second network and the third network. By obtaining the template gene sequence and the genetic gene sequence and free gene sequence of the object to be tested, the target gene site is determined and the feature data is extracted. These data are used to train the target gene prediction model for gene prediction.

Benefits of technology

It improves the accuracy and efficiency of gene sequence prediction, can accurately identify and distinguish the existence and location of specific genes, and has important application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024085303_09102025_PF_FP_ABST
    Figure CN2024085303_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a gene prediction method and apparatus, a computer device, and a computer readable storage medium. The method comprises: acquiring a template gene sequence, and a genetic gene sequence and a free gene sequence which correspond to a subject under test (step S102); determining a target gene site in the genetic gene sequence and the free gene sequence (step S104); extracting feature data corresponding to the target gene site, wherein the feature data is used for representing attribute features of the target gene site, and the feature data comprises first input data and second input data (step S106); acquiring a target gene prediction model, wherein the target gene prediction model comprises a first network, a second network, and a third network, and an output of the first network and the second network is an input of the third network (step S108); and respectively inputting the first input data and the second input data into the first network and the second network to obtain a gene prediction result outputted by the third network and corresponding to said subject (step S110).
Need to check novelty before this filing date? Find Prior Art

Description

Gene prediction method, device, computer equipment and computer readable storage medium Technical Field

[0001] The present application relates to the field of bioinformatics, and in particular to a gene prediction method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] With the development of bioinformatics, the detection results of peripheral plasma free genes based on the genetic father can effectively predict the gene sequence of the genetic offspring, which has very important research value for the experimental analysis and practical application of biological genetic characteristics.

[0003] In traditional technology, the genome of genetic offspring is usually inferred based on Bayesian probability models, but the prediction accuracy of the model is low.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a gene prediction method, apparatus, computer device, and computer-readable storage medium.

[0006] In a first aspect, the present application provides a gene prediction method, which is executed by a computer device, comprising:

[0007] Obtaining the template gene sequence and the genetic gene sequence and free gene sequence corresponding to the object to be tested, wherein the free gene sequence represents the sequence of the free gene corresponding to the object to be tested in the corresponding genetic object sample;

[0008] Determine the target gene sites in genetic gene sequences and episomal gene sequences;

[0009] Extracting feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, the feature data including the first input data and the second input data, and;

[0010] Obtaining a target gene prediction model, where the target gene prediction model includes a first network, a second network, and a third network, where the outputs of the first network and the second network serve as inputs to the third network; and

[0011] The first input data and the second input data are input into the first network and the second network respectively, and a gene prediction result corresponding to the object to be tested is obtained by the third network output.

[0012] In a second aspect, the present application further provides a method for generating a gene prediction model, which is executed by a computer device and comprises:

[0013] Obtaining a genetic gene sequence, an episomal gene sequence, a standard gene sequence, and a template gene sequence corresponding to the biological sample, wherein the episomal gene sequence represents the sequence of the episomal gene corresponding to the biological sample in the corresponding genetic object sample, and the standard gene sequence corresponds to the genetic gene sequence and the episomal gene sequence;

[0014] Determine the target gene site in the genetic gene sequence and the free gene sequence; extract feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes first input data and second input data, and;

[0015] The corresponding initial gene prediction model is trained based on the first input data and the second input data to obtain a corresponding target gene prediction model; the target gene prediction model includes a first network, a second network and a third network, the outputs of the first network and the second network are the inputs of the third network, and the target gene prediction model is used to predict the gene results of the corresponding biological sample.

[0016] In a third aspect, the present application further provides a gene prediction device, comprising:

[0017] An acquisition module is used to obtain the template gene sequence and the genetic gene sequence and free gene sequence corresponding to the object to be tested, wherein the free gene sequence represents the sequence of the free gene corresponding to the object to be tested in the corresponding genetic object sample;

[0018] An extraction module is used to determine the target gene site in the genetic gene sequence and the free gene sequence; extract the feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes the first input data and the second input data;

[0019] The prediction module is used to obtain a target gene prediction model, which includes a first network, a second network, and a third network. The outputs of the first network and the second network are input to the third network; and the first input data and the second input data are input to the first network and the second network respectively, to obtain the gene prediction result corresponding to the object to be tested output by the third network.

[0020] In a fourth aspect, the present application further provides a gene prediction model generation device, comprising:

[0021] An acquisition module is used to obtain the genetic gene sequence, free gene sequence, standard gene sequence and template gene sequence corresponding to the biological sample, the free gene sequence represents the sequence of the free gene corresponding to the biological sample in the corresponding genetic object sample, and the standard gene sequence corresponds to the genetic gene sequence and the free gene sequence;

[0022] An extraction module is used to determine the target gene site in the genetic gene sequence and the free gene sequence; extract feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes the first input data and the second input data;

[0023] The training module is used to train the corresponding initial gene prediction model based on the feature data to obtain the corresponding target gene prediction model. The gene prediction model is used to predict the gene results of the corresponding biological sample.

[0024] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of any of the above methods when executing the computer-readable instructions.

[0025] A computer-readable storage medium stores computer-readable instructions, which implement the steps of any of the above methods when executed by a processor.

[0026] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] FIG1 is a schematic diagram of a gene prediction method according to an embodiment;

[0029] FIG2 is a schematic diagram of a process for determining target gene sites in one embodiment;

[0030] FIG3 is a schematic diagram of a process for extracting characteristic data corresponding to target gene sites in one embodiment;

[0031] FIG4 is a schematic diagram of a process for screening and obtaining characteristic data in one embodiment;

[0032] FIG5 is a schematic diagram of a process for predicting a gene prediction sequence based on a target gene prediction model in one embodiment;

[0033] FIG6 is a schematic diagram of a process for determining a predicted gene sequence of a test object in one embodiment;

[0034] FIG7 is a schematic flow chart of a method for generating a gene prediction model in another embodiment;

[0035] FIG8 is a schematic diagram of a process for determining target gene sites in one embodiment;

[0036] FIG9 is a schematic diagram of a process for generating characteristic data corresponding to each target gene locus in one embodiment;

[0037] FIG10 is a schematic diagram of a process for screening and obtaining characteristic data in one embodiment;

[0038] FIG11 is a schematic diagram of a process for obtaining a gene prediction model based on feature data training in one embodiment;

[0039] FIG13 is a technical schematic diagram of a gene prediction method in a specific embodiment;

[0040] FIG14 is a neural network model used in a gene prediction method in a specific embodiment;

[0041] FIG15 is a schematic diagram of an actual prediction application of a gene prediction method according to a specific embodiment;

[0042] FIG16 is a block diagram of a gene prediction device according to an embodiment;

[0043] FIG17 is a block diagram of a device for generating a gene prediction model according to an embodiment;

[0044] FIG18 is a diagram showing the internal structure of a computer device according to one embodiment;

[0045] FIG19 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0047] In one embodiment, as shown in FIG1 , a gene prediction method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0048] Step S102: Obtain the template gene sequence and the genetic gene sequence and episomal gene sequence corresponding to the object to be tested.

[0049] The episomal sequence represents the sequence of the episomal gene corresponding to the test object in the corresponding genetic object sample (such as peripheral blood, plasma).

[0050] Specifically, the computer device can obtain the template gene sequence and the genetic gene sequence, free gene sequence and concentration corresponding to the free gene sequence of the object to be tested from the database, or obtain the above data from a third-party device, server, etc.

[0051] Step S104: determining the target gene sites in the genetic gene sequence and the episomal gene sequence.

[0052] Among them, the target gene site is a specific base pair.

[0053] Specifically, the computer device can compare / match the genetic gene sequence and the free gene sequence with the template gene sequence for each corresponding base pair, and use the base pairs in the genetic gene sequence and the free gene sequence that fail to compare / match as their target gene points, wherein the template gene sequence is the gene sequence corresponding to the object to be tested, such as the gene sequence template corresponding to a certain animal, plant or microorganism, and the genetic gene sequence is the gene sequence corresponding to the genetic parent of the object to be tested.

[0054] Step S106: extracting feature data corresponding to the target gene locus.

[0055] The feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes first input data and second input data.

[0056] Specifically, the computer equipment extracts corresponding characteristic parameters based on the chromosome number and position information of the target gene site, as well as the sequencing depth and sequencing quality of the gene sequence where the target gene site is located. The characteristic parameters can reflect the chromosome number and position information of the corresponding target gene site, as well as the sequencing depth and sequencing quality of the gene sequence where the target gene site is located.

[0057] Step S108: Obtain a target gene prediction model.

[0058] Among them, the target gene prediction model includes a first network, a second network and a third network. The outputs of the first network and the second network are the inputs of the third network. The first network, the second network and the third network can adopt any neural network or deep learning model, such as a convolutional neural network, a fully connected neural network, etc.

[0059] In step S110 , the first input data and the second input data are input into the first network and the second network respectively, and the gene prediction result target gene position corresponding to the object to be tested is obtained from the third network output.

[0060] Among them, the gene prediction result can be a genotype or a gene sequence. The gene prediction result can include at least two candidate gene sequences or genotypes. Each candidate gene prediction sequence or genotype includes a corresponding confidence level. The confidence level is used to characterize the probability that the corresponding candidate gene prediction sequence is a true gene sequence. The greater the confidence level, the greater the probability that the corresponding candidate gene prediction sequence is a true gene sequence, and vice versa.

[0061] The above-mentioned gene prediction method obtains the template gene sequence and the genetic gene sequence and free gene sequence corresponding to the object to be tested, determines the target gene site in the genetic gene sequence and the free gene sequence, extracts the characteristic data corresponding to the target gene site, and the characteristic data is used to characterize the attribute characteristics of the target gene site. The characteristic data includes first input data and second input data, obtains a target gene prediction model, and the target gene prediction model includes a first network, a second network and a third network. The outputs of the first network and the second network are inputs to the third network; the first input data and the second input data are respectively input into the first network and the second network to obtain the gene prediction result corresponding to the object to be tested output by the third network. In this way, the target gene site can be quickly and accurately identified / determined based on the genetic gene sequence and the free gene sequence, and according to the obtained target gene prediction model, the first input data and the second input data are respectively input into the first network and the second network of the target gene prediction model, and the gene prediction result corresponding to the object to be tested is obtained by the output of the third network, thereby completing the prediction of the gene sequence of the object to be tested, completing the combination of the characteristic attributes of the sample data itself with the specific model, and more accurately learning the base characteristics of each gene site in the sample gene sequence, thereby improving the prediction accuracy of the gene sequence.

[0062] In one embodiment, the genetic gene sequence corresponding to the subject to be tested is the genotype of the genetic subject corresponding to the subject to be tested.

[0063] In one embodiment, determining the target gene site in the genetic gene sequence or the episomal gene sequence includes:

[0064] The genetic gene sequence and the free gene sequence are matched with the template gene sequence respectively to determine the target gene sites in the genetic gene sequence and the free gene sequence.

[0065] Specifically, the computer device can compare / match the genetic gene sequence and the free gene sequence with the template gene sequence for each corresponding base pair, and use the base pairs in the genetic gene sequence and the free gene sequence that fail to compare / match as their target gene points, wherein the template gene sequence is the gene sequence corresponding to the object to be tested, such as the gene sequence template corresponding to a certain animal, plant or microorganism, and the genetic gene sequence is the gene sequence corresponding to the genetic parent of the object to be tested.

[0066] In this example, the target gene loci within the inherited and episomal gene sequences are determined by matching them with the template gene sequence. This matching process allows for precise identification of specific gene locations, providing a clear basis for further analysis or manipulation. This method can effectively identify and differentiate the presence and location of specific genes, possessing significant application value in areas such as genetic research, disease diagnosis, and personalized medicine.

[0067] In one embodiment, as shown in FIG2 , matching the genetic gene sequence and the episomal gene sequence with the template gene sequence respectively to determine the target gene sites in the genetic gene sequence and the episomal gene sequence includes:

[0068] Step S202 , matching the genetic gene sequence and the episomal gene sequence with the template gene sequence respectively, to obtain matching results corresponding to each base pair.

[0069] Specifically, the computer equipment matches the genetic gene sequence of the object to be tested with the template gene sequence for each base pair, thereby determining the target base pairs in the genetic gene sequence that are different from the corresponding base pairs in the template gene sequence. Similarly, the free gene sequence is matched with the template gene sequence to obtain the target base pairs in the free gene sequence that are different from the corresponding base pairs in the template gene sequence.

[0070] Step S204: determining the target base pair between the genetic gene sequence and the episomal gene sequence based on the matching result.

[0071] The target base pair is a base pair that does not match the base pair at the corresponding position in the template gene sequence in the genetic gene sequence and the free gene sequence.

[0072] Step S206: Using the target base pair as the target gene site of the genetic gene sequence and the episomal gene sequence.

[0073] In this embodiment, the genetic gene sequence and the free gene sequence are matched with the template gene sequence respectively to obtain matching results corresponding to the base pairs. Based on the matching results, the target base pairs of the genetic gene sequence and the free gene sequence are determined, and the target base pairs are used as the target gene sites of the genetic gene sequence and the free gene sequence. Therefore, the template gene sequence can be used as the comparison standard to accurately and quickly determine the target gene sites in the genetic gene sequence and the free gene sequence, thereby improving the recognition efficiency and accuracy of the gene sites.

[0074] In one embodiment, as shown in FIG3 , extracting characteristic data corresponding to target gene loci includes:

[0075] Step S302: Obtain the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site.

[0076] Among them, sequencing depth refers to the number of times a genomic region is sequenced (generally speaking, the higher the sequencing depth, the more accurate the test results, because a higher coverage depth can reduce detection errors and the probability of missed detection); sequencing quality is used to characterize the sequencing accuracy of the sequence sample data of the corresponding gene, and the position information indicates the specific position of the corresponding gene site on the corresponding chromosome.

[0077] Step S304: generating feature data corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site.

[0078] Specifically, the computer device extracts the corresponding characteristic parameters for the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the gene site determined in the above steps. The characteristic parameters can be used to characterize the above-mentioned characteristic attributes of the target gene site. The specific method can be that the computer device obtains the correspondence between the characteristic attributes of the gene site in the database (sequencing depth, sequencing quality, and chromosome number and position information) and the characteristic parameters, and then determines the target characteristic parameters corresponding to the target gene site based on the correspondence and the characteristic attributes of the target gene site. Finally, the target characteristic parameters are feature screened to determine the characteristic data corresponding to the target gene site.

[0079] In this embodiment, the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site are obtained respectively. Based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site, feature data corresponding to the target gene site is generated, thereby determining / extracting the corresponding feature data based on the characteristic attributes of the target gene site, removing a large amount of redundant data, retaining the feature data that can characterize the characteristic attributes of the target gene site, streamlining the input data of the model, reducing the computational burden of the data, and improving the computational efficiency of the data.

[0080] In one embodiment, as shown in FIG4 , based on the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene locus, feature data corresponding to the target gene locus is generated, including:

[0081] Step S402 : generating initial characteristic parameters corresponding to the target gene locus according to the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene locus.

[0082] Among them, the initial characteristic parameters can represent the characteristic attributes corresponding to the target gene site (sequencing depth, sequencing quality, and the number and position information of the chromosome where it is located). The initial characteristic parameters can be determined by a computer device based on the mapping / correspondence between the characteristic attributes of the gene site and the characteristic parameters.

[0083] Step S404: Based on the significance analysis model, each initial feature parameter is tested in sequence to obtain a significance analysis result corresponding to each initial feature parameter.

[0084] The significance analysis model is used to characterize the significance of the influence of each initial characteristic parameter on the corresponding gene prediction sequence. For example, the significance analysis model can be constructed using a linear logistic regression method.

[0085] Specifically, after the computer device constructs the significance analysis model, each initial feature parameter is added to the significance analysis model in sequence. Each time an initial feature parameter is added, the current corresponding significance analysis model is determined, and then the prediction effect (such as prediction accuracy) of the current significance analysis model is verified. In subsequent steps, the feature parameter among each initial feature parameter that can cause a significant improvement in the prediction effect of the corresponding significance analysis model is determined as the target feature parameter.

[0086] Step S406 : Based on the significance analysis results corresponding to the initial characteristic parameters, target characteristic parameters corresponding to the target gene loci are screened and determined as characteristic data corresponding to the target gene loci.

[0087] The screening method may be to determine the characteristic parameters that can significantly improve the prediction effect of the significance analysis model among the initial characteristic parameters as target characteristic parameters.

[0088] It can be understood that in this embodiment, each initial feature parameter is sequentially involved in constructing a significance analysis model, and then the significance of the feature parameter for the gene sequence prediction result is determined based on the prediction effect of the significance analysis model corresponding to each added feature parameter. For feature parameters whose significance exceeds a preset threshold, they have a higher significance impact on the gene sequence prediction result, while other initial feature parameters have a lower significance impact on the gene sequence prediction result, and are considered to be redundant data and are eliminated.

[0089] In this embodiment, initial characteristic parameters corresponding to the target gene site are generated based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site. Based on the significance analysis model, each initial characteristic parameter is tested in sequence to obtain a significance analysis result corresponding to each initial characteristic parameter. Based on the significance analysis result corresponding to each initial characteristic parameter, the target characteristic parameter corresponding to the target gene site is screened and obtained, and the target characteristic parameter is determined as the characteristic data corresponding to the target gene site. In this way, the initial characteristic parameters that have a significant impact on the prediction model can be determined based on the characteristic attributes of the target gene site, thereby eliminating redundant data and retaining the characteristic data that can characterize the characteristic attributes of the target gene site, streamlining the input data of the model, reducing the computational burden of the data, and improving the computational efficiency of the data.

[0090] In one embodiment, the first input data represents base sequence characteristics of the corresponding genetic gene sequence and the free gene sequence, and the second input data represents sequencing quality of the corresponding genetic gene sequence and the free gene sequence.

[0091] In one embodiment, as shown in FIG5 , the first input data and the second input data are input into the first network and the second network respectively, and a gene prediction result corresponding to the object to be tested is obtained by the third network output, including:

[0092] Step S502: Merge the output of the first network and the output of the second network, input them into the third network, and output at least two candidate gene prediction results.

[0093] Among them, each candidate gene prediction sequence includes a corresponding confidence level, which is used to characterize the probability that the corresponding candidate gene prediction sequence is a true gene sequence. The greater the confidence level, the greater the probability that the corresponding candidate gene prediction result is a true gene sequence, and vice versa.

[0094] Step S504: determining the corresponding target gene prediction result from each candidate gene prediction sequence based on the confidence level corresponding to each candidate gene prediction sequence.

[0095] Specifically, the computer device uses the candidate gene prediction result corresponding to the maximum confidence as the corresponding target gene prediction sequence based on the confidence corresponding to each candidate gene prediction result determined in the above steps. Optionally, when the candidate gene prediction result corresponding to the maximum confidence violates Mendel's law of inheritance (i.e., a gene mutation occurs, resulting in a new genotype) and cannot be matched to the gene sequence in the free gene sequence, the candidate gene prediction result corresponding to the maximum confidence is eliminated, and the target gene prediction result is re-determined in the remaining candidate gene prediction sequences.

[0096] Step S506: Using the target gene prediction sequence as the gene prediction result corresponding to the object to be tested.

[0097] In this embodiment, the output of the first network and the output of the second network are merged and input into the third network to output at least two candidate gene prediction sequences. Based on the confidence corresponding to each candidate gene prediction sequence, the corresponding target gene prediction result is determined from each candidate gene prediction sequence, and the target gene prediction sequence is used as the gene prediction sequence corresponding to the object to be tested. By introducing multiple output results and their corresponding confidence levels, the erroneous prediction results that violate Mendel's laws of inheritance that appear in the prediction process can be identified and eliminated, thereby effectively improving the reliability and accuracy of gene prediction.

[0098] In one embodiment, as shown in FIG6 , obtaining a target gene prediction model includes:

[0099] Step S602: Obtain the concentration corresponding to the free gene sequence.

[0100] The concentration is used to characterize the content of the free gene sequence of the object to be tested in the plasma of the corresponding genetic subject.

[0101] Specifically, the computer device may obtain the concentration corresponding to the free gene sequence from a database, or may obtain the above data from a third-party device, server, etc.

[0102] Step S604: Determine the corresponding target gene prediction model based on the gene type and concentration of the target gene locus.

[0103] Among them, the gene type of each gene locus refers to the gene type of the genetic gene sequence corresponding to each gene locus (that is, the gene sequence corresponding to the genetic parent of the object to be tested). The target gene prediction model is used to output the gene sequence prediction result of the object to be tested based on the genetic gene sequence and free gene sequence of the object to be tested. The gene sequence prediction result includes at least two candidate gene prediction sequences and their corresponding confidence levels, wherein the confidence level is used to characterize the corresponding candidate gene prediction sequence, and the sum of the confidence levels corresponding to all predicted gene sequences is 1.

[0104] Specifically, the computer device determines the corresponding target gene prediction model from each gene prediction model based on the gene type of the target gene site and the corresponding concentration of the free gene sequence, and then inputs the characteristic data into the target gene prediction model to obtain the model output, and then determines the gene prediction sequence corresponding to the object to be tested based on the model output, wherein each gene prediction model corresponds to each gene type within a different free gene sequence concentration range. For example, the concentration of the free gene sequence from the fetal part of the object to be tested is divided into 5 concentration ranges {FF<5%, 5%≤FF<10%, 10%≤FF<15%, 15%≤FF<20%, 20%≤FF}, and the gene types of the gene sequence are divided into ABAB, AAAB and ABAA, then each gene type corresponding to each concentration range corresponds to a gene prediction model.

[0105] In one embodiment, as shown in FIG7 , a method for generating a gene prediction model is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0106] Step S702 , obtaining the genetic gene sequence, episomal gene sequence, standard gene sequence and template gene sequence corresponding to the biological sample.

[0107] Among them, the free gene sequence represents the sequence of the free gene corresponding to the biological sample in the corresponding genetic object sample (such as peripheral blood, plasma), the standard gene sequence corresponds to the genetic gene sequence and the free gene sequence, and the genetic gene sequence is the gene sequence corresponding to the genetic parent of the object to be tested.

[0108] Step S704: determining target gene sites in the genetic gene sequence and the episomal gene sequence.

[0109] Among them, the target gene site is a specific base pair.

[0110] Specifically, the computer device can compare / match the genetic gene sequence and the free gene sequence with the template gene sequence for each corresponding base pair, and use the base pairs in the genetic gene sequence and the free gene sequence that fail to compare / match as their target gene sites.

[0111] Step S706: extracting feature data corresponding to the target gene locus.

[0112] Among them, feature data is used to characterize the attribute characteristics of the target gene site.

[0113] Specifically, the computer equipment extracts corresponding characteristic parameters based on the chromosome number and position information of the target gene site, as well as the sequencing depth and sequencing quality of the gene sequence where the target gene site is located. The characteristic parameters can reflect the chromosome number and position information of the corresponding target gene site, as well as the sequencing depth and sequencing quality of the gene sequence where the target gene site is located.

[0114] Step S708 : training the corresponding initial gene prediction model based on the first input data and the second input data to obtain the corresponding target gene prediction model.

[0115] Among them, the target gene prediction model includes a first network, a second network and a third network. The outputs of the first network and the second network are the inputs of the third network. The target gene prediction model is used to predict the gene results of the corresponding biological sample. The first network, the second network and the third network can adopt any neural network or deep learning model, such as a convolutional neural network, a fully connected neural network, etc.

[0116] The above-mentioned gene prediction model generation method obtains the genetic gene sequence, free gene sequence, standard gene sequence and template gene sequence corresponding to the biological sample; determines the target gene site in the genetic gene sequence and the free gene sequence; then extracts the feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, wherein the feature data includes first input data and second input data; based on the first input data and the second input data, the corresponding initial gene prediction model is trained to obtain the corresponding target gene prediction model; the target gene prediction model includes a first network, a second network and a third network, the outputs of the first network and the second network are the inputs of the third network, and the target gene prediction model is used to predict the gene result target gene site target gene site of the corresponding biological sample, so that the corresponding model can be trained more specifically according to data with different attribute characteristics, effectively improving the generalization ability of model training, and performing feature mining on the different attributes of the feature data, so that the global features and local features of the feature data can be fully learned, thereby improving the prediction efficiency and accuracy of the model.

[0117] In one embodiment, the standard gene sequence is a high-depth sequencing sequence corresponding to a biological sample, the standard gene sequence includes at least two reference gene sequences, and each reference gene sequence includes a corresponding standard confidence.

[0118] In one embodiment, determining a target gene site in a genetic gene sequence or an episomal gene sequence includes:

[0119] The genetic gene sequence and the free gene sequence are matched with the template gene sequence respectively to determine the target gene sites in the genetic gene sequence and the free gene sequence.

[0120] Specifically, the computer equipment matches the genetic gene sequence of the object to be tested with the template gene sequence for each base pair, thereby determining the target base pairs in the genetic gene sequence that are different from the corresponding base pairs in the template gene sequence. Similarly, the free gene sequence is matched with the template gene sequence to obtain the target base pairs in the free gene sequence that are different from the corresponding base pairs in the template gene sequence.

[0121] Target gene sites In this embodiment, the genetic gene sequence and the free gene sequence are matched with the template gene sequence respectively to obtain matching results corresponding to each base pair. Based on the matching results, each target base pair of the genetic gene sequence and the free gene sequence is determined, and each target base pair is used as the target gene site of the genetic gene sequence and the free gene sequence. Therefore, the template gene sequence can be used as the comparison standard to accurately and quickly determine the target gene sites in the genetic gene sequence and the free gene sequence, thereby improving the recognition efficiency and accuracy of the gene sites.

[0122] In one embodiment, as shown in FIG8 , matching the genetic gene sequence and the episomal gene sequence with the template gene sequence respectively to determine the target gene sites in the genetic gene sequence and the episomal gene sequence includes:

[0123] Step S802: Based on the matching results, determine each target base pair of the genetic gene sequence and the free gene sequence.

[0124] The target base pair is a base pair that does not match the base pair at the corresponding position in the template gene sequence in the genetic gene sequence and the free gene sequence.

[0125] Step S804: using each target base pair as a target gene site of the genetic gene sequence and the episomal gene sequence.

[0126] In this embodiment, the genetic gene sequence and the free gene sequence are matched with the template gene sequence respectively to obtain the matching results corresponding to each base pair. Based on the matching results, the target base pairs of the genetic gene sequence and the free gene sequence are determined, and each target base pair is used as the target gene site of the genetic gene sequence and the free gene sequence. Therefore, the template gene sequence can be used as the comparison standard to accurately and quickly determine the target gene site in the genetic gene sequence and the free gene sequence, thereby improving the recognition efficiency and accuracy of the gene site.

[0127] In one embodiment, as shown in FIG9 , extracting characteristic data corresponding to the target gene locus includes:

[0128] Step S902: Obtain the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site.

[0129] Among them, sequencing depth refers to the number of times a genomic region is sequenced (generally speaking, the higher the sequencing depth, the more accurate the test results, because a higher coverage depth can reduce detection errors and the probability of missed detection); sequencing quality is used to characterize the sequencing accuracy of the sequence sample data of the corresponding gene, and the position information indicates the specific position of the corresponding gene site on the corresponding chromosome.

[0130] Step S904: generating feature data corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site.

[0131] Specifically, the computer device extracts the corresponding characteristic parameters for the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the gene site determined in the above steps. The various characteristic parameters can be used to characterize the above-mentioned characteristic attributes of the target gene site. The specific method can be that the computer device obtains the correspondence between the characteristic attributes of the gene site in the database (sequencing depth, sequencing quality, and chromosome number and position information) and the characteristic parameters, and then determines the target characteristic parameters corresponding to the target gene site based on the correspondence and the characteristic attributes of the target gene site. Finally, the target characteristic parameters are feature screened to determine the characteristic data corresponding to the target gene site.

[0132] In this embodiment, the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site are obtained respectively. Based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site, feature data corresponding to the target gene site is generated, thereby determining / extracting the corresponding feature data based on the characteristic attributes of the target gene site, removing a large amount of redundant data, retaining the feature data that can characterize the characteristic attributes of the target gene site, streamlining the input data of the model, reducing the computational burden of the data, and improving the computational efficiency of the data.

[0133] In one embodiment, as shown in FIG10 , based on the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene locus, feature data corresponding to the target gene locus is generated, including:

[0134] Step S1002 : generating initial characteristic parameters corresponding to the target gene locus according to the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene locus.

[0135] Among them, the initial characteristic parameters can represent the characteristic attributes corresponding to the target gene site (sequencing depth, sequencing quality, and the number and position information of the chromosome where it is located). The initial characteristic parameters can be determined by a computer device based on the mapping / correspondence between the characteristic attributes of the gene site and the characteristic parameters.

[0136] Step S1004: Based on the significance analysis model, each initial feature parameter is tested in sequence to obtain a significance analysis result corresponding to each initial feature parameter.

[0137] The significance analysis model is used to characterize the significance of the influence of each initial characteristic parameter on the corresponding gene prediction sequence. For example, the significance analysis model can be constructed using a linear logistic regression method.

[0138] Specifically, after the computer device constructs the significance analysis model, each initial feature parameter is added to the significance analysis model in sequence. Each time an initial feature parameter is added, the current corresponding significance analysis model is determined, and then the prediction effect (such as prediction accuracy) of the current significance analysis model is verified. In subsequent steps, the feature parameter among each initial feature parameter that can cause a significant improvement in the prediction effect of the corresponding significance analysis model is determined as the target feature parameter.

[0139] Step S1006 , based on the significance analysis results corresponding to the initial characteristic parameters, target characteristic parameters corresponding to the target gene loci are screened and determined as characteristic data corresponding to the target gene loci.

[0140] The screening method may be to determine the characteristic parameters that can significantly improve the prediction effect of the significance analysis model among the initial characteristic parameters as target characteristic parameters.

[0141] It can be understood that in this embodiment, each initial feature parameter is sequentially involved in constructing a significance analysis model, and then the significance of the feature parameter for the gene sequence prediction result is determined based on the prediction effect of the significance analysis model corresponding to each added feature parameter. For feature parameters whose significance exceeds a preset threshold, they have a higher significance impact on the gene sequence prediction result, while other initial feature parameters have a lower significance impact on the gene sequence prediction result, and are considered to be redundant data and are eliminated.

[0142] In this embodiment, initial characteristic parameters corresponding to the target gene site are generated based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site. Based on the significance analysis model, each initial characteristic parameter is tested in sequence to obtain a significance analysis result corresponding to each initial characteristic parameter. Based on the significance analysis result corresponding to each initial characteristic parameter, the target characteristic parameter corresponding to the target gene site is screened and obtained, and the target characteristic parameter is determined as the characteristic data corresponding to the target gene site. In this way, the initial characteristic parameters that have a significant impact on the prediction model can be determined based on the characteristic attributes of the target gene site, thereby eliminating redundant data and retaining the characteristic data that can characterize the characteristic attributes of the target gene site, streamlining the input data of the model, reducing the computational burden of the data, and improving the computational efficiency of the data.

[0143] In one embodiment, as shown in FIG11 , the corresponding initial gene prediction model is trained based on the feature data to obtain the corresponding target gene prediction model, including:

[0144] Step S1102, dividing the matched feature data in the feature data into the same feature sub-data set based on the gene type of the target gene locus and the concentration of the free gene sequence; and

[0145] Specifically, matching feature data within the signature data are grouped into the same feature sub-dataset based on the genotype and episomal concentration of the target locus. This process involves analyzing the signature data, identifying data points that match a given genotype and episomal concentration, and organizing these data points into a separate feature sub-dataset. This approach makes the analysis and processing of relevant feature data more centralized and efficient, providing a precise data foundation for subsequent genetic analysis.

[0146] Step S1104 : Based on each feature sub-data set and its corresponding standard gene sequence, a corresponding initial gene prediction model is determined and trained to obtain a target gene prediction model corresponding to each feature sub-data set.

[0147] Based on each feature subset dataset and its corresponding standard gene sequence, a corresponding initial gene prediction model is determined and trained. This process involves using the data from each feature subset dataset and the corresponding standard gene sequence as training input to build and optimize the prediction model. In this way, each feature subset dataset will develop a specific target gene prediction model that can accurately predict the expression and variation of specific gene types in that dataset. In this way, models customized for different feature subsets can improve the accuracy and efficiency of gene prediction.

[0148] In this embodiment, matching data within the signature data are first assigned to corresponding feature sub-data sets based on the genotype and episomal concentration of the target gene loci. Subsequently, an initial gene prediction model is determined and trained for each feature sub-data set and its corresponding standard gene sequence, ultimately generating a corresponding target gene prediction model. This method enables precise processing and prediction of data based on complex genetic information, improving prediction accuracy and efficiency. This ensures the specificity and adaptability of the prediction model, particularly when processing data with varying genotypes and concentrations.

[0149] In one embodiment, as shown in FIG12 , based on each feature sub-data set and its corresponding standard gene sequence, a corresponding initial gene prediction model is determined and trained to obtain a target gene prediction model corresponding to each feature sub-data set, including:

[0150] Among them, the initial gene prediction model includes a first network, a second network and a third network, and the first network, the second network and the third network can arbitrarily adopt various neural networks, deep learning networks, such as fully connected layer networks, convolutional layer networks, etc., without specific restrictions.

[0151] Step S1202 : for the same feature sub-data set, data division is performed on the corresponding feature sub-data set according to attribute features to obtain a first input sample and a second input sample.

[0152] Among them, the first training sample represents the base sequence characteristics of the corresponding genetic gene sequence and free gene sequence; the second training sample represents the sequencing quality of the corresponding genetic gene sequence and free gene sequence; among them, the attribute characteristics include base sequence characteristics, sequencing quality characteristics, and whether the genetic gene sequence and free gene sequence are homozygous genes or heterozygous genes.

[0153] Step S1204: input the first input sample into the first network to obtain a corresponding first network output.

[0154] Step S1206, inputting the second input sample into the second network to obtain a corresponding second network output;

[0155] Step S1208: merging the first network output and the second network output, and inputting the combined outputs into a third network to obtain a third network output.

[0156] Step S1210: Calculate the difference between the output of the third network and the corresponding standard gene sequence to obtain the model loss.

[0157] Step S1212: Adjust the parameter values ​​of the first network, the second network, and the third network based on the model loss until the model training conditions are met, then stop training the corresponding initial gene prediction model to obtain the corresponding target gene prediction model.

[0158] In this embodiment, for the same feature sub-data set, data partitioning is performed to obtain a first input sample and a second input sample, and then the first input sample is input into the first network to obtain the corresponding first network output, the second input sample is input into the second network to obtain the corresponding second network output, the second input sample is input into the second network to obtain the corresponding second network output, the first network output and the second network output are merged and input into the third network to obtain the third network output, and the model loss is calculated based on the difference between the third network output and the corresponding standard gene sequence. The parameter values ​​of the first network, the second network, and the third network are adjusted based on the model loss until the model training conditions are met, and the training of the corresponding initial gene prediction model is stopped to obtain the corresponding target gene prediction model, so that the corresponding model can be trained more specifically according to data with different attribute characteristics, effectively improving the generalization ability of model training, and performing feature mining on the different attributes of the feature data, so as to fully learn the global features and local features of the feature data, thereby improving the prediction efficiency and accuracy of the model.

[0159] This application also provides an application scenario, which applies the above-mentioned gene prediction method, as shown in Figure 13, and is applied to the prediction scenario of fetal genotype point mutations. Specifically, the application of the gene prediction method in this application scenario is as follows:

[0160] 1. Construction of gene prediction model

[0161] In this example, high-depth sequencing of plasma cell-free DNA (cf-DNA, corresponding to the aforementioned cell-free gene sequence) of pregnant women from 10 families was performed in the early stage, as well as whole-genome sequencing data of the pregnant women and their husbands (corresponding to the aforementioned genetic gene sequence), and the standard gene sequence of the fetal genome was obtained through umbilical cord blood.

[0162] Based on a base pair alignment of the corresponding template gene sequence with both cf-DNA and the parental genome, each locus was identified. 67 initial feature data points were then extracted, encompassing the sequencing depth, sequencing quality, chromosome number, and position of the locus. Since there are only four bases, a single base pair point mutation can have only 4*4=16 possible alleles. Regardless of the order of arrangement (e.g., AT and TA belong to the same allele), the classification results for 10 base pair alleles were obtained, as shown in Table 1 below.

[0163] Table 1

[0164] A linear logistic regression model is used to perform a significance analysis on the 67 initial feature parameters determined in the above steps to determine the significance of each feature parameter corresponding to the prediction of the gene sequence effect. Each initial feature parameter with a significance greater than a preset threshold is used as feature data. In this embodiment, a significance analysis is performed on the 67 initial feature parameters to obtain 56 target feature parameters, which are used as feature data corresponding to each gene locus.

[0165] Then, based on the attribute characteristics of the feature data, the feature data corresponding to the 56 target feature parameters are divided into the first group of sample data (30 target feature parameters) and the second group of sample data (26 target feature parameters). The first group of sample data includes 24 target feature parameters for characterizing the base pair sequence, 3 target feature parameters for characterizing the gene type of the parental genome, and 2 target feature parameters for characterizing the data quality of the parental genome and free gene sequence. The second group of sample data characterizes the control parameters related to the sequencing quality of the corresponding parental genome and free gene sequence.

[0166] The first and second groups of sample data are then partitioned based on the concentration of the episomal sequence and the genotype of the parental genome to obtain characteristic sub-data sets, each of which includes the first and second sub-sample data sets.

[0167] For the same feature sub-data set, the first group of sub-sample data in the feature sub-data set is input into the first network, the second group of sub-sample data is input into the second network, and then the network output of the first network and the network output of the second network are merged and input into the third network to obtain the third network output, and then the model loss is constructed based on the third network output and the standard gene sequence. The target parameters of the first network, the second network and the third network are dynamically adjusted based on the model loss until the model loss meets the model training stop condition to obtain the target gene prediction model; in one embodiment, the first network and the third network are both fully connected layer networks, and the second network is a convolutional layer network. Optionally, as shown in Figure 14, the types of the first network, the second network and the third network can include fully connected layer networks, convolutional layer networks, activation layer networks, etc.

[0168] In this embodiment, the target gene site is quickly and accurately identified / determined based on the differences between the parental genome and the free gene sequence and the corresponding template gene sequence, and the corresponding gene prediction model is accurately located based on the gene type of the target gene site and the concentration characteristics of the free gene sequence, thereby completing the prediction of the gene sequence of the object to be tested, and completing the combination of the characteristic attributes of the sample data itself with the specific model, which can more accurately learn the base characteristics of each gene site in the sample gene sequence, thereby improving the prediction accuracy of the gene sequence.

[0169] 2. Application of gene prediction models

[0170] The parental genome sequences corresponding to the fetus to be tested and the peripheral blood free gene sequences of the pregnant woman are obtained to construct the input sample data.

[0171] The input sample data is then compared with the corresponding template gene sequence, and all base pairs with inconsistent comparison results in the input sample data are used as target gene sites. The initial feature parameters corresponding to each target gene site are extracted, and then the linear logistic regression model is used to screen each initial feature parameter to obtain the input feature data.

[0172] According to the concentration characteristics of the free gene sequence in the peripheral blood of the pregnant woman and the genotype of the parental genome sequence, the corresponding target gene prediction model is determined, and then the input feature data is input into the target gene prediction model to output each candidate gene prediction sequence, where each candidate gene prediction sequence corresponds to a confidence level (the sum of the confidence levels of each gene prediction sequence is 1). Finally, the candidate gene prediction sequence with the highest confidence level is used as the gene prediction sequence of the fetus to be tested, as shown in Figure 15.

[0173] In this embodiment, feature data corresponding to target gene sites in parental genome sequences and episomal gene sequences are extracted, and based on the concentration characteristics corresponding to the episomal gene sequences and the gene types of the target gene sites, the feature data are divided into subsets, and then the corresponding gene prediction models are trained based on each feature sub-data set. This allows for more targeted training of corresponding models based on data with different attribute characteristics, effectively improving the generalization ability of model training, and performing feature mining on different attributes of the feature data, thereby enabling full learning of the global and local features of the feature data and improving the prediction efficiency and accuracy of the model.

[0174] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0175] In one embodiment, as shown in FIG16 , a gene prediction device is provided. The device may be implemented as a software module or a hardware module, or a combination of both, as part of a computer device. The device specifically includes: an acquisition module 1602 , an extraction module 1604 , and a prediction module 1606 , wherein:

[0176] An acquisition module 1602 is configured to acquire a template gene sequence and a genetic gene sequence and an episomal gene sequence corresponding to the object to be tested, wherein the episomal gene sequence represents a sequence of an episomal gene corresponding to the object to be tested in a corresponding genetic object sample;

[0177] Extraction module 1604 is used to determine the target gene site in the genetic gene sequence and the free gene sequence; extract feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes first input data and second input data;

[0178] Prediction module 1606 is used to obtain a target gene prediction model, which includes a first network, a second network, and a third network. The outputs of the first network and the second network are input to the third network; and the first input data and the second input data are input to the first network and the second network respectively, to obtain the gene prediction result corresponding to the object to be tested output by the third network.

[0179] In one embodiment, the genetic gene sequence corresponding to the subject to be tested is the genotype of the genetic subject corresponding to the subject to be tested.

[0180] In one embodiment, the extraction module 1604 is further configured to match the genetic gene sequence and the episomal gene sequence with the template gene sequence, respectively, to determine the target gene sites in the genetic gene sequence and the episomal gene sequence.

[0181] In one embodiment, the extraction module 1604 is also used to match the genetic gene sequence and the free gene sequence with the template gene sequence respectively to obtain matching results corresponding to each base pair; based on the matching results, determine the target base pair of the genetic gene sequence and the free gene sequence, the target base pair being the base pair in the genetic gene sequence and the free gene sequence that does not match the base pair at the corresponding position in the template gene sequence; and use the target base pair as the target gene point of the genetic gene sequence and the free gene sequence.

[0182] In one embodiment, the extraction module 1604 is also used to obtain the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site, where the sequencing quality is used to characterize the sequencing accuracy of the sequence sample data of the corresponding gene; and based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site, generate characteristic data corresponding to the target gene site.

[0183] In one embodiment, the extraction module 1604 is also used to generate initial characteristic parameters corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site; based on the significance analysis model, each initial characteristic parameter is tested in turn to obtain the significance analysis results corresponding to each initial characteristic parameter, and the significance analysis model is used to characterize the significance of the influence of each initial characteristic parameter on the corresponding gene prediction sequence; and based on the significance analysis results corresponding to each initial characteristic parameter, the target characteristic parameters corresponding to the target gene site are screened and determined as the characteristic data corresponding to the target gene site.

[0184] In one embodiment, the first input data represents base sequence characteristics of the corresponding genetic gene sequence and the free gene sequence, and the second input data represents sequencing quality of the corresponding genetic gene sequence and the free gene sequence.

[0185] In one embodiment, the prediction module 1606 is further used to merge the first network output and the second network output, input them into the third network, and output at least two candidate gene prediction sequences, each candidate gene prediction sequence including a corresponding confidence level; based on the confidence level corresponding to each candidate gene prediction sequence, determine the corresponding target gene prediction result from each candidate gene prediction sequence; and use the target gene prediction result as the gene prediction result corresponding to the object to be tested.

[0186] In one embodiment, the prediction module 1606 is further used to obtain the concentration corresponding to the free gene sequence, which is used to characterize the content of the free gene sequence of the subject to be tested in the plasma of the corresponding genetic subject; based on the gene type and concentration of the target gene site, the corresponding target gene prediction model is determined.

[0187] The above-mentioned gene prediction device obtains the template gene sequence and the genetic gene sequence and free gene sequence corresponding to the object to be tested, determines the target gene site in the genetic gene sequence and the free gene sequence, extracts the characteristic data corresponding to the target gene site, and the characteristic data is used to characterize the attribute characteristics of the target gene site. The characteristic data includes first input data and second input data, obtains a target gene prediction model, and the target gene prediction model includes a first network, a second network and a third network. The outputs of the first network and the second network are inputs of the third network; the first input data and the second input data are respectively input into the first network and the second network to obtain the gene prediction result corresponding to the object to be tested output by the third network. In this way, the target gene site can be quickly and accurately identified / determined based on the genetic gene sequence and the free gene sequence, and according to the obtained target gene prediction model, the first input data and the second input data are respectively input into the first network and the second network of the target gene prediction model, and the gene prediction result corresponding to the object to be tested is obtained by the output of the third network, thereby completing the prediction of the gene sequence of the object to be tested, completing the combination of the characteristic attributes of the sample data itself with the specific model, and more accurately learning the base characteristics of each gene site in the sample gene sequence, thereby improving the prediction accuracy of the gene sequence.

[0188] The specific definition of the gene prediction device can be found in the definition of the gene prediction method above and will not be repeated here. The various modules in the above-mentioned gene prediction device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0189] In one embodiment, as shown in FIG17 , a gene prediction model generation device is provided. The device may be a software module or a hardware module, or a combination of both, as part of a computer device. The device specifically includes: an acquisition module 1702, an extraction module 1704, and a training module 1706, wherein:

[0190] An acquisition module 1702 is configured to acquire a genetic gene sequence, an episomal gene sequence, a standard gene sequence, and a template gene sequence corresponding to a biological sample, wherein the episomal gene sequence represents a sequence of an episomal gene corresponding to the biological sample in the corresponding genetic object sample, and the standard gene sequence corresponds to the genetic gene sequence and the episomal gene sequence;

[0191] Extraction module 1704 is used to determine the target gene site in the genetic gene sequence and the free gene sequence; extract feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes the first input data and the second input data;

[0192] The training module 1706 is used to train the corresponding initial gene prediction model based on the feature data to obtain the corresponding target gene prediction model. The gene prediction model is used to predict the gene results of the corresponding biological sample.

[0193] In one embodiment, the standard gene sequence is a high-depth sequencing sequence corresponding to a biological sample, the standard gene sequence includes at least two reference gene sequences, and each reference gene sequence includes a corresponding standard confidence.

[0194] In one embodiment, 1704 is also used to match the genetic gene sequence and the free gene sequence with the template gene sequence respectively, and determine the target gene site in the genetic gene sequence and the free gene sequence.

[0195] In one embodiment, the extraction module 1704 is also used to match the genetic gene sequence and the free gene sequence with the template gene sequence respectively to obtain matching results corresponding to each base pair; based on the matching results, determine each target base pair of the genetic gene sequence and the free gene sequence, the target base pair being the base pair in the genetic gene sequence and the free gene sequence that does not match the base pair at the corresponding position in the template gene sequence; and use each target base pair as the target gene point of the genetic gene sequence and the free gene sequence.

[0196] In one embodiment, the extraction module 1704 is also used to obtain the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site, where the sequencing quality is used to characterize the sequencing accuracy of the sequence sample data of the corresponding gene; and based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site, generate characteristic data corresponding to the target gene site.

[0197] In one embodiment, the extraction module 1704 is also used to generate initial characteristic parameters corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site; based on the significance analysis model, each initial characteristic parameter is tested in turn to obtain the significance analysis results corresponding to the initial characteristic parameters, and the significance analysis model is used to characterize the significance of the influence of each initial characteristic parameter on the corresponding gene prediction sequence; and based on the significance analysis results corresponding to each initial characteristic parameter, the target characteristic parameters corresponding to the target gene site are screened and determined as the characteristic data corresponding to the target gene site.

[0198] In one embodiment, the extraction module 1604 is also used to divide the matching feature data in the feature data into the same feature sub-data set based on the gene type of the target gene site and the concentration of the free gene sequence; and determine and train the corresponding initial gene prediction model based on each feature sub-data set and its corresponding standard gene sequence, to obtain the target gene prediction model corresponding to each feature sub-data set.

[0199] In one embodiment, the extraction module 1704 is also used to perform data division on the corresponding feature sub-data set according to the attribute characteristics for the same feature sub-data set to obtain a first input sample and a second input sample, wherein the first training sample represents the base sequence characteristics of the corresponding genetic gene sequence and the free gene sequence, and the second training sample represents the sequencing quality of the corresponding genetic gene sequence and the free gene sequence; the first input sample is input into the first network to obtain the corresponding first network output; the second input sample is input into the second network to obtain the corresponding second network output; the first network output and the second network output are merged and input into the third network to obtain the third network output; the model loss is obtained based on the difference between the third network output and the corresponding standard gene sequence; and the parameter values ​​of the first network, the second network and the third network are adjusted based on the model loss until the model training conditions are met, and the training of the corresponding initial gene prediction model is stopped to obtain the corresponding target gene prediction model, wherein the initial gene prediction model includes the first network, the second network and the third network.

[0200] The above-mentioned gene prediction model generation device obtains the genetic gene sequence, free gene sequence, standard gene sequence and template gene sequence corresponding to the biological sample; determines the target gene site in the genetic gene sequence and the free gene sequence; then extracts the feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, wherein the feature data includes first input data and second input data; based on the first input data and the second input data, the corresponding initial gene prediction model is trained to obtain the corresponding target gene prediction model; the target gene prediction model includes a first network, a second network and a third network, the outputs of the first network and the second network are the inputs of the third network, and the target gene prediction model is used to predict the gene result target gene site target gene site of the corresponding biological sample, so that the corresponding model can be trained more specifically according to data with different attribute characteristics, effectively improving the generalization ability of model training, and performing feature mining on the different attributes of the feature data, so that the global features and local features of the feature data can be fully learned, thereby improving the prediction efficiency and accuracy of the model.

[0201] For the specific definition of the gene prediction model generation device, please refer to the definition of the gene prediction model generation method above, which will not be repeated here. The various modules in the above-mentioned gene prediction model generation device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0202] In an exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be shown in Figure 18. The computer device includes a processor, memory, an input / output (I / O) interface, and a communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store feature data corresponding to gene loci. The I / O interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via a network connection. When executed by the processor, the computer program implements a gene prediction method or a gene prediction model generation method.

[0203] In an exemplary embodiment, a computer device is provided, which may be a terminal. A diagram of its internal structure may be shown in FIG19 . The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and an external device. The communication interface of the computer device is configured to communicate with an external terminal via wired or wireless communication, where the wireless communication may be achieved via Wi-Fi, a mobile cellular network, NFC (near field communication), or other technologies. When executed by the processor, the computer program implements a gene prediction method or a gene prediction model generation method. The display unit of the computer device is configured to form a visually visible image, and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0204] Those skilled in the art will understand that the structures shown in Figures 18 and 19 are merely block diagrams of partial structures related to the scheme of the present application, and do not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device may include more or fewer components than shown in the figures, or combine certain components, or have a different arrangement of components.

[0205] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0206] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0207] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of each of the above-described method embodiments.

[0208] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0209] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0210] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A gene prediction method, characterized in that: Executed by a computer device, the method includes: Obtaining a template gene sequence and a genetic gene sequence and an episomal gene sequence corresponding to the object to be tested, wherein the episomal gene sequence represents a sequence of the episomal gene corresponding to the object to be tested in the corresponding genetic object sample; Determining the target gene site in the genetic gene sequence and the free gene sequence; Extracting feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, the feature data includes first input data and second input data, and; Obtaining a target gene prediction model, the target gene prediction model comprising a first network, a second network, and a third network, wherein outputs of the first network and the second network serve as inputs of the third network; and The first input data and the second input data are input into the first network and the second network respectively, and a gene prediction result corresponding to the object to be tested is obtained which is output by the third network.

2. The method according to claim 1, characterized in that The genetic gene sequence corresponding to the object to be tested is the genotype of the genetic object corresponding to the object to be tested.

3. The method according to claim 1, characterized in that The determining of the target gene sites in the genetic gene sequence and the episomal gene sequence includes: The genetic gene sequence and the free gene sequence are matched with the template gene sequence respectively to determine the target gene sites in the genetic gene sequence and the free gene sequence.

4. The method according to claim 3, characterized in that The step of matching the genetic gene sequence and the episomal gene sequence with the template gene sequence to determine the target gene sites in the genetic gene sequence and the episomal gene sequence comprises: Matching the genetic gene sequence and the episomal gene sequence with the template gene sequence to obtain matching results corresponding to each base pair; Based on the matching result, determining a target base pair between the genetic gene sequence and the episomal gene sequence, the target base pair being a base pair between the genetic gene sequence and the episomal gene sequence that does not match a base pair at a corresponding position in the template gene sequence; and The target base pair is used as the target gene site of the genetic gene sequence and the free gene sequence.

5. The method according to claim 1, wherein The step of extracting characteristic data corresponding to the target gene locus includes: Obtaining the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site, wherein the sequencing quality is used to characterize the sequencing accuracy of the sequence sample data of the corresponding gene; and Based on the sequencing depth, sequencing quality, and chromosome number and position information corresponding to the target gene site, feature data corresponding to the target gene site is generated.

6. The method according to claim 5, characterized in that Generating characteristic data corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site includes: Generate initial characteristic parameters corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site; Based on the significance analysis model, each initial characteristic parameter is tested in turn to obtain a significance analysis result corresponding to each initial characteristic parameter, wherein the significance analysis model is used to characterize the significance of the influence of each initial characteristic parameter on the corresponding gene prediction sequence; and Based on the significance analysis results corresponding to the various initial characteristic parameters, target characteristic parameters corresponding to the target gene loci are screened and obtained, and the target characteristic parameters are determined as characteristic data corresponding to the target gene loci.

7. The method according to claim 1, characterized in that The first input data represents the base sequence characteristics of the corresponding genetic gene sequence and the free gene sequence, and the second input data represents the sequencing quality of the corresponding genetic gene sequence and the free gene sequence.

8. The method according to claim 1, characterized in that The step of inputting the first input data and the second input data into the first network and the second network respectively, and obtaining a gene prediction result corresponding to the object to be tested outputted by the third network, comprises: Merging the output of the first network and the output of the second network, inputting the results into the third network, and outputting at least two candidate gene prediction sequences, each candidate gene prediction sequence including a corresponding confidence score; Determining corresponding target gene prediction results from each candidate gene prediction sequence based on the confidence level corresponding to each candidate gene prediction sequence; and The target gene prediction result is used as the gene prediction result corresponding to the object to be tested.

9. The method according to claim 1, characterized in that The obtaining of the target gene prediction model comprises: Obtaining the concentration corresponding to the free gene sequence, wherein the concentration is used to characterize the content of the free gene sequence of the subject to be tested in the plasma of the corresponding genetic subject; Based on the gene type and the concentration of the target gene locus, a corresponding target gene prediction model is determined.

10. A method for generating a gene prediction model, characterized in that: Executed by a computer device, the method includes: Obtaining a genetic gene sequence, an episomal gene sequence, a standard gene sequence, and a template gene sequence corresponding to the biological sample, wherein the episomal gene sequence represents a sequence of an episomal gene corresponding to the biological sample in the corresponding genetic object sample, and the standard gene sequence corresponds to the genetic gene sequence and the episomal gene sequence; Determine the target gene site in the genetic gene sequence and the free gene sequence; extract feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, the feature data includes first input data and second input data, and; The corresponding initial gene prediction model is trained based on the first input data and the second input data to obtain a corresponding target gene prediction model; the target gene prediction model includes a first network, a second network and a third network, the outputs of the first network and the second network are the inputs of the third network, and the target gene prediction model is used to predict the gene results of the corresponding biological sample.

11. The method according to claim 10, characterized in that The standard gene sequence is a high-depth sequencing sequence corresponding to a biological sample. The standard gene sequence includes at least two reference gene sequences, and each reference gene sequence includes a corresponding standard confidence.

12. The method according to claim 10, characterized in that The determining of the target gene sites in the genetic gene sequence and the episomal gene sequence includes: The genetic gene sequence and the free gene sequence are matched with the template gene sequence respectively to determine the target gene sites in the genetic gene sequence and the free gene sequence.

13. The method according to claim 12, characterized in that The step of matching the genetic gene sequence and the episomal gene sequence with the template gene sequence to determine the target gene sites in the genetic gene sequence and the episomal gene sequence comprises: Matching the genetic gene sequence and the episomal gene sequence with the template gene sequence to obtain matching results corresponding to each base pair; Based on the matching results, determining each target base pair between the genetic gene sequence and the episomal gene sequence, the target base pair being a base pair between the genetic gene sequence and the episomal gene sequence that does not match a base pair at a corresponding position in the template gene sequence; and Each target base pair is used as a target gene site of the genetic gene sequence and the episomal gene sequence.

14. The method according to claim 10, characterized in that The step of extracting characteristic data corresponding to the target gene locus includes: Obtaining the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site, wherein the sequencing quality is used to characterize the sequencing accuracy of the sequence sample data of the corresponding gene; and Based on the sequencing depth, sequencing quality, and chromosome number and location information of the target gene site, Generate characteristic data corresponding to the target gene site.

15. The method according to claim 14, characterized in that Generating characteristic data corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site includes: Generate initial characteristic parameters corresponding to the target gene site based on the sequencing depth, sequencing quality, and chromosome number and location information corresponding to the target gene site; Based on the significance analysis model, each initial characteristic parameter is tested in turn to obtain a significance analysis result corresponding to the initial characteristic parameter, wherein the significance analysis model is used to characterize the significance of the influence of each initial characteristic parameter on the corresponding gene prediction sequence; and Based on the significance analysis results corresponding to the various initial characteristic parameters, target characteristic parameters corresponding to the target gene loci are screened and obtained, and the target characteristic parameters are determined as characteristic data corresponding to the target gene loci.

16. The method according to claim 15, characterized in that The training of the corresponding initial gene prediction model based on the feature data to obtain the corresponding target gene prediction model includes: Dividing the matched feature data in the feature data into the same feature sub-data set based on the gene type of the target gene locus and the concentration of the free gene sequence; and Based on each feature sub-data set and its corresponding standard gene sequence, the corresponding initial gene prediction model is determined and trained to obtain the target gene prediction model corresponding to each feature sub-data set.

17. The method according to claim 16, characterized in that The method of determining and training the corresponding initial gene prediction model based on each feature sub-data set and its corresponding standard gene sequence to obtain the target gene prediction model corresponding to each feature sub-data set includes: The initial gene prediction model includes a first network, a second network and a third network; For the same feature sub-data set, the corresponding feature sub-data set is divided according to the attribute characteristics to obtain a first input sample and a second input sample, wherein the first training sample represents the base sequence characteristics of the corresponding genetic gene sequence and the free gene sequence, and the second training sample represents the sequencing quality of the corresponding genetic gene sequence and the free gene sequence; Inputting the first input sample into the first network to obtain a corresponding first network output; Inputting the second input sample into the second network to obtain a corresponding second network output; Merging the first network output and the second network output, and inputting the combined outputs into the third network to obtain a third network output; Calculating the difference between the output of the third network and the corresponding standard gene sequence to obtain a model loss; and The parameter values ​​of the first network, the second network, and the third network are adjusted based on the model loss until the model training conditions are met, and the training of the corresponding initial gene prediction model is stopped to obtain the corresponding target gene prediction model.

18. A gene prediction device, characterized in that: The device comprises: An acquisition module is used to obtain a template gene sequence and a genetic gene sequence and an episomal gene sequence corresponding to the object to be tested, wherein the episomal gene sequence represents a sequence of the episomal gene corresponding to the object to be tested in the corresponding genetic object sample; An extraction module is used to determine the target gene site in the genetic gene sequence and the free gene sequence; extract feature data corresponding to the target gene site, wherein the feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes first input data and second input data; A prediction module is used to obtain a target gene prediction model, wherein the target gene prediction model includes a first network, a second network, and a third network, wherein the outputs of the first network and the second network are inputs to the third network; and the first input data and the second input data are input to the first network and the second network respectively, to obtain a gene prediction result corresponding to the object to be tested output by the third network.

19. A gene prediction model generation device, characterized in that: The device comprises: An acquisition module is used to obtain a genetic gene sequence, an episomal gene sequence, a standard gene sequence, and a template gene sequence corresponding to a biological sample, wherein the episomal gene sequence represents a sequence of an episomal gene corresponding to the biological sample in a corresponding genetic object sample, and the standard gene sequence corresponds to the genetic gene sequence and the episomal gene sequence; An extraction module is used to determine the target gene site in the genetic gene sequence and the free gene sequence; extract feature data corresponding to the target gene site, the feature data is used to characterize the attribute characteristics of the target gene site, and the feature data includes first input data and second input data; The training module is used to train the corresponding initial gene prediction model based on the feature data to obtain the corresponding target gene prediction model, and the gene prediction model is used to predict the gene results of the corresponding biological sample.

20. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, wherein: When the processor executes the computer-readable instructions, the steps of the method according to any one of claims 1 to 17 are implemented.

21. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 17 are implemented.

Citation Information

Patent Citations

  • Fetal genomic analysis from a maternal biological sample

    CN102770558A

  • Noninvasive detection of fetal genetic abnormality

    CN103403183A

  • Method for detecting genetic variation

    CN104204220A

  • Method and apparatus for determining fetus target area haplotype

    CN105648045A

  • Deafness haplotype gene mutation non-invasive detection method

    CN112126677A