Transgenic event discrimination method based on targeted sequencing and exogenous sequence insertion position
Through the method based on targeted sequencing and exogenous sequence insertion location, combined with MMoE model and feature parameter extraction, the problem of insufficient accuracy of existing transgenic event detection methods is solved, and efficient and accurate transgenic event discrimination and homozygous/heterozygous distinction are achieved.
Patent Information
- Application Number
- CN202411989440.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The existing methods for detecting and discriminating genetically modified event have problems such as insufficient accuracy, difficulty in selecting primers, low quantitative accuracy and false positive contamination, making it difficult to effectively distinguish homozygous and heterozygous transgenic plants.
Using a transgenic event discrimination method based on targeted sequencing and exogenous sequence insertion location, the exogenous sequence insertion location is determined through BLAST alignment, an MMoE model is constructed, feature parameters are extracted from the BAM file, gated network and expert network are trained, and weights are dynamically adjusted to improve discrimination accuracy.
It significantly improves the accuracy of discrimination of genetically modified events, can efficiently and accurately identify genetically modified events and determine whether they are homozygous or heterozygous, greatly improving the accuracy and operability of genetically modified detection.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transgenic identification, and in particular to a transgenic event identification method based on targeted sequencing and exogenous sequence insertion positions. Background Art
[0002] Transgenic technology has the characteristics of strong targeting, high breeding efficiency and short cycle, and has become one of the important means of plant molecular breeding. Determining whether the transgenic event is homozygous or heterozygous is of great significance for ensuring the genetic stability of transgenic plants, improving breeding efficiency, ensuring the consistency of trait expression, conducting accurate application research, and meeting regulatory and safety requirements.
[0003] During the growth and development of crops, they are often affected by insect pests, weeds and diseases. The current prevention and control measures mainly use a large amount of chemical pesticides and herbicides, which not only seriously affects the ecological environment and biodiversity, increases production costs and labor intensity, but also increases the probability of human poisoning. The most economical and effective method is to explore and use insect-resistant, herbicide-resistant and disease-resistant genes for variety breeding. In the breeding process, homozygous transgenic plants are more genetically stable because they contain exogenous genes on both homologous chromosomes. They can be directly used for production or further breeding work. In the commercialization and regulatory approval process of transgenic products, homozygous plants may also be more likely to be approved. Therefore, it is of great significance to distinguish pure heterozygous transgenic events. Although existing transgenic event detection and discrimination methods such as PCR and ELISA are relatively mature, they still have certain limitations, such as primer selection, quantitative accuracy, and false positive contamination of PCR. Summary of the invention
[0004] The present invention proposes a transgenic event discrimination method based on targeted sequencing and exogenous sequence insertion positions, which improves the accuracy of distinguishing pure heterozygous in existing transgenic event detection and discrimination methods.
[0005] Among them, the method for distinguishing transgenic events based on targeted sequencing and the insertion position of exogenous sequences includes the following steps:
[0006] S1. Obtain the DNA sequence data of the sample to be tested, compare the transformant sequence of the sample to be tested with the reference genome using the BLAST comparison technique, determine the specific insertion position of the exogenous sequence, and identify the specific breakpoint where the transgenic event occurred;
[0007] S2. Align the unfiltered clean reads of the sample to be tested with the reference genome through an alignment algorithm to generate a BAM file;
[0008] S3. Construct the MMoE model and extract characteristic parameters related to the transgenic event from the BAM file as model input, wherein the characteristic parameters include the read depth near the breakpoint, the missing state of the reads, and the alignment consistency of the double-end sequencing; train the gating network, the shared underlying network, and the task-specific expert network of the MMoE model through the characteristic parameters, define the corresponding loss function for the task of the MMoE model, and perform joint optimization through the loss function;
[0009] S4. Automatically adjust the weight of the gating network through the MMoE model and select the most appropriate expert network according to the task type;
[0010] S5. Input new feature parameters into the trained MMoE model, and output the output probability of whether there is a transgenic insertion event in the sample to be tested and the output probability of the homozygous or heterozygous nature of the transgenic event based on the prediction results of each task; and determine the final transgenic event result based on the prediction results of the model and the preset judgment rules.
[0011] Furthermore, in step S3, the specific steps of constructing the MMoE model are as follows:
[0012] S301. Establishing a model structure of a shared underlying network and a task-specific expert network, wherein the shared underlying network extracts common features across tasks, and the task-specific expert network performs task-specific processing on the shared features;
[0013] S302. Define subtasks in the expert network, split the requirements of each task into independent modules, and optimize task-specific goals;
[0014] S303. For each task, define a corresponding loss function according to different task objectives.
[0015] Furthermore, in step S3, the tasks of the MMoE model specifically include:
[0016] Task 1: Use the expert network and the gating network to predict whether the sample contains a transgenic insertion event and output the probability of the transgenic insertion event;
[0017] Task 2: Predict the homozygous / heterozygous status of the sample through the expert network and the gating network, and output the probability of homozygous positive.
[0018] Furthermore, in the task 2, the predicted homozygous / heterozygous status of the sample specifically includes homozygous positive, homozygous negative and heterozygous positive.
[0019] Furthermore, in step S303, the loss function is specifically:
[0020] The loss function of task one: adopt the cross entropy loss function, that is:
[0021] L insert =-[y insert ·log(p insert )+(1-y insert )·log(1-p insert )];
[0022] The loss function of task 2: cross entropy loss function is used, that is:
[0023] L homo =-[y homo ·log(p homo )+(1-y homo )·log(1-p homo )];
[0024] The joint loss function is:
[0025] L total =L insert +L homo ;
[0026] Among them, the L total represents the joint loss function, the L insert represents the loss function used to predict whether a sample contains a transgenic insertion event, the L homo represents the loss function for predicting the homozygous / heterozygous status of a sample, the y insert Indicates the true label of whether the sample has a transgenic insertion event, the y homo The true label representing the homozygous / heterozygous nature of the transgenic event, the p insert Indicates the probability value of whether a sample has a transgenic insertion event, the p homo Indicates the probability value of whether the transgenic event of the sample is homozygous positive.
[0027] Furthermore, in step S5, the prediction result of each task is specifically:
[0028] The predicted output of task one:
[0029] p insert =σ(W insert ·h final +b insert );
[0030] The predicted output of task 2 is:
[0031] p homo =σ(W homo ·h final +b homo );
[0032] Among them, the p insertIndicates the probability value of whether a sample has a transgenic insertion event, the p homo represents the probability value of whether the transgenic event of the sample is homozygous positive, σ represents the sigmoid activation function, and h final represents the final feature obtained by the shared underlying network, that is, the final output of the shared underlying network. homo represents the bias term of the transgenic nature, the b insert represents the bias term for transgenic insertion events, the W insert represents the weight matrix of the transgenic insertion event, the W homo A weight matrix representing the properties of the transgenic gene.
[0033] Furthermore, the step S4 specifically includes the following sub-steps:
[0034] S401. The gated network dynamically calculates the weight of the expert network suitable for the task based on the input parameter features of each task;
[0035] S402. The gating network dynamically selects the most appropriate expert network according to the generated weights.
[0036] Furthermore, in step S401, specifically:
[0037] g t (x i )=[g t1 (x i ), g t2 (x i ), ..., g tK (x i )]where tk ∈[0, 1];
[0038] Among them, the g t (x i ) indicates the characteristic parameter of the i-th input for
[0039] The output vector of task t, g tK represents the output vector of the Kth task, where K represents the total number of tasks, t represents the index of the task, and x i Represents the characteristic parameter of the i-th input.
[0040] Furthermore, in step S3, the alignment consistency of the double-end sequencing is specifically the matching of the two sequences near the same reference position during the alignment process. The alignment consistency of the double-end sequencing includes alignment position consistency, alignment direction consistency and alignment offset consistency.
[0041] Furthermore, in step S5, the preset judgment rule is:
[0042] When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the probability of a homozygous positive is greater than the set homozygous positive threshold, it is determined to be a homozygous positive transgenic insertion;
[0043] When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the probability of a homozygous positive is within the set homozygous positive threshold, it is determined to be a homozygous negative transgenic insertion;
[0044] When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the homozygous positive probability x is less than the set threshold homozygous positive value, it is determined to be a heterozygous positive transgenic insertion;
[0045] When the probability of a transgenic insertion event is less than the set transgenic insertion event threshold, it is determined that the transgenic event does not exist.
[0046] The beneficial effects of the invention are:
[0047] The present invention effectively integrates cross-task features and improves prediction accuracy through the dynamic expert network and shared underlying network structure of the MMoE model. The model adaptively adjusts the weights of the gating network to achieve collaborative optimization between tasks and ensure that the specific requirements of different tasks are accurately processed. The method proposed in the present invention can not only efficiently and accurately identify transgenic events, but also determine whether they are homozygous or heterozygous, greatly improving the accuracy and operability of transgenic detection. DETAILED DESCRIPTION
[0048] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention. It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0049] Moreover, the terms "comprises," "comprising," or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or machine that includes a list of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article, or machine. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or machine that includes the element.
[0050] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0051] Among them, the method for distinguishing transgenic events based on targeted sequencing and the insertion position of exogenous sequences includes the following steps: S1. Obtain the DNA sequence data of the sample to be tested, compare the transformant sequence of the sample to be tested with the reference genome using the BLAST comparison technique, determine the specific insertion position of the exogenous sequence, and identify the specific breakpoint where the transgenic event occurred; S2. Align the unfiltered clean reads of the sample to be tested with the reference genome through the alignment algorithm to generate a BAM file; S3. Construct the MMoE model and extract characteristic parameters related to transgenic events from the BAM file as model input, including read depth near the breakpoint, read missing status and alignment consistency of double-end sequencing; train the gating network, shared underlying network and task-specific expert network of the MMoE model through the characteristic parameters, define the corresponding loss function for the task of the MMoE model, and perform joint optimization through the loss function; S4. Automatically adjust the weight of the gating network through the MMoE model and select the most appropriate expert network according to the task type; S5. Input new feature parameters into the trained MMoE model, and output the output probability of whether there is a transgenic insertion event in the sample to be tested and the output probability of the homozygous or heterozygous nature of the transgenic event based on the prediction results of each task; and determine the final transgenic event result based on the prediction results of the model and the preset judgment rules.
[0055] Specifically, BLAST (Basic Local Alignment Search Tool) is a tool for rapid alignment of genomic sequences using a local alignment algorithm. Its principle is based on finding local matches between the query sequence (the transformant sequence of the sample to be tested) and the region of maximum similarity in the reference genome. The BLAST algorithm first divides the query sequence into short segments (called words), and then performs a preliminary alignment by searching for matching regions of these segments in the reference genome. Next, these matching regions are expanded, and the alignment score is calculated to determine whether there is greater local similarity. The alignment results are used to accurately locate the insertion position of the exogenous sequence in the reference genome and the breakpoint of the transgenic event. The alignment algorithm (such as BWA or Bowtie2) generates a BAM (Binary Alignment / Map) file by aligning the DNA sequence (clean reads) of the sample to be tested with the reference genome. The BAM file records the alignment information of each read with the reference genome, including alignment position, alignment quality, alignment direction, insertion / deletion, and other information. BAM files are compressed formats that take up less space than traditional text formats (such as SAM files) and can efficiently store large amounts of sequencing data. Generating BAM files is the basis for extracting features and calculating depth in subsequent analysis.
[0056] Furthermore, MMoE (Multi-gate Mixture of Experts) is a multi-task learning framework that allows the sharing of some features of the underlying network and assigns an expert network to each task. The core idea of MMoE is to extract global features (cross-task commonality) through a shared underlying network, while using task-specific expert networks to optimize for different tasks (such as predicting transgenic insertion events and homozygosity / heterozygosity). In this method, by extracting features related to transgenic events from BAM files (such as reads depth near breakpoints, reads missing, double-end sequencing consistency, etc.) as input, the shared underlying network of MMoE first processes these features, and then different expert networks perform task-specific processing. Each expert network dynamically adjusts its weights through a gating network to ensure accurate prediction of the task.
[0057] Furthermore, in step S3, the specific steps of constructing the MMoE model are as follows:
[0058] S301. Establishing a model structure of a shared underlying network and a task-specific expert network, wherein the shared underlying network extracts common features across tasks, and the task-specific expert network performs task-specific processing on the shared features;
[0059] Specifically, the shared underlying network extracts common features for all tasks, which are common information across tasks, such as the basic location of transgenic insertion events, breakpoint information, etc. Task-specific expert networks focus on processing specific information required by different tasks. For example, in task one, the focus of the expert network processing may be the deep information around the breakpoints, while in task two, the expert network focuses on the identification features of homozygous / heterozygous. The structures of the shared underlying network and the task-specific expert network work together to ensure multi-task learning while maximizing the independent optimization of each task.
[0060] S302. Define subtasks in the expert network, split the requirements of each task into independent modules, and optimize task-specific goals;
[0061] Specifically, in the expert network, the definition of subtasks is to decompose complex tasks into more specific sub-problems according to the different needs of the tasks. For example, when judging whether a transgenic event exists, the task can be decomposed into the judgment of "whether it is a transgenic insertion", and in the homozygous / heterozygous judgment, it can be further divided into the judgment of "homozygous positive" and "heterozygous positive".
[0062] S303. For each task, define a corresponding loss function according to different task objectives.
[0063] Furthermore, in step S3, the tasks of the MMoE model specifically include:
[0064] Task 1: Use the expert network and the gating network to predict whether the sample contains a transgenic insertion event and output the probability of the transgenic insertion event;
[0065] Task 2: Predict the homozygous / heterozygous status of the sample through the expert network and the gating network, and output the probability of homozygous positive.
[0066] Furthermore, in the task 2, the predicted homozygous / heterozygous status of the sample specifically includes homozygous positive, homozygous negative and heterozygous positive.
[0067] Furthermore, in step S303, the loss function is specifically:
[0068] The loss function of task one: adopt the cross entropy loss function, that is:
[0069] L insert =-[y insert ·log(p insert )+(1-y insert )·log(1-p insert )];
[0070] The loss function of task 2: cross entropy loss function is used, that is:
[0071] L homo =-[y homo ·log(p homo )+(1-y homo )·log(1-p homo )];
[0072] The joint loss function is:
[0073] L total =L insert +L homo ;
[0074] Among them, the L total represents the joint loss function, the L insert represents the loss function used to predict whether a sample contains a transgenic insertion event, the L homo represents the loss function for predicting the homozygous / heterozygous status of a sample, the y insert Indicates the true label of whether the sample has a transgenic insertion event, the y homo The true label representing the homozygous / heterozygous nature of the transgenic event, the p insert Indicates the probability value of whether a sample has a transgenic insertion event, the p homo Indicates the probability value of whether the transgenic event of the sample is homozygous positive.
[0075] Specifically, the loss function is a function used in machine learning to evaluate the difference between the model prediction and the actual label. In this method, the cross entropy loss function is defined for Task 1 and Task 2, respectively. The cross entropy loss function is usually used for binary classification tasks to calculate the difference between the probability value output by the model and the true label. In the determination of transgenic insertion events, the cross entropy loss function evaluates the error between the probability of the model predicting the transgenic insertion event and the true label. In the homozygous / heterozygous determination task, the cross entropy loss function evaluates the error of the model in predicting the homozygosity of the sample. The joint loss function adds the losses of the two tasks to ensure that the two tasks reach the optimal solution in joint training.
[0076] Furthermore, in step S5, the prediction result of each task is specifically:
[0077] The predicted output of task one:
[0078] p insert =σ(W insert ·h final +b insert );
[0079] The predicted output of task 2 is:
[0080] p homo =σ(W homo·h final +b homo );
[0081] Among them, the p insert Indicates the probability value of whether a sample has a transgenic insertion event, the p homo represents the probability value of whether the transgenic event of the sample is homozygous positive, σ represents the sigmoid activation function, and h final represents the final feature obtained by the shared underlying network, that is, the final output of the shared underlying network. homo represents the bias term of the transgenic nature, the b insert represents the bias term for transgenic insertion events, the W insert represents the weight matrix of the transgenic insertion event, the W homo A weight matrix representing the properties of the transgenic gene.
[0082] Furthermore, the step S4 specifically includes the following sub-steps:
[0083] S401. The gated network dynamically calculates the weight of the expert network suitable for the task based on the input parameter features of each task;
[0084] S402. The gating network dynamically selects the most appropriate expert network according to the generated weights.
[0085] Specifically, the function of the gating network is to dynamically select and adjust the expert network to be used according to the input of the model (such as the alignment features of the DNA sequence, the depth of reads near the breakpoint, the missing situation, etc.). Each expert network handles a specific task, and the gating network determines which expert network should be used for the task or how the task should weight different expert networks based on the input feature information. The purpose of the gating network is to convert the input features into a gating weight through a set of calculation processes (such as linear transformation, activation function, etc.). These gating weights determine which expert networks are most effective in processing specific tasks. Exemplarily, in the MMoE model, there are multiple expert networks (each expert network handles different types of tasks), but not all tasks require each expert network. Different DNA data are input (such as different genomic data, alignment situations, missing information, etc.), and the characteristics of each set of data are different. Therefore, the gating network calculates different gating weights based on these different features. If a set of features indicates that the task should use expert network A, then the gating network will select expert network A; if the features indicate that the task should use expert network B, then the gating network will select B. The gating network is not static, but dynamically calculates the weight of the expert network suitable for the task based on the input features of each task. Different tasks and data will select different expert networks for processing. For example, task 1 (such as transgenic insertion event determination) relies more on expert network A, while task 2 (such as homozygous / heterozygous determination) relies more on expert network B. If the input features of task 1 indicate that expert network A is needed for processing, the gating network assigns a larger weight to expert network A. If the input features of task 2 indicate that expert network B is needed for processing, the gating network assigns a larger weight to expert network B.
[0086] Exemplary:
[0087] Suppose there are two tasks in the MMoE model:
[0088] Task 1: Determine whether the sample has a transgenic insertion event (the presence or absence of a gene insertion event).
[0089] Task 2: Determine whether the transgenic event of the sample is homozygous or heterozygous.
[0090] And there are two expert networks:
[0091] Expert Network A: is specifically designed to address the problem of determining whether a transgenic insertion event has occurred.
[0092] Expert network B: specifically used to judge homozygous or heterozygous problems.
[0093] Assume that the input feature set is X, which includes the alignment information of the DNA sequence, the depth of reads, and the missing status. Based on these input features, the gating network calculates the gating weight of task 1 and the gating weight of task 2. If the weight of expert network A is higher, the gating network will select expert network A to process the relevant information of task 1. For task 2, if the gating network obtains a higher gating weight value for task 2, task 2 will be processed by expert network B.
[0094] Furthermore, in step S401, specifically:
[0095] g t (x i )=[g t1 (x i ), g t2 (x i ), ..., g tK (x i )]where tk ∈[0, 1];
[0096] Among them, the g t (x i ) indicates the characteristic parameter of the i-th input for
[0097] The output vector of task t, g tK represents the output vector of the Kth task, where K represents the total number of tasks, t represents the index of the task, and x i Represents the characteristic parameter of the i-th input.
[0098] Furthermore, in step S3, the alignment consistency of the double-end sequencing is specifically the matching of the two sequences near the same reference position during the alignment process. The alignment consistency of the double-end sequencing includes alignment position consistency, alignment direction consistency and alignment offset consistency.
[0099] Furthermore, in step S5, the preset judgment rule is:
[0100] When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the probability of a homozygous positive is greater than the set homozygous positive threshold, it is determined to be a homozygous positive transgenic insertion;
[0101] When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the probability of a homozygous positive is within the set homozygous positive threshold, it is determined to be a homozygous negative transgenic insertion;
[0102] When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the homozygous positive probability x is less than the set threshold homozygous positive value, it is determined to be a heterozygous positive transgenic insertion;
[0103] When the probability of a transgenic insertion event is less than the set transgenic insertion event threshold, it is determined that the transgenic event does not exist.
[0104] Specifically, during the training process, the input feature parameters include read depth around the breakpoint, missing conditions, double-end sequencing consistency, etc. The MMoE model weights and sums these features through the trained weight parameters and predicts them through the expert network of each task. Each task (such as whether it contains a transgenic insertion event, homozygous positive or heterozygous positive, etc.) outputs the corresponding probability value. The sigmoid activation function is used to compress the output probability of each task and convert it into a probability value between 0 and 1, indicating the possibility of the event.
[0105] According to the probability value output by the model, the type of transgenic insertion event (such as homozygous positive, heterozygous positive, etc.) is finally determined through the preset threshold judgment rule. For example, if the probability of the transgenic insertion event exceeds the threshold and the probability of homozygous positive exceeds the set threshold, it is judged as homozygous positive transgenic insertion. This rule accurately identifies the specific nature of the transgenic event by combining the output results of multiple tasks.
[0106] Further, as a preferred implementation of the above embodiment, a scoring-based MMoE model evaluation method is proposed, wherein the tasks of the expert network in the MMoE model include:
[0107] Task 3: Determine whether reads are missing at the breakpoint;
[0108] Task 4: Determine whether the missing reads are aligned with the transformant sequence;
[0109] Task 5: Determine the two sequences read1 and read2 at the same position;
[0110] According to tasks three, four, and five, the MMoE model outputs the corresponding prediction results. Combined with the prediction results, the positive score and the ratio of the sum of positive and negative scores are used to judge the transgenic events.
[0111] Furthermore, the judgment rule is:
[0112] A1. When reads are missing at the breakpoint, and both read1 and read2 can be aligned with the transformant sequence, the score supporting positive is increased by one;
[0113] A2. When reads are missing at the breakpoint, and only one of read1 and read2 can be aligned with the transformant sequence, and the other one is not aligned to the position near the breakpoint in the transformant sequence, a position shift occurs, then the support positive score is increased by 1;
[0114] A3. When reads are missing at the breakpoint, and read1 and read2 cannot be aligned with the transformant sequence, the score supporting positive remains unchanged;
[0115] B1. When reads are not missing at the breakpoint, and both read1 and read2 cross the breakpoint on the reference genome, the score supporting negative is increased by one;
[0116] B2. When reads are not missing at the breakpoint, and only one of read1 and read2 crosses the breakpoint, the score supporting negative is increased by one;
[0117] B3. When reads are not missing at the breakpoint, and read1 and read2 do not cross the breakpoint on the reference genome, the negative score remains unchanged.
[0118] Furthermore, the judgment process is: the ratio is calculated by positive score / (positive score+negative score), and the closer the calculated ratio is to 1, the greater the possibility of homozygous positive; the closer the ratio is to 0, the greater the possibility of homozygous negative.
[0119] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.
Claims
1. A method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions, characterized in that: The following steps are involved: S1. Obtain the DNA sequence data of the sample to be tested, compare the transformant sequence of the sample to be tested with the reference genome using the BLAST comparison technique, determine the specific insertion position of the exogenous sequence, and identify the specific breakpoint where the transgenic event occurred; S2. Align the unfiltered clean reads of the sample to be tested with the reference genome through an alignment algorithm to generate a BAM file; S3. Construct the MMoE model and extract characteristic parameters related to the transgenic event from the BAM file as model input, wherein the characteristic parameters include the read depth near the breakpoint, the missing state of the reads, and the alignment consistency of the double-end sequencing; train the gating network, the shared underlying network, and the task-specific expert network of the MMoE model through the characteristic parameters, define the corresponding loss function for the task of the MMoE model, and perform joint optimization through the loss function; S4. Automatically adjust the weight of the gating network through the MMoE model and select the most appropriate expert network according to the task type; S5. Input new feature parameters to the trained MMoE model, and output the output probability of whether the sample to be tested has a transgenic insertion event and the output probability of the homozygous or heterozygous nature of the transgenic event according to the prediction results of each task; The final outcome of the transgenic event is determined based on the model's prediction results and preset judgment rules.
2. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions according to claim 1, characterized in that: In step S3, the specific steps of constructing the MMoE model are as follows: S301. Establishing a model structure of a shared underlying network and a task-specific expert network, wherein the shared underlying network extracts common features across tasks, and the task-specific expert network performs task-specific processing on the shared features; S302. Define subtasks in the expert network, split the requirements of each task into independent modules, and optimize task-specific goals; S303. For each task, define a corresponding loss function according to different task objectives.
3. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions as described in claim 1, characterized in that: In step S3, the tasks of the MMoE model specifically include: Task 1: Use the expert network and the gating network to predict whether the sample contains a transgenic insertion event and output the probability of the transgenic insertion event; Task 2: Predict the homozygous / heterozygous status of the sample through the expert network and the gating network, and output the probability of homozygous positive.
4. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions according to claim 3, characterized in that: In the second task, the predicted homozygous / heterozygous status of the sample specifically includes homozygous positive, homozygous negative and heterozygous positive.
5. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions according to claim 3, characterized in that: In step S303, the loss function is specifically: The loss function of task one: adopt the cross entropy loss function, that is: L insert =-[y insert ·log(p insert )+(1-y insert )·log(1-p insert )]; The loss function of task 2: cross entropy loss function is used, that is: L homo =-[y homo ·log(p homo )+(1-y homo )·log(1-p homo )]; The joint loss function is: L total =L insert +L homo ; Among them, the L total represents the joint loss function, the L insert represents the loss function used to predict whether a sample contains a transgenic insertion event, the L homo represents the loss function for predicting the homozygous / heterozygous status of a sample, the y insert Indicates the true label of whether the sample has a transgenic insertion event, the y homo The true label representing the homozygous / heterozygous nature of the transgenic event, the p insert Indicates the probability value of whether a sample has a transgenic insertion event, the p homo Indicates the probability value of whether the transgenic event of the sample is homozygous positive.
6. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions according to claim 3, characterized in that: In step S5, the prediction result of each task is specifically: The predicted output of task one: p insert =σ(W insert ·h final +b insert ); The predicted output of task 2 is: p homo =σ(W homo ·h final +b homo ); Among them, the p insert Indicates the probability value of whether a sample has a transgenic insertion event, the p homo represents the probability value of whether the transgenic event of the sample is homozygous positive, σ represents the sigmoid activation function, and h final represents the final feature obtained by the shared underlying network, that is, the final output of the shared underlying network. homo represents the bias term of the transgenic nature, the b insert represents the bias term for transgenic insertion events, the W insert represents the weight matrix of the transgenic insertion event, the W homo A weight matrix representing the properties of the transgenic gene.
7. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions according to claim 1, characterized in that: The step S4 specifically includes the following sub-steps: S401. The gated network dynamically calculates the weight of the expert network suitable for the task based on the input parameter features of each task; S402. The gating network dynamically selects the most appropriate expert network according to the generated weights.
8. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions according to claim 7, characterized in that: In the step S401, specifically: g t (x i )=[g t1 (x i ),g t2 (x i ),...,g tK (x i )]whereg tk ∈[0,1]; Among them, the g t (x i ) indicates the characteristic parameter of the i-th input for The output vector of task t, g tK represents the output vector of the Kth task, where K represents the total number of tasks, t represents the index of the task, and x i Represents the characteristic parameter of the i-th input.
9. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion positions according to claim 1, characterized in that: In step S3, the alignment consistency of double-end sequencing refers specifically to the matching of two sequences near the same reference position during the alignment process. The alignment consistency of double-end sequencing includes alignment position consistency, alignment direction consistency and alignment offset consistency.
10. The method for identifying transgenic events based on targeted sequencing and exogenous sequence insertion position according to claim 1, characterized in that: In step S5, the preset judgment rule is: When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the probability of a homozygous positive is greater than the set homozygous positive threshold, it is determined to be a homozygous positive transgenic insertion; When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the probability of a homozygous positive is within the set homozygous positive threshold, it is determined to be a homozygous negative transgenic insertion; When the probability of a transgenic insertion event is greater than the set transgenic insertion event threshold and the homozygous positive probability x is less than the set threshold homozygous positive value, it is determined to be a heterozygous positive transgenic insertion; When the probability of a transgenic insertion event is less than the set transgenic insertion event threshold, it is determined that the transgenic event does not exist.