Artificial Intelligence-Based DNA Motif Prediction Method, Device, Equipment and Medium
Through an artificial intelligence-based method, the model prediction model and strategy gradient algorithm are used to solve the contradiction between efficiency and accuracy in DNA model prediction, and efficient and accurate DNA model prediction is achieved.
Patent Information
- Application Number
- CN202210814889.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-07-12
AI Technical Summary
The prior art has high efficiency but low accuracy in DNA model prediction, which is prone to local optimum, resulting in reduced efficiency of repeated searches.
Using an artificial intelligence-based method, by obtaining the base statistical results of the to-process DNA sequence, using the model prediction model to calculate the base probability distribution, combining the strategy gradient algorithm to update the model parameters, iteratively train until converges, and obtaining the trained model prediction model.
The accuracy of DNA model prediction is improved, repeated processing is avoided, prediction efficiency is ensured, and model parameters are adaptively adjusted through online learning.
Smart Images

Figure CN115148292B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based DNA motif prediction method, device, equipment and medium. Background Art
[0002] Searching for potential DNA motifs from DNA sequences with the same biological function helps to determine the potential mechanism regulating such biological function and facilitates the study of transcriptional regulation. Since DNA sequences are long, manual processing is time-consuming and labor-intensive. Currently, computer equipment is usually used to perform heuristic searches on DNA sequences, thereby improving the efficiency of DNA motif prediction.
[0003] However, due to the complex arrangement of bases in DNA sequences, heuristic search is prone to fall into the local optimum, which makes it impossible to search for the optimal result, and the accuracy of DNA motif prediction is low. If you want to avoid the heuristic search falling into the local optimum, you need to perform multiple repeated searches, which greatly reduces the efficiency of DNA motif prediction. Therefore, how to improve the accuracy of DNA motif prediction while ensuring the high efficiency of DNA motif prediction has become an urgent problem to be solved. Summary of the invention
[0004] In view of this, the embodiments of the present invention provide a DNA motif prediction method, device, equipment and medium based on artificial intelligence to solve the problem of low accuracy of DNA motif prediction while ensuring high efficiency of DNA motif prediction.
[0005] In a first aspect, an embodiment of the present invention provides a DNA motif prediction method based on artificial intelligence, the DNA motif prediction method comprising:
[0006] Obtaining a DNA sequence to be processed, and counting the number of each type of base in the DNA sequence to be processed to obtain a statistical result;
[0007] According to the statistical results, a motif prediction model is used to obtain a base probability distribution of each preset site in the motif, and an initial motif is determined according to the base probability distribution;
[0008] The initial motif is calculated using a motif evaluation function, and according to the calculation result, the gradient of the motif evaluation function is calculated in combination with a policy gradient algorithm, and the parameters of the motif prediction model are updated according to the gradient to obtain an updated motif prediction model;
[0009] According to the statistical results, the updated motif prediction model is used to obtain the updated base probability distribution of each preset site in the motif, and the updated motif is determined according to the updated base probability distribution, and the updated motif is used as the initial motif, and the step of calculating the initial motif using the motif evaluation function is performed again until the motif evaluation function converges to obtain a trained motif prediction model;
[0010] According to the statistical results, the trained motif prediction model is used to obtain the target base probability distribution of each preset site in the motif, the target motif is determined according to the target base probability distribution, and the target motif is determined to be a DNA motif prediction result.
[0011] In a second aspect, an embodiment of the present invention provides a DNA motif prediction device based on artificial intelligence, the DNA motif prediction device comprising:
[0012] A sequence statistics module is used to obtain a DNA sequence to be processed, count the number of each type of base in the DNA sequence to be processed, and obtain a statistical result;
[0013] A motif prediction module, used to obtain the base probability distribution of each preset site in the motif using a motif prediction model according to the statistical results, and determine an initial motif according to the base probability distribution;
[0014] A model updating module is used to calculate the initial motif using a motif evaluation function, calculate the gradient of the motif evaluation function according to the calculation result in combination with a policy gradient algorithm, and update the parameters of the motif prediction model according to the gradient to obtain an updated motif prediction model;
[0015] An iterative training module is used to predict the updated base probability distribution of each preset site in the motif using the updated motif prediction model according to the statistical results, and determine the updated motif according to the updated base probability distribution, take the updated motif as the initial motif, and again perform the step of calculating the initial motif using the motif evaluation function until the motif evaluation function converges to obtain a trained motif prediction model;
[0016] A motif determination module is used to obtain the target base probability distribution of each preset site in the motif based on the statistical results using the trained motif prediction model, determine the target motif based on the target base probability distribution, and determine that the target motif is a DNA motif prediction result.
[0017] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the DNA motif prediction method as described in the first aspect when executing the computer program.
[0018] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the DNA motif prediction method as described in the first aspect is implemented.
[0019] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0020] The number of each type of base in the obtained DNA sequence to be processed is counted, and according to the statistical results, the motif prediction model is used to obtain the base probability distribution of each preset site in the motif, and the initial motif is determined according to the base probability distribution, and the motif evaluation function is used in combination with the policy gradient algorithm to calculate the gradient of the initial motif, and the parameters of the motif prediction model are updated according to the gradient to obtain an updated motif prediction model, and then the updated motif prediction model is used to obtain the updated motif, and the updated motif is used as the initial motif, and the step of calculating the initial motif using the motif evaluation function is performed again until the motif evaluation function converges to obtain a trained motif prediction model, and the trained motif prediction model is used to determine the target motif, and the motif prediction model is used for online learning, and the motif prediction model parameters can be adaptively adjusted for different DNA sequences to be processed, thereby improving the accuracy of the motif prediction model for DNA motif prediction, and the motif evaluation function is used to evaluate the initial motif and provide a gradient for updating the motif prediction model parameters, so that the DNA motif prediction process and the motif prediction model training process can be synchronized to avoid repeated processing, thereby ensuring the efficiency of DNA motif prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0022] Figure 1 This is a schematic diagram of an application environment of a DNA motif prediction method based on artificial intelligence provided in Embodiment 1 of the present invention;
[0023] Figure 2 It is a schematic flow chart of a DNA motif prediction method based on artificial intelligence provided in Example 1 of the present invention;
[0024] Figure 3 It is a flow chart of a DNA motif prediction method based on artificial intelligence provided in the second embodiment of the present invention;
[0025] Figure 4 It is a structural schematic diagram of a DNA motif prediction device based on artificial intelligence provided in Embodiment 3 of the present invention;
[0026] Figure 5 It is a structural diagram of a computer device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION
[0027] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.
[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0029] It should also be understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0030] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]", depending on the context.
[0031] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" etc. described in the present specification mean that one or more embodiments of the present invention include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0033] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0034] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0035] It should be understood that the order of execution of the steps in the following embodiments does not imply a precedence of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0036] In order to illustrate the technical solution of the present invention, specific embodiments are provided below for illustration.
[0037] The first embodiment of the present invention provides a DNA motif prediction method based on artificial intelligence, which can be applied in the following aspects: Figure 1 In the application environment, the client communicates with the server. The client includes but is not limited to PDA, desktop computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cloud terminal device, personal digital assistant (PDA) and other computer devices. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0038] See also Figure 2, is a schematic diagram of a flow chart of a DNA motif prediction method based on artificial intelligence provided in Example 1 of the present invention. The above DNA motif prediction method can be applied to Figure 1 The client in the example, the computer device corresponding to the client connects to the server, obtains the DNA sequence to be processed from the server, and the client is deployed with a motif prediction model, which can be used to predict the DNA motif corresponding to the DNA sequence to be processed. Figure 2 As shown, the DNA motif prediction method may include the following steps:
[0039] Step S201, obtaining a DNA sequence to be processed, and counting the number of each type of base in the DNA sequence to be processed to obtain a statistical result.
[0040] The DNA sequence to be processed may refer to a DNA sequence with a target biological function. Usually, the length of the DNA sequence to be processed is tens to hundreds of base pairs, and the batch for DNA motif prediction usually contains thousands of DNA fragments.
[0041] The base may refer to the constituent base of a DNA sequence, and the base may include adenine (A), thymine (T), cytosine (C), and guanine (G). The statistical result may refer to the specific base arrangement in the DNA sequence to be processed.
[0042] Specifically, the statistical result may be in matrix form, and the size of the matrix corresponding to the statistical result is N*L*4, where N may refer to the number of DNA sequences to be processed, L may refer to the uniform length of the DNA sequences to be processed, and 4 may refer to the number of DNA bases.
[0043] In order to unify the matrix format, the length of the longest DNA sequence to be processed is taken as L, that is, L = max(l 1 ,l 2 ,…,l n ,…,l N ), where max can refer to the maximum value, l can refer to the actual length of the DNA sequence to be processed, and l n It can refer to the actual length of the nth DNA sequence to be processed, and the value range of n is an integer between [1, N]. It should be noted that for the DNA sequence to be processed whose actual length is less than the uniform length, when its length is expanded to the statistical length, it is expanded from the end, and the element values corresponding to the expansion sites are all 0.
[0044] For example, assume that the order of bases in the matrix corresponding to the statistical results is arranged in ATCG order. For the element with coordinates [n, l, 2] in the matrix corresponding to the statistical results, the element value of the element indicates whether the lth site of the nth DNA sequence to be processed is a T base. When the element value of the element is 0, it indicates that the lth site of the nth DNA sequence to be processed is not a T base. When the element value of the element is 1, it indicates that the lth site of the nth DNA sequence to be processed is a T base.
[0045] The above-mentioned step of obtaining the DNA sequence to be processed, counting the number of each type of base in the DNA sequence to be processed, and obtaining the statistical result, takes the specific arrangement of the DNA sequence to be processed as the statistical result, so as to provide complete information of the DNA sequence to be processed for the subsequent motif prediction model, thereby improving the accuracy of the motif prediction.
[0046] Step S202: According to the statistical results, a motif prediction model is used to obtain the base probability distribution of each preset site in the motif, and an initial motif is determined according to the base probability distribution.
[0047] Among them, the motif prediction model can refer to a neural network model, a logistic regression model, a multi-layer perceptron model, etc. The motif can refer to a shorter fragment on DNA, usually 8 to 20 base pairs in length. The motif is enriched in DNA fragments with the same biological function (such as the promoter region of the gene and around the transcription factor binding site).
[0048] In this embodiment, the length of the motif is determined to be a preset length, and accordingly, the motif includes a fixed number of preset sites. For example, the length of the motif can be preset to 10 base pairs, and the motif includes 10 preset sites.
[0049] The base probability distribution may refer to the probability of four bases constituting DNA in each preset site, and the initial motif may refer to the DNA motif to be evaluated.
[0050] Specifically, the base probability distribution may be in a matrix form, and the size of the matrix corresponding to the base probability distribution is M*4, where M may refer to a preset length of the motif, and 4 may refer to the number of DNA bases.
[0051] It should be noted that since each row in the matrix represents the probability distribution of each type of base in that row, the sum of the element values of each row should be 1. For example, the m-th row of the matrix corresponding to the base probability distribution is [0.1, 0.2, 0.4, 0.3], indicating that the probability of the m-th preset site being an A base is 0.1, the probability of being a T base is 0.2, the probability of being a C base is 0.4, and the probability of being a G base is 0.3.
[0052] Optionally, the motif prediction model includes a convolutional layer and a fully connected layer;
[0053] According to the statistical results, the motif prediction model is used to obtain the base probability distribution of each preset site in the motif, and the initial motif is determined according to the base probability distribution, including:
[0054] Input the statistical results into the convolution layer for feature extraction to obtain statistical features;
[0055] The statistical features are input into the fully connected layer for feature mapping to obtain the base probability distribution of each preset site;
[0056] For the base probability distribution of any preset site, the base corresponding to the maximum probability in the base probability distribution is determined as the initial base of the preset site, and all the initial bases are spliced into an initial motif in the order of the sites.
[0057] The statistical feature may refer to a feature tensor of the statistical result after convolution aggregation, and the initial base may refer to the most likely base category of a preset site probability.
[0058] Specifically, a single-layer convolution layer may include a convolution calculation layer, a batch normalization layer, and an activation layer. The convolution layer may be used to extract input features. In this embodiment, the number of convolution layers is set to 2. The implementer may adjust the number of convolution layers according to actual conditions. It should be noted that the number of convolution layers should not be too deep to avoid degradation. The recommended setting range of the number of layers is [1, 4]. The convolution calculation layer uses a convolution kernel to extract features by sliding. The parameters of the convolution kernel are initialized by a random generation method. The convolution calculation layer plays a role in feature aggregation. The batch normalization layer is used to normalize the features extracted by the convolution calculation layer, thereby speeding up network training and convergence. The activation layer is used to improve the nonlinear characterization capability of the model. In this embodiment, the activation layer may use a linear rectification function (ReLU function).
[0059] The fully connected layer can be used to map features to the output space. In this embodiment, the output space is also the base probability distribution space of the preset site. The fully connected layer is connected with a normalized exponential function (Softmax function), which is used to map the output value of the fully connected layer to a probability value.
[0060] After obtaining the base probability distribution of the preset sites, the maximum function (Max function) is used to determine the base category corresponding to the maximum probability of each preset site. For example, the base probability distribution of preset site 1 is [0.1, 0.2, 0.4, 0.3], then the maximum probability of preset site 1 is 0.4, and the base category corresponding to the maximum probability of preset site 1 is C base.
[0061] Splicing can refer to the concatenation (Concatenate, Concat) method, that is, splicing the part to be spliced to the end of the spliced part. During the initial splicing, the base category corresponding to the first preset site is used as the spliced part according to the preset site order, and the base category corresponding to the second preset site is used as the part to be spliced, and the splicing result is updated to the spliced part. Accordingly, the part to be spliced is determined according to the preset site order until the base categories corresponding to all preset sites are spliced. For example, the base category corresponding to the first preset site is C base, the base category corresponding to the second preset site is A base, and the base category corresponding to the third preset site is G base, then the initial motif obtained after splicing is [A, C, G].
[0062] In one implementation, the activation layer may also use an S-curve function (Sigmoid function), a hyperbolic tangent function (Tanh function), etc.
[0063] This embodiment uses a convolutional neural network to predict DNA motifs, so that the algorithm complexity only increases linearly with the amount of input data, avoiding exponential growth due to repeated searches, and improving the efficiency of DNA motif prediction.
[0064] Based on the statistical results, the motif prediction model is used to obtain the base probability distribution of each preset site in the motif, and the initial motif step is determined according to the base probability distribution. No labeling is required in advance, and the initial motif is obtained by prediction. In the subsequent steps, adjustments are made based on the predicted initial motif feedback, thereby avoiding the problem of difficult DNA motif labeling and effectively improving the efficiency of DNA motif prediction.
[0065] Step S203, using the motif evaluation function to calculate the initial motif, according to the calculation result, in combination with the policy gradient algorithm, the gradient of the motif evaluation function is calculated, and the parameters of the motif prediction model are updated according to the gradient to obtain an updated motif prediction model.
[0066] Among them, the motif evaluation function can be used to evaluate the explanatory power of the motif on the DNA sequence to be processed, and the explanatory power can be characterized by the frequency of occurrence of the motif in the DNA sequence to be processed.
[0067] The policy gradient algorithm may refer to a calculation method for calculating the gradient of the parameters of the motif prediction model, and the gradient may refer to the gradient used to guide the update of the parameters of the motif prediction model, that is, the parameters of the motif prediction model are updated by using the gradient descent method and the back propagation algorithm. The parameters may refer to the weight parameters and bias parameters of the neurons in the motif prediction model. The updated motif prediction model may refer to the motif prediction model after the parameters are updated.
[0068] Specifically, the motif evaluation function provides an evaluation of the motif. The higher the evaluation value, the greater the explanatory power of the motif for the processed DNA sequence, and the more likely the predicted motif is to be a true motif. The lower the evaluation value, the lower the explanatory power of the motif for the processed DNA sequence, and the less likely the predicted motif is to be a true motif.
[0069] The formula of the policy gradient algorithm is:
[0070]
[0071] in, is the gradient, θ is the model parameter of the motif prediction model, π(θ) is the strategy of the motif prediction model, G is the function value of the motif evaluation function, E π(θ) [G] is the expectation of the function value, A is the initial motif, and S is the statistical result.
[0072] π(θ) may refer to the output strategy adopted by the model under given model parameters. The output strategy may include a softmax strategy, a Gaussian strategy, and the like.
[0073] G may refer to the function value of the model evaluation function calculated using the current output model. Since G depends on the decision of the model when the statistical result is fixed, when the expected gradient of the function value is taken, the right side of the equation will act on the model parameters, so that the gradient of the model parameters can be obtained. The policy gradient algorithm is used to calculate the gradient of the model parameters, so that model training can be guided in the absence of labels, thereby improving the feasibility of model prediction model training.
[0074] Optionally, the initial motif is calculated using a motif evaluation function including:
[0075] Align the initial motif with each DNA sequence to be processed, and determine the best matching fragment based on the alignment results;
[0076] The base count of each site in the best matching fragment is calculated to obtain the base count matrix of the corresponding site;
[0077] According to the base count matrix, the motif evaluation function of the corresponding site is calculated to obtain the site evaluation value, and the sum of all site evaluation values is determined as the calculation result.
[0078] The comparison may refer to similarity calculation, and the similarity may be calculated using Euclidean distance, cosine similarity, etc. The best matching fragment may refer to a DNA sub-fragment in the DNA sequence to be processed that is most similar to the initial motif.
[0079] The base count matrix may refer to the statistical result of the number of bases of the best matching fragment. The site evaluation value may refer to the discrete simulation value of the K-order Kullback-Leibler (KL) divergence, where K may refer to an integer greater than or equal to zero. In this embodiment, K is 0, that is, the site is the calculation object. The calculation result may refer to the explanatory power of the best matching fragment on the processed DNA sequence.
[0080] Specifically, the size of the base count matrix is also M*4, where M refers to the preset length of the motif and 4 refers to the number of DNA bases. Since the optimal matching fragment may not be completely consistent with the initial motif, it is necessary to count the number of bases of the optimal matching fragment again to obtain the base count matrix corresponding to the optimal matching fragment. For an initial motif of M positions, the optimal matching fragment is also M positions, and each position needs to be calculated by the motif evaluation function, and then the sum of the position evaluation values of the M positions is used as the calculation result.
[0081] In this embodiment, the DNA sub-fragment most similar to the initial motif is used as the optimal matching fragment, and the base number statistics are performed to obtain a base counting matrix. When there is a small deviation in the initial motif prediction, the initial motif can still be effectively evaluated, thereby greatly reducing the number of iterative updates of the initial motif and improving the efficiency of DNA motif prediction.
[0082] The model evaluation function is:
[0083]
[0084] Among them, α is the bit-rate value, N is the number of DNA sequences to be processed, k is the type of bases contained in the DNA sequence to be processed, and x i is the number of times the i-th base appears at the site, q i is the proportion of the i-th base in all DNA sequences to be processed.
[0085] Specifically, to avoid meaningless function calculations, the logarithmic function operations in the motif evaluation function all add one after the input, for example, logN! is actually calculated as log(N!+1).
[0086] The larger the bit comment value, the greater the difference between the base distribution on the optimal matching fragment and the base distribution of the DNA sequence to be processed. At this time, the optimal matching fragment does not conform to the characteristics of the motif. Conversely, the smaller the bit comment value, the more similar the base distribution on the optimal matching fragment and the base distribution of the DNA sequence to be processed, and the optimal matching fragment conforms to the characteristics of the motif.
[0087] This embodiment uses a motif evaluation function to evaluate the initial motif, determines the evaluation value according to the explanatory power of the motif for the DNA sequence, and provides a training gradient for the motif prediction model, thereby improving the accuracy of DNA motif prediction.
[0088] Optionally, after performing base number statistics on each site in the best matching fragment to obtain a base count matrix of the corresponding site, the following further includes:
[0089] For any site, 2K neighboring sites adjacent to the site form a K-order sub-fragment with the site;
[0090] Counting the number of bases in each K-order sub-fragment in the best matching fragment to obtain a second base count matrix corresponding to the K-order sub-fragment;
[0091] Accordingly, according to the base count matrix, the motif evaluation function of the corresponding site is calculated to obtain the site evaluation value, and the sum of all site evaluation values is determined as the motif evaluation value including:
[0092] According to the base count matrix, the position evaluation value of the corresponding site is calculated;
[0093] According to the second base count matrix, the sub-fragment evaluation value corresponding to the K-order sub-fragment is calculated, and the sum of all position evaluation values and all sub-fragment evaluation values is determined as the motif evaluation value.
[0094] Wherein, K is an integer greater than or equal to zero. In the present embodiment, K is an integer greater than zero, for example, K is 1. The adjacent site may refer to a site that is closer to the site. The K-order sub-fragment may refer to a sub-fragment in the best matching fragment. The second base count matrix may refer to a matrix obtained by performing base count on the K-order sub-fragment. The sub-fragment evaluation value may refer to the calculation result of performing a motif evaluation function calculation on the K-order sub-fragment.
[0095] Specifically, when calculating the high-order motif evaluation function, the type of base not only considers the base itself, but also its adjacent bases. For the first-order motif evaluation function, let the site of the object base be C, where C represents the sequence number of the site. The corresponding first-order sub-segment is the case of considering the object base itself and a neighboring base, that is, the first-order sub-segment is (C-1, C, C+1). At this time, the first-order sub-segment is regarded as a whole, and the number of base categories M′ of the first-order sub-segment is 4^3, that is, 64.
[0096] In this embodiment, the DNA sequence to be processed, the initial motif, and the new base type of the base category number M′ in the form of a first-order sub-fragment are all converted into tensor form. Therefore, for the K-order motif evaluation function, there is no need to consider the value of K. Two-dimensional convolution processing is used for data matching, and the number of occurrences of the matching results is counted by the sum function (Sum function). Therefore, the calculation time will not increase significantly with the number of surrounding bases considered. For example, the two-dimensional tensor of the initial motif is regarded as a convolution kernel and a sliding convolution is performed on the corresponding tensor of the DNA sequence to be processed. It is known that if there is a fragment that completely matches the initial motif, its convolution result is a fixed value. In the sliding convolution, the closer the convolution result is to the fixed value, the more similar the sliding selected fragment is to the initial motif.
[0097] This embodiment uses a high-order motif evaluation function to perform motif evaluation, thereby avoiding the situation where the probability of occurrence of base combinations in the DNA sequence itself has a preference, which affects the motif evaluation. It can better distinguish the motif from the background, accelerate the training process, and improve the accuracy of DNA motif prediction.
[0098] The above-mentioned steps use the motif evaluation function to calculate the initial motif, and according to the calculation result, combined with the policy gradient algorithm, calculate the gradient of the motif evaluation function, and update the parameters of the motif prediction model according to the gradient to obtain an updated motif prediction model. In an unsupervised case, the motif evaluation function and the policy gradient algorithm are used to calculate the gradient to guide the update of the motif prediction model parameters, which can ensure the smooth update of the motif prediction model parameters, avoid the difficulty of convergence of the motif prediction model during the training process, and improve the accuracy of the updated motif prediction model.
[0099] Step S204: Based on the statistical results, the updated motif prediction model is used to obtain the updated base probability distribution of each preset site in the motif, and the updated motif is determined based on the updated base probability distribution. The updated motif is used as the initial motif, and the step of calculating the initial motif using the motif evaluation function is performed again until the motif evaluation function converges to obtain a trained motif prediction model.
[0100] The updated base probability distribution may refer to the output obtained after the statistical results are input into the updated motif prediction model, and the updated motif may refer to the motif determined according to the updated base probability distribution.
[0101] A trained motif prediction model may refer to a motif prediction model that stably outputs the same updated motif. At this time, the parameters of the motif prediction model are stable, which means that the training process has been completed.
[0102] Specifically, the updated base probability distribution may refer to the update probability of the four bases constituting DNA in each preset site. The updated base probability distribution may still be represented in matrix form, that is, the size of the matrix corresponding to the updated base probability distribution is also M*4, and the sum of the element values of each row of elements should also be 1.
[0103] According to the above statistical results, an updated motif prediction model is used to obtain the updated base probability distribution of each preset site in the motif, and the updated motif is determined according to the updated base probability distribution. The updated motif is used as the initial motif, and the step of using the motif evaluation function to calculate the initial motif is executed again until the motif evaluation function converges to obtain the trained motif prediction model step. The motif prediction model is updated in an iterative manner, and the corresponding updated base probability distribution is obtained, so as to gradually obtain the optimal motif prediction model, avoid falling into the local optimal situation, and thus obtain a more accurate DNA motif when a small amount of data is input.
[0104] Step S205 , according to the statistical results, the trained motif prediction model is used to obtain the target base probability distribution of each preset site in the motif, the target motif is determined according to the target base probability distribution, and the target motif is determined to be the DNA motif prediction result.
[0105] The target base probability distribution may refer to an optimal base probability distribution, and the target motif may refer to an optimal DNA motif.
[0106] Specifically, the target base probability distribution may refer to the target probability of the four bases constituting DNA in each preset site. The target base probability distribution may still be represented in matrix form, that is, the size of the matrix corresponding to the target base probability distribution is also M*4, and the sum of the element values of each row of elements should also be 1.
[0107] Based on the statistical results, the trained motif prediction model is used to obtain the target base probability distribution of each preset site in the motif, and the target motif step is determined according to the target base probability distribution, so that the DNA motif prediction process and the training process of the motif prediction model are carried out simultaneously to avoid repeated processing, thereby ensuring the efficiency of DNA motif prediction.
[0108] This embodiment adopts the motif prediction model for online learning, and can adaptively adjust the parameters of the motif prediction model for different DNA sequences to be processed, thereby improving the accuracy of the motif prediction model in DNA motif prediction. The motif evaluation function is used to evaluate the initial motif and provide a gradient for updating the motif prediction model parameters, which can synchronize the DNA motif prediction process with the motif prediction model training process to avoid repeated processing, thereby ensuring the efficiency of DNA motif prediction.
[0109] See also Figure 3, is a flow chart of a DNA motif prediction method based on artificial intelligence provided in Embodiment 2 of the present invention. In the DNA motif prediction method, the DNA motif can adopt a fixed preset length or an adjustable preset length.
[0110] When the DNA motif adopts a fixed preset length, the preset length is set by the implementer and cannot be adjusted after setting. The process of setting the preset length is shown in Example 1 and will not be described in detail here.
[0111] When the DNA motif adopts an adjustable preset length, the preset length is set by the implementer, but can be adjusted according to actual conditions. The DNA motif length adjustment process includes the following steps:
[0112] Step S301, comparing the convergence value of the motif evaluation function with a preset threshold;
[0113] Step S302, if the convergence value is less than a preset threshold, adjusting the number of preset sites;
[0114] Step S303 is to execute again the step of obtaining the base probability distribution of each preset site in the motif using the motif prediction model according to the statistical results, and obtaining the initial motif according to the base probability distribution.
[0115] The convergence value may refer to the function value after the motif evaluation function converges, and the preset threshold may be used to characterize whether the DNA motif has good explanatory power. For example, in this embodiment, the threshold is set to 2.
[0116] The initial number of preset sites can be set by the implementer, and the setting range is recommended to be [8, 20]. Accordingly, the adjustment method can be to increase or decrease one bit. It should be noted that the number of sites after adjustment should also meet the above setting range.
[0117] Optionally, adjusting the number of preset sites includes:
[0118] Obtain the current number of sites and a preset site value range, and determine the probability of reducing the sampling and the probability of increasing the sampling according to the current number of sites and the site value range;
[0119] Sampling is performed according to the probability of reducing sampling and increasing sampling, and the number of sites is adjusted according to the sampling results.
[0120] Among them, the current number of sites may refer to the number of sites before adjustment, the site value range may refer to a preset range interval, such as [8, 20], the probability of reduced sampling may refer to the probability of sampling to reduce one site, and the probability of increased sampling may refer to the probability of sampling to increase one site.
[0121] Specifically, assuming that the number of current sites obtained is p and the preset site value range is [min, max], the probability of reduced sampling can be expressed as The probability of increasing the sampling bit can be expressed as It should be noted that when the number of sites is adjusted to the number of sites that have participated in training, it needs to be adjusted again until the adjusted number of sites has not participated in training.
[0122] This embodiment updates the number of sites by means of probability sampling, thereby avoiding invalid updates of the number of sites and improving the updating efficiency of the number of sites.
[0123] This embodiment dynamically adjusts the number of sites of the DNA motif, thereby avoiding the situation where it is difficult to obtain the optimal DNA motif due to the fixed number of sites, and can effectively improve the accuracy of DNA motif prediction.
[0124] Corresponding to the DNA motif prediction method based on artificial intelligence in the above embodiment, Figure 4 The structural block diagram of the artificial intelligence-based DNA motif prediction device provided in the third embodiment of the present invention is shown. The above-mentioned DNA motif prediction device is applied to the client. The computer device corresponding to the client is connected to the server, and the DNA sequence to be processed is obtained from the server. The client is deployed with a motif prediction model, which can be used to predict the DNA motif corresponding to the DNA sequence to be processed. For the convenience of explanation, only the part related to the embodiment of the present invention is shown.
[0125] See also Figure 4 , the DNA motif prediction device comprises:
[0126] A sequence statistics module 41 is used to obtain a DNA sequence to be processed, count the number of each type of base in the DNA sequence to be processed, and obtain a statistical result;
[0127] A motif prediction module 42 is used to obtain the base probability distribution of each preset site in the motif based on the statistical results using the motif prediction model, and determine the initial motif based on the base probability distribution;
[0128] The model updating module 43 is used to calculate the initial motif using the motif evaluation function, calculate the gradient of the motif evaluation function according to the calculation result and in combination with the policy gradient algorithm, and update the parameters of the motif prediction model according to the gradient to obtain an updated motif prediction model;
[0129] An iterative training module 44 is used to predict the updated base probability distribution of each preset site in the motif using the updated motif prediction model according to the statistical results, and determine the updated motif according to the updated base probability distribution, use the updated motif as the initial motif, and again perform the step of calculating the initial motif using the motif evaluation function until the motif evaluation function converges to obtain a trained motif prediction model;
[0130] The motif determination module 45 is used to obtain the target base probability distribution of each preset site in the motif based on the statistical results using the trained motif prediction model, determine the target motif based on the target base probability distribution, and determine the target motif as the DNA motif prediction result.
[0131] Optionally, the motif prediction model includes a convolutional layer and a fully connected layer;
[0132] The motif prediction module 42 includes:
[0133] A feature extraction unit is used to input the statistical results into the convolution layer for feature extraction to obtain statistical features;
[0134] A feature mapping unit is used to input the statistical features into the fully connected layer for feature mapping to obtain the base probability distribution of each preset site;
[0135] The base splicing unit is used to determine the base corresponding to the maximum probability in the base probability distribution of any preset site as the initial base of the preset site, and splice all the initial bases into an initial motif in the order of the sites.
[0136] Optionally, the DNA motif prediction device further comprises:
[0137] A threshold comparison module, used to compare the convergence value of the motif evaluation function with a preset threshold;
[0138] A site number adjustment module, used for adjusting the number of preset sites if the convergence value is less than a preset threshold;
[0139] The re-prediction module is used to re-execute the steps of obtaining the base probability distribution of each preset site in the motif using the motif prediction model according to the statistical results, and obtaining the initial motif according to the base probability distribution.
[0140] Optionally, the above-mentioned site number adjustment module includes:
[0141] A probability determination unit, used to obtain the current number of sites and a preset site value range, and determine the probability of reducing the sampling and the probability of increasing the sampling according to the current number of sites and the site value range;
[0142] The probability sampling unit is used to perform sampling according to the sampling probability of reducing the bit and the sampling probability of increasing the bit, and adjust the number of sites according to the sampling result.
[0143] Optionally, the DNA motif prediction device further comprises:
[0144] A sequence alignment module is used to align the initial motif with each DNA sequence to be processed and determine the best matching fragment based on the alignment results;
[0145] The site counting module is used to count the number of bases at each site in the best matching fragment to obtain the base counting matrix of the corresponding site;
[0146] The site evaluation module is used to calculate the motif evaluation function of the corresponding site according to the base count matrix, obtain the site evaluation value, and determine the sum of all site evaluation values as the calculation result.
[0147] Optionally, the above motif evaluation function is:
[0148]
[0149] Among them, α is the bit-rate value, N is the number of DNA sequences to be processed, M is the type of bases contained in the DNA sequences to be processed, and x i is the number of times the i-th base appears at the site, q i is the proportion of the i-th base in all DNA sequences to be processed.
[0150] Optionally, the DNA motif prediction device further comprises:
[0151] A sub-fragment composition module is used to form a K-order sub-fragment with 2K neighboring sites adjacent to any site, where K is an integer greater than zero;
[0152] A sub-fragment counting module is used to count the number of bases in each K-order sub-fragment in the best matching fragment to obtain a second base counting matrix corresponding to the K-order sub-fragment;
[0153] Accordingly, the above-mentioned site evaluation module includes:
[0154] A site evaluation unit is used to calculate the site evaluation value of the corresponding site according to the base count matrix;
[0155] The motif evaluation unit is used to calculate the sub-fragment evaluation value corresponding to the K-order sub-fragment according to the second base count matrix, and determine the sum of all position evaluation values and all sub-fragment evaluation values as the motif evaluation value.
[0156] It should be noted that the information interaction, execution process and other contents between the above-mentioned modules and units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0157] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned DNA motif prediction method embodiments are implemented.
[0158] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that Figure 5 This is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than those shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.
[0159] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0160] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory may be the memory of a computer device, and the internal memory provides an environment for the operation of an operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be a hard disk of a computer device, and in other embodiments may also be an external storage device of a computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on a computer device. Further, the memory may also include both an internal storage unit of a computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as program codes of computer programs, etc. The memory may also be used to temporarily store data that has been output or is to be output.
[0161] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiment can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0162] The present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing.
[0163] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0164] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0165] In the embodiments provided by the present invention, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0166] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.
Claims
1. An artificial intelligence-based DNA motif prediction method, characterized in that, the method includes: Obtain the DNA sequence to be processed, count the number of each type of base in the DNA sequence to be processed, and obtain a statistical result; According to the statistical result, use a motif prediction model to obtain the base probability distribution of each preset site in the motif, and determine an initial motif according to the base probability distribution; Use a motif evaluation function to calculate the initial motif, according to the calculation result, combined with the policy gradient algorithm, calculate the gradient of the motif evaluation function, and update the parameters of the motif prediction model according to the gradient to obtain an updated motif prediction model; According to the statistical result, use the updated motif prediction model to obtain the updated base probability distribution of each preset site in the motif, and determine an updated motif according to the updated base probability distribution. Take the updated motif as the initial motif, and execute the step of using the motif evaluation function to calculate the initial motif again until the motif evaluation function converges to obtain a trained motif prediction model; According to the statistical result, use the trained motif prediction model to obtain the target base probability distribution of each preset site in the motif, determine a target motif according to the target base probability distribution, and determine the target motif as the DNA motif prediction result; The calculation of the initial motif by the motif evaluation function includes: comparing the initial motif with each DNA sequence to be processed, and determining the optimal matching fragment according to the comparison result; counting the number of bases at each site in the optimal matching fragment to obtain a base count matrix corresponding to the site; According to the base count matrix, calculate the motif evaluation function of the corresponding site to obtain a site evaluation value, and determine the sum of all site evaluation values as the calculation result; the motif evaluation function is: Wherein, is the site evaluation value, N is the number of DNA sequences to be processed, and M is the type of bases included in the DNA sequences to be processed, is the number of occurrences of the th base at the site, is the proportion of the th base in all DNA sequences to be processed.
2. The DNA motif prediction method according to claim 1, characterized in that, the motif prediction model includes a convolutional layer and a fully connected layer; The step of using the motif prediction model to obtain the base probability distribution of each preset site in the motif according to the statistical result and determining the initial motif according to the base probability distribution includes: Input the statistical result into the convolutional layer for feature extraction to obtain statistical features; Input the statistical features into the fully connected layer for feature mapping to obtain the base probability distribution of each preset site; For the base probability distribution of any preset site, determine the base corresponding to the maximum probability in the base probability distribution as the initial base of the preset site, and splice all the initial bases in the site order to form the initial motif.
3. The DNA motif prediction method according to claim 1, characterized in that, after the motif evaluation function converges, it further includes: Compare the convergence value of the motif evaluation function with a preset threshold; If the convergence value is less than the preset threshold, adjust the number of preset sites, and execute again the step of using the motif prediction model to obtain the base probability distribution of each preset site in the motif according to the statistical result and obtaining the initial motif according to the base probability distribution.
4. The DNA motif prediction method according to claim 3, characterized in that, adjusting the number of the preset sites includes: obtaining the current number of sites and the preset range of site values, and determining the probability of subtracting sites and the probability of adding sites according to the current number of sites and the range of site values; sampling according to the probability of subtracting sites and the probability of adding sites, and adjusting the number of sites according to the sampling result.
5. The DNA motif prediction method according to claim 4, characterized in that, after counting the number of bases at each site in the optimal matching fragment to obtain the base count matrix corresponding to the site, it further includes: for any site, forming a K-order sub-fragment with the site and 2K adjacent neighboring sites of the site, where K is an integer greater than zero; counting the number of bases in each K-order sub-fragment in the optimal matching fragment to obtain a second base count matrix corresponding to the K-order sub-fragment; correspondingly, the calculating the motif evaluation function corresponding to the site according to the base count matrix to obtain the site evaluation value, and determining that the sum of all site evaluation values is the motif evaluation value includes: calculating the site evaluation value corresponding to the site according to the base count matrix; calculating the sub-fragment evaluation value corresponding to the K-order sub-fragment according to the second base count matrix, and determining that the sum of all site evaluation values and all sub-fragment evaluation values is the motif evaluation value.
6. An artificial intelligence-based DNA motif prediction device, characterized in that, the DNA motif prediction device includes: a sequence statistics module, configured to obtain a DNA sequence to be processed, and count the number of each type of base in the DNA sequence to be processed to obtain a statistical result; a motif prediction module, configured to obtain the base probability distribution of each preset site in the motif according to the statistical result by using a motif prediction model, and determine an initial motif according to the base probability distribution; a model update module, configured to calculate the initial motif by using a motif evaluation function, calculate the gradient of the motif evaluation function according to the calculation result in combination with a policy gradient algorithm, and update the parameters of the motif prediction model according to the gradient to obtain an updated motif prediction model; an iterative training module, configured to predict the updated base probability distribution of each preset site in the motif by using the updated motif prediction model according to the statistical result, determine an updated motif according to the updated base probability distribution, use the updated motif as the initial motif, and execute the step of calculating the initial motif by using the motif evaluation function again until the motif evaluation function converges to obtain a trained motif prediction model; A motif determination module, configured to obtain, according to the statistical result, a target base probability distribution of each preset site in the motif by using the trained motif prediction model, determine a target motif according to the target base probability distribution, and determine the target motif as the DNA motif prediction result; the calculation of the initial motif by using the motif evaluation function includes: aligning the initial motif with each DNA sequence to be processed, and determining an optimal matching segment according to the alignment result; performing base number statistics on each site in the optimal matching segment to obtain a base count matrix corresponding to the site; calculating the motif evaluation function corresponding to the site according to the base count matrix to obtain a site evaluation value, and determining the sum of all site evaluation values as the calculation result; the motif evaluation function is: Wherein, is the value of the site evaluation, N is the number of the DNA sequences to be processed, and M is the type of bases included in the DNA sequences to be processed, is the number of occurrences of the th type of base at the site, is the proportion of the th type of base in all DNA sequences to be processed.
7. A computer device, characterized in that the computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the DNA motif prediction method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, the DNA motif prediction method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method, device, and apparatus for detecting genovariation point, and storage medium
CN109411016A
Method for quickly identifying single-molecule nanopore sequencing bases based on deep network
CN112183486A