A screening method and system for synthetic biological probiotics

Through artificial intelligence technology, the K-mer frequency feature matrix of DNA sequences and the neural network model of channel attention modules is used to solve the problem of high time and resource consumption in the screening of synthetic biological probiotics, and achieve a more efficient and accurate screening effect.

CN119252334BActive Publication Date: 2025-06-20SHANDONG SYNTHETIC BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411429435.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-06-20
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

The prior art has a large time and resource consumption in the gene screening of synthetic biological probiotics, making it difficult to fully cover the complexity and diversity of gene combinations, and it is difficult to efficiently and accurately screen out probiotics that are beneficial to the human body.

Method used

Using artificial intelligence technology, by obtaining the DNA sequences of beneficial bacteria and non-benefit bacteria in the training samples, searching for target regions in the DNA sequence, counting the K-mer frequency feature matrix, building a multi-channel frequency feature matrix, and using a neural network model with channel attention module for training and screening.

Benefits of technology

It improves the screening efficiency and accuracy of synthetic biological probiotics, can more accurately identify and screen probiotics that are beneficial to the human body, and reduces the consumption of experimental resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119252334B_ABST
    Figure CN119252334B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of screening synthetic biological probiotics, and specifically relates to a screening method and system for synthetic biological probiotics. It searches for target regions in DNA sequences, counts the K-mer frequencies in the target regions, puts the frequencies of K-mers with the same first k bases into a set in sequence according to the alphabetical order of the remaining bases of the K-mers, splices all the sets into a frequency feature matrix of the target region row by row according to the alphabetical order of the first k bases, forms a multi-channel frequency feature matrix from the frequency feature matrices of all target regions, and inputs the multi-channel frequency feature matrix into a neural network model with a channel attention module for training; a multi-channel frequency feature matrix of the synthetic biological probiotics to be screened is obtained from the DNA sequence of the synthetic biological probiotics to be screened, and the multi-channel frequency feature matrix is input into the trained neural network model to obtain the score of the synthetic biological probiotics. The present invention can improve the efficiency of the primary screening of synthetic biological probiotics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of synthetic biological probiotic screening, and specifically relates to a screening method and system for synthetic biological probiotics. Background Art

[0002] Synthetic biological probiotics are based on traditional probiotics and enhanced through genetic modification to improve their functions, such as enhancing their therapeutic effects on specific diseases, improving their adaptability to specific environments, and endowing them with new functions, such as producing specific bioactive substances and degrading environmental pollutants. Synthetic biological probiotics have broad application prospects in the fields of medicine, food, agriculture, and environmental protection. The genetic screening of synthetic probiotics mainly relies on traditional molecular biology methods, including genome sequencing, gene editing technology, high-throughput screening, and functional verification. Although these methods can meet the research and application needs to a certain extent, there are still some deficiencies. For example, traditional methods require a large amount of time and experimental resources, and it is difficult to comprehensively cover the complexity and diversity of gene combinations.

[0003] The rapid development of artificial intelligence provides a new solution for the screening of synthetic biological probiotics. By using artificial intelligence technologies such as machine learning and deep learning, potential functional gene combinations can be mined from massive gene data, gene editing strategies can be predicted and optimized, and the screening efficiency and accuracy can be significantly improved. However, how to efficiently and accurately screen out probiotics beneficial to the human body remains an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the present invention is to propose a screening method for synthetic biological probiotics to solve the problems mentioned in the above background art.

[0005] A screening method for synthetic biological probiotics includes the following steps:

[0006] Obtain the DNA sequences of beneficial bacteria and non-beneficial bacteria in the training samples, search for regions with the same position and the same DNA fragments in all DNA sequences, and use the remaining regions in the DNA sequences as target regions; wherein, the length of the DNA fragment is not less than a preset value;

[0007] Statistically analyze the K-mer frequencies in the target regions, place the frequencies of K-mers with the same first k bases in a set in alphabetical order according to the remaining bases of the K-mer, splice all the sets row by row in alphabetical order of the first k bases to form a frequency feature matrix of the target regions, form a multi-channel frequency feature matrix from the frequency feature matrices of all target regions, and input the multi-channel frequency feature matrix into a neural network model with a channel attention module for training;

[0008] Obtain the multi-channel frequency feature matrix of the synthetic biological probiotic to be screened from the DNA sequence of the synthetic biological probiotic to be screened, and input the multi-channel frequency feature matrix into the trained neural network model to obtain the score of the synthetic biological probiotic.

[0009] Preferably, the frequencies of K-mers with the same first k bases are sequentially placed into a set according to the alphabetical order of the remaining bases of the K-mer, specifically:

[0010] Obtain the frequencies corresponding to all K-mers with the same first k bases, sort the frequencies according to the order of the remaining base letters, and place the sorted frequencies into the set identified by the first k bases.

[0011] Preferably, the channel attention module specifically includes:

[0012] A channel weight calculation unit, which is used to count the functional elements of the target region, set different weights for different functional elements, calculate the weights of all functional elements in the target region, and normalize the weights of all target regions to obtain the normalized channel weights, so as to obtain the channel weight vector;

[0013] A channel attention calculation unit, which is used to calculate the channel attention vector through the channel attention mechanism;

[0014] A fusion unit, which is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

[0015] Preferably, the channel attention module specifically includes:

[0016] A channel weight unit, which is used to obtain the frequency of each base appearing at the same position in the target region of all beneficial bacteria in the training samples for each target region, and obtain the frequency of each base appearing at the same position in the target region of all non-beneficial bacteria in the training samples, and obtain the contribution degree of the target region to distinguish beneficial bacteria and non-beneficial bacteria based on the frequencies, and normalize the contribution degrees of all target regions to obtain the normalized channel weights, so as to obtain the channel weight vector;

[0017] A channel attention unit, which is used to calculate the channel attention vector through the channel attention mechanism;

[0018] A fusion unit, which is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

[0019] Preferably, the fusion of the channel weight vector and the channel attention vector is specifically:

[0020] The channel weight vector and the channel attention vector are input into a fully connected layer and then passed through an activation function to obtain a fusion result.

[0021] On the other hand, the present invention proposes a screening system for synthetic biological probiotics, including the following modules:

[0022] A deduplication module, configured to obtain the DNA sequences of beneficial bacteria and non-beneficial bacteria in the training samples, search for regions with the same positions and the same DNA fragments in all DNA sequences, and use the remaining regions in the DNA sequences as target regions; wherein, the length of the DNA fragment is not less than a preset value;

[0023] A model training module, configured to count the K-mer frequencies in the target regions, place the frequencies of K-mers with the same first k bases in a set in sequence according to the alphabetical order of the remaining bases of the K-mer, splice all the sets into a frequency feature matrix of the target region in rows according to the alphabetical order of the first k bases, form a multi-channel frequency feature matrix from the frequency feature matrices of all target regions, and input the multi-channel frequency feature matrix into a neural network model with a channel attention module for training;

[0024] A screening module, configured to obtain a multi-channel frequency feature matrix of the synthetic biological probiotics to be screened from the DNA sequences of the synthetic biological probiotics to be screened, and input the multi-channel frequency feature matrix into the trained neural network model to obtain the score of the synthetic biological probiotics.

[0025] Preferably, the step of placing the frequencies of K-mers with the same first k bases in a set in sequence according to the alphabetical order of the remaining bases of the K-mer is specifically:

[0026] Obtain the frequencies corresponding to all K-mers with the same first k bases, sort the frequencies according to the alphabetical order of the remaining bases, and place the sorted frequencies in the set identified by the first k bases.

[0027] Preferably, the channel attention module specifically includes:

[0028] A channel weight calculation unit, configured to count the functional elements of the target regions, set different weights for different functional elements, calculate the weights of all functional elements in the target regions, and normalize the weights of all target regions to obtain a normalized channel weight, thereby obtaining a channel weight vector;

[0029] A channel attention calculation unit, configured to calculate a channel attention vector through a channel attention mechanism;

[0030] A fusion unit, configured to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

[0031] Preferably, the channel attention module specifically includes:

[0032] A channel weight unit, which is used to obtain the frequency of each base at the same position in the target region for all beneficial bacteria in the training samples for each target region, and obtain the frequency of each base at the same position in the target region for all non-beneficial bacteria in the training samples. Based on the frequencies, obtain the contribution degree of the target region for distinguishing beneficial bacteria and non-beneficial bacteria, and normalize the contribution degrees of all target regions to obtain the normalized channel weights, thereby obtaining a channel weight vector;

[0033] A channel attention unit, which is used to calculate a channel attention vector through a channel attention mechanism;

[0034] A fusion unit, which is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

[0035] Preferably, the fusion of the channel weight vector and the channel attention vector is specifically:

[0036] Input the channel weight vector and the channel attention vector into a fully connected layer and then obtain a fusion result through an activation function.

[0037] In addition, the present invention also proposes a computer-readable storage medium, on which a computer program is stored, and the computer program realizes the method described above when executed by a processor.

[0038] Finally, the present invention also proposes a computer device, which at least includes a processor and a readable storage medium, on which a computer program is stored, and the computer program realizes the method described above when executed by the processor.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] In order to screen out synthetic biological probiotics, the present invention searches for target regions of DNA sequences, constructs a K-mer frequency feature matrix for each target region, and multiple target regions constitute a multi-channel frequency feature matrix. The multi-channel frequency feature matrix is used to identify beneficial bacteria, and at the same time, channel attention is applied to the recognition of the multi-channel frequency feature matrix. Moreover, when calculating channel attention, it not only relies on K-mer frequencies, but also calculates according to the characteristics of each target region, and at the same time utilizes the K-mer frequency features and the functions or similarities of the target region base sequences, thereby realizing more accurate identification and screening of synthetic biological probiotics. Description of the Drawings

[0041] Figure 1Flow chart of Embodiment 1;

[0042] Figure 2 Schematic diagram of the frequency feature matrix of a target region;

[0043] Figure 3 Schematic diagram of the multi-channel frequency feature matrix;

[0044] Figure 4 Structural diagram of Embodiment 2. Detailed implementation manners

[0045] In this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] In the first embodiment of the present invention, as Figure 1 shown, the present invention proposes a method for screening synthetic biological probiotics, including the following steps:

[0048] S1. Obtain the DNA sequences of beneficial bacteria and non-beneficial bacteria in the training samples, search for regions with the same positions and the same DNA fragments in all the DNA sequences, and use the remaining regions in the DNA sequences as the target regions; wherein, the length of the DNA fragment is not less than a preset value;

[0049] Bacteria are divided into beneficial bacteria and non-beneficial bacteria. Beneficial bacteria include, for example, Lactobacillus and Bifidobacterium. Non-beneficial bacteria are further divided into harmful bacteria and harmless bacteria. For example, Escherichia coli and Staphylococcus aureus belong to harmful bacteria, and Staphylococcus epidermidis and Providencia bacteria belong to harmless bacteria or neutral bacteria, which are generally harmless to humans. Synthetic biology probiotics are genetically modified probiotics that can be used to treat diseases, detect harmful substances, decompose harmful substances, etc. After designing, constructing, and genetically modifying probiotics through vectors, generally, preliminary screening and culture experiments are required to determine which are synthetic biology probiotics. When there are a large number of synthetic biology probiotics, the workload of screening out beneficial bacteria is relatively large. The present invention uses artificial intelligence to preliminarily screen out beneficial bacteria and then conducts culture experiments, which can reduce a lot of workload.

[0050] Obtain the DNA sequences in the training samples from the database. The training samples include beneficial bacteria and non-beneficial bacteria. The DAN sequences in the database come from some public databases and the DNA sequences of some self-tested bacteria. Preferably, the training samples are different strains of the same species. For example, if the goal is to synthesize new Escherichia coli, the DNA sequences of different strains of Escherichia coli are used as training samples. It is also possible to use species with relatively close genetic relationships to ensure the same length of DNA sequences. Since the bacterial DNA of species with close genetic relationships or different strains is also relatively similar, in order to reduce the amount of data for subsequent analysis, search for regions with the same position and the same DNA fragment in all DNA sequences in the training samples. These regions have the same position in the DNA sequence and the same DNA fragment, which is not helpful for distinguishing beneficial bacteria and non-beneficial bacteria. The remaining regions in the DNA sequence are used as target regions. There will be multiple target regions, and the continuous DNA sequence between every two identical regions is a target region. Among them, if the DNA fragment is one or two bases, the DNA sequence will be too disordered. In one embodiment, the length of the DNA fragment is not less than a preset value.

[0051] S2. Statistically analyze the K-mer frequencies in the target regions. Arrange the frequencies of K-mers with the same first k bases in alphabetical order according to the remaining bases of the K-mer and put them into a set. Concatenate all the sets row by row according to the alphabetical order of the first k bases to form the frequency feature matrix of the target regions. Combine the frequency feature matrices of all target regions to form a multi-channel frequency feature matrix. Input the multi-channel frequency feature matrix into a neural network model with a channel attention module for training;

[0052] For each target region of each DNA sequence in the training samples, the K-mer frequencies are counted, where a K-mer is a substring of a DNA sequence with a length of K. For example, if K = 2, the substrings are: AA, AT, AC, AG, TA, TT, TC, TG, CA, CT, CC, CG, GA, GT, GC, GG, where A is adenine, T is thymine, C is cytosine, and G is guanine. In the present invention, a K-mer represents a specific DNA substring with a length of K, and K-mers represent all DNA substrings with a length of K. For example, AA is a K-mer, and the above sixteen substrings are K-mers. Among them, K is a positive integer greater than 2. The K-mer frequency is the frequency of a substring with a length of K appearing in the DNA sequence. In a more specific embodiment, the K-mer is slid in the target region with a step size of 1. If the target region is the same as the K-mer, the K-mer appears once; otherwise, it does not appear, and the K-mer frequency is calculated accordingly.

[0053] After calculating the frequencies of all K-mers, the frequencies of K-mers with the same first k bases are sequentially placed into a set according to the alphabetical order of the remaining bases of the K-mer. Specifically, obtain the frequencies corresponding to all K-mers with the same first k bases, sort the frequencies according to the order of the remaining bases, and place the sorted frequencies into the set identified by the first k bases, where the remaining bases are the bases from the (K - k)-th position to the K-th position of the K-mer. Assume K = 3 and k = 2. Then the first two bases of the four 3-mers AAA, AAT, AAC, and AAG are the same. If their corresponding frequencies are p1, p2, p3, and p4 respectively, and the alphabetical order of the remaining bases of the 3-mer is A, C, G, T, they are placed into a set in this order as {p1, p3, p4, p2}. Similarly, the corresponding set for ACA, ACT, ACC, and ACG can be obtained. In a preferred embodiment, k is also a positive integer and k < K. Figure 2 The frequency feature matrix of a target region with k = 1 and K = 3 is shown.

[0054] After obtaining the sets corresponding to all K-mers, the sets are concatenated into a frequency feature matrix row by row according to the alphabetical order of the first k bases. Still taking the above example, when k = 2, there are 16 combinations of the first two bases. According to the alphabetical order of the first k bases, AA is in the first row, AC is in the second row, and so on. Finally, the frequency feature matrix of this target region is obtained. When there are multiple target regions, each target region corresponds to a frequency feature matrix, and the multiple frequency feature matrices of multiple target regions constitute a multi-channel frequency feature matrix. Each channel is the frequency feature matrix of a target region, where the multi-channel frequency feature matrix is arranged in the order in which the target regions appear in the DNA sequence, such as Figure 3as shown

[0055] In the multi-channel frequency feature matrix, different channels represent the features of different target regions. The positions and functions of different target regions are also different. The present invention further uses channel attention to focus on different channels, that is, target regions, dynamically adjusts the importance of each channel, and enhances the influence of different target regions on the result. In a more specific embodiment, the neural network model with a channel attention module is the SENet model or the CBAM model, etc. Of course, it is not limited to the above models, and it can also be other classification models with an attention module added.

[0056] The K-mer frequency matrix can reflect some features of the target region, but there is loss of the functional information of the target region. In a specific embodiment

[0057] The channel attention module specifically includes:

[0058] A channel weight calculation unit, which is used to count the functional elements of the target region, set different weights for different functional elements, calculate the weights of all functional elements of the target region, normalize the weights of all target regions to obtain the normalized channel weights, and thus obtain a channel weight vector;

[0059] A channel attention calculation unit, which is used to calculate a channel attention vector through a channel attention mechanism;

[0060] A fusion unit, which is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

[0061] Among them, the functional element is a subsequence or substring with a specific function in the DNA sequence, including but not limited to codons, regulatory elements, enhancers, promoters, etc. Since the functions of different functional elements are different, different weights are set for different functional elements, the weights of all functional elements of the target region are calculated, so as to obtain the weight of the target region, and the weights of all target regions are normalized to obtain the normalized channel weights. Since one target region corresponds to a frequency feature matrix, the channel attention and channel weights of the frequency feature matrix are combined to recalculate the channel attention, that is, the channel weights and the channel attention vector are fused. Suppose there are 3 target regions, then there will be 3 frequency feature matrices, and there will be three normalized channel weights. The attention of each channel is obtained through the 3-channel frequency feature matrix, that is, a channel attention vector containing 3 elements. The 3 channel weights and the channel attention vector containing 3 elements are fused, and the fusion result is used as the output of the attention module. Among them, there are various fusion methods, including but not limited to bitwise multiplication, fusion using MLP, etc.

[0062] The differences between different target regions in beneficial bacteria and non-beneficial bacteria are different. Some target regions have a large gap, and these regions are likely to be the functional regions of beneficial bacteria and non-beneficial bacteria. In one embodiment, the channel attention module specifically includes:

[0063] A channel weight unit, which is used to obtain the frequency of each base appearing at the same position in the target region for all beneficial bacteria in the training samples for each target region, and obtain the frequency of each base appearing at the same position in the target region for all non-beneficial bacteria in the training samples. Based on the frequencies, the contribution degree of the target region to distinguishing beneficial bacteria and non-beneficial bacteria is obtained, and the contribution degrees of all target regions are normalized to obtain the normalized channel weights, thereby obtaining a channel weight vector;

[0064] A channel attention unit, which is used to calculate a channel attention vector through a channel attention mechanism;

[0065] A fusion unit, which is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module;

[0066] For each target region, each position is a base, such as one of A, T, G, C. For the same position in the same target region, the probability of each base appearing in all beneficial bacteria is statistically calculated. For example, in the second target region, the probability of A appearing at the 18th position is 0.8, the probability of T appearing is 0.1, the probability of G appearing is 0.1, and the probability of C appearing is 0; Similarly, the probability of each base appearing in all non-beneficial bacteria is statistically calculated. Still taking the above position as an example, the probability of A appearing is 0.2, the probability of T appearing is 0.5, the probability of G appearing is 0.1, and the probability of C appearing is 0.2.

[0067] Based on the above two frequencies, the contribution degrees of the target region to beneficial bacteria and non-beneficial bacteria are obtained. The contribution degree refers to the contribution situation of distinguishing beneficial bacteria or non-beneficial bacteria. In one embodiment, specifically: calculate the absolute difference of the frequencies corresponding to the same base at the same position, thereby obtaining the total sum of the absolute differences of the frequencies of the four bases at the position, calculate the average value of the total sum of the absolute differences of the frequencies of all positions in the target region, and use the average value as the contribution degree. Still taking the above example, for the 18th position, the absolute differences of the frequencies corresponding to the bases ATGC are 0.6, 0.4, 0, 0.2, and the total sum of the absolute differences of the frequencies is 1.2. The absolute difference of the frequencies is the absolute value of the frequency difference. Assume that there are 20 bases in target region 2, then calculate the average value of the total sum of the absolute differences of the frequencies corresponding to the 20 bases, and use this average value as the contribution degree. The greater the difference between the target region in beneficial bacteria and non-beneficial bacteria, the greater the contribution to distinguishing beneficial bacteria and non-beneficial bacteria.

[0068] In another embodiment, the channel attention module only includes a channel weight unit, and uses the channel weight vector output by the channel weight unit as the weight value of each channel.

[0069] Among them, each target region corresponds to a weight, and a channel weight vector is obtained according to the order of appearance of the target regions in the DNA sequence. For example, if there are 10 target regions, the channel weight vector has 10 elements, and at the same time, the channel attention vector also has 10 elements. The elements in the channel attention vector are obtained by calculating the channel attention mechanism on the multi-channel frequency feature matrix, and the channel attention mechanism adopts the attention mechanisms including but not limited to those in SENet and STN.

[0070] After obtaining the fusion result, multiply the fusion result by the multi-channel frequency feature matrix to achieve attention to different channels.

[0071] S3. Obtain the multi-channel frequency feature matrix of the synthetic biological probiotic to be screened from the DNA sequence of the synthetic biological probiotic to be screened, and input the multi-channel frequency feature matrix into the trained neural network model to obtain the score of the synthetic biological probiotic.

[0072] After the neural network model is trained, for the DNA sequence of the synthetic biological probiotic to be screened, obtain the multi-channel frequency matrix in the same way as in S1 and S2. Since the neural network model includes a channel attention module, it will also calculate the channel attention of the synthetic biological probiotic to be identified. The process is the same as above and will not be elaborated here. Input the multi-channel frequency matrix into the trained neural network model to obtain the score of the synthetic biological probiotic. The higher the score, the greater the possibility that it is a probiotic.

[0073] In the second embodiment of the present invention, as Figure 4 shown, the present invention proposes a screening system for synthetic biological probiotics, including the following modules:

[0074] A deduplication module, which is used to obtain the DNA sequences of beneficial bacteria and non-beneficial bacteria in the training samples, search for regions with the same position and the same DNA fragments in all DNA sequences, and use the remaining regions in the DNA sequences as target regions; wherein, the length of the DNA fragment is not less than a preset value;

[0075] A model training module, which is used to count the K-mer frequencies in the target regions, put the frequencies of K-mers with the same first k bases into a set in the order of the remaining base letters of the K-mer, splice all the sets row by row according to the first k base letters to form the frequency feature matrix of the target region, form the multi-channel frequency feature matrix from the frequency feature matrices of all target regions, and input the multi-channel frequency feature matrix into a neural network model with a channel attention module for training;

[0076] A screening module, which is used to obtain a multi-channel frequency feature matrix of the synthetic biological probiotics to be screened from the DNA sequences of the synthetic biological probiotics to be screened, and input the multi-channel frequency feature matrix into the trained neural network model to obtain the scores of the synthetic biological probiotics.

[0077] Preferably, the frequencies of K-mers with the same first k bases are sequentially placed into a set according to the alphabetical order of the remaining bases of the K-mer. Specifically:

[0078] Obtain the frequencies corresponding to all K-mers with the same first k bases, sort the frequencies according to the alphabetical order of the remaining bases, and place the sorted frequencies into the set identified by the first k bases.

[0079] Preferably, the channel attention module specifically includes:

[0080] A channel weight calculation unit, which is used to count the functional elements of the target region, set different weights for different functional elements, calculate the weights of all functional elements in the target region, and normalize the weights of all target regions to obtain the normalized channel weights, thereby obtaining a channel weight vector;

[0081] A channel attention calculation unit, which is used to calculate a channel attention vector through a channel attention mechanism;

[0082] A fusion unit, which is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

[0083] Preferably, the channel attention module specifically includes:

[0084] A channel weight unit, which is used to obtain the frequency of each base appearing at the same position in the target region of all beneficial bacteria in the training samples for each target region, and obtain the frequency of each base appearing at the same position in the target region of all non-beneficial bacteria in the training samples. Based on the frequencies, obtain the contribution degree of the target region to distinguish beneficial bacteria and non-beneficial bacteria, and normalize the contribution degrees of all target regions to obtain the normalized channel weights, thereby obtaining a channel weight vector;

[0085] A channel attention unit, which is used to calculate a channel attention vector through a channel attention mechanism;

[0086] A fusion unit, which is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

[0087] Preferably, the fusion of the channel weight vector and the channel attention vector is specifically:

[0088] After inputting the channel weight vector and the channel attention vector into the fully connected layer, a fusion result is obtained through an activation function.

[0089] In the third embodiment of the present invention, the present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method of the first embodiment is implemented.

[0090] In the fourth embodiment of the present invention, the present invention also proposes a computer device, which at least includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the method of the first embodiment is implemented.

[0091] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0092] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Other embodiments can also be adopted; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for screening synthetic biological probiotics, characterized in that: The following steps are involved: Obtain the DNA sequences of beneficial bacteria and non-beneficial bacteria in the training samples, search for regions with the same position and the same DNA fragments in all DNA sequences, and use the remaining regions in the DNA sequences as target regions; wherein the length of the DNA fragment is not less than a preset value; Count the K-mer frequencies in the target region, and put the frequencies of K-mers with the same first k bases into a set in the alphabetical order of the remaining bases of the K-mer. Specifically, obtain the frequencies corresponding to all K-mers with the same first k bases, sort the frequencies in the order of the remaining base letters, and put the sorted frequencies into the set of the first k base identifiers; splice all the sets into the frequency feature matrix of the target region in row order according to the alphabetical order of the first k bases, form a multi-channel frequency feature matrix from the frequency feature matrices of all the target regions, and input the multi-channel frequency feature matrix into a neural network model with a channel attention module for training; Among them, the channel attention module includes a channel weight calculation unit, which is used to count the functional elements of the target area, set different weights for different functional elements, calculate the weights of all functional elements in the target area, normalize the weights of all target areas to obtain the normalized channel weights, and obtain the channel weight vector; the channel weight calculation unit is used to obtain the frequency of occurrence of each base at the same position in the target area in all beneficial bacteria in the training sample for each target area, and obtain the frequency of occurrence of each base at the same position in the target area in all non-beneficial bacteria in the training sample, obtain the contribution of the target area to distinguishing beneficial bacteria from non-beneficial bacteria based on the frequency, normalize the contribution of all target areas to obtain the normalized channel weights, and obtain the channel weight vector; A multi-channel frequency feature matrix of the synthetic biological probiotics to be screened is obtained from the DNA sequence of the synthetic biological probiotics to be screened, and the multi-channel frequency feature matrix is ​​input into the trained neural network model to obtain the score of the synthetic biological probiotics.

2. The method for screening synthetic biological probiotics according to claim 1, characterized in that: The channel attention module further includes: A channel attention calculation unit is used to calculate a channel attention vector through a channel attention mechanism; The fusion unit is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

3. The method for screening synthetic biological probiotics according to claim 2, characterized in that: The fusion of the channel weight vector and the channel attention vector is specifically as follows: The channel weight vector and the channel attention vector are input into the fully connected layer and then passed through an activation function to obtain a fusion result.

4. A screening system for synthetic biological probiotics, characterized in that: Includes the following modules: A deduplication module is used to obtain the DNA sequences of beneficial bacteria and non-beneficial bacteria in the training samples, search for regions with the same position and the same DNA fragments in all DNA sequences, and use the remaining regions in the DNA sequences as target regions; the length of the DNA fragment is not less than a preset value; The model training module is used to count the K-mer frequencies in the target area, and put the frequencies of K-mers with the same first k bases into a set in sequence according to the alphabetical order of the remaining bases of the K-mer. Specifically, the frequencies corresponding to all K-mers with the same first k bases are obtained, and the frequencies are sorted in the order of the remaining base letters, and the sorted frequencies are put into the set of the first k base identifiers; all the sets are spliced ​​row by row according to the alphabetical order of the first k bases into the frequency feature matrix of the target area, and the frequency feature matrices of all the target areas are formed into a multi-channel frequency feature matrix, and the multi-channel frequency feature matrix is ​​input into a neural network model with a channel attention module for training; Among them, the channel attention module includes a channel weight calculation unit, which is used to count the functional elements of the target area, set different weights for different functional elements, calculate the weights of all functional elements in the target area, normalize the weights of all target areas to obtain the normalized channel weights, and obtain the channel weight vector; the channel weight calculation unit is used to obtain the frequency of occurrence of each base at the same position in the target area in all beneficial bacteria in the training sample for each target area, and obtain the frequency of occurrence of each base at the same position in the target area in all non-beneficial bacteria in the training sample, obtain the contribution of the target area to distinguishing beneficial bacteria from non-beneficial bacteria based on the frequency, normalize the contribution of all target areas to obtain the normalized channel weights, and obtain the channel weight vector; The screening module is used to obtain a multi-channel frequency feature matrix of the synthetic biological probiotics to be screened from the DNA sequence of the synthetic biological probiotics to be screened, and input the multi-channel frequency feature matrix into the trained neural network model to obtain the score of the synthetic biological probiotics.

5. The screening system for synthetic biological probiotics according to claim 4, characterized in that: The channel attention module further includes: A channel attention calculation unit is used to calculate a channel attention vector through a channel attention mechanism; The fusion unit is used to fuse the channel weight vector and the channel attention vector, and use the fusion result as the output of the channel attention module.

6. A computer storage device having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method and device for identifying specific region in microorganism target fragment and application

    CN111477274A

  • Method for accurately identifying beneficial bacteria according to whole genome sequence characteristics

    CN117668693A