Rapid evaluation system for screening of biocontrol bacteria for fruit tree continuous cropping obstacle soil
By constructing a rapid evaluation system for soil biocontrol bacteria that cause fruit tree continuous cropping obstacles, and combining gene screening, activity sieving, and field potency prediction models, the problems of unstable and long screening cycles in existing biocontrol bacteria technologies have been solved. This has enabled efficient and accurate screening of biocontrol bacteria, which is suitable for rapid remediation of soils with different fruit tree continuous cropping.
Patent Information
- Application Number
- CN202511353157.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Existing methods for screening soil biocontrol bacteria that cause continuous cropping obstacles in fruit trees rely on single-function evaluation, neglecting microbial interactions and ecological adaptability. This leads to unstable field application results, long screening cycles, and high costs, failing to meet the needs of rapid orchard restoration.
A rapid evaluation system for soil biocontrol bacteria that cause fruit tree continuous cropping obstacles was adopted, including gene screening, activity screening and field potency prediction model. Biocontrol bacteria were accurately screened through microbiome analysis and machine learning model, and a three-level progressive evaluation system was constructed, which combined microbiome data, in vitro functional verification and field potency prediction.
It achieves precise and efficient screening of biocontrol bacteria, shortening the screening cycle from the traditional 6 months to 8 weeks, reducing costs by 70%, increasing screening accuracy by 32%, and providing accurate and reliable field efficacy prediction. It is applicable to fruit tree continuous cropping soils with different pH and salinity, providing an efficient technology for managing continuous cropping obstacles.
Smart Images

Figure CN120850126B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of agricultural microorganism technology, and particularly relates to a rapid evaluation system for screening of soil biocontrol bacteria for fruit tree continuous cropping obstacles. BACKGROUND
[0002] With the limitation of land resources, industrial scale, intensification, advantage, dwarf cultivation and variety update, apple orchard reconstruction and continuous cultivation have become inevitable. Most of the renewal of old orchards is carried out on site, thus leading to the increasing aggravation of apple orchard soil continuous cropping obstacles. In addition, the factors causing apple continuous cropping obstacles are many and complex, which makes the continuous cropping obstacles become an important bottleneck restricting the sustainable development of apple industry in our province.
[0003] Apple continuous cropping obstacle, also known as apple replant disease, is an important obstacle for establishing high-yield orchards on old orchard soil, which mainly manifests as weak growth of the aboveground and underground parts of young trees, decreased apple yield and quality, and even direct plant death in severe cases. The factors of apple continuous cropping obstacle mainly include biological factors and non-biological factors. Soil disinfection can effectively alleviate apple continuous cropping obstacle, so the conclusion that biological factors are the key factors of apple continuous cropping obstacle is widely recognized. During the normal growth of crops, the beneficial microorganism species and quantity in the soil microbial community maintain a balance with the pathogenic microorganisms, while the directional selection of continuous cropping on soil microorganisms destroys the rhizosphere microecological balance. Therefore, the continuous cropping obstacle is a systematic problem, and its formation and aggravation is not caused by a single or isolated factor, but by the comprehensive action of multiple factors between the "plant-soil-microorganism" in the artificial cultivation system. Therefore, by clarifying the change trend and microecological characteristics of the apple rhizosphere microbial community structure, exploring and establishing a technical system of biological resource diversity, and restoring the function of the rhizosphere microecosystem, the reconstruction of biological relationship based on energy flow characteristics becomes a green repair strategy for apple continuous cropping obstacle.
[0004] However, the existing screening methods and mechanisms of biocontrol bacteria for continuous cropping soil still have the following shortcomings:
[0005] 1) Traditional biocontrol bacteria screening relies on single function (such as inhibition rate) evaluation, ignores microbial community interaction and ecological adaptability, and leads to unstable field application effect.
[0006] 2) The microbial community structure of continuous cropping soil is complex, and the dynamic changes of pathogenic bacteria (such as Fusarium and Cylindrocarpon) and beneficial bacteria are difficult to accurately capture.
[0007] 3) The conventional screening period is long (several months) and the cost is high, which cannot meet the rapid repair needs of orchards. SUMMARY
[0008] The present application aims to solve the problems existing in the prior art, especially for the prevention and treatment of apple and other fruit tree continuous cropping obstacles, and provides a biocontrol bacteria rapid screening and evaluation system combined with microbiome analysis, function verification and field efficacy prediction, namely a fruit tree continuous cropping obstacle soil biocontrol bacteria screening rapid evaluation system.
[0009] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0010] The fruit tree continuous cropping obstacle soil biocontrol bacteria screening rapid evaluation system comprises the following steps:
[0011] S1, gene screening
[0012] The microorganisms in the continuous cropping soil sample are separated, identified and sequenced to obtain metagenomic sequencing results, including bacterial / fungal community structure, pathogen abundance and functional gene prediction;
[0013] According to the metagenomic sequencing results, the positive effects of various microorganisms on fruit tree continuous cropping are determined, the microorganisms are separated into growth-promoting related genes (such as nitrogen fixation gene nifH, phosphorus dissolving gene phoD, iron carrier producing gene) and pathogen markers (such as Fusarium-ITS specific primer), and indigenous bacteria (such as Bacillus) with growth-promoting and pathogen-inhibiting gene potential are preferentially separated;
[0014] S2, activity screening
[0015] A modular detection platform for in vitro verification of candidate strains is constructed to detect the antibacterial efficiency, growth-promoting activity and stress resistance of the candidate strains, the growth-promoting activity includes IAA yield, phosphorus dissolving amount and iron carrier content, and the stress resistance includes salt tolerance and acid and alkali tolerance;
[0016] S3, field efficacy prediction model
[0017] A machine learning model is set and imported, and reference is made to training data including historical strain field test results (such as apple young tree biomass improvement rate and disease inhibition rate), input variables including activity screening scores and prediction factor data obtained in S2, and output field efficacy grades (A / B / C grades) to obtain a field efficacy prediction model.
[0018] The present application realizes precise and efficient screening of biocontrol bacteria by constructing a "three-level progressive" rapid evaluation system, integrating microbiome data, in vitro function verification and field efficacy prediction model.
[0019] Preferably, the metagenomic sequencing results of the continuous cropping soil sample contain the following data:
[0020] 1) Bacterial community structure, represented as a relative abundance vector at genus level:
[0021] f{B}=(b1,b2…b i …b m );
[0022] b i This represents the relative abundance of the i-th bacterial genus, where m is the number of bacterial genera.
[0023] 2) Fungal community structure, represented as a relative abundance vector at the genus level:
[0024] f{F}=(f1,f2…f j …f n );
[0025] f j This represents the relative abundance of the j-th fungal genus, where n is the number of fungal genera.
[0026] 3) Abundance of specific pathogen markers:
[0027] For example, the abundance of the genus *Fusarium* is denoted as p. fus It can be obtained through amplicon sequencing using specific primers such as Fusarium-ITS;
[0028] 4) Functional gene prediction: Gene family abundance predicted by various bioinformatics tools (such as HUMAnN2).
[0029] Preferably, the genes related to growth promotion include:
[0030] ① Nitrogen fixation gene (nifH): Abundance is denoted as g nifH ;
[0031] ② Phospho-solubilizing gene (phoD): Abundance is denoted as g phoD ;
[0032] ③ Siderogenic vector genes (such as acbA, acbB, etc.): Abundance is denoted as g siderophore .
[0033] Preferably, the isolation priority of indigenous bacteria with the potential to promote growth and inhibit pathogens is determined by the following steps:
[0034] 1) Construct a positive indicator function (potential for growth)
[0035] ;
[0036] w i For gene weights, g i To represent the abundance of growth-related genes, let P be the set of growth-related genes; for each element i in set P, extract its corresponding weight wi and growth-related quantity gi, then sum their products to obtain... A comprehensive measure;
[0037] 2) Constructing a negative indicator function (pathogen risk)
[0038] ;
[0039] where p fus is the abundance of Fusarium genus;
[0040] 3) Constructing a genus isolation priority model
[0041] Constructing a target genus screening function according to 1) and 2):
[0042] Priority(τ) = α·φ growth ·δ(τ) + β·(1-ψ pathogen )·ρ(τ);
[0043] where δ(τ) is the relative abundance of the genus, ρ(τ) is the coefficient related to antagonism against pathogenic fungi, and α + β = 1;
[0044] When Priority(τ) ≥ 0, it is preferentially isolated, otherwise it is selected for isolation or not.
[0045] Preferably, the modular detection platform comprises the following processes:
[0046] 1) Constructing a test vector:
[0047] ;
[0048] t1 is the antibacterial efficacy; t2 is the IAA yield, directly determined by colorimetry; t3 is the phosphorus solubilization amount: determined by molybdenum blue colorimetry; t4 is the iron carrier production content: CAS color diameter ratio; t5 is the salt tolerance: maximum tolerated NaCl concentration; t6 is the acid tolerance, and t7 is the alkali tolerance: growth pH range;
[0049] 2) Normalization
[0050] Normalizing t to [0, 1]:
[0051] ;
[0052] where t i refers to any one of t1 to t7 in the test vector t, and the extreme value is taken from all strains tested in the same batch;
[0053] 3) Comprehensive scoring model
[0054] Weighted scoring function:
[0055] ;
[0056] w iis referred to as any one of w1-w7 in the following weight distribution matrix w, which is multiplied by t i one-to-one correspondence:
[0057] ;
[0058] will be converted through the following index conversion function (i.e., x= ):
[0059] ;
[0060] i corresponds to the i-th row in the test vector t, i.e., t1, t2, t3 after normalization are exponentially transformed, and t4-t7 after normalization do not need to be transformed;
[0061] is a slope-related parameter, representing the steepness of the discrete fitting curve of multiple point measurements; is an offset parameter between the measured value and the normalized value, used to control the shape or characteristics of the function; when Score(s)≥0.80, the candidate strain is preferentially isolated, otherwise it is selected for isolation or not isolated.
[0062] The model realizes environmental screening decisions based on biocontrol efficacy, growth-promoting ability, and environmental adaptability, highlights core functions (40% of inhibition rate), quantifies multi-functional synergy (40% of growth-promoting activity), and guarantees field adaptability (20% of stress resistance), providing high-quality candidate strains for subsequent composite microbial agent construction.
[0063] Preferably, the machine learning model is specifically a multi-modal hybrid algorithm framework that improves the prediction accuracy and robustness of the field performance of biocontrol agents. The model is named the BCE-Net model (BioControl-Efficacy Net). The BCE-Net model architecture includes:
[0064] Input layer: input prediction factors and S2 activity screening scores;
[0065] Feature engineering module:
[0066] Multi-modal fusion module;
[0067] Dual-branch prediction module;
[0068] Output layer: output the impact evaluation mechanism of different influencing factors on continuous cropping soil, i.e., the optimized screening evaluation system.
[0069] Further, the input layer data includes strain characteristics, soil microbiota perturbation, metabolomics profile, and environmental covariates, the strain characteristics include inhibition rate, IAA production, phosphorus solubilization, colonization ability; the soil microbiota perturbation includes bacteria / fungi ratio change, pathogenic bacteria abundance change rate; the metabolomics profile includes antagonistic substance peak intensity (such as iturin, fengycin); the environmental covariates include soil pH, organic matter content, average annual temperature.
[0070] Further, the machine learning model includes the following modules:
[0071] S301, feature engineering module
[0072] 1) Adversarial autoencoder (AAE) is used for metabolomics data dimension reduction:
[0073] Let the metabolomics data be , the encoder be f enc , the decoder be f dec , the latent variable be , and d be the dimension after dimension reduction, for example, 20 dimensions, retaining 95% variance information;
[0074] Encoding process: , represents the original input data, which is the input vector of the model before processing or reconstruction;
[0075] Decoding process: , represents the reconstructed output data, which is the output vector obtained after the model decodes the original input data ;
[0076] Reconstruction loss: , represents the reconstruction loss, which is used to measure the difference between the original data and the reconstructed data; is the Euclidean norm, also known as the L2 norm, is the sum of the squares of the corresponding element differences between the original data and the reconstructed data , which is used to quantify the error between the two;
[0077] Adversarial loss: through a discriminator network D, the distribution of z matches the prior distribution p(z) (usually standard normal distribution), that is:
[0078] ;
[0079] So the total loss is: ;
[0080] 2) Time series feature extraction (LSTM is used for colonization ability data)
[0081] Colony productivity data is time series data, denoted as T = {T0, T1, ..., T}. i ,…,T n}, where T i It is the measure of colonization on day i (it may be a vector, but we assume it to be a scalar here);
[0082] The iterative formula for LSTM is as follows:
[0083] Let T i Timing: Let the hidden state of the LSTM be h. t The cell state is c t ,but:
[0084] ,
[0085] ;
[0086] h t-1 The hidden state of the previous time step records the historical information of the sequence up to time (t-1), where time (t-1) corresponds to time (T-1) in array T.
[0087] c t : The input at the current time step (e.g., in a sequence model, word embeddings, feature vectors, etc. at time t);
[0088] W: Weight matrix, used to weight the "hidden state h from the previous time step". t-1 And the current input c t The linear transformation of the concatenated vector is one of the core parameters for model learning.
[0089] b: Bias vector, the offset of the linear transformation, which helps adjust the output and is also a parameter for model learning;
[0090] σ: Represents the Sigmoid function, its expression is The output value of the Sigmoid function is between (0,1);
[0091] tanh: Represents the hyperbolic tangent function, its expression is: The output value of the tanh function is between (-1, 1);
[0092] Where i t ,f t ,o t These are the input gate, forget gate, and output gate, respectively. t The candidate cell state is taken; the hidden state h at the last moment is taken. t As a representation of temporal characteristics;
[0093] 3) Standardization of environmental factors
[0094] Using Box-Cox transformation, for environmental factor e i (the i-th environmental factor), transform to:
[0095] ;
[0096] where λ is a tuning factor to adjust the type of transformation, by changing the value of λ, we can get different forms of transformation results of e i , commonly used for distribution adjustment of data, to make the data closer to normal distribution;
[0097] S302, multi-modal fusion module (cross-modal attention mechanism)
[0098] Let the strain characteristic feature vector be α∈R 15 , the soil micro-ecological disturbance feature vector be β∈R 8 , the metabolomics dimension reduction feature be γ∈R 20 , and the environmental feature be κ∈R 5 ; we concatenate the soil and metabolomics features as β′ =[soil eco ; b profile ]∈R 28 , and then perform cross-modal attention calculation with the strain characteristics;
[0099] Use linear transformation to get query, key and value:
[0100] Query vector: Q=W q ·q;
[0101] Key vector: K=W k ·k;
[0102] Value vector: V=W v ·v;
[0103] W q , W k , W v are parameter weight matrices that need to be learned in the model training process, respectively used for linear transformation of the original input vectors q, k, v to generate corresponding query vector Q, key vector K and value vector V;
[0104] Then the attention weight is:
[0105] ;
[0106] is just the transpose of K, which exchanges the rows and columns of K; d v is the dimension of the value vector; calculate is to get the similarity matrix between query (Q) and key (K), each element in this matrix reflects the matching degree of query vector and key vector, and then combined with softmax function, scaling factor and value (V) matrix, finally get the output of attention, d v is the dimension of value vector (hyperparameter);
[0107] S303, double-branch prediction module
[0108] (a) Biomass-LSTM input: fusion features F fusion (from cross-modal attention output) and environmental features E env ∈R 5 , and concatenate them as X biomass =[F fusion ;E env ];
[0109] Through an LSTM network, the output of the LSTM is:
[0110] ;
[0111] Then take the output hidden state h through a linear layer:
[0112] ;
[0113] where w b is the weight vector matrix, w b T is the transpose of w b row and column interchanging, b b is the bias term of the learnable parameter, by transposing the weight vector w b and performing matrix multiplication (inner product) with the feature vector h, and then adding the bias term b b , thus obtaining the calculation result of biomass change percentage ΔBiomass%, which is usually in the form of linear transformation (such as in the linear layer of neural network), used to weight and combine the input features h and output the predicted value;
[0114] (b) Pathogen-GCN prediction branch:
[0115] Build a graph G=(V,E), where the node set V includes three types of nodes: pathogenic bacteria (such as Fusarium), beneficial bacteria (such as Bacillus), and key microbial groups (such as Chitinophaga);
[0116] Let the number of nodes be n, and the node feature matrix H (0) ∈R n×f(initial features are the relative abundance change rates of these groups in the soil, etc.);
[0117] The adjacency matrix A is constructed by the Spearman correlation coefficient: if the absolute value of the correlation coefficient between two nodes is greater than the threshold θ, then connect an edge (the edge weight is the absolute value of the correlation coefficient);
[0118] Graph convolution operation (using GCN layers):
[0119] ;
[0120] where: =A+I (add self-loop), is the degree matrix of , that is, σ is the activation function (such as ReLU), and the node representation is obtained after several layers of GCN;
[0121] Then take the average of the features of the pathogenic bacteria nodes (assuming there are n p pathogenic bacteria nodes), and get the inhibition rate through a fully connected layer:
[0122] ;
[0123] h pathogen is the aggregated feature vector related to the pathogen, which is obtained by aggregating multiple pathogen-related features; n p is the number of pathogen-related elements participating in aggregation, used to average the sum result to obtain the aggregated feature in the mean sense; is the index set of pathogen-related elements, each i in the set corresponds to an element participating in aggregation; is the feature vector of the i-th pathogen-related element at the L-th layer, where L can be understood as a certain layer of the model (such as a certain layer of the neural network), and the feature vector is the representation of element i after calculation of the layer;
[0124] ΔPathogen% is the percentage prediction value of pathogen change, that is, the model predicts the proportion of pathogen change based on the aggregated feature h pathogen ; w p is the weight vector matrix, w p T is the transpose of w p after row and column exchange, b p is the bias term of the learnable parameter, which is obtained by performing matrix multiplication (inner product) between the transposed weight vector w p and the feature vector h pathogen , and then adding the bias term b pThus, the result of pathogen variation ΔPathogen is obtained, which is usually in the form of linear transformation (such as in the linear layer of a neural network);
[0125] S304, loss function
[0126] The total loss function is:
[0127] ;
[0128] Wherein, λ1 is the coefficient of biomass variation loss, λ2 is the coefficient of pathogen variation loss, and λ3 is the coefficient of the difference between laboratory data and field data;
[0129] is the predicted biomass variation value of the model, is the true value of biomass variation, is the mean square error, which measures the square sum of the deviation of the predicted value of biomass variation from the true value. The greater the deviation, the higher the loss of biomass variation;
[0130] is the loss of pathogen variation; is the predicted pathogen variation value of the model, is the mean square error, which measures the deviation of the predicted pathogen variation;
[0131] is the maximum mean difference, which is used to align the distribution of laboratory data and field data. Assuming that the distribution of laboratory data is P and the distribution of field data is Q, then:
[0132] ;
[0133] x i represents laboratory data, y j represents field data, where φ is the feature mapping of the reproducing kernel Hilbert space (RKHS), and the Gaussian kernel function is commonly used;
[0134] S305, data augmentation
[0135] (a) GAN generates rare strain samples
[0136] Using the generator G to generate samples G(z) from noise z, the discriminator D distinguishes between real samples and generated samples, and the training objective is:
[0137] (b) Environmental disturbance simulation
[0138] Add random noise to continuous features (such as temperature, humidity):
[0139] ;
[0140] wherein δ represents a relative change amount, and δ is located in the interval (-0.1, 0.1), and δ = 0.1 represents that e is to be increased by 10% on the original basis, so as to obtain ;
[0141] The online learning mechanism is as follows:
[0142] Every time 100 groups of field data are newly added, the model parameters are incrementally updated, and let the newly added data set be D new , the current model parameters be τ, and the learning rate be η, then the update is:
[0143] ;
[0144] wherein is the partial derivative of τ, is the loss function;
[0145] S306, interpretability analysis
[0146] Using the SHAP method (S Hapley Additiveex Planations), the contribution value φ of each feature x i to the predicted value f(x) is calculated i :
[0147] ;
[0148] represents a certain "contribution degree" or "impact value" related to element i, which is used to measure the role of element i in the overall result in the context of function f and input x; x is an input vector, and f is a function that takes input x as a parameter and outputs the corresponding result, which is used to calculate the difference in function value under different inputs here;
[0149] N is a basic set containing all possible elements participating in combination, S is a subset of set N not containing i, |N| represents the number of elements of set N, and |S| represents the number of elements of subset S; (|N|-|S|-1) is the number of elements remaining in set N after excluding the elements of subset S and element i, is a weight coefficient related to the number of combinations, which is used to weight the terms corresponding to different subsets S, reflecting the probability or importance ratio of the occurrence of different subsets;
[0150] represents an input composed of the elements of subset S and element i, which is substituted into function f to obtain ; represents an input composed only of the elements of subset S, which is substituted into function f to obtain (by marginalizing other features); represents the change in function value after adding element i, which is used to measure the marginal effect brought by element i;
[0151] The above formula constitutes the mathematical expression of the BCE-Net model, and outputs a learning-optimized screening evaluation system.
[0152] Compared with the prior art, the beneficial effects of the present application are:
[0153] 1. The present application precisely targets indigenous beneficial bacteria (such as Bacillus) through a gene priority separation model (Priority(τ) function), with a separation efficiency improved by 300% and interference from pathogenic symbiotic bacteria (such as Streptomyces) avoided;
[0154] 2. The present application quantifies the synergistic effects of growth promotion, disease resistance and stress resistance through a modular active screening platform, with the detection rate of high-quality strains with a comprehensive score >0.8 improved by 32%;
[0155] 3. The present application integrates laboratory data and field environmental variables through a BCE-Net multi-modal machine learning model, with a prediction error <±4% (RMSE≤2.9) and a screening accuracy of A-class strains (biomass improvement >20% and disease inhibition rate >60%) reaching 89%, thus accurately and reliably predicting field potency, solving the problem of disconnection between laboratory and field effects; the screening period is shortened from the traditional 6 months to 8 weeks, and the field verification cost is reduced by 70% (only A-class strains need to be verified), with a significant reduction in cost and cycle; GAN data enhancement improves the detection rate of rare strains (<1% abundance) by 22%, avoiding the omission of high-efficiency strains.
[0156] 4. The evaluation system of the present application is suitable for apple, cherry and other fruit tree continuous cropping soils, and can maintain a disease inhibition rate >60% in adversity conditions of pH 4.5~8.5 and salinity <9%; and through SHAP explainability analysis, the core feature contribution (bacteriostatic rate 32.1%, soil pH 18.3%) is clear, guiding regional bacterial agent adaptation.
[0157] 5. The present application replaces the traditional trial-and-error screening with a three-level system of "gene targeting → function verification → intelligent prediction", achieving a cycle reduction of 67% and a 30-fold increase in the output rate of high-quality strains, and providing an efficient technical engine for continuous cropping obstacle management. BRIEF DESCRIPTION OF DRAWINGS
[0158] Figure 1 The figure shows the framework of the rapid evaluation system for screening biocontrol bacteria for fruit tree continuous cropping obstacle soils proposed by the present application. DETAILED DESCRIPTION
[0159] The technical solutions in the embodiments of the present application will be described below in conjunction with the prior known technology, obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments.
[0160] I. Research basis
[0161] The present application is based on the following research:
[0162] 1) The project research found that Chitinophaga, Fusarium and other diseases were significantly negatively / positively correlated in continuous cropping soil, and Bacillus (such as Bacillus subtilis) had biocontrol potential.
[0163] 2) High-throughput sequencing technology has analyzed the characteristics of microbial community structure in continuous cropping soil (such as the decrease of bacteria / fungi ratio and the decrease of fungal diversity).
[0164] II. Evaluation method
[0165] The rapid evaluation system for screening biocontrol bacteria in fruit tree continuous cropping obstacle soil includes the following steps:
[0166] S1, gene screening
[0167] The microorganisms in the continuous cropping soil samples are isolated, identified and sequenced to obtain metagenomic sequencing results, including bacterial / fungal community structure, pathogen abundance, and functional gene prediction;
[0168] According to the metagenomic sequencing results, the positive effects of various microorganisms on fruit tree continuous cropping are determined, the microorganisms are separated into growth-promoting related genes (such as nitrogen fixation gene nifH, phosphorus solubilizing gene phoD, and iron carrier producing gene) and pathogen markers (such as Fusarium-specific primer Fusarium-ITS), and the indigenous bacteria (such as Bacillus) with growth-promoting and pathogen-inhibiting gene potential are preferentially isolated;
[0169] S2, activity screening
[0170] A modular detection platform for in vitro verification of candidate strains is constructed to detect the antibacterial efficiency, growth-promoting activity and stress resistance of the candidate strains, the growth-promoting activity includes IAA yield, phosphorus solubility and iron carrier content, and the stress resistance includes salt tolerance and acid and alkali tolerance;
[0171] S3, field efficacy prediction model
[0172] A machine learning model is set up and imported, and the training data including historical strain field test results (such as apple young tree biomass increase rate and disease inhibition rate) are referred to, the input variables include the activity screening scores and prediction factor data obtained in S2, and the output is the field efficacy grade (A / B / C grade), and the field efficacy prediction model is obtained.
[0173] The present application realizes the precise and efficient screening of biocontrol bacteria by constructing a "three-level progressive" rapid evaluation system, integrating microbial community data, in vitro functional verification and field efficacy prediction model.
[0174] Preferably, the metagenomic sequencing results of the continuous cropping soil sample comprise the following data:
[0175] 1) Bacterial community structure, represented as a relative abundance vector at genus level:
[0176] f{B}=(b1,b2…b i …b m );
[0177] b i represents the relative abundance of the i-th bacterial genus, and m is the number of bacterial genera;
[0178] 2) Fungal community structure, represented as a relative abundance vector at genus level:
[0179] f{F}=(f1,f2…f j …f n );
[0180] f j represents the relative abundance of the j-th fungal genus, and n is the number of fungal genera;
[0181] 3) Abundance of specific pathogenic bacteria markers:
[0182] For example, the abundance of Fusarium, denoted as p fus , can be obtained by amplicon sequencing of specific primers such as Fusarium-ITS;
[0183] 4) Functional gene prediction: gene family abundance predicted by various bioinformatics tools (such as HUMAnN2).
[0184] Promotion-related genes include:
[0185] ① Nitrogen fixation gene (nifH): abundance denoted as g nifH ;
[0186] ② Phosphorus solubilizing gene (phoD): abundance denoted as g phoD ;
[0187] ③ Iron carrier production gene (such as acbA, acbB, etc.): abundance denoted as g siderophore .
[0188] The isolation priority of indigenous bacteria with potential for promoting and inhibiting pathogenic genes is determined by the following steps:
[0189] 1) Construct a positive index function (promotion potential)
[0190] ;
[0191] 2) Construct a negative index function (pathogenic risk)
[0192] ;
[0193] where p fus is the abundance of Fusarium genus;
[0194] 3) Constructing isolation priority model of genus
[0195] Constructing target genus screening function according to 1) and 2):
[0196] Priority(τ) = α·φ growth ·δ(τ) + β·(1-ψ pathogen )·ρ(τ);
[0197] where δ(τ) is the relative abundance of genus, ρ(τ) is the coefficient related to antagonism of pathogenic fungi, and α + β = 1;
[0198] When Priority(τ) ≥ 0, it is preferentially isolated, otherwise it is selected for isolation or not isolated.
[0199] Modular detection platform includes the following processes:
[0200] 1) Constructing test vector:
[0201] ;
[0202] t1 is the antibacterial efficacy; t2 is the IAA yield, which is directly determined by colorimetry; t3 is the phosphorus solubilization amount, which is determined by molybdenum blue colorimetry; t4 is the iron carrier production content, which is determined by CAS color diameter ratio; t5 is the salt tolerance, which is the maximum tolerance NaCl concentration; t6 is the acid tolerance, and t7 is the alkali tolerance, which is the growth pH range;
[0203] 2) Normalization
[0204] Normalize t to [0, 1]:
[0205] ;
[0206] where t i refers to any one of t1 to t7 in the test vector t, and the extreme value is taken from all strains tested in the same batch;
[0207] 3) Comprehensive scoring model
[0208] Weighted scoring function:
[0209] ;
[0210] w i refers to any value of w1-w7 in the following weight distribution matrix w, which is in one-to-one correspondence with t i in the test vector t:
[0211] ;
[0212] will be converted by the following index conversion function (i.e., x= ):
[0213] ;
[0214] i corresponds to the ith row in the test vector t, i.e., the normalized t1, t2, t3 are exponentially transformed, and the normalized t4-t7 do not need to be transformed;
[0215] The model realizes environmental screening decisions based on biocontrol efficacy, growth promotion ability, and environmental adaptability, highlights core functions (bacteriostatic rate accounts for 40%), quantifies multi-functional synergy (growth-promoting activity accounts for 40%), and guarantees field adaptability (stress resistance accounts for 20%), providing high-quality candidate strains for subsequent composite microbial agent construction.
[0216] The machine learning model is specifically a multi-modal hybrid algorithm framework that improves the prediction accuracy and robustness of the field performance of biocontrol agents. The model is named BCE-Net model (BioControl-Efficacy Net). The BCE-Net model architecture includes:
[0217] Input layer: input prediction factors and S2 activity screening scores;
[0218] Feature engineering module:
[0219] Multi-modal fusion module;
[0220] Dual-branch prediction module;
[0221] Output layer: output the impact evaluation mechanism of different influencing factors on continuous cropping soil, i.e., the optimized screening evaluation system.
[0222] Further, the input layer data includes strain characteristics, soil microecological disturbance, metabolomics, and environmental covariates. Strain characteristics include bacteriostatic rate, IAA yield, phosphorus solubility, and colonization ability. Soil microecological disturbance includes changes in bacteria / fungus ratio and pathogen abundance change rate. Metabolomics includes peak intensity of antagonistic substances (such as iturin and fengycin). Environmental covariates include soil pH, organic matter content, and average annual temperature.
[0223] Further, the machine learning model includes the following modules:
[0224] S301, feature engineering module
[0225] 1) Adversarial autoencoder (AAE) for metabolomics data dimensionality reduction:
[0226] Let the metabolomics data be , the encoder be f enc , the decoder be f dec , the latent variable be , d be the dimension after dimension reduction, for example, 20 dimensions, and 95% variance information be retained;
[0227] Encoding process: ;
[0228] Decoding process: ;
[0229] Reconstruction loss: ;
[0230] Adversarial loss: match the distribution of z to the prior distribution p(z) (usually standard normal distribution) through a discriminator network D, that is:
[0231] ;
[0232] So the total loss is: ;
[0233] 2) Time series feature extraction (LSTM is used for colonization data)
[0234] The colonization data is time series data, denoted as T = {T0, T1, …, T i , …, T n}, where T i is the colonization measurement value on the i-th day;
[0235] The iteration formula of LSTM is as follows:
[0236] Let T i be the time: the hidden state of LSTM is h t , and the cell state is c t , then:
[0237]
[0238] ;
[0239] 3) Environmental factor standardization
[0240] Using Box-Cox transformation, for the environmental factor e i (the i-th environmental factor), transform it as:
[0241] ;
[0242] S302, multi-modal fusion module (cross-modal attention mechanism)
[0243] Let the strain property feature vector be α∈R 15 , the soil microbiome disturbance feature vector be β∈R 8 , the metabolomic reduced dimension feature be γ∈R 20 (from AAE), and the environmental feature be κ∈R 5 (after standardization); we concatenate the soil and metabolomic features as β′ =[soil eco ; b profile ]∈R 28 , and then compute cross-modal attention with strain properties;
[0244] The query, key, and value are obtained using linear transformations:
[0245] Query vector: Q=W q ·q(W q is a 15×d v matrix);
[0246] Key vector: K=W k ·k(W k is a 28×d v matrix);
[0247] Value vector: V=W v ·v(W v is a 28×d v matrix);
[0248] Then the attention weights are:
[0249] ;
[0250] S303, double-branch prediction module
[0251] (a) Biomass-LSTM input: fused features F fusion (from cross-modal attention output) and environmental features E env ∈R 5 , concatenated as X biomass =[F fusion ; E env ];
[0252] Through an LSTM network, the output of the LSTM is:
[0253] ;
[0254] Then take the output hidden state h through a linear layer:
[0255] ;
[0256] (b) Pathogen-GCN for disease suppression prediction:
[0257] Build graph G = (V, E), where node set V includes three types of nodes: pathogen (e.g. Fusarium), beneficial bacteria (e.g. Bacillus), and key microbial community (e.g. Chitinophaga);
[0258] Let the number of nodes be n, and the node feature matrix H (0) ∈R n×f (initial features are the relative abundance change rates of these microbial communities in the soil, etc.);
[0259] Adjacency matrix A is constructed by Spearman correlation coefficient: if the absolute value of the correlation coefficient between two nodes is greater than the threshold θ, then connect an edge (edge weight is the absolute value of the correlation coefficient);
[0260] Graph convolution operation (using GCN layer):
[0261] ;
[0262] Where: =A+I (add self-loop), is the degree matrix of , i.e. , σ is the activation function (such as ReLU), and after several layers of GCN, the node representation is obtained;
[0263] Then take the average of the features of the pathogen nodes (assuming there are n p The loss function is:
[0264] ;
[0265] S304, loss function
[0266] The total loss function is:
[0267] ;
[0268] is the predicted biomass change value of the model, is the true value of biomass change, is the mean square error, which measures the squared sum of the deviation between the predicted value and the true value of biomass change. The larger the deviation, the higher the loss of biomass change;
[0269] is the loss of pathogen change; is the predicted pathogen change value of the model, is the mean square error, which measures the deviation of the predicted pathogen change;
[0270] is the maximum mean discrepancy used to align the distribution of lab data and field data, let P be the distribution of lab data and Q be the distribution of field data, then:
[0271] ;
[0272] S305, data augmentation
[0273] (a) GAN generates rare strain samples
[0274] Using generator G to generate samples G(z) from noise z, discriminator D to distinguish real samples and generated samples, training objective;
[0275] (b) environmental disturbance simulation
[0276] Add random noise to continuous features (such as temperature, humidity):
[0277] ;
[0278] Online learning mechanism as follows:
[0279] Every time 100 sets of field data are added, the model parameters are updated incrementally, let the added data set be D new , the current model parameters are τ, and the learning rate is η, then update:
[0280] ;
[0281] Where is the partial derivative of τ, is the loss function;
[0282] S306, explainability analysis
[0283] Using SHAP method (S Hapley Additiveex Planations), calculate the contribution value φ i of each feature x i to the predicted value f(x):
[0284] .
[0285] Three, practical application
[0286] On the basis of II, evaluation method, design example 1, comparative example 1 and comparative example 2:
[0287] Example 1
[0288] Screening of biocontrol bacteria for apple continuous cropping obstacles
[0289] 1. Gene screening
[0290] Soil sample: 5-year-old apple continuous cropping soil in Yantai, Shandong (pH 6.8, organic matter 1.2%);
[0291] Metagenomic sequencing results:
[0292] Bacterial community: Bacillus (12.3%), Pseudomonas (8.1%);
[0293] Fungal community: Fusarium (15.7%, pathogenic fungus);
[0294] Functional genes:
[0295] Nitrogen-fixing gene nifH abundance: g nifH =0.58;
[0296] Phosphorus-dissolving gene phoD abundance: g phoD =1.24;
[0297] Iron carrier-producing gene g siderophore =0.91;
[0298] Bacterial genus isolation priority calculation:
[0299] Bacillus (τ=Bacillus):
[0300] φ growth =0.3×0.58+0.4×1.24+0.3×0.91=1.05;
[0301] ψ pathogen =0.15×15.7%=0.0236;
[0302] Priority(τ)=0.6×1.05×12.3%+0.4×(1-0.0236)×0.85 (literature antagonistic coefficient)=0.092, should be preferentially isolated;
[0303] 2. Active screening
[0304] Candidate strains: 3 Bacillus strains (B1, B2, B3) were isolated;
[0305] The results of the modular detection platform are shown in Table 1 below:
[0306] Table 1. Modular detection results
[0307]
[0308] Comprehensive score (take B3 as an example):
[0309] Normalization: t1' = (92-78) / (92-78) = 1.0 (inhibition rate);
[0310] Weighted score:
[0311] Score = 0.4x1.0 + 0.2x(52 / 60) + 0.15x(145 / 150) +... = 0.93, should be prioritized (>0.80);
[0312] 3. Field potency prediction
[0313] BCE-Net model input:
[0314] Strain properties: inhibition rate 92%, IAA 52 μg / mL, phosphorus solubilization 145 mg / L;
[0315] Soil microecology: bacteria / fungi ratio change increased by 18%, Fusarium abundance change rate decreased by 62%;
[0316] Environmental covariates: soil pH 6.8, organic matter 1.5%;
[0317] Prediction output:
[0318] Biomass improvement rate: 23.5% (Biomass-LSTM branch);
[0319] Disease inhibition rate: 67.8% (Pathogen-GCN branch);
[0320] Field potency level: A level (>20% biomass improvement and >60% disease inhibition rate).
[0321] Comparative Example 1: Traditional screening method (without machine learning)
[0322] Based on Example 1, but only relying on inhibition zone size and IAA production to screen strains;
[0323] Results: 5 strains were isolated from the same soil, and field tests showed that the biomass improvement rate was 8%-15%, the disease inhibition rate was 20%-45%, and there was no A-level strain, with the highest potency only C-level.
[0324] Comparative Example 2: Without using Priority(τ) function
[0325] Based on Example 1, but the method only uses random isolation to obtain a genus with an abundance of >10% (including Fusarium symbiotic bacteria);
[0326] Results: Streptomyces (abundance 11.2%) was isolated, and the active screening showed that the inhibition rate was only 35% (due to the presence of a weak strain), and the field test showed that the disease inhibition rate decreased by 10% (disease aggravated).
[0327] The evaluation results are shown in Table 2 below:
[0328] Table 2. Evaluation mechanism difference between the present application and traditional screening method
[0329]
[0330] As can be seen from Table 2, the BCE-Net model verification results have the following advantages:
[0331] 1) High prediction accuracy:
[0332] Biomass improvement rate prediction error: ± 2.1% (RMSE = 1.8);
[0333] Disease inhibition rate prediction error: ± 3.7% (RMSE = 2.9);
[0334] 2) High explainability analysis:
[0335] The core feature contribution can be obtained by the formula: inhibition rate (32.1%), colonization ability (25.7%), and soil pH (18.3%);
[0336] And the key metabolites are obtained: for example, if the iturin peak intensity increases by 1 unit, the disease inhibition rate increases by 9.2%.
[0337] Comparing Comparative Example 2 with Example 1 shows that the Priority(τ) function in Example 1 improves the separation efficiency of indigenous beneficial bacteria by 3 times, enabling precise targeted screening;
[0338] Comparing Comparative Example 1 with Example 1 shows that the comprehensive scoring model of Example 1 can solve the non-linear trade-off of growth promotion-disease resistance-stress resistance (e.g., growth promotion efficiency saturates when IAA production exceeds 50 μg / mL), and can realize multifunctional synergistic quantification;
[0339] It also shows that the BCE-Net of Example 1 improves the prediction accuracy from 82% to 89% after adding 100 new data sets through an online learning mechanism; and through GAN data enhancement, the screening success rate of rare strains (<1% abundance) is increased by 22%, i.e., the machine learning generalization ability is realized. This system shortens the biocontrol bacteria screening period to 1 / 3 of the traditional method, and the output rate of A-level strains is increased by more than 30%, and through the multi-modal fusion and field adaptability learning of BCE-Net, the core problem of the disconnection between laboratory and field efficacy is solved, providing an efficient technical tool for continuous cropping obstacle management.
[0340] The above merely describes preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent replacements or changes within the technical scope disclosed by the present application and according to the technical solutions and inventive concept of the present application, which should be covered within the protection scope of the present application.
Claims
1. A rapid evaluation system for screening of biocontrol bacteria for fruit tree replant barrier soils, characterized in that, Comprising the following steps: S1, gene screening Isolating, identifying and sequencing the microorganisms in the continuous cropping soil sample to obtain metagenomic sequencing results, including bacterial / fungal community structure, pathogen abundance, and functional gene prediction; According to the metagenomic sequencing results, determine the positive effects of various microorganisms on fruit tree continuous cropping, separate the microorganisms into growth-promoting related genes and pathogen markers, and preferentially isolate indigenous bacteria with growth-promoting and pathogen-inhibiting gene potential; The isolation priority of the indigenous bacteria with growth-promoting and pathogen-inhibiting gene potential is determined by the following steps: 1) Construct a positive index function ; w i is the gene weight, g i is the abundance of the growth promoting related gene, P is the set of growth promoting related genes; 2) Construct a negative index function ; where p fus is the abundance of Fusarium; 3) Construct a bacterial genus isolation priority model According to 1) and 2), construct a target bacterial genus screening function: Priority(τ) = a · φ growth · δ(τ) + β · (1 - ψ pathogen ) · ρ(τ); δ(τ) is the relative abundance of the bacterial genus, ρ(τ) is the coefficient related to the antagonism of the pathogen, and α+β=1; When Priority(τ)≥0, preferentially isolate, otherwise, select to isolate later or not to isolate; S2, activity screening Construct a modular detection platform for in vitro verification of candidate strains to detect the antibacterial efficacy, growth-promoting activity and stress resistance of the candidate strains, the growth-promoting activity includes IAA yield, phosphorus solubility and iron carrier content, the stress resistance includes salt tolerance and acid and alkali tolerance, and an activity screening score is obtained, which is required to be ≥0.8; Constructing the modular detection platform includes the following processes: 1) Construct a test vector: ; t1 is the antibacterial efficacy; t2 is the IAA yield, directly using colorimetric method to determine the value; t3 is the phosphorus solubility: molybdenum blue colorimetric method to determine the value; t4 is the iron carrier content: CAS color development diameter ratio; t5 is the salt tolerance: maximum tolerance NaCl concentration; t6 is the acid tolerance, t7 is the alkali tolerance: growth pH range; 2) Normalization Normalize t to [0, 1]: ; where t i Refers to any one of t1 to t7 in the test vector t, the extreme value is taken from all strains tested in the same batch; 3) Comprehensive scoring model Weighted scoring function: ; w i refers to any one of the values w1-w7 in the following weight distribution matrix w, which is in one-to-one correspondence with t i in the test vector t: ; Transforming by the following index transformation function: ; i corresponds to the ith row in the test vector t, that is, the exponential transformation of normalized t1, t2, t3, while the normalized t4-t7 do not need to be transformed; is a slope-related parameter, representing the steepness of the discrete fitting curve of the multi-point actual measurement; is a deviation parameter between the actual measurement value and the normalized value, used to control the shape or characteristics of the function; when Score(s)≥0.80, the candidate strain is preferentially isolated, otherwise, it is selected for later isolation or not isolated; S3, field efficacy prediction model Set and import a machine learning model, and train the data according to the training data, which includes historical strain field test results, input variables, variables include activity screening score and prediction factor data obtained in S2, output field efficacy grade, and obtain field efficacy prediction model.
2. The rapid evaluation system for screening the biocontrol bacteria against the soil obstacles of the continuous cropping of fruit trees according to claim 1, characterized in that, The metagenomic sequencing results of the continuous cropping soil sample include the following data: 1) Bacterial community structure, represented as a relative abundance vector at genus level: f{B}=(b1,b2…b i …b m ) b i represents the relative abundance of the i-th bacterial genus, m is the number of bacterial genera; 2) Fungal community structure, represented as a relative abundance vector at genus level: f{F}=(f1,f2…f j …f n ) f j Rj represents the relative abundance of the jth fungal genus, and n is the number of fungal genera. 3) Abundance of specific pathogen markers: The abundance of the pathogen Fusarium was noted as p fus ; 4) Functional gene prediction: gene family abundance predicted by various bioinformatics tools.
3. The rapid evaluation system for screening the biocontrol bacteria against the soil obstacles of the continuous cropping of fruit trees according to claim 1, characterized in that, The growth-promoting related genes include: Bacteria with the nitrogen fixation gene: abundance is denoted as g nifH ; ii) Phosphate solubilizing genes: g phoD ; iii. siderophore production genes: abundance noted as g siderophore .
4. The rapid evaluation system for screening the biocontrol bacteria against the soil obstacles of the continuous cropping of fruit trees according to claim 1, characterized in that, The machine learning model is specifically a multi-modal hybrid algorithm framework to improve the prediction accuracy and robustness of the field performance of biocontrol bacteria, which is named BCE-Net model, and the architecture includes: Input layer: input prediction factors and S2 activity screening score; Feature engineering module: Multi-modal fusion module; Double-branch prediction module; Output layer: output the influence evaluation mechanism of different influencing factors on continuous cropping soil, that is, the optimized screening evaluation system.
5. The rapid evaluation system for screening the biocontrol bacteria against the soil obstacles of the continuous cropping of fruit trees according to claim 4, characterized in that, The input layer data includes strain characteristics, soil micro-ecological disturbance, metabolome profile and environmental covariates, the strain characteristics include inhibition rate, IAA yield, phosphorus solubilization amount, colonization ability; the soil micro-ecological disturbance includes the change of bacteria / fungi ratio, the change rate of pathogen abundance; the metabolome profile includes the peak intensity of antagonistic substances; the environmental covariates include soil pH, organic matter content, average annual temperature.
6. The rapid evaluation system for screening the biocontrol bacteria against the soil obstacles of the continuous cropping of fruit trees according to claim 4, characterized in that, The machine learning model comprises the following modules: S301, a feature engineering module 1) Adversarial autoencoder for metabolome data dimensionality reduction: Let the metabolome data be , the encoder be f enc , the decoder be f dec , the latent variable be , d be the dimension after dimension reduction, and be 20 dimensions, retaining 95% variance information; Encoding process: , represents the raw input data, which is the input vector of the model before it has been processed or reconstructed; Decoding process: , The reconstructed output data represents an output vector obtained by the model after decoding the original input data . reconstruction loss: , denotes the reconstruction loss, which measures the difference between the original data and the reconstructed data; is the Euclidean norm, also known as the L2 norm, is the original data and the reconstructed data corresponding element difference, so as to quantify the error between the two. Adversarial loss: through a discriminator network D, the distribution of z is matched with the prior distribution p(z), that is: ; So the total loss is: ; 2) Time series feature extraction Colony productivity data is time series data, denoted as T = {T0, T1, ..., T}. i ,…,T n }, where T i It is the measure of fertility on day i; The iteration formula of LSTM is as follows: Let T i Time: The hidden state of the LSTM is h t The cell state is c t ,but: 、 ; h t-1 : hidden state of last time, record history information to time (t-1), time (t-1) corresponds to (T-1) in T array c t : input at the current time; W: weight matrix, used to linearly transform the "previous hidden state h t-1 and current input c t concatenated vector" is one of the core parameters of model learning; b: bias vector, offset of linear transformation, auxiliary adjustment output, and also a parameter learned by the model; σ: represents a Sigmoid function, whose expression is The output value of the Sigmoid function is between (0, 1); tanh: represents the hyperbolic tangent function, whose expression is The output values of the tanh function are between (-1, 1); where i t f t o t are input gate, forget gate, output gate, respectively, g t is the candidate cell state; the hidden state h t at the last time step is taken as the representation of the temporal feature. 3) Environmental factor standardization Using the Box-Cox transformation, for the environmental factor e i is transformed to: ; where λ is an adjustment factor used to adjust the type of transformation, by changing the value of λ, different forms of transformation results of e i are obtained, which are used for distribution adjustment of data, so that the data is closer to the normal distribution; S302, a multi-modal fusion module Let the strain property feature vector be α ∈ R 15 , the soil micro-ecological disturbance feature vector be β ∈ R 8 , the metabolome dimension reduction feature be γ ∈ R 20 , and the environmental feature be κ ∈ R 5 ; the soil and metabolome features are spliced into β' = [soil eco ; b profile ] ∈ R 28 , and then cross-modal attention calculation is performed with the strain property. The query, key and value are obtained by linear transformation: Query vector: Q = W q • q; Key vector: K = W k • k; Value vector: V = W v • v; W q 、W k 、W v are respectively the parameter weight matrices to be learned in the model training process, and are respectively used for linear transformation of the original input vectors q, k, v to generate the corresponding query vector Q, key vector K and value vector V. Then the attention weight is: ; is the transpose of K, interchanging the rows and columns of K;d v is the dimension of the value vector; S303, a double-branch prediction module (a) bio-effect prediction branch input: fusion features F fusion and environmental features E env ∈ R 5 , which are concatenated into X biomass = [F fusion ; E env ] Through an LSTM network, the output of LSTM is: ; Then take the output hidden state h through a linear layer: ; where w b is a weight vector matrix, w b T is a row-wise transposed version of w b is a column-wise transposed version of b b is a bias term for the learnable parameters; (b) Disease inhibition prediction branch: A graph G=(V,E) is constructed, wherein the node set V includes three types of nodes: pathogenic bacteria, beneficial bacteria and key bacterial community; Let the number of nodes be n, the node feature matrix H (0) ∈R n×f ; The adjacency matrix A is constructed by the Spearman correlation coefficient: if the absolute value of the correlation coefficient between two nodes is greater than a threshold θ, an edge is connected; Graph convolution operation: ; wherein: , is the degree matrix of , σ is an activation function, and the node representation is obtained after several layers of GCN. Then the features of the pathogenic bacteria nodes are averaged, and an fully connected layer is used to obtain the inhibition rate: ; h pathogen is a pathogen-related aggregation feature vector, which is obtained by aggregating a plurality of pathogen-related features; n p is the number of pathogen-related elements participating in aggregation, used to average the summation result to obtain an aggregated feature in the sense of mean value; is an index set of pathogen-related elements, each i in the set corresponds to an element participating in aggregation; is the feature vector of the ith pathogen-related element at the Lth layer; ΔPathogen is the predicted percentage change in pathogen, i.e. the model based on aggregated features h pathogen predicted percentage change in pathogen; w p is a weight vector matrix, w p T is w p transpose of the row-column interchanged, b p is a bias term for the learnable parameters; S304, loss function The total loss function is: ; Wherein, λ1 is the coefficient of biomass change loss, λ2 is the coefficient of pathogen change loss, and λ3 is the coefficient of difference between laboratory data and field data; the predicted biomass change value for the model, the true value of the biomass change, the mean squared error; loss of pathogen change; a model predicted pathogen change value, mean squared error; is the maximum mean discrepancy used to align the distributions of the lab data and field data, where the lab data distribution is P and the field data distribution is Q. ; x i representing laboratory data, y j representing field data, where φ is a characteristic map of the reproducing kernel Hilbert space; S305, data enhancement (a) GAN generates rare strain samples A generator G is used to generate samples G(z) from noise z, a discriminator D distinguishes between real samples and generated samples, and the training target is: (b) Environmental disturbance simulation Random noise is added to continuous features: ; Wherein δ represents the relative change amount, and δ is located in the interval (-0.1, 0.1); The online learning mechanism is as follows: Every 100 groups of field data are added, the model parameters are incrementally updated, and the newly added data set is D new , the current model parameters are τ, the learning rate is η, and the update is: ; wherein is the partial derivative of τ with respect to x, is a loss function; S306, explainability analysis Using the SHAP method, the contribution value φ i to the predicted value f(x) i : ; N is a basic set containing all possible elements involved in the combination, S is a subset of set N not containing i, |N| represents the number of elements of set N, |S| represents the number of elements of subset S; (|N|-|S|-1) is the number of elements remaining in set N after removing the elements of subset S and element i, is a weight coefficient related to the number of combinations, used for weighting the terms corresponding to different subsets S, reflecting the probability or importance ratio of the occurrence of different subsets; represents the input consisting of the elements of subset S and element i together, which is substituted into function f to get ; represents the input consisting of only the elements of subset S, which is substituted into function f to get ; represents the change in function value after element i is added, which is used to measure the marginal effect brought by element i; The above formula constitutes the mathematical expression of the BCE-Net model, and outputs the learning optimized screening evaluation system.
Citation Information
Patent Citations
Soil nutrient prediction and comprehensive evaluation method based on machine learning algorithm
CN109374860A
Growth-promoting strain screening method based on high-throughput sequencing and application of method
CN111793680A