Prediction Method and Device for Decomposition Cycle of Agricultural Multifilm Based on RF-Meta Model

The RF-Meta model addresses the challenge of predicting biodegradable multifilm degradation by integrating Meta-analysis and a random forest algorithm, enabling precise and generalized degradation cycle estimation for sustainable agricultural practices.

JP2025522277AActive Publication Date: 2025-07-15JIANGSU ACAD OF AGRI SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024568868
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-25
Filing Date
2024-01-29
Publication Date
2025-07-15
Estimated Expiration
2044-01-29

AI Technical Summary

Technical Problem

Existing methods struggle to accurately predict the degradation cycle of biodegradable multifilms in agriculture due to complex interactions among material, soil, and environmental factors, leading to unpredictable and region-specific degradation processes.

Method used

An RF-Meta model combining Meta-analysis and a random forest algorithm to construct a comprehensive database, fill missing values, and optimize hyperparameters for precise degradation cycle prediction.

Benefits of technology

Accurately estimates the degradation cycle, enhances model generality, accelerates decomposition, supports informed multifilm selection, reduces labor costs, and promotes sustainable agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522277000001_ABST
    Figure 2025522277000001_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the decomposition cycle of an agricultural multi-film based on an RF-Meta model. This method includes the steps of collecting the original dataset of the decomposition of the biodegradable multi-film through Meta-analysis; ranking, through Meta-analysis, the influence QM of different factors in the original dataset on the degree of decomposition; screening to leave the factors that meet the requirements in the database and filling in the missing values in the original database using the random forest algorithm to obtain a complete database; splitting the complete database into a modeling set and a verification set and constructing an initial RF-Meta model; training the initial RF-Meta model and optimizing the hyperparameters to obtain an optimized RF-Meta model; and introducing the physical and chemical data of the initial experiment regarding the decomposition of the target biodegradable multi-film into the optimized RF-Meta model to obtain a prediction curve of the complete decomposition cycle. The model proposed by the present invention can provide, for the first time, a specific numerical prediction of the degree of decomposition in the decomposition process of BDM and has broad applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agricultural informatization, and particularly to a method and device for predicting the soil burial decomposition cycle of biodegradable multifilms through a random forest algorithm combined with a Meta-analysis method.

Background Art

[0002] Multifilm is one of the important production means in agriculture. Using multifilm can effectively suppress the growth of weeds, increase crop yields, maintain soil moisture, reduce water evaporation, and save water resources. In addition, by increasing soil temperature, promoting crop growth, reducing soil erosion, and maintaining soil fertility, the life cycle of crops can be extended, the risk of pests and diseases can be reduced, the use of pesticides and chemical fertilizers can be reduced, and the quality of grains can be improved. The amount of multifilm used in China every year is as high as 137.9 million tons. However, due to ultraviolet irradiation and the action of external forces, multifilm is easily fragmented to form plastic pieces, causing serious plastic pollution in farmland and posing a potential threat to the ecosystem and human health. Currently, the impact of plastic pollution has attracted increasing attention from the scientific community and people in society. Biodegradable plastics are plastics that can be decomposed and utilized by microorganisms and ultimately converted into microbial biomass, carbon dioxide, and water. Therefore, using biodegradable plastic multifilm (BDM) instead of PE multifilm is one of the main strategies for preventing plastic pollution in agriculture. However, the decomposition cycle of BDM in the environment is still unclear, mainly because the decomposition of BDM is a complex process and is closely related to the components of the material and natural conditions, and the decomposition effects of BDM with different formulations in different regions often vary greatly.

[0003] Currently, research on the degradation cycle of specific BDM multi-film products mainly estimates the degradation cycle based on experience after using them in local farmland for 1 to 2 years. There are relevant reports on the effects of material factors, soil factors, regional factors, etc. on the degradation of BDM, but the relevant research mainly relies on experimental data and can only judge the general trend of degradation. The research means are relatively simple and limited by the number of control groups in the experiment. Therefore, it is difficult to measure the effects of relevant influencing factors on degradation and impossible to predict the degradation cycle. In existing research, for the degradation cycle of multi-film by ultraviolet irradiation (UV), a simple fitting model was constructed to evaluate the photodegradation cycle of multi-film. The drawbacks of the above research are as follows. (1) The relevant variables affecting BDM are not only UV. (2) The degradation process of BDM is complex. In many cases, non-linear interactions occur among material factors, soil factors, climate factors, and human factors. With the analysis of a small number of factors, this interaction cannot be comprehensively analyzed, making it impossible to specifically predict the degradation process. (3) Existing research cannot solve the problem of regional characteristics of multi-film degradation. Therefore, it is urgent to develop a non-linear BDM degradation prediction model that covers more variables and can be widely applied to the regions where multi-film is used.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present invention is to provide an RF-Meta model that combines a Meta-analysis method and a random forest algorithm to predict the degradation cycle of agricultural multi-films, achieve an accurate estimation of the degradation cycle, and improve the generality of the model results.

Means for Solving the Problems

[0005] A method for predicting the degradation cycle of agricultural multi-films based on the RF-Meta model, comprising (1)Collect the original dataset of the degradation of biodegradable multifilms through meta-analysis, construct an original database based on the original dataset, and verify the reliability of the original data through meta-analysis means; (2)Rank the influence Q of different factors in the original dataset on the degree of degradation through meta-analysis; M (3)According to the meta-analysis results, select the factors with Q >1 to remain in the original database, and use the random forest algorithm to fill in the missing values in the original database to obtain a complete database; M (4)Divide the complete database into a modeling set and a verification set, and construct an initial RF-Meta model; (5)Train the initial RF-Meta model and optimize the hyperparameters to obtain an optimized RF-Meta model; (6)Introduce the physical and chemical data of the initial experiment on the degradation of the target biodegradable multifilm into the optimized RF-Meta model to obtain the prediction curve of the complete degradation cycle. (6)Introduce the physical and chemical data of the initial experiment on the degradation of the target biodegradable multifilm into the optimized RF-Meta model to obtain the prediction curve of the complete degradation cycle.

[0006] Furthermore, in the above step (1), the specific steps for constructing the original database are as follows. (1.1)According to the PICOS principle, the steps for creating the search strategy, inclusion criteria, and exclusion criteria are: P: Determine that the research object is a biodegradable multifilm in the scene of agricultural production; I: Determine the intervention means, select the range of basic material components that make up the degradable multifilm, and select different degradation means, degradation environments, and pre-treatments before degradation; C: Determine the characteristics of the multifilm before and after degradation as the control means; O: Determine that the direction of the result is the degradation rate parameter of the multifilm; S: This includes limiting the research design to laboratory experiments or field experiments and excluding reviews and opinion articles.

[0007] (1.2) Select the article database, set the search terms based on the principle of (1.1), and manually further select the articles that meet the research conditions and are repeatable. (1.3) The composition of the characteristic dataset G1 to be collected in each article is G1 = {A1, A2..., A a ; C1, C2,..., C e ; SP1, SP2,..., SP p ; SB1, SB2..., SB b ; M1, M2..., M m ; E1, E2..., E e}, where A a represents the a-th item of the basic information of the article to be collected, C e represents the e-th item of the climate conditions to be collected, SP p represents the p-th item of the physical conditions of the soil to be collected, SB b represents the b-th item of the biological and chemical conditions of the soil to be collected, M m represents the m-th item of the characteristics of the multifilm material to be collected, and E e represents the e-th item of the experimental conditions to be collected.

[0008] (1.4) The composition of the decomposition result dataset G2 to be collected in each article is G2 = {E m , E v , C m , C v}, where E m represents the average value of the results of the experimental group, E v represents the variance of the results of the experimental group, C m represents the average value of the results of the control group, and C v represents the variance of the results of the control group. The formulas for the average value and standard deviation are as follows.

Number

[0009] (1.5) By observing the relationship between the magnitude of the decomposition rate and its variance, detect whether there is publication bias in the collected articles, that is, whether there is an obvious trend in the data source, order the decomposition rates of the biodegradable multi-films obtained from the experimental data of each group according to the magnitude, and obtain the position number u of each data i and order the data of each group according to the magnitude of the experimental variance, and then obtain the position number v j and calculate the occurrence probability P of the hypothesis represented by the correlation coefficient D and D between the magnitude of the decomposition rate and its experimental variance. [Number] Here, n is the total amount of data in the database, and C n represents the number of pairs that match in each pair comparison, D n represents the number of pairs that do not match in each pair comparison, and when the two variables u i and v j increase or decrease simultaneously, they are matching pairs. When one of the two variables u i and v j increases and the other variable decreases, these are non-matching pairs. cdf(D) is the cumulative distribution function value of the D value in the standard normal distribution. By referring to the D value and the P value, it is possible to know whether there is publication bias. When the absolute value of the D value is greater than 1.96 and the p value is less than 0.05, it is considered that there is an obvious trend in the data, and it is necessary to exclude the data source with an obvious trend. If there is no obvious bias, proceed to the next step.

[0010] (1.6) If publication bias is found, first use the R language to draw a funnel chart with the decomposition rate as the X-axis data and the experimental variance as the Y-axis data. The position vector of each data in the chart is d i =(xi , s i ) and,

Number

Number

[0011] (1.7) Evaluate whether there is a general rule in the data in the database, and the significance level of the data trend is calculated as follows.

Number

[0012] Furthermore, step (2) is specifically as follows. (2.1) Scan the decomposition rate data and the variance data to check whether there are missing values or null values. If a null value is found, record the position information of the null value including the row m and column index n where the null value is located, and it is necessary to fill it using the nearest neighbor linear interpolation method.

[0013] (2.2) For the database, construct a multivariate random effects model in the Meta-analysis. The basic random effects part of the model is generated by different literature sources extracted from the data. The formula for the multivariate random effects model is as follows. [Number] Here, yi is the effect size of the i-th study, Xi is the covariate / predictor variable associated with the i-th study, β is the regression coefficient of the model, ui is the random effect of the i-th study, representing the heterogeneity between studies, and εi is the error term of the model.

[0014] (2.3) Calculate the weight value of the influence of each feature on the result using the following formula. [Number] Q m is the weight value for measuring the influence of each feature on the result. j represents calculating Qm of the j-th feature, and k represents the number of data containing the feature in the database. [Number] is the weight of the j-th feature factor, and x .j is the magnitude of the decomposition rate of the j-th feature factor.

[0015] Furthermore, step (3) specifically is 1) The initial state is as follows. Some data values in the original database constructed using the Meta-analysis method are empty. The j-th feature data of the i-th data in the database is recorded as A ij and 2) Form a random forest architecture, organize the data according to the types (A, C, SP, SB, M, E) of the feature data of different collection targets in step (1.3), there is different feature data for each feature category, and for each feature category, build an independent decision tree model for each feature among them. At the same time, randomly extract features from different feature categories

Number

[0016] Furthermore, in step (4), use the random sampling method to divide the complete data set into a modeling set and a verification set.

[0017] Furthermore, the RF-Meta model is optimized by 5-fold cross-validation.

[0018] Furthermore, in the optimization of the hyperparameters, GridSearchCV is used to perform a multi-core-based fast grid search to find the optimal hyperparameters.

[0019] A prediction device for the decomposition cycle of an agricultural multi-film based on an RF-Meta model, comprising a processor and a program stored in a memory and executable on the processor, wherein when the processor executes the executable program, a method for predicting the decomposition cycle of an agricultural multi-film based on the RF-Meta model is realized.

[0020] A storage medium containing a computer-executable program, wherein when the computer-executable program is executed by a computer processor, a method for predicting the decomposition cycle of an agricultural multi-film based on the RF-Meta model is executed.

Advantages of the Invention

[0021] The beneficial effects of the present invention are as follows. 1. Accurately estimate the decomposition cycle. There are many parameters that need to be considered for the decomposition of multi-films, such as temperature, humidity, soil type, and microbial activity. The model can examine the complex interactions of these parameters and achieve an accurate estimate of the decomposition cycle. 2. Improve the generality of the model results. By collecting experimental data from major multi-film usage regions around the world, the constructed model is used in a global BDM decomposition scenario. 3. Accelerate decomposition. By using the prediction model, farmers can better understand the relationship between important parameters such as soil moisture and total soil nitrogen and the decomposition rate of the multifilm. As a result, in order to promote the decomposition of the multifilm, they can take measures to adjust these parameters, thereby promoting environmental protection and reducing plastic residues on the land. 4. Scientifically select the multifilm. The prediction model can also provide farmers with a scientific basis for selecting the type of multifilm that is most suitable for local soil and climate conditions. This helps to improve the effectiveness and benefits of the multifilm, reduce waste and resource consumption, and contribute to ensuring the healthy growth of crops. 5. Reduce labor costs. After the decomposition cycle of the multifilm is determined, farmers no longer need a large number of personnel to clean up the fragments of the multifilm on the farmland. Farmers can plan the farmland more flexibly, plan the replacement and crop rotation of the multifilm more appropriately, and achieve sustainable and environmentally friendly agricultural production.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Embodiments for Carrying Out the Invention

[0023] Hereinafter, in conjunction with the accompanying drawings, the technical solution of the present invention will be further described.

[0024] As shown in Figure 1, the model construction flowchart of the present invention is The step of collecting the original dataset of the decomposition of the biodegradable multifilm through the Meta-analysis method, (1) Obtaining a complete database by filling in missing values in the original dataset using a random forest model; (2) Splitting the complete database into a modeling set and a validation set, and constructing an initial RF-Meta model; (3) Training the initial RF-Meta model and optimizing hyperparameters to obtain an optimized RF-Meta model; (4) Optimizing the hyperparameters of the optimized random forest prediction model to obtain an optimal RF-Meta model; (5) Verifying the accuracy of the random forest prediction model.

Example

[0025] 1. Obtain the BDM decomposition dataset. Through the PRISMA workflow of the Meta-analysis method, collect and integrate the original data for constructing the model. Use the website WebofScience to search for peer-reviewed publications that investigated the impact of climate, materials, and soil environment on the degradation of biodegradable films over a 10-year period, and the China National Knowledge Infrastructure database. To maintain quality control, define the criteria for data collection. (1) The experimental conditions should approximate a natural degradation environment, and no additional strains should be added to the soil environment for cultivation. The decomposition temperature of the soil (excluding compost) should not exceed 35°C. (2) It is necessary to report the main components of the BDM materials used in the experiment, and the ratio of PLA to PBAT should be greater than 30%. (3) Repeatedly conduct the experimental data to obtain information on the degraded soil during the experiment.

[0026] The variables collected from these documents include the title of the document, the publication year, the location of the study, meteorological variables including cumulative rainfall, average temperature, and solar radiation during the decomposition cycle, soil properties including initial soil moisture content, total soil nitrogen content, soil organic matter, pH, clay content, silt content, sand content, bulk density, and soil coarse particle content, material variables including PLA content, PBAT content, material thickness, and material surface area, decomposition time, burial depth, and Chao1 index and Shannon index in soil ecological variables. Also considered were whether the BDM had experienced outdoor exposure exceeding 7 days before decomposition and whether it had been tilled during the decomposition process.

[0027] Some missing values were supplemented through known databases, and the database sources were the 40-year product of the global land reanalysis in China (CRA / Land)-monthly product (LandProduct) and the SoilGrids database. The calculation formula for the decomposition rate of BDM is as follows.

Equation

Equation

[0028] The funnel plot drawn according to the Egger regression test results of the model is shown in Figure 3. The left and right sides are symmetric, and the data points are distributed at the bottom of the funnel, indicating that there is no obvious bias in the data and no obvious asymmetry or deviation from the straight line, showing the high reliability of the results. The value of the fail-safe number calculated using the Rosenthal method is 87778, which means that about 87778 non-significant studies are required to offset the currently observed statistical significance, indicating that the data is clearly significant.

[0029] 2. Use the rma.mv function in the metafor package of the R language to fit a multivariate random effects model (meta-regression). The collected factors are sorted according to the heterogeneity QM caused by each factor. The finally remaining factors with QM>1 are cumulative rainfall, average temperature, initial soil moisture content, total soil nitrogen content, soil organic matter, pH, clay content, silt content, sand content, bulk density, soil coarse particle content, PLA content, PBAT content, burial depth, Chao1 index, Shannon index, and decomposition time.

[0030] 3. Fill the data in the database. Since the data source is relatively complex and 8% of the data is null values, before formally starting the modeling, a small amount of null value data in the database is filled using the same method as the final modeling without changing the final result. Adopt an intelligent step-by-step method. The essence of this method is to traverse all features and start filling the feature with the least missing data to ensure that the changes to the features of the original data are minimized. The filling process is as follows. (1) The initial state is as follows. The database contains multiple features, and some of them may have missing values. Before starting the filling process, first select the feature with the least missing values and fill the missing values of other features with 0. (2) Perform regression prediction. Use known data to execute random forest regression prediction, estimate possible values of missing features through the data relationships of other complete data, and the values estimated in this way are consistent with the features of the entire data. These predicted values are returned to the original feature matrix to gradually complete the data. (3) Fill in gradually. After successfully filling in one feature, continue to select the next feature with the fewest number of missing values, and repeat the above steps until all features are traversed. Finally, based on the existing information, all null values in the database were filled with highly reliable values.

[0031] 4. Construct a modeling set and an independent verification set. Use all the collected data sets above as a sample set, where the decomposition rate is the fitting target and the rest are feature data. Divide the sample set into a modeling set and an independent verification set. By random sampling, introduce 70% of the data in the data set into the modeling set and the remaining 30% of the data into the independent verification set. The total data set of this embodiment contains 732 data samples.

[0032] 5. Train the initial random forest prediction model. Use the BDM decomposition rate data and many selected variables in the modeling set as training data to train the initial random forest prediction model, and optimize the parameters of the initial random forest prediction model by 5-fold cross-validation to obtain the random forest prediction model.

[0033] The random forest prediction model provided by the present invention is a data mining algorithm that generates n datasets (usually, n is set to 500) of the same size as the training sample set by the bootstrap method (resampling) and constructs n decision trees. The environmental variables are randomly divided into multiple environmental variable subsets, and in each decision tree, the nodes are branched by randomly dividing the multiple environmental variable subsets. The final prediction result of the model is the average value of the prediction results of all decision trees.

[0034] Optimize the hyperparameters. Optimize the hyperparameters in the model to obtain the optimal fitting effect. Use GridSearchCV to perform grid high-speed search by multi-core to find the optimal hyperparameters. The finally obtained hyperparameter values are n_estimators = 186, random_state = 42, max_depth = 23.

[0035] 6. Verify the random forest prediction model. For the constructed random forest prediction model, predict an independent validation set, compare the true decomposition rate in the independent validation set, and evaluate the prediction accuracy of the random forest conversion function. In this embodiment, the prediction accuracy of the independent validation set is evaluated by the coefficient of determination (R2). The evaluation results are shown in Figure 2. As can be seen from Figure 2, the R2 of the independent validation set is 0.965, indicating a good prediction effect. In addition, further verify the root mean square error (RMSE) of the data, and its value is 6.332, which means that the average prediction error of the model is 6.3%. Considering the complexity of the agricultural scenario, the constructed model can predict the decomposition rate of BDM very well and is considered to have practical prospects.

[0036] Meta-analysis is a statistical method that integrates the results of multiple independent studies to obtain more powerful and comprehensive conclusions. Through meta-analysis, the mode, effect size, and reliability can be identified, contributing to decision-makers making more reliable conclusions. It can clarify the general rules of the studies, simultaneously consider the weight and deviation of each study, and improve the inclusiveness and reliability of the data.

[0037] Random forest is a machine learning method based on a tree structure, which can appropriately handle non-linear relationships and shows good robustness for most regression problems. Furthermore, regression forest can more appropriately avoid overfitting of the model, so it is suitable for building models with a small amount of data. Therefore, it is expected to build a decomposition prediction model of BDM by combining meta-analysis and random forest model.

[0038] The model proposed by the present invention can first provide a specific numerical prediction of the decomposition degree in the decomposition process of BDM, has wide applicability, and can be used for BDM prediction in different soil, climate, and material environments. Also, by screening the original information through meta-analysis, it is possible to prevent the dataset from giving incorrect information to the model. In a new scenario, only by providing the initial physical and chemical information of the decomposition scenario, the future decomposition ability can be predicted. Compared with the conventional test method of obtaining the BDM decomposition characteristics of local farmland through field experiments, it does not consume a lot of human and material resources, does not require a long experimental period, and provides a new idea for predicting the BDM decomposition cycle.

Claims

1. A method for predicting the degradation cycle of an agricultural multi-film based on an RF-Meta model, comprising: (1) collecting an original dataset of the degradation of a biodegradable multi-film through Meta-analysis, constructing an original database based on the original dataset, and verifying the reliability of the original data through Meta-analysis means; (2)Through Meta-analysis, step Q of ordering the influence of different factors on the degree of decomposition in the original data set M and (3)According to the Meta-analysis results, select factors that satisfy the Q M setting value range and leave them in the original database, and use the random forest algorithm to fill in the missing values in the original database to obtain a complete database; ... (4) splitting the complete database into a modeling set and a verification set, and constructing an initial RF-Meta model; (5) training the initial RF-Meta model, optimizing hyperparameters, and obtaining an optimized RF-Meta model; (6) introducing the physical and chemical data of an initial experiment on the degradation of a target biodegradable multi-film into the optimized RF-Meta model to obtain a prediction curve of the complete degradation cycle. A method for predicting the degradation cycle of an agricultural multi-film based on an RF-Meta model, characterized by including the above steps.

2. In the step (1), the specific steps for constructing the original database are as follows: (1.1) According to the PICOS principle, the steps for creating a search strategy, inclusion criteria, and exclusion criteria are: P: Determine that the research object is a biodegradable multi-film in the scene of agricultural production; I: Determine the intervention means, select the range of basic material components constituting the degradable multi-film, and screen different degradation means, degradation environments, and pre-treatments before degradation; C: Determine the characteristics of the multi-film before and after degradation as a control means; O: Determine that the direction of the result is the degradation rate parameter of the multi-film; S: Limit the research design to laboratory experiments or field experiments, and exclude reviews and opinion articles. (1.3) Feature data set G to be collected for each article 1 The composition of G 1 = {A 1 , A 2 ..., A a ; C 1 , C 2 ,..., C e ; SP 1 , SP 2 ,..., SP p ; SB 1 , SB 2 ,..., SB b ; M 1 , M 2 ,..., M m ; E 1 , E 2 ..., E e}, where A a represents the a-th item of the basic information of the article to be collected, C e represents the e-th item of the climate conditions of the collection target, SP p represents the p-th item of the physical conditions of the soil of the collection target, SB b represents the b-th item of the biological and chemical conditions of the soil of the collection target, M m represents the m-th item of the characteristics of the multifilm material of the collection target, E e represents the e-th item of the experimental conditions of the collection target, (1.4) Decomposition result dataset G to be collected for each item 2 The composition of G 2 = {E m , E v , C m , C v}, where E m represents the average value of the results of the experimental group, and E v represents the variance of the results of the experimental group. C m represents the average value of the results of the control group, and C v represents the variance of the results of the control group. The formulas for the average value and standard deviation are as follows 【Number 12】 is the average value of the decomposition rate, s is the standard deviation of the decomposition rate, and x i is the decomposition rate of the i-th group, and i belongs to 1, 2, 3,..., n. The method for predicting the decomposition cycle of an agricultural multi-film based on the RF-Meta model according to claim 1, characterized in that. (1.2) Select an article database, set search terms based on the principle of (1.1), manually further screen articles that meet the research conditions and are repeatable,

3. By observing the relationship between the magnitude of the (1.5) decomposition rate and its variance, it is detected whether there is a publication bias in the collected articles, that is, whether there is an obvious tendency in the data source, and the decomposition rates of the biodegradable multifilms obtained from the experimental data of each group are ordered according to the magnitude, and the position number u of each data i is obtained, the data of each group are ordered according to the magnitude of the experimental variance, and then the position number v j is obtained, and it is a step of calculating the occurrence probability P of the hypothesis represented by the correlation coefficients D and D between the magnitude of the decomposition rate and its experimental variance 【Number 13】 Here, n is the total amount of data in the database, and C n represents the number of pairs that match in each pair comparison, and D n represents the number of pairs that do not match in each pair comparison. When two variables u i and v j increase or decrease simultaneously, they are matching pairs. When one of the two variables u i and v j increases and the other variable decreases, they are non-matching pairs. Cdf(D) is the cumulative distribution function value of the D value in the standard normal distribution. By referring to the D value and the P value, it is possible to determine whether there is publication bias. When the absolute value of the D value is greater than 1.96 and the p value is less than 0.05, it is considered that there is an obvious trend in the data, and it is necessary to exclude the data source with an obvious trend. When there is no obvious bias, proceed to the next step, and If publication bias is found, first, using the R language, draw a funnel plot with the decomposition rate as the X-axis data and the experimental variance as the Y-axis data. The position vector of each data in the chart is d i = (x i , s i ) and 【Number 14】 With the axis as the axis, the vector d of values t =(x t , s t ), if there is no data near the corresponding point on the opposite side of the axis and the funnel chart is asymmetric, further determine whether the data vector matches the experimental design, whether the experimental environment meets the requirements, and whether the paper is published in a journal with an IF>6. If all of the above conditions are met, the vector 【Number 15】 The step of verifying the reliability of the original data through Meta-analysis means is: adding to the database to eliminate bias, if there is a special experimental environment in the experiment corresponding to the database vector or the reliability is low, deleting the data from the database, sequentially screening the data that causes the asymmetry of the funnel chart, and then repeating (1.6) to re-check the publication bias. Evaluate whether there are general rules in the data in the database (1.7), and the significance level of the data trend is calculated as follows: 【Number 16】 Significance level F of data trend n When n is greater than 0, it is considered that there is a general rule of significance in the data in the database. When Fn < 0, it includes the step of returning to step (1.1) and the step of needing to re-add intervention means to the selection of articles. The method for predicting the decomposition cycle of an agricultural multifilm based on the RF-Meta model according to claim 2, characterized in that.

4. Specifically, step (2) is (2.1) Scan the decomposition rate data and variance data to check whether there are missing values or null values. If a null value is found, record the position information of the null value including the row m and column index n where the null value is located, and it is necessary to fill it using the nearest neighbor linear interpolation method. (2.2) Construct a multivariate random effect model in Meta-analysis for the database. The basic random effect part of the model is generated by different literature sources extracted from the data. The formula of the multivariate random effect model is as follows: 【Number 17】 Here, yi is the effect size of the i-th study, Xi is the covariate / predictor variable associated with the i-th study, β is the regression coefficient of the model, ui is the random effect of the i-th study, representing the heterogeneity between studies, and εi is the error term of the model. (2.3) Calculate the weight value of the influence of each feature quantity on the result using the following formula: 【No. 18】 Q m is a weight value for measuring the influence of each feature on the result, j represents calculating Qm of the j-th feature, and k represents the number of data including the feature in the database. 【Number 19】 is the weight of the j-th characteristic factor, x .j A method for predicting the decomposition cycle of an agricultural multi-film based on the RF-Meta model according to claim 1, characterized in that .j is the magnitude of the decomposition rate of the j-th characteristic factor.

5. Specifically, the said step (3) is 1) The initial state is as follows. Some data values in the original database constructed using the Meta-analysis method are empty, and the j-th feature data of the i-th data in the database is recorded as A ij and 2) Form a random forest architecture, organize the data according to different types of feature data of different collection targets in step (1.3), and there are different feature data for each feature category. For each feature category, construct an independent decision tree model for each feature. At the same time, randomly extract s features from different feature categories and construct d independent decision tree models to simulate the interaction of features in different feature categories. Finally, construct a total of d + 6 decision tree models and integrate them into a random forest model. 3) Extract data relationship features. Each decision tree learns the data relationship under a specific feature category from the data by judging and classifying each feature j. The said data relationship features include the value range of the feature, the relevance between features, and the data schema. 4) Gradually fill, traverse all missing values in the database, and record the number of all missing values under different feature j columns as θ j and construct a matrix [θ 1 , θ 2 ,... θ j . Check whether there are values less than or equal to 0. If so, delete the data. Next, use the bubble sort method to sort the θ values in the matrix in ascending order, fill all the data under the j feature in ascending order according to the sorted result, and finally, after traversing all the features in the matrix [θ 1 , θ 2 ,... θ j , re-traverse the data in the database to ensure that there are no missing values 5) Perform regression prediction. Based on the experimental group data with complete feature values in the original database, use the random forest algorithm and the d + 6 decision trees constructed in 2) to calculate the values judged by the d + 6 decision trees, obtain the average value, and fill the missing values in order. A method for predicting the decomposition cycle of an agricultural multi-film based on the RF-Meta model according to claim 2, characterized in that.

6. The method for predicting the degradation cycle of an agricultural multi-film based on the RF-Meta model according to claim 1, characterized in that in the step (4), a complete data set is divided into a modeling set and a verification set by using a random sampling method.

7. The method for predicting the degradation cycle of an agricultural multi-film based on the RF-Meta model according to claim 1, characterized in that the RF-Meta model is optimized by 5-fold cross-validation.

8. The method for predicting the degradation cycle of an agricultural multi-film based on the RF-Meta model according to claim 1, characterized in that in the optimization of the hyperparameters, a high-speed grid search by multi-core is performed using GridSearchCV to find the optimal hyperparameters.

9. A device for predicting the degradation cycle of an agricultural multi-film based on the RF-Meta model, comprising a processor and a program stored in a memory and executable on the processor, wherein when the processor executes the executable program, the method according to any one of claims 1 to 8 is realized.

10. A storage medium containing a computer-executable program, characterized in that when the computer-executable program is executed by a processor of a computer, the method according to any one of claims 1 to 8 is executed.

Citation Information

Patent Citations

  • Prediction method, device and equipment for residual quantity of mulching film and storage medium

    CN114821371A

  • Method for predicting service life of biodegradable mulching film

    CN115711845A

  • Feedback type edema diagnosis and treatment integrated system based on random forest

    CN116616742A

  • Birth defect risk factor assessment method based on causal learning

    CN116978556A

  • Method for determining habitat descriptor values providing a target biodegradability of a polymer

    WO2023156606A1