Auxiliary chlorophyll content prediction method and system for few-shot water body
By collecting and pre-treating water body data, calculating the average impact value and performing dimensionality reduction treatment, and combining with weak regression models for training, the problems of high modeling costs and insufficient data in monitoring or prediction of chlorophyll content in the prior art are solved, and a chlorophyll prediction model with low cost and high generalization performance is realized.
Patent Information
- Application Number
- PCT/CN2024/103586
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-07-04
- Publication Date
- 2025-06-26
AI Technical Summary
The prior art has problems such as high modeling costs, neglecting the correlation of environmental impact factors, and difficulty in modeling in the monitoring or prediction of chlorophyll content.
A chlorophyll content auxiliary prediction method for small sample water bodies is adopted. By collecting water body information in different environments, data preprocessing is performed, the average impact value is calculated, and the dimensionality reduction is performed. The mixed training set and test set are used to train with weak regression models, the sample weight is updated, and the regression model with the best prediction effect is finally selected for prediction.
The modeling process requires the quantity and quality of data, reduces the modeling cost, and establishes a low-cost and high generalization performance chlorophyll prediction model, which is suitable for the prediction of chlorophyll content in small sample water bodies.
Smart Images

Figure CN2024103586_26062025_PF_FP_ABST
Abstract
Description
A chlorophyll content auxiliary prediction method and system for small sample water bodies Technical Field
[0001] The present invention relates to the field of biological environmental science and technology, and in particular to a method and system for auxiliary prediction of chlorophyll content in a small sample water body. Background Art
[0002] Measuring chlorophyll content typically involves using traditional fluorescence detection techniques to measure parameters such as fluorescence intensity to calculate chlorophyll content, or hyperspectral imaging methods. These methods often require specific physical and chemical parameters and are costly and complex to install. Therefore, it is crucial to study the impact of factors in water on chlorophyll content, develop relevant prediction models, and promote their application in carbon emission control.
[0003] Influencing factors in water bodies include water temperature, pH, dissolved oxygen, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus, and total nitrogen. Existing technologies establish data-driven chlorophyll prediction models for each of these units, which places high demands on data quantity and quality. When the collected data is small or severely contaminated, it is difficult to establish an accurate model. Furthermore, the environmental influencing factors vary from water body to water body, making it difficult to use a universal model to meet the chlorophyll prediction needs of every water body unit. Therefore, modeling each water body unit individually requires repeated data collection and processing, model training, and other tasks, resulting in high modeling costs.
[0004] In summary, the existing chlorophyll content monitoring or prediction methods have the problems of high modeling cost, neglect of the correlation of environmental image factors, and difficulty in modeling when there is insufficient data.
[0005] Summary of the Invention
[0006] To this end, an embodiment of the present invention provides an auxiliary prediction method and system for chlorophyll content in small sample water bodies, which is used to solve the problems in the existing technology of chlorophyll content monitoring or prediction methods, such as high modeling cost, neglect of the correlation of environmental image factors, and difficulty in modeling when there is insufficient data.
[0007] To solve the above problems, an embodiment of the present invention provides a method for auxiliary prediction of chlorophyll content in a small sample of water, the method comprising:
[0008] S1: Collect water body information in different environments, obtain relevant data of the water body, and construct a data set, wherein the data set includes relevant data of small sample water bodies and relevant data of auxiliary sample water bodies, and the small sample water bodies and the auxiliary sample water bodies come from different environments;
[0009] S2: preprocessing the relevant data to obtain an ideal data set;
[0010] S3: Based on the ideal data set, calculating the average influence value and performing dimensionality reduction processing on the ideal data set;
[0011] S4: According to the dataset after dimensionality reduction, set the mixed training set and test set, and set the weight and sample update coefficient for each sample;
[0012] S5: Combine the reduced-dimensional dataset and the weight vector to train the basic weak regression model and continuously update the sample weights.
[0013] S6: Repeat S4-S5 continuously, save and record the weak regression model each time, and finally select the regression model with the best prediction effect as the strong regression model;
[0014] S7: Input the test sample into the strong regression model to obtain the chlorophyll prediction value of the test sample.
[0015] Preferably, the method for collecting water body information under different environments and obtaining water body related data is:
[0016] Sensors are used to collect water information in different environments and obtain water-related data, including water temperature, pH value, dissolved oxygen content, potassium permanganate content, ammonia nitrogen content, total phosphorus content, total nitrogen content, conductivity, turbidity, and chlorophyll A content at the previous moment.
[0017] Preferably, the preprocessing of the relevant data to obtain an ideal data set specifically includes:
[0018] S21: Remove non-numerical samples from the dataset;
[0019] S22: The chlorophyll A content in each sample in the data set was evaluated using a 3σ standard to remove outliers in the sample;
[0020] S23: The Z-score method is used to standardize the data to eliminate the dimensional effects between different variables. The standardization formula is:
[0021] Where x i is the i-th variable in the original sample x, is x i The average value, σ i is x i The standard deviation of is x i The standardized value of .
[0022] Preferably, the calculating of the average influence value based on the ideal data set and performing dimensionality reduction processing on the ideal data set specifically includes:
[0023] S31: Use the original data in the ideal data set to train a BP neural network g D (x), the ideal data set is in the form of T(T={(x i ,y i )}, where x i is the training sample, y i For test samples;
[0024] S32: Increase the value of one of the independent variables by 10% in turn, and keep the other variables unchanged. After obtaining the new training sample, input it into the BP neural network for simulation to obtain the result
[0025] Where, is the value of the jth variable in the i-th sample, t represents the number of samples in the ideal training set, and k represents the number of variables in each sample;
[0026] S33: Reduce the value of one of the independent variables by 10% in turn, and keep the other variables unchanged. After obtaining a new training sample, input it into the BP neural network for simulation to obtain the result
[0027] S34: Calculate the difference of the results obtained for each independent variable, average them according to the number of training samples, and calculate the average influence value of the j-th independent variable on the network:
[0028] Where, Get the difference of the results for each independent variable, MIV j is the average influence of the j-th independent variable on the network;
[0029] S35: Determine the absolute value of the average impact value based on the calculation result of S34, screen out the independent variables with cumulative contribution values greater than 90%, and complete the selection of input variables.
[0030] Preferably, the setting of a mixed training set and a test set, and setting a weight and a sample update coefficient for each sample, specifically includes:
[0031] S41: Set the initial weight of each sample in the mixed training set T
[0032] Where n is the number of samples from the auxiliary sample set in the mixed training set, and m is the number of samples from the small sample water body in the mixed training set;
[0033] S42: Define the sample weight update coefficient β for the auxiliary sample set:
[0034] Preferably, the step of combining the reduced-dimensional dataset and the weight vector to train the basic weak regression model and continuously updating the sample weights specifically includes:
[0035] S51: Set sample weight distribution
[0036] Where, is the sample weight of the tth iteration, n is the number of samples from the auxiliary sample set in the mixed training set, and m is the number of samples from the small sample water body in the mixed training set;
[0037] S52: Multiply the sample weight distribution with the sample, and train a weak regressor f based on the multiplied data set t (x);
[0038] S53: Calculate function error
[0039] in
[0040] Where x i is the training sample, y i For test samples;
[0041] S54: Calculate the average error e t :
[0042] S55: Define the update coefficient β of the target sample set t :
[0043] β t =e t / (1-e t );
[0044] S56: Update sample weights:
[0045] S57: Based on the new sample weights, a new BP neural network is obtained.
[0046] Preferably, the root mean square error is used to evaluate the performance of the model, where the root mean square error formula is:
[0047] Where RMSE is the root mean square error, y i is the true value of chlorophyll A in the test set, is the predicted value of chlorophyll A in the test set.
[0048] The embodiment of the present invention further provides a system for auxiliary prediction of chlorophyll content in a small sample of water. The system is used to implement the above-mentioned auxiliary prediction method for chlorophyll content in a small sample of water, specifically comprising:
[0049] A data acquisition module collects water body information in different environments, obtains relevant data of the water body, and constructs a data set, wherein the data set includes relevant data of a small sample water body and relevant data of an auxiliary sample water body, and the small sample water body and the auxiliary sample water body come from different environments;
[0050] A data preprocessing module, used to preprocess the relevant data to obtain an ideal data set;
[0051] An average influence value calculation module, configured to calculate an average influence value based on the ideal data set and perform dimensionality reduction processing on the ideal data set;
[0052] The sample weight setting and sample coefficient update module is used to set the mixed training set and test set according to the data set after dimensionality reduction, and set the weight and sample update coefficient for each sample;
[0053] The sample weight update module is used to combine the reduced-dimensional dataset and the weight vector to train the basic weak regression model and continuously update the sample weights;
[0054] A prediction model selection module is used to continuously repeat the sample weight setting and sample coefficient updating module and the sample weight updating module, save and record each weak regression model, and finally select the regression model with the best prediction effect as the strong regression model;
[0055] The prediction module is used to input the test sample into the strong regression model to obtain the chlorophyll prediction value of the test sample.
[0056] An embodiment of the present invention also provides an electronic device, which includes a processor, a memory and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the above-mentioned auxiliary prediction method for chlorophyll content in small sample water bodies.
[0057] An embodiment of the present invention also provides a computer storage medium, which stores a computer software product. The computer software product includes several instructions for enabling a computer device to execute the above-mentioned auxiliary prediction method for chlorophyll content in small sample water bodies.
[0058] It can be seen from the above technical solutions that the present invention has the following advantages:
[0059] An embodiment of the present invention provides an auxiliary prediction method and system for chlorophyll content in small sample water bodies. The present invention fully utilizes the knowledge information contained in the data of adjacent purified units, reduces the requirements for data quantity and quality in the modeling process, reduces the large amount of data collection, processing and model training work in the modeling process, and establishes a low-cost and high-generalization chlorophyll prediction model for the chlorophyll prediction process. Based on the specific conditions in small sample water bodies, the present invention establishes a more effective prediction model with the assistance of large sample water bodies. Different from commonly used measurement methods such as photobiological reaction devices, the present invention uses data-driven as the main method, collects various variable information in the water body, and applies machine learning and transfer learning methods to predict chlorophyll content, providing assistance for water quality monitoring and environmental governance. This method is suitable for biological monitoring processes with known process requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the implementation cases of the present invention or the technical solutions in the prior art, the following is a brief description of the drawings required for use in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be understood as limiting the present invention in any way. Those skilled in the art can derive other drawings based on these drawings without inventive effort. Among them:
[0061] FIG1 is a flow chart of a method for auxiliary prediction of chlorophyll content in a small sample of water provided in an embodiment;
[0062] FIG2 is a brief flow chart of the method of the present invention in an embodiment;
[0063] FIG3 is a schematic diagram of the average impact values of various influencing variables in the embodiment;
[0064] FIG4 is a schematic diagram showing a comparison of RMSE values between the method of the present invention and other algorithms in an embodiment;
[0065] FIG5 is a block diagram of a chlorophyll content auxiliary prediction system for a small sample water body provided in an embodiment. DETAILED DESCRIPTION
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0067] Example 1
[0068] As shown in FIG1 and FIG2 , an embodiment of the present invention proposes an auxiliary prediction method for chlorophyll content in a small sample of water, the method comprising:
[0069] S1: Collect water body information in different environments, obtain relevant data of the water body, and construct a data set, wherein the data set includes relevant data of small sample water bodies and relevant data of auxiliary sample water bodies, and the small sample water bodies and the auxiliary sample water bodies come from different environments;
[0070] S2: preprocessing the relevant data to obtain an ideal data set;
[0071] S3: Based on the ideal data set, calculating the average influence value and performing dimensionality reduction processing on the ideal data set;
[0072] S4: According to the dataset after dimensionality reduction, set the mixed training set and test set, and set the weight and sample update coefficient for each sample;
[0073] S5: Combine the reduced-dimensional dataset and the weight vector to train the basic weak regression model and continuously update the sample weights.
[0074] S6: Repeat S4-S5 continuously, save and record the weak regression model each time, and finally select the regression model with the best prediction effect as the strong regression model;
[0075] S7: Input the test sample into the strong regression model to obtain the chlorophyll prediction value of the test sample.
[0076] From the above technical solution, it can be seen that the present invention provides an auxiliary prediction method for the chlorophyll content of a small sample water body, by obtaining water body related data; pre-processing the data to obtain the ideal data set required for modeling; calculating the average impact value, and performing dimensionality reduction operation on the ideal data set; according to the data set after dimensionality reduction, setting a mixed training set and a test set to prepare for training the neural network, and setting weights and sample update coefficients for each sample; combining the data set and the weight vector to train the basic weak regression model, and continuously updating the sample weights; repeatedly building the model, saving and recording each weak regression model, and finally selecting the regression model with the best effect as the strong regression model. The present invention utilizes the correlation between water environment influencing factors and chlorophyll, and between different water environments, and makes full use of the knowledge information contained in the adjacent purified unit data, reduces the requirements of the modeling process on the amount and quality of data, reduces the large amount of data collection, processing and model training work in the modeling process, and establishes a chlorophyll prediction model with low cost and high generalization performance for the chlorophyll prediction process.
[0077] In this embodiment, in step S1, sensors (sensor 1 and sensor 2) are used to collect water body information in different environments, obtain relevant data of the water body, and construct a data set. The data set includes relevant data (data 1) of a small sample water body (water body A) and relevant data (data 2) of an auxiliary sample water body (water body B). The small sample water body and the auxiliary sample water body come from different environments. It should be noted that the auxiliary sample water body in this embodiment is a large sample water body used to assist in the establishment of a small water body sample to establish a prediction model.
[0078] Furthermore, the relevant data of the water body include water temperature, pH value, dissolved oxygen content, potassium permanganate content, ammonia nitrogen content, total phosphorus content, total nitrogen content, conductivity, turbidity and chlorophyll A content at the previous moment.
[0079] In this embodiment, in step S2, the relevant data are preprocessed to obtain an ideal data set, specifically including:
[0080] S21: Remove non-numerical samples from the dataset;
[0081] S22: The chlorophyll A content in each sample in the data set was evaluated using a 3σ standard to remove outliers in the sample;
[0082] S23: The Z-score method is used to standardize the data to eliminate the dimensional effects between different variables. The standardization formula is:
[0083] Where x i is the i-th variable in the original sample x, is x i The average value, σ i is x i The standard deviation of is x i The standardized value of .
[0084] In this embodiment, in step S3, the average influence value is calculated based on the ideal data set, and the dimensionality reduction process is performed on the ideal data set, specifically including:
[0085] S31: Use the original data in the ideal data set to train a BP neural network g D (x), the ideal data set is in the form of T(T={(x i ,y i )}, where x i is the training sample, y i For test samples;
[0086] S32: Increase the value of one of the independent variables by 10% in turn, and keep the other variables unchanged. After obtaining the new training sample, input it into the BP neural network for simulation to obtain the result
[0087] Where, is the value of the jth variable in the i-th sample, t represents the number of samples in the ideal training set, and k represents the number of variables in each sample;
[0088] S33: Reduce the value of one of the independent variables by 10% in turn, and keep the other variables unchanged. After obtaining a new training sample, input it into the BP neural network for simulation to obtain the result
[0089] S34: Calculate the difference of the results obtained for each independent variable, average them according to the number of training samples, and calculate the average influence value of the j-th independent variable on the network:
[0090] Where, Get the difference of the results for each independent variable, MIV j is the average influence of the j-th independent variable on the network;
[0091] S35: Determine the absolute value of the average impact value based on the calculation result of S34, screen out the independent variables with cumulative contribution values greater than 90%, and complete the selection of input variables.
[0092] In this embodiment, in step S4, a mixed training set and a test set are set according to the dataset after dimensionality reduction, and a weight and a sample update coefficient are set for each sample.
[0093] The weight and sample update coefficient are set for each sample, including:
[0094] S41: Set the initial weight of each sample in the mixed training set T
[0095] Where n is the number of samples from the auxiliary sample set in the mixed training set, and m is the number of samples from the small sample water body in the mixed training set;
[0096] S42: Define the sample weight update coefficient β for the auxiliary sample set:
[0097] In this embodiment, in step S5, the dataset after dimensionality reduction is combined with the weight vector to train the basic weak regression model, and the sample weights are continuously updated, specifically including:
[0098] S51: Set sample weight distribution
[0099] Where, is the sample weight of the tth iteration, n is the number of samples from the auxiliary sample set in the mixed training set, and m is the number of samples from the small sample water body in the mixed training set;
[0100] S52: Multiply the sample weight distribution with the sample, and train a weak regressor f based on the multiplied data set t (x);
[0101] S53: Calculate function error
[0102] in
[0103] Where x i is the training sample, y i For test samples;
[0104] S54: Calculate the average error e t :
[0105] S55: Define the update coefficient β of the target sample set t :
[0106] β t =e t / (1-e t );
[0107] S56: Update sample weights:
[0108] S57: Based on the new sample weights, a new BP neural network is obtained.
[0109] In this embodiment, in step S6, S4-S5 are repeated continuously, the weak regression model of each time is saved and recorded, and finally the regression model with the best prediction effect is selected as the strong regression model.
[0110] It should be noted that, in this embodiment, both the weak regression model and the strong regression model are BP neural network models, and the only difference is the prediction accuracy.
[0111] In this embodiment, in step S7, the test sample is input into the strong regression model to obtain the chlorophyll prediction value of the test sample.
[0112] Furthermore, the present invention uses the root mean square error to evaluate the performance of the model, wherein the root mean square error formula is:
[0113] Where RMSE is the root mean square error, yi is the true value of chlorophyll A in the test set, is the predicted value of chlorophyll A in the test set.
[0114] In order to further illustrate the advantages of the method of the present invention, it is described below in conjunction with specific experiments.
[0115] This experiment analyzed raw samples from the National Automatic Integrated Water Quality Monitoring Platform, generating 851 samples from the Huaihe River Basin. The first 172 samples were from Jiangba Town, and the remaining 679 samples were from Gaoliangjian Town. We calculated the average influence values of the relevant variables for this data to determine the input for the established neural network. The absolute values of the average influence values for each variable are shown in Figure 3. When constructing the mixed dataset, the 679 samples from the auxiliary sample water body, Area B, were added to the mixed dataset without altering the number of samples. We then sampled 10%-80% of the 172 samples from Area A, the small sample water body, and added them to the mixed dataset. The remaining samples from Area A served as the test set. We compared the results using a traditional BP network and a BP network without the average influence value algorithm. To clearly demonstrate the regression effectiveness of transfer learning, we compared the RMSE of the three models, as shown in Figure 4.
[0116] Example 2
[0117] As shown in FIG5 , the present invention provides a system for auxiliary prediction of chlorophyll content in a small sample of water. The system is used to implement the auxiliary prediction method for chlorophyll content in a small sample of water in the first embodiment, specifically comprising:
[0118] The data acquisition module 10 collects water body information in different environments, obtains relevant data of the water body, and constructs a data set, wherein the data set includes relevant data of a small sample water body and relevant data of an auxiliary sample water body, and the small sample water body and the auxiliary sample water body come from different environments;
[0119] A data preprocessing module 20 is used to preprocess the relevant data to obtain an ideal data set;
[0120] An average influence value calculation module 30 is used to calculate the average influence value based on the ideal data set and perform dimensionality reduction processing on the ideal data set;
[0121] The sample weight setting and sample coefficient updating module 40 is used to set a mixed training set and a test set according to the data set after dimensionality reduction, and set a weight and a sample update coefficient for each sample;
[0122] The sample weight updating module 50 is used to combine the reduced-dimensional dataset and the weight vector to train the basic weak regression model and continuously update the sample weights;
[0123] The prediction model selection module 60 is used to continuously repeat the sample weight setting and sample coefficient updating module 40 and the sample weight updating module 50, save and record each weak regression model, and finally select the regression model with the best prediction effect as the strong regression model;
[0124] The prediction module 70 is used to input the test sample into the strong regression model to obtain the chlorophyll prediction value of the test sample.
[0125] The present embodiment provides an auxiliary prediction system for chlorophyll content in small sample water bodies, which is used to implement the aforementioned auxiliary prediction method for chlorophyll content in small sample water bodies. Therefore, the specific implementation method of the auxiliary prediction system for chlorophyll content in small sample water bodies can be found in the embodiment part of the auxiliary prediction method for chlorophyll content in small sample water bodies mentioned above. For example, the data acquisition module 10, the data preprocessing module 20, the average influence value calculation module 30, the sample weight setting and sample coefficient update module 40, the sample weight update module 50, the prediction model selection module 60, and the prediction module 70 are respectively used to implement steps S1, S2, S3, S4, S5, S6, and S7 in the aforementioned auxiliary prediction method for chlorophyll content in small sample water bodies. Therefore, its specific implementation method can refer to the description of the corresponding embodiments of each part. In order to avoid redundancy, it will not be repeated here.
[0126] Example 3
[0127] An embodiment of the present invention also provides an electronic device, which includes a processor, a memory and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the above-mentioned auxiliary prediction method for chlorophyll content in small sample water bodies.
[0128] Example 4
[0129] An embodiment of the present invention also provides a computer storage medium, which stores a computer software product. The computer software product includes several instructions for enabling a computer device to execute the above-mentioned auxiliary prediction method for chlorophyll content in small sample water bodies.
[0130] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0131] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0132] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.
[0133] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A method for auxiliary prediction of chlorophyll content in small sample water bodies, characterized in that: include: S1: Collect water body information in different environments, obtain relevant data of water bodies, and construct a data set, wherein the data set includes relevant data of small sample water bodies and relevant data of auxiliary sample water bodies, and the small sample water bodies and the auxiliary sample water bodies come from different environments; S2: preprocessing the relevant data to obtain an ideal data set; S3: Based on the ideal data set, calculate the average influence value, and perform dimensionality reduction processing on the ideal data set; S4: According to the dataset after dimension reduction, set the mixed training set and test set, and set the weight and sample update coefficient for each sample; S5: Combine the reduced-dimensional dataset and the weight vector to train the basic weak regression model and continuously update the sample weights; S6: Repeat S4-S5 continuously, save and record each weak regression model, and finally select the regression model with the best prediction effect as the strong regression model; S7: Input the test sample into the strong regression model to obtain the chlorophyll prediction value of the test sample.
2. The chlorophyll content auxiliary prediction method for a small sample water body according to claim 1 is characterized in that: The method for collecting water body information under different environments and obtaining water body related data is as follows: The sensors are used to collect water information in different environments and obtain water-related data, including water temperature, pH value, dissolved oxygen content, potassium permanganate content, ammonia nitrogen content, total Phosphorus content, total nitrogen content, conductivity, turbidity and chlorophyll A content at the last moment.
3. The chlorophyll content auxiliary prediction method for a small sample water body according to claim 1 is characterized in that: The preprocessing of the relevant data to obtain an ideal data set specifically includes: S21: Remove non-numerical samples from the data set; S22: The chlorophyll A content in each sample in the data set is evaluated using the 3σ standard to remove outliers in the sample; S23: The Z-score method is used to standardize the data to eliminate the dimensional effects between different variables. The standardization formula is: In the formula, x i is the i-th variable in the original sample x, is x i The average value, σ i is x i The standard deviation of is x i The standardized value of .
4. The chlorophyll content auxiliary prediction method for a small sample water body according to claim 1 is characterized in that: The step of calculating the average influence value based on the ideal data set and performing dimensionality reduction processing on the ideal data set specifically includes: S31: Use the original data in the ideal data set to train a BP neural network g D (x), the ideal data set is in the form of T(T={(x i ,y i )}, where x i is the training sample, y i For test samples; S32: Increase the value of one of the independent variables by 10% in turn, and keep other variables unchanged. After obtaining a new training sample, input it into the BP neural network for simulation to obtain the result In the formula, is the value of the jth variable in the ith sample, and t represents the number of samples in the ideal training set. Number, k represents the number of variables in each sample; S33: Reduce the value of one of the independent variables by 10% in turn, and keep other variables unchanged. After obtaining a new training sample, input it into the BP neural network for simulation to obtain the result S34: Calculate the difference of the results obtained for each independent variable, average them according to the number of training samples, and calculate the average influence value of the j-th independent variable on the network: In the formula, Get the difference of results for each independent variable, MIV j is the average influence of the j-th independent variable on the network; S35: Determine the absolute value of the average impact value based on the calculation result of S34, screen out the independent variables with cumulative contribution value>90%, and complete the selection of input variables.
5. The chlorophyll content auxiliary prediction method for a small sample water body according to claim 1, characterized in that: The mixed training set and test set are set, and weights and sample update coefficients are set for each sample, specifically including: S41: Set the initial weight of each sample in the mixed training set T In the formula, n is the sample from the auxiliary sample set in the mixed training set, and m is the sample from the auxiliary sample set in the mixed training set. Number of samples from small sample water bodies; S42: Define the sample weight update coefficient β for the auxiliary sample set:
6. The chlorophyll content auxiliary prediction method for a small sample water body according to claim 1, characterized in that: The reduced-dimensional data set and the weight vector are combined to train the basic weak regression model, and the sample weights are continuously updated, specifically including: S51: Set sample weight distribution In the formula, is the sample weight of the tth iteration, n is the number of samples from the auxiliary sample set in the mixed training set, and m is the number of samples from the small sample water body in the mixed training set; S52: Multiply the sample weight distribution by the sample, and train a weak regressor f based on the multiplied data set t (x); S53: Calculate function error in In the formula, x i is the training sample, y i For test samples; S54: Calculate the average error e t : S55: Define the update coefficient β of the target sample set t : β t =and t / (1-e t ); S56: Update sample weights: S57: Based on the new sample weights, a new BP neural network is obtained.
7. The chlorophyll content auxiliary prediction method for a small sample water body according to claim 1, characterized in that: The root mean square error is used to evaluate the performance of the model, where the root mean square error formula is: In the formula, RMSE is the root mean square error, y i is the true value of chlorophyll A in the test set, is the predicted value of Chlorophyll A in the test set.
8. A chlorophyll content auxiliary prediction system for small sample water bodies, characterized in that: The system is used to implement the auxiliary prediction method for chlorophyll content in a small sample water body according to any one of claims 1 to 7, specifically comprising: A data collection module collects water body information under different environments, obtains relevant data of the water body, and constructs a data set, wherein the data set includes relevant data of a small sample water body and relevant data of an auxiliary sample water body, and the small sample water body and the auxiliary sample water body come from different environments; A data preprocessing module, used for preprocessing the relevant data to obtain an ideal data set; An average influence value calculation module, used to calculate the average influence value based on the ideal data set, and perform dimensionality reduction processing on the ideal data set; The sample weight setting and sample coefficient updating module is used to set the mixed training set and test set according to the reduced dimension data set, and set the weight and sample update coefficient for each sample; The sample weight update module is used to combine the reduced-dimensional data set with the weight vector, train the basic weak regression model, and continuously update the sample weights; A prediction model selection module is used to continuously repeat the sample weight setting and sample coefficient updating module and the sample weight updating module, save and record each weak regression model, and finally select the regression model with the best prediction effect as the strong regression model; The prediction module is used to input the test sample into the strong regression model to obtain the chlorophyll prediction value of the test sample.
9. An electronic device, characterized in that: The electronic device includes a processor, a memory and a bus system, the processor and the memory are connected through the bus system, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the auxiliary prediction method for chlorophyll content in small sample water bodies as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: The computer storage medium stores a computer software product, and the computer software product includes several instructions for enabling a computer device to execute the chlorophyll content auxiliary prediction method for a small sample water body as described in any one of claims 1 to 7.
Citation Information
Patent Citations
A perturbation-based chlorophyll a content related factor sensitivity analysis method
CN109934334A
Energy consumption auxiliary prediction method and system in organic silicon monomer fractionation process
CN113705908A
Chlorophyll content auxiliary prediction method and system for small sample water body
CN117763508A
Standard error of prediction of performance in artificial intelligence model
US20210390446A1
Cited By
Aquaculture environment analysis method based on image analysis
CN121030137A
SSA optimization-based sugarcane leaf disease infection degree classification method, apparatus and device, and medium
CN121147766A
Water chlorophyll prediction model generation method and device based on space-time transmission mechanism and proxy model, and electronic equipment
CN121483414A
A method and device for generating a water chlorophyll a prediction model based on a space-time transmission mechanism and an agent model, and an electronic device
CN121483414B
End point phosphorus content prediction method and system based on oxygen top-blown converter steelmaking
CN122241170A