Water quality detection method and system based on multiple water quality spectrum parameters
By analyzing the component similarity and aggregation characteristics of historical water quality detection records, a training data set was constructed, and a water quality spectral analysis model was established, which solved the accuracy and efficiency problems of multi-parameter water quality detection, and achieved accurate detection of complex water bodies.
Patent Information
- Application Number
- CN202510518999.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-01
AI Technical Summary
It is difficult for existing water quality detection methods to accurately detect the content of multiple components at the same time, especially in natural water bodies with diverse composition types and complex concentration distribution. The detection accuracy and applicability are limited, and historical water quality detection data have not fully played a role.
By obtaining historical water quality detection records, performing component similarity analysis and aggregation feature analysis, building a training data group, and using multiple regression analysis, neural network or support vector machine to establish a water quality spectral analysis model to achieve accurate detection of multi-parameter components.
It improves the accuracy and efficiency of water quality detection, can effectively explore the potential characteristics of historical data, build representative training data sets, capture the complex relationship between the spectrum and the multi-parameter component content, and improve the accuracy of model prediction results.
Smart Images

Figure CN120404626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water quality detection, and particularly to a water quality detection method and system based on multi-parameters of water quality spectra. Background Art
[0002] With the acceleration of the industrialization process and the increase of human activities, water quality monitoring has become an important link in environmental protection and water resource management. Traditional water quality detection methods, such as chemical titration method, chromatographic analysis method, etc., although they can provide relatively accurate component content data, their operations are complex, time-consuming, and require professional equipment and reagents, making it difficult to meet the needs of large-scale and real-time monitoring. In recent years, spectral analysis technology has been widely used in the field of water quality detection due to its non-destructive and fast characteristics. However, existing spectral-based water quality detection methods mostly focus on the analysis of single parameters, such as turbidity or the concentration of a certain specific pollutant, and it is difficult to detect the content of multiple components in water bodies simultaneously. Especially in natural water bodies with diverse component types and complex concentration distributions, the detection accuracy and applicability are limited.
[0003] In addition, historical water quality detection data, as a valuable resource, contains the spatio-temporal distribution laws of water body component content. However, in existing technologies, it is often only used for simple statistical analysis and fails to fully play its role in model construction. How to effectively mine the potential features of historical data and construct a representative training data set is one of the key challenges in improving the performance of water quality detection models. At the same time, water quality spectral data is affected by the synergistic effects of multiple components and exhibits non-linear characteristics. Traditional modeling methods such as linear regression are difficult to accurately capture the complex relationships between spectra and multi-parameter component content, resulting in large errors in model prediction results. Summary of the Invention
[0004] The purpose of the present invention is to provide a water quality detection method and system that can accurately detect water bodies with complex components.
[0005] The present invention discloses a water quality detection method based on multi-parameters of water quality spectra, including:
[0006] Step S100: Obtain historical water quality detection records, determine a number of historical detection data groups. The historical detection data groups include a number of component types and the component content of each component type. Based on component similarity analysis, classify the historical detection data groups by similarity to obtain a number of historical detection data sets;
[0007] Step S200: Conduct aggregation feature analysis on the component content in each historical detection data set, determine the aggregation degree of each component type in different component content sections, and based on the performance of the aggregation degree of each component type, determine the density feature of the training data group within the preset range of the corresponding component content section of the component type;
[0008] Step S300: Based on the performance of the density features of the training data groups corresponding to each historical detection data set, determine the training data groups to be constructed. The training data groups include several training component types and the component content of each training component type.
[0009] Step S400: Based on the training data groups, construct training detection experiments, collect the water quality spectral data of each training detection experiment, and establish the association relationship between the training data groups and the water quality spectral data to construct a water quality spectral analysis model.
[0010] Step S500: Use the water quality spectral analysis model to detect the water quality of the water body.
[0011] In some embodiments disclosed by the present invention, the method for similarity classification of historical detection data sets based on component similarity analysis includes:
[0012] Step S101: Take the equivalence of component types as the first classification condition, and perform initial classification on the historical detection data sets to obtain several initial detection data sets.
[0013] Step S102: For the component types of the historical detection data groups in each initial detection data set, construct a component content mapping diagram. The horizontal axis of the component content mapping diagram includes several component type nodes, and the vertical axis is the component content. Among them, the method for determining the unit content vertical axis height ratio of the component content corresponding to each component type node includes:
[0014] Step S1021: Determine the highest value of the component content recorded for each component type in the initial detection data set, and map the highest component content values of different component types in the component content mapping diagram to make their vertical heights flush with each other, and at this time, the unit content vertical axis height ratio of the component content of each component type.
[0015] Step S103: Based on the component content of each component type in the historical detection data group and the unit content vertical axis height ratio corresponding to the component type, calculate the vertical axis height of each component type, and map the vertical axis height to the component content mapping diagram with component content nodes.
[0016] Step S104: Connect each component content node in the component content mapping diagram to obtain a component content broken line.
[0017] Step S105: Compare the component content broken lines of different content mapping diagrams to determine the difference parameter between them. If the difference parameter is less than or equal to the preset value, classify the corresponding historical detection data groups into one category.
[0018] In some embodiments disclosed by the present invention, the difference parameter of the component content broken lines of different content mapping diagrams is the cross-sectional area between the component content broken lines, or the accumulated value of the difference height between the two. Among them, the method for determining the accumulated value of the difference height between the component content broken lines includes setting a number of height detection points at a preset interval for the component content broken line points, determining the difference height between the corresponding component content broken lines for each height detection point, and calculating the sum of all the difference heights to obtain the accumulated value of the difference height.
[0019] In some embodiments disclosed by the present invention, the method for performing aggregation feature analysis on the component content in each historical detection dataset includes:
[0020] Step S201: For each component type, construct a component content reference line, map the component content of each component type in the historical detection dataset to the component content reference line, and analyze the aggregation features of the component content mapping points on the component content reference line, including:
[0021] Step S2011: Determine the number of mapping points of the component content mapping points per unit length on the component reference line, denoted as the average mapping point density;
[0022] Step S2012: Set a mapping point density analysis interval, gradually move the mapping point density analysis interval on the component reference line, and determine the real-time mapping point density after each movement of the mapping point density analysis interval;
[0023] Step S2013: Calculate the mapping point density difference between each real-time mapping point density and the average mapping point density, and mark the position nodes of the density analysis intervals corresponding to the mapping point density differences greater than or equal to the preset value, denoted as high-density position nodes;
[0024] Step S2014: Correlate the high-density position nodes with each other whose mutual distances are less than or equal to the preset value, and identify the component content section mapped by the mutually correlated high-density position nodes as an aggregation section;
[0025] Step S2015: Determine the number of mapping points of the component content mapping points in each aggregation section, and determine the section mapping point density corresponding to the aggregation section. Identify the aggregation section and the section mapping point density corresponding to each component type as the aggregation features of the component type, and identify the section mapping point density as the aggregation degree of the aggregation section.
[0026] In some embodiments disclosed by the present invention, the method for determining the density feature of the training data group within the preset range of the corresponding component content node of the component type includes:
[0027] Step S202: Determine the section lengths of several aggregation sections corresponding to each component type, and based on the section lengths, determine the section lengths of the extended sections extending outward from the aggregation sections, and determine the preset mapping point density interval to which the section mapping point density of the aggregation section belongs, determine the training data set density of the aggregation section, and based on the preset section length ratio interval to which the section length ratio of the extended section to the aggregation section belongs, determine the training data set density of the extended section.
[0028] In some embodiments disclosed by the present invention, the method for determining the training data set to be constructed includes:
[0029] Step S301: Determine the training data set density of each component type in different component content sections in each historical detection data set, and randomly select several training component contents within the component content section based on the training data set density;
[0030] Step S302: Randomly combine the training component contents of different component types to form a training data set.
[0031] In some embodiments disclosed by the present invention, the method for constructing the required training data set further includes:
[0032] Step S303: Transform the component content of each component type in each historical detection data set, and the transformed component content is less than or equal to the preset value compared with the original component content, and randomly combine the newly transformed several component contents to obtain several training data sets.
[0033] In some embodiments disclosed by the present invention, the method for constructing a water quality spectral analysis model includes:
[0034] Step S401: According to the component type and its corresponding component content in the training data set, several water samples are obtained, and the control conditions for the training detection experiment are set, including spectral detection environment parameters;
[0035] Step S402: Perform spectral detection on each water sample to obtain spectral data corresponding to the water sample, and after preprocessing the spectral data, pair it with the training data set corresponding to the water sample to form a training data set.
[0036] Step S403: Use multiple regression analysis, neural network or support vector machine to perform model training on the training data set to obtain a water quality spectral analysis model.
[0037] In some embodiments disclosed by the present invention, a water quality detection system based on water quality spectral multi-parameters is further disclosed, including:
[0038] The first module is used to obtain historical water quality detection records, determine a number of historical detection data groups. The historical detection data groups include a number of component types and the component content of each component type. Based on component similarity analysis, the historical detection data groups are classified by similarity to obtain a number of historical detection data sets;
[0039] The second module is used to perform aggregation feature analysis on the component content in each historical detection data set, determine the aggregation degree of each component type in different component content sections, and based on the performance of the aggregation degree of each component type, determine the density feature of the training data group within the preset range of the corresponding component content section;
[0040] The third module is used to determine the training data group to be constructed based on the performance of the density feature of the training data group corresponding to each historical detection data set. The training data group includes a number of training component types and the component content of each training component type;
[0041] The fourth module is used to construct a training detection experiment based on the training data group, collect the water quality spectral data of each training detection experiment, and establish an association relationship between the training data group and the water quality spectral data to construct a water quality spectral analysis model;
[0042] The fifth module is used to detect the water quality of the water body using the water quality spectral analysis model.
[0043] The present invention relates to a water quality detection method based on multi-parameters of water quality spectrum, including the following steps: First, obtain historical water quality detection records, perform component similarity analysis, and classify to obtain a number of historical detection data sets; perform aggregation feature analysis on the component content of each historical detection data set, determine the aggregation degree of the component type in different component content sections, and accordingly construct the density feature of the training data group; then, based on the density feature of the training data group, determine the training data group to be constructed, including multiple training component types and their component content; construct a training detection experiment, collect water quality spectral data and establish its association with the training data group, and finally construct a water quality spectral analysis model; use this model to detect the water quality of the water body; the above technical solution of the present invention effectively improves the accuracy and efficiency of water quality detection.
[0044] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0045] Figure 1 It is a method step diagram of the water quality detection method based on multi-parameters of water quality spectrum disclosed in the embodiment of the present invention. Detailed Embodiments
[0046] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0047] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and should not be construed as limiting the protection scope of the present invention. Skilled technicians in this field can make some non-essential improvements and adjustments based on the content of the present invention below. In the present invention, unless otherwise clearly specified and limited, the technical terms used in the present invention should have the ordinary meaning understood by those skilled in the art of the present invention.
[0048] Embodiment:
[0049] The present invention discloses a water quality detection method based on water quality spectral multi-parameters. Refer to Figure 1 , including:
[0050] Step S100, obtaining historical water quality detection records, determining a number of historical detection data groups, where the historical detection data groups include a number of component types and the component content of each component type, and based on component similarity analysis, classifying the historical detection data groups by similarity to obtain a number of historical detection data sets.
[0051] The core principle of step S100 is to systematically obtain and process historical water quality detection records, extract data groups containing various component types and their contents, and classify these data groups into several historical detection data sets by using component similarity analysis. This process is the basis of the whole method, aiming to identify sample groups with similar pollution characteristics from a large amount of historical data and provide structured input data for subsequent aggregation feature analysis and model construction. Specifically, first, historical water quality detection data need to be obtained from a database or monitoring records. These data usually include, but are not limited to, components such as heavy metals (such as copper, lead, cadmium), organic substances, nutrient salts (such as nitrogen, phosphorus), etc., and their corresponding concentration values. Then, similarity classification is carried out through the refinement steps of claim 2 (S101 to S105): taking the equivalence of component types as the first classification condition (S101), initially grouping the data containing the same component types; then constructing a component content mapping graph for each initial data set (S102), where the horizontal axis represents different component type nodes and the vertical axis represents the component content. The highest content of each component in the initial data set is determined through step S1021, and these highest values are adjusted to the same vertical height in the mapping graph, so as to calculate the vertical height ratio of the unit content of each component; subsequently, the vertical height of each component is calculated according to the actual content and this ratio and mapped as a node (S103), and these nodes are connected to form a component content broken line (S104). Finally, the similarity is judged by comparing the difference parameters between different broken lines. If the difference is less than or equal to the preset threshold, the corresponding data groups are classified into the same category (S105). For example, suppose there are two groups of data: group A contains 5 mg / L of copper (Cu) and 2 mg / L of lead (Pb), and group B contains 4.8 mg / L of copper and 2.1 mg / L of lead. The broken lines of their mapping graphs are close (the accumulated difference height value is less than 0.5), then they can be classified into the same data set. This method ensures that the classification results reflect the true law of component distribution by quantifying similarity and provides high-quality input for subsequent analysis.
[0052] Step S200: Conduct an aggregation feature analysis on the component contents in each historical detection data set, determine the aggregation degree of each component type in different component content sections, and based on the performance of the aggregation degree of each component type, determine the density feature of the training data group within the preset range of the corresponding component content section.
[0053] Step S200 explores the distribution characteristics of each component type in different content sections through aggregated feature analysis. The aggregated feature reflects the concentration degree of component content within certain specific ranges, which is closely related to the natural laws of water quality or the characteristics of pollution sources. For example, certain pollutants may frequently appear within a specific concentration range, forming an "aggregation". By calculating the aggregation degree (such as density or distribution concentration) of each component type and determining the density feature of the training data set accordingly, the principle of this process lies in identifying the key patterns of data distribution. This pattern not only reveals the statistical characteristics of water quality parameters but also provides a basis for constructing training data subsequently, ensuring that the training data can cover the typical water quality change range.
[0054] The principle of step S200 lies in deeply statistically analyzing and extracting features from the component content in each historical detection data set, determining the aggregation degree of each component type within different concentration sections, and generating the density feature of the training data set accordingly. This process reveals the concentrated areas of key pollutants in water quality by identifying the distribution patterns of component content, providing targeted data support for model training. First, a reference line of component content is constructed for each component, and all content data points in the historical data set are mapped onto this line (S201), and then the distribution characteristics of these mapped points are analyzed. Step S2011 calculates the average mapped point density on the reference line per unit length as a benchmark; step S2012 monitors the mapped point density of the interval in real time by sliding a density analysis interval of a fixed length; step S2013 compares the difference between the real-time density and the average density, and if the difference is greater than the preset value, marks this position as a high-density position node; step S2014 associates adjacent high-density nodes and defines them as an aggregation section, and calculates the section mapped point density of this section as the aggregation degree.
[0055] Step S300 determines the training data set to be constructed based on the performance of the density feature of the training data set corresponding to each historical detection data set. The training data set includes several training component types and the component content of each training component type.
[0056] Step S300 determines the training data set to be constructed based on the aforementioned density feature. Its principle lies in guiding the selection of training samples through the characteristics of data distribution. The training data set includes component types and their contents, and these data need to reflect the main laws and variabilities in historical data. The density feature, as an indicator to measure the importance of data, can help screen out the most representative component content sections, thereby generating training samples. This method avoids the sample bias that may be caused by blind random sampling. In principle, it is similar to stratified sampling in statistics, ensuring the coverage and pertinence of training data and providing high-quality input for subsequent model training.
[0057] Step S300 determines the training data set based on the density feature. The principle is to optimize the selection of training samples through distribution characteristics to ensure their representativeness. For example, if step S200 finds that the density of nitrogen content in the range of 5-10 mg / L is high, more training samples are selected in this range, while sampling is reduced in the sparse range above 20 mg / L. This strategy uses data density as an importance indicator to avoid random selection that causes samples to deviate from the actual water quality distribution, similar to selecting survey subjects from densely populated areas to reflect overall characteristics. In this way, the training data set can better capture the main patterns in historical data and provide a reliable basis for model training.
[0058] Step S400: construct a training test experiment based on the training data set, collect water quality spectrum data of each training test experiment, establish a correlation between the training data set and the water quality spectrum data, and construct a water quality spectrum analysis model.
[0059] Step S400 establishes an association between training data and spectral features by constructing a training detection experiment and collecting water quality spectral data. The principle is based on the correlation between spectral analysis and water quality parameters. Water quality spectral data (such as absorption spectrum or fluorescence spectrum) can reflect the physical and chemical properties of components in water bodies, and different component contents will cause specific changes in the spectrum. By collecting spectral data through experiments and pairing it with training data sets, this process utilizes the sensitivity and multi-parameter detection capabilities of spectral technology. The essence of building a water quality spectral analysis model is to mine the mapping relationship between spectral signals and component contents through data-driven methods (such as regression or machine learning), thereby achieving quantitative prediction of water quality.
[0060] Step S400 collects water quality spectral data and constructs an analysis model through training and detection experiments. The principle is based on the physical and chemical correlation between spectral signals and water quality components. For example, assuming that an increase in the copper content in a water sample will lead to an enhancement of the absorption peak at 620nm, by experimentally recording the spectral changes of different copper contents, a mapping relationship between the two can be established. After collecting spectral data and pairing it with the training data set, a mathematical model (such as a neural network) is used to fit this relationship to construct a water quality spectral analysis model. This is similar to training a face recognition system by taking multiple photos to record facial features, which ultimately enables the model to identify the water quality "identity" from the spectral "face".
[0061] Step S500: Detecting the water quality of the water body using a water quality spectrum analysis model.
[0062] Step S500 uses a water quality spectral analysis model to detect the actual water body, the principle of which lies in the prediction ability of the model and the efficiency of spectral technology. For example, assuming that the model has learned to recognize the relationship between the ammonia nitrogen content and the absorption intensity of the spectrum in the 200-250 nm range, when the spectral data of a certain river water sample is input, the model can quickly calculate the ammonia nitrogen concentration (such as 5.2 mg / L). This process depends on the generalization of the rules mined in the training stage to new data.
[0063] In some embodiments disclosed by the present invention, the method for classifying historical detection data groups based on component similarity analysis includes:
[0064] Step S101 takes the equivalence of component types as the first classification condition to initially classify the historical detection data group, obtaining several initial detection data sets.
[0065] Step S101 takes the equivalence of component types as the first classification condition to initially classify the historical detection data group. The principle is to simplify the complexity of data grouping by using the sameness of component types. For example, assuming that the historical data contains multiple water sample records, some include nitrogen, phosphorus, and lead, and some only include nitrogen and phosphorus. By checking whether the component types are the same, the data can be initially divided into initial detection data sets such as the "nitrogen-phosphorus-lead" group and the "nitrogen-phosphorus" group.
[0066] Step S102 constructs a component content mapping diagram for the component types of the historical detection data group in each initial detection data set. The horizontal axis of the component content mapping diagram includes several component type nodes, and the vertical axis is the component content. Among them, the method for determining the unit content vertical axis height ratio corresponding to each component type node includes:
[0067] Step S1021 determines the maximum value of the component content recorded for each component type in the initial detection data set, and maps the maximum component content values of different component types in the component content mapping diagram to make their vertical heights flush with each other, and at this time, the unit content vertical axis height ratio of the component content of each component type.
[0068] Step S102 constructs a component content mapping diagram for each initial detection data set. The principle is to quantify the distribution characteristics of component content through visualization and standardization means. The horizontal axis of the mapping diagram lists the component types (such as nitrogen, phosphorus), and the vertical axis represents the content. Step S1021 determines the unit content vertical axis height ratio. For example, assuming that the maximum content of nitrogen is 10 mg / L and that of phosphorus is 5 mg / L, in order to make a unified comparison, set their maximum values to the same height (such as 10 units) in the diagram. Then, the unit height ratio of nitrogen is 1 mg / L corresponding to 1 unit, and that of phosphorus is 1 mg / L corresponding to 2 units.
[0069] Step S103: Calculate the vertical axis height for each component type based on the component content of each component type in the historical detection data set and the vertical axis height ratio per unit content corresponding to the component type, and map the vertical axis height to the component content mapping graph at the component content nodes.
[0070] In step S103, the vertical axis height of each component is calculated according to the vertical axis height ratio per unit content and mapped to the graph. The principle is to convert the actual content into a standardized graphical representation. For example, if the nitrogen content in a water sample is 5 mg / L and the unit height ratio is 1, the vertical axis height is 5 units; if the phosphorus content is 2 mg / L and the unit height ratio is 2, the height is 4 units. These heights are marked as nodes on the mapping graph, reflecting the relative content sizes of the components.
[0071] Step S104: Connect each component content node in the component content mapping graph to obtain a component content broken line.
[0072] In step S104, the component content nodes are connected into a broken line. The principle is to visually display the change trend of the component content in the form of a broken line. For example, in the "nitrogen - phosphorus - lead" data set, the nodes of a water sample may be nitrogen (5 units), phosphorus (4 units), and lead (3 units). After connection, a broken line is formed. This broken line is like a "water quality fingerprint", reflecting the content distribution pattern of the water sample among different components. The connection order of the nodes is usually preset according to the component type (such as from nitrogen to lead).
[0073] Step S105: Compare the component content broken lines of different content mapping graphs to determine the difference parameter between the two. If the difference parameter is less than or equal to the preset value, classify the corresponding historical detection data set into one category.
[0074] In step S105, classification is performed by comparing the broken lines of different mapping graphs and calculating the difference parameter. The principle is to use the shape similarity of the broken lines to judge the proximity of the water quality patterns. For example, assume the broken line of water sample A is nitrogen (5 units), phosphorus (4 units), and lead (3 units), and that of water sample B is nitrogen (5.2 units), phosphorus (3.8 units), and lead (3.1 units). Calculate the difference parameter between the two (such as the cumulative area difference or height difference between the broken lines). If it is less than the preset threshold (such as 0.5), they are considered to belong to the same category. This method is similar to comparing the contour similarities of two faces, ensuring the accuracy of classification by quantifying the differences, and finally classifying similar water quality patterns into one category.
[0075] In some embodiments disclosed by the present invention, the difference parameter of the component content broken lines of different content mapping diagrams is the cross-sectional area between the component content broken lines, or the accumulated value of the difference height between the two. Among them, the method for determining the accumulated value of the difference height between the component content broken lines includes setting a number of height detection points at a preset interval for the component content broken line points, determining the difference height between the corresponding component content broken lines for each height detection point, and calculating the sum of all the difference heights to obtain the accumulated value of the difference height.
[0076] In some embodiments disclosed by the present invention, the method for performing aggregation feature analysis on the component content in each historical detection dataset includes:
[0077] Step S201, constructing a component content reference line for each component type, mapping the component content of each component type in the historical detection dataset to the component content reference line, and analyzing the aggregation features of the component content mapping points on the component content reference line, including:
[0078] Step S2011, determining the number of mapping points per unit length of the component content mapping points on the component reference line, denoted as the average mapping point density;
[0079] Step S2012, setting a mapping point density analysis interval, gradually moving the mapping point density analysis interval on the component reference line, and determining the real-time mapping point density after each movement of the mapping point density analysis interval;
[0080] Step S2013, calculating the mapping point density difference between each real-time mapping point density and the average mapping point density, and marking the position nodes of the density analysis intervals corresponding to the mapping point density differences greater than or equal to the preset value, denoted as high-density position nodes;
[0081] Step S2014, correlating the high-density position nodes with a mutual distance less than or equal to the preset value, and identifying the component content sections mapped by the mutually correlated high-density position nodes as aggregation sections;
[0082] Step S2015, determining the number of mapping points of the component content mapping points in each aggregation section, and determining the section mapping point density corresponding to the aggregation section. Identifying the aggregation sections and the section mapping point density corresponding to each component type as the aggregation features of the component type, and identifying the section mapping point density as the aggregation degree of the aggregation section.
[0083] In some embodiments disclosed by the present invention, the method for determining the density feature of the training data group within the preset range of the component content node corresponding to the component type includes:
[0084] Step S202: Determine the section lengths of several aggregation sections corresponding to each component type. Based on the section lengths, determine the section lengths of the extended sections extending outward from the aggregation sections, and determine the preset mapping point density interval to which the section mapping point density of the aggregation section belongs, determine the density of the training data group of the aggregation section, and determine the density of the training data group of the extended section based on the preset section length ratio interval to which the section length ratio of the extended section to the aggregation section belongs.
[0085] Step S202 aims to determine the density characteristics of the training data group within the preset range of the corresponding component content nodes for the component type. Its principle is to provide refined guidance for training data generation by analyzing the distribution characteristics of the aggregation section and the density requirements of its extended area. Specifically, first determine the length of the aggregation section. For example, assume that the aggregation section of lead (Pb) is 0 - 2 mg / L (length 2 mg / L, from the analysis of S2014), and its section mapping point density is 15 points / mg / L. Based on this length, calculate the length of the extended section extending outward (such as using the expression ($L_{\text{extended}} = k\cdot L_{\text{aggregation}}+C$), where ($k < 1$)). Assume ($k = 0.5$), ($C = 0.1$), then ($L_{\text{extended}} = 0.5\cdot 2+0.1 = 1.1$) mg / L, and the extended section may be 2 - 3.1 mg / L. Then, determine that the density of the aggregation section (15 points / mg / L) belongs to the preset density interval (such as 10 - 20 points / mg / L), and accordingly set the density of the training data group (such as taking the median value 15 points / mg / L of the interval). For the extended section, according to the ratio of its length to the length of the aggregation section ($1.1\div2 = 0.55$) falling into the preset ratio interval (such as 0.5 - 1.0), combined with the ratio characteristics, determine a lower density (such as 10 points / mg / L) to reflect its sparsity.
[0086] Among them, the expression for calculating the section length of the extended section is:
[0087] D2 = D1×ln(L×D1 + b1)
[0089] Among them, D2 is the section length of the extended section, D1 is the section length of the aggregation section, L is the aggregation section length influence adjustment coefficient, b1 is the aggregation section length influence adjustment constant, and L×D1 + b1 is less than 1.
[0090] In some embodiments disclosed by the present invention, the method for determining the training data group to be constructed includes:
[0091] Step S301: Determine the density of the training data groups for each component type in different component content ranges in each historical detection dataset, and randomly select several training component contents within the component content range based on the density of the training data groups.
[0092] Step S302: Randomly combine the training component contents of different component types to form training data groups.
[0093] In some embodiments disclosed by the present invention, the method for constructing the required training data groups further includes:
[0094] Step S303: Transform the component content of each component type in each historical detection data group, and the transformed component content is less than or equal to the preset value compared with the original component content. Randomly combine the newly transformed several component contents to obtain several training data groups.
[0095] In some embodiments disclosed by the present invention, the method for constructing a water quality spectral analysis model includes:
[0096] Step S401: Prepare several water samples according to the component types and their corresponding component contents in the training data groups, and set the control conditions for the training detection experiment, including spectral detection environment parameters.
[0097] Step S401 prepares water samples according to the component types and their contents in the training data groups and sets the control conditions for spectral detection. The principle is to generate spectral signals corresponding to the training data through a controllable experiment to provide reliable inputs for model establishment. The training data groups (such as from S300) contain component types (such as nitrogen, phosphorus) and specific contents (such as nitrogen 5 mg / L, phosphorus 2 mg / L). Based on this, water samples are prepared. For example, 5 mg / L nitrogen and 2 mg / L phosphorus are dissolved in distilled water in the laboratory. The control conditions include spectral detection environment parameters, such as light source wavelength (200 - 800 nm), temperature (25 °C), water sample volume (50 mL), etc. The stability of these parameters directly affects the quality of spectral data. For example, when detecting nitrogen content, the spectrometer may focus on the ultraviolet region (such as 220 nm) because nitrogen compounds have characteristic absorption peaks in this region.
[0098] Step S402: Perform spectral detection on each water sample to obtain spectral data corresponding to the water sample. After preprocessing the spectral data, pair it with the training data group corresponding to the water sample to form a training dataset.
[0099] Step S402 performs spectral detection and preprocessing data on each water sample, and pairs it with a training data set to form a training data set. The principle is to capture the physicochemical properties of water quality components through spectral technology, and to improve the reliability of model training through data cleaning. For example, a water sample containing 5 mg / L nitrogen and 2 mg / L phosphorus is scanned using an ultraviolet-visible spectrometer to obtain a spectral curve (such as an absorption peak intensity of 0.8 at 220 nm, indicating the presence of nitrogen). Preprocessing may include noise removal (such as smoothing filtering to remove instrument jitter), baseline correction (such as deducting blank water sample signals), and normalization (such as standardizing the absorption value to a range of 0-1) to eliminate environmental interference. Afterwards, the processed spectral data (eigenvectors, such as absorption values at multiple wavelengths) are paired with a training data set (nitrogen 5 mg / L, phosphorus 2 mg / L) to form a training sample.
[0100] Step S403: Perform model training on the training data set using multiple regression analysis, neural network or support vector machine to obtain a water quality spectrum analysis model.
[0101] Step S403 utilizes multiple regression analysis, neural network or support vector machine to carry out model training to training data set, constructs water quality spectrum analysis model, its principle is to mine the complex mapping relationship between spectral data and component content by mathematical method.For example, assuming that the training data set contains 100 groups of water sample spectra and their corresponding contents (such as nitrogen 0-10mg / L, phosphorus 0-5mg / L), multiple regression analysis can fit linear relationship (such as absorption value=a·nitrogen content+b·phosphorus content+c), which is suitable for simple scenes; neural network (such as multi-layer perceptron) captures nonlinear features through hidden layer and adapts to complex water quality changes; support vector machine optimizes classification or regression boundary by kernel function (such as radial basis function) to improve prediction accuracy.After training, the model can predict content from new spectral input, for example, output nitrogen 4.8mg / L and phosphorus 1.9mg / L after inputting a certain spectrum.
[0102] In some embodiments disclosed in the present invention, a water quality detection system based on multi-parameter water quality spectrum is also disclosed, including:
[0103] The first module is used to obtain historical water quality test records and determine several historical test data groups. The historical test data groups include several component types and the component content of each component type. Based on component similarity analysis, the historical test data groups are similarly classified to obtain several historical test data sets.
[0104] The second module is used to perform clustering feature analysis on the component content in each historical detection data set, determine the degree of clustering of each component type in different component content segments, and based on the clustering degree performance of each component type, determine the training data group density characteristics of the component type within the preset range of the corresponding component content segment.
[0105] A third module is configured to determine a training data group to be constructed based on the performance of the density features of the training data group corresponding to each historical detection data set. The training data group includes several training component types and the component content of each training component type.
[0106] A fourth module is configured to construct a training detection experiment based on the training data group, collect the water quality spectral data of each training detection experiment, establish an association relationship between the training data group and the water quality spectral data, and construct a water quality spectral analysis model.
[0107] A fifth module is configured to perform water quality detection on the water body by using the water quality spectral analysis model.
[0108] The present invention relates to a water quality detection method based on water quality spectral multi-parameters, including the following steps: First, obtain historical water quality detection records, perform component similarity analysis, and classify to obtain several historical detection data sets; perform aggregation feature analysis on the component content of each historical detection data set to determine the aggregation degree of the component type in different component content sections, and construct the density feature of the training data group accordingly; then, based on the density feature of the training data group, determine the training data group to be constructed, including multiple training component types and their component content; construct a training detection experiment, collect water quality spectral data and establish its association with the training data group, and finally construct a water quality spectral analysis model; use this model to perform water quality detection on the water body; the above technical solution of the present invention effectively improves the accuracy and efficiency of water quality detection.
[0109] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments.
[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present invention.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solution of the present invention, and these modifications or equivalent replacements cannot make the modified technical solution deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A water quality detection method based on multi-parameters of water quality spectra, characterized in that, Including: Step S100: Obtain historical water quality detection records, determine a number of historical detection data groups. Each historical detection data group includes a number of component types and the component content of each component type. Based on component similarity analysis, perform similarity classification on the historical detection data groups to obtain a number of historical detection data sets. Step S200: Perform aggregation feature analysis on the component content in each historical detection data set, determine the aggregation degree of each component type in different component content sections, and based on the performance of the aggregation degree of each component type, determine the density feature of the training data group within the preset range corresponding to the component content section of the component type. Step S300: Based on the performance of the density feature of the training data group corresponding to each historical detection data set, determine the training data group to be constructed. The training data group includes a number of training component types and the component content of each training component type. Step S400: Based on the training data group, construct a training detection experiment, collect the water quality spectral data of each training detection experiment, and establish an association relationship between the training data group and the water quality spectral data to construct a water quality spectral analysis model. Step S500: Use the water quality spectral analysis model to detect the water quality of the water body.
2. The water quality detection method based on water quality spectral multi-parameters according to claim 1, characterized in that The method for performing similarity classification on the historical detection data groups based on component similarity analysis includes: Step S101: Take the equivalence of component types as the first classification condition, perform initial classification on the historical detection data groups to obtain a number of initial detection data sets. Step S102: For each historical detection data group in the initial detection data set, construct a component content mapping diagram. The horizontal axis of the component content mapping diagram includes a number of component type nodes, and the vertical axis is the component content. Among them, the method for determining the vertical height ratio of the unit content corresponding to each component type node includes: Step S1021: Determine the highest value of the component content recorded for each component type in the initial detection data set, map the highest component content values of different component types in the component content mapping diagram so that their vertical heights are flush with each other, and at this time, determine the vertical height ratio of the unit content of the component content of each component type. Step S103: Based on the component content of each component type in the historical detection data group and the vertical height ratio of the unit content corresponding to the component type, calculate the vertical height of each component type and map the vertical height to the component content mapping diagram with the component content node. Step S104: Connect each component content node in the component content mapping diagram to obtain a component content broken line. Step S105: Compare the component content broken lines of different content mapping diagrams, determine the difference parameter between them. If the difference parameter is less than or equal to the preset value, classify the corresponding historical detection data groups into one category.
3. The water quality detection method based on water quality spectral multi-parameters according to claim 1, characterized in that The difference parameter of the component content broken lines in the mapping diagrams with different contents is the cross-sectional area between the component content broken lines, or the accumulated value of the difference height between the two. Among them, the method for determining the accumulated value of the difference height between the component content broken lines includes setting a number of height detection points at preset intervals for the points on the component content broken lines, determining the difference height between the component content broken lines corresponding to each height detection point, and calculating the sum of all the difference heights to obtain the accumulated value of the difference height.
4. The water quality detection method based on multi-parameters of water quality spectrum according to claim 1, wherein The method for performing clustering feature analysis on the component contents in each historical detection dataset includes: Step S201: Construct a component content reference line for each component type, map the component contents of each component type in the historical detection dataset to the component content reference line, and analyze the clustering features of the component content mapping points on the component content reference line, including: Step S2011: Determine the number of mapping points of the component content mapping points per unit length on the component reference line, denoted as the average mapping point density; Step S2012: Set a mapping point density analysis interval, gradually move the mapping point density analysis interval on the component reference line, and determine the real-time mapping point density after each movement of the mapping point density analysis interval; Step S2013: Calculate the mapping point density difference between each real-time mapping point density and the average mapping point density, and mark the position nodes of the density analysis intervals corresponding to the mapping point density differences greater than or equal to the preset value, denoted as high-density position nodes; Step S2014: Correlate the high-density position nodes with each other whose mutual distances are less than or equal to the preset value, and identify the component content section mapped by the mutually correlated high-density position nodes as the clustering section; Step S2015: Determine the number of mapping points of the component content mapping points in each clustering section, and determine the section mapping point density corresponding to the clustering section. Identify the clustering section and the section mapping point density corresponding to each component type as the clustering features of the component type, and identify the section mapping point density as the clustering degree of the clustering section.
5. The water quality detection method based on water quality spectral multi-parameters according to claim 4, characterized in that, The method for determining the density feature of the training data group within the preset range of the corresponding component content nodes of the component type includes: Step S202: Determine the section lengths of several clustering sections corresponding to each component type, based on the section lengths, determine the section lengths of the extended sections extended outward by the clustering sections, judge the preset mapping point density interval to which the section mapping point density of the clustering section belongs, determine the training data group density of the clustering section, and based on the preset section length ratio interval to which the section length ratio of the extended section to the clustering section belongs, determine the training data group density of the extended section.
6. The water quality detection method based on water quality spectral multi-parameters according to claim 1, wherein, The method for determining the training data group to be constructed includes: Step S301: Determine the training data group density of each component type in different component content sections in each historical detection dataset, and randomly select several training component contents within the component content section based on the training data group density; Step S302: Randomly combine the training component contents of different component types to form a training data group.
7. The water quality detection method based on water quality spectral multi-parameters according to claim 6, characterized in that, The method for constructing the required training data group also includes: Step S303: Transform the component content of each component type in each historical detection data group, and the transformed component content is less than or equal to a preset value compared with the original component content. Randomly combine the newly transformed component contents to obtain several training data groups.
8. The water quality detection method based on water quality spectral multi-parameters according to claim 1, characterized in that, The method for constructing a water quality spectral analysis model includes: Step S401: According to the component types and their corresponding component contents in the training data groups, several water samples are obtained, and the control conditions for the training detection experiment are set, including spectral detection environment parameters; Step S402: Perform spectral detection on each water sample to obtain spectral data corresponding to the water sample. After preprocessing the spectral data, pair it with the corresponding training data group of the water sample to form a training data set; Step S403: Use multiple regression analysis, neural network or support vector machine to train the training data set to obtain a water quality spectral analysis model.
9. A water quality detection system based on multi-parameters of water quality spectra, characterized in that, For performing the water quality detection method described in any one of claims 1-7, it includes: The first module is used to obtain historical water quality detection records, determine several historical detection data groups. The historical detection data group includes several component types and the component content of each component type. Based on component similarity analysis, classify the historical detection data groups by similarity to obtain several historical detection data sets; The second module is used to perform aggregated feature analysis on the component content in each historical detection data set, determine the aggregation degree of each component type in different component content sections, and based on the performance of the aggregation degree of each component type, determine the density feature of the training data group within the preset range of the corresponding component content section; The third module is used to determine the training data groups to be constructed based on the performance of the density feature of the training data group corresponding to each historical detection data set. The training data group includes several training component types and the component content of each training component type; The fourth module is used to construct a training detection experiment based on the training data group, collect the water quality spectral data of each training detection experiment, and establish the association relationship between the training data group and the water quality spectral data to construct a water quality spectral analysis model; The fifth module is used to use the water quality spectral analysis model to detect the water quality of the water body.