Seasonal river water environment quality evaluation method based on chemical research service data

By using a low-field nuclear magnetic resonance online detection device and a one-dimensional convolutional neural network model, the problem of real-time detection and automatic analysis of organic pollutants in the aquatic environment has been solved, enabling rapid and accurate water quality assessment and supporting watershed water environment management.

CN122020250APending Publication Date: 2026-05-12SHANDONG KELIN TESTING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG KELIN TESTING CO LTD
Filing Date
2026-01-30
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies cannot achieve real-time detection and automated analysis of organic pollutants in the aquatic environment. Traditional offline detection has a long cycle, existing online monitoring equipment has limited detection capabilities for complex organic pollutants, and the deployment of deep learning models on edge computing devices is not yet mature.

Method used

By employing a low-field nuclear magnetic resonance (NMR) online detection device combined with a one-dimensional convolutional neural network model, and processing the NMR spectrum through Fourier transform, phase correction, and baseline correction, a seasonal water environment quality assessment method based on matter-element extensions is constructed to achieve automatic analysis and comprehensive evaluation of multiple types of organic pollutants.

Benefits of technology

It enables rapid on-site detection and automatic spectral analysis of organic pollutants in the aquatic environment, reveals the seasonal variation patterns of water quality, provides a scientific basis for watershed water environment management, and improves the real-time performance and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020250A_ABST
    Figure CN122020250A_ABST
Patent Text Reader

Abstract

The invention discloses a seasonal river water environment quality evaluation method based on chemical research service data, and the method comprises the following steps: 1, deploying a low-field nuclear magnetic resonance online detection device on each sampling section, applying a radio frequency excitation pulse to a water sample entering a detection cavity, and collecting a free induction attenuation signal; collecting the one-dimensional nuclear magnetic resonance spectrum vectors of the sampling sections according to seasonal division to form a nuclear magnetic resonance spectrum data set; 2, constructing a one-dimensional convolutional neural network model, and forming a pollutant feature matrix corresponding to each sampling section in each season; and step 3, establishing classic domain matter elements and joint domain matter elements corresponding to each water quality grade according to a water environment quality standard, and summarizing water environment quality evaluation grades of each sampling section in each season to form a seasonal river water environment quality comprehensive evaluation result. According to the invention, on-site rapid detection and automatic spectrogram analysis of organic pollutants are realized, and the problem of poor timeliness of traditional off-line detection is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water quality monitoring using deep learning technology, and particularly to a method for assessing the seasonal river water environment quality based on chemical research service data. Background Technology

[0002] Currently, water environment quality monitoring mainly relies on two technical approaches. The first involves sending water samples to chemical research service institutions for offline analysis, using instrumental analysis methods such as gas chromatography, liquid chromatography, and gas chromatography-mass spectrometry to qualitatively and quantitatively detect organic pollutants in the water. This method offers high accuracy and reliable results, but it has limitations such as long detection cycles, high costs, and the inability to reflect real-time dynamic changes in water quality. From sampling to obtaining a test report, it typically takes several days or even weeks, making it difficult to meet the needs of emergency monitoring and real-time early warning for the water environment. The second approach involves deploying online monitoring equipment at monitoring sections to continuously and automatically monitor conventional physicochemical indicators such as pH, dissolved oxygen, conductivity, and turbidity. Online monitoring has the advantages of strong real-time performance and low labor costs, but existing online monitoring technologies mainly target inorganic indicators and a few simple organic compounds, with limited detection capabilities for complex organic pollutants such as benzene compounds, polycyclic aromatic hydrocarbons, and organochlorides.

[0003] Nuclear magnetic resonance (NMR) spectroscopy is a crucial tool for analyzing the molecular structure of organic compounds. By detecting the resonance absorption signals of atomic nuclei in a sample under an applied magnetic field, information about the molecular structure and composition can be obtained. In recent years, advancements in permanent magnet technology and radio frequency electronics have led to the miniaturization and portability of low-field NMR equipment, creating conditions for its application in on-site monitoring. However, applying NMR to online monitoring of the water environment still faces numerous challenges. Water samples are complex, and the resonance peaks of various organic compounds overlap in NMR spectra. Traditional manual spectral analysis methods are inefficient and require extensive professional knowledge, making automated online detection difficult. The rise of deep learning technology has provided new insights for the automatic analysis of complex spectra. Convolutional neural networks (CNNs) can automatically learn feature representations from raw data, achieving significant success in image recognition, speech recognition, and other fields. While some research has explored applying CNNs to spectral data analysis, deep learning methods for analyzing NMR spectra of water samples are still immature. Existing research largely focuses on the detection of single types of pollutants, lacking comprehensive analytical models that can simultaneously identify multiple types of organic pollutants. In addition, how to deploy deep learning models on edge computing devices to enable online inference is also an engineering problem that needs to be solved. Summary of the Invention

[0004] The purpose of this invention is to provide a method for assessing the seasonal river water environment quality based on chemical research service data. This method enables rapid on-site detection and automatic spectral analysis of organic pollutants, overcomes the problem of poor timeliness in traditional offline detection, and can reveal the seasonal variation patterns of water quality, providing a scientific basis for watershed water environment management.

[0005] To address the aforementioned technical problems, this invention provides a method for assessing the seasonal river water environmental quality based on chemical research service data, comprising the following steps: Step 1, Online NMR Detection and Spectrum Data Acquisition: Low-field online NMR detection devices are deployed at each sampling section. Radio frequency excitation pulses are applied to the water samples entering the detection chamber, and free induction decay signals are collected. Fourier transform, phase correction, and baseline correction are sequentially performed on the free induction decay signals to obtain NMR spectra. The NMR spectra are discretized to obtain one-dimensional NMR spectral vectors. The one-dimensional NMR spectral vectors of each sampling section are collected according to the season to form an NMR spectral dataset. Step 2, Nuclear Magnetic Resonance Spectrum Analysis and Pollutant Feature Extraction Based on Convolutional Neural Network: Construct a one-dimensional convolutional neural network model, which includes a convolutional layer, a batch normalization layer, an activation layer, a pooling layer, a global average pooling layer, a fully connected layer, and an output layer connected in sequence. Train the one-dimensional convolutional neural network model using water sample testing data provided by the chemical research service institution. Input the one-dimensional nuclear magnetic resonance spectrum vector from the nuclear magnetic resonance spectrum dataset into the trained one-dimensional convolutional neural network model to obtain pollutant feature vectors, forming pollutant feature matrices corresponding to each sampling section and each season. Step 3: Comprehensive evaluation of seasonal water environment quality based on matter-element extension: Establish classical domain matter-element and section domain matter-element corresponding to each water quality level according to the water environment quality standards. Calculate the comprehensive correlation degree of each pollutant feature vector in the pollutant feature matrix with respect to each water quality level. Determine the water environment quality evaluation level based on the comprehensive correlation degree. Summarize the water environment quality evaluation levels of each sampling section for each season to form the comprehensive evaluation result of seasonal river water environment quality.

[0006] Furthermore, in step 1, the magnetic field strength of the low-field nuclear magnetic resonance online detection device is 0.5 Tesla, and a broadband radio frequency transmitting and receiving coil is configured; the radio frequency excitation pulse is a 90-degree radio frequency excitation pulse with a duration of 10 microseconds; the acquisition time of the free induction decay signal is 2048 milliseconds, and the sampling frequency is 10000 Hz, resulting in a free induction decay signal sequence containing 20480 sampling points.

[0007] Furthermore, in step 1, the specific process of performing Fourier transform, phase correction, and baseline correction on the free-induction attenuation signal sequence is as follows: The first 16384 sampling points of the free-induction attenuation signal sequence are truncated, and 16384 zero points are added at the end to form an extended sequence of 32768 points. A fast Fourier transform is performed on the extended sequence to obtain the frequency domain complex spectrum. The real and imaginary parts of the frequency domain complex spectrum are extracted separately. The real part is used as the abscissa, and the imaginary part is used as the ordinate to plot the phase trajectory curve. The zero-order phase is determined by rotating the phase trajectory curve so that its principal axis coincides with the real axis. To correct the phase angle, a linearly increasing phase shift is applied to each frequency point along the chemical shift axis to eliminate first-order phase distortion. The real part of the corrected frequency domain complex spectrum is extracted as the absorption mode spectrum. Two regions, -1 to 0 and 9 to 10, are selected on the chemical shift axis as baseline sampling regions. The spectral intensity values ​​corresponding to each frequency point in the baseline sampling region are extracted, and the baseline function curve is obtained by least-squares fitting using a third-order polynomial. The function value of the baseline function curve at the corresponding frequency point is subtracted from the spectral intensity value of the absorption mode spectrum at each frequency point to obtain the baseline-corrected NMR spectrum.

[0008] Furthermore, in step 1, the nuclear magnetic resonance spectrum is uniformly discretized within the chemical shift range of 0 to 9, and one spectral intensity value is taken at every chemical shift interval of 0.01 to obtain a one-dimensional nuclear magnetic resonance spectral vector containing 900 intensity data points. The one-dimensional nuclear magnetic resonance spectral vector is then compiled into a nuclear magnetic resonance spectral dataset according to the division of spring (March to May), summer (June to August), autumn (September to November), and winter (December to February of the following year).

[0009] Furthermore, in step 2, the structure of the one-dimensional convolutional neural network model includes, in sequence: an input layer, a first convolutional layer, a first batch normalization layer, a first activation layer, a first pooling layer, a second convolutional layer, a second batch normalization layer, a second activation layer, a second pooling layer, a third convolutional layer, a third batch normalization layer, a third activation layer, a global average pooling layer, a fully connected layer, and an output layer; the input dimension of the input layer is 900.

[0010] Furthermore, the first convolutional layer contains 32 convolutional kernels with a kernel length of 15, a stride of 1, and 7 padding elements; the first pooling layer has a pooling window length of 4 and a pooling stride of 4; the second convolutional layer contains 64 convolutional kernels with a kernel length of 11, a stride of 1, and 5 padding elements; the second pooling layer has a pooling window length of 5 and a pooling stride of 5; and the third convolutional layer contains 128 convolutional kernels with a kernel length of 7, a stride of 1, and 3 padding elements.

[0011] Furthermore, the first, second, and third normalization layers perform the following normalization processing on the input sequence: calculate the mean and variance of all elements in the input sequence, subtract the mean from each element of the input sequence, divide by the square root of the variance, add 0.00001, multiply the sum by a learnable scaling parameter, and add a learnable offset parameter; the first, second, and third activation layers perform modified linear unit activation operations, keeping elements greater than zero unchanged and setting elements less than or equal to zero to zero; the first and second pooling layers perform max pooling operations, taking the maximum value within each pooling window as the output.

[0012] Furthermore, the global average pooling layer calculates the arithmetic mean of all elements in each of the 128 sequences output by the third activation layer to obtain a feature vector containing 128 elements; the fully connected layer maps the feature vector into an output vector containing 6 elements through a 128-row, 6-column connection matrix and a bias vector containing 6 elements; the 6 elements output by the output layer correspond to the concentration indicators of benzene series organic pollutants, phenolic organic pollutants, polycyclic aromatic hydrocarbon organic pollutants, organochlorine pollutants, petroleum hydrocarbon pollutants, and total dissolved organic carbon, respectively.

[0013] Furthermore, in step 3, the specific process of establishing classical domain matter elements and segmented domain matter elements is as follows: For each pollutant concentration indication value in the pollutant feature vector, determine the upper limit and lower limit of the value range of each pollutant concentration indication value under Class I, Class II, Class III, Class IV, and Class V water quality. Take the water quality level as the object of the matter element, take each pollutant concentration indication value as the feature of the matter element, and take the value range of each feature under the corresponding level as the value range of the corresponding feature, forming 5 classical domain matter elements corresponding to 5 water quality levels respectively; take the value range of each feature in the entire level range as the value range of the corresponding feature, forming segmented domain matter elements.

[0014] Furthermore, in step 3, the specific process of calculating the comprehensive correlation degree and determining the water environment quality assessment level is as follows: For each pollutant concentration indication value in the pollutant feature vector to be evaluated, calculate the classical domain distance between each pollutant concentration indication value and the corresponding feature value interval in the classical domain element, and the section domain distance between each pollutant concentration indication value and the corresponding feature value interval in the section domain element. If the classical domain distance is less than or equal to zero, the single-index correlation degree is equal to the negative number of the classical domain distance divided by half the length of the corresponding feature value interval in the classical domain element. If the classical domain distance is greater than zero, the single-index correlation degree is equal to the negative number of the quotient obtained by dividing the classical domain distance by the difference between the section domain distance and the classical distance. Sum the single-index correlation degrees of each pollutant concentration indication value and divide by the number of pollutant concentration indication values ​​to obtain the comprehensive correlation degree. Take the water quality level corresponding to the largest comprehensive correlation degree value as the water environment quality assessment level.

[0015] The seasonal river water environment quality assessment method based on chemical research service data of the present invention has the following beneficial effects: This invention employs a low-field nuclear magnetic resonance (NMR) online detection device to achieve rapid on-site detection of water samples, overcoming the shortcomings of traditional offline testing methods, such as long detection cycles and the inability to reflect water quality changes in real time. NMR technology can acquire molecular structural information of organic pollutants in water samples, providing far more information than traditional online monitoring equipment can only measure conventional physicochemical indicators, laying a data foundation for the subsequent identification of multiple types of organic pollutants. Through optimized signal acquisition parameters and preprocessing procedures, including zero-fill Fourier transform, phase correction, and polynomial baseline correction, high-quality NMR spectra can be obtained, improving the accuracy and reliability of spectral analysis.

[0016] The one-dimensional convolutional neural network model constructed in this invention can automatically learn feature patterns in NMR spectra, achieving end-to-end mapping from the original spectrum to concentration indicators of multiple pollutants, eliminating the need for tedious manual peak identification and integration operations. The network employs a multi-layered convolutional and pooling hierarchical structure, coupled with batch normalization and modified linear unit activation operations, enabling it to extract abstract features layer by layer from low to high levels, demonstrating excellent analytical capabilities for complex overlapping resonance peaks in the spectrum. Utilizing high-precision detection data provided by chemical research service institutions as training labels ensures a correspondence between the pollutant feature vectors output by the model and the actual concentrations, achieving a complementary advantage between rapid NMR detection and traditional precise detection methods.

[0017] This invention employs matter-element extension theory to establish a water quality assessment model. By constructing classical domain matter-element and segmental domain matter-element models to describe the standard range of each water quality level, and by calculating a comprehensive correlation metric to quantify the degree of conformity between the evaluated sample and each level, it effectively addresses incompatibility issues and ambiguity in level boundaries during multi-index evaluation. Evaluating each sampling section separately according to season reveals the dynamic patterns of water quality changes with the seasons, identifies the differences in water quality between high-water and low-water periods, and provides a more refined decision-making basis for watershed water environment management. The evaluation results are presented in the form of water quality levels, making them intuitive and easy to understand and use by environmental management departments and the public. Attached Figure Description

[0018] Figure 1 A schematic diagram of the free induction attenuation signal acquired during online nuclear magnetic resonance detection provided in an embodiment of the present invention; Figure 2 The nuclear magnetic resonance spectrum obtained after performing Fourier transform, phase correction, and baseline correction on the free induction decay signal according to an embodiment of the present invention; Figure 3 A schematic diagram illustrating the principle of convolution operation in a one-dimensional convolutional neural network model provided in an embodiment of the present invention; Figure 4 The graph shows the correlation function curves of various water quality grades in the matter-element extension evaluation method provided in this embodiment of the invention. Detailed Implementation

[0019] A seasonal river water quality assessment method based on chemical research service data includes the following steps: Step 1, Online NMR Detection and Spectrum Data Acquisition: Low-field online NMR detection devices are deployed at each sampling section. Radio frequency excitation pulses are applied to the water samples entering the detection chamber, and free induction decay signals are collected. Fourier transform, phase correction, and baseline correction are sequentially performed on the free induction decay signals to obtain NMR spectra. The NMR spectra are discretized to obtain one-dimensional NMR spectral vectors. The one-dimensional NMR spectral vectors of each sampling section are collected according to the season to form an NMR spectral dataset. Step 2, Nuclear Magnetic Resonance Spectrum Analysis and Pollutant Feature Extraction Based on Convolutional Neural Network: Construct a one-dimensional convolutional neural network model, which includes a convolutional layer, a batch normalization layer, an activation layer, a pooling layer, a global average pooling layer, a fully connected layer, and an output layer connected in sequence. Train the one-dimensional convolutional neural network model using water sample testing data provided by the chemical research service institution. Input the one-dimensional nuclear magnetic resonance spectrum vector from the nuclear magnetic resonance spectrum dataset into the trained one-dimensional convolutional neural network model to obtain pollutant feature vectors, forming pollutant feature matrices corresponding to each sampling section and each season. Step 3: Comprehensive evaluation of seasonal water environment quality based on matter-element extension: Establish classical domain matter-element and section domain matter-element corresponding to each water quality level according to the water environment quality standards. Calculate the comprehensive correlation degree of each pollutant feature vector in the pollutant feature matrix with respect to each water quality level. Determine the water environment quality evaluation level based on the comprehensive correlation degree. Summarize the water environment quality evaluation levels of each sampling section for each season to form the comprehensive evaluation result of seasonal river water environment quality.

[0020] The primary step in assessing the seasonal water environment quality of rivers lies in obtaining detection data that accurately reflects the composition of organic pollutants in the water. Traditional water quality testing methods rely on sending water samples to chemical research service institutions for offline analysis. This approach has limitations such as long testing cycles and the inability to reflect real-time dynamic changes in water quality. This invention employs low-field nuclear magnetic resonance (NMR) technology to achieve online detection of water samples. By analyzing the relaxation characteristics of hydrogen nuclei in the water sample, it obtains the molecular structure information of organic pollutants, providing a data foundation for subsequent intelligent spectral analysis and water quality assessment.

[0021] When deploying low-field NMR online detection devices at various sampling sections, factors such as magnetic field strength, probe configuration, and environmental adaptability need to be comprehensively considered. The choice of magnetic field strength directly affects the sensitivity and resolution of the NMR signal. Although high-field NMR equipment has higher sensitivity, its large size and high maintenance cost make it difficult to meet the needs of field online detection. This invention selects a permanent magnet with a magnetic field strength of 0.5 Tesla as the main magnetic field source. This magnetic field strength can achieve miniaturization and portability of the device while ensuring sufficient signal strength. In a magnetic field environment of 0.5 Tesla, the Larmor precession frequency of hydrogen nuclei is approximately 21.3 MHz. Radio frequency electronics systems within this frequency range are mature in design and cost-controllable, which is conducive to the mass deployment of the equipment.

[0022] The RF transmitting and receiving coils employ a broadband design, operating within a frequency range of 20 MHz to 23 MHz, and are adaptable to main magnetic field drift caused by ambient temperature variations. The coils feature a saddle-shaped structure with an inner diameter of 25 mm and a length of 40 mm. This size design accommodates a sufficient volume of water sample to obtain a detectable signal strength while ensuring the uniformity of the RF field within the sample area. The coil's quality factor has been optimized to exceed 150. A higher quality factor helps improve signal reception sensitivity; however, an excessively high quality factor can narrow the coil bandwidth, affecting broadband detection capabilities. Therefore, a balance must be struck between sensitivity and bandwidth.

[0023] Water samples are drawn from the river using a peristaltic pump and then filtered through a 0.45-micron filter membrane to remove suspended particulate matter before entering the nuclear magnetic resonance (NMR) detection chamber. This filtration process eliminates interference from suspended matter on the uniformity of the magnetic field and prevents particulate matter from depositing and clogging the chamber. The detection chamber is made of polytetrafluoroethylene (PTFE), a hydrogen-free material that does not generate background signal interference and possesses excellent chemical inertness, making it resistant to corrosion from various water samples. The chamber has an effective volume of 5 ml, and the water sample resides in the chamber for approximately 30 seconds, sufficient to complete a full NMR detection procedure.

[0024] Once the water sample enters the detection chamber and stabilizes, the online nuclear magnetic resonance (NMR) detection device applies a 90-degree radio frequency (RF) excitation pulse to the water sample. The purpose of applying the 90-degree pulse is to flip the hydrogen nucleus magnetization vector, which is in thermal equilibrium, from the longitudinal direction to the transverse plane, thereby generating a transverse magnetization component that can be detected by the coil. The flip angle of the 90-degree pulse needs to be precisely controlled; a deviation from 90 degrees will result in a decrease in the transverse magnetization component, thus reducing the signal strength. The pulse duration is set to 10 microseconds, which is determined by both the RF field strength and the required flip angle. With the hardware configuration of this invention, a 10-microsecond pulse duration can generate an RF magnetic field strength of approximately 25 microtesla, sufficient to achieve a 90-degree magnetization vector flip.

[0025] refer to Figure 1 When a water sample enters the detection chamber of the low-field nuclear magnetic resonance online detection device, the device applies a 90-degree radio frequency excitation pulse to the water sample, flipping the hydrogen nucleus magnetization vector, which is in thermal equilibrium, from the longitudinal direction to the transverse plane. After the pulse ends, the transverse magnetization vector precesses around the direction of the main magnetic field under the influence of the magnetic field. At the same time, due to the inhomogeneity of the magnetic field and intermolecular interactions, the transverse magnetization vector gradually decays. The voltage signal induced in the receiving coil during this process is the free induction decay signal. Figure 1 The horizontal axis represents the acquisition time in milliseconds, ranging from 0 to 2048 milliseconds; the vertical axis represents the signal strength, expressed in any unit. Figure 1 It can be observed that the free-induction decay signal has the largest amplitude at the beginning of the acquisition, and then exhibits oscillating decay characteristics, with the signal amplitude gradually decreasing over time and eventually approaching zero. The oscillating characteristics of the signal originate from the fact that hydrogen nuclei in different chemical environments in the water sample have different resonance frequencies, and the superposition of multiple frequency components forms a complex oscillation waveform. Figure 1The signal envelope curve is plotted as a dashed line, reflecting the overall attenuation law of the transverse magnetization vector. The attenuation rate of the envelope curve is closely related to the transverse relaxation time of organic matter in the water sample. Small-molecule organic matter has a longer transverse relaxation time, resulting in a slower attenuation rate; large-molecule organic matter or organic matter tightly bound to water molecules has a shorter transverse relaxation time, resulting in a faster attenuation rate. This invention sets the acquisition duration to 2048 milliseconds and the sampling frequency to 10000 Hz, obtaining a free-induction attenuation signal sequence containing 20480 sampling points per detection. The 2048-millisecond acquisition duration can completely record the relaxation process of various organic pollutants in the water sample, ensuring sufficient frequency resolution for subsequent Fourier transform processing. Figure 1 A small amount of random noise can also be observed superimposed on the signal. The noise mainly comes from the thermal noise of the receiving circuit and environmental electromagnetic interference. Subsequent signal preprocessing steps will suppress the noise to improve the spectrum quality.

[0026] After the pulse is applied, the transverse magnetization vector precesses around the magnetic field direction under the influence of the main magnetic field. Simultaneously, due to the inhomogeneity of the magnetic field and intermolecular interactions, the transverse magnetization vector gradually decays. The electromagnetic induction signal generated in this process is the free induction decay signal. This invention sets the acquisition duration of the free induction decay signal to 2048 milliseconds and the sampling frequency to 10000 Hz. The 2048 millisecond acquisition duration can completely record the relaxation process of various organic substances in the water sample, because the transverse relaxation time of most organic pollutants in water is in the range of 100 to 1500 milliseconds. The 10000 Hz sampling frequency is determined according to the Nyquist sampling theorem. This frequency can acquire NMR signals within a frequency range of ±5000 Hz without aliasing, which is sufficient to cover the chemical shift distribution range of hydrogen nuclei in organic matter in the water sample. According to the above parameter settings, each detection can obtain a free induction decay signal sequence containing 20480 sampling points.

[0027] In optional implementations, the acquisition duration can be adjusted according to the characteristics of the target pollutant. For detection scenarios primarily focusing on small-molecule organic compounds, due to their longer transverse relaxation times, the acquisition duration can be extended to 4096 milliseconds to obtain more complete relaxation information. For detection scenarios primarily focusing on large-molecule organic compounds or colloidal pollutants, due to their shorter transverse relaxation times, the acquisition duration can be shortened to 1024 milliseconds to improve detection efficiency. The sampling frequency can also be adjusted according to actual needs. In applications that do not focus on high-frequency chemical shift regions, the sampling frequency can be reduced to 5000 Hz, thereby reducing data storage and transmission pressure.

[0028] After obtaining the free induction decay signal sequence, a series of preprocessing operations are required to obtain an NMR spectrum suitable for analysis. The first step in preprocessing is to perform a Fourier transform, converting the time-domain free induction decay signal into a frequency-domain NMR spectrum. Before performing the Fourier transform, the free induction decay signal sequence is first zero-paddinged. Specifically, the first 16,384 sampling points of the free induction decay signal sequence are truncated, and then 16,384 zero-value points are added at the end, forming an extended sequence with a total length of 32,768 points. The purpose of zero padding is to improve the digital resolution of the spectrum without increasing the actual amount of information, making the peak shapes in the final spectrum smoother and the peak positions more accurately read. After expanding the number of data points to 32,768, the digital resolution of the spectrum is improved from approximately 0.49 Hz per point to approximately 0.31 Hz per point, which is beneficial for distinguishing different resonance peaks with similar chemical shifts. The reason for choosing 16,384 valid sampling points instead of all 20,480 sampling points is that the end part of the freely inductively decayed signal mainly contains noise components, and the signal amplitude has decayed to below the noise level. Retaining this part of the data would reduce the signal-to-noise ratio of the spectrum.

[0029] refer to Figure 2 , Figure 2 The horizontal axis represents the chemical shift, expressed in parts per million (ppm), ranging from 0 to 9, and is displayed in ascending order from right to left, following the convention of NMR spectroscopy. The vertical axis represents the spectral intensity, expressed in arbitrary units. Chemical shift is a dimensionless parameter in NMR spectroscopy describing the position of resonance peaks, reflecting the influence of the chemical environment of the hydrogen nucleus on its resonance frequency. Hydrogen nuclei in different types of organic compounds exhibit characteristic resonance peaks located at different chemical shift positions in the spectrum due to their different chemical environments. Figure 2 It can be observed that the NMR spectrum exhibits distinct resonance peak clusters in multiple chemical shift regions, corresponding to characteristic signals of different types of organic pollutants in the water sample. The resonance peak clusters appearing in the chemical shift region of 0.8 to 2.0 correspond to aliphatic hydrocarbons; the methylene and methyl hydrogen nuclei in these compounds generate resonance signals in this region, with broad peak shapes and high intensities, reflecting the presence of petroleum hydrocarbon pollutants. The resonance peak clusters appearing in the chemical shift region of 3.5 to 4.5 correspond to organochlorine compounds; the electron-withdrawing effect of chlorine-containing substituents shifts the chemical shift of adjacent hydrogen nuclei towards a lower field. The resonance peak clusters appearing in the chemical shift region of 6.5 to 7.0 correspond to phenolic compounds; the hydrogen nuclei adjacent to the hydroxyl group on the benzene ring generate characteristic signals in this region. The resonance peak clusters appearing in the chemical shift region of 7.0 to 7.5 correspond to benzene series compounds; the aromatic hydrogen nuclei of monosubstituted and polysubstituted benzenes exhibit complex splitting peak shapes in this region. The resonance peak clusters appearing in the chemical shift region of 7.5 to 8.5 correspond to polycyclic aromatic hydrocarbons. The aromatic hydrogen nuclei of fused-ring aromatic hydrocarbons appear at lower field positions due to the enhanced ring current effect. Figure 2 The peak area of ​​each resonance peak in the spectrum is directly proportional to the concentration of the corresponding organic matter. By analyzing the position and area of ​​each characteristic peak in the spectrum, qualitative identification and quantitative analysis of multiple organic pollutants in water samples can be achieved. This invention uniformly discretizes the nuclear magnetic resonance spectrum within the chemical shift range of 0 to 9 at intervals of 0.01, obtaining a one-dimensional nuclear magnetic resonance spectral vector containing 900 intensity data points, which serves as the input data for a subsequent one-dimensional convolutional neural network model.

[0030] Performing a Fast Fourier Transform (FFT) on the extended sequence yields the complex spectrum in the frequency domain. The FFT is an efficient implementation of the Discrete Fourier Transform (DFT), with a computational complexity of O(n log n). ,in This indicates the number of data points. For 32,768 data points, the Fast Fourier Transform (FFT) can complete the calculation in milliseconds, meeting the real-time requirements of online detection. The frequency domain complex spectrum obtained after the Fourier transform contains two components: a real part and an imaginary part. The real part corresponds to the absorption mode spectral lines, and the imaginary part corresponds to the dispersive mode spectral lines. Ideally, the absorption mode spectral lines exhibit a symmetrical Lorentz shape, which facilitates peak identification and quantitative analysis. Therefore, subsequent processing aims to extract the pure absorption mode spectrum.

[0031] Due to the unavoidable electronic delays and filter phase shifts in practical detection systems, the original complex spectrum in the frequency domain usually exhibits phase distortion, manifesting as a mixture of absorption and dispersive modes, with spectral lines deviating from the ideal Lorentz shape. To obtain a pure absorption mode spectrum, phase correction processing is required. Phase correction consists of two steps: zero-order phase correction and first-order phase correction. Zero-order phase distortion is characterized by a constant phase shift throughout the spectrum, primarily caused by the time difference between the start of signal acquisition and the end of the RF pulse. First-order phase distortion is characterized by a linear phase shift with frequency, mainly caused by the group delay characteristics of the receiver filter.

[0032] The zero-order phase correction angle is determined using the phase trajectory method. Specifically, the real part of the complex spectrum in the frequency domain is used as the abscissa and the imaginary part as the ordinate to plot a phase trajectory curve, which appears as an approximately ellipse in the complex plane. When a deviation exists in the zero-order phase, an angle exists between the principal axis and the real axis of the ellipse; this angle is the zero-order phase angle that needs correction. By rotating the phase trajectory curve to align its principal axis with its real axis, the numerical value of the zero-order phase correction angle can be determined. Mathematically, this operation is equivalent to multiplying each data point of the complex spectrum in the frequency domain by a complex factor. ,in The imaginary unit, This is the zero-order phase correction angle.

[0033] After zero-order phase correction, first-order phase correction is performed to eliminate frequency-dependent phase distortion. The method for first-order phase correction is to apply a linearly increasing phase shift along the chemical shift axis to each frequency point. Let the shift of a certain frequency point in the spectrum relative to the reference frequency be... Then the first-order phase correction amount that needs to be applied at this frequency point is: ,in This represents the first-order phase correction coefficient, measured in radians per hertz. The determination of the first-order phase correction coefficient typically optimizes the flatness of the baseline region in the spectrum through iterative adjustments. The value is calculated until the fluctuations in the baseline region are minimized. After phase correction, the real part of the corrected frequency domain complex spectrum is extracted as the absorption mode spectrum. At this point, the spectral lines exhibit symmetrical absorption peak shapes, which are suitable for subsequent qualitative and quantitative analysis.

[0034] Baseline drift is a common problem in absorption mode spectra, characterized by a slowly undulating curve rather than a horizontal straight line. Baseline drift is primarily caused by the following factors: background signal from the detection coil, contribution from trace hydrogen impurities in the probe material, and truncation effects during the Fourier transform. Baseline drift affects the accurate determination of the integral area of ​​the resonance peak, thus impacting the quantitative analysis results; therefore, baseline correction is necessary.

[0035] The core idea of ​​baseline correction is to identify baseline regions without resonance peaks in the spectrum, fit a baseline function curve using data points from these regions, and then subtract the baseline function curve from the original spectrum to obtain the corrected spectrum. In water sample NMR spectra, the regions with chemical shifts of -1 to 0 and 9 to 10 typically lack resonance peaks for organic matter and can be used as baseline sampling areas. The rationale for selecting these two regions is that the hydrogen nuclear chemical shifts of common water organic pollutants are mainly distributed in the range of 0.5 to 8.5, and the regions of -1 to 0 and 9 to 10 are outside this range, thus representing the true baseline level.

[0036] After extracting the spectral intensity values ​​corresponding to each frequency point within the baseline sampling area, a third-order polynomial is used for least-squares fitting to obtain the baseline function curve. The reason for choosing a third-order polynomial is that first- or second-order polynomials are difficult to accurately describe the curvature of the baseline, while fourth- or higher-order polynomials may overfit the noise fluctuations in the baseline sampling area, introducing new errors. Let the chemical shift be... The baseline function curve is represented as ,in , , , The coefficients are polynomials, determined by least-squares fitting of the data points in the baseline sampling region. The baseline-corrected NMR spectrum is obtained by subtracting the baseline function curve's value at the corresponding frequency from the spectral intensity value of the absorption mode spectrum at each frequency point.

[0037] In alternative implementations, when the complex composition of the water sample leads to irregular baseline morphology, a more flexible baseline correction method can be employed. For example, an asymmetric least squares smoothing algorithm can be used to automatically identify the baseline. This algorithm iteratively smooths the spectrum, gradually excluding resonance peak regions from the baseline fitting, ultimately obtaining a smooth curve containing only the baseline component. Another alternative method is to use wavelet transform for baseline correction, separating low-frequency baseline drift components from high-frequency resonance peak components by decomposing different frequency components of the spectrum.

[0038] The baseline-corrected NMR spectrum needs to be discretized for input into the subsequent neural network model. The discretization method involves uniformly sampling the NMR spectrum within the chemical shift range of 0 to 9, taking one intensity value at each chemical shift interval of 0.01, ultimately resulting in a one-dimensional NMR spectral vector containing 900 intensity data points. The selection of chemical shifts 0 to 9 as the effective range is based on the fact that this interval covers the chemical shift distribution of hydrogen nuclei in the vast majority of organic pollutants in water. A chemical shift interval of 0.01 corresponds to a frequency resolution of approximately 2.1 Hz, which is sufficient to distinguish the characteristic peaks of most organic compounds, while keeping the data dimensionality within a reasonable range, which is beneficial for the training and inference efficiency of the neural network model.

[0039] In one alternative implementation, for applications requiring higher resolution, the chemical shift interval can be reduced to 0.005, resulting in a one-dimensional NMR spectral vector containing 1800 data points; for edge computing scenarios with higher computational efficiency requirements, the chemical shift interval can be expanded to 0.02, resulting in a one-dimensional NMR spectral vector containing 450 data points. The chemical shift range can also be adjusted according to the type of target pollutant; for example, when focusing only on aromatic compounds, the range can be reduced to 6 to 9, and when focusing only on aliphatic compounds, the range can be reduced to 0 to 3.

[0040] To achieve seasonal water environment quality assessment, it is necessary to classify and organize the collected one-dimensional nuclear magnetic resonance (NMR) spectral vectors according to the season. This invention adopts the meteorologically accepted seasonal division method: spring (March to May), summer (June to August), autumn (September to November), and winter (December to February of the following year). This division method is consistent with the climatic characteristics of most parts of my country and can reflect the impact of changes in hydrological conditions on water quality in different seasons. All one-dimensional NMR spectral vectors collected from each sampling section within each season are compiled to form NMR spectral datasets corresponding to each sampling section in each season, providing data input for subsequent deep learning-based spectral analysis and matter-element extension-based comprehensive evaluation.

[0041] Nuclear magnetic resonance (NMR) spectra contain molecular structural information of various organic pollutants in water samples, with different types of organics exhibiting characteristic resonance peak patterns in the spectra. Traditional NMR spectrum interpretation relies on manual identification by professionals based on chemical shift values ​​and peak shape characteristics. This method is time-consuming, labor-intensive, and highly subjective, making it difficult to meet the automated requirements of online detection. The development of deep learning technology has provided a new approach for the automatic interpretation of NMR spectra. Convolutional neural networks can automatically learn the feature patterns in the spectra, achieving end-to-end mapping from the original spectrum to pollutant concentration indicators.

[0042] One-dimensional convolutional neural networks (CNNs) are particularly well-suited for processing sequential data such as NMR spectra. Unlike two-dimensional CNNs that process images, the kernels of one-dimensional CNNs slide only in one dimension, effectively capturing local features distributed along the chemical shift axis in the spectrum. Resonance peaks in NMR spectra typically have a certain width, and there is strong correlation between adjacent data points. Convolutional operations can utilize this local correlation to extract features such as peak shape, peak position, and peak area. By stacking multiple convolutional layers, the network can extract abstract features layer by layer from low to high levels, ultimately achieving a comprehensive understanding of the spectrum.

[0043] The one-dimensional convolutional neural network model constructed in this invention adopts a hierarchical structure design, which includes, in sequence, an input layer, a first convolutional layer, a first batch normalization layer, a first activation layer, a first pooling layer, a second convolutional layer, a second batch normalization layer, a second activation layer, a second pooling layer, a third convolutional layer, a third batch normalization layer, a third activation layer, a global average pooling layer, a fully connected layer, and an output layer. This three-stage convolutional and pooling structure design draws on the successful experience of classic convolutional neural networks, and can achieve good feature extraction capabilities while maintaining moderate model complexity.

[0044] The input layer receives the one-dimensional nuclear magnetic resonance (NMR) spectral vector obtained in the preceding steps. The input dimension is 900, corresponding to 900 spectral intensity values ​​sampled at 0.01 intervals within the chemical shift range of 0 to 9. Before being fed into the network, the input data needs to be standardized. The mean of the entire spectral vector is subtracted from the value of each data point, and then divided by the standard deviation, making the mean of the input data close to 0 and the standard deviation close to 1. Standardization accelerates the convergence process of network training and improves the model's generalization ability to water samples with different signal intensities.

[0045] The first convolutional layer is the network's first feature extraction layer, containing 32 convolutional kernels, each with a length of 15 and a stride of 1. The choice of a kernel length of 15 is based on the fact that the full width at half maximum (FWHM) of a single resonance peak in an NMR spectrum typically ranges from 0.05 to 0.15 chemical shift units, corresponding to 5 to 15 data points. A 15-point convolutional kernel can completely cover the width of a typical resonance peak, thus effectively extracting peak shape features. To maintain the consistency between the output and input lengths, seven zeros are padded at both ends of the input sequence; this padding method is called same-size padding. The 32 convolutional kernels can learn 32 different local feature patterns, such as single-peak, double-peak, and shoulder-peak morphologies at different locations.

[0046] The mathematical process of convolution can be described as follows: For each position in the input sequence, multiply the 15 input values ​​covered by the convolution kernel by the 15 parameters of the convolution kernel position by position, and then sum them to obtain the convolution output value at that position. Let the input sequence be... The convolution kernel parameters are , bias is Then the position Convolution output at the point The calculation method is as follows ,in The convolution kernel is represented by the first... One parameter, This represents the numerical value at the corresponding position in the input sequence. 32 convolutional kernels slide across the input sequence to perform computations, producing 32 convolutional output sequences of length 900. These sequences constitute the output feature map of the first convolutional layer.

[0047] The first batch normalization layer follows the first convolutional layer and normalizes the convolutional output. The purpose of batch normalization is to alleviate the internal covariate shift problem during deep network training, i.e., the training instability caused by the constantly changing distribution of input data across layers as training progresses. The specific normalization process involves calculating the mean of 900 elements for each convolutional output sequence. and variance Each element of the sequence Transform into ,in This is a very small constant, with a value of 0.00001, used to prevent the denominator from being zero. The transformed value is then multiplied by a learnable scaling parameter. And add learnable offset parameters The final output is Scaling parameters and offset parameters During training, the network learns automatically through backpropagation, enabling it to recover the original data distribution when needed.

[0048] refer to Figure 3 , Figure 3 The upper part shows a local segment of the input spectral vector, with the horizontal axis representing the position index and the vertical axis representing the spectral intensity value. The curve presents the typical multi-peak shape of an NMR spectrum. The gray-filled area in the figure marks the window range covered by the convolution kernel at the current sliding position, and two vertical dashed lines define the left and right boundaries of the convolution kernel window. The core idea of ​​convolution operation is to extract local feature information by weighted summation of local regions of the input sequence using learnable convolution kernel parameters. Figure 3 The middle section shows the parameter distribution of the convolutional kernels, with the horizontal axis representing parameter indices and the vertical axis representing parameter values. The first convolutional layer of this invention uses a kernel of length 15, and the figure displays the numerical distribution of the 15 kernel parameters in a bar chart format. From the parameter distribution, it can be observed that the kernel parameters exhibit a symmetrical distribution characteristic, with a high center and low ends. This distribution pattern is beneficial for extracting peak-shaped features from the input sequence. The kernel parameters are automatically learned during network training through the backpropagation algorithm. Different kernels can learn different types of local feature patterns, such as single-peak features, double-peak features, or shoulder-peak features. Figure 3 The lower section shows the output feature map generated by the convolution operation, with the horizontal axis representing the position index and the vertical axis representing the feature map value. The convolution output is calculated as follows: at each sliding position, the input value covered by the convolution kernel is multiplied position by position by the convolution kernel parameters, and then summed to obtain the convolution output value at that position. Let the input sequence be... The convolution kernel parameters are Then the position Convolution output at the point The calculation method is as follows ,in The convolution kernel is represented by the first... One parameter, This represents the value at the corresponding position in the input sequence. Figure 3The bottom section uses dots to indicate the output value corresponding to the current convolutional kernel window position. This output value reflects the degree of matching between the input sequence and the convolutional kernel pattern in that local region. When the local morphology of the input sequence is highly similar to the distribution of convolutional kernel parameters, the convolutional output value is larger; conversely, it is smaller. By setting multiple different convolutional kernels, the network can simultaneously extract multiple types of local features, providing rich feature representations for subsequent pollutant concentration prediction.

[0049] The first activation layer performs a modified linear unit activation operation on the output of the first batch of normalized layers. The introduction of the activation function adds nonlinear transformation capability to the network, enabling it to learn complex nonlinear mapping relationships between inputs and outputs. The calculation rule for the modified linear unit is simple and clear: each element is compared with zero; if the element is greater than zero, it remains unchanged; if the element is less than or equal to zero, it is set to zero. The mathematical expression is as follows: ,in For input values, This is the output value after activation. Compared to traditional hyperbolic tangent or logistic functions, the modified linear unit (MRU) has the advantages of simple computation and gradient non-vanishing properties, and has become the most commonly used activation function in deep learning.

[0050] The first pooling layer performs max pooling on the output of the first activation layer, with a pooling window length of 4 and a pooling stride of 4. The purpose of pooling is to reduce the spatial resolution of the feature map, decrease the computational cost of subsequent layers, and enhance the robustness of the features to small shifts in the input position. Specifically, max pooling divides each 900-character sequence into 225 non-overlapping windows, each containing four consecutive elements, and takes the maximum value within each window as the pooling output. After processing by the first pooling layer, the 32 feature sequences of length 900 are compressed into 32 feature sequences of length 225. Max pooling is chosen over average pooling because the effective information in NMR spectra is mainly concentrated at the resonance peaks, which represent the most important feature information; max pooling preserves this peak information.

[0051] The second convolutional layer further extracts higher-level features based on the output of the first pooling layer. It contains 64 convolutional kernels, each with a length of 11, a stride of 1, and 5 padding elements. The increase in the number of convolutional kernels from 32 to 64 allows the network to express richer feature combinations. The kernel length is reduced from 15 to 11 because the effective resolution of the feature sequence has decreased after downsampling by the first pooling layer; shorter kernels are sufficient to capture local features at the current scale. Each convolutional kernel covers all 32 input sequences in the depth direction, achieving cross-channel feature fusion. The second convolutional layer outputs 64 feature sequences with a length of 225.

[0052] The second batch of normalization layers and the second activation layer operate in exactly the same way as the first batch of normalization layers and the first activation layer, performing normalization processing and corrected linear unit activation operations, respectively. The pooling window length of the second pooling layer is 5, and the pooling stride is 5, compressing 64 feature sequences of length 225 into 64 feature sequences of length 45. The window length of 5 was chosen instead of 4 to ensure that 225 is divisible, avoiding the complexity of boundary processing.

[0053] The third convolutional layer is the last convolutional layer in the network, containing 128 convolutional kernels, each with a length of 7, a stride of 1, and 3 padding elements. Increasing the number of kernels to 128 allows the network to learn more abstract and complex feature representations. The kernel length is then reduced to 7 to match the resolution of the current feature sequence. The third convolutional layer outputs 128 feature sequences of length 45. The third batch normalization layer and the third activation layer then normalize and activate the convolutional outputs sequentially.

[0054] The global average pooling layer transforms the 128 feature sequences of length 45 output from the third activation layer into 128 scalar values. The transformation is achieved by calculating the arithmetic mean of all 45 elements within each feature sequence, resulting in a scalar representing the overall response intensity of that channel. Compared to fully connected layers, global average pooling has the advantages of fewer parameters, less susceptibility to overfitting, and adaptability to variations in input sequence length. After global average pooling, a feature vector containing 128 elements is obtained, representing a highly compressed representation of the entire NMR spectrum.

[0055] The fully connected layer maps a 128-element feature vector to a 6-element output vector. This mapping is achieved through matrix multiplication and bias addition: a 128x6 connection matrix is ​​used. and a bias vector containing 6 elements , to feature vector Multiplying by the connection matrix and adding the bias vector yields the output vector. ,in This represents the transpose of the connection matrix. The 768 parameters of the connection matrix and the 6 parameters of the bias vector are learned during training.

[0056] The output layer outputs six elements corresponding to the concentration indicators of six categories of organic pollutants: benzene series organic pollutants, phenolic organic pollutants, polycyclic aromatic hydrocarbons (PAHs), organochlorine pollutants, petroleum hydrocarbons, and total dissolved organic carbon (TOC). These six indicators cover the most common types of organic pollutants in river water bodies, comprehensively reflecting the organic pollution status of the water. Benzene series organic pollutants include monocyclic aromatic hydrocarbons such as benzene, toluene, and xylene, mainly originating from emissions from the petrochemical and paint industries. Phenolic pollutants include phenol and its derivatives, mainly originating from wastewater from coking, pharmaceutical, and pesticide production. PAHs include polycyclic aromatic compounds such as naphthalene, anthracene, and pyrene, mainly originating from the incomplete combustion of fossil fuels. Organochlorine pollutants include chlorinated organic compounds such as chloroform, carbon tetrachloride, and polychlorinated biphenyls (PCBs), mainly originating from chemical production and disinfection byproducts. Petroleum hydrocarbons include petroleum components such as alkanes and cycloalkanes, mainly originating from oil spills and oily wastewater discharge. Total dissolved organic carbon reflects the content level of all dissolved organic matter in a water body.

[0057] Training a one-dimensional convolutional neural network model requires labeled sample data. The labeled data comes from water sample analysis performed using gas chromatography-mass spectrometry (GC-MS) provided by the Chemical Research Service. GC-MS is the gold standard method for the quantitative analysis of organic pollutants, accurately determining the concentration of various organic compounds in water samples. A training dataset is constructed using nuclear magnetic resonance (NMR) spectral vectors as input and GC-MS detection results as labels. The training dataset should cover water samples with different levels of pollution and combinations of pollution types to ensure good generalization ability of the model. It is recommended that the training dataset contain at least 2000 samples, with 1600 used for training and 400 for validation.

[0058] The model training employed an improved version of the stochastic gradient descent optimization algorithm, the Adaptive Moment Estimation (IME) algorithm. IEM adaptively adjusts the learning rate of each parameter based on the first and second moments of the gradient, enabling rapid convergence in the early stages of training and fine-tuning in the later stages. The initial learning rate was set to 0.001, and it was decayed by multiplying by 0.1 every 50 training epochs. The loss function used was mean squared error, calculated as the average of the squared errors between the six concentration indicators output by the network and the label value. The training process lasted for 200 epochs, with each epoch iterating through all training samples once. Training was terminated early to prevent overfitting when the loss function value on the validation set stopped decreasing for 20 consecutive epochs.

[0059] In an optional implementation, transfer learning strategies can be employed to accelerate model training. First, the model is pre-trained on a large-scale, general NMR spectral dataset to learn general feature representations of the spectra. Then, it is fine-tuned on the target water sample dataset to adapt the model to the specific detection task. This approach can achieve good detection performance even with limited labeled data. Another optional approach is to augment the training dataset using data augmentation techniques, including adding random noise to the spectral vectors, performing small translations or scaling on the spectral vectors, etc., thereby improving the model's robustness.

[0060] After training, each one-dimensional NMR spectral vector in the NMR spectral dataset is input into a one-dimensional convolutional neural network model, and the corresponding six-dimensional pollutant feature vector is calculated through forward propagation. The six-dimensional pollutant feature vectors of all samples from the same sampling section within the same season are arranged in chronological order of sampling time to form the pollutant feature matrix corresponding to each sampling section for each season. Assuming that a sampling section collected 30 water samples in a certain season, the pollutant feature matrix corresponding to that sampling section for that season is a 30-row, 6-column matrix, where each row represents the pollutant feature vector of a water sample, and each column represents the concentration indicator value sequence of a class of pollutants.

[0061] The pollutant characteristic matrix provides information on the concentration of organic pollutants at each sampling section in each season, but these discrete values ​​have not yet been translated into intuitive water quality assessment results. Matter-element extension theory is an effective mathematical tool for handling contradictory and incompatible problems, capable of transforming quantitative data from multiple indicators into comprehensive qualitative assessment conclusions. Applying the matter-element extension method to water environmental quality assessment can fully consider the complex correlation between various pollutant indicators and water quality levels, providing scientifically sound and reasonable assessment results.

[0062] The matter-element (MEE) is a fundamental concept in MEE extension theory, used to describe the relationships between things, characteristics, and their values. A MEE consists of three elements: the name of the thing, the name of the characteristic, and the value, denoted as an ordered triplet. In water environmental quality assessment, the thing is the water quality grade, the characteristic is the concentration indication value of various pollutants, and the value is the range of values ​​for each characteristic within that grade. Based on the five water quality grades stipulated in my country's surface water environmental quality standards, five classical domain MEEs and one section domain MEE need to be established.

[0063] The classical domain element describes the standard range of pollutant concentration indicators for each water quality level. Taking Class I water quality as an example, this level represents water bodies with excellent quality and basically maintained in their natural state, where the pollutant concentration indicators should be at the lowest level. Based on the surface water environmental quality standards and the concentration indicator range output by the one-dimensional convolutional neural network model of this invention, the value ranges of each feature in the classical domain element for Class I water quality are determined as follows: the concentration indicator value for benzene series organic pollutants ranges from 0 to 0.1, the concentration indicator value for phenolic organic pollutants ranges from 0 to 0.08, the concentration indicator value for polycyclic aromatic hydrocarbons ranges from 0 to 0.05, the concentration indicator value for organochlorine pollutants ranges from 0 to 0.06, the concentration indicator value for petroleum hydrocarbons ranges from 0 to 0.12, and the concentration indicator value for total dissolved organic carbon ranges from 0 to 0.15. For Class II, Class III, Class IV, and Class V water quality, the upper and lower limits of the value ranges of each characteristic increase sequentially in order of increasing pollution level.

[0064] Taking the concentration indication values ​​of benzene series organic pollutants as an example, the setting of the value ranges for each level is explained as follows: the value range for Class I water quality is 0 to 0.1, for Class II water quality it is 0.1 to 0.25, for Class III water quality it is 0.25 to 0.45, for Class IV water quality it is 0.45 to 0.7, and for Class V water quality it is 0.7 to 1.0. The value ranges for the other five characteristics are set according to a similar progressive relationship to ensure that the boundaries between adjacent levels are connected without gaps.

[0065] The section element describes the maximum possible range of pollutant concentration indicators across all water quality grades. The section element represents the set of all water quality grades, and the value range of each feature is the union of its value ranges in all classical section elements. Based on the aforementioned settings, the concentration indicator value of benzene series organic pollutants in the section element ranges from 0 to 1.0; the section value ranges for other features are determined similarly. The section element serves as a reference benchmark in correlation calculations, used to determine the degree to which the sample being evaluated deviates from the standard range of each grade.

[0066] After establishing classical domain matter-element and segmented domain matter-element, matter-element correlation degree calculation is performed on each 6-dimensional pollutant feature vector in the pollutant feature matrix to determine its corresponding water environmental quality assessment level. Correlation degree is a quantitative indicator in matter-element extension theory that measures the degree of affinity between a thing and a certain level. The higher the correlation degree, the more it conforms to the characteristics of that level. The level with the highest correlation degree is the evaluation conclusion.

[0067] The first step in correlation calculation is to calculate the distance between each component of the feature vector of the pollutant to be evaluated and the interval of the classical domain value for each level. Distance is a fundamental quantity in extension theory describing the positional relationship between a point and an interval; it characterizes both whether a point is inside the interval and how close the point is to the interval. Let the concentration indication value of a certain pollutant to be evaluated be... The corresponding range of classical field values ​​at a certain level is: Then the distance between the indicated value and the range of values ​​is The calculation method is as follows: First, calculate the midpoint of the value interval. and the half length of the range of values Then calculate the distance between the indicator value and the midpoint. Finally, the distance was obtained. .when When the absolute value is greater, it indicates that the indicated value is within the range of values; the larger the absolute value, the closer it is to the center of the range. When, it indicates that the indicated value is exactly at the boundary of the value range; when When the value is outside the range, it indicates that the value is located outside the range. The larger the value, the further it is from the range.

[0068] refer to Figure 4 , Figure 4 The horizontal axis represents the pollutant concentration indicator, ranging from 0 to 1.0; the vertical axis represents the correlation degree, ranging from -2 to 1.2. The figure shows the correlation degree function curves for Class I, II, III, IV, and V water quality using five different line types. Correlation degree is a quantitative indicator in matter-element extension theory that measures the degree to which an evaluated object conforms to a certain level; a higher correlation degree indicates a stronger conformity to the characteristics of that level. Figure 4 It can be observed that each correlation curve exhibits a single-peak shape, rising first and then falling, with the peak value appearing at the midpoint of the classical range of the corresponding water quality grade. Taking Class I water quality as an example, its classical range is 0 to 0.1. The correlation curve is positive within this range and reaches its maximum value of 1 at the midpoint of the range (0.05), indicating that the sample's conformity with Class I water quality is highest when the pollutant concentration indication value is exactly in the center of the Class I water quality standard range. When the pollutant concentration indication value deviates from the classical range, the correlation gradually decreases and becomes negative; the further the deviation, the smaller the correlation. Figure 4 The boundaries of the classical range values ​​for each water quality grade are marked by vertical dashed lines, from left to right: 0.1, 0.25, 0.45, and 0.7. This divides the range of pollutant concentration indicators into five adjacent intervals, corresponding to the standard ranges for the five water quality grades. The correlation is calculated based on the classical range distance. and node distance Two intermediate variables. Classical domain distance. This indicates the positional relationship between the indicator value to be evaluated and the interval of classical domain values, calculated as follows: ,in The indicator value to be evaluated. The midpoint of the value range, This is half the length of the value interval. When... At that time, the correlation degree of a single indicator ;when At that time, the correlation degree of a single indicator .from Figure 4 It can also be observed that at the boundary between two adjacent levels, the two correlation curves intersect, and the correlation degrees of the two levels are equal at the intersection point. This reflects that the matter-element extension method can naturally handle the ambiguity of level boundaries. In actual evaluation, the correlation degree of the sample to be evaluated with respect to the five water quality levels is calculated, and the level corresponding to the largest correlation degree value is taken as the evaluation conclusion.

[0069] The distance between the indicated value of the pollutant concentration to be evaluated and the corresponding characteristic value interval in the segment element is calculated in the same way and denoted as the segment distance. The section distance reflects the position of the indicated value relative to the entire range of possible values. Since the section range includes the entire range of classical ranges, the indicated value to be evaluated should normally be located within the section range, i.e., the section distance is negative or zero.

[0070] The correlation degree of a single index is calculated based on the classical domain distance and the section domain distance. Let the classical domain distance of a pollutant concentration indication value with respect to a certain level be denoted as . The corresponding classical field value interval half-length is The node spacing is Then the correlation of a single indicator The calculation method is as follows: If ,but At this point, the indicator value is located within the classical domain interval, and the correlation is positive; the closer to the center of the interval, the greater the correlation. ,but At this point, the indicator value is outside the classical domain interval, and the correlation is negative. The further away from the classical domain interval, the lower the correlation. The correlation ranges from negative infinity to 1. The correlation reaches its maximum value of 1 when the indicator value is exactly at the midpoint of the classical domain interval, and the correlation tends to negative infinity when the indicator value is far away from the classical domain interval.

[0071] For the 6-dimensional pollutant feature vector to be evaluated, the correlation degree of each of the six pollutant concentration indicators with respect to the same water quality level is calculated. Then, the correlation degrees of the six individual indicators are summed and divided by 6 to obtain the comprehensive correlation degree of the pollutant feature vector with respect to that water quality level. The comprehensive correlation degree reflects the overall degree of conformity between the sample to be evaluated and a certain water quality level across all pollutant indicators. This calculation process is repeated to obtain the five comprehensive correlation degrees of the pollutant feature vector with respect to Class I, Class II, Class III, Class IV, and Class V water quality.

[0072] The rule for determining the water environment quality assessment level based on the comprehensive correlation degree is as follows: compare the values ​​of the five comprehensive correlation degrees, and take the water quality level corresponding to the largest comprehensive correlation degree value as the water environment quality assessment level of the water sample corresponding to the pollutant feature vector. This rule reflects the superiority evaluation concept in extension theory, that is, selecting the level that best fits the evaluation object as the evaluation conclusion. When there is a special case where multiple levels have equal comprehensive correlation degrees and all are the maximum values, the level with the poorest water quality is selected as the evaluation conclusion to reflect the principle of prudence in environmental protection.

[0073] The aforementioned matter-element correlation degree calculation and grade determination process is performed on all 6-dimensional pollutant feature vectors in the pollutant feature matrix corresponding to each sampling section for each season. Taking a sampling section with a pollutant feature matrix containing 30 samples for a certain season as an example, 30 water environmental quality assessment grades are obtained after evaluation. The frequency of each water quality grade in these 30 assessment grades is statistically analyzed. For example, Class I appears 5 times, Class II appears 12 times, Class III appears 10 times, Class IV appears 3 times, and Class V appears 0 times. Then, Class II water quality with the highest frequency is taken as the representative water environmental quality grade of the sampling section for that season. Using frequency statistics rather than simple averaging to determine the representative grade can avoid the excessive influence of individual abnormal samples on the overall evaluation results, making the evaluation conclusions more robust and reliable.

[0074] In an optional implementation, when a more refined evaluation conclusion is required, a grade characteristic value method can be used. Specifically, water quality levels I to V are assigned characteristic values ​​from 1 to 5. The evaluation level of each sample is converted into the corresponding characteristic value, and then the arithmetic mean is calculated to obtain the average grade characteristic value for that sampling section in that season. The average grade characteristic value reflects the degree of continuous change in water quality levels. For example, an average grade characteristic value of 2.3 indicates that the water quality is between Class II and Class III, but closer to Class II. Another optional approach is to use a weighted average based on comprehensive correlation, using the magnitude of the comprehensive correlation of each level as the confidence level to provide a fuzzy grade classification.

[0075] Representative water quality levels from each sampling section in spring, summer, autumn, and winter are compiled to form a comprehensive seasonal river water quality assessment. The assessment results are presented in tabular form, with rows representing each sampling section and columns for the four seasons, and the table content showing the corresponding water quality level. By comparing and analyzing the seasonal variation patterns of each section, the characteristics of seasonal fluctuations in water quality can be identified. For example, some sections show improved water quality during the high-water season and deteriorated during the low-water season, while other sections show the opposite trend. By comparing and analyzing the spatial distribution characteristics of sections in each season, the spatial patterns of pollution can be identified. For example, the water quality of upstream sections is generally better than that of downstream sections, or the water quality near the confluence of certain tributaries shows a significant decline. These analytical conclusions can provide a scientific basis for watershed water environment management decisions and guide the optimized deployment of chemical research service monitoring projects and the precise implementation of environmental protection measures.

[0076] The present invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for assessing the seasonal river water environmental quality based on chemical research service data, characterized in that: Includes the following steps: Step 1, Online NMR Detection and Spectrum Data Acquisition: Low-field online NMR detection devices are deployed at each sampling section. Radio frequency excitation pulses are applied to the water samples entering the detection chamber, and free induction decay signals are collected. Fourier transform, phase correction, and baseline correction are sequentially performed on the free induction decay signals to obtain NMR spectra. The NMR spectra are discretized to obtain one-dimensional NMR spectral vectors. The one-dimensional NMR spectral vectors of each sampling section are collected according to the season to form an NMR spectral dataset. Step 2, Nuclear Magnetic Resonance Spectrum Analysis and Pollutant Feature Extraction Based on Convolutional Neural Network: Construct a one-dimensional convolutional neural network model, which includes a convolutional layer, a batch normalization layer, an activation layer, a pooling layer, a global average pooling layer, a fully connected layer, and an output layer connected in sequence. Train the one-dimensional convolutional neural network model using water sample testing data provided by the chemical research service institution. Input the one-dimensional nuclear magnetic resonance spectrum vector from the nuclear magnetic resonance spectrum dataset into the trained one-dimensional convolutional neural network model to obtain pollutant feature vectors, forming pollutant feature matrices corresponding to each sampling section and each season. Step 3: Comprehensive evaluation of seasonal water environment quality based on matter-element extension: Establish classical domain matter-element and section domain matter-element corresponding to each water quality level according to the water environment quality standards. Calculate the comprehensive correlation degree of each pollutant feature vector in the pollutant feature matrix with respect to each water quality level. Determine the water environment quality evaluation level based on the comprehensive correlation degree. Summarize the water environment quality evaluation levels of each sampling section for each season to form the comprehensive evaluation result of seasonal river water environment quality.

2. The method according to claim 1, characterized in that, In step 1, the magnetic field strength of the low-field nuclear magnetic resonance online detection device is 0.5 Tesla, and a broadband radio frequency transmitting and receiving coil is configured. The radio frequency excitation pulse is a 90-degree radio frequency excitation pulse with a duration of 10 microseconds; the acquisition time of the free induction attenuation signal is 2048 milliseconds, the sampling frequency is 10000 Hz, and a free induction attenuation signal sequence containing 20480 sampling points is obtained.

3. The method according to claim 2, characterized in that, In step 1, the specific process of performing Fourier transform, phase correction, and baseline correction on the free-induction attenuated signal sequence is as follows: The first 16384 sampling points of the free-induction attenuated signal sequence are truncated, and 16384 zero points are added at the end to form an extended sequence of 32768 points. A fast Fourier transform is performed on the extended sequence to obtain the frequency domain complex spectrum. The real and imaginary parts of the frequency domain complex spectrum are extracted separately. The real part is used as the abscissa, and the imaginary part as the ordinate to plot the phase trajectory curve. The zero-order phase correction is determined by rotating the phase trajectory curve so that its principal axis coincides with the real axis. A positive angle is used to apply a linearly increasing phase shift along the chemical shift axis to each frequency point to eliminate first-order phase distortion. The real part of the corrected frequency domain complex spectrum is extracted as the absorption mode spectrum. Two regions, -1 to 0 and 9 to 10, are selected on the chemical shift axis as baseline sampling regions. The spectral intensity values ​​corresponding to each frequency point in the baseline sampling region are extracted. The baseline function curve is obtained by least-squares fitting using a third-order polynomial. The function value of the baseline function curve at the corresponding frequency point is subtracted from the spectral intensity value of the absorption mode spectrum at each frequency point to obtain the baseline-corrected NMR spectrum.

4. The method according to claim 1, characterized in that, In step 1, the nuclear magnetic resonance spectrum is uniformly discretized within the chemical shift range of 0 to 9, and one spectral intensity value is taken at every chemical shift interval of 0.01 to obtain a one-dimensional nuclear magnetic resonance spectral vector containing 900 intensity data points. The one-dimensional nuclear magnetic resonance spectral vector is then compiled into a nuclear magnetic resonance spectral dataset according to the division of spring (March to May), summer (June to August), autumn (September to November), and winter (December to February of the following year).

5. The method according to claim 1, characterized in that, In step 2, the structure of the one-dimensional convolutional neural network model includes, in sequence: input layer, first convolutional layer, first batch normalization layer, first activation layer, first pooling layer, second convolutional layer, second batch normalization layer, second activation layer, second pooling layer, third convolutional layer, third batch normalization layer, third activation layer, global average pooling layer, fully connected layer, and output layer; the input dimension of the input layer is 900.

6. The method according to claim 5, characterized in that, The first convolutional layer contains 32 convolutional kernels with a kernel length of 15, a stride of 1, and 7 padding elements. The first pooling layer has a pooling window length of 4 and a pooling stride of 4. The second convolutional layer contains 64 convolutional kernels with a kernel length of 11, a stride of 1, and 5 padding elements. The second pooling layer has a pooling window length of 5 and a pooling stride of 5. The third convolutional layer contains 128 convolutional kernels with a kernel length of 7, a stride of 1, and 3 padding elements.

7. The method according to claim 5, characterized in that, The first, second, and third normalization layers perform the following normalization processing on the input sequence: calculate the mean and variance of all elements in the input sequence, subtract the mean from each element of the input sequence, divide by the square root of the variance, add 0.00001, multiply the sum by a learnable scaling parameter, and add a learnable offset parameter; the first, second, and third activation layers perform modified linear unit activation operations, keeping elements greater than zero unchanged and setting elements less than or equal to zero to zero; the first and second pooling layers perform max pooling operations, taking the maximum value within each pooling window as the output.

8. The method according to claim 5, characterized in that, The global average pooling layer calculates the arithmetic mean of all elements in each of the 128 sequences output by the third activation layer to obtain a feature vector containing 128 elements. The fully connected layer maps the feature vector into an output vector containing 6 elements through a 128-row, 6-column connection matrix and a bias vector containing 6 elements. The 6 elements output by the output layer correspond to the concentration indicators of benzene series organic pollutants, phenolic organic pollutants, polycyclic aromatic hydrocarbons, organochlorine pollutants, petroleum hydrocarbons, and total dissolved organic carbon, respectively.

9. The method according to claim 1, characterized in that, In step 3, the specific process of establishing classical domain matter elements and segmented domain matter elements is as follows: For each pollutant concentration indication value in the pollutant feature vector, determine the upper limit and lower limit of the value range of each pollutant concentration indication value under Class I, Class II, Class III, Class IV and Class V water quality. Take the water quality level as the object of the matter element, take each pollutant concentration indication value as the feature of the matter element, and take the value range of each feature under the corresponding level as the value range of the corresponding feature, to form 5 classical domain matter elements corresponding to 5 water quality levels; take the value range of each feature in the entire range of levels as the value range of the corresponding feature to form segmented domain matter elements.

10. The method according to claim 9, characterized in that, In step 3, the specific process of calculating the comprehensive correlation degree and determining the water environment quality assessment level is as follows: For each pollutant concentration indication value in the pollutant feature vector to be evaluated, calculate the classical domain distance between each pollutant concentration indication value and the corresponding feature value interval in the classical domain element, and the section domain distance between each pollutant concentration indication value and the corresponding feature value interval in the section domain element. If the classical domain distance is less than or equal to zero, the single-index correlation degree is equal to the negative number of the classical domain distance divided by half the length of the corresponding feature value interval in the classical domain element. If the classical domain distance is greater than zero, the single-index correlation degree is equal to the negative number of the quotient obtained by dividing the classical domain distance by the difference between the section domain distance and the classical distance. Sum the single-index correlation degrees of each pollutant concentration indication value and divide by the number of pollutant concentration indication values ​​to obtain the comprehensive correlation degree. Take the water quality level corresponding to the largest comprehensive correlation degree value as the water environment quality assessment level.