A method and system for detecting antibiotic pollution in water bodies based on spectral characteristics

By combining surface-enhanced Raman spectroscopy and redox potential data in a multidimensional feature fusion method, the problems of anti-interference and accuracy in antibiotic detection in complex aquatic environments were solved, and precise quantification of antibiotics was achieved.

CN121902066BActive Publication Date: 2026-05-26SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610361516.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-24
Publication Date
2026-05-26
Estimated Expiration
2046-03-24

AI Technical Summary

Technical Problem

Existing technologies for antibiotic detection in complex aquatic environments have weak anti-interference capabilities and low quantitative accuracy, making it difficult to meet the high requirements of emergency response.

Method used

By simultaneously acquiring surface-enhanced Raman spectroscopy and redox potential data, a feature set of characteristic peak intensity ratio and energy entropy is constructed. Combined with a support vector projection regression model, multi-dimensional feature fusion and nonlinear mapping are achieved.

Benefits of technology

It significantly enhances the noise resistance under complex matrix interference, enables accurate identification and quantification of antibiotics in sudden water pollution, and improves the accuracy and stability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121902066B_ABST
    Figure CN121902066B_ABST
Patent Text Reader

Abstract

This application provides a method and system for detecting antibiotic pollution in water bodies based on spectral characteristics, belonging to the field of water environment monitoring and analysis technology. First, this application simultaneously acquires the surface-enhanced Raman spectrum and redox potential time-series data of the water sample to be tested. Next, the ratio of the characteristic peak intensities of the target antibiotic is extracted from the spectrum to generate a first feature set. Simultaneously, wavelet packet decomposition is performed on the potential data, and energy entropy is calculated to construct a second feature set. Subsequently, the two feature sets are concatenated, and a mutual information maximization algorithm is used to filter out the target feature subset highly correlated with concentration. Finally, using a support vector projection regression model trained based on a structural risk minimization mechanism, and through kernel function mapping and hyperplane projection operations, the accurate detection of the target antibiotic concentration in the water sample is achieved. This application improves the anti-interference capability and quantitative accuracy of antibiotic detection in complex water body backgrounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of water environment monitoring and analysis technology, and in particular relates to a method and system for detecting antibiotic pollution in water bodies based on spectral characteristics. Background Technology

[0002] Spectral feature-based detection technology, due to its speed, non-destructive nature, and fingerprint recognition capabilities, shows broad application prospects in monitoring antibiotic pollution in water bodies, particularly in responding to sudden environmental pollution incidents and providing early warnings for drinking water source safety. It can achieve rapid on-site screening of trace antibiotics. This technology, by capturing specific Raman scattering or fluorescence signals of antibiotic molecules and combining them with chemometric algorithms, holds promise for replacing the cumbersome laboratory operations of traditional chromatography methods, meeting the needs of real-time online monitoring.

[0003] In existing technologies, surface-enhanced Raman spectroscopy (SERS) is commonly used for antibiotic detection in water. This typically involves directly reading spectral peak intensities or establishing the relationship between spectral signal and concentration using a simple linear regression model. Some improved methods attempt to introduce internal standards or preprocess the spectrum to reduce the impact of substrate signal fluctuations, aiming to extract effective antibiotic characteristic information in complex real-world aquatic environments.

[0004] However, existing technologies still face challenges in handling complex aquatic matrix interference. Relying solely on spectral intensity features is easily affected by changes in the redox environment and background noise in the water, resulting in low signal-to-noise ratios and poor reproducibility. Furthermore, traditional feature extraction and regression models struggle to fully exploit the nonlinear correlation between spectral and environmental parameters, exhibiting insufficient generalization ability under small sample conditions, leading to detection accuracy and stability that fail to meet the high requirements of emergency response. Therefore, existing technologies suffer from insufficient antibiotic detection due to weak anti-interference capabilities and low quantitative accuracy in complex aquatic environments. Summary of the Invention

[0005] The purpose of this application is to provide a method and system for detecting antibiotic pollution in water bodies based on spectral characteristics, in order to solve the problem of insufficient antibiotic detection in the prior art.

[0006] To address the aforementioned technical problems, in a first aspect, this application provides a method for detecting antibiotic pollution in water bodies based on spectral characteristics, comprising:

[0007] Simultaneously acquire the surface-enhanced Raman spectrum of the water sample to be tested, as well as the potential time series data of the redox potential of the water sample to be tested within a preset time window;

[0008] The positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities are determined by the peak-finding algorithm. The first feature set is generated based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions.

[0009] By performing wavelet packet decomposition on the potential time series data, sub-signals of different frequency bands are obtained, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all energy entropies.

[0010] The first feature set and the second feature set are concatenated by vectors to construct the target feature set. The target feature set is then filtered by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic.

[0011] The target feature subset is input as an input vector to the trained support vector projection regression model. By using kernel function mapping and hyperplane projection, the concentration value of the target antibiotic in the water sample to be tested is obtained. The support vector projection regression model is trained based on the structural risk minimization mechanism and is trained on training samples including historical input vectors and corresponding antibiotic concentrations.

[0012] Optionally, the method further includes:

[0013] Calculate the volatility of potential time series data;

[0014] Based on the preset correspondence between potential volatility and peak finding threshold, determine the target peak finding threshold corresponding to the volatility;

[0015] The positions of multiple Raman peaks and their corresponding characteristic peak intensities in the spectrum of the target antibiotic are determined using a peak-finding algorithm. Based on the ratios between the characteristic peak intensities corresponding to these multiple Raman peak positions, a first feature set is generated, including:

[0016] Based on the target peak finding threshold, the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities are determined by the peak finding algorithm. The first feature set is generated based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions.

[0017] Optionally, based on the target peak-finding threshold, a peak-finding algorithm is used to determine the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities. A first feature set is generated based on the ratios between the characteristic peak intensities corresponding to the multiple Raman peak positions, including:

[0018] Data points with signal intensities greater than the target peak-finding threshold are extracted from the spectrum as candidate peak points, and the horizontal coordinate values ​​corresponding to the candidate peak points are used as the candidate peak positions.

[0019] Calculate the absolute difference between all candidate peak positions and the corresponding preset standard peak positions, and determine the candidate peak positions whose absolute differences are within the preset error range as the Raman peak positions of the target antibiotic.

[0020] Extract the ordinate value of the candidate peak point corresponding to each Raman peak position from the spectrum as the characteristic peak intensity, and select the characteristic peak intensity with the largest value from all characteristic peak intensities as the first intensity value, and use the remaining characteristic peak intensities as the second intensity value.

[0021] Calculate the ratio of each second intensity value to the first intensity value, and form a first feature set based on all ratios.

[0022] Optionally, the method further includes:

[0023] Calculate the signal-to-noise ratio of the spectrum;

[0024] Based on the preset correspondence between signal-to-noise ratio and decomposition layer number, determine the target decomposition layer number corresponding to the signal-to-noise ratio;

[0025] By performing wavelet packet decomposition on the potential time series data, sub-signals of different frequency bands are obtained, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. Based on all energy entropies, a second feature set is constructed, including:

[0026] Based on the target decomposition level, wavelet packet decomposition is performed on the potential time series data to obtain sub-signals of different frequency bands, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all energy entropies.

[0027] Optionally, based on the target decomposition level, wavelet packet decomposition is performed on the potential time series data to obtain sub-signals of different frequency bands, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all energy entropies, including:

[0028] The potential time series data is decomposed hierarchically using orthogonal wavelet basis functions until the target decomposition level is reached, and sub-signals of different frequency bands are obtained.

[0029] The amplitude of all data points in the sub-signal of each frequency band is squared and summed to obtain the corresponding energy value. The energy entropy is obtained by calculating the ratio of the energy value to all energy values ​​and performing a logarithmic operation.

[0030] The second feature set is obtained by arranging and combining all energy entropies in order of frequency band.

[0031] Optionally, the first feature set and the second feature set are concatenated as vectors to construct a target feature set, and the target feature set is filtered by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic, including:

[0032] The first feature set and the second feature set are merged into a high-dimensional feature vector according to a preset order to obtain the target feature set;

[0033] Calculate the mutual information value between each feature dimension in the target feature set and the preset antibiotic concentration label;

[0034] A subset of target features is obtained by selecting all feature dimensions whose mutual information values ​​are greater than a preset correlation threshold from the target feature set and combining them.

[0035] Optionally, a subset of the target features is input as an input vector to a trained support vector projection regression model. By utilizing kernel function mapping and hyperplane projection, the concentration values ​​of the target antibiotic in the water sample are obtained, including:

[0036] By using the kernel function in the support vector projection regression model, the target feature subset is mapped to a high-dimensional feature space to obtain the corresponding high-dimensional mapping vector;

[0037] Calculate the distance value in the normal direction of the high-dimensional mapping vector projected onto the pre-defined target regression hyperplane in the support vector projection regression model, where the target regression hyperplane is constructed in the high-dimensional feature space based on the training samples;

[0038] Based on the distance value, the concentration of the target antibiotic in the water sample to be tested is calculated by utilizing the mapping relationship between the distance value obtained from the training samples and the antibiotic concentration.

[0039] Secondly, this application provides a water antibiotic pollution detection system based on spectral characteristics, comprising:

[0040] The acquisition module is used to simultaneously acquire the surface-enhanced Raman spectrum of the water sample to be tested, as well as the potential time series data of the redox potential of the water sample to be tested within a preset time window;

[0041] The generation module is used to determine the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities through a peak-finding algorithm, and to generate a first feature set based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions.

[0042] The generation module is also used to obtain sub-signals of different frequency bands by performing wavelet packet decomposition on potential time series data, and to calculate the energy entropy corresponding to the sub-signals of different frequency bands, and to construct a second feature set based on all energy entropies;

[0043] The filtering module is used to concatenate the first feature set and the second feature set into vectors to construct the target feature set, and to filter the target feature set by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic.

[0044] The input module is used to input the target feature subset as an input vector into the trained support vector projection regression model. By using kernel function mapping and hyperplane projection, the concentration value of the target antibiotic in the water sample to be tested is obtained. The support vector projection regression model is based on the structural risk minimization mechanism and is trained on training samples including historical input vectors and corresponding antibiotic concentrations.

[0045] Thirdly, this application provides an electronic device, comprising:

[0046] Memory, used to store computer programs;

[0047] A processor is configured to execute the computer program to implement the steps of the water antibiotic pollution detection method based on spectral characteristics as described in the first aspect above.

[0048] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the water antibiotic pollution detection method based on spectral characteristics as described in the first aspect above.

[0049] The method for detecting antibiotic pollution in water bodies based on spectral characteristics provided in this application simultaneously acquires surface-enhanced Raman spectroscopy and redox potential data, and constructs a first feature set using the ratio of characteristic peak intensities, effectively offsetting errors caused by substrate signal fluctuations. Simultaneously, it innovatively performs wavelet packet decomposition on the potential time-series data and calculates energy entropy, deeply mining a second feature set reflecting the complexity of the water body's chemical microenvironment in the time-frequency domain, introducing crucial environmental auxiliary information for antibiotic detection.

[0050] Subsequently, redundant noise was removed and multimodal joint features highly correlated with antibiotic concentration were retained through vector concatenation and mutual information maximization filtering. Finally, a support vector projection regression model based on structural risk minimization was combined, kernel function mapping was used to resolve nonlinear relationships, and hyperplane projection was used to improve generalization performance under small sample sizes. This multidimensional feature fusion strategy significantly enhances the method's robustness to noise under complex matrix interference, achieving accurate identification and quantification of antibiotics in sudden water pollution events. Therefore, this application effectively solves the problem of insufficient antibiotic detection caused by weak anti-interference ability and poor generalization performance in existing technologies.

[0051] Furthermore, this application quantifies the dynamic disturbance level of the aquatic environment by calculating the volatility of potential time series data, and establishes an adaptive mapping relationship with the peak-finding threshold accordingly, thereby realizing dynamic control of the spectral peak-finding process. When the redox environment of the aquatic body fluctuates drastically, the peak-finding threshold is automatically adjusted, effectively avoiding false peak misjudgment or weak signal omission caused by increased baseline noise, and ensuring the accuracy and reliability of the extracted Raman peak position and intensity characteristics.

[0052] This mechanism, which uses electrochemical environmental parameters to guide optical signal processing, improves the quality of feature extraction from the source and further enhances the adaptability of the detection method to complex and variable aquatic environments. Therefore, this application effectively solves the problem of inaccurate feature extraction due to fixed thresholds in environments with strong interference, which in turn affects the accuracy of antibiotic detection. Attached Figure Description

[0053] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 A schematic flowchart of a water antibiotic pollution detection method based on spectral characteristics provided in this application embodiment;

[0055] Figure 2 A flowchart illustrating a method for generating antibiotic concentration values ​​provided in an embodiment of this application;

[0056] Figure 3 A schematic flowchart of another method for detecting antibiotic pollution in water based on spectral characteristics provided in this application embodiment;

[0057] Figure 4 A schematic diagram of a water antibiotic pollution detection system based on spectral characteristics provided in this application embodiment;

[0058] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0059] In the field rapid detection technology of antibiotic pollution in water bodies, although the existing methods have achieved fingerprint recognition by using surface-enhanced Raman spectroscopy, their detection mode, which relies solely on spectral intensity, has obvious defects when facing complex water bodies: the drastic redox environment fluctuations and background noise in actual water bodies seriously interfere with the stability of spectral signals, resulting in a reduced signal-to-noise ratio and poor reproducibility.

[0060] Meanwhile, traditional linear regression models struggle to capture the complex nonlinear relationships between spectral and environmental parameters, and their generalization ability is insufficient under the small sample conditions often encountered in emergency detection. This results in the accuracy and stability of the detection results failing to meet the high standards required for handling sudden pollution incidents. This bottleneck stems from the insufficient physical shielding capability of existing technologies against environmental interference factors and the limitations of data mining depth, necessitating a novel multidimensional fusion detection method that is resistant to environmental interference and possesses strong robustness.

[0061] To address the aforementioned issues, this application proposes a method for detecting antibiotic pollution in water bodies based on spectral characteristics. The core of this method lies in achieving precise analysis of antibiotic pollution through cross-modal fusion of spectral fingerprints and electrochemical environmental characteristics. Specifically, this method first simultaneously acquires time-series data of surface-enhanced Raman spectra and redox potentials of the water sample to be tested. On the one hand, it eliminates substrate fluctuation interference by calculating the ratio of characteristic peak intensities; on the other hand, it extracts energy entropy features reflecting the complexity of the microscopic chemical environment from the potential data using wavelet packet decomposition.

[0062] Subsequently, the most discriminative subset of joint features is selected by maximizing mutual information and input into a support vector projection regression model based on structural risk minimization for nonlinear mapping and quantitative inversion. This method overcomes the limitations of single-spectral detection by introducing electrochemical time-frequency features as environmental correction factors and combining them with a projection regression algorithm with high generalization ability, ensuring accurate identification of trace antibiotics even in complex matrix backgrounds.

[0063] This method overcomes signal distortion caused by environmental fluctuations, solves the problem of model overfitting under small sample size, and addresses the detection deficiencies caused by weak anti-interference ability and low quantitative accuracy of existing technologies, thus significantly improving the reliability and intelligence level of emergency response to water pollution.

[0064] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0065] To address the problems of existing technologies, embodiments of this application provide a method, apparatus, device, computer storage medium, and computer program product for detecting antibiotic pollution in water bodies based on spectral characteristics. The method for detecting antibiotic pollution in water bodies based on spectral characteristics provided in this application embodiment will be described first below.

[0066] Figure 1A schematic flowchart of a water antibiotic pollution detection method based on spectral characteristics, according to an embodiment of this application, is shown. Figure 1 As shown, the method includes:

[0067] S101. Simultaneously acquire the surface-enhanced Raman spectrum of the water sample to be tested, as well as the potential time series data of the redox potential of the water sample to be tested within a preset time window.

[0068] The water sample to be tested refers to a water sample containing potential antibiotic contamination collected from the site of a sudden environmental pollution incident. The surface-enhanced Raman spectroscopy spectrum refers to the spectral data acquired using surface-enhanced Raman scattering technology, including the relationship between Raman scattering intensity and Raman frequency shift. The redox potential time-series data refers to the data sequence of the redox potential of the water sample to be tested changing over time, continuously recorded within a preset time window; this sequence reflects the dynamic redox activity of the water's chemical environment.

[0069] Specifically, firstly, a portable sampling device is used to obtain water samples from the site of the sudden pollution incident, and these samples are then mixed with a pre-fabricated nano-reinforced substrate in a specific ratio. Next, a portable Raman spectrometer is used to perform laser-excited scanning on the mixed sample, acquiring surface-enhanced Raman scattering signals to obtain a spectrum including antibiotic characteristic fingerprints. ,in Indicates Raman frequency shift, Indicates intensity.

[0070] Simultaneously, the oxidation-reduction potential sensor is inserted into the water sample to be tested, and a preset time window is set. Within this time window, the redox potential values ​​of the water body are continuously collected at a fixed sampling frequency to obtain potential time-series data reflecting the dynamic changes in the water body's chemical environment. ,in This represents the potential value at different times.

[0071] S102. Determine the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities using a peak-finding algorithm. Generate a first feature set based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions.

[0072] Optionally, the method further includes:

[0073] Calculate the volatility of potential time series data.

[0074] The volatility of potential time series data refers to a statistical index used to quantify the dispersion or oscillation amplitude of redox potential values ​​within a preset time window. This index reflects the stability of the chemical environment of the water sample during the detection period. A larger volatility value indicates that the redox reaction in the water is more intense or that there is unstable electrochemical interference. In the specific embodiments of this application, the standard deviation is used as a quantitative index to characterize the volatility of potential time series data.

[0075] Specifically, based on the potential timing data obtained in step S101 The standard deviation method is used to quantify its fluctuation characteristics. First, the arithmetic mean of all potential values ​​within the time window is calculated. Then, the potential value at each sampling point is calculated. Compared with the average The sum of squares of the differences. Finally, divide this sum of squares by the number of sampling points. The volatility is obtained by taking the square root. The specific calculation process is shown in the following formula (1):

[0076] (1)

[0077] Based on the preset correspondence between potential volatility and peak finding threshold, the target peak finding threshold corresponding to the volatility is determined.

[0078] The target peak threshold refers to the minimum intensity limit or signal-to-noise ratio limit used in spectral data processing to determine whether a local signal bulge is a valid Raman characteristic peak.

[0079] The preset correspondence between potential fluctuation rate and peak-finding threshold refers to a pre-established lookup rule or mapping table. This relationship is based on physical laws: when the electrochemical environment fluctuates greatly, the spectral baseline noise is usually high, requiring a higher threshold to suppress false peaks; when the environment is stable, the noise is low, and a lower threshold can be used to identify weak signals. The preset correspondence between potential fluctuation rate and peak-finding threshold is based on the baseline noise standard deviation under pure water conditions. As shown in Table 1 below:

[0080] Table 1: Correspondence between preset potential volatility and peak finding threshold

[0081]

[0082] in, In order to be in Standard deviation of potential fluctuations continuously measured in deionized water. This represents the average noise intensity of the spectrometer under dark current. As shown in Table 1, this correspondence divides the potential volatility into different intervals and matches a corresponding target peak-finding threshold for each interval. For example, when the volatility is less than... At that time, the environment was considered extremely stable, and a low threshold was set. To preserve subtle characteristics; when volatility is to When in between, set a medium threshold. When volatility is greater than or equal to At that time, if environmental interference was deemed severe, a higher threshold was set. To filter noise.

[0083] Specifically, based on the calculated volatility Based on the correspondence shown in Table 1, assuming volatility Falling in the range Therefore, the corresponding target peak-finding threshold is determined. for .

[0084] Step S102 uses a peak-finding algorithm to determine the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities. Based on the ratios between the characteristic peak intensities corresponding to the multiple Raman peak positions, a first feature set is generated, including:

[0085] Based on the target peak finding threshold, the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities are determined by the peak finding algorithm. The first feature set is generated based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions.

[0086] Peak-finding algorithms are computer programs or mathematical methods used to automatically identify local maxima in a spectral data sequence and extract their location and intensity information. The first feature set refers to the data set consisting of the ratios of the intensities of multiple Raman characteristic peaks.

[0087] Specifically, firstly, the determined target peak threshold is... As a key parameter input into the peak-finding algorithm, the spectrum is analyzed. Perform a scan. Next, traverse the spectral data, only selecting signals with intensity exceeding [a certain threshold]. Furthermore, peaks that meet the peak shape criteria are identified as valid peaks, thereby pinpointing the locations of multiple Raman peaks of the target antibiotic. and the corresponding characteristic peak intensity Next, the intensity of a specific characteristic peak is divided by the intensity of the main peak to calculate a series of intensity ratios. Finally, the vectors containing these intensity ratios are combined to form the first feature set. .

[0088] This embodiment quantifies the dynamic disturbance level of the aquatic environment by calculating the volatility of potential time-series data, and establishes an adaptive mapping relationship with the peak-finding threshold accordingly, thereby realizing dynamic control of the spectral peak-finding process. This effectively avoids false peak misjudgment or missed weak signals caused by increased baseline noise, thus significantly improving the accuracy of antibiotic detection in complex and variable aquatic environments.

[0089] Optionally, based on the target peak-finding threshold, a peak-finding algorithm is used to determine the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities. A first feature set is generated based on the ratios between the characteristic peak intensities corresponding to the multiple Raman peak positions, including:

[0090] Data points with signal intensity greater than the target peak-finding threshold are extracted from the spectrum as candidate peak points, and the horizontal coordinate value corresponding to the candidate peak point is used as the candidate peak position.

[0091] Candidate peaks are potential signal points in a spectral data sequence whose signal intensity exceeds the background noise suppression line and exhibits a convex shape within a local area. The position of a candidate peak refers to the Raman frequency shift value corresponding to the aforementioned potential signal point on the horizontal axis of the spectrum, usually expressed in wavenumber.

[0092] Specifically, firstly, based on the spectrum obtained in step S101... and the determined target peak finding threshold The spectral data is traversed using a sliding window scan or the first derivative zero-crossing method. Secondly, all intensity values ​​greater than [a certain value] are identified. Local maxima are identified and marked as candidate peaks. Finally, the x-coordinates of these points are used as the set of candidate peak locations. .

[0093] For example, suppose the spectrum Includes multiple data points, target peak threshold for After scanning, the Raman frequency shift was found to be... , and The data point intensities at are respectively , and And all three intensity values ​​are greater than Then , , Record the candidate peak position.

[0094] Calculate the absolute difference between the positions of all candidate peaks and the corresponding preset standard peak positions, and determine the positions of candidate peaks whose absolute differences are within the preset error range as the Raman peak positions of the target antibiotic.

[0095] The preset standard peak position refers to the wavenumber position of the inherent Raman characteristic peak measured under standard experimental conditions based on pure antibiotics. It serves as a fingerprint for identifying antibiotic types. The preset error range refers to the allowable tolerance interval that takes into account minor shifts in spectral peak positions caused by spectrometer hardware resolution drift, changes in ambient temperature, or water matrix effects.

[0096] The Raman peak position of the target antibiotic refers to the actual signal peak position confirmed as belonging to the antibiotic after comparison and verification with the standard library. The preset standard peak positions are set according to different target antibiotics, as shown in Table 2 below. The preset error range is set according to different detection modes, as shown in Table 3 below.

[0097] Table 2: Reference Table of Preset Standard Peak Positions

[0098]

[0099] As shown in Table 2, Table 2 stores the theoretical fingerprint peak positions of different target antibiotics. For example, for quinolone A antibiotics, it has three key characteristic peaks, located at... , and These values ​​serve as a benchmark for eliminating irrelevant impurity peaks from the candidate peaks.

[0100] Table 3: Preset Error Range Comparison Table

[0101]

[0102] As shown in Table 3, Table 3 defines the peak position matching tolerance under different detection modes. For example, in the standard mode, an error range is allowed between the measured peak position and the standard peak position. The offset within this range ensures that even with slight red or blue shifts in the peak position, it can still be correctly identified in complex field environments.

[0103] Specifically, first, traverse the set of candidate peak positions. For each element in the table, calculate the absolute value of the difference between it and the positions of all preset standard peaks of the target antibiotic in Table 2. Then, check if any absolute value of the difference is less than or equal to the preset error range determined in Table 3. If a match is found, the candidate peak position is determined to be a successful match and identified as the Raman peak position of the target antibiotic; if the absolute value of the difference between all standard peak positions is greater than 1, the candidate peak position is determined to be a different position. If so, the candidate peak is considered an interfering peak and is removed.

[0104] For example, suppose the target antibiotic to be tested is quinolone A, and its set of standard peak positions in Table 2 is as follows: Candidate peak location set The allowable error range corresponding to the current detection mode is: Then regarding the candidate peak position... Calculate separately , and Assume the calculation reveals... Then confirm Matched the standard main peak Record it as a valid Raman peak location. .

[0105] Similarly, for candidate peak positions Continue calculating its distance from the three standard peaks, assuming that it is found Then confirm Matched standard secondary peak Recorded as Finally, a set of verified Raman peak positions were determined. .

[0106] The ordinate values ​​of the candidate peak points corresponding to the position of each Raman peak are extracted from the spectrum as characteristic peak intensities. The characteristic peak intensity with the largest value is selected as the first intensity value, and the remaining characteristic peak intensities are used as the second intensity values.

[0107] The intensity of a characteristic peak refers to the signal amplitude of the effective Raman peak on the vertical axis of the spectrum after position verification. The first intensity value is the peak with the strongest signal among all effective characteristic peaks, which usually corresponds to the vibrational mode with the largest Raman scattering cross-section in the antibiotic molecule. The second intensity value refers to the intensity values ​​of the other effective characteristic peaks besides the first intensity value.

[0108] Specifically, firstly, based on the determined location of Raman Peak... Retrospectral diagram Read the ordinate values ​​corresponding to these positions to obtain the set of characteristic peak intensities. Then, for the set The values ​​in the range are compared, and the maximum value is selected as the first intensity value. The remaining values ​​are used to form a second set of intensity values. For example, suppose the validated set of characteristic peak intensities is... ,in The value is greater than Comparison shows that... The maximum value is used to determine the first intensity value. The remaining This is the second intensity value, at which point... .

[0109] Calculate the ratio of each second intensity value to the first intensity value, and form a first feature set based on all ratios.

[0110] Specifically, first, traverse the set of second intensity values. For each element in the equation, divide it by the first intensity value. First, dimensionless intensity ratios are obtained. Second, all calculated ratios are arranged in order of their corresponding peak positions to construct a first feature set in vector form. .

[0111] This embodiment dynamically suppresses environmental noise interference by establishing an adaptive correlation between potential fluctuation rate and spectral peak-finding threshold, and uses a standard fingerprint database and error tolerance mechanism to perform secondary verification on candidate peaks, eliminating spurious peaks. Simultaneously, the intensity ratio feature constructed based on the self-internal standard method effectively eliminates errors caused by SERS substrate enhancement inhomogeneity and laser power fluctuations, significantly improving the accuracy of antibiotic identification.

[0112] S103. By performing wavelet packet decomposition on the potential time series data, sub-signals of different frequency bands are obtained, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all energy entropies.

[0113] Optionally, the method further includes:

[0114] Calculate the signal-to-noise ratio of the spectrum.

[0115] The signal-to-noise ratio (SNR) of a spectrum refers to the ratio of the intensity of the characteristic signal of a target antibiotic to the intensity of background noise in a surface-enhanced Raman spectrum. It is a key indicator for evaluating the quality of spectral data. This indicator can indirectly reflect the degree of interference in the current aquatic environment. The lower the SNR, the more interference there is from suspended particles, fluorescent substances, etc., and the greater the environmental noise; conversely, it indicates that the water matrix is ​​relatively clean.

[0116] Specifically, firstly, based on the spectrum obtained in step S101... A blank band without Raman characteristic peaks is selected as the noise estimation region, and the standard deviation of the signal intensity within this region is calculated as the noise level. Simultaneously, the height of the strongest characteristic peak identified in the spectrum is taken as the signal level. Then, using the formula The signal-to-noise ratio of the calculated spectrum .

[0117] Based on the preset correspondence between signal-to-noise ratio and decomposition layer number, determine the target decomposition layer number corresponding to the signal-to-noise ratio.

[0118] The target decomposition level refers to the depth level at which the signal is split in the time-frequency domain when performing wavelet packet decomposition on potential time-series data. The preset signal-to-noise ratio (SNR) and decomposition level correspondence refers to an adaptive parameter adjustment rule. This rule is based on the logic that the more complex the environment, the finer the decomposition needs to be. When the spectral SNR is low, i.e., when environmental interference is high, the decomposition level is increased to separate noise and micro-features in a higher frequency subspace; when the SNR is high, the level is reduced to reduce computational redundancy. The preset SNR and decomposition level correspondence is shown in Table 4 below:

[0119] Table 4: Preset Relationship between Signal-to-Noise Ratio and Number of Decomposition Layers

[0120]

[0121] As shown in Table 4, this correspondence divides the signal-to-noise ratio (SNR) into several intervals and specifies the optimal number of wavelet packet decomposition layers for each interval. For example, when the SNR is extremely low, i.e., less than... At that time, set a higher number of decomposition layers. When the signal-to-noise ratio is moderate, that is... arrive In between, set a medium number of layers. When the signal-to-noise ratio is good, that is, greater than At that time, set a lower number of floors. .

[0122] Specifically, based on the spectral signal-to-noise ratio calculated in the previous step Search in Table 4. Assume the signal-to-noise ratio... Falling in the range Then, based on the correspondence shown in Table 4, the target decomposition layer for processing potential data is determined. for .

[0123] Step S103 involves performing wavelet packet decomposition on the potential time series data to obtain sub-signals of different frequency bands, calculating the energy entropy corresponding to the sub-signals of different frequency bands, and constructing a second feature set based on all energy entropies, including:

[0124] Based on the target decomposition level, wavelet packet decomposition is performed on the potential time series data to obtain sub-signals of different frequency bands, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all energy entropies.

[0125] Wavelet packet decomposition is a sophisticated time-frequency analysis method that decomposes a signal into a series of sub-band signals with different frequency ranges to reveal the microscopic chemical reaction kinetics hidden in potential fluctuations. Energy entropy refers to the measure of the disorder or complexity of the energy distribution of each sub-band signal based on information entropy theory. The second feature set refers to the vector data formed by sequentially concatenating the calculated energy entropies of each frequency band according to frequency from low to high or the index order of wavelet packet nodes.

[0126] Specifically, firstly, using selected orthogonal wavelet basis functions such as the Daubechies wavelet, the decomposition is performed according to the determined target number of layers. The potential timing data obtained in step S101 Perform hierarchical recursive decomposition. The decomposition process will generate... The system identifies sub-signal nodes in different frequency bands. Next, for each sub-signal node, its signal energy is calculated, and the proportion of that node's energy in the total energy is determined. Based on these proportions, Shannon entropy or norm entropy is calculated, resulting in a set of energy entropies. Finally, these energy entropies are arranged in frequency band order, forming the second feature set. .

[0127] This embodiment achieves cross-modal collaboration between optical sensing and electrochemical processing, ensuring that the extracted energy entropy features include rich chemical kinetic information while avoiding noise contamination and computational waste, thus significantly improving the efficiency and accuracy of multi-dimensional feature fusion detection.

[0128] Optionally, based on the target decomposition level, wavelet packet decomposition is performed on the potential time series data to obtain sub-signals of different frequency bands, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all energy entropies, including:

[0129] The potential time series data is decomposed hierarchically using orthogonal wavelet basis functions until the target decomposition level is reached, resulting in sub-signals in different frequency bands.

[0130] Orthogonal wavelet basis functions refer to a set of orthogonal functions used in wavelet transform to construct a basis. They can generate a series of mutually orthogonal basis functions through translation and scaling.

[0131] Hierarchical iterative decomposition refers to the process of wavelet packet transform. In each level of decomposition, not only are the low-frequency components (approximate coefficients) decomposed, but the high-frequency components (detail coefficients) are also further decomposed, thus forming a complete binary tree structure for frequency band subdivision. The sub-signals of different frequency bands are the time-series coefficients included in the leaf nodes of the target layer at the end of the binary tree.

[0132] Specifically, firstly, the Daubechies wavelet system, which has compact support and good regularity, is selected as the basis function, for example... Wavelet. Its scaling function and wavelet function Satisfies the two-scale equation: , .in, These are the coefficients of the low-pass filter. These are the coefficients of the high-pass filter. This is a discrete-time index.

[0133] Secondly, based on this orthogonal wavelet basis function, the potential time series data obtained in step S101 are processed. Perform a fast Mallat algorithm decomposition. In the first layer, Through the filter and Low-frequency approximate components were obtained respectively. and high-frequency detail components On the second layer, for and The same filtering and downsampling operations were performed again to obtain four sub-components. .

[0134] Finally, this process is iterated until the target decomposition level is reached. Output the first The sequence of all node coefficients of the layer is... Sub-signals in different frequency bands. For example, assume the target decomposition layer number... After three iterations, the potential timing data is decomposed into Each sub-signal is labeled as follows: to .

[0135] The amplitude of all data points in the sub-signal of each frequency band is squared and summed to obtain the corresponding energy value. The energy entropy is obtained by calculating the ratio of the energy value to all energy values ​​and performing a logarithmic operation.

[0136] Energy value refers to the total signal power integral of a sub-signal in a certain frequency band, which reflects the weight of that frequency band in the overall signal composition.

[0137] Specifically, for the first Sub-signals of each frequency band First, calculate its energy value. Next, the sum of energy values ​​for all frequency bands is calculated. Then, calculate the first... Energy percentage probability of each frequency band Finally, the corresponding energy entropy is calculated using the standard Shannon entropy formula. The calculation formula is: It should be noted that, in order to preserve the independent characteristics of each frequency band, the weighted logarithmic term of each frequency band is used as the energy entropy of that frequency band.

[0138] The second feature set is obtained by arranging and combining all energy entropies in order of frequency band.

[0139] Specifically, the calculated values ​​corresponding to each sub-signal Energy entropy The sub-signals are arranged according to their frequency index order to construct a one-dimensional feature vector, which serves as the final second feature set. .

[0140] This embodiment achieves cross-modal collaboration between optical sensing and electrochemical processing. It ensures that the extracted energy entropy features include rich chemical kinetic information while avoiding noise contamination and wasted computing power, significantly improving the efficiency and accuracy of multi-dimensional feature fusion detection.

[0141] S104. The first feature set and the second feature set are concatenated into vectors to construct the target feature set. The target feature set is then filtered by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic.

[0142] Optionally, step S104, which involves concatenating the first and second feature sets into vectors to construct a target feature set, and then using mutual information to maximize feature maximization to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic, may specifically include:

[0143] S1041. Merge the first feature set and the second feature set into a high-dimensional feature vector according to a preset order to obtain the target feature set.

[0144] The target feature set refers to a multi-dimensional joint vector formed by physically fusing a first feature set including spectral fingerprint information and a second feature set including electrochemical environmental information.

[0145] Specifically, according to the data generated in step S102, including The first feature set of each intensity ratio element and the product generated in step S103 including The second characteristic set of energy entropy elements Following preset splicing rules, such as first spectroscopy then electrochemistry, the data is connected end-to-end. The merging operation generates a data set with a length of [length missing]. The high-dimensional feature vectors are the target feature set. .

[0146] S1042. Calculate the mutual information value between each feature dimension in the target feature set and the preset antibiotic concentration label.

[0147] Mutual information is a metric in information theory used to measure the degree of statistical dependence between two random variables. In this application, it quantifies how much information about the true concentration of the target antibiotic is contained in each individual feature dimension of the target feature set, such as a spectral ratio or potential energy entropy. A higher mutual information value indicates that the feature is more important for predicting antibiotic concentration; a mutual information value of zero indicates that the feature is completely irrelevant to concentration, i.e., noise. The preset antibiotic concentration label refers to the true concentration value of the samples known during the model training phase.

[0148] Specifically, for the target feature set Each feature dimension and the corresponding antibiotic concentration label First, using the formula Calculate the mutual information value between the two. The natural logarithm is used here, and the unit of information measurement is nat. Representing characteristic variables With concentration variables The joint probability distribution, and These represent the marginal probability distributions of the feature variable and the concentration variable, respectively.

[0149] This formula uses summation in the discrete case, while in the case of continuous variables it is usually converted into an integral form or estimated using the k-nearest neighbor algorithm. Secondly, by traversing all dimensions, a mutual information vector is obtained. .

[0150] S1043. By selecting all feature dimensions with mutual information values ​​greater than a preset correlation threshold from the target feature set, a subset of target features is obtained.

[0151] The target feature subset refers to the streamlined set of features retained after screening and deemed to substantially contribute to the detection results. The preset relevance threshold is the minimum mutual information limit used to determine the effectiveness of a feature. This threshold is set based on the accuracy requirements and noise tolerance of the detection task; features below this threshold are considered redundant information or interfering noise and are discarded. It is determined based on the significance level of the mutual information distribution between features and random noise, as shown in Table 5 below:

[0152] Table 5: Preset Correlation Threshold Comparison Table

[0153]

[0154] As shown in Table 5, recommended threshold settings are provided for different pollution scenarios. For example, when processing highly complex wastewater samples, a higher threshold is set to filter out a large amount of background noise features. When processing cleaner samples, a lower threshold is set in order to preserve weak auxiliary information. .

[0155] Specifically, the mutual information vector calculated in step S1042 Each value in the table corresponds to a preset correlation threshold determined according to Table 5. Compare the values. Retain those with mutual information values ​​greater than [value missing]. The feature dimensions are extracted and recombine in their original order to form a subset of the target features. .

[0156] For example, assuming the current scenario has medium complexity, the threshold is selected according to Table 5. Target feature set Corresponding mutual information vector By comparing the hypotheses, we found Greater than Therefore, from Extract Combine to obtain a subset of target features .

[0157] This embodiment quantifies the contribution of each feature dimension to antibiotic concentration, eliminates redundant noise and invalid variables based on an adaptive threshold, and finally generates a simplified target feature subset that retains key physicochemical fingerprints while minimizing the risk of dimensionality curse, significantly improving the computational efficiency and accuracy of detection.

[0158] S105. Input the target feature subset as the input vector into the trained support vector projection regression model. By using kernel function mapping and hyperplane projection, the concentration value of the target antibiotic in the water sample to be tested is obtained. The support vector projection regression model is trained based on the structural risk minimization mechanism and is trained on training samples including historical input vectors and corresponding antibiotic concentrations.

[0159] Optionally, step S105, which inputs the target feature subset as an input vector into the trained support vector projection regression model and obtains the concentration value of the target antibiotic in the water sample by utilizing kernel function mapping and hyperplane projection, may specifically include:

[0160] Figure 2 A flowchart illustrating a method for generating antibiotic concentration values ​​according to an embodiment of this application is shown. Figure 2 As shown, the method includes:

[0161] S1051. Using the kernel function in the support vector projection regression model, the target feature subset is mapped to a high-dimensional feature space to obtain the corresponding high-dimensional mapping vector.

[0162] A high-dimensional feature space refers to a reproducing kernel Hilbert space to which the original low-dimensional data is mapped through nonlinear transformations, in which samples that were originally linearly inseparable become linearly separable. A high-dimensional mapping vector is the implicit representation of the target feature subset in this high-dimensional space.

[0163] Specifically, first, a training sample set is constructed. Standard antibiotic aqueous solutions are prepared using a gradient dilution method, covering a concentration range that... Prepared A total of 10 training samples were used. Simultaneously, different proportions of humic acid and kaolin were added to the samples to simulate real-world water matrix interference. The spectral feature set of each sample was obtained. and electrochemical feature set As historical input vector And use the configured concentration as the real label. .

[0164] Subsequently, build based on - Support vector regression optimization objective with insensitive loss function. Based on the principle of structural risk minimization, optimize the objective function. As shown in formula (2) below:

[0165]

[0166] in, For the weight vector, The penalty coefficient is... As slack variables, The threshold for insensitive loss is set. Finally, the sequential minimum optimization algorithm is used to solve the above convex quadratic programming problem to obtain the optimal Lagrange multipliers. and bias This allows us to determine the final regression decision function.

[0167] Subsequently, the target feature subset output in step S104 is... as input vector Substitute this into the pre-defined radial basis function (RBF) within the support vector projection regression model. The RBF kernel function typically takes the form of... ,in For the kernel width parameter, five-fold cross-validation was used. A grid search is performed within the specified range to determine the location. These are the support vectors stored in the model. Logically, the mapping from the input space to the feature space is completed through computation; in the actual algorithm, only the kernel function, i.e., the high-dimensional vector, needs to be calculated. The inner product without explicit derivation The specific coordinates.

[0168] S1052. Calculate the distance value in the normal direction of the high-dimensional mapping vector projected onto the pre-defined target regression hyperplane in the support vector projection regression model, wherein the target regression hyperplane is constructed in the high-dimensional feature space based on the training samples.

[0169] The target regression hyperplane refers to the hyperplane formed by trained weight vectors in a high-dimensional feature space. and bias terms The defined decision surface. The projection distance value refers to the algebraic distance of the vector points of the target feature subset after kernel function mapping, projected onto the hyperplane along the normal vector direction in high-dimensional space. In this application, the projection distance value is mathematically equivalent to the output prediction value of the regression model, that is, the stability of geometric projection is used to characterize the response intensity of antibiotic concentration. The preset target regression hyperplane parameters are shown in Table 6 below:

[0170] Table 6: Preset Target Regression Hyperplane Parameters

[0171]

[0172] As shown in Table 6, Table 6 stores the core model parameters trained for different antibiotic types. For each type of antibiotic, the model corresponds to a specific set of support vector coefficient vectors. and bias terms These parameters uniquely determine the target regression hyperplane in the high-dimensional space.

[0173] Specifically, firstly, based on the preliminarily identified antibiotic types in the water sample, such as those determined through comparison with a spectral fingerprint database, the corresponding model parameters are indexed from Table 6. Then, a regression decision function is used. Calculate distance value .in, It is the number of support vectors. These are the Lagrange multipliers, i.e., the support vector coefficients. These are the training sample labels. It is the kernel function value calculated in step S1051.

[0174] S1053. Using the mapping relationship between the distance value obtained from the training samples and the antibiotic concentration, the concentration value of the target antibiotic in the water sample to be tested is calculated.

[0175] The mapping relationship between distance values ​​and antibiotic concentration refers to the inversion function that converts the dimensionless projected distance output by the model into specific physical concentration units. Since the decision values ​​output in support vector regression have already fitted the concentration trend with the training samples, this relationship is a denormalization process.

[0176] Specifically, the distance value calculated in step S1052 Substitute into the preset inversion formula Among them, the inversion slope and intercept During the model training phase, the known concentration labels of the training sample set are fitted with the projected distance values ​​generated by the model using least squares linear regression, which is predetermined. The calculated results... This is the final concentration value of the target antibiotic in the water sample to be tested.

[0177] This embodiment effectively solves the complex nonlinear coupling problem between spectral and electrochemical data, not only achieving accurate quantification of the antibiotic to be tested, but also exhibiting strong generalization ability and anti-overfitting performance under small sample training conditions, significantly improving the reliability of quantitative analysis in emergency detection scenarios.

[0178] Figure 3 This illustration shows a flowchart of another method for detecting antibiotic pollution in water based on spectral characteristics, provided in one embodiment of this application. Figure 3 As shown, firstly, data acquisition and preprocessing are performed in step S1, covering the acquisition of spectral data of water samples and basic processing work such as denoising and baseline correction. Subsequently, in step S2, key features are extracted from the spectral data and dimensionality reduction methods are applied to extract the feature information that best characterizes the pollution situation.

[0179] Next, in step S3, a suitable prediction model is selected for training and optimization, and model parameters are adjusted using historical sample data. In step S4, the model performance is evaluated and validated using a test set to ensure that the accuracy and recall of the detection results meet practical requirements. Finally, in step S5, the mature model is deployed to the online monitoring system to achieve automated real-time analysis and early warning of antibiotic pollution in water bodies.

[0180] Figure 4 This is a schematic diagram illustrating a specific implementation of a water antibiotic pollution detection system based on spectral characteristics, as provided in this application. (Refer to...) Figure 4 The system may include:

[0181] 410 Acquisition Module is used to simultaneously acquire the surface-enhanced Raman spectrum of the water sample to be tested, as well as the potential time series data of the redox potential of the water sample to be tested within a preset time window;

[0182] The 420 generation module is used to determine the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities through a peak-finding algorithm, and to generate a first feature set based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions.

[0183] The 420 generation module is also used to obtain sub-signals of different frequency bands by performing wavelet packet decomposition on potential time series data, and to calculate the energy entropy corresponding to the sub-signals of different frequency bands, and to construct a second feature set based on all energy entropies;

[0184] The 430 filtering module is used to concatenate the first feature set and the second feature set into vectors to construct the target feature set, and to filter the target feature set by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic.

[0185] The 440 input module is used to input the target feature subset as an input vector into the trained support vector projection regression model. By using kernel function mapping and hyperplane projection, the concentration value of the target antibiotic in the water sample to be tested is obtained. The support vector projection regression model is based on the structural risk minimization mechanism and is trained on training samples including historical input vectors and corresponding antibiotic concentrations.

[0186] The water antibiotic pollution detection system based on spectral features in this application is used to implement the aforementioned water antibiotic pollution detection method based on spectral features. Therefore, the specific implementation of the water antibiotic pollution detection system based on spectral features can be found in the embodiment section of the water antibiotic pollution detection method based on spectral features above. The specific implementation can be referred to the description of the corresponding embodiments, and will not be repeated here.

[0187] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application is shown.

[0188] The electronic device may include a processor 510 and a memory 520 storing computer program instructions.

[0189] Specifically, the processor 510 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0190] Memory 520 may include mass storage for data or instructions. For example, and not limitingly, memory 520 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 520 may include removable or non-removable (or fixed) media. Where appropriate, memory 520 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 520 is non-volatile solid-state memory.

[0191] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.

[0192] The processor 510 reads and executes computer program instructions stored in the memory 520 to implement any of the above embodiments of a water antibiotic pollution detection method based on spectral characteristics.

[0193] In one example, the electronic device may also include a communication interface 530 and a bus 540. Wherein, such as Figure 5 As shown, the processor 510, memory 520, and communication interface 530 are connected through bus 540 and complete communication with each other.

[0194] The communication interface 530 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0195] Bus 540 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 540 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0196] The electronic device can execute the water antibiotic pollution detection method based on spectral characteristics in the embodiments of this application, thereby realizing the water antibiotic pollution detection method based on spectral characteristics described in conjunction with the accompanying drawings.

[0197] Furthermore, in conjunction with the spectral feature-based water antibiotic pollution detection method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the spectral feature-based water antibiotic pollution detection methods in the above embodiments.

[0198] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0199] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0200] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0201] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0202] The above provides a detailed description of a method and system for detecting antibiotic pollution in water based on spectral characteristics, as provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for detecting antibiotic pollution in water bodies based on spectral characteristics, characterized in that, The method includes: Simultaneously acquire the surface-enhanced Raman spectrum of the water sample to be tested, as well as the potential time series data of the redox potential of the water sample to be tested within a preset time window; The positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities are determined by a peak-finding algorithm. A first feature set is generated based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions. By performing wavelet packet decomposition on the potential time series data, sub-signals of different frequency bands are obtained, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all the energy entropies. The first feature set and the second feature set are concatenated by vectors to construct a target feature set. The target feature set is then filtered by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic. The target feature subset is used as an input vector to be fed into a trained support vector projection regression model. By using kernel function mapping and hyperplane projection, the concentration value of the target antibiotic in the water sample to be tested is obtained. The support vector projection regression model is trained based on a structural risk minimization mechanism and is trained on training samples including historical input vectors and corresponding antibiotic concentrations. The method further includes: Calculate the volatility of the potential time series data; Based on the preset correspondence between potential volatility and peak finding threshold, determine the target peak finding threshold corresponding to the volatility; The method involves determining the positions of multiple Raman peaks and their corresponding characteristic peak intensities in the spectrum of the target antibiotic using a peak-finding algorithm, and generating a first feature set based on the ratios between the characteristic peak intensities corresponding to the multiple Raman peak positions, including: Based on the target peak-finding threshold, the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities are determined by a peak-finding algorithm. A first feature set is generated based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions. The step involves determining the positions of multiple Raman peaks and their corresponding characteristic peak intensities of the target antibiotic in the spectrum using a peak-finding algorithm based on the target peak threshold, and generating a first feature set based on the ratios between the characteristic peak intensities corresponding to the multiple Raman peak positions, including: Data points with signal intensities greater than the target peak-finding threshold are extracted from the spectrum as candidate peak points, and the horizontal coordinate values ​​corresponding to the candidate peak points are used as candidate peak positions. Calculate the absolute difference between all the candidate peak positions and the corresponding preset standard peak positions, and determine the candidate peak positions whose absolute differences are within a preset error range as the Raman peak positions of the target antibiotic. The ordinate value of the candidate peak point corresponding to the position of each Raman peak is extracted from the spectrum as the characteristic peak intensity, and the characteristic peak intensity with the largest value is selected from all the characteristic peak intensities as the first intensity value, and the remaining characteristic peak intensities are used as the second intensity value. Calculate the ratio of each second intensity value to the first intensity value, and form the first feature set based on all the ratios.

2. The method according to claim 1, characterized in that, The method further includes: Calculate the signal-to-noise ratio of the spectrum; Based on the preset correspondence between signal-to-noise ratio and decomposition layer number, determine the target decomposition layer number corresponding to the signal-to-noise ratio; The process involves obtaining sub-signals of different frequency bands by performing wavelet packet decomposition on the potential time series data, calculating the energy entropy corresponding to the sub-signals of different frequency bands, and constructing a second feature set based on all the energy entropies, including: Based on the target decomposition level, wavelet packet decomposition is performed on the potential time series data to obtain sub-signals of different frequency bands, and the energy entropy corresponding to the sub-signals of different frequency bands is calculated. A second feature set is constructed based on all the energy entropies.

3. The method according to claim 2, characterized in that, The step involves obtaining sub-signals of different frequency bands by performing wavelet packet decomposition on the potential time series data according to the target decomposition level, calculating the energy entropy corresponding to the sub-signals of different frequency bands, and constructing a second feature set based on all the energy entropies, including: The potential time series data is decomposed hierarchically using orthogonal wavelet basis functions until the target decomposition level is reached, resulting in sub-signals of different frequency bands. The amplitude of all data points in the sub-signal of each frequency band is squared and summed to obtain the corresponding energy value. The energy entropy is obtained by calculating the ratio of the energy value to all the energy values ​​and performing a logarithmic operation. The second feature set is obtained by arranging and combining all the energy entropies in frequency band order.

4. The method according to claim 1, characterized in that, The step of concatenating the first feature set and the second feature set into vectors to construct a target feature set, and then performing feature filtering on the target feature set by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic, includes: The first feature set and the second feature set are merged into a high-dimensional feature vector according to a preset order to obtain the target feature set; Calculate the mutual information value between each feature dimension in the target feature set and the preset antibiotic concentration label; The target feature subset is obtained by selecting all feature dimensions whose mutual information value is greater than a preset correlation threshold from the target feature set and combining them.

5. The method according to claim 1, characterized in that, The step of inputting the target feature subset as an input vector into a trained support vector projection regression model, and obtaining the concentration value of the target antibiotic in the water sample by utilizing kernel function mapping and hyperplane projection, includes: Using the kernel function in the support vector projection regression model, the target feature subset is mapped to a high-dimensional feature space to obtain the corresponding high-dimensional mapping vector; Calculate the distance value between the high-dimensional mapping vector projected onto the normal vector direction of the preset target regression hyperplane in the support vector projection regression model, wherein the target regression hyperplane is constructed based on the training samples in the high-dimensional feature space; Based on the distance value, the concentration of the target antibiotic in the water sample to be tested is calculated by utilizing the mapping relationship between the distance value obtained from the training samples and the antibiotic concentration.

6. A water antibiotic pollution detection system based on spectral characteristics, characterized in that, include: The acquisition module is used to simultaneously acquire the surface-enhanced Raman spectrum of the water sample to be tested, as well as the potential time series data of the redox potential of the water sample to be tested within a preset time window; The generation module is used to determine the positions of multiple Raman peaks of the target antibiotic in the spectrum and the corresponding characteristic peak intensities through a peak-finding algorithm, and to generate a first feature set based on the ratio between the characteristic peak intensities corresponding to the multiple Raman peak positions. The generation module is also used to calculate the volatility of the potential time series data; Based on the preset correspondence between potential volatility and peak-finding threshold, a target peak-finding threshold corresponding to the volatility is determined; data points with signal intensities greater than the target peak-finding threshold are extracted from the spectrum as candidate peak points, and the abscissa values ​​corresponding to the candidate peak points are used as candidate peak positions; the absolute difference between all candidate peak positions and the corresponding preset standard peak positions is calculated, and the candidate peak positions with absolute differences within a preset error range are determined as the Raman peak positions of the target antibiotic; the ordinate values ​​of the candidate peak points corresponding to each Raman peak position are extracted from the spectrum as characteristic peak intensities, and the characteristic peak intensity with the largest value is selected from all characteristic peak intensities as the first intensity value, and the remaining characteristic peak intensities are used as the second intensity values; the ratio of each second intensity value to the first intensity value is calculated respectively, and a first feature set is formed based on all the ratios; the generation module is also used to obtain sub-signals of different frequency bands by performing wavelet packet decomposition on the potential time series data, and to calculate the energy entropy corresponding to the sub-signals of different frequency bands, and to construct a second feature set based on all the energy entropies; The filtering module is used to concatenate the first feature set and the second feature set into vectors to construct a target feature set, and to filter the target feature set by maximizing mutual information to generate a target feature subset that is statistically correlated with the concentration of the target antibiotic. The input module is used to input the target feature subset as an input vector into the trained support vector projection regression model. By using kernel function mapping and hyperplane projection, the concentration value of the target antibiotic in the water sample to be tested is obtained. The support vector projection regression model is trained based on the structural risk minimization mechanism and is trained on training samples including historical input vectors and corresponding antibiotic concentrations.

7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the method for detecting antibiotic pollution in water based on spectral characteristics as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the water antibiotic pollution detection method based on spectral characteristics as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and system for determining peak area of Raman spectrum characteristic peak without deducting Raman background

    CN116933056A

  • Spectral feature and artificial intelligence fused grain heavy metal detection method and system

    CN121275686A