Aquatic tail water antibiotic concentration inversion method based on spectrum-machine learning coupling
Through the spectral-machine learning coupling method, the representativeness and accuracy problems of antibiotic concentration monitoring in aquaculture effluent were solved, and efficient and accurate antibiotic concentration inversion was achieved, which adapted to complex aquaculture environments and ensured data security.
Patent Information
- Application Number
- CN202510725173.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies make it difficult to accurately monitor antibiotic concentrations in aquaculture effluents. Traditional methods lack representative sampling, have low spectral data processing accuracy, and weak model generalization capabilities, making them unable to meet the needs of complex and changing aquaculture environments.
A method based on spectral-machine learning coupling is adopted to obtain representative samples through multi-dimensional sampling. High-precision spectral instruments and adaptive noise removal models are used to construct an improved neural network model. Inversion is performed in combination with environmental correction factors, and intelligent sampling equipment and encrypted storage technology are used.
It achieves efficient and accurate monitoring of antibiotic concentrations in aquatic effluent, improves the representativeness of sample data and the comprehensiveness of analysis, enhances the learning efficiency and generalization performance of the model, and ensures the accuracy and security of the inversion results.
Smart Images

Figure CN120632298A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aquaculture antibiotic concentration supervision, and in particular to a method for inverting antibiotic concentration in aquaculture tail water based on spectral-machine learning coupling. Background Art
[0002] With the booming aquaculture industry, antibiotics are widely used to prevent and treat aquatic animal diseases. However, large amounts of unabsorbed antibiotics are discharged into surrounding water bodies with aquaculture tailwater, posing a serious threat to the ecological environment and human health. Accurately monitoring antibiotic concentrations in aquaculture tailwater is crucial for effectively controlling aquaculture pollution and maintaining ecological balance. Traditional tailwater antibiotic concentration monitoring methods are unable to meet the current complex and changing aquaculture environment and the high-precision monitoring requirements. Therefore, the development of innovative and efficient monitoring technologies is urgent.
[0003] Existing technologies have significant shortcomings in sample collection. Conventional sampling methods are often limited to specific areas and time periods, and the samples collected lack broad representativeness. Antibiotics are unevenly distributed in aquaculture waters across different regions due to differences in aquaculture density, water flow conditions, and other factors. Antibiotic concentrations also vary at different depths and time points, depending on factors such as water circulation and biodegradation. Limited samples make it difficult to fully reflect the true status of antibiotics in tailwater, resulting in significant deviations in subsequent analytical results and an inability to provide a reliable basis for pollution control.
[0004] There are also many problems with spectral data processing and concentration inversion models. Traditional spectral analysis instruments have poor precision and are easily affected by noise interference and baseline drift when collecting spectral data. The accuracy of the acquired data is greatly reduced, making it difficult to accurately extract spectral features related to antibiotic concentrations. In the concentration inversion process, the models relied on are mostly based on simple linear relationships and cannot fully consider the impact of complex environmental factors such as water temperature, salinity, and dissolved oxygen on the spectrum-antibiotic concentration relationship. These models have weak generalization capabilities. When faced with tail water from different aquaculture environments, the prediction results fluctuate greatly and have high errors, making it difficult to meet the accuracy and stability requirements in actual monitoring. For this reason, a method for inverting antibiotic concentrations in aquatic tail water based on spectral-machine learning coupling has emerged, aiming to break through the bottleneck of existing technologies and achieve efficient and accurate monitoring of antibiotic concentrations in aquatic tail water. Summary of the Invention
[0005] In order to overcome the shortcomings and deficiencies of the existing technology, the present invention provides an inversion method for antibiotic concentration in aquatic tail water based on spectral-machine learning coupling.
[0006] A method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling includes:
[0007] Step 1: Systematically collect aquatic tailwater samples from multiple aquaculture areas, at different depths, and at different time points, according to preset sampling specifications and standard procedures, to obtain representative aquatic tailwater samples and construct a sample dataset;
[0008] Step 2: Using a spectrum analyzer, under a set spectral range and preset measurement conditions, collect spectral data of each of the collected aquatic tail water samples to obtain original spectral data;
[0009] Step 3: Preprocess the raw spectral data to remove noise, baseline drift, and other abnormal fluctuations in the spectral data through a set filtering algorithm and data calibration method;
[0010] Step 4: Using the set feature extraction algorithm from the preprocessed spectral data, the spectral feature parameters that are correlated with the antibiotic concentration in aquatic tail water are mined to form a feature vector set;
[0011] Step 5: Build a machine learning model and select an appropriate machine learning algorithm architecture, including a neural network model based on a specific neuron connection method and activation function. Divide the feature vector set into a training set and a test set. Use the training set to perform multiple rounds of training and optimization on the machine learning model and adjust the model's internal parameters.
[0012] Step 6: Use the test set to evaluate the performance of the trained and optimized machine learning model. Quantify the accuracy and stability of the model based on the evaluation indicators. If the model performance does not meet the preset standards, return to step 5 and retrain and optimize until the model performance meets the requirements.
[0013] Step 7: Apply the verified machine learning model to the actual aquatic tailwater spectral data. Based on the output of the model and the pre-set concentration conversion rules, the antibiotic concentration in the aquatic tailwater is inverted.
[0014] Furthermore, in the preprocessing operation of step 3, an innovative adaptive noise removal model is adopted, and the model formula is: Among them, Y(i) is the value of the original spectrum data at the i-th sampling point, Y new (i) is the value of the spectral data at the i-th sampling point after noise removal, α is the adaptive adjustment coefficient, which ranges from [0.3, 0.7] and is dynamically adjusted according to the degree of fluctuation of the spectral data, and n is the neighborhood window size, which ranges from [3, 7] and is set according to the resolution of the spectral data.
[0015] Furthermore, in the feature extraction process of step 4, a new spectral feature enhancement model is introduced. The model formula is: Among them, F(k) is the kth spectral feature value originally extracted, F enhanced (k) is the enhanced spectral characteristic value, β is the characteristic weight coefficient, the value range is [0.6, 0.9], γ is the fluctuation enhancement coefficient, the value range is [0.1, 0.3], is the average value of all original extracted spectral feature values, and m is the total number of spectral features.
[0016] Furthermore, in the machine learning model constructed in step 5, an improved neural network structure is adopted, and the connection weight update formula between the hidden layer neurons is: in, is the updated connection weight, is the current connection weight, is the connection weight before the last update, η is the learning rate, and its value range is [0.001, 0.01], δ j is the error term of neuron j, O i is the output value of neuron i, λ is the weight attenuation coefficient, and its value range is [0.0001, 0.001].
[0017] Furthermore, in the model training process of step 5, a new training sample selection strategy is introduced. By calculating the feature discreteness of the sample and the similarity with other samples, samples with large feature discreteness and low similarity with the selected samples are preferentially selected for training. The sample selection formula is: Among them, S(k) is the selection priority of the kth sample, σ F(k) is the standard deviation of the kth sample feature, sim(F(k), F(l)) is the similarity between the kth sample and the lth sample feature, calculated using cosine similarity, and N is the total number of samples.
[0018] Furthermore, in the model performance evaluation in step 6, a comprehensive evaluation index is used, which combines the model's prediction error, model complexity, and the model's adaptability to samples in different concentration ranges. The evaluation index formula is:
[0019]
[0020] Among them, E is the comprehensive evaluation index value, is the predicted antibiotic concentration of the i-th sample, is the true antibiotic concentration of the i-th sample, M is the number of test set samples, ω1, ω2, ω3 are weight coefficients, and their value ranges are [0.4, 0.6], [0.2, 0.3], [0.2, 0.3], respectively, L is the number of model parameters, W jis the j-th model parameter, P is the number of intervals divided by antibiotic concentration, var represents variance, and mean represents mean.
[0021] Furthermore, in the concentration inversion process of step seven, considering the influence of different breeding environments on the spectrum-antibiotic concentration relationship, an environmental correction factor is introduced, and the inversion formula is: Among them, C final is the final inversion-derived antibiotic concentration in aquaculture tail water, C model is the antibiotic concentration output by the machine learning model, θ q is the correction coefficient of the qth environmental factor, and its value range is determined according to experimental data. q is the measured value of the qth environmental factor, Q is the number of environmental factors, and the environmental factors include water temperature, salinity, and dissolved oxygen.
[0022] Furthermore, in the entire inversion method, the spectral data and sample information are encrypted and stored using a new encryption algorithm. The encryption formula is: Among them, D encrypted (r) is the encrypted data, D(r) is the original data, is an XOR operation, H is a hash function, timestamp is the data collection timestamp, key is the encryption key, the key length is 128 bits and is changed regularly.
[0023] Furthermore, during the sample collection process in step 1, an intelligent sampling device is used. The device monitors the physical and chemical parameters of the aquaculture water in real time through built-in sensors and automatically adjusts the sampling position and time according to the pre-set sampling rules. The sampling rule formula is: T sample =T base +ΔT×f(pH, DO, Temperature), where T sample is the actual sampling time, T base is the basic sampling time, ΔT is the time adjustment step, f(pH, DO, Temperature) is the adjustment function determined according to the water pH, dissolved oxygen and temperature parameters, and the function form is obtained by fitting the experimental data.
[0024] A method for inverting antibiotic concentrations in aquatic tail water based on spectral-machine learning coupling is implemented through the following units:
[0025] The sample collection multi-element control unit is used to collect aquatic tail water samples under various environmental conditions according to a multi-dimensional sampling strategy, and is connected to the sample transmission buffer unit through a high-speed data transmission line to transmit the collected sample information to the sample transmission buffer unit;
[0026] The sample transmission buffer unit is used to temporarily store the sample information transmitted by the sample acquisition multi-element control unit, and to preliminarily organize and cache the information. It is connected to the spectral data acquisition and precision analysis unit through a stable data interface, and transmits the organized sample information to the spectral data acquisition and precision analysis unit in an orderly manner;
[0027] The spectral data acquisition and precision analysis unit is used to collect high-precision spectral data of samples under a strictly controlled spectral measurement environment, and perform preliminary analysis and processing on the original spectral data. It is connected to the spectral data preprocessing intelligent optimization unit through a dedicated data channel and transmits the pre-processed spectral data to the spectral data preprocessing intelligent optimization unit;
[0028] The spectral data preprocessing intelligent optimization unit is used to perform noise removal and baseline calibration preprocessing operations on the spectral data using multiple intelligent algorithms, and to evaluate and optimize the preprocessing effect in real time. It is connected to the spectral feature deep mining unit via a high-speed data bus and transmits the optimized spectral data to the spectral feature deep mining unit;
[0029] The spectral feature deep mining unit uses innovative feature extraction models and algorithms to deeply mine characteristic parameters related to antibiotic concentration from preprocessed spectral data. It connects to the machine learning model construction and training unit through a data interaction interface and transmits the extracted feature vector set to the machine learning model construction and training unit.
[0030] The machine learning model construction and training unit is used to select the appropriate machine learning algorithm architecture, build and train the model, optimize and adjust the model based on the model performance evaluation results, and connect to the antibiotic concentration inversion result generation unit through a stable data output port to transmit the trained and optimized model to the antibiotic concentration inversion result generation unit;
[0031] The antibiotic concentration inversion result generation unit is used to input actual spectral data into the trained machine learning model, combine it with environmental correction factors, invert the antibiotic concentration in aquatic tail water, store and output the inversion results, and form a feedback connection with the sample collection multi-control unit to adjust the subsequent sample collection strategy according to the inversion results.
[0032] Beneficial effects: The present invention proposes a method for inverting the antibiotic concentration in aquatic tail water based on spectral-machine learning coupling. In the sample collection link, this method adopts a multi-dimensional sampling strategy to widely collect samples in different breeding areas, depths and time nodes to ensure that the constructed sample data set is highly representative, which greatly improves the comprehensiveness and reliability of subsequent analysis. In the process of spectral data collection and processing, professional and high-precision instruments and innovative adaptive noise removal models are used to effectively filter out noise interference and calibrate baseline drift. The acquired accurate spectral data provides a solid foundation for feature extraction. The introduction of a new spectral feature enhancement model can deeply explore the potential correlation between the spectrum and the antibiotic concentration, significantly enhance the correlation between the feature parameters and the target concentration, and make the model input more valuable. In terms of machine learning model construction, the improved neural network structure and unique training sample selection strategy have greatly improved the learning efficiency and fitting ability of the model. By dynamically adjusting the connection weights, the model can better capture the complex relationship between spectral features and antibiotic concentrations, and give priority to training samples with large discreteness and low similarity to avoid model overfitting and enhance its generalization performance. The comprehensive evaluation index system considers multiple dimensions such as prediction error, model complexity, and adaptability to different concentration ranges to ensure the comprehensiveness and accuracy of the model performance evaluation. In the concentration inversion stage, an environmental correction factor is introduced to fully consider the impact of environmental factors such as water temperature and salinity on the spectrum-antibiotic concentration relationship, so that the final inversion result is more in line with the actual situation. In addition, the method encrypts and stores the data to ensure data security, and the intelligent sampling equipment automatically adjusts the sampling according to the real-time parameters of the water body, further improving the scientificity and pertinence of the sample collection. The present invention provides an efficient, accurate, and intelligent solution for monitoring the concentration of antibiotics in aquatic tail water, which effectively promotes the progress and development of aquaculture environment monitoring technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flow chart of the method steps of the present invention;
[0034] Figure 2 This is a composition diagram of the method operation unit of the present invention. DETAILED DESCRIPTION
[0035] It should be noted that, unless there is a conflict, the embodiments in this application and the features described in the embodiments can be combined with each other. The application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] like Figure 1 As shown in the figure, the inversion method of antibiotic concentration in aquatic tail water based on spectral-machine learning coupling includes the following steps:
[0037] Step 1: Systematically collect aquatic tailwater samples from multiple aquaculture areas, at different depths, and at different time points, according to preset sampling specifications and standard procedures, to obtain representative aquatic tailwater samples and construct a sample dataset;
[0038] Specifically, to comprehensively obtain aquatic tailwater samples, collection efforts must be conducted in multiple different aquaculture areas. For example, in a pond aquaculture scenario, sampling points are selected at the four corners and center of the pond. These areas have varying antibiotic distributions due to differences in water flow and the distribution of aquatic organisms. Regarding depth, sampling is conducted starting at 0.2 meters below the water surface and at 0.5-meter intervals until near the bottom. Antibiotic concentrations vary at different depths due to factors such as light, microbial activity, and water exchange. Regarding timing, samples are collected in the early morning, midday, and evening. In the early morning, antibiotics may accumulate in the tailwater after aquatic animals metabolize the water overnight. At noon, higher water temperatures affect microbial activity and water convection, affecting antibiotic distribution. In the evening, after aquatic animals have ingested food, the tailwater composition also changes. Sampling procedures are strictly adhered to, using calibrated sterile sampling bottles and collecting 1 liter of water at a time to ensure sample purity and representativeness.
[0039] In areas with high aquaculture density, antibiotic use is relatively high, and tailwater antibiotic concentrations may be higher. Different feeding practices can lead to varying absorption and excretion of antibiotics by aquatic animals, which in turn affects tailwater concentrations. Flow patterns and microbial degradation capabilities vary at different depths of the water column. Furthermore, the feeding and metabolic states of aquatic animals vary at different times, causing fluctuations in tailwater antibiotic levels. By collecting extensive and targeted samples, we can comprehensively capture the various factors influencing antibiotic concentrations, providing a solid data foundation for subsequent, precise analysis. Only with comprehensive and accurate data can subsequent analysis be more reliable.
[0040] For example: In a large aquaculture pond covering an area of 10 mu, staff took samples according to the above method. Water samples were collected at 6 a.m., 12 noon, and 6 p.m. at the four corners and center of the pond, at depths of 0.2 meters, 0.7 meters, 1.2 meters, and 1.7 meters below the water surface. After analysis, it was found that the antibiotic concentration in the deeper waters in the center of the pond was relatively high in the early morning. This may be due to the weak water exchange at night, which causes the accumulation of antibiotics. This multi-dimensional collection of samples can comprehensively reflect the changes in the antibiotic concentration in the tail water of the entire pond at different times and spaces, providing rich and valuable data for subsequent research.
[0041] Step 2: Using a spectrum analyzer, under a set spectral range and preset measurement conditions, collect spectral data of each of the collected aquatic tail water samples to obtain original spectral data;
[0042] Specifically, a UV-visible spectrophotometer with a wavelength range of 200-1100 nanometers is used to collect spectral data. This wavelength range can cover the characteristic absorption peaks of most common antibiotics, making it possible to identify and quantify antibiotics. The scanning speed is set to 100 nanometers per minute. This speed can ensure that the scan is completed within a reasonable time without missing important information due to being too fast. The integration time is set to 0.5 seconds, which can effectively reduce the impact of instrument noise on the data and make the collected data more accurate. The collected water sample is carefully placed in a quartz cuvette with a 1 cm optical path and placed in the instrument for spectral scanning. The quartz cuvette has good light transmittance and chemical stability, which can ensure that the water sample is not disturbed during the measurement process.
[0043] Different antibiotics have unique absorption characteristics at specific wavelengths. By measuring the absorption of light at different wavelengths by a water sample, spectral data containing antibiotic information can be obtained. Proper scanning speed and integration time settings are key to ensuring both efficient data acquisition and high data quality. If the scanning speed is too fast, subtle changes in absorption peaks may not be accurately captured; if the integration time is too short, noise interference may lead to inaccurate data. Accurate spectral data is the basis for subsequent feature extraction and analysis, and its quality directly affects the accuracy of the entire inversion method.
[0044] For example, a well-known UV-Vis spectrophotometer was used to scan a sample of aquaculture wastewater containing tetracycline antibiotics according to the aforementioned parameters. During the scan, the instrument accurately recorded the absorbance of the water sample at different wavelengths. The results revealed a distinct absorption peak at wavelengths between 350 and 400 nanometers, which closely matched the characteristic absorption wavelength of tetracycline antibiotics. This finding provided key clues for the subsequent identification and quantification of tetracycline antibiotics in the sample and demonstrated the importance of accurately collecting spectral data.
[0045] Step 3: Preprocess the raw spectral data. By using a set filtering algorithm and data calibration method, noise interference, baseline drift, and other abnormal fluctuations in the spectral data are removed to obtain smooth and accurate preprocessed spectral data.
[0046] Specifically, a median filter algorithm was used to remove noise, with a filter window size of 5 data points. Median filtering can effectively suppress isolated noise points and preserve the true characteristics of the spectral data. Baseline calibration was performed using polynomial fitting, with a cubic polynomial being chosen here. Polynomial fitting can effectively simulate the trend of baseline drift and correct for baseline drift caused by insufficient instrument preheating and uneven cuvette transmittance. In actual operation, the raw spectral data was corrected sequentially using the median filter algorithm and the cubic polynomial fitting algorithm to obtain the preprocessed spectral data.
[0047] During spectral data acquisition, environmental electromagnetic interference and internal instrument electronic noise can cause noise fluctuations in the collected data. This noise can obscure the true spectral characteristics of antibiotics, leading to deviations in subsequent analysis. Baseline drift can shift the entire spectral curve, affecting the position and intensity of characteristic absorption peaks and, in turn, the accuracy of feature extraction. Spectral data preprocessing can purify the data, remove noise, and correct the baseline, ensuring that the subsequently extracted spectral features more accurately reflect the information about the antibiotic, improving the reliability of the entire analysis process.
[0048] For example, when analyzing raw spectral data, significant noise fluctuations were observed at 450 nanometers, with data points exhibiting irregular jitter. After processing with the median filter algorithm, the data at this location became smooth, consistent with the changing trend of surrounding data points. Furthermore, upward baseline drift at 500-600 nanometers was restored to a stable state after calibration using a cubic polynomial fit, and some faint spectral features previously obscured by the drift were now clearly visible. This preprocessed spectral data lays a solid foundation for subsequent accurate spectral feature extraction.
[0049] Step 4: Using the set feature extraction algorithm from the preprocessed spectral data, the spectral feature parameters that are correlated with the antibiotic concentration in aquatic tail water are mined to form a feature vector set;
[0050] Specifically, the successive projection algorithm (SPA) was used to extract spectral features, with a maximum number of iterations set to 50. The number of iterations was set to ensure that the algorithm could fully search for the optimal feature combination while avoiding the waste of computing resources and the risk of overfitting due to excessive iterations. Ten characteristic variables were selected based on extensive experiments and data analysis. These variables effectively represent the relationship between spectral data and antibiotic concentration, reduce data dimensionality, and improve the efficiency of subsequent model training. In practice, the preprocessed spectral data was input into the SPA algorithm, and after multiple iterative calculations, the ten characteristic variables with the strongest correlation with antibiotic concentration were selected.
[0051] Raw spectral data contains a vast amount of information from the ultraviolet to the visible bands, with extremely high dimensionality and a significant amount of redundant information. This redundant information not only increases the computational burden of subsequent model training but can also interfere with the model's learning of key information, leading to reduced performance. The goal of feature extraction is to select the features most relevant to antibiotic concentration from complex spectral data, reducing the data dimensionality and highlighting key information. This allows subsequent model training to focus more on learning the relationship between useful features and antibiotic concentration, improving the model's accuracy and generalization, enabling it to accurately predict antibiotic concentrations in aquaculture effluent samples from diverse aquaculture environments.
[0052] For example, for a sample of spectral data from aquatic tailwater containing multiple antibiotics, the SPA algorithm extracted features and identified absorbance values at 10 specific wavelengths, such as 250 nm, 380 nm, and 520 nm, as key characteristic variables. Further analysis revealed that these characteristic variables were much more highly correlated with the antibiotic concentration in the sample than data at other wavelengths. For example, when the concentration of a particular antibiotic in the sample increased, the absorbance value at 250 nm showed a clear upward trend. These key characteristic variables effectively represent the concentration of antibiotics in the tailwater, providing high-quality input data for subsequent machine learning model training.
[0053] Step 5: Build a machine learning model and select an appropriate machine learning algorithm architecture, such as a neural network model based on a specific neuron connection method and activation function. Divide the feature vector set into a training set and a test set. Use the training set to train and optimize the machine learning model over multiple rounds, adjusting the model's internal parameters to improve the model's ability to fit the relationship between spectral features and antibiotic concentration.
[0054] Specifically, a multi-layer perceptron (MLP) neural network model was selected to construct the concentration inversion model. The number of hidden layer nodes was set to 30. The choice of the number of hidden layer nodes will affect the learning ability and complexity of the model. 30 nodes can give the model sufficient nonlinear fitting ability without causing overfitting due to too many nodes. The ReLU activation function was used. The ReLU function can effectively solve the gradient vanishing problem in neural network training and improve the efficiency of model training. The learning rate was set to 0.001. The learning rate determines the step size of the model parameter update during the training process. A suitable learning rate can enable the model to converge stably during the training process. The number of training times was 500. Through multiple trainings, the model was able to fully learn the complex relationship between spectral features and antibiotic concentration. The feature vector set obtained after feature extraction was divided into 70% as the training set and 30% as the test set. The training set is used to learn and adjust the model parameters, and the test set is used to evaluate the performance of the model.
[0055] MLP neural networks possess powerful nonlinear mapping capabilities, capable of learning the complex nonlinear relationship between spectral features and antibiotic concentration. Appropriate settings for the number of hidden layer nodes, activation function, learning rate, and number of training sessions are key to ensuring effective model convergence during training and avoiding overfitting or underfitting. Overfitting can cause a model to perform well on the training set, but significantly degrade performance on the test set and in real-world applications. Underfitting, on the other hand, prevents the model from fully learning the patterns in the data, resulting in low prediction accuracy. Dividing the training and test sets allows for the assessment of the model's generalization capabilities, ensuring that the model can accurately predict antibiotic concentrations even when faced with new, unseen data.
[0056] For example, during training, the model's mean squared error (MSE) on the training set gradually decreased with increasing training cycles. After approximately 300 training cycles, the MSE stabilized, indicating that the model had gradually converged and begun to learn the intrinsic connection between spectral features and antibiotic concentrations. On the test set, the model's predictions for different antibiotic concentrations were relatively close to the true values. For example, for a set of test samples with true antibiotic concentrations between 0.2 and 0.8 mg / L, the model's predictions were mostly within an error range of ±0.1 mg / L of the true values, demonstrating the model's excellent learning and generalization capabilities, enabling it to accurately predict antibiotic concentrations in aquaculture effluents.
[0057] Step 6: Use the test set to evaluate the performance of the trained and optimized machine learning model. Quantify the model's accuracy, stability, and other performance based on the evaluation indicators. If the model performance does not meet the preset standards, return to step 5 and retrain and optimize until the model performance meets the requirements.
[0058] Specifically, the root mean square error (RMSE), mean absolute error (MAE) and coefficient of determination (R 2 ) as an evaluation metric. RMSE is used to measure the average magnitude of the deviation between the predicted value and the true value. When calculating, first find the square of the difference between each predicted value and the true value, then find the average of these squared values, and finally take the square root. MAE measures the average absolute value of the prediction error, that is, calculate the average of the absolute values of the difference between each predicted value and the true value. 2 Used to evaluate the goodness of fit of the model, which indicates the proportion of data variation that the model can explain. In practice, the model's predicted concentrations for the test set are compared with the known true concentrations to calculate RMSE, MAE, and R respectively. 2 .
[0059] RMSE can directly reflect the average deviation between the model prediction value and the true value. The smaller the value, the closer the prediction value is to the true value. MAE measures the size of the model prediction error in the form of mean absolute error from another perspective. Similarly, the smaller the value, the better. 2 The closer the value is to 1, the better the model fits the data, indicating that the model is able to explain the relationship between spectral features and antibiotic concentrations. These three indicators allow for a comprehensive and quantitative assessment of the model's accuracy, stability, and other performance characteristics, determining whether the model meets the accuracy and reliability requirements for practical applications. Only well-performing models can be effective in monitoring antibiotic concentrations in aquaculture tailwater.
[0060] For example, the performance of a trained model is evaluated on the test set. After calculation, the RMSE is 0.05 mg / L, the MAE is 0.03 mg / L, and the R 2The RMSE and MAE values are small, indicating that the deviation between the model prediction value and the true value is small and the prediction accuracy is high. 2 A value close to 1 indicates that the model fits the test data very well and can well capture the relationship between spectral features and antibiotic concentration. Taking these three indicators into consideration, we can conclude that the model performs well and can be applied to the actual inversion of antibiotic concentrations in aquaculture tailwater, providing reliable technical support for monitoring antibiotic contamination in aquaculture tailwater.
[0061] Step 7: Apply the verified machine learning model to the actual aquatic tailwater spectral data. Based on the output of the model and the pre-set concentration conversion rules, the antibiotic concentration in the aquatic tailwater is inverted.
[0062] Specifically, the spectral data of aquatic tailwater collected, preprocessed, and feature extracted are used to generate corresponding feature vectors, which are then fed into a trained machine learning model. Based on the learned relationship between spectral features and antibiotic concentration, the model outputs a predicted value. This value is then converted to the actual antibiotic concentration using a pre-established concentration conversion rule. This concentration conversion rule, derived from extensive experimental data and theoretical analysis, accurately converts the model output value into an antibiotic concentration that corresponds to physical meaning.
[0063] After completing the previous steps, a reliable machine learning model was established. Concentration inversion involves applying this model to the analysis of actual aquaculture tailwater samples. By converting spectral information into intuitive antibiotic concentration data, this provides critical data support for monitoring and remediation of antibiotic contamination in aquaculture tailwater. Accurate concentration data enables farmers to promptly understand the status of antibiotic contamination in tailwater and implement appropriate remediation measures, thereby reducing pollution to the surrounding aquatic environment and protecting ecological balance.
[0064] For example, a newly collected aquatic tailwater sample is processed according to the previous steps to obtain a feature vector, which is then input into the trained model. The model output value is 0.3. According to the pre-established concentration conversion rules, this output value corresponds to an antibiotic concentration of 0.6 mg / L. Through this concentration inversion, the antibiotic concentration in the tailwater sample can be quickly and accurately determined, providing an important basis for determining whether the tailwater meets discharge standards and for subsequent treatment decisions. For example, if the aquaculture area stipulates that the tailwater antibiotic concentration discharge standard is 0.5 mg / L, then based on the inversion results, the aquaculture farmer will need to take appropriate treatment measures to reduce the tailwater antibiotic concentration to protect the environment.
[0065] Preferably, in the preprocessing operation of step 3, an innovative adaptive noise removal model is adopted, and the model formula is: Among them, Y(i) is the value of the original spectrum data at the i-th sampling point, Y new (i) is the value of the spectral data at the i-th sampling point after noise removal, α is the adaptive adjustment coefficient, which ranges from [0.3, 0.7] and is dynamically adjusted according to the degree of fluctuation of the spectral data, and n is the neighborhood window size, which ranges from [3, 7] and is set according to the resolution of the spectral data.
[0066] The adaptive noise removal model provides an innovative and effective noise removal method for spectral data preprocessing. It flexibly adjusts parameters based on the actual fluctuation level and resolution of the spectral data, accurately filtering out noise interference. This ensures that subsequent spectral feature extraction is based on pure and accurate data, improving the data foundation quality of the entire inversion method.
[0067] Preferably, in the feature extraction process of step 4, a new spectral feature enhancement model is introduced, and the model formula is: Among them, F(k) is the kth spectral feature value originally extracted, F enhanced (k) is the enhanced spectral characteristic value, β is the characteristic weight coefficient, the value range is [0.6, 0.9], γ is the fluctuation enhancement coefficient, the value range is [0.1, 0.3], is the average value of all original extracted spectral feature values, and m is the total number of spectral features.
[0068] The new spectral feature enhancement model strengthens the correlation between original spectral features and antibiotic concentration. Through a unique computational approach, it highlights key features, making the model input feature vector more representative. This helps the machine learning model learn the relationship between spectra and antibiotic concentration more efficiently and accurately, thereby improving inversion accuracy.
[0069] Preferably, in the machine learning model constructed in step 5, an improved neural network structure is adopted, and the connection weight update formula between the hidden layer neurons is: in, is the updated connection weight, is the current connection weight, is the connection weight before the last update, η is the learning rate, and its value range is [0.001, 0.01], δ j is the error term of neuron j, O i is the output value of neuron i, λ is the weight attenuation coefficient, and its value range is [0.0001, 0.001].
[0070] The improved neural network connection weight update formula provides a more rational parameter adjustment mechanism for machine learning model training. Dynamic weight adjustment enables the model to better capture the complex nonlinear relationship between spectral features and antibiotic concentration, enhancing the model's learning ability and fitting effect, and improving the model's adaptability to different samples.
[0071] Preferably, in the model training process of step 5, a new training sample selection strategy is introduced. By calculating the feature discreteness of the sample and the similarity with other samples, samples with large feature discreteness and low similarity with the selected samples are preferentially selected for training. The sample selection formula is: Among them, S(k) is the selection priority of the kth sample, σ F(k) is the standard deviation of the kth sample feature, sim(F(k), F(l)) is the similarity between the kth sample and the lth sample feature, calculated using cosine similarity, and N is the total number of samples.
[0072] A unique training sample selection strategy changes the traditional random sample selection method. It prioritizes samples with large feature dispersion and low similarity to existing samples for training, preventing the model from falling into local optimal solutions. This effectively enhances the model's generalization ability, enabling it to maintain good predictive performance when dealing with aquaculture tailwater samples from various complex aquaculture environments.
[0073] Preferably, in the model performance evaluation in step 6, a comprehensive evaluation index is used, which combines the prediction error of the model, the complexity of the model and the adaptability of the model to samples in different concentration ranges. The evaluation index formula is:
[0074]
[0075] Among them, E is the comprehensive evaluation index value, is the predicted antibiotic concentration of the i-th sample, is the true antibiotic concentration of the i-th sample, M is the number of test set samples, ω1, ω2, ω3 are weight coefficients, and their value ranges are [0.4, 0.6], [0.2, 0.3], [0.2, 0.3], respectively, L is the number of model parameters, W j is the j-th model parameter, P is the number of intervals divided by antibiotic concentration, var represents variance, and mean represents mean.
[0076] Comprehensive evaluation measures model performance from multiple dimensions, rather than being limited to a single metric. This comprehensive approach more accurately assesses model performance in terms of prediction error, complexity, and adaptability to samples across different concentration ranges. This provides a more scientific and comprehensive basis for model optimization and selection, ensuring the reliability of the inversion model.
[0077] Preferably, in the concentration inversion process of step seven, the influence of different breeding environments on the spectrum-antibiotic concentration relationship is taken into account, and an environmental correction factor is introduced. The inversion formula is: Among them, C final is the final inversion-derived antibiotic concentration in aquaculture tail water, C model is the antibiotic concentration output by the machine learning model, θ q is the correction coefficient of the qth environmental factor, and its value range is determined according to experimental data. q is the measured value of the qth environmental factor, Q is the number of environmental factors, and environmental factors include water temperature, salinity, dissolved oxygen, etc.
[0078] The inclusion of an environmental correction factor in the concentration inversion process fully accounts for the complexity of the actual aquaculture environment. Environmental factors such as water temperature and salinity can affect the spectrum-antibiotic concentration relationship. This claim makes the inversion results more realistic, improving the accuracy and practicality of the inversion method in practical applications.
[0079] Preferably, in the entire inversion method, the spectral data and sample information are encrypted and stored using a new encryption algorithm. The encryption formula is: encrypted (r)=D(r)⊕H(timestamp⊕key), where D encrypted (r) is the encrypted data, D(r) is the original data, ⊕ is the XOR operation, H is the hash function, timestamp is the data collection timestamp, key is the encryption key, the key length is 128 bits, and it is changed regularly.
[0080] Data encryption and storage ensure the security of spectral data and sample information. In environments where data is vulnerable to leakage, a new encryption algorithm uses a timestamp and key combined with a hash function, with regular key changes, to prevent unauthorized access and tampering, ensuring data confidentiality and integrity throughout the inversion process.
[0081] Preferably, in the sample collection process of step 1, an intelligent sampling device is used. The device monitors the physical and chemical parameters of the aquaculture water in real time through built-in sensors, and automatically adjusts the sampling position and time according to the pre-set sampling rules. The sampling rule formula is: T sample =T base +ΔT×f(pH, DO, Temperature), where T sample is the actual sampling time, T base is the basic sampling time, ΔT is the time adjustment step, f(pH, DO, Temperature) is the adjustment function determined according to parameters such as water pH, dissolved oxygen and temperature, and the function form is obtained by fitting the experimental data.
[0082] Intelligent sampling equipment automatically adjusts sampling based on the real-time physical and chemical parameters of the water, achieving intelligent and precise sample collection. This overcomes the shortcomings of traditional manual sampling and can promptly adjust sampling strategies based on real-time changes in the water, further improving the scientific nature and representativeness of samples and providing higher-quality data for subsequent analysis.
[0083] like Figure 2 As shown in the figure, the inversion method of antibiotic concentration in aquatic tail water based on spectral-machine learning coupling is implemented by the following units, including:
[0084] The sample collection multi-element control unit is used to collect aquatic tail water samples under various environmental conditions according to a multi-dimensional sampling strategy, and is connected to the sample transmission buffer unit through a high-speed data transmission line to transmit the collected sample information to the sample transmission buffer unit;
[0085] The sample transmission buffer unit is used to temporarily store the sample information transmitted by the sample acquisition multi-element control unit, and to preliminarily organize and cache the information. It is connected to the spectral data acquisition and precision analysis unit through a stable data interface, and transmits the organized sample information to the spectral data acquisition and precision analysis unit in an orderly manner;
[0086] The spectral data acquisition and precision analysis unit is used to collect high-precision spectral data of samples under a strictly controlled spectral measurement environment, and perform preliminary analysis and processing on the original spectral data. It is connected to the spectral data preprocessing intelligent optimization unit through a dedicated data channel and transmits the pre-processed spectral data to the spectral data preprocessing intelligent optimization unit;
[0087] The spectral data preprocessing intelligent optimization unit is used to perform preprocessing operations such as noise removal and baseline calibration on the spectral data using a variety of intelligent algorithms, and to evaluate and optimize the preprocessing effect in real time. It is connected to the spectral feature deep mining unit via a high-speed data bus and transmits the optimized spectral data to the spectral feature deep mining unit;
[0088] The spectral feature deep mining unit uses innovative feature extraction models and algorithms to deeply mine characteristic parameters related to antibiotic concentration from preprocessed spectral data. It connects to the machine learning model construction and training unit through a data interaction interface and transmits the extracted feature vector set to the machine learning model construction and training unit.
[0089] The machine learning model construction and training unit is used to select the appropriate machine learning algorithm architecture, build and train the model, optimize and adjust the model based on the model performance evaluation results, and connect to the antibiotic concentration inversion result generation unit through a stable data output port to transmit the trained and optimized model to the antibiotic concentration inversion result generation unit;
[0090] The antibiotic concentration inversion result generation unit is used to input actual spectral data into the trained machine learning model, combine it with factors such as environmental correction, invert the antibiotic concentration in aquatic tail water, store and output the inversion results, and form a feedback connection with the sample collection multi-element control unit to adjust the subsequent sample collection strategy according to the inversion results.
[0091] The inversion method based on spectral-machine learning coupling shows significant advantages. In the sample collection process, traditional methods are often limited to a small number of areas and specific times, and the samples lack representativeness. This method collects samples every 0.5 meters from 0.2 meters below the water surface at different depths in multiple aquaculture areas, covering the four corners and center of the pond, and at different time points such as early morning, noon, and evening, comprehensively covering the spatiotemporal factors that affect the concentration of antibiotics. At the same time, the use of intelligent sampling equipment automatically adjusts the sampling location and time according to the real-time physical and chemical parameters of the water body, greatly improving the comprehensiveness and scientific nature of the samples, and overcoming the shortcomings of the single sample of existing technologies.
[0092] Traditional methods for spectral data processing suffer from limited instrument precision, susceptibility to noise interference and baseline drift, resulting in poor data accuracy and difficulty in accurately extracting spectral features. This method utilizes specialized, high-precision spectral analysis instruments, combined with an innovative adaptive noise removal model, to effectively filter out noise and precisely calibrate the baseline through methods such as cubic polynomial fitting. During feature extraction, a novel spectral feature enhancement model is introduced to deeply mine features highly correlated with antibiotic concentration, resulting in data quality and feature validity far exceeding traditional techniques, addressing the issue of crude spectral data processing in existing technologies.
[0093] In terms of machine learning model construction and application, traditional models are mostly based on simple linear relationships, cannot adapt to complex aquaculture environments, and have weak generalization capabilities. This method constructs a multi-layer perceptron (MLP) neural network model by improving the way the hidden layer neuron connection weights are updated, such as using a specific formula to dynamically adjust the weights, setting reasonable parameters such as the number of hidden layer nodes and learning rate, and using a unique training sample selection strategy to prioritize training samples with large feature discreteness and low similarity to selected samples, thereby significantly improving the model's learning efficiency and fitting ability. When inverting concentrations, an environmental correction factor is introduced to fully consider the impact of environmental factors such as water temperature and salinity on the spectrum-antibiotic concentration relationship. In addition, a comprehensive evaluation index is used to evaluate model performance from multiple dimensions, including prediction error, model complexity, and adaptability to different concentration ranges, to ensure model accuracy and stability. At the same time, data is encrypted and stored, and a new encryption algorithm is used to ensure data security. These innovative measures enable this method to accurately invert antibiotic concentrations, overcoming the shortcomings of existing technical models such as poor adaptability, low accuracy, and lack of data security.
[0094] While embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling, characterized by: The method includes: Step 1: Systematically collect aquatic tailwater samples from multiple aquaculture areas, at different depths, and at different time points, according to preset sampling specifications and standard procedures, to obtain representative aquatic tailwater samples and construct a sample dataset; Step 2: Using a spectrum analyzer, under a set spectral range and preset measurement conditions, collect spectral data of each of the collected aquatic tail water samples to obtain original spectral data; Step 3: Preprocess the raw spectral data to remove noise, baseline drift, and other abnormal fluctuations in the spectral data through a set filtering algorithm and data calibration method; Step 4: Using the set feature extraction algorithm from the preprocessed spectral data, the spectral feature parameters that are correlated with the antibiotic concentration in aquatic tail water are mined to form a feature vector set; Step 5: Build a machine learning model and select an appropriate machine learning algorithm architecture, including a neural network model based on a specific neuron connection method and activation function. Divide the feature vector set into a training set and a test set. Use the training set to perform multiple rounds of training and optimization on the machine learning model and adjust the model's internal parameters. Step 6: Use the test set to evaluate the performance of the trained and optimized machine learning model. Quantify the accuracy and stability of the model based on the evaluation indicators. If the model performance does not meet the preset standards, return to step 5 and retrain and optimize until the model performance meets the requirements. Step 7: Apply the verified machine learning model to the actual aquatic tailwater spectral data. Based on the output of the model and the pre-set concentration conversion rules, the antibiotic concentration in the aquatic tailwater is inverted.
2. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the preprocessing operation of step 3, an innovative adaptive noise removal model is used. The model formula is: Among them, Y(i) is the value of the original spectrum data at the i-th sampling point, Y new (i) is the value of the spectral data at the i-th sampling point after noise removal, α is the adaptive adjustment coefficient, which ranges from [0.3, 0.7] and is dynamically adjusted according to the degree of fluctuation of the spectral data, and n is the neighborhood window size, which ranges from [3, 7] and is set according to the resolution of the spectral data.
3. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the feature extraction process of step 4, a new spectral feature enhancement model is introduced. The model formula is: Among them, F(k) is the kth spectral feature value originally extracted, F enhanced (k) is the enhanced spectral characteristic value, β is the characteristic weight coefficient, the value range is [0.6, 0.9], γ is the fluctuation enhancement coefficient, the value range is [0.1, 0.3], is the average value of all original extracted spectral feature values, and m is the total number of spectral features.
4. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the machine learning model constructed in step 5, an improved neural network structure is adopted, and the connection weight update formula between the hidden layer neurons is: in, is the updated connection weight, is the current connection weight, is the connection weight before the last update, η is the learning rate, and its value range is [0.001, 0.01], δ j is the error term of neuron j, O i is the output value of neuron i, λ is the weight attenuation coefficient, and its value range is [0.0001, 0.001].
5. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the model training process of step 5, a new training sample selection strategy is introduced. By calculating the feature discreteness of the sample and the similarity with other samples, samples with large feature discreteness and low similarity with the selected samples are preferentially selected for training. The sample selection formula is: Among them, S(k) is the selection priority of the kth sample, σ F(k) is the standard deviation of the kth sample feature, sim(F(k), F(l)) is the similarity between the kth sample and the lth sample feature, calculated using cosine similarity, and N is the total number of samples.
6. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the model performance evaluation in step 6, a comprehensive evaluation index is used, which combines the model's prediction error, model complexity, and the model's adaptability to samples in different concentration ranges. The evaluation index formula is: Among them, E is the comprehensive evaluation index value, is the predicted antibiotic concentration of the i-th sample, is the true antibiotic concentration of the i-th sample, M is the number of test set samples, ω1, ω2, ω3 are weight coefficients, and their value ranges are [0.4, 0.6], [0.2, 0.3], [0.2, 0.3], respectively, L is the number of model parameters, W j is the j-th model parameter, P is the number of intervals divided by antibiotic concentration, var represents variance, and mean represents mean.
7. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the concentration inversion process of step 7, considering the influence of different breeding environments on the spectrum-antibiotic concentration relationship, an environmental correction factor is introduced, and the inversion formula is: Among them, C final is the final inversion-derived antibiotic concentration in aquaculture tail water, C model is the antibiotic concentration output by the machine learning model, θ q is the correction coefficient of the qth environmental factor, and its value range is determined according to experimental data. q is the measured value of the qth environmental factor, Q is the number of environmental factors, and the environmental factors include water temperature, salinity, and dissolved oxygen.
8. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the entire inversion method, the spectral data and sample information are encrypted and stored using a new encryption algorithm. The encryption formula is: Among them, D encrypted (r) is the encrypted data, D(r) is the original data, is an XOR operation, H is a hash function, timestamp is the data collection timestamp, key is the encryption key, the key length is 128 bits and is changed regularly.
9. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to claim 1 is characterized in that: In the sample collection process of step 1, an intelligent sampling device is used. The device monitors the physical and chemical parameters of the aquaculture water in real time through built-in sensors and automatically adjusts the sampling position and time according to the pre-set sampling rules. The sampling rule formula is: T sample =T base +ΔT×f(pH, DO, Temperature), where T sample is the actual sampling time, T base is the basic sampling time, ΔT is the time adjustment step, f(pH, DO, Temperature) is the adjustment function determined according to the water pH, dissolved oxygen and temperature parameters, and the function form is obtained by fitting the experimental data.
10. The method for inverting antibiotic concentration in aquatic tail water based on spectral-machine learning coupling according to any one of claims 1 to 9, characterized in that: The method is implemented by the following units, including: The sample collection multi-element control unit is used to collect aquatic tail water samples under various environmental conditions according to a multi-dimensional sampling strategy, and is connected to the sample transmission buffer unit through a high-speed data transmission line to transmit the collected sample information to the sample transmission buffer unit; The sample transmission buffer unit is used to temporarily store the sample information transmitted by the sample acquisition multi-element control unit, and to preliminarily organize and cache the information. It is connected to the spectral data acquisition and precision analysis unit through a stable data interface, and transmits the organized sample information to the spectral data acquisition and precision analysis unit in an orderly manner; The spectral data acquisition and precision analysis unit is used to collect high-precision spectral data of samples under a strictly controlled spectral measurement environment, and perform preliminary analysis and processing on the original spectral data. It is connected to the spectral data preprocessing intelligent optimization unit through a dedicated data channel and transmits the pre-processed spectral data to the spectral data preprocessing intelligent optimization unit; The spectral data preprocessing intelligent optimization unit is used to perform noise removal and baseline calibration preprocessing operations on the spectral data using multiple intelligent algorithms, and to evaluate and optimize the preprocessing effect in real time. It is connected to the spectral feature deep mining unit via a high-speed data bus and transmits the optimized spectral data to the spectral feature deep mining unit; The spectral feature deep mining unit uses innovative feature extraction models and algorithms to deeply mine characteristic parameters related to antibiotic concentration from preprocessed spectral data. It connects to the machine learning model construction and training unit through a data interaction interface and transmits the extracted feature vector set to the machine learning model construction and training unit. The machine learning model construction and training unit is used to select the appropriate machine learning algorithm architecture, build and train the model, optimize and adjust the model based on the model performance evaluation results, and connect to the antibiotic concentration inversion result generation unit through a stable data output port to transmit the trained and optimized model to the antibiotic concentration inversion result generation unit; The antibiotic concentration inversion result generation unit is used to input actual spectral data into the trained machine learning model, combine it with environmental correction factors, invert the antibiotic concentration in aquatic tail water, store and output the inversion results, and form a feedback connection with the sample collection multi-control unit to adjust the subsequent sample collection strategy according to the inversion results.