Quality parameter detection model construction method based on incremental learning driving

By using an incremental learning-driven method to build a quality parameter detection model, we have solved the problems of cumbersome operation, high cost, and model forgetting in bauxite quality detection. This method enables efficient adaptive updates and accurate detection in dynamic environments, reducing deployment costs and time.

CN120932777AActive Publication Date: 2025-11-11CHINA CERTIFICATION & INSPECTION (GROUP) CO LTD HEBEI BRANCH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511460516.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-11-11
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Traditional bauxite quality testing methods are cumbersome, time-consuming, and costly. Furthermore, portable near-infrared spectrometers face catastrophic amnesia and limited data access under complex working conditions, resulting in decreased predictive performance and failing to meet the needs of rapid and efficient on-site testing.

Method used

We adopt an incremental learning-driven method for constructing quality parameter detection models. By building a basic parameter prediction model and training it using incremental learning, combined with a generative adversarial reference network model and a feature extractor, we can achieve adversarial knowledge updates, alleviate data distribution drift in dynamic environments, and reduce deployment costs and computation time.

Benefits of technology

It achieves efficient adaptive updates in dynamic environments, extends the model lifecycle, improves the reliability and accuracy of quality parameter detection, reduces costs, and solves the problems of model forgetting and limited data access in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932777A_ABST
    Figure CN120932777A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of spectral analysis and intelligent detection, in particular to a quality parameter detection model construction method based on incremental learning driving. The method comprises the following steps: constructing a parameter prediction basic model for performing quality parameter prediction on a to-be-detected substance, providing a training data set corresponding to the to-be-detected substance, and performing incremental learning driving-based model training on the parameter prediction basic model by utilizing a training data set in the training data set, and generating a quality parameter detection model after model training. According to the method, the quality parameter detection model can be effectively constructed and formed, the deployment time and calculation cost of the quality parameter detection model in practical application are reduced, the deployed quality parameter detection model can effectively relieve disastrous forgetting during reasoning, efficient self-adaptive updating in a dynamic environment is achieved, and the method is suitable for popularization and application. The life cycle of the quality parameter detection model is obviously prolonged, the cost of quality parameter detection by adopting the near infrared spectrum is reduced, and the reliability and precision of quality parameter detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a quality parameter detection method, and more particularly to a quality parameter detection model construction method based on incremental learning. Background Technology

[0002] Bauxite, as a primary raw material for the aluminum industry, holds significant economic and industrial value globally. Its superior properties, including low density, high ductility, good electrical conductivity, strong corrosion resistance, and high recyclability, make it widely used in defense technology, aerospace, automobile manufacturing, and emerging green industries. Bauxite is primarily composed of minerals such as gibbsite (AlO(OH)) and gibbsite (Al(OH)3), and contains varying amounts of silicon, iron, titanium, and trace amounts of other elemental compounds.

[0003] Iron oxide (Fe2O3) content is one of the key indicators determining the quality of bauxite. Its changes not only affect ore classification but also directly influence the ore's thermal properties and phase transformation behavior, while also indicating the presence of associated metal resources. Therefore, accurate determination of Fe2O3 content in bauxite plays an irreplaceable role in optimizing beneficiation ratios, improving resource utilization efficiency, and ensuring the stable operation of industrial production processes. While traditional chemical analysis methods offer high accuracy, they suffer from drawbacks such as cumbersome operation, long cycles, and the need to destroy samples, making them unsuitable for rapid and efficient on-site testing. Furthermore, they require a high level of expertise from analysts.

[0004] Modern detection technologies have demonstrated high efficiency and accuracy in bauxite quality testing, showing great application potential. These include methods such as atomic absorption spectrometry, X-ray fluorescence spectrometry, and laser-induced breakdown spectroscopy. However, these methods require stringent sample pretreatment and involve high instrument costs, which limits their widespread application to some extent. For example, patent application CN112179930A describes a method for determining the content of Al2O3, SiO2, and Fe2O3 in high-sulfur bauxite using X-ray fluorescence spectrometry. This application involves preparing standard sample glass slides by uniformly mixing sulfur-containing flux and bauxite standard samples, followed by pre-oxidation treatment. Subsequently, the fluorescence intensity of the glass slides is measured using an X-ray fluorescence spectrometer, and a working curve is established with fluorescence intensity as the ordinate and the corresponding mass concentrations of Al2O3, SiO2, and Fe2O3 as the abscissa, thereby calculating the corresponding contents of Al2O3, SiO2, and Fe2O3 in the glass slide. It is understandable that while the detection method used in this patent application has advantages in accuracy, it requires complex sample preparation and high instrument costs.

[0005] Near-infrared spectroscopy, with its advantages of low cost, no need for complex sample preparation, and rapid, non-destructive operation, shows enormous application potential. Traditional near-infrared spectrometers are generally bulky and expensive, limiting their field deployment capabilities. In contrast, portable near-infrared spectrometers, with their compact size and ease of operation, are gradually becoming an important platform for spectral detection in complex working conditions, and are expected to bring revolutionary changes to the effective utilization of bauxite resources and the intelligent quality detection of mineral products.

[0006] However, when applying portable near-infrared spectrometers to complex mining sites, the traditional one-time static model construction method faces a catastrophic forgetting problem in practical applications. That is, the model completely forgets the old knowledge, resulting in a sharp decline in its predictive performance. This is mainly due to changes in batch source, environment, sampling conditions or instrument status, which cause the collected near-infrared spectral data to show obvious distribution drift, weakening the predictive stability of the static model.

[0007] In practical applications, near-infrared spectral data of samples are usually acquired in batches, which may lead to data access restrictions: historical samples may not be retained or shared due to storage costs, commercial confidentiality or privacy compliance (such as across factories or institutions). This means that the traditional strategy of relying on full retraining faces high time and computing costs in actual deployment and is difficult to meet the requirements of data security and long-term adaptability. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for constructing a quality parameter detection model based on incremental learning. This method can effectively construct a quality parameter detection model, reduce the time and computational cost of deploying the quality parameter detection model in practical applications, and effectively alleviate catastrophic forgetting during inference. It enables efficient adaptive updates in dynamic environments, significantly extends the life cycle of the quality parameter detection model, reduces the cost of using near-infrared spectroscopy for quality parameter detection, and improves the reliability and accuracy of quality parameter detection.

[0009] According to the technical solution provided by this invention, a method for constructing a quality parameter detection model based on incremental learning is provided, the method comprising: A basic parameter prediction model is constructed for predicting quality parameters of the substance to be tested, and a training data set corresponding to the substance to be tested is provided, wherein... The training data group includes at least two training datasets, and the spectral distribution characteristics of all training datasets are at least incompletely matched. Each training dataset includes several training samples related to the substance to be tested. For any training sample, the training sample includes training near-infrared spectral data and training parameter labels characterizing the quality parameter state of the substance under test, wherein the training parameter labels are positively correlated with the training near-infrared spectral data; The basic parameter prediction model is trained using the training dataset within the training data set, driven by incremental learning. This trained model then generates a quality parameter detection model. When performing incremental learning-driven model training, the following are included: Configure one training dataset as the base set, and use the remaining training datasets as incremental sets. The basic model for parameter prediction is trained using the base set. Then, incremental learning is performed once using the incremental set. After all incremental learning is performed, a quality parameter detection model is generated.

[0010] During basic training, the parameter prediction basic model and the generative adversarial network basic model are trained using the basic set, respectively, so as to generate the parameter prediction basic model and the generative adversarial network basic model after training. When performing the first incremental learning, the parameter prediction base model, the generative adversarial base network model, and the basic label information of the base training set are configured as the parameter prediction reference model, the generative adversarial reference network model, and the reference label information, respectively. The basic label information is generated based on the base set. After performing incremental learning once using an incremental set, a parameter prediction reference model, an adversarial reference network model, and reference label information are generated for the next incremental learning. After performing the last incremental learning, a quality parameter detection model is generated based on the parameter prediction reference model after the incremental learning.

[0011] When performing incremental learning, the following are included: Determine the incremental set to be used for this incremental learning, and obtain the latest generated parameter prediction reference model, generative adversarial reference network model, and reference label information before executing this incremental learning; When generating the parameter prediction reference model for the next incremental learning based on the determined incremental set and the obtained parameter prediction reference model, the following steps are included: The parameter prediction reference model is incrementally trained using the determined increment set, wherein the incremental training includes several rounds of incremental iteration. In each round of incremental iteration, the parameter prediction reference model is trained using the incremental set to generate a parameter prediction intermediate model after training. The regression prediction head in the parameter prediction intermediate model is then jointly optimized to form the parameter prediction intermediate model for the next round of incremental iteration after joint optimization. Once the incremental iteration reaches the specified number of rounds, the intermediate parameter prediction model that has reached the specified number of rounds will be configured as the parameter prediction reference model for the next incremental learning. When generating the generative adversarial reference network (GARNR) model required for the next incremental learning, based on the determined incremental set and the GARNR model, the following steps are included: The acquired generative adversarial reference network model is trained using the incremental set, so as to generate the generative adversarial reference network model required for the next incremental learning. When generating reference label information, the training parameter labels of all training samples in the determined incremental set are obtained, and the label mean and label variance corresponding to all training parameter labels are calculated, so as to form the reference label information required for the next incremental learning based on the calculated label mean and label variance.

[0012] When jointly optimizing the regression prediction heads within the intermediate model for parameter prediction, the following is included: Generate a feature-label set required for joint optimization of the regression prediction head, wherein the feature-label set includes several pseudo-feature-pseudo-label pairs and several real feature-real-label pairs; When generating true feature-true label pairs, the following steps are included: The feature extractor within the intermediate model of parameter prediction is used to extract features from a training near-infrared spectral data to obtain the true features; Based on the training near-infrared spectral data corresponding to the real features, the training parameter labels that correspond to the near-infrared spectral data are determined, and the determined training parameter labels are configured as real labels. The training near-infrared spectral data belongs to a training sample in the incremental set used to perform this incremental learning. When generating pseudo-feature-pseudo-label pairs, the following is included: Based on the reference label information obtained from this incremental learning process, pseudo-labels are generated through sampling. The pseudo-labels are loaded into the Generative Adversarial Reference Network (GARN) model, which then generates pseudo-features corresponding to the pseudo-labels. The regression prediction head is optimized and updated using the feature-label set to generate an optimized regression prediction head.

[0013] When training the obtained parameter prediction reference model using the determined increment set, the incremental training loss function used is:

[0014] in, For incremental training loss function, For the average absolute error loss, For distillation loss, For offset penalty loss, Mean absolute error loss The weight, Distillation loss weights, Offset penalty loss The weight, The training near-infrared spectral data for a training sample within the increment set. To utilize the parameters obtained from incremental iteration during this incremental learning process to predict the spectral features extracted by the intermediate model, To utilize the latest parameters generated before this incremental learning to predict the spectral features extracted from the reference model, Spectral characteristics With spectral characteristics The square of the L2 norm of the difference, Spectral characteristics With spectral characteristics The expected value of the square of the L2 norm of the difference; This is the first intermediate model for parameter prediction during incremental iteration in this incremental learning process. Each model parameter This is the first parameter prediction reference model generated before this incremental learning. Each model parameter For the first Model parameters The estimated value of the Fisher information matrix, This is a hyperparameter.

[0015] When generating training near-infrared spectral data for each training sample, the following steps are included: Near-infrared spectral source data corresponding to each training near-infrared spectral data is acquired, and the near-infrared spectral source data is preprocessed to generate training near-infrared spectral data after preprocessing. The near-infrared spectral source data is generated through near-infrared spectral acquisition; Preprocessing of near-infrared spectral source data includes standard normal transformation and / or Huber-based adaptive asymmetric least squares baseline correction. When the preprocessing of near-infrared spectral source data includes both standard normal transformation and Huber-based adaptive asymmetric least squares baseline correction, then the near-infrared spectral source data is subjected to standard normal transformation and Huber-based adaptive asymmetric least squares baseline correction sequentially. When performing Huber-regularized adaptive asymmetric least squares baseline correction, we have:

[0016] in, For the final baseline estimation data, To train near-infrared spectral data, the total number of wavenumber points corresponding to near-infrared light spectral source data, The residual at the nth wavenumber point, This represents the weight corresponding to the nth wavenumber point. For the smoothing constraint term corresponding to the near-infrared light spectral source data, The absorbance at the nth wavenumber point within the near-infrared spectral correction base data. The absorbance at the nth wavenumber point within the initial baseline estimation data generated based on near-infrared spectral correction baseline data. For absorbance The second-order difference was performed; When the preprocessing performed only includes Huber regularization-based adaptive asymmetric least squares baseline correction, the near-infrared spectral correction basis data is the near-infrared spectral source data. When the preprocessing includes standard normal transformation and Huber regularization-based adaptive asymmetric least squares baseline correction, the near-infrared spectral source data is generated by performing standard normal transformation on the near-infrared spectral source data. Polynomial fitting is performed on the final baseline estimate data to generate the final baseline fit data after polynomial fitting; The near-infrared spectral correction baseline data is subtracted from the final baseline fitting data to form a data difference, and the data difference is configured as training near-infrared spectral data.

[0017] When generating initial baseline estimation data based on near-infrared spectral correction baseline data, the following are included: The near-infrared spectral correction baseline data were sequentially processed by median filtering and polynomial fitting to generate initial baseline estimation data after polynomial fitting. For weights and smoothing constraint terms Then we have:

[0018] in, b The median of all residuals. or , This represents the standard deviation of all residuals in the current near-infrared spectral correction baseline data. , All parameters are over-parameterized.

[0019] The feature extractor within the intermediate model for parameter prediction includes a channel augmentation feature extraction unit and a main feature extraction unit connected in sequence. The main feature extraction unit adopts a U-Net network architecture. The feature extraction main unit includes an encoder network and a decoder network adapted and connected to the encoder network. The encoder network includes several encoder layers connected in sequence, and the channel expansion feature extraction unit is connected to the encoder layer located at the opening of the U-Net network within the encoder network; The decoder network includes several decoder layers connected in sequence. The number of decoder layers in the decoder network is consistent with the number of encoder layers in the encoder network. The decoder layer at the opening of the U-Net network is adapted to connect with the regression prediction head. For any encoder layer, the encoder layer receives the feature to be encoded and first performs wavelet filtering decomposition on the feature to be encoded to generate coded detail components and coded approximate components after wavelet filtering decomposition. Then, the coded detail components and coded approximate components are concatenated by channels and downsampled after channel concatenation to generate the encoded feature. For any decoder layer, the decoder layer receives the feature to be decoded, and first performs upsampling and merging operations on the feature to be decoded, and then performs feature fusion processing after the merging operation to generate the decoded feature after feature fusion processing; Within the main feature extraction unit, the bottom-most decoder layer directly loads the encoded features into the corresponding decoder layer. In addition to the decoder layer located at the bottom layer, the input of the other decoder layers is connected to the output of a feature extraction splicer, which receives the decoded features output from the previous decoder layer and the encoded features output from the encoder layer that is in the same layer as the current decoder layer.

[0020] The encoder layer includes a multi-scale wavelet decomposition module, an encoder splicer, and an encoder double convolution module, wherein... The multi-scale wavelet decomposition module is used to perform wavelet filtering decomposition on the features to be coded, and generate coded detail components and coded approximation components. The encoded detail component and the encoded approximation component are concatenated through a channel by an encoder splicer to generate encoded spliced ​​component features; The encoding double convolution module downsamples the features of the encoded concatenated components and outputs the encoded features.

[0021] The decoder layer includes a multi-scale wavelet reconstruction module, a decoding linear adder, and a decoding double convolution module, wherein... The multi-scale wavelet reconstruction module is used to upsample the features to be decoded, and the low-frequency component and high-frequency component of the decoded component are generated after the upsampling operation. A decoding linear adder is used to linearly add the decoded low-frequency components and the decoded high-frequency components to form the characteristics of the decoded combined components; The decoding dual convolution module is used to perform feature fusion processing on the decoded and merged component features to generate decoded features.

[0022] The advantages of this invention are: It introduces an incremental learning mechanism for generating feature replay into the field of near-infrared spectroscopy analysis, providing an innovative solution to the problem of data distribution drift under dynamic operating conditions. When changes in sampling conditions, instrument status, or batch source lead to inconsistencies between the distribution of new and historical data, it eliminates the need to access and store massive amounts of historical raw data. Instead, it utilizes old task data and a trained generative adversarial reference network model to replay the spectral feature distribution of the old task data. By using the incremental set and generated pseudo-features together to optimize the parameter prediction reference model, the model can effectively consolidate and memorize old knowledge while learning new knowledge, thus successfully mitigating the catastrophic forgetting problem during inference. This achieves efficient adaptive updating of the quality parameter detection model in dynamic environments, extending the lifespan of the analysis model.

[0023] The feature extractor within the quality parameter detection model integrates wavelet decomposition and wavelet reconstruction to comprehensively mine the hierarchical features of the spectrum, enhancing the feature capture capability of the quality parameter detection model, thereby improving the prediction accuracy of the quality parameters of the tested material and the robustness of the model inference.

[0024] To address the random noise interference introduced by portable devices during near-infrared spectral acquisition, an adaptive asymmetric least squares baseline correction based on Huber regularization is employed. This enables refined processing of spectra at different noise levels, significantly improving the stability of near-infrared spectral data. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of an embodiment of the construction of the quality parameter detection model of the present invention.

[0026] Figure 2 This is a block diagram of an embodiment of the quality parameter detection model constructed by the present invention.

[0027] Figure 3 This is a schematic diagram of an example of obtaining near-infrared spectral data of the same sample at different sampling distances.

[0028] Figure 4 This is a schematic diagram of an embodiment for collecting near-infrared spectral data of different samples at different sampling distances.

[0029] Figure 5 for Figure 4 A schematic diagram of the probability density function corresponding to the content of quality parameters in mid-to-near infrared spectral data.

[0030] Figure 6 This is a schematic diagram of an embodiment of preprocessing using standard normal transformation and asymmetric least squares baseline correction.

[0031] Figure 7 This is a schematic diagram of an embodiment of the present invention that utilizes standard normal transformation processing and Huber regularization-based adaptive asymmetric least squares baseline correction processing.

[0032] Figure 8 This is a schematic diagram of an embodiment of the dual convolution module of the present invention.

[0033] Figure 9 This is a schematic diagram of an embodiment of the present invention for joint optimization of regression prediction heads.

[0034] Figure 10 This is a schematic diagram of an embodiment of the present invention for training a generative adversarial reference network model.

[0035] Figure 11 This is a schematic diagram of an embodiment of the feature discriminator of the present invention.

[0036] Figure 12 This is a schematic diagram of an embodiment of the feature generator of the present invention.

[0037] Figure 13 This is a schematic diagram illustrating an embodiment of the relationship between predicted and standard values ​​when predicting Fe2O3 content in bauxite according to the present invention.

[0038] Figure 14 This is a schematic diagram comparing the results of different methods of the present invention on a stage task test set.

[0039] Figure 15 This is a schematic diagram illustrating a comparison of the difference between the final model result at each stage of this invention and Oracle. Detailed Implementation

[0040] The present invention will be further described below with reference to specific accompanying drawings and embodiments.

[0041] To effectively construct a quality parameter detection model, reduce the deployment time and computational cost of the model in practical applications, effectively mitigate catastrophic forgetting during inference, and achieve efficient adaptive updates in dynamic environments, this invention provides a method for constructing a quality parameter detection model based on incremental learning. Specifically, the method includes: A basic parameter prediction model is constructed for predicting quality parameters of the substance to be tested, and a training data set corresponding to the substance to be tested is provided, wherein... The training data group includes at least two training datasets, and the spectral distribution characteristics of all training datasets are at least incompletely matched. Each training dataset includes several training samples related to the substance to be tested. For any training sample, the training sample includes training near-infrared spectral data and training parameter labels characterizing the quality parameter state of the substance under test, wherein the training parameter labels are positively correlated with the training near-infrared spectral data; The basic parameter prediction model is trained using the training dataset within the training data set, driven by incremental learning. This trained model then generates a quality parameter detection model. When performing incremental learning-driven model training, the following are included: Configure one training dataset as the base set, and use the remaining training datasets as incremental sets. The basic model for parameter prediction is trained using the base set. Then, incremental learning is performed once using the incremental set. After all incremental learning is performed, a quality parameter detection model is generated.

[0042] It should be understood that the quality parameter detection model constructed in this invention is mainly used to predict the quality parameters of the substance to be tested. The type of substance to be tested can be selected as needed, specifically to meet the requirements for quality parameter prediction using near-infrared spectroscopy. For example, the substance to be tested can be bauxite or other types of substances for which near-infrared spectral data can be collected. When the substance to be tested is bauxite, the quality parameter can be the Fe2O3 content in the bauxite. Of course, it can also be the alumina content or silicon dioxide content in the bauxite. The specific quality parameter can be determined according to the needs. In addition, when the substance to be tested is of other types, the corresponding quality parameter type can be determined accordingly, which will not be illustrated here.

[0043] To construct a quality parameter detection model, by Figure 1 Therefore, a basic parameter prediction model should be constructed first. This model should be capable of processing near-infrared spectral data to predict the quality parameters of the substance under test. After constructing the basic model, it is trained using the training data set to generate a quality parameter detection model. Generally, the quality parameter detection model shares the same network architecture as the basic parameter prediction model. The network architecture of the quality parameter detection model can be found in the following description.

[0044] To effectively mitigate catastrophic forgetting during inference and achieve efficient adaptive updates in dynamic environments, in one embodiment of this invention, when constructing the generative quality parameter detection model, the provided training data set should include at least two training datasets. Incremental learning can be driven by at least two training datasets. In specific implementation, the number of training datasets in the training data set can be selected as needed to meet the requirements of incremental learning. It should be noted that the training datasets in the training data set can be near-infrared spectral data acquired in batches, and the spectral distribution characteristics of the training datasets in the training data set should be at least not perfectly matched. For example, if two training datasets have the same spectral distribution characteristics, or if the spectral distribution characteristics of one training dataset are completely covered by the spectral distribution characteristics of another training dataset, even if the two training datasets are acquired in batches, only one training dataset should be used for incremental learning.

[0045] Generally, each training dataset may include several training samples. Each training sample should include a training near-infrared spectral data set and a training parameter label. Each training parameter label characterizes the quality parameter corresponding to the training near-infrared spectral data set within its respective training sample. For example, if the substance to be tested is bauxite, as mentioned above, training near-infrared spectral data can be generated by sampling the bauxite sample using near-infrared spectroscopy and performing necessary processing. After measuring the Fe2O3 content in the bauxite sample, the obtained Fe2O3 content can be configured as the corresponding training parameter label. That is, both the training parameter label and the training near-infrared spectral data correspond to the Fe2O3 content of the bauxite sample. Therefore, within the same training sample, the training parameter label and the training near-infrared spectral data are positively correlated.

[0046] It should be noted that the training samples in each training dataset are related to the substance to be tested. Specifically, the training samples in each training dataset represent the near-infrared spectral data and corresponding parameter labels of the same substance to be tested. For example, when preparing a training dataset, multiple samples of the substance to be tested can be prepared first. Then, near-infrared spectral data of each sample can be collected and processed in the same sampling scenario to form the training near-infrared spectral data in each training sample. In addition, the training parameter labels in different training samples should not be completely the same.

[0047] In practice, a portable near-infrared spectrometer can be used to acquire near-infrared spectral data for each sample. After processing the acquired near-infrared spectral data, corresponding training near-infrared spectral data can be generated. It should be understood that when preparing a training dataset, the portable near-infrared spectrometer should be placed in the same sampling environment as much as possible. In this case, all training near-infrared spectral data can be considered to be subject to the same external interference, thus minimizing the impact of external interference. Generally, the multiple samples prepared should belong to the same type of analyte, but the quality parameters within different samples may differ.

[0048] Since each training dataset contains several training samples, and each training sample includes one training near-infrared spectral data point, a spectral map can be constructed for each training dataset. The constructed spectral map can be as follows: Figure 3 , Figure 4 , Figure 6 and Figure 7 As shown in the figure, the horizontal axis represents the wavelength of the near-infrared spectrum, and the vertical axis represents the absorbance. It should be noted that the "wavelength" on the horizontal axis represents the arrangement of spectral sampling points. By statistically analyzing the absorbance of all training near-infrared spectral data within the training dataset, the spectral distribution characteristics of each training dataset can be obtained. Subsequently, it can be determined whether the spectral distribution characteristics of any two training datasets match. The specific methods for determining the spectral distribution characteristics and judging whether they are the same can be consistent with existing technologies. If two spectral distribution characteristics are identical, or one spectral distribution characteristic is covered by another, then the two spectral distribution characteristics match; otherwise, the two spectral distribution characteristics do not match. Figure 4 An example of a mismatch in spectral distribution characteristics is shown in the figure.

[0049] As described above, this invention utilizes a training data set to train a basic parameter prediction model based on incremental learning, thereby obtaining a quality detection model. Specifically, during model training, one training dataset should be configured as the base set, and the remaining training datasets should each be configured as an incremental set. For example, if the training data set includes three training datasets, one training dataset can be configured as the base set, and the other two training datasets can be configured as incremental sets. In addition, the first training dataset can be configured as the base set, and training datasets with mismatched spectral distribution characteristics encountered in subsequent work or inference can be configured as incremental sets. Of course, the selection of the base set and incremental sets can also be determined according to the actual situation, which will not be elaborated here.

[0050] After selecting the base set and the increment set, the basic model for parameter prediction should first be trained using the base set. Then, incremental learning should be performed once using each increment set to generate the quality parameter detection model after all incremental learning operations are performed. As explained above, when there are two increment sets, the first increment set should be used for incremental learning first, and after the incremental learning is completed, the second increment set should be used for incremental learning. When performing incremental learning using the second increment set, it should be performed based on the incremental learning performed on the first increment set. When the increment set has other quantities, please refer to the explanation here, which will not be illustrated in detail here.

[0051] It is understandable that after generating the quality parameter detection model, the construction of the quality parameter detection model is completed. Subsequently, the quality parameter detection model can be deployed and used to predict the quality parameters of the substance under test. When predicting the quality parameters of the substance under test, the near-infrared spectral data of the substance under test should be obtained and loaded into the quality parameter detection model. The method for obtaining the near-infrared spectral data can refer to the above description of the method for obtaining training near-infrared spectral data. After loading the near-infrared spectral data into the quality parameter detection model, the model can process the near-infrared spectral data and output the corresponding predicted quality parameter values, thus achieving quality parameter prediction. The method for processing the near-infrared spectral data of the quality parameter detection model can be combined with... Figure 2 The model architecture for the quality parameter detection model is determined, and the following explanation can be used as a reference.

[0052] To ensure high accuracy in the predicted quality parameters, the spectral distribution characteristics of the near-infrared spectral data of the substance under test should match at least one spectral distribution characteristic belonging to the training data set. The matching of two spectral distribution characteristics can be found in the explanation of spectral distribution characteristic matching above. It should be understood that if the spectral distribution characteristics of the near-infrared spectral data do not match any of the spectral distribution characteristics of the training dataset, the accuracy of the predicted quality parameters will be low when using the current quality parameter detection model to process the near-infrared spectral data, making it difficult to meet actual prediction requirements.

[0053] To achieve efficient adaptive updates in dynamic environments, a new training dataset can be generated based on the near-infrared spectral data to be tested. The generated training dataset is then used to incrementally learn the quality parameter detection model. In this case, the generated training dataset should be used as an increment set. After performing one incremental learning operation, a new quality parameter detection model can be generated. The incremental learning method can be referred to the above description of incremental learning. That is, after generating the quality parameter detection model, when encountering new spectral distribution features, a training dataset can be prepared as an increment set. Subsequently, the increment set is used to incrementally learn the quality parameter detection model.

[0054] As explained above, incremental learning in this scenario can update the quality parameter detection model without accessing historical data. Since the updated model covers more spectral distribution features, it improves the model's adaptability and generalization in predicting the quality parameters of the tested substance. Furthermore, the historical data that does not require access specifically refers to the base set and corresponding incremental set used to generate the quality parameter detection model.

[0055] Generally, when predicting quality parameters, a batch of near-infrared spectral data to be tested is acquired. Based on this batch of data, the spectral distribution characteristics of the batch can be determined, and the matching of these spectral distribution characteristics can then be assessed. When incremental learning is required, each piece of near-infrared spectral data to be tested is configured as training data, and existing techniques are used to determine the corresponding training parameter label for each piece of data. This constitutes a training sample, and all training samples form a training dataset, which can then be used for incremental learning. Therefore, this incremental learning is compatible with existing quality parameter prediction operations. However, the difference lies in the need to use commonly used techniques in this field to measure and determine the quality parameter value corresponding to each piece of near-infrared spectral data to be tested. In other words, the original inference operation should be transformed into the operation of creating a training dataset to enable further incremental learning.

[0056] It should be noted that the creation of a new training dataset based on the near-infrared spectral data to be tested is primarily to form an incremental set, and to further incrementally learn the existing quality parameter detection model to generate a new quality parameter detection model. Compared to directly using a training dataset to train a model and obtain a quality parameter detection model, the incremental learning-driven quality parameter detection model of this invention can have broad coverage and generalization ability as incremental learning continues, thus meeting the needs of subsequent quality parameter prediction. Furthermore, as incremental learning continues, it can eventually cover the quality parameter prediction needs of many scenarios, reducing the time and computational cost of deploying the quality parameter detection model in practical applications, effectively mitigating catastrophic forgetting during inference, and achieving efficient adaptive updates in dynamic environments.

[0057] As explained above, the training datasets within the aforementioned training data set can be acquired all at once, or generated during the work process based on batches of near-infrared spectral data to be inspected. The method of creating the training datasets can be selected as needed. When multiple training datasets are created in a similar manner, and incremental learning is used to generate new quality parameter detection models, the quality parameter detection models can cover more spectral distribution features. This can meet the needs of predicting the quality parameters of substances to be inspected in common working scenarios and improve the accuracy and reliability of the output quality parameter prediction values.

[0058] In one embodiment of the present invention, generating training near-infrared spectral data for each training sample includes: Near-infrared spectral source data corresponding to each training near-infrared spectral data is acquired, and the near-infrared spectral source data is preprocessed to generate training near-infrared spectral data after preprocessing. The near-infrared spectral source data is generated through near-infrared spectral acquisition; Preprocessing of near-infrared spectral source data includes standard normal transformation and / or Huber-based adaptive asymmetric least squares baseline correction. When the preprocessing of near-infrared spectral source data includes both standard normal transformation and Huber-based adaptive asymmetric least squares baseline correction, then the near-infrared spectral source data is subjected to standard normal transformation and Huber-based adaptive asymmetric least squares baseline correction sequentially. When performing Huber-regularized adaptive asymmetric least squares baseline correction, we have:

[0059] in, For the final baseline estimation data, To train near-infrared spectral data, the total number of wavenumber points corresponding to near-infrared light spectral source data, The residual at the nth wavenumber point, This represents the weight corresponding to the nth wavenumber point. For the smoothing constraint term corresponding to the near-infrared light spectral source data, The absorbance at the nth wavenumber point within the near-infrared spectral correction base data. The absorbance at the nth wavenumber point within the initial baseline estimation data generated based on near-infrared spectral correction baseline data. For absorbance The second-order difference was performed; When the preprocessing performed only includes Huber regularization-based adaptive asymmetric least squares baseline correction, the near-infrared spectral correction basis data is the near-infrared spectral source data. When the preprocessing includes standard normal transformation and Huber regularization-based adaptive asymmetric least squares baseline correction, the near-infrared spectral source data is generated by performing standard normal transformation on the near-infrared spectral source data. Polynomial fitting is performed on the final baseline estimate data to generate the final baseline fit data after polynomial fitting; The near-infrared spectral correction baseline data is subtracted from the final baseline fitting data to form a data difference, and the data difference is configured as training near-infrared spectral data.

[0060] Specifically, the method for acquiring near-infrared spectral source data can be consistent with existing technologies. For example, a portable near-infrared spectrometer can be used to acquire near-infrared spectra of the sample of the substance to be tested, thereby obtaining the corresponding near-infrared spectral source data. During the near-infrared spectral acquisition process, the data acquired by the portable near-infrared spectrometer is often affected by factors such as sampling conditions and batch origin, exhibiting significant distribution drift. Furthermore, due to random errors (such as noise) and systematic errors (such as baseline drift), the acquired spectral data may deviate significantly from the ideal spectral data, affecting the overall data quality and analytical effect, and thus weakening the prediction accuracy and stability of the static quality parameter detection model.

[0061] In order to generate training near-infrared spectral data, the near-infrared spectral source data should be preprocessed. The preprocessing is the necessary processing mentioned above. The purpose of preprocessing is mainly to reduce the noise interference introduced during the near-infrared spectral acquisition process, so as to improve the reliability of the quality parameter detection model and the accuracy and reliability of subsequent quality parameter prediction.

[0062] Specifically, the preprocessing of near-infrared spectral source data includes standard normal transformation and / or Huber-based adaptive asymmetric least squares baseline correction; the choice of preprocessing method can be made as needed. It should be noted that when preprocessing includes both standard normal transformation and Huber-based adaptive asymmetric least squares baseline correction, the near-infrared spectral source data should be processed sequentially using both methods. The specific processing method and procedure for standard normal transformation (SNV) of near-infrared spectral source data can be consistent with existing technologies and will not be elaborated here.

[0063] As explained above, when performing Huber-regularized adaptive asymmetric least squares baseline correction, initial baseline estimation data of the near-infrared spectral correction baseline data should be generated. In one embodiment of the present invention, generating initial baseline estimation data based on the near-infrared spectral correction baseline data includes: The near-infrared spectral correction baseline data were sequentially processed by median filtering and polynomial fitting to generate initial baseline estimation data after polynomial fitting. For weights and smoothing constraint terms Then we have:

[0064] in, b The median of all residuals. or , This represents the standard deviation of all residuals in the current near-infrared spectral correction baseline data. , All parameters are over-parameterized.

[0065] When generating initial baseline estimation data using median filtering and polynomial fitting, commonly used methods can be employed, such as: Here, This represents the median filtering operation. This represents a polynomial fitting operation. Statement No. Near-infrared spectral correction basis data, This indicates the window size for median filtering; here, the window size for median filtering is 11. This indicates the degree of the polynomial in the polynomial fitting; in this case, the degree of the polynomial in the polynomial fitting is 3.

[0066] Specifically, the total number of wavenumber points corresponding to the training near-infrared spectral data and near-infrared spectral source data. The hyperparameters can be determined based on the settings for near-infrared spectral acquisition using a portable near-infrared spectrometer. Hyperparameters All settings can be selected based on experience, such as hyperparameters. It can be set to a positive number less than 10. The specific value can be selected as needed, and examples will not be given here.

[0067] It should be noted that weight This invention utilizes Huber weights and is based on adaptive asymmetric least squares baseline correction processing with Huber regularization. It generates initial baseline estimation data through median filtering and polynomial fitting, and then introduces weights. By using flexible suppression of the residual squared term, robust modeling of noise is achieved. As can be seen from the above calculation method for generating the final baseline estimation data, the adaptive asymmetric least squares baseline correction processing based on Huber regularization of this invention has an adaptive regularization attenuation mechanism based on the residual standard deviation, which can dynamically adjust the smoothing constraint strength and achieve refined processing of spectra at different noise levels. By performing polynomial fitting on the final baseline estimation baseline, small oscillations caused by error propagation can be suppressed, thereby effectively eliminating scattering effects and suppressing random noise interference. Dynamically processing different noise levels and adjusting the smoothing strength can significantly improve the stability of the obtained training near-infrared spectral data / near-infrared spectral data to be tested.

[0068] It should be noted that during inference, each near-infrared spectral data to be tested should generally also be generated after the above preprocessing. At this time, the predicted value of the quality parameter output by the quality parameter detection model has high accuracy and reliability. In addition, when performing polynomial fitting on the final baseline estimation data, the degree of polynomial fitting is consistent with the degree of polynomial fitting performed on the initial baseline estimation data.

[0069] Figure 6 The diagram illustrates an embodiment of preprocessing using SNV processing and asymmetric least squares baseline correction processing. Figure 7 The figure shows a schematic diagram of an embodiment of the present invention using SNV processing and Huber regularization-based adaptive asymmetric least squares baseline correction processing. As can be seen from the figure, the preprocessing method of the present invention can effectively eliminate scattering effects and suppress random noise interference, and dynamically process different noise to adjust the smoothing intensity, thereby significantly improving the stability of near-infrared spectral data.

[0070] In one embodiment of the present invention, during basic training, the parameter prediction basic model and the generative adversarial network basic model are trained using a basic set, so as to generate the parameter prediction basic model and the generative adversarial network respectively after training. When performing the first incremental learning, the parameter prediction base model, the generative adversarial base network, and the base label information of the base training set are configured as the parameter prediction reference model, the generative adversarial reference network model, and the reference label information, respectively. The base label information is generated based on the base set. After performing incremental learning once using an incremental set, a parameter prediction reference model, an adversarial reference network model, and reference label information are generated for the next incremental learning. After performing the last incremental learning, a quality parameter detection model is generated based on the parameter prediction reference model after the incremental learning.

[0071] As explained above, when implementing the incremental learning driven by this invention, basic training should be performed first, followed by at least one incremental learning iteration. During basic training, the basic parameter prediction model is trained using a base set. The method of training the basic parameter prediction model using the base set is the same as the training method and process for neural networks in the prior art, and will not be elaborated further here.

[0072] It should be understood that some necessary training conditions should be configured when performing basic training. For example, when the basic parameter prediction model is built based on the PyTorch framework, the necessary training conditions may include: using an Adam optimizer with a regularization weight of 0.0002, a batch size of 32, a maximum number of iterations of 100, a learning rate of 0.003, a learning rate decay of 0.9 every 50 epochs, and the loss function used for basic training can be the mean absolute error. Based on the configured training conditions, the basic parameter prediction model is trained using the base set so that the basic parameter prediction model can be obtained after training reaches the target state.

[0073] Since the training parameter labels for each training sample in the base set are known, the mean and variance corresponding to all training parameter labels in the base set can be determined using commonly used statistical methods. These determined mean and variance can then be configured as base label information. Therefore, the base label information can characterize the statistical state of the training parameter labels within the base set. The methods for generating reference label information described below can all be referenced here.

[0074] As explained above, incremental learning may encounter situations where the base set and past incremental sets are inaccessible. To address the catastrophic forgetting problem during inference, this invention employs a generative feature replay mechanism. The core idea is to learn and simulate the feature distribution of old task data through an independently trained generative adversarial network (GAN) model without accessing historical raw data. When new task data arrives, the feature generator within the GAN model replays a large number of pseudo-features consistent with the spectral feature distribution of the old task data. These pseudo-features are then mixed with the real features of the new task data and used together to update the parameter prediction reference model during incremental learning. It should be noted that the "old task data" here refers to the corresponding training samples of the base set and incremental set used before this incremental learning, while the "new task data" refers to the incremental set used during this incremental learning process. The terms "old task data" and "new task data" have the same meaning in the following explanations and can be understood with reference to this description.

[0075] To achieve the aforementioned generation of pseudo-features via a generative adversarial network (GAN), during basic training, in addition to constructing a basic parameter prediction model, a basic GAN model should also be built. After training the basic parameter prediction model using a base set, the basic GAN model is then trained using the base set to obtain the final GAN. The training of the basic GAN model specifically involves a game-like adversarial training process. The basic GAN model should include a feature generator and a feature discriminator. The following section will discuss this in detail. Figures 10-12 The case of generative adversarial basic network model is explained in detail.

[0076] A feature generator typically consists of at least three sequentially connected linear feature generation layers, with a LeakyReLU (Leaky Rectified Linear Unit) activation function placed between adjacent linear feature generation layers. Figure 12 The figure shows a schematic diagram of an embodiment of a feature generator. In the figure, the three feature generation linear layers are feature generation linear layer SL1, feature generation linear layer SL2 and feature generation linear layer SL3. A Leaky ReLU activation function is set between feature generation linear layer SL1 and feature generation linear layer SL2, and a Leaky ReLU activation function is also set between feature generation linear layer SL2 and feature generation linear layer SL3.

[0077] During operation, the feature generator maps the input noise features and sampling labels to pseudo-features, striving to generate pseudo-features that are indistinguishable from real features. The details regarding noise features and sampling labels are explained below. The following section combines... Figure 12 The structure of the feature generator is shown, and its operation is explained in detail. First, the noise features and sampling labels are concatenated into a fused vector, which is then input into the feature generation linear layer SL1. SL1 projects the received fused vector and expands it into a higher-dimensional "hidden feature space," preparing for subsequent complex processing. The LeakyReLU activation function following SL1 introduces crucial non-linear computational capabilities, enabling the feature generator to learn complex patterns that simple linear transformations cannot express.

[0078] The feature generation linear layer SL2 typically serves as the hidden space layer. SL2 deeply refines and reorganizes the features calculated by the Leaky ReLU activation function, learning the inherent and subtle relationships between features. Then, the Leaky ReLU activation function after passing through SL2 enters the feature generation linear layer SL3, which serves as the output layer. SL3 precisely maps these highly refined internal features to the target dimension required for the final "pseudo-features," completing the final shaping and output. Through this alternating stacking of feature generation linear layers and Leaky ReLU activation functions, the feature generator can efficiently map a simple combination of 'noisy features + sampling labels' into a complex pseudo-feature space that is highly similar to the distribution of real features, thus achieving its purpose of generating pseudo-features. The feature discriminator generally includes at least two sequentially connected feature discrimination linear layers, with a Leaky ReLU activation function placed between adjacent feature discrimination linear layers. Figure 11 The figure shows an embodiment of a feature discriminator, in which the two linear layers are feature discrimination linear layer PL1 and feature discrimination linear layer PL2.

[0079] When training a Generative Adversarial Network (GAN) basic model, the feature generator and the feature discriminator compete with each other and jointly update their parameters. The feature generator strives to improve its forgery skills to deceive the feature discriminator, while the feature discriminator continuously improves its ability to detect forgeries. To balance this game process, in one embodiment of the present invention, an asymmetric training strategy is adopted, such as training the feature generator for 5 rounds and then training the feature discriminator for 1 round. In this case, during the 5-round training cycle of the GAN basic model, the feature generator's parameters are updated in each round of training, but only in the 5th round of the training cycle. The training situation in each round is consistent with the prior art, and will not be described in detail here.

[0080] Figure 10 The document illustrates one embodiment of training a basic generative adversarial network model. When training the generative adversarial network during basic training, then... Figure 10The feature extractor Fn in the model is a feature extractor within the basic parameter prediction model. Since no generative adversarial networks are available, Figure 10 In this context, the feature generator Gn-1 and the feature generator Gn should be the same constructible feature generator, that is, the feature generator of the initial state. Furthermore, Figure 10 In this context, feature generator Gn-1 refers to the initialization of feature generator Gn, specifically referring to the network parameter configuration based on feature generator Gn-1 as the initial parameters of feature generator Gn.

[0081] When training the generative adversarial network basic model, a noise feature should be randomly initialized. This noise feature should be consistent with the training near-infrared spectral data of this invention, such as having the same wavenumber points. The absorbance of each wavenumber point can be randomly generated. The randomly initialized noise feature and the training parameter labels from a training sample within the base set are loaded into the feature generator, which generates a training feature. Simultaneously, the training near-infrared spectral data from the current training sample is loaded into the parameter prediction basic model, where the feature extractor extracts the corresponding ground truth features. Afterward, the training parameter labels-training features and training parameter labels-ground truth features are loaded into the feature discriminator. Figure 10 The label information in the data is the training parameter label.

[0082] When training a generative adversarial network (GAN) basic model adversarially, it is essentially a cyclical competition and co-evolution between the feature generator and the feature discriminator. In a complete training cycle, the feature discriminator is trained first. Specifically, the feature discriminator receives both "real feature-training parameter label" pairs based on the base set and "pseudo-feature-training parameter label" pairs generated by the feature generator. It scores these two batches of data using the feature discrimination linear layer and label projection structure inside the feature discriminator, and then uses the WGAN loss function to maximize the difference between the real score and the fake score, thereby improving its ability to "distinguish between real and fake". The two batches of data mentioned above specifically refer to the "real feature-training parameter label" pairs and the "pseudo-feature-training parameter label" pairs.

[0083] It should be noted that the above "real feature-training parameter label" pairs and "pseudo feature-training parameter label" pairs are both generated based on a training sample within the base set. For real features and pseudo features, please refer to the corresponding explanations below.

[0084] When employing the aforementioned asymmetric training strategy, with the discriminator's parameters temporarily "frozen," the feature generator is trained five times consecutively. In each training iteration, the feature generator creates a batch of new pseudo-features based on noise features and training parameter labels, using three linear feature generation layers (SL1, SL2, SL3), and feeds them into the discriminator for scoring. Understandably, the feature generator's goal is opposite to the discriminator's; it strives to adjust its parameters to maximize the scores of its pseudo-features, aiming to "fool" the discriminator. Through this asymmetric cyclical adversarial process of "the discriminator learning once, the feature generator training five times," the discriminator's discrimination criteria become increasingly stringent, forcing the feature generator to continuously improve its forgery skills. Ultimately, a dynamic equilibrium is reached, enabling the feature generator to create high-quality pseudo-features that are indistinguishable from genuine features based on given sampling labels and random noise features.

[0085] In practice, when training the generative adversarial basic network model, the number of training rounds can be set. When the number of training rounds is reached, it is considered that the training has reached the target state. At this time, the generative adversarial basic network model that has reached the set number of training rounds can be configured as the generative adversarial basic network model.

[0086] As explained above, after basic training, incremental learning should be performed at least once. This means that during the first incremental learning run, the parameter prediction base model, the generative adversarial network (GAN) base network, and the basic label information of the basic training set should be configured as the parameter prediction reference model, GAN reference network model, and reference label information required for the first incremental learning run, respectively. After performing incremental learning once using an incremental set, the parameter prediction reference model, GAN reference network model, and reference label information required for the next incremental learning run will be generated. For example, after the first incremental learning run, the parameter prediction reference model, GAN reference network model, and reference label information required for the second incremental learning run can be generated. Therefore, as incremental learning continues, the parameter prediction reference model and GAN reference network model can be continuously updated. Of course, since different incremental learning runs use different incremental sets, the corresponding reference label information will also be different.

[0087] Specifically, after performing the final incremental learning, a quality parameter detection model is generated based on the parameter prediction reference model after the incremental learning. It should be understood that the number of incremental learning operations should be consistent with the number of incremental sets. Of course, after performing the final incremental learning, a generative adversarial reference network model and reference label information still need to be formed for use in subsequent incremental learning. Here, "final" specifically refers to the number of incremental sets that have been determined.

[0088] In one embodiment of the present invention, incremental learning includes: Determine the incremental set to be used for this incremental learning, and obtain the latest generated parameter prediction reference model, generative adversarial reference network model, and reference label information before executing this incremental learning; When generating the parameter prediction reference model for the next incremental learning based on the determined incremental set and the obtained parameter prediction reference model, the following steps are included: The parameter prediction reference model is incrementally trained using the determined increment set, wherein the incremental training includes several rounds of incremental iteration. In each round of incremental iteration, the parameter prediction reference model is trained using the incremental set to generate a parameter prediction intermediate model after training. The regression prediction head in the parameter prediction intermediate model is then jointly optimized to form the parameter prediction intermediate model for the next round of incremental iteration after joint optimization. Once the incremental iteration reaches the specified number of rounds, the intermediate parameter prediction model that has reached the specified number of rounds will be configured as the parameter prediction reference model for the next incremental learning. When generating the generative adversarial reference network (GARNR) model required for the next incremental learning, based on the determined incremental set and the GARNR model, the following steps are included: The acquired generative adversarial reference network model is trained using the incremental set, so as to generate the generative adversarial reference network model required for the next incremental learning. When generating reference label information, the training parameter labels of all training samples in the determined incremental set are obtained, and the label mean and label variance corresponding to all training parameter labels are calculated, so as to form the reference label information required for the next incremental learning based on the calculated label mean and label variance.

[0089] When multiple increment sets exist, the order in which incremental learning is performed using each increment set can be selected as needed. For example, the order in which the increment sets are acquired can be used to perform incremental learning. Each increment set can only be used in one incremental learning session. Based on the order of incremental learning, the latest generated parameter prediction reference model, generative adversarial network (GAN) reference model, and reference label information before this incremental learning session can be determined. If this is the first time incremental learning is performed, the latest generated parameter prediction reference model, GAN reference model, and reference label information before this incremental learning session should be the parameter prediction base model, GAN base network, and base label information of the base training set obtained through basic training. If this is the second time incremental learning is performed, the latest generated parameter prediction reference model, GAN reference model, and reference label information should be the parameter prediction reference model, GAN reference model, and reference label information obtained after the first incremental learning session. Other cases follow the same principle, and will not be illustrated here.

[0090] As explained above, after performing incremental learning, a parameter prediction reference model is generated for the next incremental learning iteration. Specifically, when generating the parameter prediction reference model for the next incremental learning iteration, the obtained parameter prediction reference model should be trained using the incremental set from the current incremental learning iteration, and an intermediate parameter prediction model should be generated in each round of incremental iteration. The method of using the incremental set to perform incremental iterations on the obtained parameter prediction reference model in each round is consistent with the method of training the basic parameter prediction model using the base set, the only difference being the loss function used. Figure 9 The image shows an example of training a reference model for parameter prediction.

[0091] In one embodiment of the present invention, when training the obtained parameter prediction reference model using the determined increment set, the incremental training loss function used is:

[0092] in, For incremental training loss function, For the average absolute error loss, For distillation loss, For offset penalty loss, Mean absolute error loss The weight, Distillation loss weights, Offset penalty loss The weight, The training near-infrared spectral data for a training sample within the increment set. To utilize the parameters obtained from incremental iteration during this incremental learning process to predict the spectral features extracted by the intermediate model, To utilize the latest parameters generated before this incremental learning to predict the spectral features extracted from the reference model, Spectral characteristics With spectral characteristics The square of the L2 norm of the difference, Spectral characteristics With spectral characteristics The expected value of the square of the L2 norm of the difference; This is the first intermediate model for parameter prediction during incremental iteration in this incremental learning process. Each model parameter This is the first parameter prediction reference model generated before this incremental learning. Each model parameter For the first Model parameters The estimated value of the Fisher information matrix, This is a hyperparameter.

[0093] It should be noted that each parameter prediction reference model generally includes a feature extractor and a regression prediction head. The feature extractor extracts the true features of the training near-infrared spectral data, and the regression prediction head outputs the predicted values ​​of the quality parameters. Based on the expression of the loss function above, when training the parameter prediction reference model, feature extraction needs to be performed using both the current parameter prediction reference model and the latest parameter prediction reference model obtained during this incremental learning. For details on using the feature extractor for feature extraction, please refer to the corresponding explanation below.

[0094] In one embodiment of the present invention, distillation loss is introduced into the incremental training loss function. and offset penalty loss This can enhance the stability and robustness of the parameter prediction reference model during incremental learning. Specifically, it is achieved through distillation loss. As can be seen from the expression, by calculating the Euclidean distance between the spectral features extracted by the current parameter prediction reference model and the spectral features extracted by the acquired parameter prediction reference model, and by introducing the distillation loss function constraint, the spectral features extracted by the current parameter prediction reference model are kept consistent with the spectral features extracted by the acquired parameter prediction reference model in the representation space. This not only preserves the effective representation of the old task data and avoids the catastrophic forgetting of the feature extractor itself within the parameter prediction reference model, but also provides a stable feature baseline for incremental learning and a stable feature foundation for incremental learning.

[0095] Introducing offset penalty loss Subsequently, knowledge protection of old task data can be strengthened. Specifically, during the parameter update process of the parameter prediction reference model, offset penalty loss can be used. Constraints can be imposed on key weights to prevent them from being over-modified during incremental learning training. Here, key weights specifically refer to the parameters most important to task performance; the details of these key weight parameters are consistent with existing techniques and will not be elaborated upon here.

[0096] It should be noted that hyperparameters Generally, empirical values ​​are used, the first... Model parameters Fisher information matrix estimate It can be used to measure the first Model parameters The regularization coefficients, which assess the importance of old task data, can also be used to balance the relative importance of new and old task data. Calculate the estimated Fisher information matrix. Then:

[0097] in, This represents the total number of training samples within the old task data. It is the training near-infrared spectral data of the q-th training sample within the old task data. The training parameter labels are the training parameters for the q-th training sample within the old task data. This is the latest parameter prediction reference model generated before this incremental learning, applied to the training near-infrared spectral data under the model parameter set θ. Predict the training parameter labels. The conditional probability, This is the set of model parameters for the latest parameter prediction reference model generated before this incremental learning. For the model parameter set Calculate the log-likelihood. Indicates the first Model parameters The partial derivatives of .

[0098] It should be noted that the above calculation of the Fisher information matrix estimate... In this context, "old task data" refers only to one set in the base set or incremental set before the current incremental learning. For example, when performing incremental learning for the first time, the old task data is the base set. Other cases can be referred to here for explanation. Once the old task data is determined, the value of the total number Q can be determined.

[0099] In practical implementation, during incremental learning, the necessary training conditions can be consistent with the training conditions for the basic parameter prediction model using the base set mentioned above. Under the training conditions and the incremental training loss function, the parameter prediction reference model in this incremental learning can be incrementally trained to obtain the parameter prediction reference model required for the next incremental learning. During incremental training, the specified number of incremental iterations can be selected as needed, such as being consistent with the number of iterations set in the training conditions mentioned above. In each incremental iteration, an intermediate parameter prediction model is first obtained. Then, the regression prediction heads within the intermediate parameter prediction model are jointly optimized, resulting in a new intermediate parameter prediction model. It can be understood that in the next incremental iteration, the intermediate parameter prediction model finally formed in the current incremental iteration should be trained, that is, the training operation of each incremental iteration is repeated until the specified number of incremental iterations is reached.

[0100] In one embodiment of the present invention, joint optimization of the regression prediction head within the intermediate model for parameter prediction includes: Generate a feature-label set required for joint optimization of the regression prediction head, wherein the feature-label set includes several pseudo-feature-pseudo-label pairs and several real feature-real-label pairs; When generating true feature-true label pairs, the following steps are included: The feature extractor within the intermediate model of parameter prediction is used to extract features from a training near-infrared spectral data to generate true features; Based on the training near-infrared spectral data corresponding to the real features, the training parameter labels that correspond to the near-infrared spectral data are determined, and the determined training parameter labels are configured as real labels. The training near-infrared spectral data belongs to a training sample in the incremental set used to perform this incremental learning. When generating pseudo-feature-pseudo-label pairs, the following is included: Based on the reference label information obtained from this incremental learning process, pseudo-labels are generated through sampling. The pseudo-labels are loaded into the Generative Adversarial Reference Network (GARN) model, which then generates pseudo-features corresponding to the pseudo-labels. The regression prediction head is optimized and updated using the feature-label set to generate an optimized regression prediction head.

[0101] To enable the regression prediction head to map the spectral distribution features learned in this incremental learning, the regression prediction heads within the intermediate parameter prediction model should be jointly optimized. Figure 9 The figure illustrates one embodiment of joint optimization of regression prediction heads. Figure 9 The domain category in the text refers to the spectral distribution characteristics mentioned above, that is... Figure 9This illustrates an embodiment where different sampling distances in a portable near-infrared spectrometer result in different spectral distribution characteristics. Figure 9 As can be seen, when performing joint optimization, a feature-label set should be generated. The following explains the process of generating the feature-label set.

[0102] Figure 9 In the diagram, the regression prediction head Rn-1 is the regression prediction head obtained from the parameter prediction reference model during this incremental learning process, while the regression prediction head Rn is the regression prediction head from the intermediate model used for parameter prediction during this incremental learning process. The feature generator in the diagram is the feature generator obtained from the generative adversarial reference network model during this incremental learning process.

[0103] It should be noted that when generating true feature-true label pairs, training samples from the incremental set of this incremental learning should be used. For example, a training near-infrared spectral data from the training samples can be loaded into the intermediate parameter prediction model, and the feature extractor within the intermediate parameter prediction model can be used to extract features to obtain true features. Figure 9 In this context, feature extractor Fn is the feature extractor of the intermediate parameter prediction model, and feature extractor Fn-1 is the feature extractor within the reference parameter prediction model obtained during this incremental learning process. Generally, the feature extractor of the intermediate parameter prediction model can be used to extract the true features of all training near-infrared spectral data in the current incremental set. Subsequently, the corresponding training parameter labels are used as true labels, thus forming true feature-true label pairs.

[0104] When generating pseudo-feature-pseudo-label pairs, the reference label information obtained in this incremental learning should be used. As explained above, the reference label information includes the mean and variance of the corresponding training dataset. Therefore, a normal distribution of quality parameters can be generated based on the mean and variance in the reference label information. Subsequently, a pseudo-label can be obtained by randomly sampling the formed normal distribution. Figure 9 In this context, the mean and variance loaded into the feature generator Gn-1 specifically refer to the pseudo-labels generated based on the reference label information. After generating the pseudo-labels, they should be loaded into the feature generator within the generative adversarial reference network model, such as... Figure 9 As shown above, in order to generate pseudo-features, as explained above, noisy features should also be added to the feature generator. Figure 9 Z~[0,1] represents the noise feature formed using the [0,1] normal distribution. For details on the noise feature, please refer to the corresponding explanation above. Based on the noise feature and the pseudo-label, the feature generator can generate corresponding pseudo-features. Using the pseudo-features and their corresponding pseudo-labels, a corresponding pseudo-feature-pseudo-label can be generated.

[0105] It is understandable that the joint optimization here specifically refers to simultaneously optimizing and updating the regression prediction head using both pseudo-feature-pseudo-label pairs and true feature-true label pairs, so that an optimized regression prediction head can be obtained after the update. The following example illustrates the optimization and update of the regression prediction head: First, the real features and pseudo features are concatenated, and the real labels and pseudo labels are concatenated. During concatenation, the number of real features and pseudo features is the same as the batch size set in the training conditions mentioned above. For example, if the batch size of 32 is used, the concatenated real features and pseudo features will form a total of 64 mixed features. Similarly, a total of 64 mixed labels will be formed. It should be noted that the arrangement of real labels and pseudo labels in the mixed labels corresponds directly to the real features and pseudo features in the mixed features. That is, each corresponds to a pseudo feature-pseudo label pair and a real feature-real label pair.

[0106] The batch of mixed features formed above is completely fed into the regression prediction head. After calculation, the corresponding predicted value for each input feature is output. In the subsequent loss calculation stage, the L1 regression loss function can be used to accurately measure the overall difference or error between this batch of predicted values ​​and the concatenated "true / false" label values. Based on the quantized error value, the backpropagation process is driven, which calculates the gradient of the loss relative to all parameters (such as weights and biases) within the regression prediction head along the network path of the regression prediction head.

[0107] Finally, in the parameter update step, the optimizer uses this gradient information to iteratively update all parameters of the regression prediction head, optimizing them in a direction that better fits the feature-label set, thereby reducing future prediction errors.

[0108] It should be understood that, based on the feature extractor within the parameter prediction intermediate model and the optimized regression prediction head, a reference prediction model can be generated after this incremental learning process, which will serve as the reference prediction model for the next incremental learning iteration. The reference label information generated after this incremental learning process can be referenced in the above-mentioned instructions for generating reference label information; the difference is that the corresponding reference label information should be generated using the incremental set from this incremental learning process.

[0109] Furthermore, in this incremental learning, the method of using the incremental set to train the generative adversarial reference network model and generate the generative adversarial reference network model required for the next incremental learning can refer to the above description of training the basic generative adversarial network model using the base set. The difference is that the training dataset used is different, and the specific training process will not be repeated here.

[0110] In one embodiment of the present invention, the feature extractor within the intermediate model for parameter prediction includes a channel augmentation feature extraction unit and a feature extraction master unit connected in sequence. The feature extraction master unit adopts a U-Net network structure. The feature extraction main unit includes an encoder network and a decoder network adapted and connected to the encoder network. The encoder network includes several encoder layers connected in sequence, and the channel expansion feature extraction unit is connected to the encoder layer located at the opening of the U-Net network within the encoder network; The decoder network includes several decoder layers connected in sequence. The number of decoder layers in the decoder network is consistent with the number of encoder layers in the encoder network. The decoder layer at the opening of the U-Net network is adapted to connect with the regression prediction head. For any encoder layer, the encoder layer receives the feature to be encoded and first performs wavelet filtering decomposition on the feature to be encoded to generate coded detail components and coded approximate components after wavelet filtering decomposition. Then, the coded detail components and coded approximate components are concatenated by channels and downsampled after channel concatenation to generate the encoded feature. For any decoder layer, the decoder layer receives the feature to be decoded, and first performs upsampling and merging operations on the feature to be decoded, and then performs feature fusion processing after the merging operation to generate the decoded feature after feature fusion processing; Within the main feature extraction unit, the bottom-most decoder layer directly loads the encoded features into the corresponding decoder layer. In addition to the decoder layer located at the bottom layer, the input of the other decoder layers is connected to the output of a feature extraction splicer, which receives the decoded features output from the previous decoder layer and the encoded features output from the encoder layer that is in the same layer as the current decoder layer.

[0111] Figure 2 The diagram illustrates an embodiment of the basic parameter prediction model architecture. It is understood that the basic parameter prediction model, the intermediate parameter prediction model, and the quality parameter prediction detection model all share the same architecture. The channel augmentation feature extraction unit can employ a dual convolutional module. Figure 2 In this context, the dual convolutional module DC0 is the channel augmentation feature extraction unit. Figure 2 In this context, NIR refers to either the training near-infrared spectral data or the near-infrared spectral data to be tested. Since the feature dimension of both the training and testing near-infrared spectral data is N×1, the number of channels should be expanded using a channel expansion feature extraction unit to meet the subsequent processing requirements of the feature extractor. Furthermore, when the channel expansion feature extraction unit uses a dual convolution module, it can simultaneously extract abstract features from both the training and testing near-infrared spectral data.

[0112] Figure 2 The figure illustrates an embodiment of a feature extractor using a U-Net network. The encoder network in the figure includes four encoder layers and four decoder layers. The four encoder layers are Encoder1, Encoder2, Encoder3, and Encoder4; the four decoder layers are Decoder1, Decoder2, Decoder3, and Decoder4. Figure 2 In the U-Net network, the encoder layer Encoder1 and the decoder layer Decoder1 are located at the opening position, the encoder layer Encoder4 is the bottom encoder layer, and the decoder layer Decoder4 is the bottom decoder layer.

[0113] In practice, each encoder layer adopts the same structure, and each decoder layer also adopts the same structure. The encoded features output by the encoder layer Encoder4 are directly loaded into the decoder layer Decoder4 as the features to be decoded received by the decoder layer Decoder4. Figure 2 In the above, for encoder layer Encoder1, the feature to be encoded should be the feature output by the channel augmentation feature extraction unit; for encoder layer Encoder2, the feature to be encoded should be the encoded feature output by encoder layer Encoder1. Other cases can be referred to [reference needed]. Figure 2 And this explanation is not listed here, so I will not give examples of each one.

[0114] Figure 2 In the example, Cat5~Cat7 are used as feature extraction and splicing devices, such as... Figure 2 In this code, the feature channel splicer Cat7 can concatenate the encoded features output from the encoder layer (Encoder3) and the decoded features output from the decoder layer (Decoder4). Cat7 outputs the concatenated encoded / decoded features, which are then used as the features to be decoded by the decoder layer (Decoder3). Other cases can be found in [reference needed]. Figure 2 And this is an explanation, which will not be listed here again.

[0115] In one embodiment of the present invention, the encoder layer includes a multi-scale wavelet decomposition module, an encoder splicer, and an encoder double convolution module, wherein, The multi-scale wavelet decomposition module is used to perform wavelet filtering decomposition on the features to be coded, and generate coded detail components and coded approximation components. The encoded detail component and the encoded approximation component are concatenated through a channel by an encoder splicer to generate encoded spliced ​​component features; The encoding double convolution module downsamples the features of the encoded concatenated components and outputs the encoded features.

[0116] Figure 2 The diagram illustrates an embodiment where the encoder layer employs the same structure. Within the Encoder1 layer, MD1 is a multi-scale wavelet decomposition module, Cat1 is an encoder concatenator, and DC1 is an encoder dual convolution module. The multi-scale wavelet decomposition module can achieve wavelet filtering decomposition through at least one one-dimensional convolutional layer. The kernel weights of this one-dimensional convolutional layer are set to predetermined wavelet filter coefficients to decompose the input spectral features. In practice, the kernel weights of the one-dimensional convolutional layer are configured using the wavelet filter coefficients stored in the pywt library of the Python library. The details for other encoder layers are as follows. Figure 2 And this explanation.

[0117] For each encoder layer, the features to be encoded are first decomposed by wavelet filtering using a multi-scale wavelet decomposition module, generating encoded detail components and encoded approximation components. The encoded approximation components reflect the slowly changing trend signal in the long wavelength range, usually corresponding to the background response of stable group composition features or dominant chemical structures, and have a high energy retention rate. The encoded detail components characterize the drastic absorption fluctuations at local wavelength points, which can reflect information of real microstructure features such as trace element perturbations, and may also reflect random noise introduced by changes in environment and sampling conditions.

[0118] The encoder splicer can perform channel splicing of the encoded detail component and the encoded approximation component to generate encoded spliced ​​component features. The encoded detail component and the encoded approximation component can then be passed down to provide parallel and complementary information to the feature extractor, allowing subsequent encoder layers to adaptively learn and weigh the importance of information at different scales according to task requirements.

[0119] A dual convolutional module can consist of two cascaded one-dimensional convolutional modules, batch normalization, and a ReLU activation function. Figure 8 The diagram illustrates one embodiment of a dual convolutional module. As shown, Conv1 and Conv2 are two cascaded one-dimensional convolutional modules. Batch normalization and ReLU activation functions are applied after both one-dimensional convolutional modules Conv1 and Conv2. Figure 8 In this context, BN stands for batch normalization, and ReLU stands for ReLU activation function. Figure 8 In this example, the kernel size of the one-dimensional convolution module is 3.

[0120] The following example, using the dual convolutional module employed in the channel expansion feature extraction unit, illustrates the working process of the dual convolutional module. Specifically, during operation, the one-dimensional convolutional module Conv1 first performs a sliding window operation on the training near-infrared spectral data / the near-infrared spectral data to be tested; the batch normalization layer standardizes the convolutional features, thereby accelerating model convergence and enhancing its stability; finally, the ReLU activation function enhances the nonlinear expressive power, enabling the network to learn and fit the complex nonlinear mapping relationship between spectral data and substance content. By repeating this process twice, the dual convolutional module can construct more expressive deep local features.

[0121] When the channel expansion feature extraction unit is working, the one-dimensional convolution module Conv1 first standardizes or compresses the information of any number of input channels (e.g., adjusts them to the baseline number of channels C). Then, the one-dimensional convolution module Conv2 receives this intermediate result and maps and transforms it to a target output channel with a larger number and higher dimension (e.g., r*C), thus completing a two-step strategic feature expansion and feature extraction.

[0122] The working method of the corresponding double convolution modules in the encoder layer and decoder layer can be referred to here, and will not be illustrated here one by one.

[0123] In one embodiment of the present invention, the decoder layer includes a multi-scale wavelet reconstruction module, a decoding linear adder, and a decoding double convolution module, wherein, The multi-scale wavelet reconstruction module is used to upsample the features to be decoded, and the low-frequency component and high-frequency component of the decoded component are generated after the upsampling operation. A decoding linear adder is used to linearly add the decoded low-frequency components and the decoded high-frequency components to form the characteristics of the decoded combined components; The decoding dual convolution module is used to perform feature fusion processing on the decoded and merged component features to generate decoded features.

[0124] Figure 2The diagram illustrates an embodiment where the decoder layer employs the same structure. Within the decoder layer Decoder1, MR1 is a multi-scale wavelet reconstruction module, AD1 is a decoding linear adder, and DC5 is a decoding dual convolution module. The multi-scale wavelet reconstruction module can achieve wavelet reconstruction through at least one one-dimensional transposed convolutional layer. The kernel weights of this one-dimensional transposed convolutional layer are set to predetermined, inverted wavelet reconstruction filter coefficients to upsample the input sub-components. This process can faithfully recover the signal's dimension and structure, laying a high-quality data foundation for subsequent fusion with shallow features. In specific implementations, the kernel weights of the one-dimensional transposed convolutional layer are configured using wavelet filter coefficients stored in the pywt library of the Python library. For other encoder layers, refer to [reference needed]. Figure 2 And this explanation.

[0125] It should be noted that the multi-scale wavelet reconstruction submodule is a key unit in the decoder layer that performs upsampling operations. It forms a mathematically precise inverse operation with the wavelet decomposition module in the encoder layer. As explained above, within the decoder network of this invention, different decoder layers can achieve skip connections to the corresponding encoder layers through the feature channel splicer, thereby enabling precise tracing from abstract concepts to specific details.

[0126] Skip connections serve as the information bridge between the encoder and decoder layers, and are the essence of the entire U-Net architecture. Their core function is to address the severe information loss problem caused by continuous downsampling in deep networks, particularly high-resolution local details. During encoding, as the network depth increases, the feature resolution continuously decreases, and the model learns increasingly abstract semantic information. However, high-resolution details such as the precise location, shape, and width of absorption peaks in the near-infrared spectrum are gradually blurred or lost during downsampling. Skip connections directly copy and transmit uncompressed features from specific layers in the encoding path to the corresponding layers in the decoding path. In each reconstruction step, the decoder network simultaneously receives two types of information: features from the previous decoder layer reconstructed by wavelet convolution, and features from the corresponding encoder layer transmitted via skip connections. The decoding dual-convolution module in the decoder layer then fuses these two concatenated pieces of information, enabling the feature extractor to utilize the fine features of the shallow layers to correct and optimize the abstract features of the deeper layers, ultimately achieving accurate localization and extraction of the target features.

[0127] As explained above, to fully explore the hierarchical features of the near-infrared spectrum of the sample under test, this invention proposes a feature extractor that integrates wavelet decomposition and reconstruction. In the encoding stage, wavelet decomposition is used to decompose the near-infrared spectral signal into coded approximate components and coded detail components, which are then used as input to the multi-channel information network for downsampling. This allows the quality parameter detection model to effectively focus on key response frequency bands related to the content of quality parameters. In the decoding stage, wavelet reconstruction accurately recovers signal details, more comprehensively exploring the hierarchical features of the spectrum, enhancing the feature capture capability of the quality parameter detection model, and thus improving the prediction accuracy of quality parameters within the sample under test and the robustness of model inference.

[0128] Figure 2 The figure illustrates one embodiment of a regression prediction head. As shown in the figure, the regression prediction head may include two sequentially connected linear regression prediction layers. Figure 2 In this diagram, HL1 and HL2 are two sequentially connected linear regression layers. Of course, the regression prediction head can take other forms, which will not be illustrated here.

[0129] To verify the necessity and effectiveness of incremental learning in this invention, the following verification is performed using sampling distance as the basis for distinguishing spectral distribution features. Specifically, the verification steps include: 1331 bauxite samples with a particle size of 0.15 mm were collected after drying, crushing, grinding, and sieving. The spectral data of the bauxite samples were acquired using a MicroNIR Pro handheld near-infrared spectrometer manufactured by VIAVI, with a wavelength range of 908-1676 nm, a resolution of 6.24 nm, and 125 wavelength sampling points. The sampling distances were 5 mm, 15 mm, and 25 mm. Figure 3 The image shows near-infrared spectra of the same sample at different sampling distances. Figure 4 Near-infrared spectra of different samples at different sampling distances are shown. Near-infrared spectroscopy is a typical one-dimensional data. The Fe2O3 content in bauxite samples was determined using X-ray fluorescence spectrometry based on the standard "Methods for Chemical Analysis of Bauxite" as a label.

[0130] In real-world industrial applications, the distribution of near-infrared spectral data can drift due to changes in raw material batches, environmental conditions, and even instrument status. To simulate this complex phenomenon in a controllable and reproducible manner, during validation, sampling distance was selected as the core physical variable to partition the dataset. Specifically, three non-repeating data subsets (D0, D1, D2) were created at different sampling distances, corresponding to 5mm, 15mm, and 25mm, respectively. Figure 3The diagram shows the probability density function of the iron oxide content in each subset, revealing the distribution drift problem in the task. Therefore, the setting of the data subsets here provides an ideal benchmark for evaluating the ability of the incremental learning algorithm to handle complex concept drift in real industrial scenarios. Specifically, each subset constitutes three main task stages, Task (1, 2, 3), and incremental learning is performed sequentially. The method and process of incremental learning can be referred to the above description and will not be repeated here.

[0131] After training for each stage of the task, this invention evaluates the performance of the parameter prediction reference model on the current task and all historical tasks, using RMSE, MSE, and R... 2 The results, used as evaluation metrics, are shown in Table 1. In Task 1, the parameter prediction reference model, trained and evaluated only on the data subset D0, performed excellently with R... 2 The RSI was 0.9707, validating the effectiveness of the basic model architecture in feature extraction and providing a benchmark for subsequent incremental learning. Moving to Task 2, incremental learning was performed using data subset D1, and RSI was tested on the test set of data subset D1. 2 The R value reached 0.9724, and was also higher on the test set of the backtest data subset D0. 2 With an R-value of 0.9294, it effectively mitigated catastrophic forgetting; in Task 3, the parameter prediction reference model achieved an R-value of 0.9632 on the test set of data subset D2. 2 This demonstrates that their learning ability did not significantly decline due to knowledge accumulation, and in Task 3, the parameter prediction reference model showed R on the joint test set (ALL). 2 The value reached 0.9429, indicating that the parameter prediction reference model has good distance robustness and knowledge preservation ability in incremental learning scenarios.

[0132] Figure 13 The scatter distribution of the model's predicted and true values ​​for ALL in the final stage is given, which can be used to intuitively evaluate the model's generalization performance on ALL. Among them, the samples with prediction errors of |yx|≤1.5 account for 93% of all data, indicating that the parameter prediction reference model can make accurate predictions on most samples.

[0133] Table 1. Results of this example method on the multi-stage task test set.

[0134] Figure 14 and Figure 15 The figure shows a comparison of the test set results of the model stage at the final stage of different methods. As can be seen from the figure, the effectiveness and necessity of the method (Ours) for constructing a quality parameter detection model using incremental learning in this invention are demonstrated. Figure 14The middle section shows the results of the final stage model on the test set for each stage task. Figure 15 The difference between each model's result and the Oracle model is shown. The Oracle model is specifically the model trained using only a subset of training data, as described above, which is the parameter prediction baseline model obtained through basic training. First, on the ALL test set, which measures the model's generalization ability, the quality parameter prediction model constructed in this invention significantly outperforms benchmark strategies such as traditional fine-tuning and incremental learning with an R² of 0.9429, demonstrating the superiority of the overall framework. Furthermore, the construction method of this invention not only achieves the highest prediction accuracy among all comparison models, but its performance is also closest to the ideal Oracle model (trained and tested separately for each task), thus proving that the quality parameter detection model constructed in this invention possesses excellent stability and strong generalization ability.

Claims

1. A method for constructing a quality parameter detection model based on incremental learning, characterized in that, The method for constructing the quality parameter detection model includes: A basic parameter prediction model is constructed for predicting quality parameters of the substance to be tested, and a training data set corresponding to the substance to be tested is provided, wherein... The training data group includes at least two training datasets, and the spectral distribution characteristics of all training datasets are at least incompletely matched. Each training dataset includes several training samples related to the substance to be tested. For any training sample, the training sample includes training near-infrared spectral data and training parameter labels characterizing the quality parameter state of the substance under test, wherein the training parameter labels are positively correlated with the training near-infrared spectral data; The basic parameter prediction model is trained using the training dataset within the training data set, driven by incremental learning. This trained model then generates a quality parameter detection model. When performing incremental learning-driven model training, the following are included: Configure one training dataset as the base set, and use the remaining training datasets as incremental sets. The basic model for parameter prediction is trained using the base set. Then, incremental learning is performed once using the incremental set. After all incremental learning is performed, a quality parameter detection model is generated.

2. The method for constructing a quality parameter detection model based on incremental learning as described in claim 1, characterized in that: During basic training, the parameter prediction basic model and the generative adversarial network basic model are trained using the basic set, respectively, so as to generate the parameter prediction basic model and the generative adversarial network basic model after training. When performing the first incremental learning, the parameter prediction base model, the generative adversarial base network model, and the basic label information of the base training set are configured as the parameter prediction reference model, the generative adversarial reference network model, and the reference label information, respectively. The basic label information is generated based on the base set. After performing incremental learning once using an incremental set, a parameter prediction reference model, an adversarial reference network model, and reference label information are generated for the next incremental learning. After performing the last incremental learning, a quality parameter detection model is generated based on the parameter prediction reference model after the incremental learning.

3. The method for constructing a quality parameter detection model based on incremental learning as described in claim 2, characterized in that: When performing incremental learning, the following are included: Determine the incremental set to be used for this incremental learning, and obtain the latest generated parameter prediction reference model, generative adversarial reference network model, and reference label information before executing this incremental learning; When generating the parameter prediction reference model for the next incremental learning based on the determined incremental set and the obtained parameter prediction reference model, the following steps are included: The parameter prediction reference model is incrementally trained using the determined increment set, wherein the incremental training includes several rounds of incremental iteration. In each round of incremental iteration, the parameter prediction reference model is trained using the incremental set to generate a parameter prediction intermediate model after training. The regression prediction head in the parameter prediction intermediate model is then jointly optimized to form the parameter prediction intermediate model for the next round of incremental iteration after joint optimization. Once the incremental iteration reaches the specified number of rounds, the intermediate parameter prediction model that has reached the specified number of rounds will be configured as the parameter prediction reference model for the next incremental learning. When generating the generative adversarial reference network (GARNR) model required for the next incremental learning, based on the determined incremental set and the GARNR model, the following steps are included: The acquired generative adversarial reference network model is trained using the incremental set, so as to generate the generative adversarial reference network model required for the next incremental learning. When generating reference label information, the training parameter labels of all training samples in the determined incremental set are obtained, and the label mean and label variance corresponding to all training parameter labels are calculated, so as to form the reference label information required for the next incremental learning based on the calculated label mean and label variance.

4. The method for constructing a quality parameter detection model based on incremental learning as described in claim 3, characterized in that: When jointly optimizing the regression prediction heads within the intermediate model for parameter prediction, the following is included: Generate a feature-label set required for joint optimization of the regression prediction head, wherein the feature-label set includes several pseudo-feature-pseudo-label pairs and several real feature-real-label pairs; When generating true feature-true label pairs, the following steps are included: The feature extractor within the intermediate model of parameter prediction is used to extract features from a training near-infrared spectral data to obtain the true features; Based on the training near-infrared spectral data corresponding to the real features, the training parameter labels that correspond to the near-infrared spectral data are determined, and the determined training parameter labels are configured as real labels. The training near-infrared spectral data belongs to a training sample in the incremental set used to perform this incremental learning. When generating pseudo-feature-pseudo-label pairs, the following is included: Based on the reference label information obtained from this incremental learning process, pseudo-labels are generated through sampling. The pseudo-labels are loaded into the Generative Adversarial Reference Network (GARN) model, which then generates pseudo-features corresponding to the pseudo-labels. The regression prediction head is optimized and updated using the feature-label set to generate an optimized regression prediction head.

5. The method for constructing a quality parameter detection model based on incremental learning as described in claim 3, characterized in that: When training the obtained parameter prediction reference model using the determined increment set, the incremental training loss function used is: in, For incremental training loss function, For the average absolute error loss, For distillation loss, For offset penalty loss, Mean absolute error loss The weight, Distillation loss weights, Offset penalty loss The weight, The training near-infrared spectral data for a training sample within the increment set. To utilize the parameters obtained from incremental iteration during this incremental learning process to predict the spectral features extracted by the intermediate model, To utilize the latest parameters generated before this incremental learning to predict the spectral features extracted from the reference model, Spectral characteristics With spectral characteristics The square of the L2 norm of the difference, Spectral characteristics With spectral characteristics The expected value of the square of the L2 norm of the difference; This is the first intermediate model for parameter prediction during incremental iteration in this incremental learning process. Each model parameter This is the first parameter prediction reference model generated before this incremental learning. Each model parameter For the first Model parameters The estimated value of the Fisher information matrix, This is a hyperparameter.

6. The method for constructing a quality parameter detection model based on incremental learning according to any one of claims 1 to 5, characterized in that, When generating training near-infrared spectral data for each training sample, the following steps are included: Near-infrared spectral source data corresponding to each training near-infrared spectral data is acquired, and the near-infrared spectral source data is preprocessed to generate training near-infrared spectral data after preprocessing. The near-infrared spectral source data is generated through near-infrared spectral acquisition; Preprocessing of near-infrared spectral source data includes standard normal transformation and / or Huber-based adaptive asymmetric least squares baseline correction. When the preprocessing of near-infrared spectral source data includes both standard normal transformation and Huber-based adaptive asymmetric least squares baseline correction, then the near-infrared spectral source data is subjected to standard normal transformation and Huber-based adaptive asymmetric least squares baseline correction sequentially. When performing Huber-regularized adaptive asymmetric least squares baseline correction, we have: in, For the final baseline estimation data, To train near-infrared spectral data, the total number of wavenumber points corresponding to near-infrared light spectral source data, The residual at the nth wavenumber point, This represents the weight corresponding to the nth wavenumber point. For the smoothing constraint term corresponding to the near-infrared light spectral source data, The absorbance at the nth wavenumber point within the near-infrared spectral correction base data. The absorbance at the nth wavenumber point within the initial baseline estimation data generated based on near-infrared spectral correction baseline data. For absorbance The second-order difference was performed; When the preprocessing performed only includes Huber regularization-based adaptive asymmetric least squares baseline correction, the near-infrared spectral correction basis data is the near-infrared spectral source data. When the preprocessing includes standard normal transformation and Huber regularization-based adaptive asymmetric least squares baseline correction, the near-infrared spectral source data is generated by performing standard normal transformation on the near-infrared spectral source data. Polynomial fitting is performed on the final baseline estimate data to generate the final baseline fit data after polynomial fitting; The near-infrared spectral correction baseline data is subtracted from the final baseline fitting data to form a data difference, and the data difference is configured as training near-infrared spectral data.

7. The method for constructing a quality parameter detection model based on incremental learning as described in claim 6, characterized in that, When generating initial baseline estimation data based on near-infrared spectral correction baseline data, the following are included: The near-infrared spectral correction baseline data were sequentially processed by median filtering and polynomial fitting to generate initial baseline estimation data after polynomial fitting. For weights and smoothing constraint terms Then we have: in, b The median of all residuals. or , This represents the standard deviation of all residuals in the current near-infrared spectral correction baseline data. , All parameters are over-parameterized.

8. The method for constructing a quality parameter detection model based on incremental learning according to any one of claims 3 to 5, characterized in that, The feature extractor within the intermediate model for parameter prediction includes a channel augmentation feature extraction unit and a main feature extraction unit connected in sequence. The main feature extraction unit adopts a U-Net network architecture. The feature extraction main unit includes an encoder network and a decoder network adapted and connected to the encoder network. The encoder network includes several encoder layers connected in sequence, and the channel expansion feature extraction unit is connected to the encoder layer located at the opening of the U-Net network within the encoder network; The decoder network includes several decoder layers connected in sequence. The number of decoder layers in the decoder network is consistent with the number of encoder layers in the encoder network. The decoder layer at the opening of the U-Net network is adapted to connect with the regression prediction head. For any encoder layer, the encoder layer receives the feature to be encoded and first performs wavelet filtering decomposition on the feature to be encoded to generate coded detail components and coded approximate components after wavelet filtering decomposition. Then, the coded detail components and coded approximate components are concatenated by channels and downsampled after channel concatenation to generate the encoded feature. For any decoder layer, the decoder layer receives the feature to be decoded, and first performs upsampling and merging operations on the feature to be decoded, and then performs feature fusion processing after the merging operation to generate the decoded feature after feature fusion processing; Within the main feature extraction unit, the bottom-most decoder layer directly loads the encoded features into the corresponding decoder layer. In addition to the decoder layer located at the bottom layer, the input of the other decoder layers is connected to the output of a feature extraction splicer, which receives the decoded features output from the previous decoder layer and the encoded features output from the encoder layer that is in the same layer as the current decoder layer.

9. The method for constructing a quality parameter detection model based on incremental learning as described in claim 8, characterized in that: The encoder layer includes a multi-scale wavelet decomposition module, an encoder splicer, and an encoder double convolution module, wherein... The multi-scale wavelet decomposition module is used to perform wavelet filtering decomposition on the features to be coded, and generate coded detail components and coded approximation components. The encoded detail component and the encoded approximation component are concatenated through a channel by an encoder splicer to generate encoded spliced ​​component features; The encoding double convolution module downsamples the features of the encoded concatenated components and outputs the encoded features.

10. The method for constructing a quality parameter detection model based on incremental learning as described in claim 8, characterized in that: The decoder layer includes a multi-scale wavelet reconstruction module, a decoding linear adder, and a decoding double convolution module, wherein... The multi-scale wavelet reconstruction module is used to upsample the features to be decoded, and the low-frequency component and high-frequency component of the decoded component are generated after the upsampling operation. A decoding linear adder is used to linearly add the decoded low-frequency components and the decoded high-frequency components to form the characteristics of the decoded combined components; The decoding dual convolution module is used to perform feature fusion processing on the decoded and merged component features to generate decoded features.

Citation Information

Patent Citations

  • Method for determining contents of nine substances in high-sulfur bauxite by X-ray fluorescence spectrometry

    CN112179930A

  • Spectral adaptive incremental learning modeling method and system

    CN119026057A

  • Multi-quality parameter cooperative detection method and system based on near infrared spectrum

    CN120142228A

  • Method and apparatus for class incremental learning

    US20230145919A1