Multi-scale anti-decoupling near infrared spectrum model transfer method and system

By employing a multi-scale adversarial decoupling near-infrared spectroscopy model transfer method, and utilizing a dual-branch CNN and an improved CycleGAN to generate target instrument-specific noise data, the cross-instrument model applicability problem was solved. This method achieves high robustness and model transfer under low sample size, thereby improving the model's adaptability and accuracy.

CN120995133APending Publication Date: 2025-11-21CHINESE ACAD OF AGRI MECHANIZATION SCI GRP CO LTD

Patent Information

Application Number
CN202511092288.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, near-infrared spectroscopy models cannot be applied across instruments, resulting in high costs of repeated modeling and difficulties in cross-category promotion. Traditional methods cannot handle nonlinear instrument responses and high-frequency noise coupling problems, and the initial sample size of the target instrument is insufficient to support the training of deep learning models.

Method used

A multi-scale adversarial decoupling near-infrared spectroscopy model transfer method is adopted. The instrument-related and irrelevant features are separated by a dual-branch CNN, and the target instrument-specific noise data is generated by an improved CycleGAN. The model is trained and fine-tuned by a dynamic weighted sampling strategy to achieve cross-instrument model transfer.

Benefits of technology

It achieves highly robust model transfer with only 10% target small sample size, solves the problems of nonlinear instrument response and high-frequency noise coupling, and improves the model's cross-instrument generalization ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995133A_ABST
    Figure CN120995133A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale adversarial decoupling near infrared spectrum model transfer method. The method comprises the following steps: acquiring near infrared spectrum data sets of a source domain and a target domain, executing a spatio-temporal data distribution strategy in time and space dimensions, carrying out proportion division on the spectrum data sets, and carrying out near infrared spectrum model pre-training; respectively extracting instrument irrelevant features and instrument relevant features through a neural network, performing adversarial training to minimize domain difference loss, and outputting a domain invariant subset; inputting the related characteristics of the instrument, the near infrared spectrum data set of the target domain and the domain invariant subset into a generator to generate virtual spectrum data; and adopting a dynamic weighted sampling strategy, adaptively adjusting the sampling probability of the virtual spectrum data and the real spectrum data, performing model training fine tuning, deploying the fine-tuned model to a target instrument, and realizing cross-instrument transfer of the near infrared spectrum model. According to the method, the cross-instrument generalization ability of the model is greatly improved, and the robustness of the spectrum model in the space-time dimension is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of instrument calibration and chemometrics model transfer technology, and particularly relates to a near-infrared spectrum model transfer method and system. BACKGROUND

[0002] In recent years, the near-infrared spectrum technology has been widely applied in the detection of material characteristics in the fields of pharmaceuticals, agricultural products, and chemical industry due to its characteristics of rapidness, non-destructiveness, and environmental protection. However, due to the spectral nonlinear distortion caused by the light source attenuation of different devices, the sensitivity difference of detectors, and the distribution difference of spectral characteristics between different instruments, it is necessary to repeatedly model for each instrument, which increases the cost, and makes the current still stay in the architecture of "one variety, one set of model, and single machine operation". The model is fixed and cannot be updated, which cannot adapt to the variety replacement or demand change, resulting in high cost of repeated modeling, difficult promotion across categories, and inability to meet the multi-scene and multi-target detection requirements. If the model of one instrument is to be applied to another device, model transfer must be performed to ensure the accuracy and consistency of the prediction results of the models of multiple instruments.

[0003] The best method for near-infrared model transfer is the method of spectral correction. However, the traditional linear correction methods (such as direct correction method (PDS) and direct standardization-partial least squares (DS-PLS)) cannot handle the problems of nonlinear instrument response and high-frequency noise coupling; the initial sample size of the target instrument is insufficient to support the training of deep learning models; and the virtual spectra generated by the traditional data enhancement method and the virtual spectra synthesized by the conventional generative adversarial network (GAN) lack instrument-specific noise characteristics, resulting in overfitting of the model during fine-tuning and a negative transfer risk of >20%.

[0004] Therefore, in view of the problems existing in the prior art, a new multi-scale anti-decoupling and virtual enhancement near-infrared spectrum model transfer method architecture is needed. How to separate the instrument-related features and irrelevant features through a double-branch CNN, improve the CycleGAN to synthesize target instrument-specific noise data, and realize high-robustness model transfer with only 10% of the target small sample has become a key difficulty in current research. SUMMARY

[0005] In order to solve the problems of nonlinear instrument response and high-frequency noise coupling and insufficient data of the target instrument, improve the model cross-instrument generalization ability, and enhance the robustness of the model in the time and space dimensions, the present application provides a multi-scale anti-decoupling near-infrared spectrum model transfer method and system.

[0006] In a first aspect, the embodiments of the present application provide a multi-scale anti-decoupling near-infrared spectrum model transfer method, which comprises the following steps:

[0007] The model pre-training step: obtaining the near-infrared spectrum data set of the source domain and the target domain from the instrument, performing a time-space data allocation strategy, proportionally dividing the spectrum data set in the time and space dimensions, and pre-training the near-infrared spectrum model based on the divided spectrum data set;

[0008] The multi-scale adversarial feature decoupling network construction step: based on the near-infrared spectrum data set of the source domain, instrument-independent features and instrument-related features are extracted through a neural network, adversarial training is performed to minimize the domain difference loss, and a domain-invariant subset is output;

[0009] The spectrum virtual enhancement step: inputting the instrument-related features, the near-infrared spectrum data set of the target domain and the domain-invariant subset into a generator to generate virtual spectrum data;

[0010] The dynamic weighted sampling fine-tuning step: adopting a dynamic weighted sampling strategy to adaptively adjust the sampling probability of the virtual spectrum data and the real spectrum data, performing model training fine-tuning, deploying the fine-tuned model to the target instrument, and realizing the cross-instrument transfer of the near-infrared spectrum model.

[0011] In the embodiment of the present application, the above-mentioned model pre-training step further comprises:

[0012] The near-infrared spectrum data set of the source domain is divided in the time dimension according to the data acquisition sequence, and is clustered in the space dimension by a dimension reduction algorithm, and after uniform distribution of samples, it is divided into a training set, a validation set and a test set in turn according to a self-defined proportion, and model pre-training is performed.

[0013] In the embodiment of the present application, the above-mentioned multi-scale adversarial feature decoupling network construction step further comprises:

[0014] The instrument-independent feature extraction step: inputting the near-infrared spectrum data set of the source domain into the first branch of the double-branch neural network to extract instrument-independent features; the first branch adopts a spectral-spatial domain decomposition residual block;

[0015] The related feature extraction step: inputting the near-infrared spectrum data set of the source domain into the second branch of the double-branch neural network to extract instrument-related features; the second branch adopts a noise perception residual module;

[0016] The feature adversarial training step: inputting the extracted instrument-independent features and instrument-related features into a gradient inversion layer, and performing adversarial training through the domain discriminator and the feature extractor in the gradient inversion layer to minimize the domain difference loss.

[0017] In the embodiment of the present application, the above-mentioned instrument-independent feature extraction step further comprises:

[0018] The first branch carries out low-pass filtering on near-infrared spectrum data to obtain a high-dimensional signal, a projection matrix is obtained by performing principal component analysis on the spectrum data, a preset number of principal components in sequence are extracted, and the filtered high-dimensional signal is converted into spectrum data of a preset number of dimensions.

[0019] By applying a gated attention mechanism to the spectrum, wavelength points without material characteristic information in the converted spectrum data are masked, an effective wavelength segment area is automatically identified, global pooling is performed, and the results after processing the spectrum dimension and the spatial dimension are connected by jumping to obtain an instrument-independent feature that retains original spectrum information.

[0020] In the embodiment of the present application, the above-mentioned related feature extraction step further comprises:

[0021] The second branch carries out second-order differentiation on the near-infrared spectrum data to extract instrument-related noise characteristics caused by instrument aging;

[0022] The noise characteristics after second-order differentiation are reduced in dimension by residual convolution to a low-dimensional noise matrix containing high-frequency noise information;

[0023] The near-infrared spectrum data is encoded by an instrument fingerprint encoding function to extract inherent characteristics of the instrument;

[0024] The instrument noise characteristics and inherent characteristics are fused by feature splicing to realize extraction of instrument-related characteristics.

[0025] In the embodiment of the present application, the above-mentioned feature adversarial training step further comprises:

[0026] The feature extractor learns the input instrument-related characteristics, corrects its own parameters, generates noise information and integrates it into the instrument-independent characteristics;

[0027] The domain discriminator learns the input instrument-independent characteristics, corrects its own parameters to improve the discrimination ability of the source domain and the target domain spectrum;

[0028] The feature extractor and the domain discriminator continuously learn through continuous confrontation, the feature extractor continuously maximizes the adversarial training loss function, and the domain discriminator continuously minimizes the adversarial training loss function, when the dynamic balance of the maximized adversarial training loss function and the minimized adversarial training loss function is reached, the output domain-invariant subset is obtained.

[0029] In the embodiment of the present application, the above-mentioned spectrum virtual enhancement step further comprises:

[0030] The instrument-related characteristics, the target domain unlabeled spectrum data set and the domain-invariant subset are input into the generator, dimension splicing is performed, and the instrument-related characteristics are expanded;

[0031] The characteristics of the wavelength dimension are reserved, the generator output value after splicing is encoded and decoded, the noise characteristics in the instrument-related characteristics are injected into the target domain unlabeled spectrum data set and the domain-invariant subset, and the virtual spectrum is output.

[0032] In the embodiment of the application, the dynamic weighted sampling fine-tuning step further comprises:

[0033] The sampling weights of the full connection layer are fine-tuned by using the dynamic weighted sampling strategy, and the formula of the weight is:

[0034] The sampling weight = alpha(t) * virtual data sampling probability + (1-alpha(t)) * real data sampling probability, wherein alpha(t) is a virtual data dynamic weight coefficient, t is the training round number, and the value of alpha(t) is determined by the dynamic weighted sampling strategy.

[0035] In the embodiment of the application, the dynamic weighted sampling strategy is:

[0036] When t is in the first value interval, alpha(t) = 0.9-0.02t, more than or equal to 90% of the samples in the virtual spectrum data are collected in each round, and less than or equal to 10% of the real spectrum data are collected;

[0037] When t is in the second value interval, alpha(t) = 0.7-0.02(t-10), the sampling proportion of the virtual spectrum gradually decreases from 70% to 30% in each round, and the sampling proportion of the real fine-tuning set gradually increases from 30% to 70%.

[0038] When t is in the third value interval, alpha(t) = 0.3, the sampling proportion of the virtual spectrum and the real fine-tuning spectrum is determined according to the preset threshold.

[0039] In the second aspect, the embodiment of the application provides a multi-scale anti-decoupling near-infrared spectrum model transfer system, which adopts the multi-scale anti-decoupling near-infrared spectrum model transfer method, and the system comprises:

[0040] The model pre-training module is used for acquiring the near-infrared spectrum data set of the source domain and the target domain from the main instrument and the target instrument, performing the time-space data allocation strategy, performing the proportion division in the time and space dimensions on the spectrum data set, and pre-training the near-infrared spectrum model based on the divided spectrum data set.

[0041] The multi-scale anti-decoupling feature network module is used for extracting the instrument-independent features and the instrument-related features by using the neural network based on the near-infrared spectrum data set of the source domain, performing the anti-training, minimizing the domain difference loss, and outputting the domain-invariant subset.

[0042] The virtual enhancement controller module is used for inputting the instrument-related features, the near-infrared spectrum data set of the target domain and the domain-invariant subset into the generator to generate the virtual spectrum data.

[0043] Dynamic partitioner module: for adopting dynamic weighted sampling strategy, adaptively adjusting the sampling probability of virtual spectral data and real spectral data, fine-tuning the model training, deploying the fine-tuned model to the target instrument, realizing the cross-instrument transfer of near-infrared spectral model.

[0044] In a third aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the steps of the multi-scale anti-decoupling near-infrared spectral model transfer method.

[0045] In a fourth aspect, the embodiments of the present application provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the multi-scale anti-decoupling near-infrared spectral model transfer method when executing the program.

[0046] Compared with the related prior art, the present application has the following outstanding beneficial effects:

[0047] (1) The spatio-temporal data allocation strategy proposed in the method of the present application divides the spectral data set, ensures that the training set contains full concentration and matrix, and avoids the generation of distribution drift;

[0048] (2) The method of the present application proposes to extract instrument-independent features and instrument-dependent features through a double-branch convolutional neural network (CNN), and then perform anti-training through the domain discriminator and feature extractor in the gradient inversion layer, to realize the domain invariance of the features. The problems of nonlinear instrument response and high-frequency noise coupling are solved;

[0049] (3) The method of the present application proposes to adaptively adjust the sampling probability of virtual data and real data through a dynamic weighted sampling strategy, effectively reducing the risk of insufficient initial sample size of the target instrument, which is difficult to support the training of the deep learning model;

[0050] (4) The method of the present application proposes a new architecture of multi-scale anti-decoupling and virtual enhancement near-infrared spectral model transfer method. The instrument-dependent features and instrument-independent features are separated through a double-branch CNN, the CycleGAN is improved to synthesize target instrument-specific noise data, and high-robustness model transfer is realized with only 10% of the target small sample. BRIEF DESCRIPTION OF DRAWINGS

[0051] The accompanying drawings used to provide further understanding of the present application and form a part of the present application, and the illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0052] Figure 1 It is a schematic diagram of the multi-scale anti-decoupling near-infrared spectral model transfer method of the present application.

[0053] Figure 2 A schematic diagram of a multi-scale adversarial decoupling near-infrared spectroscopy model transfer method according to an embodiment of the present application;

[0054] Figure 3 A schematic diagram of a multi-scale adversarial decoupling near-infrared spectroscopy model transfer system according to an embodiment of the present application;

[0055] Figure 4 A schematic diagram of a computer hardware according to an embodiment of the present application. DETAILED DESCRIPTION

[0056] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0057] It should also be understood that the term "and / or" herein merely describes an associated relationship between associated objects, and means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but can also represent an "and / or" relationship. The specific meaning can be understood according to the context before and after it.

[0058] It should also be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0059] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other manners. For example, the described device embodiments are merely schematic. Taking the division of the units as an example, the division can be changed in actual implementation, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0060] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment of the present application according to actual needs.

[0061] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0062] If the functions are realized in the form of software functional units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various program code storage media.

[0063] In order to make the above features and effects of the present application more clear and easy to understand, the following embodiments are described in detail below, and the accompanying drawings are described as follows. The present application discloses one or more embodiments containing the features of the present application. The disclosed embodiments are only used for illustration. The protection scope of the present application is not limited to the disclosed embodiments, and the present application is defined by the appended claims.

[0064] The following is a system embodiment corresponding to the above method embodiment, and the present embodiment can be implemented in cooperation with the above embodiments. The related technical details mentioned in the above embodiments are still valid in the present embodiment. In order to reduce repetition, they will not be described here. Correspondingly, the related technical details mentioned in the present embodiment can also be applied in the above embodiments.

[0065] The method of the present application aims to propose a new architecture of multi-scale adversarial decoupling and virtual enhancement near-infrared spectral model transfer method. By separating the instrument related features and irrelevant features through a double branch CNN, the CycleGAN is improved to synthesize target instrument specific noise data, and a high robustness model transfer method is realized only with 10% target small sample.

[0066] The application is suitable for cross-instrument adaptive migration and inter-instrument consistency calibration technology of near-infrared spectrum models in the fields of pharmaceuticals, agricultural product detection and the like.

[0067] The method of the embodiment of the application will be described in detail below in combination with specific embodiments:

[0068] Embodiment one

[0069] As shown in Figure 1 and Figure 2 , the embodiment of the application provides a multi-scale adversarial decoupling near-infrared spectrum model transfer method, which comprises the following steps:

[0070] A model pre-training step 101: acquiring near-infrared spectrum data sets of source and target domains from instruments, performing a time-space data allocation strategy in time and space dimensions, proportionally dividing the spectrum data sets, and pre-training near-infrared spectrum models based on the divided spectrum data sets;

[0071] Specifically, in the embodiment of the application, the time-space data allocation strategy is performed, the data is strictly divided in the time dimension according to the data acquisition sequence, and the samples are uniformly allocated after t-SNE clustering in the space dimension. In the dynamic data division strategy, the time sequence segmentation simulates the influence of instrument aging, and the feature space clustering ensures the consistency of the distribution of chemical component concentrations in the training set, the validation set and the test set.

[0072] A multi-scale adversarial feature decoupling network construction step 102: based on the near-infrared spectrum data set of the source domain, instrument-independent features and instrument-dependent features are respectively extracted through a neural network, adversarial training is performed, domain difference loss is minimized, and a domain-invariant subset is output;

[0073] Specifically, in the embodiment of the application, a multi-scale adversarial feature decoupling network is constructed. First, a double-branch convolutional neural network (CNN) is used to extract instrument-independent features and instrument-dependent features, respectively. The first branch uses a spectral-spatial domain decomposition residual block (SSD-ResBlock) to process input spectrum data, outputs instrument-independent features, and the second branch extracts instrument-dependent features by a noise perception residual module (NA-ResBlock) to process original spectrum, and outputs instrument-dependent features. Then, the output features of the double branches are fused, connected to a domain discriminator through a gradient reversal layer (GRL) for adversarial training, and the domain difference loss is minimized.

[0074] A spectrum virtual enhancement step 103: inputting the instrument-dependent features, the near-infrared spectrum data set of the target domain and the domain-invariant subset into a generator to generate virtual spectrum data;

[0075] Specifically, in the embodiment of the present application, spectral virtual enhancement migration is performed. Improved cycle generative adversarial network (CycleGAN) is used to simulate the target instrument noise distribution to generate virtual spectral data based on the pre-trained master instrument model.

[0076] Dynamic weighted sampling fine-tuning step 104: using a dynamic weighted sampling strategy, the sampling probability of virtual spectral data and real spectral data is adaptively adjusted, the model is fine-tuned, the fine-tuned model is deployed to the target instrument, and the cross-instrument transfer of the near-infrared spectral model is realized.

[0077] Specifically, in the embodiment of the present application, the virtual spectral data and 10% target instrument real data are mixed, and a dynamic weighted sampling strategy is used to fine-tune the full connection layer. The fine-tuned model is deployed to the target instrument to verify the migration effect, and the "one model, multiple instrument reuse" is realized.

[0078] In the embodiment of the present application, the above model pre-training step 101 further comprises:

[0079] The near-infrared spectral data set of the source domain is divided in the time dimension according to the data acquisition order, and in the spatial dimension by the dimension reduction algorithm clustering, and after uniformly distributing the samples, it is divided into training set, verification set and test set according to the self-defined proportion in turn, and the model is pre-trained.

[0080] More specifically, in the embodiment of the present application, the source domain spectral data set is divided by executing a time-space data allocation strategy. In the time dimension, the source domain spectral data set is strictly arranged in ascending order of time stamp, and is divided into a training set (60% of the first period data), a verification set (20% of the middle period data) and a test set (20% of the last period data) in turn, so as to ensure that the test set covers the future state, avoids the risk of time leakage, and makes the model maintain relative stability in a long-time continuous dynamic prediction environment. In the spatial dimension, t-SNE clustering uniform sampling is adopted: the full-waveband high-dimensional spectral data is arranged into a high-dimensional vector matrix (for N spectra of X wavelength points, the high-dimensional vector matrix is N X X), the similarity between samples is calculated to form a distance matrix (N X N), t-SNE nonlinear dimension reduction is performed on the full-waveband spectral data, density clustering is performed, chemical component similarity sub-clusters are formed, samples in each cluster are randomly extracted as a training set, a verification set and a test set according to a ratio of 6:2:2, and full concentration and matrix are ensured to be contained in the training set to avoid the generation of distribution drift. Then, the data samples in the time dimension and the spatial dimension are finally mixed and input into the model for model pre-training.

[0081] In the embodiment of the present application, the above step of constructing a multi-scale adversarial feature decoupling network 102 further comprises:

[0082] The instrument-independent feature extraction step includes the following steps: inputting a near-infrared spectrum data set of a source domain into a first branch of a double-branch neural network to extract instrument-independent features; and using a spectral-spatial domain decomposition residual block in the first branch.

[0083] The instrument-dependent feature extraction step includes the following steps: inputting the near-infrared spectrum data set of the source domain into a second branch of the double-branch neural network to extract instrument-dependent features; and using a noise perception residual module in the second branch.

[0084] The feature adversarial training step includes the following steps: inputting the extracted instrument-independent features and instrument-dependent features into a gradient inversion layer; and performing adversarial training through a domain discriminator and a feature extractor in the gradient inversion layer to minimize domain difference loss.

[0085] In the embodiment of the present application, the instrument-independent feature extraction step further includes the following operations of the PCA-constrained residual block: performing principal component dimension reduction on the input spectrum slice; learning non-linear mapping of the reduced features through a convolution layer; and adding a skip connection to retain original spectrum information.

[0086] In the first branch, the near-infrared spectrum data is low-pass filtered into a high-dimensional signal, a projection matrix is obtained by performing principal component analysis on the spectrum data, a preset number of principal components are extracted in order, and the filtered high-dimensional signal is converted into spectrum data of a preset number of dimensions.

[0087] The wavelength points without material feature information in the converted spectrum data are masked by applying a gated attention mechanism to the spectrum, the effective wavelength segment region is automatically identified, global pooling is performed, the results after processing the spectrum dimension and the spatial dimension are connected by a skip connection, and the instrument-independent features retained in the original spectrum information are obtained.

[0088] More specifically, in the embodiment of the present application, the spectral-spatial domain decomposition residual block (SSD-ResBlock) is used to extract instrument-independent features F inv , and the formula is as follows:

[0089]

[0090] Spectral-spatial coupling convolution layer:

[0091] Principal component attention gate: M spec = sigmoid(W T att U k );

[0092] wherein X represents an input spectrum, represents spectrum dimension convolution, represents spatial dimension convolution, K low represents a low-pass filter, W pca is a projection matrix, and Mspec P is a spectral attention mask, P is a pooling operation, U k is the first k PCA eigenvector. W T pca : the transpose of the projection matrix, the high-dimensional feature after spectral convolution is reduced to low-dimensional, and the main spectral information is retained.

[0093] Y: the fusion result of spectral features and spatial features, used for subsequent feature extraction. W T att : the transpose of the attention projection matrix, the principal component U k mapped to the attention space generates a weight vector matching the number of bands.

[0094] The following will be explained in detail in combination with the formula:

[0095] First, the near-infrared spectrum X is low-pass filtered, the intrinsic signal in the near-infrared spectrum is a low-frequency signal, and the instrument noise is a high-frequency signal. The low-frequency intrinsic signal is retained by filtering the high-frequency signal through a low-pass filter, and the formula is The filtered low-frequency signal is still a high-dimensional signal, and the projection matrix is obtained by performing PCA principal component analysis on the spectral data of the instrument, and the first d principal components are extracted, and the high-dimensional signal is converted into d-dimensional signal, and the formula is In the full-band spectral data, not all wavelength points contain material characteristic information. By applying a gated attention mechanism to the spectrum, the wavelength points in the spectrum that do not contain material characteristic information are masked, the effective wavelength segment area is automatically identified, and global pooling is performed, and the formula is P(X⊙M spec ); The results of spectral dimension and spatial dimension processing are connected by jumping to obtain the instrument-independent features that retain the original spectral information.

[0096] Wherein, the value range of d (the number of the first d principal components) is usually 10-100, and the optimal value needs to be determined according to the "cumulative variance contribution rate", that is, the number of the least principal components with cumulative variance contribution rate≥95% (or 98%, 99%). The selection of the number of PCA principal components takes the "cumulative variance contribution rate" as the core index, which means that the first d principal components explain the proportion of the variance of the original data. In near-infrared spectral analysis, the cumulative variance contribution rate≥95% is the conventional threshold (most essential information can be retained), and≥98% is more stringent (suitable for scenes with high precision requirements).

[0097] If the cumulative variance contribution rate requirement is 95%, d is usually 20-50; if the requirement is 98%, d is usually 30-70; if the requirement is 99%, d is usually 50-100.

[0098] In the embodiment of the application, the above-mentioned related feature extraction steps further comprise:

[0099] In the second branch, the near-infrared spectrum data is subjected to second-order differentiation to extract instrument-related noise characteristics caused by instrument aging;

[0100] The noise characteristics after second-order differentiation are subjected to residual convolution to reduce high-dimensional noise characteristics to a low-dimensional noise matrix containing high-frequency noise information;

[0101] The near-infrared spectrum data is encoded by an instrument fingerprint encoding function to extract inherent characteristics of the instrument;

[0102] The instrument noise characteristics and inherent characteristics are fused by feature splicing to realize extraction of instrument-related characteristics.

[0103] More specifically, in the embodiment of the present application, the instrument-related characteristics F dep are extracted by the noise-aware residual module (NA-ResBlock) as follows:

[0104]

[0105] wherein, is a second-order differentiation operator, H res is residual convolution, is feature splicing, and δ is an instrument fingerprint encoding function. The above is explained in detail in combination with the formula:

[0106] The instrument noise in the near-infrared spectrum is usually reflected in the curvature change of the spectrum, and the size of the curvature reflects the aging rate of the instrument. The instrument-related characteristics caused by the instrument noise are extracted by second-order differentiation of the spectrum X. The noise characteristics after second-order differentiation contain a large amount of high-frequency noise information. The high-dimensional noise characteristics are reduced to a low-dimensional noise matrix containing high-frequency noise information by residual convolution, and the formula is In addition to the noise characteristics caused by instrument aging, there are also inherent characteristics of the instrument. The spectrum X is encoded by an instrument fingerprint encoding function to extract the inherent characteristics of the instrument, and the formula is δ(X). The noise characteristics caused by instrument aging and the inherent characteristics of the instrument are fused by feature splicing to realize extraction of instrument-related characteristics.

[0107] In the embodiment of the present application, the above feature adversarial training step further includes:

[0108] The feature extractor learns the input instrument-related characteristics, corrects its own parameters, generates noise information, and integrates it into instrument-independent characteristics;

[0109] The domain discriminator learns the input instrument-independent characteristics, corrects its own parameters, and improves the discrimination ability of the source domain and the target domain spectrum;

[0110] The feature extractor and the domain discriminator are continuously learned through the continuous confrontation, the feature extractor continuously maximizes the confrontation training loss function, the domain discriminator continuously minimizes the confrontation training loss function, and when the maximization and minimization of the confrontation training loss function reach a dynamic balance, the domain-invariant subset is output.

[0111] More specifically, in the embodiment of the present application, the above-mentioned instrument-independent features and instrument-dependent features extracted from the source domain and the target domain spectra are input into the gradient reversal layer through spectral-spatial domain decomposition residual block and spectral-spatial domain decomposition residual block, respectively, and feature confrontation training is performed. The feature extractor in the gradient reversal layer continuously learns the input instrument-dependent features, continuously corrects its own parameters to generate noise information that is difficult for the domain discriminator to distinguish and integrates into the instrument-independent features;

[0112] The domain discriminator continuously learns the input instrument-independent features, continuously corrects its own parameters to improve its own discrimination ability for the source domain and the target domain spectra; through continuous confrontation learning between the two, the feature extractor continuously maximizes the confrontation training loss function, and the domain discriminator continuously minimizes the confrontation training loss function, and when the two reach a dynamic balance, the domain-invariant subset is output The confrontation training loss function L is adv The formula is:

[0113] L adv = E X-S [logD(F(X))]+E X-T [log(1-D(F(X)))]

[0114] Where F is the feature extractor, D is the domain discriminator, X is the input spectrum, S is the main instrument data, and T is the target instrument data. X-S represents that the input spectrum X comes from the source domain S (i.e. the spectrum data collected by the main instrument); X-T represents that the input spectrum X comes from the target domain T (i.e. the spectrum data collected by the target instrument); E X-S : expectation of source domain samples; E X-T : expectation of target domain samples.

[0115] In the embodiment of the present application, the above-mentioned spectral virtual enhancement step 103 further comprises:

[0116] The instrument-dependent features, the target domain unlabeled spectrum data set, and the domain-invariant subset are input into the generator, dimension splicing is performed, and the instrument-dependent features are expanded;

[0117] The features in the wavelength dimension are reserved, the spliced generator output value is encoded and decoded, the noise features in the instrument-dependent features are injected into the target domain unlabeled spectrum data set and the domain-invariant subset, and the virtual spectrum is output.

[0118] More specifically, in the specific embodiment of the present application, the structured noise feature F dep is injected into the target domain unlabeled data X ′ , The domain-invariant subset S→T is input into the generator G dep , and the dimension splicing is performed first, and F ′ is expanded to the same dimension as X dep through a fully connected layer; the spliced input is encoded-decoded using a U-Net structure (which is good at preserving the wavelength dimension of the feature) to inject the noise feature in F ′ into the spectrum X ′ , , and output the virtual spectrum

[0119] wherein the target pre-unlabeled data X S→T is the spectrum data collected by the target instrument, but the corresponding label is not obtained through the laboratory standard method. The role is to provide the generator with the "spectrum structure information of the target domain" so that the generated virtual spectrum not only retains the essential features of the source domain (such as the chemical composition of the material), but also conforms to the instrument characteristics of the target domain (such as baseline drift and noise mode).

[0120] The formula is:

[0121]

[0122] wherein G ′ is the generator, X is the unlabeled target data, is the domain-invariant subset.

[0123] In the embodiment of the present application, the above-mentioned dynamic weighted sampling fine-tuning step 104 further comprises:

[0124] The sampling weight of the fully connected layer is fine-tuned by using the dynamic weighted sampling strategy, and the formula of the weight is:

[0125] Sampling weight = alpha(t) * virtual data sampling probability + (1-alpha(t)) * real data sampling probability; wherein alpha(t) is a dynamic weight coefficient of the virtual data, t is the number of training rounds, and the value of alpha(t) is determined by the dynamic weighted sampling strategy.

[0126] In the embodiment of the present application, the above-mentioned dynamic weighted sampling strategy is:

[0127] When t is in the first value interval, alpha(t) = 0.9-0.02t, and each round randomly collects more than or equal to 90% of the samples from the virtual spectrum data, and less than or equal to 10% of the real spectrum data.

[0128] When t is located in the second value interval, alpha(t) = 0.7-0.02(t-10), the value collected in each round of virtual spectrum is gradually reduced from 70% to 30%, and the real fine tuning set collection is gradually increased from 30% to 70%;

[0129] When t is located in the third value interval, alpha(t) = 0.3, the sampling ratio of virtual spectrum and real fine tuning spectrum is taken according to the preset threshold.

[0130] More specifically, in the embodiment of the application, the sampling probability of virtual data (i.e. virtual spectrum ) and real data is adaptively adjusted through training rounds, and the formula is:

[0131] Sampling weight = alpha(t) * virtual data sampling probability + (1-alpha(t)) * real data sampling probability

[0132] Wherein alpha(t) is a dynamic weight coefficient of virtual data, which decreases with the training round t, and the decay rate of alpha(t) is dynamically adjusted by the validation set RMSE: if the validation set RMSE of a round decreases by more than 5%, the decay rate of alpha(t) is accelerated; if the RMSE increases, the decay is suspended, and the current alpha(t) is maintained to avoid overfitting caused by too high weight of real data. The design is as follows:

[0133] When t is located in the first value interval, i.e. t = 1-10: alpha(t) = 0.9-0.02t, at this time, virtual spectrum data is dominant, 90% samples are collected from virtual spectrum set in each round, and 10% samples are collected from real fine tuning set, and noise characteristics are learned quickly.

[0134] When t is located in the second value interval, i.e. t = 10-30, alpha(t) = 0.7-0.02(t-10), at this time, the weight of virtual spectrum data gradually weakens, and the virtual spectrum set is collected from 70% to 30% in each round, and the real fine tuning set is collected from 30% to 70%.

[0135] When t is located in the third value interval, i.e. t > 30, alpha(t) = 0.3, at this time, real spectrum data is dominant, and the virtual spectrum is fixed: the real fine tuning spectrum is sampled at 3:7 to prevent overfitting.

[0136] During the fine tuning process, Radam is selected as the optimizer to dynamically adjust the learning rate variance and relieve the gradient fluctuation in small sample training. The initial learning rate of the optimizer RAdam is 1x10-3.

[0137] As described above, the method of the application can be better implemented.

[0138] In summary, the application provides a new architecture of a multi-scale adversarial decoupling and virtual enhancement near-infrared spectral model transfer method.

[0139] Embodiment two

[0140] As shown in Figure 3 The embodiment of the application provides a multi-scale adversarial decoupling near-infrared spectral model transfer system, which adopts the multi-scale adversarial decoupling near-infrared spectral model transfer method, and the system comprises:

[0141] The model pre-training module 201 is configured to acquire source domain and target domain near-infrared spectral data sets from a main instrument and a target instrument, perform a time-space data allocation strategy, proportionally divide the spectral data sets in the time and space dimensions, and pre-train a near-infrared spectral model based on the divided spectral data sets;

[0142] The multi-scale adversarial feature decoupling network module 202 is configured to extract instrument-independent features and instrument-dependent features based on the source domain near-infrared spectral data set through a neural network, perform adversarial training, minimize domain difference loss, and output a domain-invariant subset;

[0143] The virtual enhancement controller module 203 is configured to input the instrument-dependent features, the target domain near-infrared spectral data set, and the domain-invariant subset into a generator to generate virtual spectral data;

[0144] The dynamic divider module 204 is configured to adopt a dynamic weighted sampling strategy, adaptively adjust the sampling probabilities of the virtual spectral data and the real spectral data, perform model training fine-tuning, deploy the fine-tuned model to the target instrument, and realize cross-instrument transfer of the near-infrared spectral model.

[0145] The multi-scale adversarial feature decoupling network module comprises a double-branch feature decoupling module, a multi-scale input preprocessing module, and a gradient reversal layer; the virtual enhancement controller module integrates a CycleGAN and a mixed data fine-tuning module; and the dynamic divider module fine-tunes a fully connected layer.

[0146] Embodiment three

[0147] The embodiment of the application provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize the steps of the multi-scale adversarial decoupling near-infrared spectral model transfer method.

[0148] Embodiment four

[0149] The embodiment of the present application provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, when the processor executes the program, the steps of the multi-scale anti-decoupling near-infrared spectrum model transfer method are realized.

[0150] In addition, in combination with Figure 1 The multi-scale anti-decoupling near-infrared spectrum model transfer method of the embodiment of the present application can be realized by an electronic device, such as a computer device. Figure 4 The hardware structure of the computer device according to the embodiment of the present application is shown.

[0151] In some embodiments, the computer device can further include a communication interface 83 and a bus 80. Among them, as shown in the figure, the processor 81, the memory 82, the communication interface 83 are connected through the bus 80 and complete the communication between each other. Figure 4 As shown in the figure, the processor 81, the memory 82, the communication interface 83 are connected through the bus 80 and complete the communication between each other.

[0152] Specifically, the processor 81 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured to implement one or more integrated circuits of the embodiment of the present application.

[0153] The memory 82 can be used to store or cache various data files required for processing and / or communication, and the computer program instructions executed by the processor 81.

[0154] The processor 81 realizes any one of the multi-scale anti-decoupling near-infrared spectrum model transfer methods in the above embodiment by reading and executing the computer program instructions stored in the memory 82.

[0155] The technical features of the above-mentioned embodiments can be combined in any way, in order to make the description simple, not all possible combinations of the technical features in the above-mentioned embodiments are described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0156] The above-mentioned embodiments only express several implementation manners of the present application, the description is more specific and detailed, but it cannot be understood as the limitation of the patent scope of the present application. It should be pointed out that, for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which belong to the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.

Claims

1. A multi-scale counter-decoupling near-infrared spectroscopy model transfer method, characterized in that, The method comprises: A model pre-training step: obtaining a near-infrared spectrum dataset of a source domain and a target domain from an instrument, performing a time-space data allocation strategy in time and space dimensions, proportionally dividing the spectrum dataset, and pre-training a near-infrared spectrum model based on the divided spectrum dataset; A multi-scale adversarial feature decoupling network construction step: based on the near-infrared spectrum dataset of the source domain, instrument-independent features and instrument-dependent features are extracted through a neural network, adversarial training is performed, domain difference loss is minimized, and a domain-invariant subset is output; A spectrum virtual enhancement step: inputting the instrument-dependent features, the near-infrared spectrum dataset of the target domain and the domain-invariant subset into a generator to generate virtual spectrum data; A dynamic weighted sampling fine-tuning step: using a dynamic weighted sampling strategy to adaptively adjust the sampling probability of the virtual spectrum data and real spectrum data, performing model training fine-tuning, deploying the fine-tuned model to a target instrument, and realizing cross-instrument transfer of the near-infrared spectrum model. 2.The method of claim 1, wherein, The model pre-training step further comprises: The near-infrared spectrum dataset of the source domain is divided in time dimension according to data acquisition sequence and in space dimension by dimension reduction algorithm, and after uniform distribution of samples, the dataset is divided into a training set, a validation set and a test set in turn according to a self-defined proportion for model pre-training.

3. The multi-scale counter-decoupling near-infrared spectroscopy model transfer method according to claim 1, characterized in that, The multi-scale adversarial feature decoupling network construction step further comprises: An independent feature extraction step: inputting the near-infrared spectrum dataset of the source domain into a first branch of a double-branch neural network to extract instrument-independent features; the first branch uses a spectral-spatial domain decomposition residual block; A related feature extraction step: inputting the near-infrared spectrum dataset of the source domain into a second branch of the double-branch neural network to extract instrument-dependent features; the second branch uses a noise perception residual module; A feature adversarial training step: inputting the extracted instrument-independent features and instrument-dependent features into a gradient inversion layer, performing adversarial training through a domain discriminator and a feature extractor in the gradient inversion layer, and minimizing domain difference loss.

4. The multi-scale counter-decoupling near-infrared spectroscopy model transfer method according to claim 3, characterized in that, The independent feature extraction step further comprises: In the first branch, low-pass filtering is performed on the near-infrared spectrum data to obtain high-dimensional signals, a projection matrix is obtained by performing principal component analysis on the spectrum data, a preset number of principal components are extracted, and the filtered high-dimensional signals are converted into spectrum data of signals of the preset number of dimensions; A gating attention mechanism is applied to the spectrum data to mask wavelength points that do not contain material feature information, automatically identify effective wavelength segment regions, and perform global pooling; the results after processing the spectrum dimension and the space dimension are connected by jumping, and instrument-independent features are obtained which retain original spectrum information.

5. The multi-scale counter-decoupling near-infrared spectroscopy model transfer method according to claim 3, characterized in that, The related feature extraction step further comprises: In the second branch, second-order differentiation is performed on the near-infrared spectrum data to extract instrument-dependent noise features caused by instrument aging; The noise features after second-order differentiation are reduced in dimension by residual convolution to a low-dimensional noise matrix containing high-frequency noise information; The near-infrared spectrum data is encoded by an instrument fingerprint coding function to extract inherent characteristics of the instrument; The inherent characteristics and the noise characteristics of the instrument are fused by feature splicing to realize extraction of instrument-related characteristics.

6. The multi-scale counter-decoupling near-infrared spectroscopy model transfer method according to claim 3, characterized in that, The feature adversarial training step further comprises: The feature extractor learns the instrument-related characteristics of the input, corrects its own parameters, generates noise information, and integrates it into the instrument-independent characteristics; The domain discriminator learns the instrument-independent characteristics of the input, corrects its own parameters, and improves the discrimination ability of the source domain and the target domain spectrum; The feature extractor and the domain discriminator continuously learn by confrontation, the feature extractor continuously maximizes the adversarial training loss function, and the domain discriminator continuously minimizes the adversarial training loss function. When the dynamic balance of maximizing the adversarial training loss function and minimizing the adversarial training loss function is reached, the output domain invariant subset is output.

7. The multi-scale counter-decoupling near-infrared spectroscopy model transfer method according to claim 1, characterized in that, The spectrum virtual enhancement step further comprises: The instrument-related characteristics, the target domain unlabeled spectrum data set, and the domain invariant subset are input into the generator for dimension splicing, and the instrument-related characteristics are expanded; The features in the wavelength dimension are retained, the generator output value after splicing is encoded and decoded, the noise characteristics in the instrument-related characteristics are injected into the target domain unlabeled spectrum data set and the domain invariant subset, and the virtual spectrum is output.

8. The multi-scale counter-decoupling near-infrared spectroscopy model transfer method according to claim 1, characterized in that, The dynamic weighted sampling fine-tuning step further comprises: The dynamic weighted sampling strategy is used to fine-tune the sampling weight of the fully connected layer, and the formula of the weight is: Sampling weight = α(t) × virtual data sampling probability + (1-α(t)) × real data sampling probability; Wherein, α(t) is a virtual data dynamic weight coefficient, t is the number of training rounds, and the value of α(t) is determined by the dynamic weighted sampling strategy.

9. The multi-scale counter-decoupling near-infrared spectroscopy model transfer method according to claim 8, characterized in that, The dynamic weighted sampling strategy is: When t is in the first value interval, α(t) = 0.9-0.02t, more than or equal to 90% of the samples are collected from the virtual spectrum data in each round, and less than or equal to 10% of the samples are collected from the real spectrum data; When t is in the second value interval, α(t) = 0.7-0.02(t-10), the sampling value in the virtual spectrum set gradually decreases from 70% to 30% in each round, and the real fine-tuning set gradually increases from 30% to 70%; When t is in the third value interval, α(t) = 0.3, the sampling ratio of virtual spectrum and real fine-tuning spectrum is determined according to the preset threshold.

10. A multi-scale adversarial decoupled near-infrared spectroscopy model transfer system, adopting the multi-scale adversarial decoupled near-infrared spectroscopy model transfer method according to any one of claims 1-9, characterized in that, The system comprises: A model pre-training module is configured to obtain source domain and target domain near-infrared spectrum data sets from a main instrument and a target instrument, perform a time-space data allocation strategy, proportionally divide the spectrum data sets in time and space dimensions, and pre-train a near-infrared spectrum model based on the divided spectrum data sets; A multi-scale adversarial feature decoupling network module is configured to extract instrument-independent characteristics and instrument-related characteristics based on the near-infrared spectrum data set of the source domain, perform adversarial training, minimize domain difference loss, and output a domain invariant subset. A virtual augmented controller module is configured to input the instrument-related features, the near-infrared spectroscopy dataset of the target domain and the domain-invariant subset into a generator to generate virtual spectroscopy data; A dynamic partitioner module is configured to adopt a dynamic weighted sampling strategy to adaptively adjust the sampling probability of the virtual spectroscopy data and the real spectroscopy data, perform model training fine-tuning, deploy the fine-tuned model to a target instrument, and achieve cross-instrument transfer of the near-infrared spectroscopy model.

11. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the multi-scale adversarial decoupling near-infrared spectroscopy model transfer method of any one of claims 1-9.

12. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps of the multi-scale adversarial decoupling near-infrared spectroscopy model transfer method of any one of claims 1-9.

Citation Information

Patent Citations

  • Method for transferring on-line near infrared spectrum model of assumed standard sample

    CN113158575A

  • Near-infrared spectroscopy-based method for chemical pattern recognition of authenticity of traditional chinese medicine gleditsiae spina

    US20210025815A1

  • Systems and methods for synthesizing data for training statistical models on different imaging modalities including polarized images

    US20220215266A1

Cited By

  • Soybean producing area traceability detection method based on near infrared spectrum technology

    CN122016714A