Network architecture search synchronous transfer learning method oriented to multi-modal map data
The network architecture search transfer learning method addresses the inefficiencies in multi-modal data fusion by using CNN-RNN architectures and attention mechanisms, enhancing model precision and reducing annotation dependency, thus improving agricultural spectroscopy analysis.
Patent Information
- Application Number
- CN202510819528.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The fusion efficiency of near-infrared spectroscopy, Raman spectroscopy, and image data is low, with significant heterogeneity in data dimensions and feature spaces, leading to insufficient utilization of complementary information and dimensionality and semantic gaps, while existing methods heavily rely on large annotated datasets, posing a challenge in constructing high-precision models for agricultural spectroscopy.
A network architecture search (NAS) based transfer learning method is employed, combining CNN and RNN architectures, utilizing self- and cross-attention mechanisms for feature extraction and fusion, along with semi-supervised training and active learning to generate pseudo-labels, and loss weight optimization to reduce feature correlation.
Enhances model generalization by effectively fusing multi-modal data, reducing reliance on extensive annotations, and optimizing feature representation, thereby improving precision in agricultural spectroscopy under data scarcity.
Smart Images

Figure CN120317331A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of multimodal data fusion analysis. Specifically, it relates to a network architecture search synchronous transfer learning method for multimodal atlas data. Background Art
[0002] In the agricultural field, near-infrared spectroscopy, Raman spectroscopy, and microscopic image detection technologies have been widely studied and applied due to their advantages of portability and high efficiency.
[0003] However, currently, the fusion efficiency of spectra and images is low, and the model methods are poor. There are significant heterogeneities in data dimensions and feature spaces between the two. Traditional fusion methods are difficult to break through the modality gap, resulting in insufficient utilization of complementary information and double problems of dimensional disaster and semantic fault in feature expression, making it difficult to meet the actual requirements of high-precision and low annotation cost in agricultural spectrum detection.
[0004] In addition, mainstream machine learning and deep learning methods rely heavily on large-scale labeled data during model training and verification. However, the shortage of labeled data has become the main bottleneck in constructing high-precision models.
[0005] Therefore, in the case of low atlas fusion efficiency, poor methods, and insufficient labeled sample data, how to efficiently construct an effective model has become a difficult problem in the field of multimodal atlas fusion modeling analysis. Summary of the Invention
[0006] This application provides a network architecture search synchronous transfer learning method for multimodal atlas data to at least solve the technical problems of low atlas fusion efficiency, poor methods, and the inability to construct an effective model in the case of insufficient labeled sample data in related technologies.
[0007] According to one aspect of this application, there is provided a network architecture search synchronous transfer learning method for multimodal atlas data. The transfer method includes the following steps: Extract and fuse the collected near-infrared spectra, Raman spectra, and microscopic images; In the constructed SSL_NST module, the fused dataset is divided into a source domain and a target domain through transfer learning. The neural network architecture search NAS adopts a three-stage optimization strategy including network search, fine-tuning, and feedback; where: In the network search stage, replace the default search space of neural architecture search (NAS) with the combined architecture space of CNN and RNN, and use neural architecture search (NAS) to search for the source domain model on the source domain data; in the fine-tuning stage, use the semi-supervised self-training algorithm to assign pseudo-labels to the target domain data based on the source domain model, and use the maximum entropy sampling in active learning to screen out the pseudo-label data. Freeze the network structure of the source domain model, and use the screened pseudo-label data to fine-tune the source domain model. By operating on unfreezing network layers layer by layer, adjusting the optimizer and learning rate, minimize the target domain loss; in the feedback stage, use the loss of the fine-tuned model on the target domain task as a feedback signal to act on the search strategy of network search and guide the next round of search process; after a preset number of iterations, the searched network structure is used as the target domain model for the target domain task.
[0008] Optionally, the steps of feature extraction and fusion for the collected near-infrared spectra, Raman spectra, and microscopic images include:
[0009] Normalize the obtained near-infrared spectra, Raman spectra, and microscopic images using Min-Max normalization and preprocess them using multiplicative scatter correction (MSC);
[0010] Use the self-attention mechanism to extract 128-dimensional optimal features F-nir and F-raman from the preprocessed near-infrared spectra and Raman spectra respectively;
[0011] Form a 256-dimensional fusion feature F-nir-raman from the 128-dimensional optimal features of the near-infrared spectra and Raman spectra using the concatenation method;
[0012] Adjust the image size of the preprocessed microscopic image to 224×224 pixels and use the self-attention mechanism to extract 256-dimensional optimal features F-micro image;
[0013] Use the cross-attention mechanism to perform feature fusion on the 256-dimensional fusion feature F-nir-raman and the 256-dimensional optimal feature F-micro image to form 512-dimensional optimal fusion feature F-fusion.
[0014] Optionally, the step of using the self-attention mechanism to extract 256-dimensional optimal features F-micro image includes:
[0015] Set the input feature sequence X∈R n×d , when extracting 128-dimensional features from the near-infrared spectra and Raman spectra, first pass through the learnable weight matrices W Q , W K , W V ∈R d×128, the input feature X is respectively mapped to a query vector Q = XW of 128 dimensions Q , a key vector K = XW K and a value vector V = XW V ; then calculate the dot product of the query vector and all key vectors and divide by the square root of the dimension for scaling, and the scaling factor is , and then generate an attention weight matrix through the Softmax function, which characterizes the correlation degree of the features at each position in the input sequence; finally, multiply the attention weights by the value vectors to obtain a 128-dimensional self-attention output feature that fuses the internal dependencies of the sequence ,
[0016] When the microscopic image extracts 256-dimensional features, the input features are mapped to 256-dimensional query, key, and value vectors; the dimension scaling factor is adjusted to , and a 256-dimensional self-attention output feature is generated , expressed as:
[0017] 128-dimensional self-attention output feature and 256-dimensional self-attention output feature are respectively expressed as:
[0018]
[0019]
[0020] In the formula, X represents the input feature, represents the weight matrix of the value vector V, represents the weight matrix of the query vector Q, represents the weight matrix of the key vector K, and T represents the transpose of the matrix.
[0021] Optionally, the step of using the cross-attention mechanism to perform feature fusion to form a 512-dimensional optimal fusion feature F-fusion includes:
[0022] Set as the 256-dimensional fusion feature after splicing the near-infrared spectrum and the Raman spectrum, as the 256-dimensional feature extracted from the microscopic image, first through the learnable weight matrices , , , map the spectral fusion feature to a query vector , map the image feature to a key vector and a value vector ;
[0023] Then, calculate the dot product of the spectral query vector and the image key vector, scale it, and generate cross-modal attention weights through Softmax. These weights represent the dependence of spectral features on image features;
[0024] Multiply the attention weights by the image value vector to obtain the response features of the image features to the spectral features;
[0025] Finally, through the Concat operation, concatenate the original spectral fusion feature F3 and the cross-modal response feature along the feature dimension to generate a 512-dimensional cross-modal fusion feature F5. The 512-dimensional fusion feature F5 output by the cross-attention mechanism can be expressed as:
[0026]
[0027] In the formula, F3 represents the spectral fusion feature, represents the learnable weight matrix of the query vector, F4 represents the image feature, represents the learnable weight matrix of the key vector, represents the learnable weight matrix of the value vector, and T represents the transpose of the matrix.
[0028] Optionally, the step of using the semi-supervised self-training algorithm to assign pseudo-labels to the target domain data and using the maximum entropy sampling in active learning to screen out the pseudo-label data includes:
[0029] Set the fusion dataset as , where x i represents the sample of the fusion dataset, y i represents the label corresponding to the sample, and n is the number of labeled samples; the source domain dataset is , where m is the number of source domain samples, and x j represents the sample of the source domain dataset;
[0030] The neural network model generated by the NAS architecture is used as the classification model S. Based on fully learning the data distribution P(x, y) of the source domain data, the target domain samples are screened through the confidence threshold mechanism in the self-training algorithm. Among them, first use the model C to predict each sample x u in the unlabeled sample set D j of the target domain, and obtain the prediction result and its corresponding prediction confidence confidence(x j ); then set a confidence threshold τ, and by comparing the prediction confidence of each sample with this threshold, screen out the samples with a confidence higher than the threshold. The screened target domain sample set D t consists of all samples x that satisfy the condition confidence(x j ) > τ jComposed of the selected target domain sample set D t It is expressed as:
[0031] ;
[0032] In the formula, x j represents a sample, D u represents the sample set, and τ represents the confidence threshold;
[0033] Then, calculate the information entropy of the selected target domain samples. During the calculation of the information entropy, first calculate the conditional probability p(y c |x) of the target domain sample x belonging to each category c, then take the natural logarithm of the conditional probability of each category, multiply it by the conditional probability, and take the negative value to obtain the uncertainty contribution of a single category; finally, sum the uncertainty contributions corresponding to all categories to obtain the information entropy H(x) of the sample x. Among them, the information entropy of the target domain sample x j is expressed as:
[0034]
[0035] In the formula, c represents the category of the target domain sample x, c = 1, 2,..., C, and C is the total number of categories; p(y c |x) represents the conditional probability, that is, the posterior probability that the sample x belongs to the category c under the given features;
[0036] Finally, in each iteration t, combine the idea of maximum entropy sampling based on uncertainty in active learning to sample the target domain data set D t to obtain the required sample subset D t = S(D v ) according to the sampling function S(D t ), and hand these samples to the source domain model for annotation and assign them pseudo-labels .
[0037] Preferably, the step of using the neural network architecture search NAS to search for the source domain model on the source domain data includes:
[0038] During the entire search process, in the first stage of the search, model the prediction loss of the candidate network architecture on the source domain data D u through Gaussian process, which is expressed as:
[0039]
[0040] Among them, u(·) represents the mean of the Gaussian process prediction loss, σ(·) represents the standard deviation of the Gaussian process prediction loss, β represents the balance factor, D u represents the source domain data set, Cost(f, D u) represents the prediction loss of the network architecture f on the source domain data D u The degree of attention to the mean and uncertainty of the model performance is adjusted by the balance factor β:
[0041] When β is large, the search process tends to explore architectures with high uncertainty;
[0042] When β is small, the search process pays more attention to the verified low-loss architectures;
[0043] On the basis of generating candidate architectures in the first stage by minimizing the prediction loss of the candidate architectures on the target domain data D t the most suitable network architecture for the target domain is screened out This optimization objective guides the NAS process to search in a direction with stronger cross-domain adaptability by directly relating to the target domain performance, realizing the synchronous and stable migration of the model structure. The optimization objective is expressed as: wherein, in the formula, represents the prediction loss of the candidate architecture on the target domain data D t represents the candidate architecture generated in the first stage, represents the screened network architecture, and argmin represents the function for calculating the minimum value.
[0044] Preferably, when fine-tuning the model, the optimal network structure f is selected from the set of candidate network structures F CR such that its prediction loss t-test on the target domain test set D is minimized. The process of minimizing the target domain loss is expressed as:
[0045]
[0046] where, is the optimal parameter obtained by pre-training or previous stage learning of the network structure f. This process screens out the network architecture that best adapts to the feature distribution of the target domain by evaluating the performance of different network structures on the target domain test data, ensuring the generalization ability of the model in cross-domain scenarios;
[0047] During the fine-tuning process, the Adam optimizer is used as the optimization algorithm, the activation function is set to the Sigmoid function, and the loss function is the binary cross-entropy function.
[0048] Preferably, the steps of this migration method further include: Introduce a loss-weighted optimization strategy into the obtained target domain model to reduce the correlation between fused features, adjust the ratio between the classification loss value and the sample weight, and optimize the training process of the fused data model; Among them:
[0049] Use the random Fourier feature (RFF) to map the features to a high-dimensional feature space;
[0050] During the weight learning process, by concatenating the global feature sets (F G1 , F G2 ,..., F Gi , F L ) and the current batch of local features F L , form the comprehensive feature F0 for the next batch optimization; similarly, concatenate the global weight sets (W G1 , W G2 ,..., W Gi , W L ) and the current batch of local weights W L to obtain the comprehensive weight W0. Reduce the storage and calculation costs by the method of iterative saving and reloading weights. The formulas for the features F0 and weights W0 used to optimize the sample weights of the next batch in each batch are as follows:
[0051]
[0052]
[0053] Among them, the global feature F Gi and the global weight W Gi represent the global information accumulated in historical batches, which remain unchanged during the training process of each batch. Only the current batch of local features F L and local weights W L can be trained and updated under the condition of minimizing the feature dependence relationship;
[0054] Then, at the end of each iteration, through the smoothing parameter perform weighted fusion on the global information (F Gi , W Gi ) and the current batch of local information (F L , W L ) to obtain the global information for the next round, which is expressed as follows:
[0055]
[0056]
[0057] In the formula, is used to adjust the memory length of the global information: when When approaching 1, the model relies more on historical global information; when approaching 0, the model pays more attention to the local information of the current batch; this mechanism realizes the dynamic balance between global knowledge and real-time features, avoiding over-reliance on historical information or local noise;
[0058] Then, the maximum reduction of the dependence between features is achieved by minimizing the Frobenius norm of the cross-covariance matrix between weighted sample features. The square of the Frobenius norm of the covariance matrix of the weighted features The expression is as follows:
[0059] where represents the weighted mutual information index between feature groups A and B, which is used to measure the correlation strength between the two. The smaller its value, the weaker the correlation between features; is the square of the Frobenius norm, that is, the sum of the squares of the matrix elements;
[0060] Furthermore, the weighted covariance matrix is expressed as follows:
[0061] In the formula, u(·) and v(·) are random Fourier feature RFF mapping functions that map the original features to a high-dimensional space to simplify linear calculations; w i is the sample weight, which satisfies the normalization condition. This norm is used as an independence test statistic. The smaller its value, the lower the linear correlation between features A and B. Therefore, by optimizing the weight w, the model is forced to learn low-correlated fusion features, improving the independence of feature representation;
[0062] Finally, the weight matrix W(*) that achieves the optimal feature independence is element-wise multiplied by the classification loss cot(*) of the model to form the final training loss function, expressed as: Loss = cot(*) ⊙ W(*), where ⊙ represents element-wise multiplication.
[0063] Compared with the prior art, the technical advantages of this application are as follows:
[0064] First, the present invention performs optimal feature fusion on near-infrared spectroscopy data, Raman spectroscopy data, and microscopic image data. The features fused by their respective self-attention mechanisms are further fused through a cross-attention mechanism to form optimal fusion features, which provides effective help for comprehensively analyzing the condition of crops. The complementarity of multi-source data can effectively enhance the generalization ability of the model;
[0065] Second, the present invention further combines semi-supervised active learning, neural network architecture search (NAS), and transfer learning techniques, and uses a loss-stabilizing weighting strategy to significantly reduce the dependence on a large amount of labeled data during model training, improve the applicability of the model in data-scarce scenarios, and capture more effective features by optimizing the loss function, thereby reducing the correlation between features, optimizing the model training process, and improving the transfer effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application.
[0067] In the drawings:
[0068] Figure 1 is a flowchart of the implementation of the network architecture search and synchronous transfer learning method for multi-modal graph data according to an embodiment of the present application;
[0069] Figure 2 is a system architecture diagram of feature fusion in the network architecture search and synchronous transfer learning method for multi-modal graph data according to an embodiment of the present application;
[0070] Figure 3 is a system architecture diagram of network architecture search in the network architecture search and synchronous transfer learning method for multi-modal graph data according to an embodiment of the present application;
[0071] Figure 4 is a system architecture diagram of loss-stabilizing weighting in the network architecture search and synchronous transfer learning method for multi-modal graph data according to an embodiment of the present application;
[0072] Figure 5 is a schematic structural diagram of the network architecture search and synchronous transfer learning system for multi-modal graph data according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0074] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0075] According to an embodiment of this application, a method embodiment of a network architecture search synchronous transfer learning method for multi-modal atlas data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0076] Figure 1 is a flowchart of a network architecture search synchronous transfer learning method for multi-modal atlas data according to an embodiment of this application, as Figure 1 shown, the synchronous transfer learning method includes the following steps: S101: Extract and fuse features from the collected near-infrared spectra, Raman spectra and microscopic images; Specifically, after preprocessing the collected near-infrared spectra, Raman spectra and microscopic images, feature extraction is performed respectively, and the extracted features are fused through a cross-attention mechanism to form a new optimal fused feature;
[0077] As Figure 1 and Figure 2 shown, in the above step S101, by fusing multi-source data with information complementarity such as near-infrared spectra, Raman spectra and microscopic images, the features of the three types of data are optimized and fused, providing effective support for comprehensively analyzing the crop state;
[0078] As Figure 1 and Figure 3 shown, the synchronous transfer learning method provided by the embodiments of this disclosure further includes the following steps:
[0079] S102: In the constructed SSL_NST module, the fused data set is divided into a source domain and a target domain through transfer learning technology, and the neural network architecture search NAS adopts a three-stage optimization strategy including network search, fine-tuning and feedback;
[0080] Specifically, in the network search stage, replace the default search space of neural architecture search (NAS) with the combined architecture space of CNN and RNN, and use neural architecture search (NAS) to search for the source domain model on the source domain data;
[0081] Among them, this application replaces the default search space of neural architecture search (NAS) with the combined architecture (CR) space of CNN and RNN, which not only combines the ability of CNN to extract local and abstract features from the original spectrum, but also combines the advantages of RNN in learning various dependencies of sequence features; in the task of fusing data, a complex model structure may have a high fitting ability on the fused features. Therefore, the convolutional neural network combined with the recurrent neural network continues to show good classification performance;
[0082] In the fine-tuning stage, based on the source domain model, use the semi-supervised self-training algorithm to assign pseudo-labels to the target domain data, and use the maximum entropy sampling in active learning to screen out the pseudo-labeled data. Freeze the network structure of the source domain model, and use the screened pseudo-labeled data to fine-tune the source domain model. By operating layer by layer to unfreeze the network layers, adjust the optimizer and learning rate, minimize the target domain loss;
[0083] In the feedback stage, use the loss of the fine-tuned model on the target domain task as a feedback signal to act on the search strategy of network search and guide the next round of search process; after a preset number of iterations, the searched network structure is used as the target domain model for the target domain task;
[0084] In step S102 of the present disclosure, the SSL_NST module divides the fused dataset into the source domain and the target domain through transfer learning technology. The neural architecture search (NAS) adopts a three-stage optimization strategy, mainly including three stages: network search, fine-tuning, and feedback. In the fine-tuning mechanism, a semi-supervised active learning method is introduced to assign pseudo-labels to the target domain data, so that the searched model architecture can better adapt to the target task during the optimization process;
[0085] As Figure 1 and Figure 4 shown, the synchronous transfer learning method provided by the embodiments of the present disclosure further includes the following steps:
[0086] S103: Introduce a loss weighting optimization strategy in the obtained target domain model to reduce the correlation between the fused features, adjust the ratio between the classification loss value and the sample weight, and optimize the training process of the fused data model.
[0087] The network architecture search synchronous transfer learning method of this application combines semi-supervised active learning and transfer learning techniques, and then uses a loss-stable weighting strategy to transfer the knowledge of related fields to the target task, enabling the model to maintain stable learning ability under the condition of a small amount of labeled training data. Moreover, by optimizing the loss function to capture more effective features, the correlation between features is reduced, the training process of the model is optimized, the transfer effect is improved, the data scarcity problem of asymptomatic samples is effectively alleviated, and early diagnosis of crop diseases is achieved.
[0088] In the above embodiment, the steps of feature extraction and fusion for the collected near-infrared spectrum, Raman spectrum, and microscopic image include: normalizing the obtained near-infrared spectrum, Raman spectrum, and microscopic image using Min-Max normalization and preprocessing them using multiplicative scatter correction (MSC); respectively extracting 128-dimensional optimal features F-nir and F-raman from the preprocessed near-infrared spectrum and Raman spectrum using the self-attention mechanism; forming a 256-dimensional fusion feature F-nir-raman from the 128-dimensional optimal features of the near-infrared spectrum and Raman spectrum using the concatenation method; adjusting the image size of the preprocessed microscopic image to 224×224 pixels and extracting 256-dimensional optimal features F-micro image using the self-attention mechanism; and performing feature fusion on the 256-dimensional fusion feature F-nir-raman and the 256-dimensional optimal feature F-micro image using the cross-attention mechanism to form a 512-dimensional optimal fusion feature F-fusion.
[0089] Further, the steps of extracting 256-dimensional optimal features F-microimage using the self-attention mechanism provided by the embodiments of the present disclosure include: setting the input feature sequence , when extracting 128-dimensional features from the near-infrared spectrum and Raman spectrum, first passing through learnable weight matrices W Q , W K , W V ∈ R d×128 to map the input features to 128-dimensional query vectors , key vectors , and value vectors respectively; then calculating the dot product of the query vector and all key vectors and dividing by the square root of the dimension for scaling, with the scaling factor being , and then generating an attention weight matrix through the Softmax function, which characterizes the correlation degree of the features at each position in the input sequence; finally, multiplying the attention weight by the value vector to obtain a 128-dimensional self-attention output feature that fuses the internal dependencies of the sequence, realizing the context semantic modeling of the spectral modal features;
[0090] When extracting 256-dimensional features from microscopic images, for the high-dimensional feature requirements of microscopic images, through a learnable weight matrix , , the input features are mapped into query, key, and value vectors of 256 dimensions; adopt a scaled dot-product attention calculation process similar to the spectral modality, but adjust the dimensional scaling factor to , and finally generate self-attention output features of 256 dimensions , adapt to the high-dimensional feature representation requirements of the image modality, strengthen the long-distance dependence modeling ability of image features, and generate self-attention output features of 256 dimensions :
[0091] Therefore, the 128-dimensional self-attention output features and the 256-dimensional self-attention output features are respectively expressed as:
[0092]
[0093]
[0094] In the formula, X represents the input features, represents the weight matrix of the value vector V, represents the weight matrix of the query vector Q, represents the weight matrix of the key vector K, and T represents the transpose of the matrix.
[0095] The steps of using the cross-attention mechanism to perform feature fusion to form the 512-dimensional optimal fusion feature F-fusion include:
[0096] Set as the 256-dimensional fusion feature after splicing the near-infrared spectrum and Raman spectrum, as the 256-dimensional feature extracted from the microscopic image. First, through the learnable weight matrices , , , map the spectral fusion feature into the query vector , map the image feature into the key vector and the value vector ; then calculate the dot product of the spectral query vector and the image key vector and scale it, and generate cross-modal attention weights through Softmax. These weights represent the dependence relationship of spectral features on image features;
[0097] Multiply the attention weights by the image value vector to obtain the response features of the image features to the spectral features;
[0098] Finally, through the Concat operation, the original spectral fusion feature F3 and the cross-modal response feature are concatenated along the feature dimension to generate a 512-dimensional cross-modal fusion feature F5, thus realizing the multi-modal information interaction and feature fusion of near-infrared spectroscopy, Raman spectroscopy, and microscopic images. The 512-dimensional fusion feature F5 output by the cross-attention mechanism can be expressed as:
[0099]
[0100] In the formula, F3 represents the spectral fusion feature, represents the learnable weight matrix of the query vector, F4 represents the image feature, represents the learnable weight matrix of the key vector, represents the learnable weight matrix of the value vector, and T represents the transpose of the matrix.
[0101] Optionally, the step of using the semi-supervised self-training algorithm to assign pseudo-labels to the target domain data and using the maximum entropy sampling in active learning to screen out the pseudo-label data includes:
[0102] Set the fusion dataset as , where x i represents the sample of the fusion dataset, y i represents the label corresponding to the sample, and n is the number of labeled samples; the source domain dataset is , where m is the number of source domain samples;
[0103] The neural network model generated by the NAS architecture is used as the classification model S. On the basis of fully learning the data distribution P(x, y) of the source domain data (here P(x, y) represents the joint probability distribution), the target domain samples are screened through the confidence threshold mechanism in the self-training algorithm. Among them, first use the model C to predict each sample x u in the unlabeled sample set D j of the target domain, and obtain the prediction result and its corresponding prediction confidence confidence(x j ). Then set a confidence threshold τ, and by comparing the prediction confidence of each sample with this threshold, screen out the samples with a confidence higher than the threshold. The screened target domain sample set is composed of all samples x j that satisfy the condition confidence(x j ) > τ, that is, only retain the samples with higher credibility of the model prediction results for subsequent tasks to improve the performance of the model in the target domain;
[0104] The screened target domain sample set is expressed as:
[0105]
[0106] where x j represents a sample, D u represents a sample set, and τ represents a confidence threshold;
[0107] Then, calculate the information entropy of the filtered target domain samples. The larger the information entropy, the more information the sample contains. During the calculation of the information entropy, first calculate the conditional probability p(y c |x) of the target domain sample x belonging to each category c. Then take the natural logarithm of the conditional probability of each category, multiply it by the conditional probability, and take the negative value to obtain the uncertainty contribution of a single category. Finally, sum the uncertainty contributions corresponding to all categories to obtain the information entropy H(x) of the sample x. The larger the information entropy, the more ambiguous the category attribution of the sample and the higher the potential information value it contains;
[0108] where the information entropy of the target domain sample x j is expressed as:
[0109]
[0110] where c represents the category of the target domain sample x, c = 1, 2,..., C, and C is the total number of categories; p(y c |x) represents the conditional probability, that is, the posterior probability that the sample x belongs to category c under the given features;
[0111] Finally, in each iteration combine the idea of maximum entropy sampling based on uncertainty in active learning to sample the target domain data set D t . Let the sampling function be S(D t ), and find the most valuable sample subset D v = S(D t ) through this function, and hand these samples to the source domain model for annotation and assign them pseudo-labels .
[0112] Preferably, the step of searching for the source domain model on the source domain data using the Neural Architecture Search NAS includes:
[0113] During the entire search process, in the first stage of the search, model the prediction loss of the candidate network architecture f on the source domain data D u using a Gaussian process, which is expressed as:
[0114]
[0115] Among them, u(·) represents the mean of the Gaussian process prediction loss, reflecting the average performance of the model on the source domain; σ(·) represents the standard deviation of the Gaussian process prediction loss, used to characterize the uncertainty of the model performance;
[0116] The degree of attention to the mean and uncertainty of the model performance is adjusted by the balance factor β:
[0117] When β is large, the search process is more inclined to explore architectures with high uncertainty, which helps to discover potentially better model structures;
[0118] When β is small, the search process pays more attention to the verified low-loss architectures, accelerating convergence to the local optimal solution;
[0119] Generate candidate architectures in the first stage Based on this, by minimizing the candidate architectures On the target domain data D t The prediction loss on Filter out the network architecture most suitable for the target domain In this process, a small amount of labeled data in the target domain is used to evaluate and select candidate architectures, ensuring that the finally obtained model structure can be effectively migrated to the target domain and alleviating the domain difference between the source domain and the target domain; this optimization objective guides the NAS process to search in a direction with stronger cross-domain adaptability by directly associating with the target domain performance, realizing the synchronous and stable migration of the model structure, and the optimization objective is expressed as: , where Represents the prediction loss of the candidate architecture On the target domain data D t On, Represents the candidate architecture generated in the first stage, Represents the selected network architecture, and argmin represents the function used to calculate the minimum value.
[0120] Preferably, when fine-tuning the model, select the optimal network structure f from the set of candidate network structures F CR To minimize its prediction loss t-test On the target domain test set D The process of minimizing the target domain loss is expressed as:
[0121]
[0122] Among them, Is the optimal parameter obtained by the network structure f through pre-training or previous stage learning. This process evaluates the performance of different network structures on the target domain test data, filters out the network architecture most adaptable to the target domain feature distribution, and ensures the generalization ability of the model in the cross-domain scenario;
[0123] During the fine-tuning process, the Adam optimizer is used for the optimization algorithm, the activation function is set to the Sigmoid function, and the loss function is the binary cross-entropy function.
[0124] Preferably, the step of introducing a loss-weighted optimization strategy into the obtained target domain model to reduce the correlation between fusion features and adjust the ratio between the classification loss value and the sample weight to optimize the training process of the fusion data model includes:
[0125] The random Fourier features (RFF) are used to map the features to a high-dimensional feature space to reduce feature complexity and enable linear calculations;
[0126] During the weight learning process, by concatenating the global feature sets (F G1 , F G2 ,..., F Gi , F L ) with the current batch of local features F L , a comprehensive feature F0 for the next batch of optimization is formed; similarly, by concatenating the global weight sets (W G1 , W G2 ,..., W Gi , W L ) with the current batch of local weights W L , the comprehensive weight W0 is obtained. The method of iteratively saving and reloading weights is used to reduce storage and computational costs. The formulas for the features F0 and weights W0 used to optimize the sample weights for the next batch in each batch are as follows:
[0127]
[0128]
[0129] Among them, the global feature F Gi and the global weight W Gi represent the global information accumulated in historical batches. Specifically, F Gi represents the global feature accumulated in the i-th group of history, and W Gi represents the corresponding accumulated weight; they remain unchanged during the training process of each batch, and only the local features and local weights of the current batch can be trained and updated under the condition of minimizing feature dependence;
[0130] Then, at the end of each iteration, through the smoothing parameter , the global information (F Gi , W Gi ) and the local information of the current batch (F L , W LPerform weighted fusion to obtain the global information of the next round, which is expressed as follows:
[0131]
[0132]
[0133] In the formula, is used to adjust the memory length of the global information: when approaches 1, the model relies more on historical global information; when approaches 0, the model pays more attention to the local information of the current batch; this mechanism realizes the dynamic balance between global knowledge and real-time features, avoiding over-reliance on historical information or local noise; represents the global feature after weighted fusion, represents the global weight after weighted fusion, F L and W L respectively represent the local feature and local weight of the current batch, represents the smoothing parameter, F Gi and W Gi respectively represent the global feature and global weight;
[0134] Then, minimize the Frobenius norm of the cross-covariance matrix between the weighted sample features to minimize the dependence between features as much as possible. The square of the Frobenius norm of the covariance matrix of the weighted features is expressed as follows:
[0135] ;
[0136] Among them, represents the weighted mutual information index between feature groups A and B, which is used to measure the correlation strength between the two. The smaller its value, the weaker the correlation between features; is the square of the Frobenius norm, that is, the sum of the squares of the matrix elements;
[0137] Furthermore, the weighted covariance matrix is expressed as follows:
[0138]
[0139] Among them, Each element represents the covariance between a certain feature dimension of A and a certain feature dimension of B; n represents the total number of samples, A i and B idenote the relevant and irrelevant features of sample i respectively, T represents the transpose of a matrix or vector, which is used to calculate the correlation between features; u(·) and v(·) are random Fourier feature (RFF) mapping functions that map the original features to a high-dimensional space; w i is the sample weight, satisfying the normalization condition, which weakens the influence of spurious associations by adaptively adjusting the contribution of samples to the correlation calculation; and are weighted mean vectors, ensuring that the correlation measure is not interfered by the mean shift;
[0140] This norm serves as an independence test statistic. The smaller its value, the lower the linear correlation between features A and B. Thus, by optimizing the weight w, the model is forced to learn low-correlated fused features, enhancing the independence of feature representation;
[0141] Finally, the weight matrix W(*) that achieves the optimal feature independence is element-wise multiplied with the classification loss cot(*) of the model to form the final training loss function, expressed as: Loss = cot(*) ⊙ W(*), where ⊙ represents element-wise multiplication.
[0142] In the embodiments of the present disclosure, the weight matrix W(*) that achieves the optimal feature independence is element-wise multiplied with the classification loss cot(*) of the model to form the final training loss function; on the one hand, it retains the constraint of the traditional classification loss on the prediction accuracy, and on the other hand, it introduces a feature independence regularization term through the weight matrix, forcing the model to automatically learn a weight allocation strategy that can minimize the feature correlation while reducing the classification error, achieving the joint optimization of classification performance and feature quality.
[0143] Figure 5 is the structural diagram of a network architecture search synchronous transfer learning system for multi-modal atlas data according to an embodiment of the present application. As Figure 5 shown, the system includes:
[0144] A feature fusion module 201, which is used to extract and fuse features from the collected near-infrared spectra, Raman spectra, and microscopic images;
[0145] An SSL_NST module 202, which is used to divide the fused dataset into a source domain and a target domain through transfer learning technology. The neural network architecture search NAS adopts a three-stage optimization strategy including network search, fine-tuning, and feedback;
[0146] A loss stability weighting module 203, which is used to introduce a loss weighting optimization strategy into the obtained target domain model to reduce the correlation between fused features, adjust the ratio between the classification loss value and the sample weight, and optimize the training process of the fused data model.
[0147] Preferably, the SSL_NST module 202 is further configured to perform the following operations: during the network search phase, replace the default search space of neural architecture search (NAS) with the combined architecture space of CNN and RNN, and use NAS to search for a source domain model on source domain data; during the fine-tuning phase, use the semi-supervised self-training algorithm to assign pseudo-labels to target domain data based on the source domain model, and use the maximum entropy sampling in active learning to screen out the pseudo-labeled data. Freeze the network structure of the source domain model, and use the screened pseudo-labeled data to fine-tune the source domain model. By gradually unfreezing network layers, adjusting the optimizer and learning rate, minimize the target domain loss; during the feedback phase, use the loss of the fine-tuned model on the target domain task as a feedback signal to act on the search strategy of network search and guide the next round of search process; after a preset number of iterations, the searched network structure is used as the target domain model for the target domain task.
[0148] It should be noted that the above Figure 5 each module can be a program module (for example, a set of program instructions that implement a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the above each module is presented as a processor, or the functions of the above each module are implemented by a processor.
[0149] It should be noted that Figure 5 the preferred implementation manners of the illustrated embodiments can be referred to Figure 1 the relevant descriptions of the illustrated embodiments, which will not be elaborated here.
[0150] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.
Claims
1. A network architecture search synchronous transfer learning method for multi-modal graph data, characterized in that It includes the following steps: Extract features and fuse the collected near-infrared spectra, Raman spectra, and microscopic images; In the constructed SSL_NST module, through transfer learning, the fused dataset is divided into a source domain and a target domain. The neural network architecture search NAS adopts a three-stage optimization strategy including network search, fine-tuning, and feedback. Among them: In the network search stage, replace the default search space of the neural network architecture search NAS with a combined architecture space of CNN and RNN, and use the neural network architecture search NAS to search for the source domain model on the source domain data; in the fine-tuning stage, based on the source domain model, use the semi-supervised self-training algorithm to assign pseudo-labels to the target domain data, and use the maximum entropy sampling in active learning to screen out the pseudo-labeled data. Freeze the network structure of the source domain model, and use the screened pseudo-labeled data to fine-tune the source domain model. By operating on thawing network layers layer by layer, adjusting the optimizer and learning rate, minimize the target domain loss; In the feedback stage, use the loss of the fine-tuned model on the target domain task as a feedback signal to act on the search strategy of the network search and guide the next round of search process; after a preset number of iterations, the searched network structure is used as the target domain model for the target domain task.
2. The network architecture search synchronous transfer learning method for multi-modal atlas data according to claim 1, characterized in that The steps of extracting features and fusing the collected near-infrared spectra, Raman spectra, and microscopic images include: Normalize the obtained near-infrared spectra, Raman spectra, and microscopic images using Min-Max normalization and preprocess them using multiplicative scatter correction MSC; Use the self-attention mechanism to extract 128-dimensional optimal features F-nir and F-raman from the preprocessed near-infrared spectra and Raman spectra respectively; Form a 256-dimensional fusion feature F-nir-raman by concatenating the 128-dimensional optimal features of the near-infrared spectra and Raman spectra; Adjust the image size of the preprocessed microscopic images to 224×224 pixels and use the self-attention mechanism to extract 256-dimensional optimal features F-micro image; Use the cross-attention mechanism to fuse the 256-dimensional fusion feature F-nir-raman and the 256-dimensional optimal feature F-micro image to form a 512-dimensional optimal fusion feature F-fusion.
3. The network architecture search synchronous transfer learning method for multi-modal atlas data according to claim 2, characterized in that, The steps of using the self-attention mechanism to extract 256-dimensional optimal features F-micro image include: Set the input feature sequence X ∈ R n×d , when extracting 128-dimensional features from near-infrared spectroscopy and Raman spectroscopy, first pass through the learnable weight matrices W Q , W K , W V ∈ R d×128 , map the input feature X to the 128-dimensional query vector Q = XW Q , key vector K = XW K and value vector V = XW V respectively; then calculate the dot product of the query vector and all key vectors and divide by the square root of the dimension for scaling, the scaling factor is , and then generate the attention weight matrix through the Softmax function, which characterizes the correlation degree of the features at each position in the input sequence; finally, multiply the attention weights by the value vectors to obtain the 128-dimensional self-attention output features that fuse the internal dependencies of the sequence ; When extracting 256-dimensional features from microscopic images, map the input features to query, key, and value vectors of 256 dimensions; adjust the dimensionality scaling factor to to generate self-attention output features of 256 dimensions ; 128-dimensional self-attention output features and 256-dimensional self-attention output features are respectively represented as: ; ; Wherein, X represents an input feature, and W V represents a weight matrix of the value vector V, and W Q represents a weight matrix of the query vector Q, and W K represents a weight matrix of the key vector K, and T represents the transpose of a matrix.
4. The network architecture search synchronous transfer learning method for multi-modal atlas data according to claim 3, characterized in that The steps of using the cross-attention mechanism to fuse features to form a 512-dimensional optimal fusion feature F-fusion include: Settings is the 256-dimensional fusion feature after splicing near-infrared spectroscopy and Raman spectroscopy, is the 256-dimensional feature extracted from microscopic images. First, through the learnable weight matrix , map the spectral fusion feature F3 to the query vector , map the image feature F4 to the key vector and the value vector ; Then calculate the dot product of the spectral query vector and the image key vector and scale it, and generate cross-modal attention weights through Softmax; Multiply the attention weights by the image value vector to obtain the response features of the image features to the spectral features; Through the Concat operation, concatenate the original spectral fusion feature F3 and the cross-modal response features along the feature dimension to generate a 512-dimensional cross-modal fusion feature F5. The 512-dimensional fusion feature F5 output by the cross-attention mechanism can be expressed as: ; Wherein, F3 represents the spectral fusion feature, represents the learnable weight matrix of the query vector, F4 represents the image feature, represents the learnable weight matrix of the key vector, represents the learnable weight matrix of the value vector, and T represents the transpose of the matrix.
5. The network architecture search synchronous transfer learning method for multi-modal atlas data according to claim 4, characterized in that The step of assigning pseudo-labels to the target domain data using the semi-supervised self-training algorithm and screening out the pseudo-labeled data using the maximum entropy sampling in active learning includes: Set the fused dataset as , where x i represents the samples of the fused dataset, y i represents the labels corresponding to the samples, and n is the number of labeled samples; The source domain dataset is , where m is the number of source domain samples, and x j represents the samples in the source domain dataset; The neural network model generated by the NAS architecture is used as the classification model S. This model fully learns the data distribution P(x, y) of the source domain data and uses the confidence threshold mechanism in the self-training algorithm to screen the target domain samples. First, the model C is used to classify the target domain unlabeled sample set D. u For each sample x in j Make predictions and get prediction results and its corresponding prediction confidence(x j ); Then set a confidence threshold τ, and compare the prediction confidence of each sample with the threshold to filter out samples with confidence higher than the threshold. The filtered target domain sample set D t From all the confidence(x j )>τ conditional sample x j Composition, the selected target domain sample set D t It is expressed as: ; where x j represents a sample, D u represents a sample set, and τ represents a confidence threshold; Then, calculate the information entropy of the screened target domain samples. During the calculation of the information entropy, first calculate the conditional probability p(y c ∣x) that the target domain sample x belongs to each category c, then take the natural logarithm of the conditional probability of each category, multiply it by the conditional probability, and take the negative value to obtain the uncertainty contribution of a single category; finally, sum up the uncertainty contributions corresponding to all categories to obtain the information entropy H(x) of the sample x. Among them, the information entropy of the target domain sample x j is expressed as: ; where \(c\) represents the class of the target domain sample \(x\), \(c = 1, 2,\cdots,C\), and \(C\) is the total number of classes; \(p(y c |x)\) represents the conditional probability, that is, the posterior probability that the sample \(x\) belongs to class \(c\) given the features; In each iteration t, combined with the idea of maximum entropy sampling based on uncertainty in active learning, sample the target domain data set D t to obtain the required sample subset D t = S(D v ) through the sampling function S(D t ), and submit these samples to the source domain model for annotation and assign them pseudo-labels.
6. The network architecture search synchronous transfer learning method for multimodal atlas data according to claim 5, characterized in that The step of searching for the source domain model on the source domain data using the neural network architecture search NAS includes: During the entire search process, the first stage of the search models the prediction loss of candidate network architectures on the source domain data D through a Gaussian process, expressed as: u as follows: ; Among them, u(·) represents the mean of the Gaussian process prediction loss, σ(·) represents the standard deviation of the Gaussian process prediction loss, β represents the balance factor, and D u represents the source domain dataset, and Cost(f, D u ) represents the prediction loss of the network architecture f on the source domain data D u ; Adjusting the degree of attention to the model performance mean and uncertainty through the balance factor β; Generate candidate architectures in the first stage Based on this, by minimizing the candidate architectures on the target domain data D t the prediction loss to screen out the network architecture most suitable for the target domain The optimization objective is expressed as: ; In the formula, represents the candidate architecture the prediction loss on the target domain data D t and represents the candidate architecture generated in the first stage, represents the selected network architecture, and argmin represents the function used to calculate the minimum value.
7. The network architecture search synchronous transfer learning method for multi-modal atlas data according to claim 6, characterized in that When fine-tuning the model, select the optimal network structure f from the set of candidate network structures to minimize the prediction loss on the target domain test set. The process of minimizing the target domain loss is expressed as: ; Among them, represents the prediction loss, is the optimal parameter obtained by the network structure f through pre-training or learning in the previous stage, F CR represents the set of candidate network structures, D t-test represents the target domain test set; During the fine-tuning process, the Adam optimizer is used as the optimization algorithm, the activation function is set to the Sigmoid function, and the loss function is the binary cross-entropy function.
8. The network architecture search synchronous transfer learning method for multi-modal atlas data according to claim 7, characterized in that, It also includes: Introducing a loss weighting optimization strategy in the obtained target domain model to reduce the correlation between the fusion features, adjusting the ratio between the classification loss value and the sample weight, and optimizing the training process of the fusion data model; Where: Using the random Fourier features RFF to map the features to a high-dimensional feature space; During the weight learning process, by concatenating the global feature sets (F G1 , F G2 ,..., F Gi , F L ) with the current batch of local features F L , a comprehensive feature F0 for the optimization of the next batch is formed; Similarly, the global weight set (W G1 , W G2 ,..., W Gi , W L ) is concatenated with the local weight W of the current batch L to obtain the comprehensive weight W0. The calculation formulas for the feature F0 and the weight W0 used to optimize the weights of the next batch of samples in each batch are as follows: ; ; Among them, the global feature F Gi and the global weight W Gi represent the global information accumulated in historical batches, remaining unchanged during the training process of each batch. Only the local feature F L and the local weight W L can be trained and updated under the condition of minimizing the feature dependency relationship; At the end of each iteration, through the smoothing parameter weighted fusion is performed on the global information (F Gi , W Gi ) and the local information of the current batch (F L , W L ) to obtain the global information of the next round, which is expressed as follows: ; ; Wherein, is used to adjust the memory length of the global information, represents the global feature after weighted fusion, represents the global weight after weighted fusion, F L and W L respectively represent the local feature and the local weight of the current batch, represents the smoothing parameter, F Gi and W Gi respectively represent the global feature and the global weight; Minimize the Frobenius norm of the cross-covariance matrix between weighted sample features, which is the square of the Frobenius norm of the covariance matrix of the weighted features The expression is as follows: ; Among them, represents the weighted mutual information index between feature groups A and B, and is used to measure the correlation strength between the two; is the square of the Frobenius norm; Weighted covariance matrix is expressed as follows: ; Among them, represents the covariance between a certain feature dimension of A and a certain feature dimension of B for each element; n represents the total number of samples, A i and B i respectively represent the relevant features and irrelevant features of sample i, T represents the transpose of a matrix or vector, which is used to calculate the correlation between features; u(·) and v(·) are random Fourier feature (RFF) mapping functions that map the original features to a high-dimensional space; w i is the sample weight, satisfying the normalization condition; and are weighted mean vectors; Performing an element-wise product of the weight matrix W(*) that achieves the optimal feature independence and the classification loss cot(*) of the model to form the final training loss function, expressed as: Loss = cot(*) ⊙ W(*), where ⊙ represents the element-wise product.
Citation Information
Patent Citations
Image recognition method and device and computer readable storage medium
CN113076963A
Image classification network generation method and device, electronic equipment and storage medium
CN118570602A
Neural architecture search method, neural architecture search device and readable recording medium
CN119323239A
Method and system for determining objects depicted in images
US20190095764A1
Hybrid and Hierarchical Multi-Trial and OneShot Neural Architecture Search on Datacenter Machine Learning Accelerators
US20230297580A1
Cited By
Cross-working-condition multivariable time sequence anomaly detection method based on stage perception migration diffusion
CN121980476A
Cross-condition multivariate time series anomaly detection method based on phase-aware migration diffusion
CN121980476B