Identification model training method, data identification method and electronic equipment
By combining a dual-branch neural network and a meta-learning optimizer, the problems of feature distribution differences and coupling in cross-scenario tasks are solved, achieving higher recognition accuracy and adaptability.
Patent Information
- Application Number
- CN202510722168.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing adaptive algorithms have reduced recognition accuracy in cross-scenario tasks due to feature distribution differences and coupling problems, making them difficult to effectively apply in new scenarios.
A dual-branch neural network structure is used to extract scene-related and irrelevant features, and is combined with a meta-learning optimizer for training. Feature decoupling is optimized through reconstruction, adversarial and contrast loss errors, and internal and external loop training is used to improve model adaptability.
The recognition accuracy and adaptability of the model in new scenarios are improved, and the recognition ability of cross-scenario tasks is enhanced.
Smart Images

Figure CN120632449A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a recognition model training method, a data recognition method, and an electronic device. Background Art
[0002] Artificial intelligence can achieve functions such as classification or recognition, but as application scenarios become increasingly complex and diverse, existing adaptive algorithms have exposed many problems in cross-scenario tasks. For example, since the features to be recognized are distributed in different scenarios, recognition is difficult, which leads to a decrease in recognition accuracy. Summary of the Invention
[0003] The purpose of this application is to provide a recognition model training method, a data recognition method and an electronic device, which can improve the recognition accuracy of the trained model.
[0004] In the first aspect, the present invention provides a recognition model training method, comprising: inputting a training data set into a scene feature decoupling module for a first training; wherein the scene feature decoupling module comprises a first branch neural network unit and a second branch neural network unit, the first branch neural network unit being used to extract scene-related features in the training data, and the second branch neural network unit being used to extract scene-independent features in the training data; the training data in the training data set comprises one or more of image data and speech data; calculating a first loss error for the first training; inputting the scene-independent features extracted by the scene feature decoupling module into a meta-learning optimizer for a second training; calculating a second loss error for the second training; correcting parameters of the scene feature decoupling module and the meta-learning optimizer based on the first loss error and the second loss; repeating the above process until the first loss error and the second loss error are less than a set value, or the number of training times reaches a set number, to obtain a target recognition model; wherein the target recognition model is used to identify data to be identified, and the data to be identified comprises one or more of image data and speech data.
[0005] In the above implementation, the final recognition model consists of two parts. One part is used by the scene feature decoupling module to identify the scene features, so that the part of the data used to represent the object to be identified can be extracted more accurately. Furthermore, the scene-independent features can be input into the meta-learning optimizer, so that recognition based on scene-independent features can obtain more accurate recognition results.
[0006] In an optional embodiment, the first loss error includes a reconstruction loss error, and the scene feature decoupling module also includes a decoding unit, which is used to reconstruct scene-independent features and scene-related features to obtain reconstructed training data; the calculation of the first loss error of the first training includes: calculating the reconstruction loss error based on the training data and the reconstructed training data.
[0007] In the above implementation method, the scene feature decoupling module may also include a decoding unit, which realizes reconstruction of the extracted features based on the decoding unit, further determines the reconstruction loss error, and implements scene feature decoupling module training based on the reconstruction loss error. This allows the trained scene feature decoupling module to better realize effective decoupling of scene features while extracting scene-irrelevant features.
[0008] In an optional embodiment, the first loss error includes an adversarial loss error; the scene feature decoupling module also includes a discrimination module, the input data of the discrimination module is the reconstructed training data, and the discrimination module is used to identify whether the reconstructed training data is a scene-related feature; the calculation of the first loss error of the first training includes: calculating the adversarial loss error based on the output result of the discrimination module.
[0009] In the above implementation method, when training the scene feature decoupling module, the adversarial loss error can also be combined to realize the correction of the parameters of the scene feature decoupling module, thereby improving the decoupling effectiveness of the trained scene feature decoupling module.
[0010] In an optional embodiment, the first loss error includes a contrast loss error; the calculation of the first loss error of the first training includes: screening out one scene-related feature as a positive sample and taking other scene-related features as negative samples from the training results obtained from the first training of each training data in the same scene; and calculating the contrast loss error based on the positive sample and the negative sample.
[0011] In the above implementation method, when training the scene feature decoupling module, the contrast loss error can also be combined to realize the correction of the parameters of the scene feature decoupling module, thereby improving the decoupling effectiveness of the trained scene feature decoupling module.
[0012] In an optional embodiment, the first branch neural network unit includes multiple layers of convolution layers, pooling layers, normalization layers, and noise perception layers; the second branch neural network unit includes multiple layers of convolution layers and pooling layers.
[0013] In the above implementation, the two branch units can use different structures to achieve different processing focuses. For example, the noise perception layer simulates the distribution characteristics of different noises, allowing the convolution kernel to learn the characteristic patterns of noise. In this way, the feature vector ultimately output by this branch can accurately reflect relevant factors such as lighting and noise in the scene, effectively extracting specific scene attributes.
[0014] In an optional embodiment, the scene-independent features extracted by the scene feature decoupling module are input into the meta-learning optimizer for a second training, comprising: using a first scene-independent feature set to perform inner-loop training on the meta-learning optimizer; and using a second scene-independent feature set to perform outer-loop training on the meta-learning optimizer; wherein, the first scene-independent feature set and the second scene-independent feature set are scene-independent feature sets for different scenarios, and the number of scene-independent features contained in the first scene-independent feature set is smaller than the number of scene-independent features in the second scene-independent feature set.
[0015] In the above implementation, inner-loop training and outer-loop training can be combined. Inner-loop training focuses on practicing specific tasks in the meta-learning optimizer. The outer-loop training aims to evaluate model performance on multiple different tasks and optimize meta-parameters, enabling the model to adapt more quickly to new tasks.
[0016] In an optional embodiment, before inputting the training data set into the scene feature decoupling module for the first training, the method further includes: obtaining original training data, wherein the original training data includes features to be identified; performing transformation processing on the original training data to obtain training data; wherein the transformation processing includes one or more transformation processing of transforming the scene features in the original training data and adding noise to the original training data.
[0017] In the above implementation, the original image data may also be transformed to diversify the training data, thereby making the model trained based on the training data more adaptable to changes in the scene.
[0018] In the second aspect, the present invention provides a data recognition method, comprising: inputting the data to be recognized into a target recognition model for recognition to obtain a recognition result; wherein, the data to be recognized includes one or more of image data and voice data, and the target recognition model is trained using the recognition model training method described in any one of the aforementioned embodiments.
[0019] In a third aspect, the present invention provides an electronic device comprising: a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the machine-readable instructions are executed by the processor to perform the steps of the method described in any one of the aforementioned embodiments.
[0020] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which executes the steps of the method described in any one of the aforementioned embodiments when the computer program is executed by a processor.
[0021] In a fifth aspect, the present invention provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method described in any one of the aforementioned embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 A block diagram of an electronic device provided in an embodiment of the present application;
[0024] Figure 2 A flowchart of the recognition model training method provided in an embodiment of the present application;
[0025] Figure 3 A flowchart of a data identification method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application.
[0027] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0028] In the fields of artificial intelligence and machine learning, the performance of adaptive algorithms plays a key role in determining system performance in different scenarios. As AI application scenarios become increasingly complex and diverse, existing adaptive algorithms have exposed numerous issues in cross-scenario tasks such as image and speech recognition. The inventors' research has revealed that the main challenges with adaptive algorithms in these cross-scenario tasks lie in the differences in data distribution between the target data and the training data, as well as in the coupling of features within the data.
[0029] For example, the distribution of new scene data that needs to be recognized is often significantly different from that of the training data. In practical applications, such as in the field of image recognition, different shooting environments of images (such as lighting conditions, shooting angles, background complexity, etc.) will cause the distribution of image data to change; for example, in speech recognition scenarios, speech data in different noise environments (such as street noise, indoor echo, etc.) and different speaker characteristics (such as accents, speaking speed, etc.) will also cause the distribution of speech data to be significantly different from the training data used during training. Traditional learning methods usually require a large amount of labeled data for retraining. However, obtaining a large amount of labeled data is not only costly, but may also be infeasible in some cases, which seriously limits the application efficiency and adaptability of the algorithm in new scenarios.
[0030] Feature coupling in data manifests as a high degree of coupling between scene-independent features (such as object shape and semantic category) and scene-dependent features (such as light intensity, noise level, and resolution). Taking image data as an example, when lighting conditions change, the appearance of objects changes accordingly, causing features originally used to identify objects to become intermingled with lighting-dependent features. During model training, this feature coupling makes it difficult for the model to distinguish between truly discriminative, cross-scene universal features and those applicable only to specific scenarios. Consequently, when the model encounters new scenarios, its inability to effectively utilize scene-independent features significantly impacts its generalization ability, making it difficult to accurately classify or recognize data in these new scenarios. Taking image data as an example, when the noise in the speech environment changes, the noise becomes superimposed and mixed with the speech content being recognized, making it difficult to accurately recognize data in these new noise environments.
[0031] Based on the above analysis, the embodiments of the present application can provide a recognition model training method, data recognition method, and electronic device that can improve the adaptability of the trained model in new scenarios, thereby better realizing data recognition. The recognition model training method and data recognition method provided by the present application are described below in conjunction with multiple embodiments.
[0032] The following first explains some concepts involved in the method provided in the embodiment of the present application:
[0033] Accuracy: In classification tasks, accuracy is one of the most commonly used evaluation metrics. It indicates the proportion of samples correctly predicted by the model to the total number of samples. For example, in an image classification task, if the model correctly classifies 80 of 100 test images, the accuracy is 80%. Accuracy can intuitively reflect the model's classification performance in new scenarios.
[0034] F1-score: F1-score is an indicator that takes into account both precision and recall. It is more effective when dealing with imbalanced datasets. Precision represents the proportion of samples predicted as positive by the model that are actually positive, while recall represents the proportion of samples that are actually positive that are correctly predicted by the model. The formula for calculating F1-score is In some practical applications, such as disease diagnosis and anomaly detection, the dataset often has the problem of class imbalance. At this time, the F1-score can more comprehensively evaluate the performance of the model.
[0035] Mean Average Precision (mAP): mAP is a commonly used evaluation metric in object detection tasks. It is calculated by calculating the average precision (AP) across different categories and then averaging it. AP is the area between the precision and recall curves at different recall rates. mAP comprehensively reflects the detection performance of a model across multiple categories.
[0036] First, if Figure 1 , which is a block diagram of an electronic device. The electronic device 100 may include a memory 111 and a processor 113. A person skilled in the art will understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the electronic device 100. For example, the electronic device 100 may further include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0037] The memory 111 and processor 113 are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines. The processor 113 is used to execute the executable modules stored in the memory.
[0038] The memory 111 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory 111 is used to store programs, and the processor 113 executes the programs after receiving an execution instruction. The method executed by the electronic device 100 defined by the process disclosed in any embodiment of the present application can be applied to the processor 113 or implemented by the processor 113.
[0039] The above-mentioned processor 113 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 113 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. Optionally, the processor 113 may also be a graphics processing unit (GPU).
[0040] The electronic device 100 in this embodiment can be used to execute each step in each method provided in the embodiments of the present application. The following describes in detail the implementation process of the recognition model training method and the data recognition method through several embodiments.
[0041] See also Figure 2, is a flow chart of the recognition model training method provided by the embodiment of the present application. The recognition model training method provided by the embodiment of the present application can be applied to an electronic device, and the steps in the recognition model training method are performed by the electronic device. Figure 2 The specific process shown is described in detail.
[0042] Step 210: Input the training data set into the scene feature decoupling module for first training.
[0043] In this embodiment, the scene feature decoupling module can have a dual-branch structure, with each branch used to extract a class of features. The scene feature decoupling module can include a first-branch neural network unit and a second-branch neural network unit. The first-branch neural network unit is used to extract scene-related features from the training data, and the second-branch neural network unit is used to extract scene-independent features from the training data. The training data in the training dataset includes one or more of image data and speech data. This training data can be data collected in fields such as security, autonomous driving, and intelligent voice interaction.
[0044] Among them, the scene-related features can be features that are unrelated to the object to be identified. Taking the training data as image data as an example, the scene-related features can include scene attributes of the missing image, such as light intensity, noise level, etc., and also include background features, lighting features, shooting angle, color and other features in the image. Scene-independent features can be features related to the object to be identified in the image, such as the shape and texture of the object to be identified. Taking the training data as speech data as an example, the scene-related features can be features such as noise, tone, accent and other features in the speech data. Scene-independent features can be features related to the content to be identified in the speech.
[0045] Step 220: Calculate the first loss error of the first training.
[0046] Exemplarily, a pre-designed loss function may be used to calculate the first loss error.
[0047] Step 230: Input the scene-independent features extracted by the scene feature decoupling module into the meta-learning optimizer for second training.
[0048] Optionally, the meta-learning optimizer can be built based on a model-independent meta-learning (Model-AgnosticMeta-Learning, MAML for short) framework. In this embodiment, the model-independent meta-learning framework can be improved, and the scene-independent features can be processed by multi-scale feature fusion. In the image recognition task, convolution layers of different scales are set. For example, a small-scale convolution layer of 3×3 can be set to capture local detail features, and a large-scale convolution layer of 7×7 can be set to obtain global features. Then, the features extracted at different scales can be fused through features, so that the meta-learning optimizer can more comprehensively grasp the information of scene-independent features and enhance its adaptability to tasks in different scenes.
[0049] The output layer of the meta-learning optimizer is designed specifically for each task. For example, for classification tasks, the output layer can be designed as a softmax function that outputs the probability of each category. For regression tasks, the output layer can be designed as a linear layer that outputs continuous numerical results.
[0050] Step 240: Calculate the second loss error of the second training.
[0051] Exemplarily, a pre-designed loss function may be used to calculate the second loss error.
[0052] Step 250: Correct parameters of the scene feature decoupling module and the meta-learning optimizer based on the first loss error and the second loss.
[0053] Repeat the above process until the first loss error and the second loss error are less than the set value, or the number of training times reaches the set number, and the target recognition model is obtained.
[0054] In this embodiment, the meta-learning optimizer and the scene feature decoupling module do not operate independently, but work closely together. Figure 2 The order shown is for illustrative purposes only; the training of the scene feature decoupling module and the meta-learning optimizer can be performed simultaneously, that is, steps 210 and 230 can be run in parallel. The second loss error of the meta-learning optimizer can also be used to guide parameter adjustment of the scene feature decoupling module, and the scene-independent features output by the scene feature decoupling module can serve as input data for training the meta-learning optimizer.
[0055] During the joint training process, the scene feature decoupling module continuously optimizes the effect of feature decoupling, providing the meta-learning optimizer with purer scene-independent features. In the process of adapting to new scenes, the meta-learning optimizer will also feed back relevant information, such as the second loss error and which information causes the error, to the scene feature decoupling module, prompting it to further optimize the feature decoupling method of the scene feature decoupling module. For example, when the recognition accuracy of the meta-learner in a new scene is low, if analysis shows that the scene-independent features obtained may contain some scene interference information due to incomplete scene feature decoupling, the parameters of the scene feature decoupling module can be adjusted. For example, the weight of the contrast loss can be increased to enhance the feature decoupling effect of the scene feature decoupling module.
[0056] In addition, when the meta-learning optimizer updates the meta-parameters, it will also affect the training direction of the scene feature decoupling module. The two will continuously improve the overall performance of the algorithm through mutual collaboration and achieve more efficient cross-scene rapid adaptation.
[0057] During the joint training process, the adjustment of hyperparameters has a key impact on the training effect. Among them, the learning rate is one of the important hyperparameters, which is related to the scene feature decoupling module and the meta-learning optimizer. In this embodiment, a learning rate decay strategy can be adopted. For example, a cosine annealing learning rate scheduler can be used to set a large learning rate at the beginning of training to accelerate the model convergence. As training progresses, the learning rate is gradually reduced to prevent the model from skipping the optimal solution.
[0058] The target recognition model is used to recognize the data to be recognized, and the data to be recognized includes one or more of image data and voice data.
[0059] After joint training, the object recognition model can be used to predict data in new scenarios. In image classification tasks, the object recognition model can input the probability of an image belonging to each category and ultimately select the category with the highest probability as the prediction result; in speech recognition tasks, the object recognition model can output the recognized speech text.
[0060] Optionally, the object recognition model can also output the confidence level of the prediction, which is an assessment of the credibility of the prediction result. Confidence can be measured by calculating the entropy value of the predicted probability, with smaller entropy values indicating higher confidence levels. For example, in image classification, if the model predicts that the probability that an image belongs to the "cat" category is 0.9, then the prediction can be considered to have a high confidence level. If the predicted probabilities are more evenly distributed, such as the predicted probabilities for categories such as "cat," "dog," and "bird" are all around 0.3, then the prediction confidence level is lower.
[0061] In one usage scenario, the target recognition model can be used to identify the object to be found. The object can be a species such as a person or an animal. The training data can be image data containing features of the object to be found and scene features. The scene feature decoupling module described above can be used to identify scene-independent features and scene-dependent features. Scene-independent features can include the shape of an object; the object may be the object to be found, may be a suspected object to be found, or may be an object of the same type as the object to be found. Scene-dependent features can include the light intensity, noise level, resolution, etc. in the scene where the object is located.
[0062] In one use case, the object recognition model can be used to translate speech data, extract text, and more. The training data can be speech data. The aforementioned scene feature decoupling module can be used to identify scene-independent and scene-dependent features. Scene-independent features can be speech features; scene-dependent features can include noise levels in the speech data, such as white noise and ambient noise.
[0063] In an embodiment of the present application, in each round of iteration, the parameters of the scene feature decoupling module can be updated, and then the parameters of the meta-learning optimizer can be updated. When updating the parameters of the scene feature decoupling module, the first loss error can be calculated based on the current training data set, and the parameters of the scene feature decoupling module can be updated by the back propagation algorithm. When updating the parameters of the meta-learning optimizer, the parameters are updated according to the dual training of inner loop training and outer loop training. Through continuous alternating optimization, the scene feature decoupling module and the meta-learning optimizer can cooperate with each other and optimize together. As the training progresses, the model gradually learns more effective scene feature decoupling methods and rapid adaptation strategies, thereby improving performance in cross-scene tasks.
[0064] In some optional embodiments, the first loss error includes a reconstruction loss error, and the scene feature decoupling module further includes a decoding unit, which is used to reconstruct the scene-independent features and the scene-dependent features to obtain reconstructed training data.
[0065] Optionally, the decoding unit can also be constructed based on a neural network, for example, a deconvolutional neural network (DeCNN) can be used. Through the deconvolutional neural network, low-dimensional feature vectors can be gradually restored to high-dimensional vectors, so that the reconstructed data is as similar as possible to the original training data in terms of visual or auditory perception.
[0066] The role of the decoding unit is to convert the decoupled scene-independent feature z s and scene-related features z c Recombining and reconstructing the original training data, we get the reconstructed training data: x^=D(zs , z c ).
[0067] The above-mentioned step 220 may include: calculating a reconstruction loss error based on the training data and the reconstructed training data.
[0068] In order to ensure that the training data can be restored after feature decoupling, the reconstruction loss error L is defined recon The reconstruction loss error can be measured by calculating the difference between the original training data x and the reconstructed training data x^. Optionally, the reconstruction error loss can be measured using the mean square error (MSE), i.e., L recon =E x ||xD(E s (x), E c (x))|| 2 Among them, the scene-independent feature is represented by z s =E s (x), scene-related features are represented as z c =E c (x).
[0069] During the training process, by minimizing the reconstruction loss, the encoder formed by the first branch neural network unit and the second branch neural network unit and the decoder formed by the decoding unit are able to learn effective feature representation and feature reconstruction methods, so that the decoupled features can accurately restore the original training data.
[0070] In some optional embodiments, the above-mentioned first loss error may include an adversarial loss error; the scene feature decoupling module also includes a discrimination module, the input data of the discrimination module is the reconstructed training data, and the discrimination module is used to identify whether the reconstructed training data is a scene-related feature.
[0071] The above-mentioned step 220 may include: calculating the adversarial loss error based on the output result of the discriminant module.
[0072] In order to further optimize the effect of feature decoupling, the scene feature decoupling module introduces an adversarial training mechanism. Optionally, by adding the above-mentioned discriminant module, the discriminant module can constitute an adversarial network together with the two branches (the first branch neural network unit and the second branch neural network unit) and the decoding unit. Exemplarily, the discriminant module may include a multilayer perceptron (MLP) structure, the input of which may be a vector composed of scene-independent features and scene-related features, and the output is a judgment result of whether the combined vector is scene-related or irrelevant.
[0073] During the training process, the goal of the two branches (the first branch neural network unit and the second branch neural network unit) and the decoding unit is to minimize various losses. Among them, the adversarial loss is defined as the opposite of the cross-entropy loss of the discriminant module to correctly identify the feature type. The goal of the discriminant module is to maximize the adversarial loss. Through this continuous game process, the first branch neural network unit and the second branch neural network unit can gradually learn to generate features that are more difficult for the discriminant module to distinguish, thereby significantly improving the effect of feature decoupling, so that scene-independent features and scene-related features can be more clearly separated.
[0074] In some optional embodiments, the first loss error may include a contrastive loss error. Step 220 may include: selecting one scene-related feature from the training results obtained from the first training of each training data under the same scene as a positive sample and selecting the other scene-related features as negative samples; and calculating the contrastive loss error based on the positive sample and the negative sample.
[0075] In order to maximize the scene-related features z of different scenes c while minimizing the difference between z c The difference between the two can introduce the contrast loss error L contrast The contrast loss error can be calculated based on cosine similarity, and the formula of the contrast loss error can be designed as Where τ represents the temperature coefficient, which is used to adjust the distribution of similarity, and sim represents the cosine similarity function. When calculating the contrast loss error, the training data in the same batch can be divided into one or more scene groups. The training data in each scene group corresponds to different scenes. The scene-related features in each training data can be expressed as Scene-related features for selecting a sample in different scene groups As positive samples, select scene-related features of other samples in the same scene group By minimizing the contrast loss error, the scene feature decoupling module can better distinguish scene-related features of different scenes, thus enhancing the effect of feature decoupling.
[0076] In the calculation of the contrast loss error, in order to improve the calculation efficiency, the method of comparing the training data within the batch can be adopted. Each batch of training data can be grouped according to the scene label. In this embodiment, the temperature coefficient τ in the formula of the contrast loss error can be in an adjustable state during the training process, and it is not fixed. In the early stage of training, τ can be set to a larger value, such as 0.5. At this time, the similarity distribution is relatively flat, and the scene feature decoupling module can more easily learn the approximate relationship between features; as the training progresses, the value of τ is gradually reduced, such as to 0.1. This can enhance the discrimination between features, make the scene feature decoupling module more sensitive to the differences in related features of different scenes, further optimize the effect of contrast loss, and thus better achieve the decoupling of scene features.
[0077] Optionally, the reconstruction loss error and the contrast loss error can be combined to obtain a total loss, which can be expressed as the first loss error. The total loss error can be expressed as L disentangle =L recon +λL contrast (λ = 0.5), where λ = 0.5 is a balancing coefficient used to adjust the weights of the two loss errors in the total loss. During the training process, the total loss is minimized while optimizing the parameters in the scene feature decoupling module to achieve effective decoupling of scene features.
[0078] Optionally, the reconstruction loss error, the adversarial loss error, and the contrast loss error may be combined to obtain a total loss, which may be represented as the first loss error. The weights of the loss errors may be set as needed.
[0079] In this embodiment, the weights of each loss function, such as the weights of the reconstruction loss, the contrastive loss, and the adversarial loss, as well as the weights of each loss function in the meta-learning optimizer, are also important hyperparameters. Experimentation and cross-validation can be used to determine the appropriate combination of loss function weights.
[0080] In some optional embodiments, the first branch neural network unit may include multiple layers of convolutional layers, pooling layers, normalization layers, and noise perception layers.
[0081] Exemplarily, the first branch neural network unit for realizing scene-related feature extraction can be based on the basic structure of convolution layer, pooling layer and ReLU activation function, and can also add a light-sensitive layer and a noise-perception layer. The light-sensitive layer can utilize the brightness channel information of the image and enhance the sensitivity to lighting changes by designing special convolution kernels. For example, the convolution kernel in the light-sensitive layer can adjust the parameters according to the changing pattern of the image pixel value under different light intensities, so as to more accurately capture the lighting-related features. The noise-perception layer can simulate the distribution characteristics of different noises so that the convolution kernel in the noise-perception layer can learn the characteristic pattern of the noise. Based on this design, the feature vector output by the first branch neural network unit can accurately reflect the scene-related factors such as lighting and noise in the scene, and realize the effective extraction of scene-specific attributes.
[0082] Optionally, the second branch neural network unit includes multiple layers of convolutional layers and pooling layers. Exemplarily, for the second branch neural network unit extraction branch for extracting scene-independent features, it can be constructed by including multiple convolutional layers, pooling layers and ReLU activation functions. Optionally, the convolutional layer sorted in front can use a smaller convolution kernel, for example, with a size of 3×3, a step size of 1, and a padding of 1. Such a setting is conducive to capturing local details of the image, such as texture information of the object. The subsequent pooling layer selects the maximum pooling method, with a pooling kernel size of 2×2 and a step size of 2, which can effectively reduce the data dimension and reduce the amount of calculation. The addition of the ReLU activation function can enhance the nonlinear expression ability of the model, enabling the model to learn more complex feature relationships. By stacking multiple layers of such convolutional layers, pooling layers, and ReLU combinations, stable and representative scene-independent feature vectors can be output. These features can remain relatively stable in different scenarios, thereby achieving effective extraction of the essential features of the object.
[0083] In this embodiment, step 230 may be performed using a combination of inner-loop training and outer-loop training. Step 230 may include: performing inner-loop training on the meta-learning optimizer using the first scenario-independent feature set; and performing outer-loop training on the meta-learning optimizer using the second scenario-independent feature set.
[0084] The first scene-independent feature set and the second scene-independent feature set are scene-independent feature sets for different scenes, and the number of scene-independent features included in the first scene-independent feature set is smaller than that in the second scene-independent feature set.
[0085] Parameter update is performed on the first scene-independent feature set. Assuming that the initial parameters of the meta-learning optimizer are θ, the loss function Lτ is calculated on the first scene-independent feature set. i (f θ), and update the parameters according to the gradient descent method Where α is the learning rate of the inner loop training, with a default value of 0.01. Through the inner loop update, the meta-learning optimizer can quickly adjust parameters on the first scene-independent feature set to adapt to the scene characteristics of the specific task.
[0086] Update the meta-parameters on the second scene-independent feature set. Calculate the loss function Lτ on the second scene-independent feature set i (f θ′ ), and update the meta-parameters according to the gradient descent method Where β is the learning rate of the outer loop training, with a default value of 0.001. The purpose of the outer loop training update is to optimize the meta-parameter θ by evaluating the model performance on a second scene-independent feature set of multiple different tasks, so that the model can adapt faster when faced with new tasks.
[0087] Optionally, the inner loop training phase of the meta-learning optimizer can use an optimization algorithm with an adaptive learning rate. For example, it can be the Adam algorithm. The Adam algorithm can adaptively adjust the learning rate of each parameter based on the first-order moment estimate (mean) and second-order moment estimate (variance) of the gradient. In this implementation, for parameters with large changes, the learning rate can automatically decrease; for parameters with small changes, the learning rate can be increased accordingly, thereby accelerating the convergence speed of the meta-learning optimizer and improving the optimization effect.
[0088] Optionally, in order to prevent gradient explosion during inner loop training updates, a gradient clipping mechanism can be introduced to limit the maximum value of the gradient.
[0089] When updating the meta-parameters in the outer loop training, the update can be performed based on the loss of the second scene-independent feature set. Optionally, a regularization term is added. For example, L1 regularization or L2 regularization can be added. Taking L2 regularization as an example, adding The term λ is used as a regularization coefficient to constrain the size of the meta-parameters, thereby avoiding overfitting of the model.
[0090] Optionally, a momentum mechanism can be introduced to stabilize the meta-parameter updates of the meta-learning optimizer. This mechanism not only considers the current gradient but also incorporates previous gradient information when updating the meta-parameters, reducing gradient fluctuations and making the meta-parameter updates of the meta-learning optimizer smoother, preventing the optimal solution from being skipped during the update process.
[0091] In the above implementation, the training of the meta-learning optimizer can include inner-loop training and outer-loop training. During the inner-loop training, scene-independent features can be used for rapid parameter updates. To avoid the gradient explosion problem, an optimization algorithm with adaptive learning rate (such as Adam) is adopted, and a gradient clipping mechanism is introduced. Through multiple iterative updates, the model is enabled to quickly adapt to the task on the support set. The outer-loop training calculates the loss on the training set and updates the meta-parameters. In this process, in addition to considering the loss on the training set, regularization terms and momentum mechanisms are introduced to make the meta-parameter updates more stable and reliable. Moreover, the meta-learning optimizer will pass the feedback information during the training process to the scene feature decoupling module, thereby adjusting its parameters to achieve coordinated optimization of the two.
[0092] In order to better implement model training, the training data used for training can be determined in the following way.
[0093] Before executing the training process from step 210 to step 250 , the recognition model training method may further include: obtaining original training data, where the original training data includes features to be recognized; and performing transformation processing on the original training data to obtain training data.
[0094] The transformation processing includes one or more transformation processings of transforming scene features in the original training data and adding noise to the original training data.
[0095] Optionally, the transformation process may also include random cropping, flipping, rotation, etc. Taking the image classification task as an example, the original training data is randomly flipped horizontally to simulate object images under different perspectives.
[0096] Taking speech recognition as an example, sampling rate conversion and normalization are performed before extracting features such as MFCC and FBank. To simulate a realistic and complex speech environment, noise, reverberation, and other artifacts can also be added to the speech data. For example, in speech recognition tasks, varying degrees of white noise can be added to clean speech samples.
[0097] Optionally, the training data may be obtained by a stratified cluster sampling method.
[0098] Taking image and speech tasks as examples, we first stratify the scenes based on their key features. For example, in image tasks, we can stratify scenes based on factors such as light intensity and background complexity. The original training dataset can be divided into lighting strata such as strong light, low light, and normal light. We can also stratify the original training dataset into strata with different background complexity levels. In speech tasks, we can stratify scenes based on factors such as noise type and speaker accent. The original training dataset can be divided into noise strata such as white noise and ambient noise.
[0099] Furthermore, clustering algorithms can be used to group samples with similar features into one category for the original training data within a layer. In actual sampling tasks, samples can be evenly selected from different layers and scene categories as training data, allowing each task to cover a rich variety of scenes and feature information. For example, the training data obtained from each sampling of an image classification task can include multiple object categories under different lighting conditions and backgrounds of varying complexity; the training data obtained from sampling of a speech recognition task can include speech samples from speakers with different accents and in different noise environments, thereby improving the model's generalization capabilities in various scenarios.
[0100] In this embodiment, before executing the training process of steps 210 to 250, the training data may also be preprocessed. The preprocessing method is different for different types of data. For image data, the pixel values are first normalized to the interval [0, 1] or [-1, 1], which can speed up the convergence of the model. Exemplarily, the training data may also be text data, and the preprocessing may include word segmentation and stop word removal of the text data, and then using a pre-trained word vector model (such as Word2Vec, GloVe) through word embedding operations to convert the text into a vector form so that each word has a corresponding fixed-length vector representation.
[0101] In order to fully train and verify the training effect of the recognition model training method provided by this embodiment, training can be performed based on training data sets of multiple scenarios. For example, the training data set can include at least five different scenarios. For example, in an image classification task, it can include image data with different light intensities (such as strong light, weak light, indoor light, outdoor light, etc.), different noise levels (such as Gaussian noise, salt and pepper noise, etc.), and different resolutions (such as high resolution, low resolution).
[0102] In one example, a speech recognition task can include speech data from different noise environments (e.g., various noise types with a signal-to-noise ratio (SNR) of 0 to 20dB) and different speakers (e.g., speakers of different genders, ages, and accents). The training data for each scene can be no less than 1,000 pieces to ensure the model can learn the characteristic information of that scene; the test data can be no less than 200 pieces to evaluate the model's generalization ability in that scene. For example, in an image classification task, a training dataset with multiple illumination versions can be expanded, with image dimensions of 32×32×3 and a pixel value range of [0, 255]. By expanding the dataset under various illumination conditions, the model can be provided with a rich set of image data from different scenes. It should be understood that the above specific values are merely illustrative and can be adjusted appropriately to better suit the needs of the specific recognition task under varying circumstances.
[0103] In one example, in a speech recognition task, a data set with a multi-noise environment (SNR = 0 ~ 20dB) can be added, with a sampling rate of 16kHz and a speech duration of 1 to 5 seconds. In the above implementation method, the obtained training data set can simulate the situation in which the speech signal in the real scene is interfered with by different degrees of noise, which is helpful for training and evaluating the performance of the model in a complex speech environment. It can be understood that the above specific values are only for illustration. Under the changes in the actual task, the above values can be appropriately changed to better meet the needs of the specific recognition task.
[0104] After the training is completed, verification can be performed based on the trained target recognition model to confirm the possible effect of the obtained target recognition model.
[0105] The accuracy of the scene feature decoupling module in decoupling and reconstruction can be evaluated by calculating the mean squared error (MSE) between the original training data and the reconstructed training data. The smaller the reconstruction loss error, the better the scene feature decoupling module can restore the training data after feature decoupling, that is, the scene feature decoupling module can effectively extract and combine features. In image tasks, the reconstruction loss error can intuitively reflect the quality of image reconstruction; in speech tasks, the reconstruction error reflects the degree of restoration of the speech signal.
[0106] Optionally, mutual information (MI) can be used to measure the independence between scene-independent and scene-dependent features. Mutual information represents the degree of dependence between two random variables. A smaller mutual information value indicates a greater independence between the two features, i.e., a more thorough feature decoupling by the scene feature decoupling module. By calculating the mutual information between scene-independent and scene-dependent features, it is possible to verify whether the feature decoupling module has successfully separated the two features.
[0107] To verify the recognition performance of the object recognition model, you can select a validation dataset from the training data or re-determine a validation dataset using the same method as the training data. Ensure that the validation dataset covers a variety of scenarios. Input the validation dataset into the trained target model to obtain the model's prediction results. Based on the prediction results and the true labels, calculate metrics such as accuracy, F1 score, and mean average prediction (mAP) to determine the training performance of the object recognition model.
[0108] Based on the calculated indicator values, evaluate the model's performance in new scenarios. If each indicator reaches or exceeds the expected value, it indicates that the model has good adaptability and generalization ability.
[0109] During training, the model needs to be regularly evaluated and monitored to ensure performance and convergence. Metrics such as reconstruction loss, feature separation, accuracy, F1-score, and mAP are used to evaluate model performance.
[0110] Optionally, you can record information such as loss error and evaluation metrics during training and plot loss and metric curves. By observing the changing trends of these curves, you can determine whether the model has converged and whether overfitting or underfitting is occurring. The changes in these curves can help you identify potential problems in training and promptly adjust model hyperparameters or improve the model structure, thereby ensuring smooth training and good model performance.
[0111] See also Figure 3 , is a flow chart of the data identification method provided by the embodiment of the present application. The data identification method provided by the embodiment of the present application can be applied to an electronic device, and the steps in the data identification method are performed by the electronic device. Figure 3 The specific process shown is described in detail.
[0112] Step 310: input the data to be recognized into the target recognition model for recognition to obtain a recognition result.
[0113] The data to be identified includes one or more of image data and voice data.
[0114] The target recognition model can be trained using the aforementioned recognition model training method. For other details about the target recognition model, please refer to the description in the aforementioned embodiment and will not be repeated here.
[0115] In actual application scenarios, in order to enable the target recognition model to adapt to changes in new scenarios in real time, an online update mechanism can also be provided.
[0116] When the object recognition model encounters a small number of samples from a new scene, it first feeds these samples into the scene feature decoupling module to extract the corresponding scene-independent and scene-dependent features. These newly extracted features are then used to update the meta-learning optimizer. For example, based on the existing meta-learning optimizer, the small number of samples can also be divided into a first sample set and a second sample set, with parameter updates performed for the inner and outer loops, respectively.
[0117] In this embodiment, in order to avoid excessive modification of the original meta-parameters of the meta-learning optimizer, a smaller learning rate can be used during online updating.
[0118] To avoid frequent updates to the target recognition model, a first update threshold can be pre-set. Only when the difference between the new small number of samples and the training data used in the original training phase exceeds the first update threshold will the online update operation be performed. In this implementation, the stability of the target recognition model can be improved while also improving its adaptability.
[0119] In this embodiment, a second update threshold can also be set, and the online update operation is performed only when the number of new small samples is greater than the second update threshold. In this implementation, the stability of the target recognition model can be improved while also improving the adaptability of the target recognition model.
[0120] Compared with the prior art, the training method and recognition method provided in the embodiments of the present application effectively solve the problem of feature coupling by decoupling scene features, so that the model can better utilize features shared across scenes and improve the generalization ability of the model. Secondly, the meta-learning optimizer combines the idea of meta-learning and can quickly adapt to new scenes with only a small number of samples, greatly reducing the dependence on a large amount of labeled data and reducing the application cost. Finally, the joint training of the scene feature decoupling module and the meta-learning optimizer reduces the error accumulation and improves the overall performance and training efficiency of the model.
[0121] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the recognition model training method and the data recognition method in the above method embodiment are executed.
[0122] The computer program product of the recognition model training method or data recognition method provided in the embodiments of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the recognition model training method and data recognition method described in the above method embodiments. For details, please refer to the above method embodiments and will not be repeated here.
[0123] In the several embodiments provided in this application, it should be understood that the disclosed methods can also be implemented in other ways. The method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the methods and computer program products according to the multiple embodiments of the application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a portion of code, and the module, program segment, or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.
[0124] In addition, each method step in each embodiment of the present application can be integrated together to form an independent part for execution, or each method step can be executed by a separate module, or two or more steps can be formed into an independent part for execution.
[0125] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0126] The foregoing description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0127] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A recognition model training method, characterized in that: include: Inputting the training data set into the scene feature decoupling module for first training; wherein the scene feature decoupling module includes a first branch neural network unit and a second branch neural network unit, the first branch neural network unit is used to extract scene-related features in the training data, and the second branch neural network unit is used to extract scene-independent features in the training data; the training data in the training data set includes one or more of image data and speech data; Calculating a first loss error of the first training; Inputting the scene-independent features extracted by the scene feature decoupling module into the meta-learning optimizer for a second training; Calculating a second loss error of the second training; Correcting parameters of the scene feature decoupling module and the meta-learning optimizer based on the first loss error and the second loss; Repeat the above process until the first loss error and the second loss error are less than the set value, or the number of training times reaches the set number, to obtain a target recognition model; wherein, the target recognition model is used to identify the data to be identified, and the data to be identified includes one or more of image data and voice data.
2. The method according to claim 1, characterized in that The first loss error includes a reconstruction loss error, and the scene feature decoupling module further includes a decoding unit, wherein the decoding unit is used to reconstruct the scene-independent features and the scene-dependent features to obtain reconstructed training data; The calculating the first loss error of the first training includes: A reconstruction loss error is calculated based on the training data and the reconstructed training data.
3. The method according to claim 2, characterized in that The first loss error includes an adversarial loss error; the scene feature decoupling module further includes a discrimination module, the input data of the discrimination module is the reconstructed training data, and the discrimination module is used to identify whether the reconstructed training data is a scene-related feature; The calculating the first loss error of the first training includes: Based on the output result of the discrimination module, the adversarial loss error is calculated.
4. The method according to claim 1, wherein The first loss error includes a contrast loss error; The calculating the first loss error of the first training includes: In the training results obtained from the first training of each training data under the same scene, one scene-related feature is selected as a positive sample, and the other scene-related features are selected as negative samples; A contrast loss error is calculated based on the positive sample and the negative sample.
5. The method according to any one of claims 1 to 4, characterized in that The first branch neural network unit includes multiple layers of convolutional layers, pooling layers, normalization layers and noise perception layers; The second branch neural network unit includes multiple layers of convolutional layers and pooling layers.
6. The method according to claim 1, characterized in that Inputting the scene-independent features extracted by the scene feature decoupling module into the meta-learning optimizer for second training includes: performing inner-loop training on the meta-learning optimizer using the first scene-independent feature set; The meta-learning optimizer is trained in an outer loop using a second scene-independent feature set; wherein the first scene-independent feature set and the second scene-independent feature set are scene-independent feature sets for different scenarios, and the number of scene-independent features included in the first scene-independent feature set is smaller than the number of scene-independent features in the second scene-independent feature set.
7. The method according to claim 1, characterized in that Before inputting the training data set into the scene feature decoupling module for the first training, the method further includes: Obtaining original training data, wherein the original training data includes features to be identified; The original training data is transformed to obtain training data; wherein the transformation includes one or more transformations of transforming scene features in the original training data and adding noise to the original training data.
8. A data identification method, characterized in that: include: The data to be recognized is input into the target recognition model for recognition to obtain a recognition result; wherein, the data to be recognized includes one or more of image data and voice data, and the target recognition model is trained using the recognition model training method described in any one of claims 1-7.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the machine-readable instructions are executed by the processor to perform the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method according to any one of claims 1 to 7.
11. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Scene recognition method and device, computer equipment and storage medium
CN113033507A
Model training method, scene recognition method and related equipment
CN115187824A
Target domain information search method and model training method and device
CN115827732A
Model training method, scene recognition method, and related device
US20240169687A1