Classification model training method and device for multi-heterogeneous model fusion and classification method and device
Through the classification model training method of multiheterogeneous model fusion, the image classification model is fused and distilled using DS evidence theory and generative adversarial network, which solves the problems of high cost of sample data acquisition and labeling and uneven data distribution in the existing technology, and improves the prediction performance and accuracy of the image classification model.
Patent Information
- Application Number
- CN202510011216.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In the prior art, sample data acquisition and labeling are high costs and data distribution is uneven, resulting in low prediction performance of the trained image classification model.
The classification model training method of multi-heterogeneous model fusion is adopted, and the output representations of multiple teacher models are fused through DS evidence theory to obtain the fused output representation, and the probability distribution vectors of each category are determined based on it. Then, the probability distribution vector guidance is used to generate adversarial network GAN for training, and the student model is distilled and trained to obtain the target student model to perform the image classification task.
It improves the prediction performance of the image classification model, enhances the accuracy of image classification results, and solves the problems of high cost of sample data acquisition and labeling and uneven data distribution.
Smart Images

Figure CN119992248A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a classification model training method, a classification method and a device for fusing multiple heterogeneous models. Background Art
[0002] In recent years, with the development and maturity of artificial intelligence (AI) technology, its application in the field of computer vision (CV) has become more and more extensive; in particular, the development of deep learning (DL) methods in the field of image recognition and classification has achieved gratifying results.
[0003] Image recognition and classification is one of the core tasks in the field of computer vision. The model needs to identify specific objects or scenes from an input image and then classify it into one of the predefined categories. With the rise of artificial intelligence technology, deep learning methods such as convolutional neural networks (CNN) and Transformer models have greatly improved the performance of image recognition and classification.
[0004] In related technologies, image classification models require a large amount of effective sample data for training. However, due to the high cost of sample data acquisition and annotation and the problem of uneven data distribution, the trained models often show obvious performance degradation when applied to actual scenarios, resulting in reduced accuracy of image reasoning results such as image recognition, image classification and image segmentation. Summary of the invention
[0005] The present invention provides a classification model training method, a classification method and a device for fusing multiple heterogeneous models, which are used to solve the defects in the prior art that the cost of sample data acquisition and annotation is high, and the prediction performance of the trained image classification model is low due to uneven data distribution, thereby improving the accuracy of image classification.
[0006] The present invention provides a classification model training method for fusion of multiple heterogeneous models, comprising: Based on the DS evidence theory, the output representations of multiple teacher models are fused to obtain a fused output representation, and the probability distribution vectors of each category in the category set corresponding to the multiple teacher models are determined according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; Based on the probability distribution vector, the generative adversarial network GAN is trained to obtain the trained GAN, and the student model is distilled and trained according to the probability distribution vector and the trained GAN to obtain the target student model to perform the image classification task.
[0007] According to a classification model training method for multi-heterogeneous model fusion provided by the present invention, the training of a generative adversarial network GAN is guided based on the probability distribution vector to obtain the trained GAN, which includes: Based on the GAN, the probability distribution vector is mapped to the image space, and the GAN is iteratively trained according to a first joint loss function, and the trained GAN is obtained when the GAN converges or reaches a maximum number of iterations; wherein the first joint loss function is determined based on the probability value of the category with the largest probability after fusion, the random sampling vector determined based on the probability space after fusion, and the unlabeled image sample.
[0008] According to a classification model training method for multi-heterogeneous model fusion provided by the present invention, the student model is distilled and trained according to the probability distribution vector and the trained GAN to obtain a target student model, which includes: Determine a KL divergence loss function based on the KL divergence between the output representation of the student model and the probability distribution vector; determine a distance loss function based on the distance between the output representation of the generator in the trained GAN and the mapping representation; wherein the mapping representation is obtained by mapping the output representation of the generator to the image space based on the generator; A second joint loss function is determined based on the KL divergence loss function and the distance loss function, and the student model is iteratively trained according to the second joint loss function, and the target student model is obtained when the student model converges or reaches a maximum number of iterations.
[0009] According to a classification model training method for multi-heterogeneous model fusion provided by the present invention, the KL divergence loss function is expressed by the following formula: ; in, is the KL divergence loss, is an image generated by the trained GAN according to the image after sampling noise, is the weighting coefficient, is the number of samples for a single training, is the output representation after fusion, is the output representation of the student model; The distance loss function is expressed by the following formula: ; in, is the distance loss function, is the temperature coefficient, G For the generator, It is a paradigm operation.
[0010] According to a classification model training method for multi-heterogeneous model fusion provided by the present invention, determining the probability distribution vector of each category in the category set corresponding to the multiple teacher models according to the fused output representation includes: The probability distribution vector is determined based on the fused output representation and the number of categories in the union of multiple category sets; the data length of the probability distribution vector is the number of categories in the union plus one.
[0011] The present invention also provides a classification method, comprising: Obtain image data to be classified; The image data to be classified is processed based on the target student model to obtain a target classification result; wherein the target student model is trained based on the classification model training method of multi-heterogeneous model fusion.
[0012] The present invention also provides a classification model training device for fusion of multiple heterogeneous models, comprising: A model fusion module is used to fuse the output representations of multiple teacher models based on the DS evidence theory to obtain a fused output representation, and determine the probability distribution vectors of each category in the category set corresponding to the multiple teacher models according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; The model distillation training module is used to guide the training of the generative adversarial network GAN based on the probability distribution vector to obtain the trained GAN, and to perform distillation training on the student model according to the probability distribution vector and the trained GAN to obtain a target student model to perform an image classification task.
[0013] The present invention also provides a classification device, comprising: An image acquisition module, used to acquire image data to be classified; A classification module is used to process the image data to be classified based on a target student model to obtain a target classification result; wherein the target student model is trained based on the classification model training method of the multi-heterogeneous model fusion.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements a classification model training method or classification method for fusing multiple heterogeneous models as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a classification model training method or classification method for fusion of multiple heterogeneous models as described in any one of the above.
[0016] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements a classification model training method or a classification method for fusing multiple heterogeneous models as described above.
[0017] The classification model training method, classification method and device for multi-heterogeneous model fusion provided by the present invention fuse the output representations of multiple teacher models through the DS evidence theory to obtain the fused output representation, and determine the probability distribution vectors of each category in the category set corresponding to the multiple teacher models based on the fused output representation, and then use the probability distribution vector to guide the training of the generative adversarial network GAN, and perform distillation training on the student model based on the probability distribution vector and the trained GAN to obtain the target student model, thereby realizing the fusion and distillation of the image classification model using the DS evidence theory and the generative model, improving the prediction performance of the image classification model, and thereby improving the accuracy of the image classification results. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 It is one of the flow charts of the classification model training method for multi-heterogeneous model fusion provided by the present invention.
[0020] Figure 2 This is the second flow chart of the classification model training method for multi-heterogeneous model fusion provided by the present invention.
[0021] Figure 3 It is a schematic diagram of the mapping relationship of the GAN network provided by the present invention.
[0022] Figure 4 It is a flowchart of the GAN network training method provided by the present invention.
[0023] Figure 5 This is the third flow chart of the classification model training method for multi-heterogeneous model fusion provided by the present invention.
[0024] Figure 6 It is a flow chart of the classification method provided by the present invention.
[0025] Figure 7 It is a structural schematic diagram of a classification model training device for fusion of multiple heterogeneous models provided by the present invention.
[0026] Figure 8 It is a schematic diagram of the structure of the classification device provided by the present invention. Fig. 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] Combine the following Figure 1-Figure 8 The present invention describes the classification model training method, classification method and device for fusion of multiple heterogeneous models.
[0029] Figure 1 This is one of the flow charts of the classification model training method for multi-heterogeneous model fusion provided by the present invention, such as Figure 1 As shown, the classification model training method for multi-heterogeneous model fusion includes the following steps: Step 110: Based on the DS evidence theory, the output representations of multiple teacher models are fused to obtain a fused output representation, and the probability distribution vectors of each category in the category set corresponding to the multiple teacher models are determined according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks.
[0030] It should be noted that the teacher model can be defined as a fused model. The student model acquires the task capabilities of each teacher through the knowledge of each teacher model, and at the same time achieves the ability to generalize to actual application scenarios, making its task performance in application scenarios better than that of the teacher model.
[0031] In this step, a model that has been trained to recognize and classify images is used as the teacher model, such as a large commercial model hosted by a cloud server or other open source models; however, due to its training data, internal structure and composition, connections between layers, parameters of the teacher model and gradients used for back propagation are not visible.
[0032] For example, teacher models include but are not limited to Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory Networks (LSTM) or Convolutional Recurrent Neural Networks (CRNN).
[0033] In this embodiment, when selecting a teacher model, there is no restriction on the differences in the internal structure of the model. Multiple models with different tasks (category sets) can be selected as teacher models. These models can provide valuable probability outputs for actual application scenarios. For example, The category set of contains some categories of actual application scenarios, then the model As a teacher model, ensure that the teacher model has good accuracy and robustness.
[0034] In this embodiment, the multiple teacher models may be deep learning models with exactly the same model architecture, for example, the multiple teacher models are all convolutional neural networks.
[0035] In this embodiment, the two teacher models may be deep learning models with different model architectures. For example, the two teacher models may be a convolutional neural network and a recurrent neural network, respectively.
[0036] In this embodiment, when performing image recognition or image segmentation tasks, the corresponding teacher model also includes a trained image recognition model or a trained image segmentation model.
[0037] In this embodiment, the results of image reasoning by multiple teacher models are used as input to train the final student model. The reasoning results are expressed as follows: ; in, is a normalized probability distribution, including out-of-domain probabilities, represented by a vector with a length equal to the number of categories plus 1, representing the input image The corresponding probability of each class and the probability of non-class set, Represents the teacher model's response to the input image reasoning process.
[0038] It should be noted that when executing the classification model training method of multi-heterogeneous model fusion, it is not necessary to load the teacher model into the local. , then the subsequent steps can be carried out.
[0039] In this embodiment, after selecting the teacher model, the output of the teacher model is first integrated to obtain the correct probability representation; for example, the probability of the teacher output is integrated using the DS evidence theory. The specific integration process is as follows: Assume that A teacher model, For the The category set of the model. The fusion formula is as follows: ; ; in, is the normalization coefficient, which is used to ensure the normalized category probability. ; For the The category set of the model, Representative of non The categories included in the model are Indicated by and The power set of the set, is the confidence or probability assigned to class C after fusion, express an element of and They are reciprocals of each other.
[0040] According to the actual image classification scenario, the above formula is abbreviated as follows: ; ; ; in, , express The number of categories contained in; through the above formula, we will eventually get an incomplete recognition framework The basic probability distribution of The basic probability distribution of (elements not in the set are assigned the value zero); h A conditional variable.
[0041] Finally, after the above calculation process, the fused probability distribution is obtained, that is, the probability of each category in the category set and out-of-domain probability .
[0042] In this embodiment, a length equal to the number of categories plus Vector Represents the probability distribution after fusion, where the number of categories is the number of categories in the union of all teacher model category sets.
[0043] Specifically, determining the probability distribution vectors of each category in the category set corresponding to multiple teacher models according to the fused output representation includes: determining the probability distribution vector based on the fused output representation and the number of categories in the union of multiple category sets; the data length of the probability distribution vector is the number of categories in the union plus one.
[0044] In this embodiment, the vector With the above probability and out-of-domain probability The relationship between them is expressed by the following formula: ; in, is the intra-class probability vector and out-of-class probability The concatenated vector is is the number of categories after fusion.
[0045] Step 120: Based on the probability distribution vector, the generative adversarial network GAN is trained to obtain the trained GAN, and the student model is distilled and trained according to the probability distribution vector and the trained GAN to obtain a target student model to perform the image classification task.
[0046] In this step, in view of the differences between actual application scenarios and teacher training scenarios, as well as the difficulty in collecting high-quality labeled data in actual application scenarios, Generative Adversarial Networks (GAN) are used to expand data and generalize the domain.
[0047] In this embodiment, by applying the scene picture and the probability distribution vector The generative adversarial network (GAN) is trained to achieve unlabeled data expansion and deprivatization of the teacher model to achieve better generalization effects for actual application scenarios.
[0048] In this embodiment, the student model includes but is not limited to one of a convolutional neural network, a recurrent neural network, a long short-term memory network, or a convolutional recurrent neural network.
[0049] In this embodiment, after obtaining the trained GAN, the student model is distilled and trained using the trained GAN network and the probability distribution vector to obtain a target student model suitable for actual application scenarios; then, the target student model is tested using a small amount of labeled data to verify the model prediction accuracy.
[0050] Specifically, after the test image is input into the trained target student model, the model will output the corresponding probability distribution, select the category with the highest probability as the classification answer, and compare it with the labeled answer to statistically calculate the accuracy of the target student model.
[0051] Figure 2 This is the second flow chart of the classification model training method for multi-heterogeneous model fusion provided by the present invention. Figure 2 In the illustrated embodiment, the classification model training method of multi-heterogeneous model fusion is also implemented through the following steps: S1, obtaining a teacher model, S2: integrating teacher output, training a generated model, and training a student model; S3: evaluating the student model.
[0052] The present invention provides a classification model training method for multi-heterogeneous model fusion, which fuses the output representations of multiple teacher models through the DS evidence theory to obtain a fused output representation, and determines the probability distribution vectors of each category in the category set corresponding to the multiple teacher models based on the fused output representation, and then uses the probability distribution vector to guide the training of the generative adversarial network GAN, and performs distillation training on the student model based on the probability distribution vector and the trained GAN to obtain a target student model, thereby realizing the fusion and distillation of the image classification model using the DS evidence theory and the generative model, improving the prediction performance of the image classification model, and thereby improving the accuracy of the image classification results.
[0053] In some embodiments, training of a generative adversarial network (GAN) is guided based on a probability distribution vector to obtain a trained GAN, including: mapping the probability distribution vector to an image space based on the GAN, and iteratively training the GAN according to a first joint loss function, to obtain the trained GAN when the GAN converges or reaches a maximum number of iterations; wherein the first joint loss function is determined based on the probability value of the category with the highest probability after fusion, a random sampling vector determined based on the probability space after fusion, and an unlabeled image sample.
[0054] In this embodiment, random sampling is performed from the probability distribution vector to map the fused probability distribution back to the image space, so as to train the GAN network to generate corresponding application scenario images, thereby ensuring that the images generated by the GAN network are more consistent with the classification results of the fused probability, so as to improve the number and quality of samples.
[0055] In this embodiment, during the training process, the probability distribution vector can be obtained based on the output of the pre-trained teacher model; the loss function corresponding to the pre-trained teacher model is expressed by the following formula: ; in, , is the probability value of the category with the highest probability after fusion; is a random sampling vector determined based on the fused probability space, The length and same; is an unlabeled image sample for actual application scenarios. and Respectively represent the discriminator and generator of the trained GAN; For the generator according to Generated pseudo image.
[0056] The loss function of the GAN discriminator is expressed as follows: ; in, For the discriminator Determined judgment result.
[0057] The loss function of the GAN generator is expressed as follows: ; In this embodiment, the first joint loss function L It is expressed by the following formula: ; in, , and They are all weight factors not greater than 1, and the values of the weight factors can be set according to user needs.
[0058] In this embodiment, the GAN is iteratively trained using the first joint loss function. In each iteration, the parameters of the generator and the discriminator are updated until the GAN converges or reaches the maximum number of iterations, thereby obtaining the trained GAN.
[0059] Figure 3 is a schematic diagram of the mapping relationship of the GAN network provided by the present invention. Figure 3 In the embodiment shown, They represent GAN, the fused teacher model and the student model respectively. The fused teacher model and the student model calculate the distribution probability of the category to which the image belongs through the input image, that is, the mapping from the image space to the category set probability space is realized through the fused teacher model and the student model. The GAN training is guided according to the determined distribution probability to generate more realistic images, that is, the mapping from the category set probability space to the image space is realized through GAN.
[0060] Figure 4 is a flow chart of the GAN network training method provided by the present invention. Figure 4 In the embodiment shown, a pseudo image is generated by the generator G according to the input image Z, and the pseudo image is discriminated by the discriminator D to achieve sample expansion. In this process, the output representations of multiple teacher models (T1, ...Tn) are fused through the DS evidence theory to obtain a probability distribution vector, and then the GAN is iteratively trained according to the DS guidance to obtain the trained GAN, where the loss function of each teacher model is L T , the loss function of GAN is L GAN .
[0061] The present invention provides a classification model training method for multi-heterogeneous model fusion, which maps a probability distribution vector to an image space through a GAN, and iteratively trains the GAN according to a first joint loss function to obtain a trained GAN, thereby improving the robustness and accuracy of GAN-generated image samples, thereby improving sample quality and generation efficiency.
[0062] In some embodiments, performing distillation training on the student model according to the probability distribution vector and the trained GAN to obtain the target student model includes: (1) A KL divergence loss function is determined based on the KL divergence between the output representation of the student model and the probability distribution vector; a distance loss function is determined based on the distance between the output representation of the generator in the trained GAN and the mapping representation; wherein the mapping representation is obtained by mapping the output representation of the generator to the image space based on the generator.
[0063] In this embodiment, the KL divergence loss function is expressed by the following formula: ; in, is the KL divergence loss, is the image generated by the trained GAN based on the image after sampling noise, is the weighting coefficient, is the number of samples for a single training, is the output representation after fusion, is the output representation of the student model.
[0064] The distance loss function is expressed as follows: ; in, is the distance loss function, is the temperature coefficient, G For the generator, It is a paradigm operation.
[0065] (2) Determine a second joint loss function based on the KL divergence loss function and the distance loss function, and iteratively train the student model according to the second joint loss function. When the student model converges or reaches the maximum number of iterations, the target student model is obtained.
[0066] In this embodiment, the second joint loss function L 2 is represented by the following formula: ; in, and They are all weight factors not greater than 1, and the values of the weight factors can be set according to user needs.
[0067] In this embodiment, the student model is iteratively trained using the second joint loss function. In each iteration, the network parameters of the student model are updated until the student model converges or reaches the maximum number of iterations, thereby obtaining a target student model.
[0068] Figure 5 This is the third flow chart of the classification model training method for multi-heterogeneous model fusion provided by the present invention. Figure 5 In the embodiment shown, the output representations of multiple teacher models (T1, ...Tn) are fused through the DS evidence theory to obtain a probability distribution vector (used to represent the probability of each category), and the KL divergence is calculated in combination with the output representation of the student model S to determine the KL divergence loss function L KL ; Use the generator G to inversely map the output representation to the image space and use the distance function of the two as the distance loss function L F ; Finally use L KL and L F The second joint loss is determined to complete the iterative training of the student model.
[0069] The present invention provides a classification model training method for multi-heterogeneous model fusion, which determines the KL divergence loss function through the KL divergence between the output representation of the student model and the probability distribution vector; determines the distance loss function through the distance between the output representation of the generator and the mapping representation, and then obtains the second joint loss function and iteratively trains the student model to obtain the target student model, thereby improving the prediction performance of the image classification model and further improving the image classification accuracy.
[0070] The classification method provided by the present invention is described below. The classification method described below and the classification model training method for fusion of multiple heterogeneous models described above can be referenced to each other.
[0071] Figure 6 It is a schematic diagram of the process of the classification method provided by the present invention, such as Figure 6 As shown, the method includes the following: Step 610: Obtain image data to be classified.
[0072] In this step, the image data to be classified may be images captured in real time, or may be existing images obtained from an image database.
[0073] Step 620: Process the image data to be classified based on the target student model to obtain a target classification result; wherein the target student model is trained based on a classification model training method that integrates multiple heterogeneous models.
[0074] In this step, the target student model is trained through the following steps: (1) Based on the DS evidence theory, the output representations of multiple teacher models are fused to obtain a fused output representation, and the probability distribution vectors of each category in the category set corresponding to the multiple teacher models are determined based on the fused output representation; the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; (2) Based on the probability distribution vector, the generative adversarial network (GAN) is trained to obtain the trained GAN, and the student model is distilled and trained according to the probability distribution vector and the trained GAN to obtain the target student model to perform the image classification task.
[0075] It should be noted that the implementation method of each training step of the target student model corresponds one-to-one to the above-mentioned steps 110 to 120, and will not be repeated in this embodiment.
[0076] In this embodiment, the target student model obtained through the above training processes the image data to be classified, and can obtain a more accurate classification result to meet the image classification task requirements in different scenarios.
[0077] The classification method provided by the embodiment of the present invention processes the image data to be classified by the target student model trained by the classification model training method of multi-heterogeneous model fusion, thereby improving the image classification accuracy in a variety of complex scenarios.
[0078] The classification model training device for multi-heterogeneous model fusion provided by the present invention is described below. The classification model training device for multi-heterogeneous model fusion described below and the classification model training method for multi-heterogeneous model fusion described above can be referenced to each other.
[0079] Figure 7 : is a schematic diagram of the structure of the classification model training device for multi-heterogeneous model fusion provided by the present invention, such as Figure 7 As shown, the device includes: a model fusion module 710 and a model distillation training module 720.
[0080] The model fusion module 710 is used to fuse the output representations of multiple teacher models based on the DS evidence theory to obtain a fused output representation, and determine the probability distribution vectors of each category in the category set corresponding to the multiple teacher models according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; The model distillation training module 720 is used to guide the training of the generative adversarial network GAN based on the probability distribution vector to obtain the trained GAN, and to perform distillation training on the student model according to the probability distribution vector and the trained GAN to obtain the target student model to perform the image classification task.
[0081] The present invention provides a classification model training device for multi-heterogeneous model fusion, which fuses the output representations of multiple teacher models through the DS evidence theory to obtain a fused output representation, and determines the probability distribution vectors of each category in the category set corresponding to the multiple teacher models based on the fused output representation, and then uses the probability distribution vector to guide the training of the generative adversarial network GAN, and performs distillation training on the student model based on the probability distribution vector and the trained GAN to obtain the target student model, thereby realizing the fusion and distillation of the image classification model using the DS evidence theory and the generative model, improving the prediction performance of the image classification model, and thereby improving the accuracy of the image classification results.
[0082] The classification device provided by the present invention is described below. The classification device described below and the classification method described above can be referenced to each other.
[0083] Figure 8 It is a structural schematic diagram of the classification device provided by the present invention, such as Figure 8 As shown, the device includes: an image acquisition module 810 and a classification module 820.
[0084] An image acquisition module 810 is used to acquire image data to be classified; The classification module 820 is used to process the image data to be classified based on the target student model to obtain a target classification result; wherein the target student model is trained based on the classification model training method of multi-heterogeneous model fusion as claimed in any one of claims 1-5.
[0085] The classification device provided by the embodiment of the present invention processes the image data to be classified by the target student model trained by the classification model training method of multi-heterogeneous model fusion, thereby improving the image classification accuracy in a variety of complex scenarios.
[0086] Fig. 9 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Fig. 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930 and a communication bus 940, wherein the processor 910, the communication interface 920 and the memory 930 communicate with each other through the communication bus 940. The processor 910 may call the logic instructions in the memory 930 to execute a classification model training method for fusion of multiple heterogeneous models, the method comprising: fusing the output representations of multiple teacher models based on the DS evidence theory to obtain the fused output representation, and determining the probability distribution vectors of each category in the category set corresponding to the multiple teacher models according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; based on the probability distribution vector, the generative adversarial network GAN is trained to obtain the trained GAN, and the student model is distilled and trained according to the probability distribution vector and the trained GAN to obtain the target student model to perform the image classification task.
[0087] Or execute a classification method, the method comprising: obtaining image data to be classified; processing the image data to be classified based on a target student model to obtain a target classification result; wherein the target student model is trained based on a classification model training method of multi-heterogeneous model fusion.
[0088] In addition, the logic instructions in the above-mentioned memory 930 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0089] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the classification model training method of multi-heterogeneous model fusion provided by the above-mentioned methods, and the method includes: fusing the output representations of multiple teacher models based on the DS evidence theory to obtain a fused output representation, and determining the probability distribution vectors of each category in the category set corresponding to the multiple teacher models according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; based on the probability distribution vector, the generative adversarial network GAN is trained to obtain the trained GAN, and the student model is distilled and trained according to the probability distribution vector and the trained GAN to obtain the target student model to perform the image classification task.
[0090] Or execute a classification method, the method comprising: obtaining image data to be classified; processing the image data to be classified based on a target student model to obtain a target classification result; wherein the target student model is trained based on a classification model training method of multi-heterogeneous model fusion.
[0091] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the classification model training method for the fusion of multiple heterogeneous models provided by the above-mentioned methods, the method comprising: fusing the output representations of multiple teacher models based on the DS evidence theory to obtain a fused output representation, and determining the probability distribution vectors of each category in the category set corresponding to the multiple teacher models according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; training a generative adversarial network GAN is guided based on the probability distribution vector to obtain the trained GAN, and distilling the student model according to the probability distribution vector and the trained GAN to obtain a target student model to perform image classification tasks.
[0092] Or execute a classification method, the method comprising: obtaining image data to be classified; processing the image data to be classified based on a target student model to obtain a target classification result; wherein the target student model is trained based on a classification model training method of multi-heterogeneous model fusion.
[0093] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0094] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A classification model training method for multi-heterogeneous model fusion, characterized in that: include: Based on the DS evidence theory, the output representations of multiple teacher models are fused to obtain a fused output representation, and the probability distribution vectors of each category in the category set corresponding to the multiple teacher models are determined according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; Based on the probability distribution vector, the generative adversarial network GAN is trained to obtain the trained GAN, and the student model is distilled and trained according to the probability distribution vector and the trained GAN to obtain the target student model to perform the image classification task.
2. The classification model training method for multi-heterogeneous model fusion according to claim 1 is characterized in that: The training of the generative adversarial network GAN is guided based on the probability distribution vector to obtain the trained GAN, which includes: Based on the GAN, the probability distribution vector is mapped to the image space, and the GAN is iteratively trained according to a first joint loss function, and the trained GAN is obtained when the GAN converges or reaches a maximum number of iterations; wherein the first joint loss function is determined based on the probability value of the category with the largest probability after fusion, the random sampling vector determined based on the probability space after fusion, and the unlabeled image sample.
3. The classification model training method for multi-heterogeneous model fusion according to claim 1 is characterized in that: Performing distillation training on the student model according to the probability distribution vector and the trained GAN to obtain a target student model includes: Determine a KL divergence loss function based on the KL divergence between the output representation of the student model and the probability distribution vector; determine a distance loss function based on the distance between the output representation of the generator in the trained GAN and the mapping representation; wherein the mapping representation is obtained by mapping the output representation of the generator to the image space based on the generator; A second joint loss function is determined based on the KL divergence loss function and the distance loss function, and the student model is iteratively trained according to the second joint loss function, and the target student model is obtained when the student model converges or reaches a maximum number of iterations.
4. The classification model training method for multi-heterogeneous model fusion according to claim 3 is characterized in that: The KL divergence loss function is expressed by the following formula: ; in, is the KL divergence loss, is an image generated by the trained GAN according to the image after sampling noise, is the weighting coefficient, is the number of samples for a single training, is the output representation after fusion, is the output representation of the student model; The distance loss function is expressed by the following formula: ; in, is the distance loss function, is the temperature coefficient, G For the generator, It is a paradigm operation.
5. The classification model training method for multi-heterogeneous model fusion according to claim 1 is characterized in that: Determining the probability distribution vectors of each category in the category set corresponding to the multiple teacher models according to the fused output representation includes: The probability distribution vector is determined based on the fused output representation and the number of categories in the union of multiple category sets; the data length of the probability distribution vector is the number of categories in the union plus one.
6. A classification method, characterized in that: include: Obtain image data to be classified; The image data to be classified is processed based on a target student model to obtain a target classification result; wherein the target student model is trained based on the classification model training method for fusion of multiple heterogeneous models as described in any one of claims 1 to 5.
7. A classification model training device for multi-heterogeneous model fusion, characterized in that: include: A model fusion module is used to fuse the output representations of multiple teacher models based on the DS evidence theory to obtain a fused output representation, and determine the probability distribution vectors of each category in the category set corresponding to the multiple teacher models according to the fused output representation; wherein the fused output representation is used to represent the probability and out-of-domain probability of each category in the category set; different teacher models are used to perform different image classification tasks; The model distillation training module is used to guide the training of the generative adversarial network GAN based on the probability distribution vector to obtain the trained GAN, and to perform distillation training on the student model according to the probability distribution vector and the trained GAN to obtain a target student model to perform an image classification task.
8. A classification device, characterized in that: include: An image acquisition module, used to acquire image data to be classified; A classification module is used to process the image data to be classified based on a target student model to obtain a target classification result; wherein the target student model is trained based on the classification model training method for fusion of multiple heterogeneous models as described in any one of claims 1 to 5.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image classification method and device based on model distillation, storage medium and terminal
CN113408571A
Double-layer evidence fusion learning method, classification evaluation method and device
CN116992344A