Dynamic zero sample image recognition method and system based on multi-modal fusion

Through the dynamic zero-sample image recognition method of multimodal fusion, the generator generates no-series visual features. Combined with the discriminator and embedded space representation function, it solves the problem of data imbalance and insufficient embedded space discriminant ability in zero-sample image recognition, and improves the discriminant ability and robustness of the model.

CN120561871AInactive Publication Date: 2025-08-29JIANGSU VARIABLE SUPERCOMP TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511054998.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing zero-sample image recognition method has a high error rate of unclassified categories and difficult to effectively solve the domain deviation problem when the data distribution is unbalanced and the embedding space discrimination ability is insufficient.

Method used

Through the dynamic zero-sample image recognition method of multimodal fusion, a generator is used to generate unseen visual features of categories, combined with the discriminator, embedded spatial representation function and learning comparator network, the multi-channel classification sub-problem algorithm and the minimized cross-entropy loss algorithm are used to enhance the discriminant ability and robustness of the model.

Benefits of technology

Effectively reduce the error detection rate, improve the generalization ability of model, enhance the distinction ability of embedded space, alleviate data imbalance and category deviation, and improve the identification performance of unknown categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561871A_ABST
    Figure CN120561871A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion dynamic zero sample image recognition method and system in the technical field of image recognition, and the method comprises the steps: inputting first data into a generator, employing a first fusion algorithm to drive the generator to generate unseen category visual features, and enabling the first data to comprise Gaussian noise, semantic descriptors and text descriptions; and generating a sample if no category visual feature is seen. According to the method, the feature generation and embedding model and the supervision information of the category level and the instance level are fused, and the generator and the embedding model are alternately optimized, so that the distinguishing capability of the embedding space is improved on the premise of ensuring the quality of the generated image, and the image quality is improved. And meanwhile, by combining multi-source information, cooperating with a feature generation and embedding model and introducing a dynamic adjustment mechanism, the problems of unbalanced training data of a known category and a non-known category, category deviation and insufficient embedding space discrimination capability in zero sample learning in the prior art can be effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a multimodal fusion dynamic zero-sample image recognition method and system. Background Art

[0002] Because some classes in many datasets contain only a few images, deep learning networks struggle to learn the features they need to extract, hindering classification. Chinese patent application number CN202411429974.1 discloses a "zero-shot image recognition model and training method, image recognition method, program product, medium, and device." This method employs an indirect domain adaptation approach, mapping both the generated and real features of an image into an embedding space. This achieves distribution registration between the generated and real features, and uses this registered distribution to train a classifier.

[0003] Although existing methods can solve the problem of unbalanced data distribution in zero-shot image recognition to a certain extent, since most existing semantic embedding methods focus on the supervision information at the category level, the original embedding space lacks sufficient discriminative ability and fails to fully mine the distinguishing information in the data, resulting in a high error rate in the classification of unclassified categories, which is not conducive to effectively solving the problem of effective distinction of domain deviation. Summary of the Invention

[0004] The purpose of the present invention is to provide a multimodal fusion dynamic zero-shot image recognition method and system, which can enhance the model's discrimination ability, effectively classify different categories, reduce the false detection rate, improve the model's generalization ability, and improve the model's robustness by bringing samples of the same category closer and samples of different categories farther apart, thereby helping to solve the problem of domain deviation that is difficult to effectively distinguish.

[0005] To achieve the above object, the present invention provides the following technical solutions: In a first aspect, the present invention provides a multimodal fusion dynamic zero-shot image recognition method, characterized by comprising: Input the first data into the generator G, and use the first fusion algorithm to drive the generator G to generate visual features of unseen categories , wherein the first data includes Gaussian noise , semantic descriptors , text description , the visual features of unseen categories are used to generate samples; Collect real samples , and use the discriminator D to distinguish the real samples and generate samples , and generate the first loss function; Collect visual features x as the input of the embedding mapping function, and output the embedding space representation function; Get instance samples , based on example samples , for each instance sample Construct a multi-way classification sub-problem algorithm to obtain instance samples The representation result and the second loss function in the embedding space are calculated using a learning comparator network F(h,a) to calculate the representation result and the semantic descriptor The third loss function is obtained by using the third loss function to obtain the category-level supervision information of the instance sample; Using a classifier, adjusting parameters of the classifier based on a cross entropy loss minimization algorithm to obtain a target classifier, and using the target classifier to classify features of the input sample; Based on the first loss function, the second loss function and the third loss function, the total loss data of the input sample is obtained, the target loss is obtained by minimizing the total loss, and the target classification result of the input sample is obtained based on the target loss.

[0006] As a further solution of the present invention: the algorithm for the multi-way classification sub-problem is Algorithm for the road classification subproblem.

[0007] As a further solution of the present invention: the first fusion algorithm is: ,in, is the visual feature of the unseen category, is Gaussian noise, is a semantic descriptor, is the text description, G is the fusion process function of the generator; The first fusion algorithm is expanded as follows: ,in, 、 、 Gaussian noise , semantic descriptors and text description The nonlinear transformation of 、 、 Gaussian noise , semantic descriptors and text description The weight coefficients learned through the attention mechanism, is a non-linear activation function.

[0008] As a further solution of the present invention: the discriminator D is used to distinguish the real samples and generate samples ,include: The calculation formula of the discriminator is: Distinguishing real samples and generate samples , and obtain the discrimination result, where represents the visual features of the input, Represents the discrimination result output by the discriminator.

[0009] As a further solution of the present invention: generating the first loss function includes using a discriminator to generate the first loss function, wherein the first loss function is an adversarial loss function; The adversarial loss function is: ,in, is the expected operation function, is the real data distribution, is the distribution function of input noise and semantic descriptor.

[0010] As a further solution of the present invention: the embedding space representation function is: , where E(x) is the embedding function. By using the embedding function to construct the embedding space representation function, high-dimensional visual features can be mapped to a low-dimensional embedding space while retaining the key information and category discriminant features of the sample.

[0011] As a further solution of the present invention: the acquisition of example samples , based on example samples , for each instance sample structure Road classification sub-problem algorithm, get instance samples The representation result in the embedding space and the second loss function include: Based on nonlinear projection head , get the nonlinear projection head The loss function is: , where T is the transposition symbol; For the The embedded features of samples after nonlinear projection; is with Positive samples belonging to the same category are embedded; Is from other categories negative sample embeddings; is the number of negative samples; is the temperature parameter, which is used to adjust the sharpness of the similarity distribution in contrastive learning; And the second loss function: Calculate the instance sample The representation result in the embedding space is, For all positive sample pairs; For the generator; is the embedded function; For the projection head, the second loss function is the contrastive embedding loss function.

[0012] As a further solution of the present invention: the initial loss function of the learning comparator network F(h,a) is: in, For example samples The corresponding positive semantic expression, is the semantic descriptor of all seen categories, is the temperature parameter for class-level contrast embedding, and >0, is the number of seen categories.

[0013] As a further solution of the present invention: based on the initial loss function, a third loss function is obtained: ,in, For the generator, is the embedding function, is a semantic comparison network, For the embedding vectors, For the correct semantic representation, the third loss function is to optimize the category-level contrast embedding loss function.

[0014] As a further solution of the present invention: the adjusting of the parameters of the classifier based on the cross entropy loss minimization algorithm includes: Use the minimized cross entropy loss function: Optimize the first parameter W and the second parameter b of the classifier, where N is the number of training samples, is the total number of categories, which includes the number of seen categories and the number of unseen categories. By optimizing the first parameter W and the second parameter b of the classifier, the target classifier can better learn the mapping relationship between the embedding space representation and the category label, thereby improving the classification accuracy.

[0015] In a second aspect, the present invention provides a system for use in the multimodal fusion dynamic zero-sample image recognition method described in the above scheme, wherein the system comprises an unseen category visual feature generation module, a sample collection and differentiation module, a category-level supervision module, and a feature classification module; the unseen category visual feature generation module is configured to input the first data into a generator G, and use a first fusion algorithm to drive the generator G to generate unseen category visual features. , wherein the first data includes Gaussian noise , semantic descriptors , text description , the visual features of the unseen categories are used to generate samples; the sample collection and differentiation module is configured to collect real samples , and use the discriminator D to distinguish the real samples and generate samples , and generate the first loss function, collect visual features x as the input of the embedding mapping function, and output the embedding space representation function; the category level supervision module is configured to obtain instance samples , based on example samples , for each instance sample structure Road classification sub-problem algorithm, get instance samples The representation result and the second loss function in the embedding space are calculated using a learning comparator network F(h,a) to calculate the representation result and the semantic descriptor correlation, obtain a third loss function, and adopt the third loss function to obtain the category-level supervision information of the instance sample; the feature classification module is configured to adopt a classifier, adjust the parameters of the classifier based on the minimization cross entropy loss algorithm, obtain a target classifier, adopt the target classifier to classify the features of the input sample, based on the first loss function, the second loss function and the third loss function, obtain the total loss data of the input sample, minimize the total loss to obtain the target loss, and obtain the target classification result of the input sample based on the target loss.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. In the present invention, by fusing feature generation and embedding models, as well as supervision information at the category level and instance level, and by alternately optimizing the generator and embedding model, the distinguishing ability of the embedding space is improved while ensuring the quality of the generated image. At the same time, by combining multi-source information, coordinating feature generation and embedding models, and introducing a dynamic adjustment mechanism, the problems of imbalanced training data between seen and unseen categories, category bias, and insufficient discriminative ability of the embedding space in zero-shot learning in the prior art can be effectively alleviated.

[0017] 2. In the present invention, by minimizing the total loss, the model can more accurately map real samples and synthetic samples into the embedding space and perform effective classification in the space; enhance the model's ability to distinguish different categories in the embedding space. By setting a comparator network and a category-level loss function, the model's semantic alignment and category discrimination capabilities in the embedding space can be effectively enhanced. In the case of a lack of samples in unseen categories, the feature distribution of the unseen categories can be learned solely by relying on semantic descriptions, thereby effectively alleviating the data imbalance problem and improving the recognition performance of unseen categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a step diagram of the method of the present invention; Figure 2 It is a flowchart of the method of the present invention; Figure 3 This is a system diagram of the present invention.

[0019] In the figure: 1. Unseen category visual feature generation module; 2. Sample collection and differentiation module; 3. Category-level supervision module; 4. Feature classification module. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0021] Example: See also Figure 1-Figure 3 In an embodiment of the present invention, a multimodal fusion dynamic zero-shot image recognition method is provided, comprising the following steps: S1: Input the first data into the generator G, and use the first fusion algorithm to drive the generator G to generate the visual features of the unseen category , wherein the first data includes Gaussian noise , semantic descriptors , text description , the visual features of unseen categories are used to generate samples; S2: Collect real samples , and use the discriminator D to distinguish the real samples and generate samples , and generate the first loss function; S3: Collect visual features x as the input of the embedding mapping function, and output the embedding space representation function; S4: Get instance samples , based on example samples , for each instance sample Construct a multi-way classification sub-problem algorithm to obtain instance samples The representation results in the embedding space and the second loss function are calculated using the learning comparator network F(h,a) to calculate the representation results and semantic descriptors The third loss function is obtained by using the third loss function to obtain the category-level supervision information of the instance sample; S5: Using a classifier, adjusting the parameters of the classifier based on the cross entropy loss minimization algorithm to obtain a target classifier, and using the target classifier to classify the features of the input sample; S6: Based on the first loss function, the second loss function and the third loss function, the total loss data of the input sample is obtained, the target loss is obtained by minimizing the total loss, and the target classification result of the input sample is obtained based on the target loss.

[0022] In this embodiment, when the dynamic zero-sample image recognition process of multimodal fusion is started, the model and the optimizer are initialized, wherein the model is a model composed of the first fusion algorithm, and the optimizer is a generator. By inputting the first data containing Gaussian noise, semantic descriptors, and text descriptions into the first fusion algorithm, the generator is driven to generate visual features of unseen categories, and the collected real samples and generated samples are loaded as data sets. The real samples and generated samples are distinguished through discriminator cyclic training, and the discriminator is updated. The discriminator is used to distinguish the real samples from the generated samples and the first loss function, the second loss function, and the third loss function are used in turn. , it can further calculate the loss of real samples, the loss of generated samples, the contrast loss and classification loss, and optimize the discriminator based on the loss calculation results, and judge whether the number of iterations of the discriminator reaches the expected number. If not, the discriminator is updated, and the new first loss function, second loss function and third loss function are generated. If so, the generator is updated, and the contrast loss and classification loss are calculated to judge whether the number of training times meets the expectations. If not, the loop training step is returned. If it meets the expectations, the model is evaluated, the classifier is trained after generating synthetic features, and the classifier performance is evaluated, and the performance indicators are printed out.

[0023] In this embodiment, the pseudo code of the loss function is shown in Table 1 below: Table 1

[0024] In this embodiment, the pseudo code of the cyclic training step is shown in Table 2 below: Table 2

[0025] In this embodiment, the pseudo code of the model evaluation is shown in Table 3 below: Table 3

[0026] Preferably, the algorithm for the multi-way classification sub-problem is Algorithm for the road classification subproblem.

[0027] Preferably, the first fusion algorithm is: ,in, is the visual feature of the unseen category, is Gaussian noise, is a semantic descriptor, is the text description, G is the fusion process function of the generator; The first fusion algorithm is expanded as: ,in, 、 、 Gaussian noise , semantic descriptors and text description The nonlinear transformation of 、 、 Gaussian noise , semantic descriptors and text description The weight coefficients learned through the attention mechanism, is a non-linear activation function.

[0028] Preferably, a discriminator D is used to distinguish real samples and generate samples ,include: The calculation formula of the discriminator is: Distinguishing real samples and generate samples , and obtain the discrimination result, where represents the visual features of the input, Represents the discrimination result output by the discriminator.

[0029] Preferably, generating the first loss function includes using a discriminator to generate the first loss function, wherein the first loss function is an adversarial loss function; The adversarial loss function is: ,in, is the expected operation function, is the real data distribution, is the distribution function of input noise and semantic descriptor.

[0030] Preferably, the embedding space representation function is: , where E(x) is the embedding function. By using the embedding function to construct the embedding space representation function, high-dimensional visual features can be mapped to a low-dimensional embedding space while retaining the key information and category discriminant features of the sample.

[0031] Preferably, obtain an instance sample , based on example samples , for each instance sample structure Road classification sub-problem algorithm, get instance samples The representation results in the embedding space and the second loss function include: Based on nonlinear projection head , get the nonlinear projection head The loss function is: , where T is the transposition symbol; For the The embedded features of samples after nonlinear projection; is with Positive samples belonging to the same category are embedded; Is from other categories negative sample embeddings; is the number of negative samples; is the temperature parameter, which is used to adjust the sharpness of the similarity distribution in contrastive learning; And the second loss function: Calculate the instance sample The representation results in the embedding space are: For all positive sample pairs; For the generator; is the embedded function; For the projection head, the second loss function is the contrastive embedding loss function.

[0032] Preferably, the initial loss function for learning the comparator network F(h,a) is: in, For example samples The corresponding positive semantic expression, is the semantic descriptor of all seen categories, is the temperature parameter for class-level contrast embedding, and >0, is the number of seen categories.

[0033] Preferably, based on the initial loss function, a third loss function is obtained: ,in, For the generator, is the embedding function, is a semantic comparison network, For the embedding vectors, For the correct semantic representation, the third loss function is to optimize the category-level contrast embedding loss function.

[0034] Preferably, the parameters of the classifier are adjusted based on a cross entropy loss minimization algorithm, including: Use the minimized cross entropy loss function: Optimize the first parameter W and the second parameter b of the classifier, where N is the number of training samples, is the total number of categories, which includes the number of seen categories and the number of unseen categories. By optimizing the first parameter W and the second parameter b of the classifier, the target classifier can better learn the mapping relationship between the embedding space representation and the category label, thereby improving the classification accuracy.

[0035] like Figure 3 As shown, the present invention provides a system for a dynamic zero-sample image recognition method of multimodal fusion as in the above scheme, the system includes an unseen category visual feature generation module 1, a sample collection and differentiation module 2, a category level supervision module 3 and a feature classification module 4; the unseen category visual feature generation module 1 is configured to input the first data into the generator G, and use the first fusion algorithm to drive the generator G to generate the unseen category visual feature , wherein the first data includes Gaussian noise , semantic descriptors , text description , the visual features of the unseen categories are used to generate samples; the sample collection and differentiation module 2 is configured to collect real samples , and use the discriminator D to distinguish the real samples and generate samples , and generate the first loss function, collect visual features x as the input of the embedding mapping function, and output the embedding space representation function; the category level supervision module 3 is configured to obtain instance samples , based on example samples , for each instance sample structure Road classification sub-problem algorithm, get instance samples The representation results in the embedding space and the second loss function are calculated using the learning comparator network Fh,a to calculate the representation results and semantic descriptors The correlation between the first and second loss functions is obtained, and the third loss function is used to obtain the category-level supervision information of the instance sample; the feature classification module 4 is configured to adopt a classifier, adjust the parameters of the classifier based on the cross entropy loss algorithm to obtain a target classifier, and adopt the target classifier to classify the features of the input sample, and obtain the total loss data of the input sample based on the first loss function, the second loss function and the third loss function, minimize the total loss to obtain the target loss, and obtain the target classification result of the input sample based on the target loss.

[0036] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A multimodal fusion dynamic zero-shot image recognition method, characterized in that: include: Inputting first data into a generator, and using a first fusion algorithm to drive the generator to generate unseen visual features, wherein the first data includes Gaussian noise, a semantic descriptor, and a text description, and the unseen visual features are generated samples; Collect real samples, use a discriminator to distinguish real samples from generated samples, and generate the first loss function; Collect visual features as input to the embedding mapping function, and output the embedding space representation function; Obtain instance samples, construct a multi-way classification sub-problem algorithm for each instance sample based on the instance samples, obtain the representation result of the instance sample in the embedding space and the second loss function, and use the learning comparator network to calculate the representation result and semantic descriptor The third loss function is obtained by using the third loss function to obtain the category-level supervision information of the instance sample; Using a classifier, adjusting parameters of the classifier based on a cross entropy loss minimization algorithm to obtain a target classifier, and using the target classifier to classify features of the input sample; Based on the first loss function, the second loss function and the third loss function, the total loss data of the input sample is obtained, the target loss is obtained by minimizing the total loss, and the target classification result of the input sample is obtained based on the target loss.

2. The multimodal fusion dynamic zero-shot image recognition method according to claim 1, characterized in that: The first fusion algorithm is: ,in, is the visual feature of the unseen category, is Gaussian noise, is a semantic descriptor, is the text description, G is the fusion process function of the generator; The first fusion algorithm is expanded as follows: ,in, 、 、 Gaussian noise , semantic descriptors and text description The nonlinear transformation of 、 、 Gaussian noise , semantic descriptors and text description The weight coefficients learned through the attention mechanism, is a non-linear activation function.

3. The multimodal fusion dynamic zero-shot image recognition method according to claim 2, characterized in that: The discriminator is used to distinguish between real samples and generated samples, including: The calculation formula of the discriminator is: Distinguish the real samples from the generated samples and get the discrimination results, where represents the visual features of the input, Represents the discrimination result output by the discriminator.

4. The multimodal fusion dynamic zero-shot image recognition method according to claim 3, characterized in that: Generating the first loss function includes using a discriminator to generate the first loss function, wherein the first loss function is an adversarial loss function; The adversarial loss function is: ,in, is the expected operation function, is the real data distribution, is the distribution function of input noise and semantic descriptor.

5. The multimodal fusion dynamic zero-shot image recognition method according to claim 4, characterized in that: The embedding space representation function is: , where E(x) is the embedding function.

6. The multimodal fusion dynamic zero-shot image recognition method according to claim 5, characterized in that: The acquiring of instance samples, constructing a multi-way classification sub-problem algorithm for each instance sample based on the instance samples, and obtaining a representation result of the instance sample in the embedding space and a second loss function, includes: Based on nonlinear projection head , get the nonlinear projection head The loss function is: , where T is the transposition symbol; For the The embedded features of samples after nonlinear projection; is with Positive samples belonging to the same category are embedded; Is from other categories negative sample embeddings; is the number of negative samples; is the temperature parameter, which is used to adjust the sharpness of the similarity distribution in contrastive learning; And the second loss function: The representation result of the instance sample in the embedding space is calculated, where: For all positive sample pairs; For the generator; is the embedded function; For the projection head, the second loss function is the contrastive embedding loss function.

7. The multimodal fusion dynamic zero-shot image recognition method according to claim 6, characterized in that: The initial loss function of the learning comparator network is: in, For example samples The corresponding positive semantic expression, is the semantic descriptor of all seen categories, is the temperature parameter for class-level contrast embedding, and >0, is the number of seen categories.

8. The multimodal fusion dynamic zero-shot image recognition method according to claim 7, characterized in that: Based on the initial loss function, the third loss function is obtained: ,in, For the generator, is the embedding function, is a semantic comparison network, For the embedding vectors, For the correct semantic representation, the third loss function is to optimize the category-level contrast embedding loss function.

9. The multimodal fusion dynamic zero-shot image recognition method according to claim 8, characterized in that: The adjusting of the parameters of the classifier based on the cross entropy loss minimization algorithm includes: Use the minimized cross entropy loss function: Optimize the first parameter W and the second parameter b of the classifier, where N is the number of training samples, is the total number of categories, which includes the number of seen categories and the number of unseen categories.

10. A system, characterized in that: The dynamic zero-shot image recognition method for multimodal fusion according to any one of claims 1 to 9, wherein the system comprises: An unseen category visual feature generation module (1), wherein the unseen category visual feature generation module (1) is configured to input first data into a generator and drive the generator to generate unseen category visual features using a first fusion algorithm, wherein the first data includes Gaussian noise, a semantic descriptor, and a text description, and the unseen category visual features are generated samples; A sample collection and differentiation module (2), wherein the sample collection and differentiation module (2) is configured to collect real samples, and use a discriminator to distinguish real samples from generated samples, and generate a first loss function, collect visual features as input of an embedding mapping function, and output an embedding space representation function; A category-level supervision module (3) is configured to obtain instance samples, construct a multi-way classification sub-problem algorithm for each instance sample based on the instance samples, obtain a representation result of the instance sample in the embedding space and a second loss function, use a learning comparator network to calculate the correlation between the representation result and the semantic descriptor, obtain a third loss function, and use the third loss function to obtain category-level supervision information of the instance sample; A feature classification module (4) is configured to use a classifier, adjust the parameters of the classifier based on a cross entropy loss minimization algorithm to obtain a target classifier, use the target classifier to classify the features of the input sample, obtain total loss data of the input sample based on a first loss function, a second loss function, and a third loss function, minimize the total loss to obtain a target loss, and obtain a target classification result of the input sample based on the target loss.

Citation Information

Patent Citations

  • Zero sample image recognition model and training method, image recognition method, program product, medium and equipment

    CN119323692A

  • Mixed generalized zero sample learning method based on feature optimization

    CN115761355A

  • Contrast diffusion zero sample learning method

    CN118504656A