Model Training Method, Sample Generation Method, Device, Equipment and Storage Medium

By constructing a sample generation network model composed of encoder, generator and discriminator, generating adversarial samples through adversarial training, optimizing the product recognition network structure, solving the recognition error problem caused by factors such as light and angle during the product recognition process, and achieving higher recognition accuracy.

CN114693978BActive Publication Date: 2025-07-11广州市玄瞳科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210358641.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-07-11
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

During the product recognition process, due to factors such as image light, background color, photo angle, product rotation and occlusion, the identification results are prone to errors, and it is urgent to improve the accuracy of the identification results.

Method used

A sample generation network model consisting of the first encoder, generator, second encoder and discriminator is constructed, and an adversarial sample that is closer to natural types is generated through adversarial training, and the commodity identification network structure is optimized.

Benefits of technology

It improves the accuracy of product identification results, and the generated samples can effectively maintain the original visual information, enhancing the reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693978B_ABST
    Figure CN114693978B_ABST
Patent Text Reader

Abstract

The present invention discloses a model training method, a sample generation method, a device, a device and a storage medium. Among them, the model training method includes: obtaining a training image and inputting the training image into a first encoder to obtain a first hidden vector output by the first encoder; inputting the first hidden vector into a generator to obtain a reconstructed image output by the generator; the reconstructed image contains a target noise generated according to a preset noise generation method; inputting the reconstructed image into a second encoder to obtain a second hidden vector output by the second encoder; inputting the second hidden vector into a discriminator to obtain a discrimination result output by the discriminator; performing adversarial training on the generator and the discriminator according to the discrimination result until the training end condition is satisfied, and then determining the second encoder and the discriminator as a sample generation model. The above method can generate more reliable adversarial samples closer to the natural type through adversarial training of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a model training method, a sample generation method, a device, a device and a storage medium. Background Art

[0002] With the expansion and in-depth application of product fingerprints in the fast-moving consumer goods field, it has gradually become a trend to apply product fingerprints on a large scale to product recognition. The identification feature of product fingerprints is that after locking a product, the closest feature can be matched in the product feature library through feature search, thus completing the product identification process.

[0003] However, during the product identification process, affected by factors such as image light, background color, photographing angle, product rotation, and occlusion, the identification result is often prone to errors. Therefore, there is an urgent need to provide a solution that can accurately identify product image information under complex conditions. Summary of the Invention

[0004] The purpose of the present invention is to at least solve one of the technical problems existing in the prior art, and provide a model training method, a sample generation method, a device, a device and a storage medium. Among them, the method generates more natural-type adversarial samples through adversarial training, and then can efficiently optimize the product recognition network structure to improve the accuracy of the recognition result.

[0005] In a first aspect, the present invention provides a model training method, including:

[0006] Obtain a training image and input the training image into a first encoder to obtain a first hidden vector output by the first encoder;

[0007] Input the first hidden vector into a generator to obtain a reconstructed image output by the generator; the reconstructed image contains target noise generated according to a preset noise generation method;

[0008] Input the reconstructed image into a second encoder to obtain a second hidden vector output by the second encoder; the error between the second hidden vector and the first hidden vector is less than a preset threshold;

[0009] Input the second hidden vector into a discriminator to obtain a discrimination result output by the discriminator;

[0010] Perform adversarial training on the generator and the discriminator according to the discrimination result until the training end condition is met, and then determine the second encoder and the discriminator as a sample generation model.

[0011] As a further improvement, the preset noise generation method includes:

[0012] Set an initial noise range and iteratively obtain the second hidden vectors corresponding to each noise value within the noise range;

[0013] Gradually narrow down the initial noise range based on the obtained second hidden vectors until the target noise is determined.

[0014] As a further improvement, the preset noise generation method further includes:

[0015] Obtain the classification result of the second hidden vector output by the second encoder through a pre - constructed classification model;

[0016] Determine the noise range based on the classification result of the second hidden vector and the classification label of the training image;

[0017] Determine the target noise according to the error between the second hidden vectors corresponding to each noise value within the noise range and the first hidden vector.

[0018] As a further improvement, the pre - constructed classification model is specifically a DenseNet network model.

[0019] In a second aspect, the present invention further provides a sample generation method, including:

[0020] Obtain an initial image, input the initial image into a sample generation model, and the discriminator in the sample generation model generates a target sample of the initial image by adding target noise; wherein,

[0021] The target noise is generated by the preset noise generation method;

[0022] The sample generation model is a model trained by using the model training method as described in the first aspect.

[0023] In a third aspect, the present invention further provides a model training device, including:

[0024] A first encoding module, configured to obtain a training image and input the training image into a first encoder to obtain a first hidden vector output by the first encoder;

[0025] A generation module, configured to input the first hidden vector into a generator to obtain a reconstructed image output by the generator; the reconstructed image contains the target noise generated according to the preset noise generation method;

[0026] A second encoding module, configured to input the reconstructed image into a second encoder to obtain a second hidden vector output by the second encoder; the error between the second hidden vector and the first hidden vector is less than a preset threshold;

[0027] A discrimination module, configured to input the second hidden vector into a discriminator to obtain a discrimination result output by the discriminator;

[0028] A training module, configured to perform adversarial training on the generator and the discriminator according to the discrimination result until a training end condition is satisfied, and then determine the second encoder and the discriminator as a sample generation model.

[0029] As a further improvement, in the generation module, the preset noise generation method includes:

[0030] Set an initial noise range, and iteratively obtain second hidden vectors corresponding to respective noise values within the noise range;

[0031] Gradually narrow down the initial noise range based on the obtained second hidden vectors until a target noise is determined.

[0032] As a further improvement, in the generation module, the preset noise generation method further includes:

[0033] Obtain a classification result of the second hidden vector output by the second encoder through a pre-constructed classification model;

[0034] Determine a noise range based on the classification result of the second hidden vector and the classification label of the training image;

[0035] Determine the target noise according to the error between the second hidden vector corresponding to each noise value within the noise range and the first hidden vector.

[0036] In a fourth aspect, the present invention provides a data processing device, including a processor, the processor being coupled to a memory, the memory storing a program, the program being executed by the processor, so that the data processing device executes the model training method described in the first aspect or the sample generation method described in the second aspect.

[0037] In a fifth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the model training method described in the first aspect or the sample generation method described in the second aspect is implemented.

[0038] Compared with the prior art, a model training method provided by the present invention at least has the following beneficial effects:

[0039] Construct a sample generation network model composed of a first encoder, a generator, a second encoder, and a discriminator, and train model parameters through an adversarial training method. At the same time, add noise during the training process, so that the samples generated by the model can effectively retain the original visual information and have higher reliability.

[0040] Furthermore, the reliable samples generated by using the above sample generation model can be used to efficiently optimize the commodity recognition network structure, so as to improve the accuracy of commodity recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for the implementation manners will be briefly introduced below. Obviously, the drawings in the following description are only some implementation manners of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is a schematic flowchart of the model training method provided by an embodiment of the present invention;

[0043] Figure 2 is a schematic structural diagram of the sample generation model provided by an embodiment of the present invention;

[0044] Figure 3 is a schematic structural diagram of the model training device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] This part will describe the specific embodiments of the present invention in detail. The preferred embodiments of the present invention are shown in the drawings. The drawings are used to supplement the description of the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present invention. However, it should not be construed as a limitation on the protection scope of the present invention.

[0046] As Figure 1 shown, in a first aspect, an embodiment of the present invention provides a model training method, including the following steps S1 - S5.

[0047] S1: Obtain training images and input the training images into a first encoder to obtain a first hidden vector output by the first encoder.

[0048] S2: Input the first hidden vector into a generator to obtain a reconstructed image output by the generator; the reconstructed image contains target noise generated according to a preset noise generation method.

[0049] S3: Input the reconstructed image into a second encoder to obtain a second hidden vector output by the second encoder; the error between the second hidden vector and the first hidden vector is less than a preset threshold.

[0050] S4: Input the second hidden vector into a discriminator to obtain a discrimination result output by the discriminator.

[0051] S5: Perform adversarial training on the generator and the discriminator according to the discrimination result until the training end condition is met, and then determine the second encoder and the discriminator as the sample generation model.

[0052] It should be noted that in S2, the noise is generated by the jitter of pixel values, and specifically, interference is carried out by setting a certain jitter range.

[0053] In one example, the preset noise generation method can iteratively obtain the second hidden vectors corresponding to each noise value within the noise range by setting an initial noise range; gradually narrow the initial noise range based on the obtained second hidden vectors, and perform a search from coarse to fine until the target noise is determined.

[0054] In another example, the classification result of the second hidden vector output by the second encoder can also be obtained through a pre-constructed classification model. Based on the classification result of the second hidden vector and the classification label of the training image, the noise range is determined, and then the target noise is determined according to the error between the second hidden vector and the first hidden vector corresponding to each noise value within the noise range.

[0055] Specifically, N perturbations can be randomly sampled in each iteration process, and the search range is gradually increased according to the set search step size until the label of the corresponding generated data changes, that is: the classification result of the second hidden vector is inconsistent with the classification label of the training image, and then the one with the highest similarity to the original sample is selected from the generated adversarial samples to determine the target noise. Among them, the similarity can be determined by the error between the first hidden vector and the second hidden vector.

[0056] In this embodiment, the pre-constructed classification model can be a DenseNet network model.

[0057] Specifically, a DenseNet 202 classifier can be trained with 2000-class beverage data. This classifier is a standard DenseNet, and a compression rate of 0.25 and a growth rate coefficient of 12 are set. The loss function is the standard cross-entropy function, the learning rate is 0.001, and the decay coefficient is 0.00001. After a total of 60 epochs of training, it basically converges.

[0058] The present invention will describe the implementation process of the above model training method through a specific embodiment.

[0059] Please refer to Figure 2 The structural schematic diagram of the provided sample generation model. In this embodiment, the sample generation model is composed of a first encoder F, a generator G, a second encoder E, and a discriminator D. Among them, the first encoder F and the generator G form the first set of encoder-decoders, and the second encoder E and the discriminator D form the second set of encoder-decoders. The two sets of encoder-decoders are trained in an adversarial manner.

[0060] Specifically, the first group of codecs is used to re-encode the input original samples to obtain the first latent vector h. During the re-decoding process, the target noise n needs to be added to obtain the reconstructed image G(h, n).

[0061] It can be understood that the class label of the reconstructed image G(h, n) with the target noise n can be consistent with the label of the training image.

[0062] Further, the second group of codecs is used to re-encode the reconstructed image G(h, n) output by the first group of codecs to obtain the second latent vector h ^ and the error between the first latent vector h and the second latent vector h ^ is made less than a preset threshold, that is: the latent variable distributions generated by the two groups of codecs are made consistent.

[0063] Further, the adversarial training is performed on the latent variables of the second group of codecs by using the adversarial generative network until the training end condition is satisfied.

[0064] In this embodiment, the feature distribution of the latent space is matched to an arbitrary prior distribution. For example, if the input is x and the latent space outputs h, the codec results are q(h|x) and q(x|h). Based on the distribution q(h) of the latent space feature h to match a certain prior p(h), then the distribution q(h) will become the integral of p(h) and q(h|x).

[0065] Since the latent space can always find a suitable matching distribution in the prior, therefore, the model can naturally find the class that is closest to it or has the closest appearance and texture features in the learned classes for mapping.

[0066] During the model training process, classification training needs to be performed on the second latent vector h ^ specifically, the cross-entropy loss can be used to train the classifier so that the second group of codecs has classification ability.

[0067] This embodiment trains the network in a multi-task manner, which can not only ensure that the classification ability is not affected, but also make the network more robust to noise.

[0068] In the above embodiment of the present invention, a sample generation network model composed of a first encoder, a generator, a second encoder, and a discriminator is constructed, and the model parameters are trained based on the adversarial training method. At the same time, noise is added during the training process, so that the samples generated by the model can effectively retain the original visual information and have higher reliability.

[0069] Further, the reliable samples generated by using the above sample generation model can be used to efficiently optimize the commodity recognition network structure to improve the accuracy of the commodity recognition result.

[0070] In a second aspect, the present invention further provides a sample generation method, including: obtaining an initial image, inputting the initial image into a sample generation model, and generating a target sample of the initial image by a discriminator in the sample generation model by adding target noise. Wherein, the target noise is generated by a preset noise generation method, and the sample generation model is a model trained by the model training method described in the first aspect.

[0071] In a third aspect, the present invention further provides a model training device, including a first encoding module 101, a generation module 102, a second encoding module 103, a discrimination module 104, and a training module 105.

[0072] The first encoding module 101 is configured to obtain a training image and input the training image into a first encoder to obtain a first hidden vector output by the first encoder.

[0073] The generation module 102 is configured to input the first hidden vector into a generator to obtain a reconstructed image output by the generator; the reconstructed image includes target noise generated according to a preset noise generation method.

[0074] The second encoding module 103 is configured to input the reconstructed image into a second encoder to obtain a second hidden vector output by the second encoder; the error between the second hidden vector and the first hidden vector is less than a preset threshold.

[0075] The discrimination module 104 is configured to input the second hidden vector into a discriminator to obtain a discrimination result output by the discriminator.

[0076] The training module 105 is configured to perform adversarial training on the generator and the discriminator according to the discrimination result until a training end condition is met, and then determine the second encoder and the discriminator as a sample generation model.

[0077] Regarding the information interaction, execution process, etc. among the above-mentioned modules in the device, since they are based on the same concept as the embodiment of the model training method of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0078] In a fourth aspect, the present invention provides a data processing device, including a processor, the processor is coupled to a memory, the memory stores a program, and the program is executed by the processor, so that the data processing device executes the model training method described in the first aspect or the sample generation method described in the second aspect.

[0079] In a fifth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the model training method described in the first aspect or the sample generation method described in the second aspect is implemented.

[0080] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0081] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

Claims

1. A model training method, characterized in that, Including: Obtain a training image and input the training image into a first encoder to obtain a first hidden vector output by the first encoder; Input the first hidden vector into a generator to obtain a reconstructed image output by the generator; the reconstructed image contains target noise generated according to a preset noise generation method; Input the reconstructed image into a second encoder to obtain a second hidden vector output by the second encoder; the error between the second hidden vector and the first hidden vector is less than a preset threshold; Input the second hidden vector into a discriminator to obtain a discrimination result output by the discriminator; Perform adversarial training on the generator and the discriminator according to the discrimination result until the training end condition is satisfied, and then determine the second encoder and the discriminator as a sample generation model; Wherein, the preset noise generation method further includes: obtaining a classification result of the second hidden vector output by the second encoder through a pre-constructed classification model; determining a noise range based on the classification result of the second hidden vector and the classification label of the training image; determining the target noise according to the error between the second hidden vector corresponding to each noise value within the noise range and the first hidden vector.

2. The model training method according to claim 1, wherein The preset noise generation method includes: Set an initial noise range and iteratively obtain the second hidden vector corresponding to each noise value within the noise range; Gradually narrow down the initial noise range based on the obtained second hidden vector until the target noise is determined.

3. The model training method according to claim 1, wherein The pre-constructed classification model is specifically a DenseNet network model.

4. A method for generating a sample, characterized in that, Including: Obtain an initial image, input the initial image into a sample generation model, and the discriminator in the sample generation model generates a target sample of the initial image by adding target noise; wherein, The target noise is generated by a preset noise generation method; The sample generation model is a model trained by using the model training method described in any one of claims 1 to 3.

5. A model training device, characterized in that, Including: A first encoding module, configured to obtain a training image and input the training image into a first encoder to obtain a first hidden vector output by the first encoder; A generation module, configured to input the first hidden vector into a generator to obtain a reconstructed image output by the generator; the reconstructed image contains target noise generated according to a preset noise generation method; A second encoding module, configured to input the reconstructed image into a second encoder to obtain a second hidden vector output by the second encoder; the error between the second hidden vector and the first hidden vector is less than a preset threshold; A discrimination module, configured to input the second hidden vector into a discriminator to obtain a discrimination result output by the discriminator; A training module, configured to perform adversarial training on the generator and the discriminator according to the discrimination result until the training end condition is satisfied, and then determine the second encoder and the discriminator as a sample generation model; Wherein, the preset noise generation method further includes: obtaining a classification result of a second hidden vector output by a second encoder through a pre-constructed classification model; determining a noise range based on the classification result of the second hidden vector and the classification label of the training image; and determining a target noise according to an error between the second hidden vector corresponding to each noise value within the noise range and the first hidden vector.

6. The model training device according to claim 5, wherein In the generation module, the preset noise generation method includes: setting an initial noise range, and iteratively obtaining second hidden vectors corresponding to each noise value within the noise range; gradually narrowing down the initial noise range based on the obtained second hidden vectors until the target noise is determined.

7. A data processing device, characterized in that, including: a processor, the processor is coupled to a memory, the memory stores a program, and the program is executed by the processor to enable the data processing device to execute the model training method according to any one of claims 1 to 3 or the sample generation method according to claim 4.

8. A computer storage medium, characterized in that, The computer storage medium stores computer instructions for executing the model training method according to any one of claims 1 to 3 or the sample generation method according to claim 4.

Citation Information

Patent Citations

  • Color image reconstruction method and device and image classification method and device

    CN111080727A

  • Image defect detection model training method and device, and computer equipment

    CN113610787A

  • Image reconstruction method and system, terminal equipment and storage medium

    CN113674187A

  • Image processing method, image processing model training method and device, and storage medium

    CN113963087A