Image processing method, device, computer equipment and storage medium

Through the method of feature extraction and optimization parameter update, the problem of difficulty in improving the effect of traditional image recognition models is solved, and the performance and effect of image recognition models are significantly improved.

CN114492734BActive Publication Date: 2025-06-06SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111651438.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-06-06
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The traditional image recognition effect based on image recognition models is difficult to improve, which limits the ability of image recognition processing.

Method used

By extracting the received images, generating identity identification features, determining optimization parameters based on feature distance, and updating the output results of the loss function of the image recognition model to improve the performance of the image recognition model.

Benefits of technology

The image recognition processing capability of the image recognition model has been significantly improved, and the image recognition effect has been further improved, which is suitable for improving the performance of the lightweight image recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492734B_ABST
    Figure CN114492734B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, computer equipment and storage medium for image processing, wherein the method for image processing may include: performing feature extraction processing on a received first image to obtain multiple original image features of the first image; using multiple original image features to generate identity identification features of the first image, determining the feature distance between each original image feature and the identity identification feature, and determining optimization parameters according to the feature distance; using the optimization parameters to update the output result of the loss function used for iterative training of an image recognition model, and the trained image recognition model is used to perform image recognition processing on a second image to be identified. The device may include an original feature extraction module, an identification feature generation module, an optimization parameter determination module and a loss output update module. Compared with the prior art, the present invention can significantly improve the image recognition processing capability of the image recognition model, and the effect of image recognition is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology. More specifically, the present invention can provide an image processing method, an apparatus, a computer device and a storage medium. Background Art

[0002] At present, the technology of obtaining image recognition models by training neural networks has been widely used. However, in order to improve the performance of image recognition models, the complexity of the neural network structure is increasing, resulting in a decrease in the real-time responsiveness of image processing. Therefore, someone introduced the method of knowledge distillation. The main idea of ​​knowledge distillation is to train a small-scale model to imitate a large-scale model, so that the small-scale model has a certain accuracy and a simple network. This method can be used to obtain an image recognition model with a low neural network complexity. Although a small-scale image recognition model with certain performance is obtained in this way, the disadvantage of this method is that it limits the image recognition model's ability to process image recognition, and the effect of image recognition is difficult to further improve, and it is urgently needed to be improved or optimized. Summary of the invention

[0003] In order to solve the problem that the image recognition effect based on traditional image recognition models is difficult to improve, the present invention can provide an image processing method, device, computer equipment and storage medium, so as to achieve technical goals such as further improving image recognition processing performance.

[0004] To achieve the above technical objectives, the present invention can provide an image processing method, which may include but is not limited to one or more of the following steps.

[0005] Perform feature extraction processing on the received first image to obtain multiple original image features of the first image.

[0006] The identity identification features of the first image are generated using the multiple original image features.

[0007] The feature distance between each of the original features of the image and the identity identification feature is determined, and the optimization parameter is determined according to the feature distance.

[0008] The optimization parameters are used to update the output result of the loss function used for iterative training of the image recognition model, and the trained image recognition model is used to perform image recognition processing on the second image to be recognized.

[0009] In order to achieve the above technical objectives, the present invention can also provide an image processing device, which includes but is not limited to an original feature extraction module, an identification feature generation module, an optimization parameter determination module and a loss output update module.

[0010] The original feature extraction module is used to perform feature extraction processing on the received first image to obtain multiple original image features of the first image.

[0011] The identification feature generation module is used to generate the identity identification feature of the first image by using the multiple original features of the images.

[0012] The optimization parameter determination module is used to determine the feature distance between each of the original features of the image and the identity identification feature, and determine the optimization parameter according to the feature distance.

[0013] The loss output updating module is used to update the loss function output result used for iterative training of the image recognition model using the optimization parameters, and the trained image recognition model is used to perform image recognition processing on the second image to be recognized.

[0014] In order to achieve the above-mentioned technical objectives, the present invention can also provide a computer device, which includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the image processing method described in any embodiment of the present invention.

[0015] To achieve the above-mentioned technical objectives, the present invention may also provide a storage medium storing computer-readable instructions, which, when executed by one or more processors, enables the one or more processors to perform the steps of the image processing method described in any embodiment of the present invention.

[0016] To achieve the above technical objectives, the present invention can also provide a computer program product. When the instructions in the computer program product are executed by a processor, the steps of the image processing method described in any embodiment of the present invention are executed.

[0017] The beneficial effects of the present invention are as follows: based on the original image features obtained through image feature extraction and the identity features obtained after reprocessing, the present invention updates the loss function used for iterative training of the image recognition model according to the optimization parameters determined by the feature distance between the original image features and the identity features, thereby avoiding the limitation of the loss function itself and the large-scale model on the performance of the image recognition model. The technical solution provided by the present invention can significantly improve the image recognition processing capability of the image recognition model, and the image recognition effect is further improved. The present invention is suitable for improving the performance of lightweight image recognition models, and is particularly suitable for lightweight and high-precision deep neural networks obtained based on image distillation. The present invention specifically describes the image feature distribution through the relationship between the original image features and the identity features, thereby obtaining the optimization parameters for updating the output results of the loss function according to the image feature distribution, thereby achieving the purpose of improving the recognition performance of the image recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic flow chart of an image processing method in one or more embodiments of the present invention is shown.

[0019] Figure 2 A schematic diagram of a process for generating relaxation coefficients for distillation training in one or more embodiments of the present invention is shown.

[0020] Figure 3 A schematic diagram of the process of distillation training of an image recognition model in one or more embodiments of the present invention is shown.

[0021] Figure 4 A schematic diagram showing the relationship between characteristic distance and frequency in one or more embodiments of the present invention is shown.

[0022] Figure 5 A schematic diagram of generating identity identification features based on original image features in one or more embodiments of the present invention is shown.

[0023] Figure 6 A schematic diagram showing a picture query test based on the image recognition model obtained in one or more embodiments of the present invention is shown.

[0024] Figure 7 A schematic diagram showing the structural composition of an image processing device in one or more embodiments of the present invention is shown.

[0025] Figure 8 A schematic diagram of the internal structure of a computer device in one or more embodiments of the present invention is shown. DETAILED DESCRIPTION

[0026] The following is a detailed explanation and description of an image processing method, apparatus, computer equipment and storage medium provided by the present invention in conjunction with the accompanying drawings.

[0027] like Figure 1 As shown, and can be combined Figure 2 and Figure 3 One or more embodiments of the present invention may provide an image processing method. The image processing method includes but is not limited to one or more of the following steps, which are described in detail as follows.

[0028] Step 100, feature extraction is performed on the received first image to obtain a plurality of original image features of the first image. It should be understood that the first image in the embodiment of the present invention is generally a plurality of images, such as Figure 5 The first image shown in the upper middle part may include picture 0, picture 1, picture 2, picture 3, or the first picture may include Figure 5For the pictures 0, 1, and 2 in the middle and lower part, a feature extraction model can be selected to realize feature extraction, for example, feature extraction is performed on picture 0 to obtain the features of picture 0, feature extraction is performed on picture 1 to obtain the features of picture 1, feature extraction is performed on picture 2 to obtain the features of picture 2, feature extraction is performed on picture 3 to obtain the features of picture 3, and so on. Among them, the feature extraction model can be, for example, a feature extraction model in a teacher network; the first image size in this embodiment can be 128 high × 112 wide, and the feature dimension can be 512, but it is certainly not limited thereto.

[0029] Optionally, the embodiment of the present invention may include: receiving the first image after data augmentation, and performing feature extraction on the first image after data augmentation. The data augmentation process of the present invention may include a variety of different processing methods, including but not limited to: (1) grayscale conversion, converting the three-channel RGB color image into a GRAY grayscale image according to a certain probability; (2) mirror flipping, randomly flipping the image left and right; (3) scale transformation, scaling the image scale to a preset scale (such as 0.5 of the original width and height) and then scaling it back to the original scale; (4) adding noise, randomly generating a black block of a preset size to replace a block in the image; (5) random erasing, using random values ​​to erase image pixels; (6) Gaussian blur processing of the image; (7) random flipping of the image; so that one image can obtain corresponding multiple different image features.

[0030] Step 200, using multiple original features of the image to generate an identity feature of the first image. The identity feature of the present invention is an ID (Identity document, identity feature), which can be used as a target for deep learning. For example, Figure 5 The image 0 feature, image 1 feature, image 2 feature and image 3 feature shown in the upper part generate the identity 0 feature, or can be obtained by Figure 5 The picture 0 feature, picture 1 feature and picture 2 feature shown in the lower part generate the identity identification n feature, but of course it is not limited to this.

[0031] like Figure 5 As shown, in the embodiment of the present invention, using multiple original image features to generate the identity identification feature of the first image includes: performing mean processing on the multiple original image features to obtain the identity identification feature. The mean processing in the embodiment of the present invention can be arithmetic mean calculation, and each ID feature is the average value of all original image features in the image, and the original image features include but are not limited to the original image features of the image after data enhancement processing. Based on the original image features, the embodiment of the present invention can specifically determine the identity identification feature through the following formula.

[0032]

[0033] in, The feature value representing the identity feature, The eigenvalue represents the original feature of the image, and n represents the number of images with the same identity.

[0034] Step 300, determine the feature distance between each original feature of the image and the identity feature, and determine the optimization parameter based on the feature distance. The present invention is used for distillation training of image recognition models. The present invention adds a relaxation strategy to the image feature distillation process, so the optimization parameter in this embodiment is the relaxation coefficient. It should be understood that a lightweight and high-precision deep network can be trained by image distillation, or a network structure with a certain accuracy can be quickly initialized.

[0035] Optionally, the embodiment of the present invention determines the feature distance between each original feature of the image and the identity identification feature, which may include: determining the second difference between the feature value of the original feature of the image and the feature value of the identity identification feature in each feature dimension, and determining the feature distance between the original feature of the image and the identity identification feature based on the sum of the squares of the second difference in all feature dimensions. The embodiment of the present invention may use the square root of the sum of the squares of the second difference in all feature dimensions (specifically the square root) as the feature distance between the original feature of the image and the identity identification feature, such as calculating using the following formula.

[0036]

[0037] Among them, L d represents the feature distance, The feature value representing the identity feature, The feature value represents the original feature of the image, dim represents the feature dimension, and the feature dimension dim in this embodiment is specifically 512.

[0038] Specifically, the embodiment of the present invention determines the optimization parameter according to the feature distance, including: sorting the obtained multiple feature distances in order from small to large or from large to small to obtain a sorting result of the multiple feature distances; selecting the feature distance at a preset position in the sorting result as the optimization parameter. This embodiment determines the image feature distribution by sorting, such as Figure 4 The relationship between the feature distance between the original image feature and the identity identification feature and the frequency of occurrence of the feature distance is shown.

[0039] The embodiment of the present invention sorts all the obtained feature distances in ascending order. The preset position may be, for example, 10%, 12.5%, or 20%. In this embodiment, the preset position is 10%, corresponding to Figure 4The shadow area on the left side of the vertical line in FIG. 1 accounts for 10% of the total shadow area. relax Represents the optimization parameters, that is, through L relax represents the relaxation coefficient. In the figure, the relaxation coefficient corresponding to the first 10% may be 0.3825, for example.

[0040] Step 400, using the optimized parameters to update the output result of the loss function used for iterative training of the image recognition model, the trained image recognition model is used to perform image recognition processing on the second image to be recognized. The present invention uses the obtained optimized parameters to assist the training of the image recognition model, that is, uses the relaxation coefficient to optimize the distillation training process of the image recognition model. The present invention specifically adds a relaxation coefficient in the image feature distillation process. Among them, the loss function in the embodiment of the present invention may include but is not limited to the L2 regression loss function.

[0041] The image feature distillation method of the present invention may include: identifying the image features for distillation through a teacher network, and then using these features as learning targets, and then using RMSE, L2 loss and other supervised methods for training; wherein RMSE refers to Root Mean Square Error, L2 loss refers to L2 regression loss function, and L2 refers to L2 norm (Euclidean norm). In the process of image distillation, the present invention adds a relaxation coefficient based on the feature distribution on the basis of the L2 regression loss function to improve the distillation effect and the image recognition level.

[0042] like Figure 3As shown, after calculating the current output result of the loss function, the present invention determines the loss after adding the relaxation coefficient, and then performs cyclic iterative training. Specifically, in this embodiment, the output result of the loss function for iterative training of the image recognition model using the optimization parameter is updated, including: obtaining the first recognition result of the student network for the first image and the second recognition result of the teacher network for the first image, wherein the student network is used to form an image recognition model, and the teacher network is used to perform distillation training on the student network; processing the first recognition result and the second recognition result by the loss function to obtain the current output result of the loss function; updating the current output result of the loss function using the optimization parameter, that is, the present invention can optimize the output result of the loss function by the relaxation coefficient, and it can be seen that the present invention can provide an image feature distillation method adding a relaxation strategy. The teacher network in the embodiment of the present invention can be, for example, resnet101 (residual network 101), and the student network can be, for example, resnet50 (residual network 50), but it is certainly not limited thereto. The distillation training process of this embodiment can use 2.6 million identity identifiers, and use 10 million pictures as a training set, and each identity identifier contains at least two pictures. For the specific distillation training process, it can be reasonably selected according to needs, and this embodiment will not be repeated.

[0043] Optionally, the embodiment of the present invention may use the optimization parameters to update the output result of the loss function for iterative training of the image recognition model, which may include: determining the first difference between the current output result of the loss function and the optimization parameters, and if the first difference is greater than or equal to zero, updating the output result of the loss function to the first difference; if the first difference is less than zero, updating the output result of the loss function to zero. The final output result of the loss function of the present invention may be the L2 feature distance between the predicted features of the student network and the target features of the teacher network, which is represented by the following formula in this embodiment.

[0044] L = max{0, L dist -L relax}

[0045]

[0046] Among them, L represents the final output result of the loss function, L dist Represents the current output result of the loss function, L relax represents the relaxation coefficient; represents the output of the student network, Represents the output of the teacher network.

[0047] like Figure 3As shown, the present invention can perform data augmentation processing on the input image data. Specifically, the embodiment of the present invention can obtain the first recognition result of the student network for the first image and the second recognition result of the teacher network for the first image, which may include: using the student network to perform image recognition processing on the first image after data augmentation processing to obtain the first recognition result; using the teacher network to perform image recognition processing on the first image after data augmentation processing to obtain the second recognition result; and then obtaining the first recognition result and the second recognition result. Among them, the data augmentation processing may include a variety of different augmentation processing methods, including but not limited to: (1) grayscale conversion, converting the three-channel RGB color image into a GRAY grayscale image according to a certain probability; (2) mirror flipping, randomly flipping the image left and right; (3) scale transformation, scaling the image scale to a preset scale (such as 0.5 of the original width and height) and then scaling it back to the original scale; (4) adding noise, randomly generating a black block of a preset size to replace a block in the image; (5) random erasing, using random values ​​to erase image pixels; (6) Gaussian blur processing image; (7) random flipping image. By combining data enhancement in the process of image feature distillation, the present invention can further improve the distillation effect of image features so that the image recognition model has stronger image recognition performance.

[0048] like Figure 6 As shown, it shows a schematic diagram of the search accuracy calculation under 1 million interference base libraries applying the present invention. 1. Each test set is divided into an identity identification (ID) image and a search (query) image, wherein each identity identification may correspond to one identity identification image and multiple search images. 2. The images in the interference base library images are identity identification images, and the interference base library identities have no intersection with the identity identifications of the test set. 3. Specific search process: before searching a certain search image, the corresponding identity identification image is added to the interference base library as the search base library of the search image. After the search of the search image is completed, the identity identification image is deleted from the interference base library, and the above process is repeated to search the next search image; after all the search images of a test set are completed, the top1 (ranked first) hit rate of the test set is calculated.

[0049]

[0050] After determining the top1 hit rate of each test set, the top1 hit rates of all test sets are averaged to obtain the final search accuracy.

[0051] Tests have shown that the student network (image recognition model) obtained based on the relaxation coefficient distillation of the present invention has an average top 1 (ranked first) search ACC (accuracy) of 80.12 under the condition of 1 million interference base images. It can be seen that the present invention can significantly improve the performance of the image recognition model.

[0052] Based on the technical solution provided by the present invention, the present invention describes the image feature distribution based on the relationship between the original image features and the identity identification features, so as to obtain the optimization parameters for updating the output results of the loss function according to the image feature distribution, so as to avoid the limitations of the traditional loss function itself on the performance of the image recognition model, thereby achieving the purpose of improving the recognition performance of the image recognition model, and effectively solving the problem that the currently achievable image recognition effect is difficult to further improve.

[0053] Compared with conventional knowledge distillation strategies, the present invention is suitable for improving the performance of lightweight image recognition models, and is particularly suitable for lightweight and high-precision deep neural networks obtained based on image distillation.

[0054] like Figure 7 As shown, based on the same inventive technical concept as the image processing method, one or more embodiments of the present invention can also provide an image processing device.

[0055] The image processing device may specifically include but is not limited to an original feature extraction module 501, an identification feature generation module 502, an optimization parameter determination module 503 and a loss output update module 504, which are described in detail below.

[0056] The original feature extraction module 501 is used to perform feature extraction processing on the received first image to obtain multiple original image features of the first image.

[0057] Optionally, the original feature extraction module 501 is used to receive the first image that has been processed by data enhancement, and to perform feature extraction processing on the first image that has been processed by data enhancement.

[0058] The identification feature generation module 502 is used to generate an identity identification feature of the first image using multiple original image features.

[0059] Optionally, the identification feature generation module 502 is used to perform mean processing on multiple original image features to obtain identity identification features.

[0060] The optimization parameter determination module 503 is used to determine the feature distance between each original feature of the image and the identity identification feature, and determine the optimization parameter according to the feature distance.

[0061] Optionally, the optimization parameter determination module 503 can be used to determine the second difference between the feature value of the original feature of the image and the feature value of the identity identification feature in each feature dimension, and the optimization parameter determination module 503 can be used to determine the feature distance between the original feature of the image and the identity identification feature based on the sum of the squares of the second difference in all feature dimensions.

[0062] Optionally, the optimization parameter determination module 503 is used to sort the obtained multiple feature distances in order from small to large or from large to small to obtain a sorting result of the multiple feature distances; the optimization parameter determination module 503 is used to select the feature distance at a preset position in the sorting result as the optimization parameter.

[0063] The loss output updating module 504 is used to update the loss function output result used for iterative training of the image recognition model using the optimized parameters. The trained image recognition model is used to perform image recognition processing on the second image to be recognized.

[0064] Optionally, the loss output update module 504 is used to determine a first difference between a current output result of the loss function and the optimization parameter; the loss output update module 504 is used to update the output result of the loss function to the first difference based on the first difference being greater than or equal to zero; the loss output update module 504 is used to update the output result of the loss function to zero based on the first difference being less than zero.

[0065] Optionally, the loss output update module 504 is used to obtain a first recognition result of the student network for the first image and a second recognition result of the teacher network for the first image, the loss output update module 504 is used to process the first recognition result and the second recognition result through the loss function to obtain a current output result of the loss function, and the loss output update module 504 is used to update the current output result of the loss function using the optimization parameters. Among them, the student network in the embodiment of the present invention is used to form an image recognition model, and the teacher network is used to perform distillation training on the student network.

[0066] Optionally, the loss output update module 504 can be used to perform image recognition processing on the first image that has been processed with data enhancement using the student network to obtain a first recognition result, and the loss output update module 504 can be used to perform image recognition processing on the first image that has been processed with data enhancement using the teacher network to obtain a second recognition result; the loss output update module 504 can be used to obtain the first recognition result and the second recognition result.

[0067] like Figure 8 As shown, based on the same inventive technical concept as the image processing method, one or more embodiments of the present invention can also provide a computer device, the computer device includes a memory and a processor, the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the image processing method in any embodiment of the present invention. Among them, the specific implementation process of the image processing method has been described in detail in this specification and will not be repeated here.

[0068] like Figure 8As shown, based on the same inventive technical concept as the image processing method, one or more embodiments of the present invention can also provide a storage medium storing computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the image processing method in any embodiment of the present invention. Among them, the specific implementation process of the image processing method has been described in detail in this specification and will not be repeated here.

[0069] Based on the same technical concept as the image processing method, one or more embodiments of the present invention can also provide a computer program product. When the instructions in the computer program product are executed by a processor, the steps of the image processing method described in any embodiment of the present invention are executed. The specific implementation process of the image processing method of this embodiment has been described in detail in this specification and will not be repeated here.

[0070] The logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, apparatus, or device and execute instructions), or in conjunction with such instruction execution system, apparatus, or device. For purposes of this specification, "computer-readable storage medium" may be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device. More specific examples of computer-readable storage media (a non-exhaustive list) include the following: an electrical connection with one or more wirings (electronic devices), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory, or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM, Compact Disc Read-Only Memory). In addition, the computer-readable storage medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0071] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA, Programmable Gate Array), a field programmable gate array (FPGA, Field Programmable Gate Array), etc.

[0072] In the description of this specification, the description with reference to the terms "this embodiment", "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0073] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0074] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and simple improvements made to the essential contents of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for image processing, It is characterized in that include: Performing feature extraction processing on the received first image to obtain a plurality of original image features of the first image; Generate an identity feature of the first image using the multiple original features of the images; Determining the feature distance between each of the original features of the image and the identity identification feature, and determining the optimization parameter according to the feature distance, including: sorting the obtained multiple feature distances in an order from small to large or from large to small to obtain a sorting result of the multiple feature distances, and selecting the feature distance at a preset position in the sorting result as the optimization parameter; The optimization parameters are used to update the output result of the loss function used for iterative training of the image recognition model, including: determining a first difference between a current output result of the loss function and the optimization parameters; if the first difference is greater than or equal to zero, updating the output result of the loss function to the first difference; if the first difference is less than zero, updating the output result of the loss function to zero; the trained image recognition model is used to perform image recognition processing on a second image to be recognized.

2. The image processing method according to claim 1, It is characterized in that Determining the feature distance between each of the original features of the image and the identity identification feature comprises: Determine a second difference between a feature value of the original feature of the image and a feature value of the identity feature in each feature dimension; The feature distance between the original feature of the image and the identity identification feature is determined based on the sum of squares of the second difference in all feature dimensions.

3. The image processing method according to claim 1, It is characterized in that The step of generating the identity identification feature of the first image by using the multiple original features of the images comprises: The plurality of original image features are processed by averaging to obtain the identity identification feature.

4. The image processing method according to claim 1, It is characterized in that The output result of the loss function for iterative training of the image recognition model using the optimization parameters includes: Obtaining a first recognition result of the first image by a student network and a second recognition result of the first image by a teacher network; The student network is used to form the image recognition model, and the teacher network is used to perform distillation training on the student network; Processing the first recognition result and the second recognition result by the loss function to obtain a current output result of the loss function; The optimization parameters are used to update the current output result of the loss function.

5. The image processing method according to claim 4, It is characterized in that The performing feature extraction processing on the received first image comprises: receiving the first image processed by data enhancement, and performing feature extraction processing on the first image processed by data enhancement; The obtaining of a first recognition result of the first image by the student network and a second recognition result of the first image by the teacher network comprises: Perform image recognition processing on the first image that has been processed with data enhancement using the student network to obtain a first recognition result; perform image recognition processing on the first image that has been processed with data enhancement using the teacher network to obtain a second recognition result; obtain the first recognition result and the second recognition result.

6. An image processing device, It is characterized in that include: An original feature extraction module, used to perform feature extraction processing on the received first image to obtain a plurality of original image features of the first image; An identification feature generation module, used to generate an identity identification feature of the first image using the multiple original features of the images; An optimization parameter determination module is used to determine the feature distance between each of the original features of the image and the identity identification feature, and determine the optimization parameter according to the feature distance; the optimization parameter determination module is used to sort the obtained multiple feature distances in order from small to large or from large to small to obtain a sorting result of the multiple feature distances; the optimization parameter determination module is used to select the feature distance at a preset position in the sorting result as the optimization parameter; A loss output updating module, used to update the loss function output result used for iterative training of the image recognition model using the optimization parameters, and the trained image recognition model is used to perform image recognition processing on the second image to be recognized; A loss output updating module, used to determine a first difference between a current output result of the loss function and the optimization parameter; The loss output updating module is used to update the loss function output result to the first difference value according to the first difference value being greater than or equal to zero; The loss output updating module is used to update the loss function output result to zero according to the first difference being less than zero.

7. A computer device, It is characterized in that The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the image processing method according to any one of claims 1 to 5.

8. A storage medium storing computer-readable instructions, It is characterized in that When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the image processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Relationship-guided pedestrian attribute recognition method

    CN112733602A

  • Similar face retrieval method, device and storage medium

    US20200250226A1