An attribute recognition method, device, equipment and medium
By performing LAB color space correction processing on the image before image recognition, the problem of low accuracy caused by uncorrected color data in the existing technology is solved, and higher attribute recognition accuracy and stability are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CLOUD TECH CO LTD
- Filing Date
- 2022-08-10
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, attribute recognition methods based on convolutional neural networks do not perform uniform correction on color data, which causes the recognition accuracy to depend on the quality of the image acquisition device, thus reducing the accuracy of attribute recognition.
Before recognition, the image is corrected in the LAB color space. By determining the mean and variance ratio of each pixel, combined with the pre-saved standard reference pixels, the corrected value is calculated and converted to the RGB color space, and then input into the pre-trained attribute recognition model for recognition.
It improves the accuracy of attribute recognition, reduces dependence on the quality of image acquisition equipment, and enhances the stability of recognition.
Smart Images

Figure CN115294394B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an attribute recognition method, apparatus, device, and medium. Background Technology
[0002] Pedestrians in everyday surveillance videos possess rich attribute information, which has wide-ranging applications in image retrieval, smart security, and human-computer interaction. Therefore, human attribute recognition has become a hot research area in computer vision. Currently, improving the accuracy of human attribute recognition, increasing the number of attribute types that can be recognized, and enhancing the stability of attribute recognition in videos have become key research priorities.
[0003] In existing technologies, methods for human attribute recognition are mainly based on general convolutional neural network classification, background suppression, or classification networks improved by loss functions. Although these methods have achieved some results, they have not performed uniform correction for color data, which leads to the color recognition accuracy relying too much on the quality of image data acquired by the image acquisition device, thus reducing the accuracy of attribute recognition.
[0004] Therefore, improving the accuracy of attribute recognition has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides an attribute recognition method, apparatus, device, and medium to solve the problem of low accuracy in attribute recognition in the prior art.
[0006] Firstly, this application provides an attribute identification method, the method comprising:
[0007] Determine the first mean and first variance of each channel of the first image to be recognized in the LAB color space;
[0008] For each pixel in the first image to be identified, the first corrected value of the pixel in each channel of the LAB color space is determined based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the standard control pixel in each channel and the first mean in the pre-saved standard control pixel set.
[0009] Based on the first corrected value of each pixel in each channel of the first image to be identified in the LAB color space, the first target image to be identified in the corrected RGB color space of the first image to be identified is determined.
[0010] The first target image to be identified is input into a pre-trained attribute recognition model, and the attribute values of the first target image to be identified are determined based on the output of the attribute recognition model.
[0011] Secondly, this application provides an attribute recognition device, the device comprising:
[0012] The determination module is used to determine the first mean and first variance of each channel of the acquired first image to be recognized in the LAB color space; and to determine the first corrected value of each pixel in the first image to be recognized in each channel of the LAB color space based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the corresponding standard control pixel in each channel and the first mean in the pre-saved standard control pixel set.
[0013] The correction module determines the first target image to be identified in the RGB color space after correction, based on the first corrected value of each pixel in the first image to be identified in each channel of the LAB color space.
[0014] The recognition module is used to input the first target image to be recognized into a pre-trained attribute recognition model, and determine the attribute value of the first image to be recognized based on the output of the attribute recognition model.
[0015] Thirdly, this application also provides an electronic device, which includes at least a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of any of the attribute recognition methods described above.
[0016] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the attribute recognition methods described above.
[0017] This application provides an attribute recognition method, apparatus, device, and medium. The method determines the first mean and first variance of each channel in the LAB color space of a first image to be recognized. For each pixel in the first image to be recognized, based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference values of the corresponding standard control pixels in each channel and the first mean in a pre-saved set of standard control pixels, a first corrected value corresponding to each channel in the LAB color space is determined. Based on the first corrected value of each pixel in the first image to be recognized corresponding to each channel in the LAB color space, a first target image to be recognized in the corrected RGB color space of the first image to be recognized is determined. This first target image to be recognized is input into a pre-trained attribute recognition model, and the attribute value of the first image to be recognized is determined based on the output of the attribute recognition model. In this embodiment, before performing attribute recognition based on the pre-trained attribute recognition model, for each pixel in the acquired first image to be recognized, the first corrected value corresponding to each channel in the LAB color space is determined based on the ratio of the first variance of the first image to be recognized in each channel to the pre-saved standard control variance, the pre-saved control values of the standard control pixels corresponding to the pixel in each channel, and the first mean of the first image to be recognized in each channel in the LAB color space. This achieves the correction of each pixel in the first image to be recognized, and based on the first corrected value of each pixel in the first image to be recognized in each channel in the LAB color space, the first target image to be recognized in the RGB color space is obtained. Attribute recognition is then performed based on the corrected first target image to be recognized, effectively improving the accuracy of attribute recognition. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the attribute recognition process provided in an embodiment of this application;
[0020] Figure 2 A schematic diagram illustrating the attribute recognition process of the attribute recognition model provided in this application embodiment;
[0021] Figure 3 This is a schematic diagram illustrating another attribute recognition process provided in an embodiment of this application;
[0022] Figure 4 This is a schematic diagram of the structure of the attribute recognition device provided in the embodiments of this application;
[0023] Figure 5 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art are within the scope of protection of this application.
[0025] This application provides an attribute recognition method, apparatus, device, and medium. The method determines the first mean and first variance of each channel in the LAB color space of a first image to be recognized. For each pixel in the first image to be recognized, based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference values of the corresponding standard control pixels in each channel and the first mean in a pre-saved set of standard control pixels, a first corrected value corresponding to each channel in the LAB color space is determined. Based on the first corrected value of each pixel in the first image to be recognized corresponding to each channel in the LAB color space, a first target image to be recognized in the corrected RGB color space of the first image to be recognized is determined. This first target image to be recognized is input into a pre-trained attribute recognition model, and the attribute value of the first image to be recognized is determined based on the output of the attribute recognition model.
[0026] Example 1:
[0027] Figure 1 This is a schematic diagram of the attribute recognition process provided in an embodiment of this application. The process specifically includes the following steps:
[0028] S101: Determine the first mean and first variance of each channel of the first image to be identified in the LAB color space.
[0029] The attribute recognition process provided in this application is applicable to electronic devices, such as servers, PCs, image acquisition devices, etc.
[0030] Image acquisition devices are susceptible to interference from the external environment when acquiring images. For example, when the ambient light intensity is high, the image acquired by the device will be relatively bright. If attribute recognition is directly performed on this image, the accuracy of attribute recognition will be low due to the poor image quality. Therefore, to improve the accuracy of attribute recognition, in this embodiment, the electronic device acquires the first image to be recognized from the image acquisition device and performs uniform correction on this first image to avoid the problem of the image quality of the first image affecting the accuracy of attribute recognition.
[0031] In this embodiment, the electronic device can acquire a first image to be identified and determine the first mean and first variance of each channel of the first image to be identified in the LAB color space. The first image to be identified can be an image acquired by an image acquisition device connected to the electronic device, a frame image corresponding to a video frame acquired by the same device, or an image input by the user of the electronic device. If the electronic device is an image acquisition device, then the first image to be identified is an image acquired by the electronic device itself.
[0032] When determining the first mean and first variance of each channel of the first image to be recognized in the LAB color space, the value of each pixel of the first image to be recognized in the L channel of the LAB color space is obtained, and the first mean of the first image to be recognized in the L channel is determined based on the value of each pixel in the L channel. and first variance Obtain the value of each pixel in the LAB color space corresponding to the A channel of the first image to be recognized, and determine the first mean value of the first image to be recognized in the A channel based on the value of each pixel in the A channel. and first variance Obtain the value of each pixel in the B channel of the first image to be recognized in the LAB color space, and determine the first mean value of the first image to be recognized in the B channel based on the value of each pixel in the B channel. and first variance
[0033] S102: For each pixel in the first image to be identified, based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the standard control pixel corresponding to the pixel in each channel and the first mean in the pre-saved standard control pixel set, determine the first corrected value of the pixel in each channel in the LAB color space.
[0034] In this embodiment of the application, when correcting the first image to be identified, the first corrected value of each pixel in the first image to be identified can be determined for each channel in the LAB color space.
[0035] Specifically, for each pixel in the first image to be identified, and for each channel in the LAB color space, the ratio of the first variance of the pixel in that channel to the pre-saved standard control variance is first determined. Then, based on the reference value of the standard control pixel in that channel in the pre-saved set of standard control pixels and the first mean of the first image to be identified in that channel, the first corrected value of the pixel in that channel is determined.
[0036] In this embodiment, a standard reference variance can be pre-stored for each channel in the LAB color space. This standard reference variance can be determined by the operator based on a reference image with good shooting results. In this embodiment, the standard reference pixel set stores the reference values for multiple standard reference pixels corresponding to each channel in the LAB color space. That is, for each standard reference pixel in the standard reference pixel set, three standard reference values are stored, namely the reference values corresponding to the L channel, A channel, and B channel in the LAB color space. The number of standard reference pixels stored in the standard reference pixel set can be consistent with the number of pixels in the first image to be identified. To make the attribute recognition method applicable to first images to be identified with different pixel sizes, in this embodiment, multiple standard reference pixel sets can be pre-stored according to the possible pixel sizes of the first image to be identified. Before correcting the first image to be identified, it can be determined which pre-stored standard reference pixel set will be used to correct the first image to be identified based on the pixel size of the first image to be identified.
[0037] Specifically, in this embodiment, when determining the first corrected value of any pixel in the first image to be recognized in the L channel, the first variance of the first image to be recognized in the L channel under the LAB color space can be determined. variance compared with pre-saved standard control For ease of description, this ratio can be expressed as: In this embodiment of the application, the first product of the first preset parameter c1 and the ratio can be determined. After determining the first product, we can also determine the first sum of the first product and the second preset parameter c2. Based on the determined first sum, the corresponding standard reference pixel in the L channel of the pixel in the pre-saved set of standard reference pixels is compared with the first sum. The second product The second product is combined with the first mean of the L channel of the first image to be identified in the LAB color space. The second sum The value is determined as the first corrected value of the pixel in the L channel of the LAB color space. The method for determining the first corrected values of the pixel in the A and B channels is similar to the method for determining the first corrected value of the pixel in the L channel. The first corrected value of the pixel in the A channel... It can be represented as:
[0038]
[0039] Where c1 represents the first preset parameter, This represents the first variance of the first image to be identified in channel A. c2 represents the pre-saved standard control variance of channel A in the LAB color space, and c2 represents the second preset parameter. This represents the reference value in channel A of the standard reference pixel corresponding to this pixel within a pre-saved set of standard reference pixels. This represents the first mean value of the first image to be identified in channel A.
[0040] The first corrected value of this pixel in the B channel. It can be represented as:
[0041]
[0042] Where c1 represents the first preset parameter, This represents the first variance of the first image to be identified in the B channel. c2 represents the pre-saved standard control variance of the B channel in the LAB color space, and c2 represents the second preset parameter. This represents the B-channel reference value of the corresponding standard reference pixel in a pre-saved set of standard reference pixels. This represents the first mean value of the first image to be identified in the B channel.
[0043] In this embodiment, the sum of the first preset parameter c1 and the second preset parameter c2 is 1, i.e., c1 + c2 = 1. In this embodiment, c1 and c2 can be used to control the color depth and brightness of the corrected image. Those skilled in the art can set them as needed. In this embodiment, the default is c1 = 0.5 and c2 = 0.5.
[0044] S103: Based on the first corrected value of each pixel in the first image to be identified corresponding to each channel in the LAB color space, determine the first target image to be identified in the corrected RGB color space of the first image to be identified.
[0045] After determining the first corrected value of each pixel in each channel of the first image to be identified in the LAB color space, in this embodiment of the application, the first corrected value of each pixel in each channel of the LAB color space can be converted into the corresponding value in the RGB color space, thereby obtaining the first target image to be identified in the RGB color space after correction.
[0046] S104: Input the first target image to be identified into the pre-trained attribute recognition model, and determine the attribute value of the first image to be identified based on the output of the attribute recognition model.
[0047] After determining the first target image to be identified in the RGB color space after correction, this first target image can be input into a pre-trained attribute recognition model for attribute recognition. The attribute values of the identified first target image are then determined based on the output of the attribute recognition model. The attribute values of the first target image can be the probability of identifying a certain preset attribute. For example, the attribute recognition model can output the probability of each preset color corresponding to the color of a person's upper body clothing in the identified first target image. Alternatively, the attribute values of the first target image can be the specific color of the person's upper body clothing identified by the attribute recognition model.
[0048] In this embodiment, before performing attribute recognition based on the pre-trained attribute recognition model, for each pixel in the acquired first image to be recognized, the first corrected value corresponding to each channel in the LAB color space is determined based on the ratio of the first variance of the first image to be recognized in each channel to the pre-saved standard control variance, the pre-saved control values of the standard control pixels corresponding to the pixel in each channel, and the first mean of the first image to be recognized in each channel in the LAB color space. This achieves the correction of each pixel in the first image to be recognized, and based on the first corrected value of each pixel in the first image to be recognized in each channel in the LAB color space, the first target image to be recognized in the RGB color space is obtained. Attribute recognition is then performed based on the corrected first target image to be recognized, effectively improving the accuracy of attribute recognition.
[0049] Example 2:
[0050] To further improve the accuracy of attribute recognition, based on the above embodiments, in this embodiment, the process of determining the reference value of the standard reference pixel in each channel of the standard reference pixel set includes:
[0051] Each first standard control image in the first sample set is converted to the LAB color space to obtain each second standard control image;
[0052] Based on each of the second standard control images in the first sample set, determine the first control average value of each channel of the first sample set in the LAB color space;
[0053] For each pixel, based on the value of each channel corresponding to the pixel in each of the second standard reference images, the second reference average value corresponding to each channel in the LAB color space is determined; and for each channel in the LAB color space, the difference between the second reference average value corresponding to the channel and the first reference average value corresponding to the channel is calculated, and the difference is used as the reference value of the standard reference pixel corresponding to the pixel in the standard reference pixel set for that channel.
[0054] To further improve the accuracy of attribute recognition, when determining the reference value of the standard reference pixel in each channel, it can also be determined based on multiple images. In this embodiment, it can be determined based on a first sample set. The first sample set stores multiple pre-captured first standard reference images. Each first standard reference image in the first sample set can be a pre-selected image with distinct colors. The image with distinct colors can be understood as one that does not have problems such as overexposure or ghosting. Furthermore, the size of each first standard reference image in the first sample set is equal.
[0055] In this embodiment, each first standard control image in the first sample set can be converted to the LAB color space to obtain each second standard control image. Based on each second standard control image in the first sample set, the first control average value of the first sample set in each channel of the LAB color space is determined.
[0056] Specifically, the first sample set G = {g1, g2, ..., g...} n The first sample set G contains n first standard control images g. These n first standard control images can be converted to the LAB color space to obtain n second standard control images. Assuming n = 10, and each second standard control image contains 50 pixels, then the average value of the first control images in the L channel of the first sample set G in the LAB color space is determined. At that time, the values of 50 pixels in the L channel of 10 second standard control images can be obtained, that is, the values of 50*10 L channels can be obtained. The average value of the obtained 500 L channel values can be determined, and this average value is the first control average value of the first sample set G in the L channel.
[0057] After determining the first control average value of each channel in the LAB color space for the first sample set, in this embodiment of the application, the second control average value of each pixel in the LAB color space can be determined based on the value of each channel corresponding to the pixel in each second standard control image.
[0058] Specifically, assuming the first sample set G includes n second standard control images, the value of each channel corresponding to pixel 1 in the n second standard control images can be obtained, and the average value of the second control corresponding to each channel in the LAB color space can be determined based on the value of each channel.
[0059] For ease of description, the second average value of the L channel for each pixel in the LAB color space can be represented as L. G The average value of the second control group corresponding to channel A can be expressed as A. G The second control average value corresponding to channel B can be expressed as B. G .
[0060] In this embodiment of the application, the second comparative average value L of the L channel of each pixel in the LAB color space is... G The process of determining can be expressed as:
[0061]
[0062] in, This represents the value of the pixel in the L channel in the j-th second standard control image in the first sample set G.
[0063] The second control average value A of the first sample set G in the LAB color space for channel A. G The process of determining can be expressed as:
[0064]
[0065] in, This represents the value of the pixel in channel A in the j-th second standard control image in the first sample set G.
[0066] The second control average value B of the first sample set G in the LAB color space. G The process of determining can be expressed as:
[0067]
[0068] in, This represents the value of the pixel in the B channel in the j-th second standard control image in the first sample set G.
[0069] After determining the second reference average value corresponding to each channel of the pixel in the LAB color space, the difference between the second reference average value corresponding to each channel in each LAB color space and the first reference average value corresponding to that channel can be calculated. This difference can be used as the reference value of the standard reference pixel corresponding to the pixel in the standard reference pixel set for that channel.
[0070] Specifically, for the L channel in the LAB color space, the second reference average value L can be calculated for that pixel in the L channel. G The first control mean of the first sample set in the L channel The difference is used as the reference value in the L channel for the corresponding standard reference pixel in the standard reference pixel set. For ease of description, The process of determining can be expressed as:
[0071]
[0072] Among them, L G This represents the second average value of the corresponding L channel for that pixel. This represents the first control mean of the first sample set in channel L.
[0073] Determine the reference value in channel A of the reference pixel corresponding to the standard reference pixel in the standard reference pixel set. The process can be represented as:
[0074]
[0075] Among them, A G This represents the second average value of the corresponding pixel in channel A. This represents the first control mean of the first sample set in channel A.
[0076] Determine the standard control pixel set and the corresponding standard control pixel value in the B channel. The process can be represented as:
[0077]
[0078] Among them, B GThis represents the second average value of the corresponding B channel for that pixel. This represents the first control mean of the first sample set in channel B.
[0079] To further improve the accuracy of attribute recognition, based on the above embodiments, the method in this application embodiment further includes:
[0080] Based on each of the second standard control images in the first sample set, determine the first control variance for each channel of the first sample set in the LAB color space, and update the standard control variance using the first control variance.
[0081] To further improve the accuracy of attribute recognition, the first control variance can be determined based on each second standard control image in the first sample set.
[0082] Specifically, the first sample set G = {g1, g2, ..., g...} n The dataset contains n second-standard control images, assuming n = 10, and each second-standard control image contains 50 pixels. The first control variance of the L channel of the first sample set G in the LAB color space is determined. At that time, the values of 50 pixels in the L channel of 10 second standard control images can be obtained, that is, the values of 50*10 L channels can be obtained. The variance of the obtained 500 L channel values can be determined, and this variance is the first control variance of the first sample set G in the L channel.
[0083] Specifically, in this embodiment, after determining the first reference variance of each channel of the first sample set in the LAB color space, the pre-saved standard reference variance of the corresponding channel can be updated using the determined first reference variance of each channel of the first sample set in the LAB color space. For example, the determined reference variance A1 of channel L can be used to update the pre-saved standard reference variance for channel L. Update, even
[0084] Example 3:
[0085] To further improve the accuracy of attribute recognition, based on the above embodiments, in this embodiment, the training process of the attribute recognition model includes:
[0086] For each sample image in the second sample set, each sample image has a corresponding label, which is used to identify the target attribute value corresponding to each preset attribute of the human image contained in the sample image;
[0087] The sample image is input into a pre-trained feature extraction network. The feature extraction layer of the feature extraction network extracts the feature image of the sample image and inputs the feature image into a feature pyramid network. The feature pyramid network extracts each first image feature vector of the feature image and selects any first image feature vector to input into a projection layer. The projection layer performs convolution processing on the input first image feature vector to obtain a mapped second image feature vector, and inputs the second image feature vector into a convolution layer. The convolution layer performs dilated convolution processing on the second image feature vector to obtain a third image feature vector, and performs fusion processing on the second image feature vector and the third image feature vector to obtain a fused fourth image feature vector, and inputs the fourth image feature vector into a pooling layer. The pooling layer performs average pooling processing on the fourth image feature vector to obtain each pooled fifth image feature vector, and inputs each fifth image feature vector into a fully connected layer. The fully connected layer determines the attribute value corresponding to each preset attribute of the sample image based on each fifth image feature vector.
[0088] Based on the attribute value corresponding to each preset attribute and the target attribute value corresponding to each preset attribute, the loss value corresponding to the sample image is determined, and the attribute recognition model is adjusted based on the loss value corresponding to the sample image.
[0089] To obtain an attribute recognition model with high accuracy, a second sample set is pre-configured in this embodiment. This second sample set contains multiple sample images, and the attribute recognition model can be trained based on each sample image in the second sample set. In this embodiment, in addition to the sample images, the second sample set also contains labels corresponding to each sample image. These labels are used to identify the target attribute value corresponding to each preset attribute of the person in the sample image. The preset attributes can be attributes such as the gender, age, whether the person is wearing a hat, hairstyle, whether they are wearing glasses, the color of their upper body clothing, and the color of their lower body clothing. The target attribute value corresponding to each attribute can be a specific attribute value. For example, the target attribute value for gender is female, the attribute value for age can be childhood, adolescence, youth, middle age, old age, etc., the attribute value for age can also be an age range, such as 0-12 years old, 12-20 years old, 20-40 years old, etc., and the target attribute value for upper body clothing color is red, etc.
[0090] In this embodiment of the application, for each attribute, the possible attribute values of each attribute can be saved in advance. When annotating the sample image, the target attribute value corresponding to the sample image can be selected from the possible attribute values of each attribute saved in advance, and the target attribute value can be used for annotation. For example, for the attribute of upper body clothing color, the possible attribute values of this attribute may be red, orange, yellow, green, blue, and white. Assuming that the upper body clothing color of the person in the sample image is white, when annotating the sample image, it is only necessary to mark the white attribute value corresponding to the upper body clothing color attribute as 1, and mark the other attribute values as 0.
[0091] To further improve the accuracy of the attribute recognition model, in this embodiment, the labels of the sample images can be further divided. For example, based on the inherent characteristics of human attributes, the labels can be divided into global attributes, local attributes, body structure attributes, and color attributes. Global attributes may include age, behavior, clothing style, etc.; local attributes may be attributes corresponding to the human head, such as whether a hat is worn, whether glasses are worn, hairstyle, etc.; body structure attributes may include upper body clothing, lower body clothing, shoe type, whether socks are worn, etc.; and color attributes may include upper body color, lower body color, and hair color, etc.
[0092] In order to train the attribute recognition model, in this embodiment of the application, after obtaining the second sample set, each sample image in the second sample set can be sequentially input into the feature extraction network in the original attribute recognition model, and the feature extraction layer of the feature extraction network extracts the feature image of the input sample image.
[0093] Specifically, the feature extraction network can be a backbone network TResNet-M. The feature extraction layer of the TResNet-M network can encode the sample image at different scales and output feature images of different sizes. In this embodiment, assuming the resolution of the sample image is 256x192, the feature extraction layers of the TResNet-M network extract feature images of different sizes of the sample image, namely feature1, feature2, and feature3, where feature1 has a size dimension of 8x6x256, feature2 has a size dimension of 16x12x256, and feature3 has a size dimension of 32x24x256.
[0094] In this embodiment, the feature extraction layer of the feature extraction network can input each extracted feature image into the Feature Pyramid Network (FPN), and the Feature Pyramid Network extracts each first image feature vector of the feature image.
[0095] In this embodiment, to balance the timeliness and efficiency of attribute recognition, attribute recognition can be performed based on a Single-Input Single-Output (SiSo) mechanism. This involves selecting any first image feature vector and inputting it into the projection layer of the feature pyramid network. The projection layer can then perform convolution processing on the input first image feature vector to obtain a mapped second image feature vector. Specifically, in this embodiment, the projection layer can perform convolution processing on the input first image feature vector using 1×1 or 3×3 convolution kernels.
[0096] To increase the receptive field of the attribute recognition model, in this embodiment, the projection layer can input the mapped second image feature vector into the convolutional layer. The convolutional layer performs dilated convolution on the obtained second image feature vector to obtain a third image feature vector. Specifically, in this embodiment, dilated convolution can be performed on the obtained second image feature vector using four 1×1, 3×3, and 1×1 convolutional kernels. After determining the third image feature vector after dilated convolution, the second and third image feature vectors can be fused to obtain a fused and enhanced fourth image feature vector.
[0097] In this embodiment, the convolutional layer can input the fourth image feature vector into the pooling layer, and the pooling layer can perform average pooling on the received fourth image feature vector to determine each pooled fifth image feature vector, wherein the fifth image feature vector can be the pooled feature vector corresponding to each preset attribute.
[0098] In this embodiment, the pooling layer can input each fifth image feature vector into the fully connected layer. The fully connected layer classifies each received fifth image feature vector to obtain the attribute value corresponding to each preset attribute of the sample image and outputs it. Specifically, in this embodiment, each received fifth image feature vector can be classified based on the weighted cross-entropy loss function to obtain the attribute value corresponding to each preset attribute of the sample image.
[0099] When training the attribute recognition model, a sample image is input into the model, and the model outputs the attribute value corresponding to each preset attribute of the sample image. Since the target attribute value corresponding to each preset attribute of the sample image is known, the loss value corresponding to the sample image can be determined based on the target attribute value and the recognized attribute value, and the parameters of the attribute recognition model can be adjusted based on the determined loss value.
[0100] In the embodiments of this application, a convergence condition is preset. The convergence condition may be that the number of times the attribute value corresponding to each preset attribute obtained by training the attribute recognition model on the sample images in the second sample set is consistent with the target attribute value corresponding to each preset attribute is greater than a preset number; or it may be that the number of iterations of the attribute recognition model training reaches the set maximum number of iterations, etc. The specific embodiments of this application do not limit this.
[0101] The following describes the attribute recognition process of the attribute recognition model using a specific embodiment. Figure 2 A schematic diagram illustrating the attribute recognition process of the attribute recognition model provided in this application embodiment, as shown below. Figure 2As shown, the first target image containing a person is input into the feature extraction network TResNet-M. The feature extraction layer of the TResNet-M network extracts feature images of the first target image at different dimensions, namely 32x24x256, 16x12x256 and 8x6x256. The TResNet-M network's feature extraction layer inputs each extracted feature image into the feature pyramid network. The feature pyramid network extracts each first image feature vector from the feature image. Based on a one-in-one-out mechanism, any image feature vector extracted from each first image feature vector by the feature pyramid network is selected and input into the projection layer. Specifically, a first image feature vector with dimensions of 8*6*256 can be selected and input into the projection layer. The projection layer performs convolution processing on the received first image feature vector using 1×1 and 3×3 convolution kernels to obtain the mapped second image feature vector. The convolutional layer performs dilation convolution processing on the mapped second image feature vector using four 1×1, 3×3, and 1×1 convolution kernels to obtain the third image feature vector. Finally, the second and third image feature vectors are fused to obtain a feature vector with a receptive field. A pooling layer performs average pooling on the fourth image feature vector to obtain multiple pooled fifth image feature vectors, Task1, Task2, ..., Task5. Task1, Task2, ..., Task5 are fifth image feature vectors corresponding to pre-defined sets of different attribute types. Task1 corresponds to a global attribute set, which may include age, behavior, clothing style, etc.; Task2 corresponds to a local attribute set, which may include whether a hat is worn, whether glasses are worn, hairstyle, etc.; Task3 corresponds to a body structure attribute set, which may include upper body clothing, lower body clothing, etc.; Task4 corresponds to a foot feature set, which may include shoe type, whether socks are worn; and Task5 corresponds to a color attribute set, which may include upper body color, lower body color, hair color, etc. A fully connected layer classifies each received fifth image feature vector to determine and output the attribute value corresponding to each preset attribute of the first target image to be identified.
[0102] In this embodiment of the application, the trained attribute recognition model can identify each attribute in a set of multiple different types of attributes of the first target image to be recognized, and determine the attribute value corresponding to each preset attribute of the first target image to be recognized based on the identified multiple different types of attributes, which effectively improves the accuracy of attribute recognition.
[0103] To further improve the accuracy of the attribute recognition model, based on the above embodiments, in this embodiment, before inputting each sample image in the second sample set into the pre-trained feature extraction network, the method further includes:
[0104] For each sample image in the second sample set, determine the second mean and second variance of each channel in the LAB color space; for each pixel in the sample image, based on the ratio of the second variance of each channel in the LAB color space to the variance of the standard control, and the control values of the corresponding standard control pixels in each channel and the second mean in the standard control pixel set, determine the second corrected value of each channel in the LAB color space for that pixel; based on the second corrected value of each pixel in each channel in the LAB color space for that sample image, determine the target sample image in the corrected RGB color space for that sample image; and update the sample image using the target sample image.
[0105] To further improve the accuracy of the attribute recognition model, each sample image in the second sample set can be corrected during the training of the attribute recognition model.
[0106] In this embodiment, for each sample image in the second sample set, a LAB color space conversion can be performed, that is, the sample image in RGB color space can be converted to a sample image in LAB color space. The second mean and second variance of each channel of the sample image in LAB color space are determined. For each pixel in the sample image, based on the ratio of the second variance of each channel in LAB color space to the variance of the standard control, and the control values of the corresponding standard control pixels in each channel and the second mean in the standard control pixel set, the second corrected value corresponding to the pixel in each channel is determined, and then an RGB color space conversion is performed, that is, the corrected sample image in LAB color space is converted to a target sample image in RGB color space. In this embodiment, the target sample image can be used to update the sample image, that is, the corrected sample image can be used to train the attribute recognition model.
[0107] Since the process of determining the target sample image corresponding to each sample image in the second sample set is similar to the process of determining the target image to be identified corresponding to the first image to be identified in the above embodiments, and has been described in detail in the above embodiments, it will not be repeated in the embodiments of this application.
[0108] Example 4:
[0109] To improve the stability of attribute recognition, based on the above embodiments, in this embodiment of the application, if the first image to be recognized is a frame image in a video, the method further includes:
[0110] Obtain a preset number of second images to be identified in the video that are adjacent to the first image to be identified;
[0111] For each second image to be identified, the third mean and third difference of each channel in the LAB color space are determined. For each pixel in the second image to be identified, based on the ratio of the third difference of each channel in the LAB color space to the standard control variance, and the control values of the corresponding standard control pixels in each channel and the third difference in the pre-saved set of standard control pixels, the third corrected value of each pixel in each channel in the LAB color space is determined. Based on the third corrected value of each pixel in each channel in the LAB color space, the second target image to be identified in the corrected RGB color space is determined. The second target image to be identified is input into the pre-trained attribute recognition model, and the attribute values of the identified second image to be identified are determined based on the output of the attribute recognition model.
[0112] The target attribute value of the first image to be identified is determined based on the attribute value of the first image to be identified, the attribute value of each second image to be identified, and a preset weight.
[0113] Since each frame in a video corresponds to a continuous image, and the content contained in the images is generally related, in order to improve the stability of attribute recognition and thus improve the accuracy of attribute recognition, in this embodiment of the application, when recognizing the attributes of people contained in the first image to be recognized, if the first image to be recognized is a frame image in the video, the target attribute value of the first image to be recognized can be determined by combining the attributes of people contained in other second images to be recognized adjacent to the first image to be recognized.
[0114] In this embodiment of the application, a preset number of second images to be identified that are adjacent to the first image to be identified in the video can be obtained.
[0115] Specifically, in this embodiment, the frame position of the first image to be identified in the video can be determined, and a preset number of other second images to be identified after that frame position can be obtained, or a preset number of other second images to be identified before that frame position can be obtained, or other second images to be identified including that frame position can be obtained. In this embodiment, when obtaining a preset number of second images to be identified adjacent to the first image to be identified in the video, sliding window sampling can be performed in the video sequence of the video, with each sliding window interval being 1, and the size of the window being [width, height, t], where width is the width of the second image to be identified, height is the height of the second image to be identified, and t is a preset number. Starting from the position corresponding to the first image to be identified in the video sequence, the sliding window can be moved t times, and the image content corresponding to the sliding window each time it is moved is the obtained second image to be identified.
[0116] In this embodiment, for each acquired second image to be identified, each pixel of the second image to be identified can be corrected, and a third corrected value for each pixel in each channel can be obtained, thereby determining the second target image to be identified in the corrected RGB color space. The attribute values of the identified second image to be identified are then determined based on the output of the attribute recognition model.
[0117] After determining the attribute values of a preset number of second images to be identified, in this embodiment of the application, the target attribute value of the first image to be identified can be determined comprehensively based on the attribute values of the first image to be identified, the attribute values of each second image to be identified, and the preset weight.
[0118] Specifically, by combining the attribute values of the first image to be recognized and the attribute values of each second image to be recognized, the first attribute value corresponding to any preset attribute of the first image to be recognized can be determined to be A. wci Then the first attribute value A wci The process of determining can be expressed as:
[0119] A wci =W (c-t / 2+1)i A (c-t / 2+1)i +W (c-t / 2+2)i A (c-t / 2+2)i +...+W ci A ci +W (c+1)i A (c+1)i +...+W (c+t / 2) i A (c+t / 2)i
[0120] Where 'c' represents the identifier of the image to be identified, which can be the sorting number of the image in the video or an identifier defined according to any rule; 'i' represents the identifier corresponding to the preset attribute, for example, the identifier corresponding to preset attribute A is 1, and the identifier corresponding to preset attribute B is 2, A ci W represents the attribute value corresponding to attribute i in the c-th image to be recognized. This attribute value can be the probability of recognizing the image corresponding to that attribute. ci The weight of attribute i in the c-th image to be identified is the preset weight, where W (c-t / 2+1)i +W (c-t / 2+2)i +...+W ci +W (c+1)i +...+W (c+t / 2)i =1.
[0121] In this embodiment of the application, after determining the first attribute value corresponding to each preset attribute of the first image to be identified, it is possible to determine whether the first attribute value is greater than a preset threshold thresh for each preset attribute. If so, the target attribute value corresponding to the preset attribute can be determined as the target attribute value of the first image to be identified.
[0122] The attribute recognition process will be explained below with reference to a specific embodiment. Figure 3 This is a schematic diagram of another attribute recognition process provided in an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps:
[0123] S301: Obtain the first image to be recognized, correct the first image to be recognized, and obtain the first target image to be recognized.
[0124] S302: Based on the pre-trained attribute recognition model, perform feature extraction on the first image to be recognized and determine the attribute values of the first image to be recognized.
[0125] S303: Obtain a preset number of second images to be identified that are adjacent to the first image to be identified in the video, and identify the attribute value of each second image to be identified. Based on the attribute value of the first image to be identified and the attribute value of each second image to be identified, determine the target attribute value of the first image to be identified.
[0126] Example 5:
[0127] Figure 4 This is a schematic diagram of the structure of the attribute recognition device provided in the embodiments of this application, as shown below. Figure 4 As shown, the device includes:
[0128] The determining module 401 is used to determine the first mean and first variance of each channel of the acquired first image to be recognized in the LAB color space; for each pixel in the first image to be recognized, based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the standard control pixel corresponding to the pixel in each channel and the first mean in the pre-saved standard control pixel set, the first corrected value of the pixel in each channel in the LAB color space is determined;
[0129] The correction module 402 is used to determine the first target image to be identified in the RGB color space after correction of the first image to be identified based on the first corrected value of each pixel in each channel of the LAB color space.
[0130] The recognition module 403 is used to input the first target image to be recognized into a pre-trained attribute recognition model, and determine the attribute value of the first image to be recognized based on the output of the attribute recognition model.
[0131] In one possible implementation, the determining module 401 is further configured to perform LAB color space conversion on each first standard reference image in the first sample set to obtain each second standard reference image; determine the first reference average value of each channel in the first sample set in the LAB color space based on each second standard reference image in the first sample set; for each pixel, determine the second reference average value corresponding to each channel in the LAB color space based on the value of each channel corresponding to the pixel in each second standard reference image; and for each channel in the LAB color space, calculate the difference between the second reference average value corresponding to the channel and the first reference average value corresponding to the channel, and use the difference as the reference value of the standard reference pixel corresponding to the pixel in the standard reference pixel set in that channel.
[0132] In one possible implementation, the determining module 401 is further configured to determine the control variance of each channel of the first sample set in the LAB color space based on each of the second standard control images in the first sample set, and update the standard control variance using the control variance.
[0133] In one possible implementation, the device further includes:
[0134] Training module 404 is used to train each sample image in the second sample set, wherein each sample image corresponds to a label, the label being used to identify the target attribute value corresponding to each preset attribute of the human image contained in the sample image; inputting the sample image into a pre-trained feature extraction network, the feature extraction layer of the feature extraction network extracts the feature image of the sample image, and inputs the feature image into a feature pyramid network; the feature pyramid network extracts each first image feature vector of the feature image, and selects any first image feature vector to input into a projection layer; the projection layer performs convolution processing on the input first image feature vector to obtain a mapped second image feature vector, and inputs the second image feature vector into a convolutional layer; the convolutional layer processes the second image feature vector... The vectors are subjected to dilated convolution to obtain a third image feature vector. The second and third image feature vectors are then fused to obtain a fused fourth image feature vector, which is then input into a pooling layer. The pooling layer performs average pooling on the fourth image feature vector to obtain each pooled fifth image feature vector, which is then input into a fully connected layer. The fully connected layer determines the attribute value corresponding to each preset attribute of the sample image based on each fifth image feature vector. Based on the attribute value corresponding to each preset attribute and the target attribute value corresponding to each preset attribute, the loss value corresponding to the sample image is determined, and the attribute recognition model is adjusted based on the loss value corresponding to the sample image.
[0135] In one possible implementation, the training module 404 is further configured to: determine, for each sample image in the second sample set, a second mean and a second variance for each channel of the sample image in the LAB color space; for each pixel in the sample image, based on the ratio of the second variance of each channel in the LAB color space to the variance of the standard control, and the control values of the corresponding standard control pixels in each channel and the second mean in the standard control pixel set, determine a second corrected value for each pixel in the LAB color space; based on the second corrected value for each pixel in the sample image in each channel of the LAB color space, determine a target sample image in the corrected RGB color space of the sample image; and update the sample image using the target sample image.
[0136] In one possible implementation, the device further includes:
[0137] The acquisition module 405 is used to acquire a preset number of second images to be identified in the video that are adjacent to the first image to be identified;
[0138] The determining module 401 is further configured to, for each second image to be identified, determine the third mean and third difference of each channel of the second image to be identified in the LAB color space; and for each pixel in the second image to be identified, determine the third corrected value of each channel of the second image to be identified in the LAB color space based on the ratio of the third difference of each channel of the second image to be identified in the LAB color space to the standard control variance, and the control value of the standard control pixel corresponding to the pixel in each channel and the third difference in the pre-saved set of standard control pixels.
[0139] The correction module 402 is further configured to determine the second target image to be identified in the RGB color space after correction based on the third corrected value corresponding to each channel of each pixel point in the LAB color space of the second image to be identified;
[0140] The recognition module 403 is further configured to input the second target image to be recognized into the pre-trained attribute recognition model, and determine the attribute value of the recognized second image to be recognized based on the output of the attribute recognition model;
[0141] The determining module 401 is further configured to determine the target attribute value of the first image to be identified based on the attribute value of the first image to be identified, the attribute value of each of the second images to be identified, and a preset weight.
[0142] Example 6:
[0143] Figure 5 This application provides a schematic diagram of an electronic device structure as an embodiment of the present application. Based on the above embodiments, the present application also provides an electronic device, such as... Figure 5 As shown, it includes: processor 501, communication interface 502, memory 503 and communication bus 504, wherein processor 501, communication interface 502 and memory 503 communicate with each other through communication bus 504.
[0144] The memory 503 stores a computer program, which, when executed by the processor 501, causes the processor 501 to perform the following steps:
[0145] Determine the first mean and first variance of each channel of the first image to be recognized in the LAB color space;
[0146] For each pixel in the first image to be identified, the first corrected value of the pixel in each channel of the LAB color space is determined based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the standard control pixel in each channel and the first mean in the pre-saved standard control pixel set.
[0147] Based on the first corrected value of each pixel in each channel of the first image to be identified in the LAB color space, the first target image to be identified in the corrected RGB color space of the first image to be identified is determined.
[0148] The first target image to be identified is input into a pre-trained attribute recognition model, and the attribute values of the first target image to be identified are determined based on the output of the attribute recognition model.
[0149] In one possible implementation, the process of determining the reference value of the standard reference pixel in each channel of the standard reference pixel set includes:
[0150] Each first standard control image in the first sample set is converted to the LAB color space to obtain each second standard control image;
[0151] Based on each of the second standard control images in the first sample set, determine the first control average value of each channel of the first sample set in the LAB color space;
[0152] For each pixel, based on the value of each channel corresponding to the pixel in each of the second standard reference images, the second reference average value corresponding to each channel in the LAB color space is determined; and for each channel in the LAB color space, the difference between the second reference average value corresponding to the channel and the first reference average value corresponding to the channel is calculated, and the difference is used as the reference value of the standard reference pixel corresponding to the pixel in the standard reference pixel set for that channel.
[0153] In one possible implementation, the method further includes:
[0154] Based on each of the second standard control images in the first sample set, determine the control variance of each channel of the first sample set in the LAB color space, and update the standard control variance using the control variance.
[0155] In one possible implementation, the training process of the attribute recognition model includes:
[0156] For each sample image in the second sample set, each sample image has a corresponding label, which is used to identify the target attribute value corresponding to each preset attribute of the human image contained in the sample image;
[0157] The sample image is input into a pre-trained feature extraction network. The feature extraction layer of the feature extraction network extracts the feature image of the sample image and inputs the feature image into a feature pyramid network. The feature pyramid network extracts each first image feature vector of the feature image and selects any first image feature vector to input into a projection layer. The projection layer performs convolution processing on the input first image feature vector to obtain a mapped second image feature vector, and inputs the second image feature vector into a convolution layer. The convolution layer performs dilated convolution processing on the second image feature vector to obtain a third image feature vector, and performs fusion processing on the second image feature vector and the third image feature vector to obtain a fused fourth image feature vector, and inputs the fourth image feature vector into a pooling layer. The pooling layer performs average pooling processing on the fourth image feature vector to obtain each pooled fifth image feature vector, and inputs each fifth image feature vector into a fully connected layer. The fully connected layer determines the attribute value corresponding to each preset attribute of the sample image based on each fifth image feature vector.
[0158] Based on the attribute value corresponding to each preset attribute and the target attribute value corresponding to each preset attribute, the loss value corresponding to the sample image is determined, and the attribute recognition model is adjusted based on the loss value corresponding to the sample image.
[0159] In one possible implementation, before inputting each sample image in the second sample set into the pre-trained feature extraction network, the method further includes:
[0160] For each sample image in the second sample set, determine the second mean and second variance of each channel in the LAB color space; for each pixel in the sample image, based on the ratio of the second variance of each channel in the LAB color space to the variance of the standard control, and the control values of the corresponding standard control pixels in each channel and the second mean in the standard control pixel set, determine the second corrected value of each pixel in each channel in the LAB color space; based on the second corrected value of each pixel in each channel in the LAB color space, determine the target sample image in the corrected RGB color space of the sample image; and update the sample image using the target sample image.
[0161] In one possible implementation, if the first image to be identified is a frame image in a video, the method further includes:
[0162] Obtain a preset number of second images to be identified in the video that are adjacent to the first image to be identified;
[0163] For each second image to be identified, the third mean and third difference of each channel in the LAB color space are determined. For each pixel in the second image to be identified, based on the ratio of the third difference of each channel in the LAB color space to the standard control variance, and the control values of the corresponding standard control pixels in each channel and the third difference in the pre-saved set of standard control pixels, the third corrected value of each pixel in each channel in the LAB color space is determined. Based on the third corrected value of each pixel in each channel in the LAB color space, the second target image to be identified in the corrected RGB color space is determined. The second target image to be identified is input into the pre-trained attribute recognition model, and the attribute values of the identified second image to be identified are determined based on the output of the attribute recognition model.
[0164] The target attribute value of the first image to be identified is determined based on the attribute value of the first image to be identified, the attribute value of each second image to be identified, and a preset weight.
[0165] Since the principle of solving the problem by the above-mentioned electronic device is similar to that of the attribute recognition method, the implementation of the above-mentioned electronic device can refer to the above embodiments, and the repeated parts will not be described again.
[0166] The communication bus mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. Communication interface 502 is used for communication between the above-mentioned electronic device and other devices. The memory can include random access memory (RAM), or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor. The aforementioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processing unit (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0167] Example 7:
[0168] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by a processor. When the program is run on the processor, the processor executes the following steps:
[0169] Determine the first mean and first variance of each channel of the first image to be recognized in the LAB color space;
[0170] For each pixel in the first image to be identified, the first corrected value of the pixel in each channel of the LAB color space is determined based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the standard control pixel in each channel and the first mean in the pre-saved standard control pixel set.
[0171] Based on the first corrected value of each pixel in each channel of the first image to be identified in the LAB color space, the first target image to be identified in the corrected RGB color space of the first image to be identified is determined.
[0172] The first target image to be identified is input into a pre-trained attribute recognition model, and the attribute values of the first target image to be identified are determined based on the output of the attribute recognition model.
[0173] In one possible implementation, the process of determining the reference value of the standard reference pixel in each channel of the standard reference pixel set includes:
[0174] Each first standard control image in the first sample set is converted to the LAB color space to obtain each second standard control image;
[0175] Based on each of the second standard control images in the first sample set, determine the first control average value of each channel of the first sample set in the LAB color space;
[0176] For each pixel, based on the value of each channel corresponding to the pixel in each of the second standard reference images, the second reference average value corresponding to each channel in the LAB color space is determined; and for each channel in the LAB color space, the difference between the second reference average value corresponding to the channel and the first reference average value corresponding to the channel is calculated, and the difference is used as the reference value of the standard reference pixel corresponding to the pixel in the standard reference pixel set for that channel.
[0177] In one possible implementation, the method further includes:
[0178] Based on each of the second standard control images in the first sample set, determine the control variance of each channel of the first sample set in the LAB color space, and update the standard control variance using the control variance.
[0179] In one possible implementation, the training process of the attribute recognition model includes:
[0180] For each sample image in the second sample set, each sample image has a corresponding label, which is used to identify the target attribute value corresponding to each preset attribute of the human image contained in the sample image;
[0181] The sample image is input into a pre-trained feature extraction network. The feature extraction layer of the feature extraction network extracts the feature image of the sample image and inputs the feature image into a feature pyramid network. The feature pyramid network extracts each first image feature vector of the feature image and selects any first image feature vector to input into a projection layer. The projection layer performs convolution processing on the input first image feature vector to obtain a mapped second image feature vector, and inputs the second image feature vector into a convolution layer. The convolution layer performs dilated convolution processing on the second image feature vector to obtain a third image feature vector, and performs fusion processing on the second image feature vector and the third image feature vector to obtain a fused fourth image feature vector, and inputs the fourth image feature vector into a pooling layer. The pooling layer performs average pooling processing on the fourth image feature vector to obtain each pooled fifth image feature vector, and inputs each fifth image feature vector into a fully connected layer. The fully connected layer determines the attribute value corresponding to each preset attribute of the sample image based on each fifth image feature vector.
[0182] Based on the attribute value corresponding to each preset attribute and the target attribute value corresponding to each preset attribute, the loss value corresponding to the sample image is determined, and the attribute recognition model is adjusted based on the loss value corresponding to the sample image.
[0183] In one possible implementation, before inputting each sample image in the second sample set into the pre-trained feature extraction network, the method further includes:
[0184] For each sample image in the second sample set, determine the second mean and second variance of each channel in the LAB color space; for each pixel in the sample image, based on the ratio of the second variance of each channel in the LAB color space to the variance of the standard control, and the control values of the corresponding standard control pixels in each channel and the second mean in the standard control pixel set, determine the second corrected value of each pixel in each channel in the LAB color space; based on the second corrected value of each pixel in each channel in the LAB color space, determine the target sample image in the corrected RGB color space of the sample image; and update the sample image using the target sample image.
[0185] In one possible implementation, if the first image to be identified is a frame image in a video, the method further includes:
[0186] Obtain a preset number of second images to be identified in the video that are adjacent to the first image to be identified;
[0187] For each second image to be identified, the third mean and third difference of each channel in the LAB color space are determined. For each pixel in the second image to be identified, based on the ratio of the third difference of each channel in the LAB color space to the standard control variance, and the control values of the corresponding standard control pixels in each channel and the third difference in the pre-saved set of standard control pixels, the third corrected value of each pixel in each channel in the LAB color space is determined. Based on the third corrected value of each pixel in each channel in the LAB color space, the second target image to be identified in the corrected RGB color space is determined. The second target image to be identified is input into the pre-trained attribute recognition model, and the attribute values of the identified second image to be identified are determined based on the output of the attribute recognition model.
[0188] The target attribute value of the first image to be identified is determined based on the attribute value of the first image to be identified, the attribute value of each second image to be identified, and a preset weight.
[0189] Since the principle of solving the problem using the computer-readable medium provided above is similar to the attribute identification method, the steps implemented after the processor executes the computer program in the computer-readable medium can be referred to the above embodiments, and repeated parts will not be described again.
[0190] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0191] For system / device embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments.
[0192] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0193] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0194] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0195] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0196] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. An attribute recognition method, characterized in that, The method includes: Determine the first mean and first variance of each channel of the first image to be recognized in the LAB color space; For each pixel in the first image to be identified, the first corrected value of the pixel in each channel of the LAB color space is determined based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the standard control pixel in each channel and the first mean in the pre-saved standard control pixel set. Based on the first corrected value of each pixel in each channel of the first image to be identified in the LAB color space, the first target image to be identified in the corrected RGB color space of the first image to be identified is determined. The first target image to be identified is input into a pre-trained attribute recognition model, and the attribute values of the first target image to be identified are determined based on the output of the attribute recognition model. The process of determining the reference value of the standard reference pixel in each channel of the standard reference pixel set includes: Each first standard control image in the first sample set is converted to the LAB color space to obtain each second standard control image; For each channel in the LAB color space, obtain the value of each pixel in each second standard reference image in that channel; based on each obtained value, determine the first reference average value of the first sample set in that channel; For each pixel, based on the value of each channel corresponding to the pixel in each of the second standard reference images, the second reference average value corresponding to each channel in the LAB color space is determined; and for each channel in the LAB color space, the difference between the second reference average value corresponding to the channel and the first reference average value corresponding to the channel is calculated, and the difference is used as the reference value of the standard reference pixel corresponding to the pixel in the standard reference pixel set for that channel.
2. The method as described in claim 1, characterized in that, The method further includes: Based on each of the second standard control images in the first sample set, determine the control variance of each channel of the first sample set in the LAB color space, and update the standard control variance using the control variance.
3. The method as described in claim 1, characterized in that, The training process of the attribute recognition model includes: For each sample image in the second sample set, each sample image has a corresponding label, which is used to identify the target attribute value corresponding to each preset attribute of the human image contained in the sample image; The sample image is input into a pre-trained feature extraction network. The feature extraction layer of the feature extraction network extracts the feature image of the sample image and inputs the feature image into a feature pyramid network. The feature pyramid network extracts each first image feature vector of the feature image and selects any first image feature vector to input into a projection layer. The projection layer performs convolution processing on the input first image feature vector to obtain a mapped second image feature vector, and inputs the second image feature vector into a convolution layer. The convolution layer performs dilated convolution processing on the second image feature vector to obtain a third image feature vector, and performs fusion processing on the second image feature vector and the third image feature vector to obtain a fused fourth image feature vector, and inputs the fourth image feature vector into a pooling layer. The pooling layer performs average pooling processing on the fourth image feature vector to obtain each pooled fifth image feature vector, and inputs each fifth image feature vector into a fully connected layer. The fully connected layer determines the attribute value corresponding to each preset attribute of the sample image based on each fifth image feature vector. Based on the attribute value corresponding to each preset attribute and the target attribute value corresponding to each preset attribute, the loss value corresponding to the sample image is determined, and the attribute recognition model is adjusted based on the loss value corresponding to the sample image.
4. The method as described in claim 3, characterized in that, Before inputting each sample image in the second sample set into the pre-trained feature extraction network, the method further includes: For each sample image in the second sample set, determine the second mean and second variance of each channel in the LAB color space; for each pixel in the sample image, based on the ratio of the second variance of each channel in the LAB color space to the variance of the standard control, and the control values of the corresponding standard control pixels in each channel and the second mean in the standard control pixel set, determine the second corrected value of each pixel in each channel in the LAB color space; based on the second corrected value of each pixel in each channel in the LAB color space, determine the target sample image in the corrected RGB color space of the sample image; and update the sample image using the target sample image.
5. The method as described in claim 1, characterized in that, If the first image to be identified is a frame image in a video, the method further includes: Obtain a preset number of second images to be identified in the video that are adjacent to the first image to be identified; For each second image to be identified, the third mean and third difference of each channel in the LAB color space are determined. For each pixel in the second image to be identified, based on the ratio of the third difference of each channel in the LAB color space to the standard control variance, and the control values of the corresponding standard control pixels in each channel and the third difference in the pre-saved set of standard control pixels, the third corrected value of each pixel in each channel in the LAB color space is determined. Based on the third corrected value of each pixel in each channel in the LAB color space, the second target image to be identified in the corrected RGB color space is determined. The second target image to be identified is input into the pre-trained attribute recognition model, and the attribute values of the identified second image to be identified are determined based on the output of the attribute recognition model. The target attribute value of the first image to be identified is determined based on the attribute value of the first image to be identified, the attribute value of each second image to be identified, and a preset weight.
6. An attribute recognition device, characterized in that, The device includes: The determination module is used to determine the first mean and first variance of each channel of the acquired first image to be recognized in the LAB color space; and to determine the first corrected value of each pixel in the first image to be recognized in each channel of the LAB color space based on the ratio of the first variance of each channel in the LAB color space to the pre-saved standard control variance, and the reference value of the corresponding standard control pixel in each channel and the first mean in the pre-saved standard control pixel set. The correction module determines the first target image to be identified in the RGB color space after correction, based on the first corrected value of each pixel in the first image to be identified in each channel of the LAB color space. The recognition module is used to input the first target image to be recognized into a pre-trained attribute recognition model, and determine the attribute value of the first image to be recognized based on the output of the attribute recognition model. The determining module is further configured to perform LAB color space conversion on each first standard reference image in the first sample set to obtain each second standard reference image; for each channel in the LAB color space, obtain the value of each pixel in each second standard reference image in that channel; determine the first reference average value of the first sample set in that channel based on the obtained value; for each pixel, determine the second reference average value corresponding to each channel in the LAB color space based on the value of the pixel in each second standard reference image in each channel; and for each channel in the LAB color space, calculate the difference between the second reference average value corresponding to that channel and the first reference average value corresponding to that channel, and use the difference as the reference value of the standard reference pixel corresponding to that pixel in the standard reference pixel set in that channel.
7. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of the attribute recognition method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the attribute recognition method according to any one of claims 1-5.