A method for identifying target attributes, electronic device and computer-readable storage medium
By dividing the image area into multiple image blocks in image attribute recognition and enhancing processing based on the characteristics of these image blocks, the problem of interference information influence in image attribute recognition is solved, and the accuracy of recognition is improved.
Patent Information
- Application Number
- CN202510295811.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-13
AI Technical Summary
When the prior art recognizes attributes through images, it is susceptible to interference information, resulting in low recognition accuracy.
By obtaining the target image area where the target object in the target image is located, and dividing the area into multiple image blocks. The spatial correlation features, label information and global image features of each image block are enhanced to obtain the target image features. The target object is then attributed based on these features.
By introducing spatial correlation features and label information of image blocks, the target image features pay more attention to effective information, thereby improving the accuracy of target attribute recognition.
Smart Images

Figure CN119810897B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a target attribute recognition method, an electronic device and a computer-readable storage medium. Background Art
[0002] Attribute recognition is an important task in the field of computer science, which aims to automatically detect and understand the target attributes of individuals by analyzing information from different sources (such as text, sound, image, etc.). Among them, when performing attribute recognition through images, the target attributes of the target object are mainly obtained through image features, but there is a lot of interference information in the image, which may affect the accuracy of attribute recognition.
[0003] Based on this, how to reduce the impact of interference information and improve the accuracy of attribute recognition has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] The main technical problem solved by the present application is to provide a target attribute recognition method, an electronic device and a computer-readable storage medium, which can improve the recognition accuracy of target attributes.
[0005] To solve the above technical problems, a technical solution adopted in the present application is: a target attribute recognition method is provided, and the target attribute recognition method includes: obtaining a target image area where a target object is located in a target image, and the target image area includes at least one image block; enhancing the image blocks according to spatial correlation features, label information and global image features of each image block to obtain target image features of each image block; performing attribute recognition processing on the target object based on the target image features of each image block to obtain the target attributes of the target object.
[0006] In one embodiment, the step of obtaining a target image area in a target image where a target object is located, wherein the target image area includes at least one image block, comprises: performing recognition processing on the target image to obtain an initial image area of the target object in the target image; dividing the initial image area according to a preset ratio to obtain a first image sub-area and a second image sub-area; and determining the target image area based on the first image sub-area and the second image sub-area.
[0007] In one embodiment, the image block includes a first image block and a second image block, the target image feature includes a target image feature of the first image block and a target image feature of the second image block, and the step of performing attribute recognition processing on the target object based on the target image features of each image block to obtain the target attribute of the target object includes: performing splicing processing on the target image feature of the first image block and the target image feature of the second image block to obtain a spliced feature; and performing attribute recognition processing on the target object according to the spliced feature to obtain the target attribute of the target object.
[0008] In one embodiment, the step of performing enhancement processing on the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block includes: performing full connection processing on the spatial correlation features and label information of each image block to obtain the label space correlation features of each image block; and determining the target image features of each image block based on the label space correlation features of each image block and the global image features of the corresponding image block.
[0009] In one embodiment, the step of determining the target image features of each image block based on the label space association features of each image block and the global image features of the corresponding image block includes: decoding the label space association features of each image block and the global image features of the corresponding image block to obtain the image features of each image block after decoding; and performing feature extraction on the image features of each image block after decoding to obtain the target image features of each image block.
[0010] In one embodiment, the step of decoding the label space association features of each image block and the global image features of the corresponding image block to obtain the image features of each image block after decoding includes: determining a query vector, a key vector and a value vector according to the global image features of the image block; obtaining the vector similarity between the query vector and the key vector; enhancing the corresponding vector similarity according to the label space association features of the image block to obtain the target vector similarity; weighting the value vector based on the target vector similarity to obtain an attention value; and determining the sum of the attention value and the global image feature as the image feature of the image block after decoding.
[0011] In one embodiment, before the step of performing enhancement processing on the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block, the method further includes: determining the spatial correlation features of each image block according to the spatial relationship between each image block and other image blocks; and performing feature extraction processing on each image block to obtain the global image features of each image block.
[0012] In one embodiment, the step of performing feature extraction processing on each image block to obtain the global image features of each image block includes: determining the product of the height of each image block and the width of the corresponding image block as the number of pixel rows of the corresponding image block; serializing each pixel point in the image block according to the number of pixel rows of each image block to obtain the serialized image block; and performing full connection processing on the serialized image block to obtain the global image features of the corresponding image block.
[0013] To solve the above technical problems, another technical solution adopted in the present application is: to provide an electronic device, including a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the above target attribute recognition method.
[0014] In order to solve the above technical problems, another technical solution adopted by the present application is: providing a computer-readable storage medium including program data stored therein, wherein the program data is used to implement the above target attribute identification method when executed by a processor.
[0015] The above scheme obtains the target image region where the target object is located in the target image, and the target image region includes at least one image block; the image blocks are enhanced according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block; the target object is attributed based on the target image features of each image block to obtain the target attribute of the target object. Thus, the spatial correlation features and label information of the image blocks are introduced, so that the target image features pay more attention to the effective information in the target image region, thereby improving the accuracy of target attribute recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work, among which:
[0017] Figure 1 is a flowchart of an exemplary embodiment of a target attribute recognition method shown in the present application;
[0018] Figure 2 It is a schematic diagram of a framework of an exemplary embodiment of a target attribute recognition method shown in the present application;
[0019] Figure 3 yes Figure 2 A schematic diagram of a framework of an exemplary embodiment of a grid enhancement module shown in FIG.
[0020] Figure 4 is a schematic structural diagram of an exemplary embodiment of a target attribute recognition device shown in the present application;
[0021] Figure 5 It is a structural schematic diagram of an embodiment of an electronic device provided by the present application;
[0022] Figure 6 It is a structural schematic diagram of an embodiment of a computer-readable storage medium provided by the present application. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be appreciated that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some but not all structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the art without making creative work are within the scope of protection of the present application.
[0024] First of all, it should be noted that the extraction of image features has an important impact on the identification of accurate target attributes. When the acquired image features are incomplete or there is too much interference information, there will be errors between the identified target attributes and the actual attributes of the target object. Based on this, the embodiment of the present application adds spatial correlation features and label information of image blocks as auxiliary information, so that the enhanced target image features can pay more attention to the effective area in the target image and improve the recognition accuracy of the target attributes.
[0025] For details, please refer to Figure 1 , Figure 1 It is a flowchart of an exemplary embodiment of a target attribute recognition method shown in the present application.
[0026] The target attribute identification method may be executed by a terminal device or a server or other processing device, wherein the terminal device may be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The target attribute identification method may also be executed by a target attribute identification device. In some possible implementations, the target attribute identification method may be implemented by a processor calling a computer-readable instruction stored in a memory.
[0027] In the embodiment of the present application, the target attribute recognition device is used as the execution subject for description. Specifically, the target attribute recognition method of this embodiment includes the following steps:
[0028] S110: Acquire a target image region where a target object is located in a target image, wherein the target image region includes at least one image block.
[0029] The target image may be an image selected from the image set, or an image captured by an image acquisition device, or a video frame extracted from a video stream captured by a video acquisition device. For example, it may be any image in the image set, or an image with the best image quality in the image set, which is not limited in the embodiments of the present application.
[0030] The target object in the target image refers to an object instance or a region of interest in the target image. Exemplarily, the target object may be an animal, a vehicle, or a person, etc. Exemplarily, the target object in the target image may be detected by a target detection algorithm, wherein the target detection algorithm may include but is not limited to YOLO-V5, YOLO-V4, YOLO-V7, PP-YOLOv2, etc., for accurately identifying and locating the target object from an image or video.
[0031] The target image region may be an image region where the target object is located in the target image. Exemplarily, a target frame obtained after a target detection algorithm detects the target object in the target image may be used as the target image region, or the detected target frame may be optimized to obtain the target image region. Specifically, the target attribute recognition device may detect the target region in the target object to obtain the target image region, such as using a facial region or a body region as the target image region.
[0032] An image block is a small block into which the target image area is cut. For example, the image block can be obtained by a grid, a sliding window, or a translation window. Each image block can be processed and analyzed independently.
[0033] The target attribute recognition device obtains a target image from an acquisition device, an image collection, etc.; then performs target detection on the target image to obtain a target image area where the target object is located; and performs gridding processing on the target image area to obtain at least one image block in the target image area.
[0034] S120: performing enhancement processing on the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain target image features of each image block.
[0035] The spatial correlation features include the position information of the image block in the target image area, the spatial relationship between the image block and other image blocks, etc. Exemplarily, a two-dimensional coordinate system corresponding to the target image area is established, and the spatial correlation features of each image block are determined according to the position information of the image block in the two-dimensional coordinate system.
[0036] The label information may be metadata associated with each image block. For example, the label information of each image block may be obtained by using manual or automated annotation tools, etc., which is not limited in the embodiments of the present application. The label information may include, but is not limited to, image quality score, whether the target object is blocked, whether the target object is wearing a mask, whether the target object is wearing a hat, whether the target object is wearing a uniform, the angle of the target object, and other features.
[0037] Global image features are used to describe the overall structure, color distribution, and texture changes of image blocks. In some embodiments, the image block can be input into a fully connected layer for full connection processing to obtain the global image features of the image block. In other embodiments, convolutional neural networks, such as VGG (Visual Geometry Group), ResNet (Residual Network), etc., can also be used to process the image block to obtain the global image features of the image block. Global image features focus on the image block as a whole, are easily disturbed by invalid information, and ignore key areas in the image. Therefore, the embodiment of the present application introduces label information and spatial correlation features to enhance the image blocks corresponding to the valid area, focus on the image blocks corresponding to the valid area during attribute recognition, and improve the accuracy of attribute recognition.
[0038] The target image feature is the image feature after the enhancement processing of each image block. Exemplarily, the target attribute recognition device first obtains the global image feature, spatial correlation feature and label information of each image block, and enhances the global image feature using the spatial correlation feature and label information to obtain the target image feature of each image block.
[0039] S130: Performing attribute recognition processing on the target object based on the target image features of each image block to obtain target attributes of the target object.
[0040] As an example, the attribute to be identified by the embodiment of the present application may be emotion, which may include happiness, sadness, anger, surprise, disgust, and fear. These emotions are not only basic emotions widely recognized in psychological research, but also have important application value for understanding individual behavior and reactions. Being prepared to identify emotions can not only provide an effective emotion analysis tool for the field of mental health, but also warn of possible dangerous situations and improve safety. As another example, the attribute to be identified by the embodiment of the present application may also be behavioral intention, which may include the movement trend of an individual, such as in which direction to move.
[0041] The target attribute is a specific category of the target object under a certain attribute. For example, when the attribute identified in the embodiment of the present application is emotion, the target attribute can be one of the emotion categories, such as happiness.
[0042] Emotion recognition pays great attention to the subtle expressions of the target object. For example, the curvature of the mouth corners and the curvature of the eyes may reflect different emotions. First, the effectiveness of each image block can be determined through the spatial correlation features of the image block and the label information of each image block. For example, when the target object wears a mask, the eye area of the target object is the effective area. The label information of each image block can make attribute recognition pay more attention to the eye area of the target object; then, combining the global image features of each image block can improve the feature extraction effect, thereby improving the accuracy of the target attribute.
[0043] It can be seen that the target attribute recognition method of the embodiment of the present application obtains the target image area where the target object is located in the target image, and the target image area includes at least one image block; the image blocks are enhanced according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block; the target object is attributed based on the target image features of each image block to obtain the target attribute of the target object. The spatial correlation features and label information of the image blocks are introduced, so that the target image features pay more attention to the effective information in the target image area, thereby improving the accuracy of target attribute recognition.
[0044] Among them, the above-mentioned step S110 may further include: performing recognition processing on the target image to obtain an initial image area of the target object in the target image; dividing the initial image area according to a preset ratio to obtain a first image sub-area and a second image sub-area; and determining the target image area based on the first image sub-area and the second image sub-area.
[0045] The initial image region is an image region obtained by overall recognition of the target object. As an example, when the target object is a dog, the initial image region includes the dog's facial region and body region. Exemplarily, the target image can be recognized by a pre-trained target detection model to obtain an image region where at least one object in the target image is located, and the image region where one of the objects is located is determined as the initial image region where the target object is located.
[0046] Afterwards, the target attribute recognition device can divide the initial image area according to a preset ratio to obtain a first image sub-area and a second image sub-area. The preset ratio can be determined according to the body proportion of the target object and the shooting angle of the target object. As an example, when the target object is a dog, the shooting angle is sideways, the preset ratio can be 1:5, and the facial area and body area of the dog are respectively extracted from the initial image area to obtain the first image sub-area and the second image sub-area. The facial area of the target object can also be extracted as the first image sub-area, and the initial image area of the target object is used as the second image sub-area. The first image sub-area and the second image sub-area are determined as the target image area.
[0047] In other embodiments, the target attribute recognition device can identify the first image sub-region where the target object is located through a first recognition module, and identify the second image sub-region where the target object is located through a second recognition module; and then determine the target image region based on the first image sub-region and the second image sub-region.
[0048] When the target image region is obtained based on the first image sub-region and the second image sub-region, the target attribute recognition device divides the first image sub-region to obtain the first image block in the first image sub-region, and divides the second image sub-region to obtain the second image block in the second image sub-region; enhances the first image blocks according to the spatial correlation features, label information and global image features of each first image block to obtain the target image features of each first image block; enhances the second image blocks according to the spatial correlation features, label information and global image features of each second image block to obtain the target image features of each second image block; and recognizes the target object based on the target image features of each first image block and the target image features of each second image block to obtain the target attribute of the target object.
[0049] In this embodiment, when the image block includes the first image block and the second image block, and the target image feature includes the target image feature of the first image block and the target image feature of the second image block, the target object is identified based on the target image features of each first image block and the target image features of each second image block, and the step of obtaining the target attribute of the target object may include: splicing the target image features of the first image block and the target image features of the second image block to obtain the splicing features; performing attribute identification processing on the target object according to the splicing features to obtain the target attribute of the target object. Thus, by combining different regions for multimodal recognition, the accuracy of attribute recognition can be improved. In particular, through the combination of facial regions and body regions, the facial region provides detailed information of expressions and micro-expressions, and the body region can show the posture, movement and overall situation of the target object. By utilizing the correlation and complementarity of the two, the accuracy and robustness of attribute recognition can be significantly improved.
[0050] Among them, the target image features of each first image block and the splicing features of each second image block can be processed by a neural network model to obtain the target attributes of the target object. Exemplarily, the neural network model can be a transformer model or BERT (Bidirectional Encoder Representations from Transformers). In other embodiments, in addition to inputting the target image features of each first image block and the target image features of each second image block into the neural network model, the label information corresponding to each image block can also be input into the neural network model; the target image features of each first image block, the target image features of each second image block and the label information of each image block are fused by the neural network model to obtain fused features; the fused features are attribute classified by the neural network model to obtain the target attributes of the target object.
[0051] In order to fully understand the above-mentioned method of performing attribute recognition through the first image sub-region and the second image sub-region, the following is described in detail as an example:
[0052] In this example, the first image sub-region is determined as a facial region, and the second image sub-region is determined as a body region; the emotion of the attribute to be identified.
[0053] First, the model training phase, the model includes a facial region feature extraction module, used to extract the target image features of the facial region; a body region feature extraction module, used to extract the target image features of the body region; a feature fusion module, used to fuse the target image features of the facial region and the target image features of the body region, and perform attribute recognition to obtain the target attributes of the target object. The facial region feature extraction module and the facial region feature extraction module are modules of the same structure, but the parameters vary with the optimization target.
[0054] The facial region feature extraction module is trained using a large dataset with labeled facial IDs to improve its ability to express facial features. Then, the facial region feature extraction module is fine-tuned for emotion classification using large facial data with emotion annotations to improve its ability to perceive emotional information in facial regions.
[0055] The facial region feature extraction module is trained as a model, and a large dataset with body IDs annotated is used to train the body region feature extraction module to improve its ability to express body features. Then, a large dataset with emotion annotated body data is used to fine-tune the emotion classification of the body region feature extraction module to improve its ability to perceive emotions in body regions.
[0056] There is no order in which to train the models of the facial region feature extraction module and the body region feature extraction module.
[0057] Finally, the entire model is trained by using the dataset associated with the facial region and the body region, loading the facial region feature extraction module and the model parameters trained by the facial region feature extraction module, and extracting the target image features of the facial region and the target image features of the body region. The loss value calculation formula through joint training is:
[0058]
[0059] in, represents the total loss value; Represents a constant greater than 0, which can be set based on experience; Represents the loss value of the facial area, Represents the loss value of the body area. Thus, through a three-stage training method, the facial feature and body feature extraction model reduces the demand for large-scale face and body joint annotation data, reduces the difficulty of training convergence, and improves training efficiency.
[0060] After obtaining the trained facial region feature extraction module, body region feature extraction module and feature fusion module, a target image is acquired, and the facial region of the target object is identified by the first recognition module, and the body region of the target object is identified by the second recognition module; the facial region is input into the facial region feature extraction module, and the facial region is gridded to obtain each first image block of the facial region, and the first image block is enhanced according to the spatial correlation features, label information and global image features of each first image block in the facial region to obtain the target image features of each first image block; the body region is input into the body region feature extraction module, and the body region is gridded to obtain each second image block of the body region, and the second image block is enhanced according to the spatial correlation features, label information and global image features of each second image block in the body region to obtain the target image features of each second image block; the target image features of each first image block extracted by the facial region feature extraction module are expressed as ,The target image features of each second image block extracted by the body region feature extraction module are expressed as ,The label information of face and body is expressed as , Represents the sum of the dimensions of the face and body.
[0061] The target image features of each first image block, the target image features of each second image block and the attribute information of each image block are spliced together and input into the feature fusion module for feature fusion processing to obtain the fusion feature. The fusion feature is expressed as ; Then, attribute recognition processing is performed based on the fused features to obtain the target attributes of the target object.
[0062] The above is only an example, and the target image area of the embodiment of the present application may also include only one, for example, only the face area or only the body area. The target attribute recognition device performs feature extraction processing on each image block in the target image area through the feature extraction module to obtain the target image feature of each image block; then the target image feature of each image block is input into the feature fusion module for attribute recognition processing to obtain the target attribute of the target object.
[0063] Furthermore, the process of the above step S120 may further include: performing full connection processing on the spatial correlation features and label information of each image block to obtain the label spatial correlation features of each image block; determining the target image features of each image block based on the label spatial correlation features of each image block and the global image features of the corresponding image block. Compared with the traditional image block processing that inputs the image blocks in a serialized manner, loses the spatial correlation information between the image blocks, thereby limiting the expressiveness of the features, the embodiment of the present application combines the spatial correlation features and label information to ensure that the target image features of each image block can be closely connected with other image blocks, thereby improving the accuracy of the features.
[0064] In other embodiments, the target attribute recognition device can also input the spatial correlation features, label information and global image features of each image block into a pre-trained network model for fusion processing to obtain fusion features; perform attribute recognition based on the fusion features to obtain the target attributes of the target object.
[0065] The target attribute recognition device first needs to obtain the spatial correlation features, label information and global image features of each image block, and then enhance the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block. Exemplarily, the spatial correlation features of each image block can be determined according to the spatial relationship between each image block and other image blocks; and feature extraction processing can be performed on each image block to obtain the global image features of each image block. In other embodiments, the spatial correlation features, label information and global image features of each image block can also be known information.
[0066] Exemplarily, the spatial relationship between each image block and other image blocks can be determined based on the position information of each image block and other image blocks. and image blocks For example, the image block The coordinates of the upper left corner of the target image area are expressed as , the coordinates of the lower right corner are expressed as ( ), image block The coordinates of the upper left corner of the target image area are expressed as , the coordinates of the lower right corner are expressed as . Image Block The calculation of the center point coordinates, height and width satisfies the following formula:
[0067]
[0068] in, represents the x-coordinate of the center point, represents the y coordinate of the center point, Represents an image block The width of Represents an image block The height of the image block can also be obtained according to the above formula The center point coordinates ( ),high And width .
[0069] That image block and image blocks The spatial relationship between them can be expressed as:
[0070]
[0071] in, Represents an image block and image blocks The spatial relationship between Represents an image block The x coordinate of the center point, Represents an image block The y coordinate of the center point, Represents an image block The width of Represents an image block Height; Represents an image block The x coordinate of the center point, Represents an image block The y coordinate of the center point, Represents an image block The width of Represents an image block height.
[0072] The spatial relationship is processed by feature extraction to obtain the spatial correlation features of each image block.
[0073] In order to reduce parameter complexity and improve computational efficiency, when extracting the global image features of each image block, the target attribute recognition device determines the product of the height of each image block and the width of the corresponding image block as the number of pixel rows of the corresponding image block; serializes each pixel point in the image block according to the number of pixel rows of each image block to obtain the serialized image block; and performs full connection processing on the serialized image block to obtain the global image features of the corresponding image block. In this way, each image block is flattened and serialized into one-dimensional data, simplifying the input format. In the fully connected layer, each neuron is connected to all neurons in the previous layer. Therefore, if high-dimensional input is used directly, the number of parameters will increase sharply. The flattened one-dimensional data reduces the number of such connections, thereby reducing the complexity of the model.
[0074] As an example, the size information of each image block, including height and width, can be expressed as , represents the height of the image block, Represents the width of the image block; flattens the two-dimensional image data of the image block into one-dimensional data, and sets the number of pixel rows , whereby the pixels of each image block are arranged into a column according to the number of pixel rows; if the number of image blocks is N, the image block after serialization can be represented as I , each image block is input into the fully connected layer for full connection processing to obtain the global image features of the corresponding image block. Exemplarily, the full connection processing formula is as follows:
[0075]
[0076] in, represents the global image features of each image block, represents the dimension of the fully connected layer, and the formula represents the image block After the fully connected layer The linear transformation is then performed through the ReLU activation function to get the output.
[0077] In other embodiments, the target attribute recognition device may also perform full connection processing on the two-dimensional data of each image block to obtain the global image features of each image block.
[0078] After obtaining the spatial correlation features, label information and global image features of each image block, the target attribute recognition device performs full connection processing on the spatial correlation features and label information of each image block to obtain the label spatial correlation features of each image block. In some embodiments, the spatial relationship and label information are first spliced to obtain the spliced spatial relationship and label information. Exemplarily, the spliced spatial relationship and label information can be expressed as:
[0079]
[0080] in, Represents an image block and image blocks The spatial relationship between Represents an image block The x coordinate of the center point, Represents an image block The y coordinate of the center point, Represents an image block The width of Represents an image block Height; Represents an image block The x coordinate of the center point, Represents an image block The y coordinate of the center point, Represents an image block The width of Represents an image block Height, Represents an image block and image blocks The tag information of Tag information.
[0081] The spliced spatial relationship and label information are input into the fully connected layer for full connection processing to obtain label space association features. For example, the calculation of label space association features satisfies the following formula:
[0082]
[0083] in, express After the fully connected layer The linear transformation of is output; Represents the label space correlation feature, the output obtained by the ReLU activation function; is the high-dimensional feature representation of r, are the training parameters in the model, .
[0084] Then, the target image features of each image block are determined based on the label space association features of each image block and the global image features of the corresponding image block. Exemplarily, the label space association features of each image block and the global image features of the corresponding image block are decoded to obtain the image features of each image block after decoding; the image features of each image block after decoding are subjected to feature extraction to obtain the target image features of each image block. In other embodiments, the image features of each image block after decoding can be directly used as the target image features of each image block. In other embodiments, the label space association features of each image block can also be used to perform weighted processing on the global image features of the corresponding image block to obtain the target image features of each image block.
[0085] The purpose of the decoding process is to enhance the global image features of the corresponding image blocks through the label association features of each image block, so as to obtain different importance levels of each image block. As an example, the decoding process can be performed by a decoder based on the self-attention mechanism. Suppose the training parameters of the decoder are , , the decoder decodes the label association features and global image features of each image block, and the specific steps of the decoding process may include: determining the query vector, key vector and value vector according to the global image features of the image block; obtaining the vector similarity between the query vector and the key vector; enhancing the corresponding vector similarity according to the label space association features of the image block to obtain the target vector similarity; weighting the value vector based on the target vector similarity to obtain the attention value; and determining the sum of the attention value and the global image feature as the image feature after the image block is decoded.
[0086] The query vector, key vector, and value vector are the three core concepts of the attention mechanism. The query vector represents the location or content of the information you want to obtain, the key vector represents the features or content at each location, and the value vector contains the actual information to be transmitted. The query vector, key vector, and value vector are generated based on the global image features of the input image block. The global image features can be mapped to different vector spaces through the transformation matrix to obtain the query vector, key vector, and value vector. As an example, the calculation of the query vector, key vector, and value vector satisfies the following formula:
[0087]
[0088] in, represents the query vector, represents the key vector, represents a value vector, represents the global image features of the image patch, , , Represents the training parameters of the decoder.
[0089] Then, the vector similarity between the query vector and the key vector is obtained. Commonly used similarity measurement methods include dot product, scaled dot product, etc.
[0090] Compared with the traditional attention calculation method, the embodiment of the present application adds label space association features for guidance, and the calculation formula of its attention value is transformed into:
[0091]
[0092] in, represents the attention value, represents the activation function, represents the query vector, represents the key vector, represents a value vector, represents the dimension of the attention layer in the decoder, Represents label space association features.
[0093] After obtaining the attention value, the sum of the attention value and the global image feature is determined as the image feature after the image block is decoded. In other embodiments, the target attribute recognition device may also determine the product of the attention value and the global image feature as the image feature after the image block is decoded.
[0094] After obtaining the image features after decoding of each image block, the target attribute recognition device can use the ViT network (Vision Transformer, visual transformation network) to perform feature extraction processing on the image features after decoding of each image block to obtain the target image features of each image block.
[0095] In order to elaborate on the target attribute recognition method of the present application, Figure 2 and Figure 3 The framework diagram shown further illustrates it, as follows:
[0096] It should be noted that Figure 2 is a schematic diagram of a framework of an exemplary embodiment of a target attribute recognition method shown in the present application, Figure 3 yes Figure 2 A schematic diagram of a framework of an exemplary embodiment of a grid enhancement module is shown in FIG. This description is only used as an example. In this example, the target image region includes a first image sub-region and a second image sub-region. The first image sub-region is a face region, and the second image sub-region is a body region.
[0097] The target attribute recognition device obtains a target image, performs recognition processing on the target image, and obtains the facial region and body region of the target object. It should be noted that if only the facial region or body region of the target object exists in the target image, only the facial region or body region can be used as the target image region. Then, the facial region is gridded to obtain a first image block in the facial region; the body region is gridded to obtain a second image block in the body region.
[0098] For details, please refer to Figure 3 , determine the spatial correlation features of each image block according to the spatial relationship between each first image block and other image blocks in the facial area; extract annotation information of the facial area to obtain annotation information of each image block in the facial area; input the spatial correlation features and annotation information into the fully connected layer for full connection processing to obtain the annotated spatial correlation features. Flatten and serialize each first image block into one-dimensional data, and the number of image blocks in the facial area is N; input each first image block into the fully connected layer for full connection processing to obtain the global image features of each first image block; then input the annotated spatial correlation features and the global image features into the decoder for decoding processing to obtain the image features of each first image block after decoding. The image features of each image block in the body area after decoding can be obtained by referring to the above method.
[0099] Each first image block in the facial region is input into the ViT network for feature extraction to obtain the target image features of each first image block in the facial region; each second image block in the body region is input into the ViT network for feature extraction to obtain the target image features of each second image block in the body region.
[0100] The target image features of each first image block in the facial area, the target image features of each second image block in the body area, and the label information corresponding to each image block are spliced and input into the transformer for feature fusion processing to obtain the fused features. Then, attribute classification processing is performed based on the fused features to obtain the target attributes of the target object.
[0101] The above scheme uses a two-dimensional image as input, and ordinary image acquisition equipment can be applied. It does not require the cooperation of the object to be identified, and can achieve seamless attribute recognition. It does not need to extract and fuse video features, which reduces the computational complexity and time consumption, and can achieve real-time attribute recognition.
[0102] See also Figure 4 , Figure 4 It is a structural diagram of an exemplary embodiment of a target attribute recognition device shown in the present application. The target attribute recognition device 400 includes an acquisition module 410, an enhancement module 420 and a recognition module 430. The acquisition module 410 is used to acquire a target image region where a target object is located in a target image, and the target image region includes at least one image block; the enhancement module 420 is used to enhance the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block; the recognition module 430 is used to perform attribute recognition processing on the target object based on the target image features of each image block to obtain the target attribute of the target object.
[0103] In the above scheme, the target attribute recognition device obtains the target image region where the target object is located in the target image, and the target image region includes at least one image block; the image blocks are enhanced according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block; the target object is attributed based on the target image features of each image block to obtain the target attribute of the target object. Thus, the spatial correlation features and label information of the image blocks are introduced, so that the target image features pay more attention to the effective information in the target image region, thereby improving the accuracy of target attribute recognition.
[0104] The functions of each module can be found in the target attribute identification method implementation example, and will not be repeated here.
[0105] In order to implement the target attribute recognition method of the above embodiment, the present application proposes another electronic device, which is specifically referred to as Figure 5 , Figure 5 It is a structural schematic diagram of an embodiment of an electronic device provided by the present application.
[0106] The electronic device 500 includes a memory 510 and a processor 520 , wherein the memory 510 and the processor 520 are coupled.
[0107] The memory 510 is used to store program data, and the processor 520 is used to execute the program data to implement the target attribute recognition method of the above embodiment.
[0108] In this embodiment, the processor 520 may also be referred to as a CPU (Central Processing Unit). The processor 520 may be an integrated circuit chip having signal processing capabilities. The processor 520 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. A general-purpose processor may be a microprocessor or the processor 520 may also be any conventional processor, etc.
[0109] The present application also provides a computer-readable storage medium, such as Figure 6 As shown, the computer-readable storage medium 600 is used to store program data 610. When the program data 610 is executed by the processor, it is used to implement the target attribute recognition method in the method embodiment of the present application.
[0110] The method involved in the target attribute identification method embodiment of the present application, when implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a device, such as a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0111] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A target attribute recognition method, characterized in that: The target attribute recognition method comprises: Acquire a target image region where a target object is located in a target image, wherein the target image region includes at least one image block; Performing enhancement processing on the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain target image features of each image block; Performing attribute recognition processing on the target object based on the target image features of each image block to obtain the target attribute of the target object; The step of performing enhancement processing on the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block comprises: Perform full connection processing on the spatial correlation features and label information of each image block to obtain the label spatial correlation features of each image block; Determining a query vector, a key vector, and a value vector according to the global image features of the image block; Obtaining vector similarity between the query vector and the key vector; Performing enhancement processing on the corresponding vector similarity according to the label space association feature of the image block to obtain the target vector similarity; Performing weighted processing on the value vector based on the target vector similarity to obtain an attention value; Determine the sum of the attention value and the global image feature as the image feature after the image block is decoded; The image features after decoding of each image block are subjected to feature extraction processing to obtain target image features of each image block.
2. The target attribute recognition method according to claim 1, characterized in that: The step of obtaining a target image region where a target object is located in a target image, wherein the target image region includes at least one image block, comprises: Performing recognition processing on the target image to obtain an initial image region of the target object in the target image; Dividing the initial image region according to a preset ratio to obtain a first image sub-region and a second image sub-region; The target image region is determined based on the first image sub-region and the second image sub-region.
3. The target attribute recognition method according to claim 1, characterized in that: The image block includes a first image block and a second image block, the target image feature includes a target image feature of the first image block and a target image feature of the second image block, and the step of performing attribute recognition processing on the target object based on the target image features of each image block to obtain the target attribute of the target object includes: Performing splicing processing on the target image feature of the first image block and the target image feature of the second image block to obtain a splicing feature; Attribute recognition processing is performed on the target object according to the splicing features to obtain target attributes of the target object.
4. The target attribute recognition method according to claim 1, characterized in that: Before the step of performing enhancement processing on the image blocks according to the spatial correlation features, label information and global image features of each image block to obtain the target image features of each image block, the method further includes: Determine the spatial correlation feature of each image block according to the spatial relationship between each image block and other image blocks; Perform feature extraction processing on each image block to obtain the global image features of each image block.
5. The target attribute recognition method according to claim 4, characterized in that: The step of performing feature extraction processing on each image block to obtain the global image features of each image block includes: The product of the height of each image block and the width of the corresponding image block is determined as the number of pixel rows of the corresponding image block; Performing serialization processing on each pixel point in the image block according to the number of pixel rows of each image block to obtain a serialized image block; The serialized image blocks are fully connected to obtain global image features of the corresponding image blocks.
6. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that: include: Program data is stored, and when the program data is executed by a processor, it is used to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image recognition method, electronic equipment and storage medium
CN117173764A
Pedestrian attribute recognition method and device, electronic equipment and storage medium
CN119206769A