Pedestrian attribute recognition method and related equipment

By extracting feature, mapping and building graph structures of pedestrian image data, combined with multi-layer perceptron decoding, the problem of insufficient recognition accuracy of existing pedestrian attribute recognition methods is solved, and efficient and accurate pedestrian attribute recognition is achieved.

CN114359961BActive Publication Date: 2025-05-16SHENZHEN HARZONE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111607874.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-05-16
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

The existing pedestrian attribute recognition methods are not highly accurate in the field of intelligent video surveillance and security, resulting in additional workload when looking for pedestrians.

Method used

By obtaining human image data, feature extraction is performed to obtain feature maps, feature mapping and graph structure is constructed. Each node corresponds to an attribute identification of one dimension, and the nodes of the graph structure are decoded using a multi-layer perceptron to obtain multiple attribute information.

Benefits of technology

The accuracy of pedestrian attribute recognition is improved, the structural features are aggregated and extracted through graph convolution, and the nodes are decoded using multi-layer perceptrons, so as to reasonably understand and identify the attributes of each part and obtain efficient and accurate prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359961B_ABST
    Figure CN114359961B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a pedestrian attribute recognition method and related equipment, the method comprising: obtaining human image data; extracting features from the human image data to obtain a feature map; performing feature mapping on the feature map, and constructing a graph structure based on the result after feature mapping, each graph structure comprising multiple nodes, each node corresponding to an attribute identifier of a dimension; and decoding each node of the graph structure using a multi-layer perceptron to obtain multiple attribute information. The embodiment of the present application can improve the accuracy of pedestrian attribute recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a pedestrian attribute recognition method and related equipment. Background Art

[0002] In the existing technology, pedestrian attribute recognition methods have a significant application value in the field of intelligent video surveillance and security. The appearance and movements of pedestrians, including gender, age, clothing style, walking and cycling, etc., bring obvious convenience, accuracy and efficiency to the needs of finding lost persons. However, due to the low recognition accuracy, it often brings extra workload to the search. Therefore, the problem of how to accurately identify pedestrian attributes needs to be solved urgently. Summary of the invention

[0003] The embodiments of the present application provide a method and related equipment for pedestrian attribute recognition, which can improve the accuracy of pedestrian attribute recognition.

[0004] In a first aspect, an embodiment of the present application provides a method for identifying pedestrian attributes, the method comprising:

[0005] Acquiring human body image data;

[0006] Performing feature extraction on the human body image data to obtain a feature map;

[0007] Performing feature mapping on the feature graph, and constructing a graph structure based on the result of feature mapping, each graph structure includes a plurality of nodes, and each node corresponds to an attribute identifier of a dimension;

[0008] A multi-layer perceptron is used to decode each node of the graph structure to obtain multiple attribute information.

[0009] In a second aspect, an embodiment of the present application provides a pedestrian attribute recognition device, the device comprising: an acquisition unit, an extraction unit, a mapping unit and a decoding unit, wherein:

[0010] The acquisition unit is used to acquire human body image data;

[0011] The extraction unit is used to extract features from the human body image data to obtain a feature map;

[0012] The mapping unit is used to perform feature mapping on the feature graph and construct a graph structure based on the result of feature mapping, each graph structure includes a plurality of nodes, and each node corresponds to an attribute identifier of a dimension;

[0013] The decoding unit is used to decode each node of the graph structure using a multi-layer perceptron to obtain multiple attribute information.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program includes instructions for executing the steps in the first aspect of the embodiment of the present application.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, wherein the computer program enables a computer to execute part or all of the steps described in the first aspect of the embodiment of the present application.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application. The computer program product may be a software installation package.

[0017] The implementation of the embodiments of the present application has the following beneficial effects:

[0018] It can be seen that the pedestrian attribute recognition method and related equipment described in the embodiments of the present application obtain human image data, perform feature extraction on the human image data, obtain a feature map, perform feature mapping on the feature map, and construct a graph structure based on the result after feature mapping. Each graph structure includes multiple nodes, each node corresponds to an attribute identifier of one dimension, and a multi-layer perceptron is used to decode each node of the graph structure to obtain multiple attribute information. Since each node represents the semantic features of each part of the human body that can be learned, after the structural features are aggregated and extracted by graph convolution, the multi-layer perceptron is used to decode each node, so that the attributes of each part are reasonably decoupled and identified, and finally the prediction results are obtained efficiently and accurately, which helps to improve the accuracy of pedestrian attribute recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1A It is a flowchart of a pedestrian attribute recognition method provided in an embodiment of the present application;

[0021] Figure 1B It is a demonstration schematic diagram of a data synthesis model provided in an embodiment of the present application;

[0022] Figure 1C It is a demonstration schematic diagram of an algorithm reasoning model provided in an embodiment of the present application;

[0023] Figure 2 It is a flowchart of another pedestrian attribute recognition method provided in an embodiment of the present application;

[0024] Figure 3 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application;

[0025] Figure 4 This is a block diagram of the functional units of a pedestrian attribute recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0027] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0028] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0029] The electronic devices described in the embodiments of the present application may include smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, PDAs, driving recorders, traffic control platforms, servers, laptops, mobile Internet devices (MID, Mobile Internet Devices) or wearable devices (such as smart watches, Bluetooth headsets), etc. The above are only examples and not exhaustive, including but not limited to the above electronic devices.

[0030] The following is a detailed description of the embodiments of the present application.

[0031] See also Figure 1A , Figure 1A : is a flowchart of a pedestrian attribute recognition method provided by an embodiment of the present application. As shown in the figure, the pedestrian attribute recognition method includes:

[0032] 101. Obtain human body image data.

[0033] In the embodiment of the present application, the human body image data includes an image of the area where the human body is located. In a specific implementation, the human body can be photographed to obtain a first image, and then the foreground of the first image can be extracted to obtain the human body image data.

[0034] Optionally, the following steps may also be included:

[0035] A1. Acquire a first image, where the first image includes a person;

[0036] A2. Acquire a human body image, clothing parameters, and posture parameters in the first image;

[0037] A3, inputting the human body image, the clothing parameters and the posture parameters into a preset condition human body analysis network to obtain a part segmentation result map;

[0038] A4. Input the part segmentation result map, the clothing parameters and the posture parameters into an adaptive generative adversarial network to obtain an adaptive synthetic map, and use the adaptive synthetic map and / or the human body image as the human body image data.

[0039] Among them, the preset conditional human body analysis network can be pre-set or system default, and the preset conditional human body analysis network can be used to realize human body segmentation. For example, the preset conditional human body analysis network can be a human body analysis network CE2P. For example, the preset conditional human body analysis network can include a residual network (Residual Blocks). Clothing parameters can include clothing images or clothing feature parameters. The clothing feature parameters can include at least one of the following: feature points, feature textures, feature vectors, etc., which are not limited here. Posture parameters can include at least one of the following: posture contours, posture labels, etc., which are not limited here. In an embodiment of the present application, at least one of the adaptive synthetic image and the human body image can be used as human body image data.

[0040] Among them, in the embodiment of the present application, the parts may include at least one of the following: head, left hand, right hand, neck, stomach, chest, left leg, right leg, waist, buttocks, etc., which are not limited here.

[0041] In a specific implementation, a first image can be obtained, which includes a person. The first image can also be segmented to obtain a human body image, clothing parameters, and posture parameters in the first image. The human body image, clothing parameters, and the posture parameters are then input into a preset conditional human body parsing network to obtain a part segmentation result map. The part segmentation result map, clothing parameters, and posture parameters are input into an adaptive generative adversarial network to obtain an adaptive synthetic map. The adaptive synthetic map and / or the human body image are used as human body image data. In this way, human body data enhancement can be achieved.

[0042] In the embodiment of the present application, adaptive clothing virtual synthesis based on the source image (human body image in the first image) can be performed, such as Figure 1B As shown in the figure, first, the human body image and the target clothing and its posture can be used as input, and the designed conditional parsing network can be used to predict the synthetic map of human body parts and their texture areas. Then, the inherent body shape information and posture characteristics of the human body are introduced, and the conditions such as the human body measurements and clothing posture key points are input into the adaptive generative adversarial learning. In this way, the distorted clothing texture can be adaptively adjusted while retaining the inherent body shape and posture characteristics of the human body, and a satisfactory clothing synthesis result can be output, thereby enhancing the diversity of source data.

[0043] Optionally, the following steps may also be included:

[0044] Obtaining the three-dimensional data of the character;

[0045] Then step A3, inputting the human body image, the clothing parameters and the posture parameters into a preset condition human body analysis network to obtain a part segmentation result map, can be implemented as follows:

[0046] The human body image, the clothing parameters, the posture parameters, and the three-dimensional data are input into a preset condition human body analysis network to obtain a part segmentation result map.

[0047] In the specific implementation, the three-dimensional data of the person can also be obtained. The image of the person can be obtained through the camera, and the three-dimensional data can be obtained by performing three-dimensional recognition. Then the human image, clothing parameters, posture parameters, and three-dimensional data are input into the preset condition human body analysis network to obtain the part segmentation result map.

[0048] For example, in the specific implementation, data enhancement can be achieved according to the following steps:

[0049] 1. Randomly select augmented clothing, input human source image, and randomly generate human body measurements;

[0050] 2. This step mainly performs the synthesis and segmentation of human body parts and locates the position and range of the clothing to be synthesized. Figure 1B As shown in the figure, given the target human image I (source data I), the input image of the selected clothing C (target clothing C) and the target posture P (i.e. clothing key point P), these three pieces of information are used as conditional input and sent to the conditional human parsing network (Residual Blocks). The prediction process of the part segmentation (part segmentation image) S can be expressed as maximizing the posterior probability p(S′ t |(M h , M f , M b , C, P)) ,) Finally, the adaptive synthetic graph r can be obtained. The conditional parsing is based on the conditional generative adversarial network (CGAN), and the maximized posterior probability p can be expressed as:

[0051] p(S′ t |(M h , M f , M b , C, P)) = G(M h , M f , M b , C, P)

[0052] Among them, the generative adversarial network uses a network similar to ResNet as the generator G to build a conditional parsing model. On the one hand, the L1 function can be used as the supervision loss to further improve the performance, thereby generating smoother results. On the other hand, the pixel-level softmax loss is used to encourage the generator to synthesize high-quality synthetic segmentation maps of human body parts. Therefore, the learning process of the conditional human body parsing network is expressed as:

[0053]

[0054] 3. According to the synthetic segmentation map of human body parts, clothing texture distortion synthesis can be performed. Since pixel misalignment will lead to blurred results, an adaptive generative adversarial network is introduced to distort the desired clothing texture appearance into the human body analysis synthesis map, thereby alleviating the misalignment problem between the input human body posture and the desired clothing posture. Affine and TPS (thin plate spline interpolation) transformations are used at the same time. In addition, in order to improve the generalization ability of the model, the pre-trained model is directly used to estimate the transformation mapping analysis between the target analysis and the synthetic analysis, and then this transformation mapping is used to distort the target human image. In this structure, it is composed of the designed deep adaptive distortion GAN generator Gwarp and the adaptive distortion GAN discriminator Dwarp respectively. In addition, the target clothing image is distorted using a geometric matching module. In this process, the distorted clothing image Cw, the target human body image Iw / o clothing, the target clothing posture P and the synthesized human body analysis St0 can be used as the input of the deep adaptive distortion network GAN, and finally the human clothing distortion result is synthesized:

[0055]

[0056] Furthermore, during the network learning process, perceptual loss is applied to measure the distance between high-level features in the pre-trained model, which promotes the generator to synthesize high-quality and realistic images. The perceptual loss can be expressed as:

[0057]

[0058] Among them, φ i (i) represents the i-th (i=0-4) layer feature map in the pre-trained network φ of the real image i. In the embodiment of the present application, the pre-trained VGG19 can be used as φ, and the L1 norm of the last five layers of feature maps in φ is weighted and summed to represent the perceptual loss between images, α i Control the loss weight of each layer. In this way, the model recognition accuracy can be further improved because the generator synthesizes high-quality and realistic images.

[0059] 102. Perform feature extraction on the human body image data to obtain a feature map.

[0060] In a specific implementation, the human body image data can be subjected to feature extraction through a preset neural network model to obtain a feature map. The preset neural network model can include at least one of the following: a convolutional neural network model, a recurrent neural network model, and a fully connected neural network model.

[0061] In a specific implementation, human image data may include source data and synthetic data, where the source data is the human image existing in the image itself, and the synthetic data is the data synthesized using the human image. In a specific implementation, the source data and the synthetic data are mixed and sampled, and the human image X to be identified (dimensions H*W*C, HWC are image length, width and channel respectively) is input into the algorithm model for feature extraction. For example, a deep convolutional network structure obtains image X, performs preliminary feature learning, and outputs a feature map M (dimensions H'*W'*C'), where the deep convolutional network can be four residual convolution modules with a step size of 2.

[0062] 103. Perform feature mapping on the feature graph, and construct a graph structure based on the result of feature mapping, each graph structure includes multiple nodes, and each node corresponds to an attribute identifier of a dimension.

[0063] The attribute identifier may include at least one of the following: gender, color, clothing type, body shape, age, whether to wear makeup, posture, action, etc., which are not limited here.

[0064] In the specific implementation, when performing feature mapping, an undirected graph G = (V, E) can be defined, where V is the feature vector (dimension N*D, N represents the number of parts of the human body, and D is the dimension of the semantic feature representation of each part), and E is the edge connecting each node. The feature map mapping process can be expressed as: Z = Sigma (X, W), where the high-level graph represents Z as the output of the feature mapping, dimension N*D, W is a trainable transposed matrix, which converts the dimension C' of the intermediate feature X' into dimension D, and Sigma is a nonlinear activation function. Among them, if Figure 1C As shown in Figure 1, Z is the input feature stream of node decoding, that is, the output of the constructed graph structure, that is, the high-level graph representation, which can also be understood as the graph output. The transposed matrix can be understood as the feature information transmission between nodes in the graph structure using the graph convolution operation, for example, W is the weight of the graph convolution.

[0065] In the specific implementation, first, learn a mapping parameter matrix P with a dimension of C'*N, and convert the feature dimension C' of the feature graph M into dimension N according to the number of nodes in the graph. This process can be expressed as X1=X(H'W'*C')*P(C'*N). Then, a transition feature matrix X2=X1'(N*H'W')*X(H'W'*C') can be calculated; finally, using the transition feature matrix X2 and the trainable transposed matrix W, the feature graph M is mapped into N D-dimensional graph nodes representing the semantics of various parts of the human body, which can be expressed as: Z(N,D)=Sigma(X,W)=P T (N*C')*X T (C'*H'W')*X(H'W'*C')*W(C'*D).

[0066] Optionally, the above step 103, performing feature mapping on the feature graph, may include the following steps:

[0067] 31. Obtain the preset mapping parameter matrix;

[0068] 32. Convert the feature dimension of the feature graph into a preset dimension according to the preset mapping parameter matrix and the number of nodes in the feature graph to obtain a conversion matrix;

[0069] 33. Determine a transition characteristic matrix according to the conversion matrix;

[0070] 34. According to the transition feature matrix and the preset trainable transposed matrix, the feature map is mapped into graph nodes of multiple dimensions representing the semantics of various parts of the human body, and the graph nodes of multiple dimensions are used as the result after feature mapping.

[0071] In a specific implementation, the preset mapping parameter matrix can be preset or the system defaults, and the preset trainable transposed matrix can also be preset or the system defaults. Specifically, the preset mapping parameter matrix can be obtained, and then the feature dimension of the feature map is converted into the preset dimension according to the preset mapping parameter matrix and the number of nodes in the feature map to obtain a conversion matrix, which can be expressed as an equation. Then, through equation transformation, the transition feature matrix can be determined according to the conversion matrix. Finally, according to the transition feature matrix and the preset trainable transposed matrix, the feature map can be mapped into graph nodes of multiple dimensions representing the semantics of various parts of the human body, and the graph nodes of the multiple dimensions are used as the result after feature mapping.

[0072] Optionally, the above step 103, constructing a graph structure based on the result after feature mapping, can be implemented as follows:

[0073] According to the inherent properties of the human body, the connection between the parts of the human body is used to encode the relationship between two nodes to obtain a node edge set, and a graph structure is constructed based on the node edge set.

[0074] In the specific implementation, the graph convolution method can be used to iteratively train each node of the graph to enhance the feature semantic information. Specifically, the semantic constraints in the knowledge of human body parts can be used to construct a global graph reasoning based on the high-level graph representation Z. According to the inherent properties of the human body, the connection between the parts of the human body can be introduced to encode the relationship between two nodes, and the node edge set E can be obtained.

[0075] For example, the torso usually appears with the head, so these two nodes are linked. The head node and the leg node are disconnected because they have no association. Then, according to the graph convolution method, matrix multiplication is used to perform graph propagation on the representation Z of all part nodes to obtain the enhanced feature Z e :Z e=Sigma(A e ZW e ), where W e is a trainable weight matrix with dimension D*D and node adjacency weight A e According to the definition of the node edge set E, which is a normalized symmetric adjacency matrix (dimension N*N), according to the experimental results, four graph convolution operations can be performed to obtain the optimal result.

[0076] 104. Decode each node of the graph structure using a multi-layer perceptron to obtain multiple attribute information.

[0077] In the embodiment of the present application, a multi-layer perceptron (MLP) is used to decode each enhanced feature semantic node Ze in the graph structure to obtain the prediction result of the corresponding attribute of each part. Each attribute information can correspond to an attribute set, which can be a class of attributes, or can also be the attribute of a part.

[0078] That is, the main idea of ​​the embodiment of the present application is to extract features from portraits, map the feature flow to the nodes of the graph structure (the mapping relationship is the graph structure modeling), and use graph convolution operations in the graph structure to communicate and transmit information between nodes. The graph structure outputs a high-level graph representation (such as Z mentioned above) for subsequent node decoding, and finally generates prediction results for each attribute.

[0079] In the specific implementation, an end-to-end supervised learning strategy can be adopted, such as Figure 1C As shown in the figure, first, the human image data X is input into the deep convolutional network to extract the feature map M (feature encoding). After feature mapping, the feature map is constructed into a graph G (graph structure modeling). Each node in the graph G represents the semantic features of each part of the human body that can be learned. After the structural features are aggregated and extracted by graph convolution, each node is decoded by a multi-layer perceptron, so that the attributes of each part can be reasonably decoupled and identified. For example, after identification, the global attributes and the attribute information of each part can be obtained, and finally the prediction results can be obtained efficiently and accurately.

[0080] It can be seen that the pedestrian attribute recognition method and related equipment described in the embodiments of the present application obtain human image data, perform feature extraction on the human image data, obtain a feature map, perform feature mapping on the feature map, and construct a graph structure based on the result after feature mapping. Each graph structure includes multiple nodes, each node corresponds to an attribute identifier of one dimension, and a multi-layer perceptron is used to decode each node of the graph structure to obtain multiple attribute information. Since each node represents the semantic features of each part of the human body that can be learned, after the structural features are aggregated and extracted by graph convolution, the multi-layer perceptron is used to decode each node, so that the attributes of each part are reasonably decoupled and identified, and finally the prediction results are obtained efficiently and accurately, which helps to improve the accuracy of pedestrian attribute recognition.

[0081] With the above Figure 1A For the embodiments shown, please refer to Figure 2 , Figure 2 : is a flowchart of a pedestrian attribute recognition method provided by an embodiment of the present application, which is applied to an electronic device. As shown in the figure, the pedestrian attribute recognition method includes:

[0082] 201. Acquire a first image, where the first image includes a person.

[0083] 202. Acquire a human body image, clothing parameters, and posture parameters in the first image.

[0084] 203. Input the human body image, the clothing parameters and the posture parameters into a preset conditional human body analysis network to obtain a part segmentation result map.

[0085] 204. Input the part segmentation result image, the clothing parameters, and the posture parameters into an adaptive generative adversarial network to obtain an adaptive synthetic image, and use the adaptive synthetic image and / or the human body image as human body image data.

[0086] 205. Acquire the human body image data.

[0087] 206. Perform feature extraction on the human body image data to obtain a feature map.

[0088] 207. Perform feature mapping on the feature graph, and construct a graph structure based on the result of feature mapping, where each graph structure includes multiple nodes, and each node corresponds to an attribute identifier of a dimension.

[0089] 208. Decode each node of the graph structure using a multi-layer perceptron to obtain multiple attribute information.

[0090] The specific description of the above steps 201 to 208 can refer to the above Figure 1A The corresponding steps of the described pedestrian attribute recognition method will not be repeated here.

[0091] It can be seen that the pedestrian attribute recognition method described in the embodiment of the present application, on the one hand, uses human body images and target clothing and its posture as input, predicts the synthetic map of human body parts and its texture area through the designed conditional parsing network, introduces the inherent body information and posture characteristics of the human body itself, and inputs clothing, posture key points and other conditions into adaptive generative adversarial learning, so that while retaining the inherent body characteristics and posture characteristics of the human body, it can adaptively adjust the distorted clothing texture, output satisfactory clothing synthesis results, and enhance the diversity of source data. On the other hand, since each node represents the semantic characteristics of each part of the human body that can be learned, after the structural features are aggregated and extracted by graph convolution, each node is decoded by a multi-layer perceptron, so that the attributes of each part are reasonably decoupled and identified, and finally the prediction results are obtained efficiently and accurately, which helps to improve the accuracy of pedestrian attribute recognition.

[0092] In accordance with the above embodiment, please refer to Figure 3 , Figure 3 : is a structural diagram of an electronic device provided in an embodiment of the present application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs, which are applied to the electronic device. The one or more programs are stored in the memory and configured to be executed by the processor. In the embodiment of the present application, the program includes instructions for executing the following steps:

[0093] Acquiring human body image data;

[0094] Performing feature extraction on the human body image data to obtain a feature map;

[0095] Performing feature mapping on the feature graph, and constructing a graph structure based on the result of feature mapping, each graph structure includes a plurality of nodes, and each node corresponds to an attribute identifier of a dimension;

[0096] A multi-layer perceptron is used to decode each node of the graph structure to obtain multiple attribute information.

[0097] Optionally, in the aspect of performing feature mapping on the feature graph, the program includes instructions for performing the following steps:

[0098] Get the preset mapping parameter matrix;

[0099] Convert the feature dimension of the feature graph into a preset dimension according to the preset mapping parameter matrix and the number of nodes in the feature graph to obtain a conversion matrix;

[0100] Determine a transition characteristic matrix according to the conversion matrix;

[0101] According to the transition feature matrix and the preset trainable transposed matrix, the feature map is mapped into graph nodes of multiple dimensions representing the semantics of various parts of the human body, and the graph nodes of the multiple dimensions are used as the result after feature mapping.

[0102] Optionally, in the aspect of constructing the graph structure based on the result after feature mapping, the program includes instructions for executing the following steps:

[0103] According to the inherent properties of the human body, the connection between the parts of the human body is used to encode the relationship between two nodes to obtain a node edge set, and a graph structure is constructed based on the node edge set.

[0104] Optionally, the program further includes instructions for executing the following steps:

[0105] Acquire a first image, where the first image includes a person;

[0106] Acquire a human body image, clothing parameters, and posture parameters in the first image;

[0107] Inputting the human body image, the clothing parameters and the posture parameters into a preset condition human body analysis network to obtain a part segmentation result map;

[0108] The part segmentation result map, the clothing parameters and the posture parameters are input into an adaptive generative adversarial network to obtain an adaptive synthetic map, and the adaptive synthetic map and / or the human body image are used as the human body image data.

[0109] Optionally, the program further includes instructions for executing the following steps:

[0110] Obtaining the three-dimensional data of the character;

[0111] In the aspect of inputting the human body image, the clothing parameters and the posture parameters into a preset conditional human body analysis network to obtain a part segmentation result map, the above program includes instructions for executing the following steps:

[0112] The human body image, the clothing parameters, the posture parameters, and the three-dimensional data are input into a preset condition human body analysis network to obtain a part segmentation result map.

[0113] It can be seen that the electronic device described in the embodiment of the present application obtains human image data, performs feature extraction on the human image data, obtains a feature map, performs feature mapping on the feature map, and constructs a graph structure based on the result after feature mapping. Each graph structure includes multiple nodes, each node corresponds to an attribute identifier of a dimension, and a multi-layer perceptron is used to decode each node of the graph structure to obtain multiple attribute information. Since each node represents the semantic features of each part of the human body that can be learned, after the structural features are aggregated and extracted by graph convolution, the multi-layer perceptron is used to decode each node, so as to reasonably decouple and identify the attributes of each part, and finally obtain the prediction result efficiently and accurately, which helps to improve the accuracy of pedestrian attribute recognition.

[0114] Figure 4 4 is a functional unit block diagram of a pedestrian attribute recognition device 400 involved in an embodiment of the present application. The pedestrian attribute recognition device 400 is applied to an electronic device, and the device 400 includes: an acquisition unit 401, an extraction unit 402, a mapping unit 403 and a decoding unit 404, wherein:

[0115] The acquisition unit 401 is used to acquire human body image data;

[0116] The extraction unit 402 is used to extract features from the human body image data to obtain a feature map;

[0117] The mapping unit 403 is used to perform feature mapping on the feature graph and construct a graph structure based on the result of feature mapping, each graph structure includes a plurality of nodes, and each node corresponds to an attribute identifier of a dimension;

[0118] The decoding unit 404 is used to decode each node of the graph structure using a multi-layer perceptron to obtain multiple attribute information.

[0119] Optionally, in the aspect of performing feature mapping on the feature graph, the mapping unit 403 is specifically used for:

[0120] Get the preset mapping parameter matrix;

[0121] Convert the feature dimension of the feature graph into a preset dimension according to the preset mapping parameter matrix and the number of nodes in the feature graph to obtain a conversion matrix;

[0122] Determine a transition characteristic matrix according to the conversion matrix;

[0123] According to the transition feature matrix and the preset trainable transposed matrix, the feature map is mapped into graph nodes of multiple dimensions representing the semantics of various parts of the human body, and the graph nodes of the multiple dimensions are used as the result after feature mapping.

[0124] Optionally, in the aspect of constructing the graph structure based on the result after feature mapping, the mapping unit 403 is specifically used for:

[0125] According to the inherent properties of the human body, the connection between the parts of the human body is used to encode the relationship between two nodes to obtain a node edge set, and a graph structure is constructed based on the node edge set.

[0126] Optionally, the device 400 is further specifically used for:

[0127] Acquire a first image, wherein the first image includes a person;

[0128] Acquire a human body image, clothing parameters, and posture parameters in the first image;

[0129] Inputting the human body image, the clothing parameters and the posture parameters into a preset condition human body analysis network to obtain a part segmentation result map;

[0130] The part segmentation result map, the clothing parameters and the posture parameters are input into an adaptive generative adversarial network to obtain an adaptive synthetic map, and the adaptive synthetic map and / or the human body image are used as the human body image data.

[0131] Optionally, the device 400 is further specifically used for:

[0132] Obtaining the three-dimensional data of the character;

[0133] In the aspect of inputting the human body image, the clothing parameters and the posture parameters into the preset condition human body analysis network to obtain the part segmentation result map, the device 400 is further specifically used for:

[0134] The human body image, the clothing parameters, the posture parameters, and the three-dimensional data are input into a preset condition human body analysis network to obtain a part segmentation result map.

[0135] It can be seen that the pedestrian attribute recognition device described in the embodiment of the present application obtains human image data, performs feature extraction on the human image data, obtains a feature map, performs feature mapping on the feature map, and constructs a graph structure based on the result after feature mapping. Each graph structure includes multiple nodes, each node corresponds to an attribute identifier of a dimension, and a multi-layer perceptron is used to decode each node of the graph structure to obtain multiple attribute information. Since each node represents the semantic features of each part of the human body that can be learned, after the structural features are aggregated and extracted by graph convolution, the multi-layer perceptron is used to decode each node, so that the attributes of each part are reasonably decoupled and identified, and finally the prediction results are obtained efficiently and accurately, which helps to improve the accuracy of pedestrian attribute recognition.

[0136] It can be understood that the functions of each program module of the pedestrian attribute recognition device of this embodiment can be specifically implemented according to the method in the above method embodiment, and its specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.

[0137] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, wherein the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments, and the above computer includes an electronic device.

[0138] The embodiment of the present application also provides a computer program product, the computer program product includes a non-transitory computer-readable storage medium storing a computer program, the computer program is operable to cause a computer to execute some or all of the steps of any method described in the method embodiment. The computer program product may be a software installation package, and the computer includes an electronic device.

[0139] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0140] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0141] In the several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the above-mentioned units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0142] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0143] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0144] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the above-mentioned methods of each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or CD-ROM and other media that can store program codes.

[0145] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (English: Read-Only Memory, abbreviated as: ROM), a random access memory (English: Random Access Memory, abbreviated as: RAM), a magnetic disk or an optical disk, etc.

[0146] The embodiments of the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A pedestrian attribute recognition method, characterized in that: The method comprises: Acquiring human body image data; Performing feature extraction on the human body image data to obtain a feature map; The feature graph is feature mapped, and a graph structure is constructed based on the result of feature mapping, each graph structure includes multiple nodes, and each node corresponds to an attribute identifier of a dimension; the graph structure uses a graph convolution operation to communicate and transmit information between nodes, and the graph structure outputs a high-level graph representation for subsequent node decoding, and finally generates prediction results for each attribute; Using a multi-layer perceptron to decode each node of the graph structure to obtain multiple attribute information; Wherein, performing feature mapping on the feature map includes: Get the preset mapping parameter matrix; Convert the feature dimension of the feature graph into a preset dimension according to the preset mapping parameter matrix and the number of nodes in the feature graph to obtain a conversion matrix; Determine a transition characteristic matrix according to the conversion matrix; According to the transition feature matrix and the preset trainable transposed matrix, the feature map is mapped into graph nodes of multiple dimensions representing the semantics of various parts of the human body, and the graph nodes of the multiple dimensions are used as the result after feature mapping.

2. The method according to claim 1, characterized in that The step of constructing a graph structure based on the result of feature mapping includes: According to the inherent properties of the human body, the connection between the parts of the human body is used to encode the relationship between two nodes to obtain a node edge set, and a graph structure is constructed based on the node edge set.

3. The method according to claim 1 or 2, characterized in that: The method further comprises: Acquire a first image, where the first image includes a person; Acquire a human body image, clothing parameters, and posture parameters in the first image; Inputting the human body image, the clothing parameters and the posture parameters into a preset condition human body analysis network to obtain a part segmentation result map; The part segmentation result map, the clothing parameters and the posture parameters are input into an adaptive generative adversarial network to obtain an adaptive synthetic map, and the adaptive synthetic map and / or the human body image are used as the human body image data.

4. The method according to claim 3, characterized in that The method further comprises: Obtaining the three-dimensional data of the character; The step of inputting the human body image, the clothing parameters and the posture parameters into a preset condition human body analysis network to obtain a part segmentation result graph includes: The human body image, the clothing parameters, the posture parameters, and the three-dimensional data are input into a preset condition human body analysis network to obtain a part segmentation result map.

5. A pedestrian attribute recognition device, characterized in that: The device comprises: an acquisition unit, an extraction unit, a mapping unit and a decoding unit, wherein: The acquisition unit is used to acquire human body image data; The extraction unit is used to extract features from the human body image data to obtain a feature map; The mapping unit is used to perform feature mapping on the feature graph and construct a graph structure based on the result after feature mapping, each graph structure includes multiple nodes, and each node corresponds to an attribute identifier of a dimension; the graph structure uses a graph convolution operation to communicate and transmit information between nodes, and the graph structure outputs a high-level graph representation for subsequent node decoding, and finally generates prediction results for each attribute; The decoding unit is used to decode each node of the graph structure using a multi-layer perceptron to obtain multiple attribute information; Wherein, in the aspect of performing feature mapping on the feature map, the mapping unit is specifically used for: Get the preset mapping parameter matrix; Convert the feature dimension of the feature graph into a preset dimension according to the preset mapping parameter matrix and the number of nodes in the feature graph to obtain a conversion matrix; Determine a transition characteristic matrix according to the conversion matrix; According to the transition feature matrix and the preset trainable transposed matrix, the feature map is mapped into graph nodes of multiple dimensions representing the semantics of various parts of the human body, and the graph nodes of the multiple dimensions are used as the result after feature mapping.

6. The device according to claim 5, characterized in that In terms of constructing a graph structure based on the result after feature mapping, the mapping unit is specifically used to: According to the inherent properties of the human body, the connection between the parts of the human body is used to encode the relationship between two nodes to obtain a node edge set, and a graph structure is constructed based on the node edge set.

7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store one or more programs and is configured to be executed by the processor, wherein the program comprises instructions for executing the steps in the method according to any one of claims 1 to 4.

8. A computer-readable storage medium, characterized in that: A computer program for electronic data exchange is stored, wherein the computer program enables a computer to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Pedestrian attribute identification and positioning method and convolutional neural network system

    US20200272902A1