Target object detection method and device, electronic equipment and storage medium

By performing attention analysis and fusion processing with the contour map on the detection feature map used in gait recognition, a more critical information fusion data map is generated, which solves the problem of low accuracy of existing gait recognition methods and achieves higher gait recognition accuracy.

CN119942633AActive Publication Date: 2025-05-06TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411838111.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-06
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing gait recognition methods use contour or skeleton information as model input, resulting in low accuracy of detection results.

Method used

By obtaining the detection feature maps corresponding to multiple consecutive detection images of the target object, the analysis map is obtained by performing attention analysis, combining the object outline map for fusion processing, generating a fusion data map, and inputting the gait recognition network for identification.

Benefits of technology

The accuracy of gait recognition detection of the target object is improved. By retaining and enhancing the location characteristics of the target object, a fusion data map with more critical information is generated, thereby improving the accuracy of gait recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942633A_ABST
    Figure CN119942633A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a target object detection method and device, electronic equipment and a storage medium, and the method comprises the steps: firstly, obtaining detection feature maps corresponding to a plurality of continuous detection images of a target object, carrying out the attention analysis of the detection feature maps, and obtaining an analysis map which comprises at least one part feature of the target object; then, obtaining an object contour diagram corresponding to each detection image, and carrying out fusion processing on the object contour diagram and the analysis diagram to obtain a fusion data diagram; and finally, inputting the fusion data graph into a gait recognition network for gait recognition to obtain a gait recognition result of the target object, thereby obtaining the fusion data graph with more key information, and effectively improving the accuracy of the gait recognition result when the gait recognition is performed based on the fusion data graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a target object detection method, device, electronic device and storage medium. Background Art

[0002] Gait recognition is a long-range biometric recognition technology that authenticates an individual by capturing and analyzing their unique walking patterns. Compared with other biometric technologies such as face recognition and fingerprint recognition, gait recognition has some unique advantages. For example, gait data can be acquired from a long distance without approaching the collection device, and data can be collected discretely without attracting attention. With these advantages, gait recognition has gradually been widely used in security, identity authentication, social security and other fields.

[0003] In the related art, gait recognition detection for an object usually uses the object's contour or skeleton information as a model input for gait recognition. However, since the contour or skeleton information contains limited key information, the detection results obtained by this gait recognition method have low accuracy. Summary of the invention

[0004] The embodiments of the present application provide a target object detection method, device, electronic device and storage medium, which can improve the detection accuracy of gait recognition of the target object.

[0005] To achieve the above object, a first aspect of an embodiment of the present application provides a target object detection method, the method comprising:

[0006] Acquire detection feature maps corresponding to a plurality of consecutive detection images of the target object, respectively, and perform attention analysis on the detection feature maps to obtain analysis maps, wherein the analysis maps include at least one part feature of the target object;

[0007] Acquire an object contour map corresponding to each of the detection images, and fuse the object contour map with the analytical map to obtain a fused data map;

[0008] The fused data graph is input into a gait recognition network for gait recognition to obtain a gait recognition result of the target object.

[0009] In some embodiments, performing attention parsing on the detection feature graph to obtain a parsed graph includes:

[0010] Flattening the detection feature map to obtain a flattened detection feature;

[0011] Adding a position code to the flattened detection feature to obtain a position detection feature;

[0012] The position detection features are subjected to attention analysis to obtain the analysis graph.

[0013] In some embodiments, adding a position code to the flattened detection feature to obtain a position detection feature includes:

[0014] Get position parameters, index parameters, and feature dimension parameters;

[0015] Based on the ratio of the index parameter to the feature dimension parameter, and based on the preset parameter, exponential processing is performed to obtain the feature parameter;

[0016] Obtaining a position coding feature based on a sine value corresponding to a ratio of the position parameter to the feature parameter;

[0017] The position detection feature is obtained by performing a superposition process based on the position coding feature and the flattening detection feature.

[0018] In some embodiments, performing attention parsing on the position detection features to obtain the parsing graph includes:

[0019] Obtaining a query vector and a mask matrix, and performing probability distribution processing based on the query vector, the position detection feature and the mask matrix to obtain a mask attention parameter;

[0020] Parsing the mask attention parameters to obtain a position mask and a category prediction corresponding to each position detection feature;

[0021] A position mask whose category prediction matches the part of the target object is selected from the plurality of position masks as a target position mask, and the parsing graph corresponding to the target position mask is generated based on matrix multiplication.

[0022] In some embodiments, the position detection feature includes a first position detection feature and a second position detection feature, and the probability distribution processing based on the query vector, the position detection feature and the mask matrix to obtain the mask attention parameter includes:

[0023] Obtaining a feature vector dimension of the position detection feature;

[0024] Based on the product of the first position detection feature and the query vector, divided by the square root of the feature vector dimension, and then added to the mask matrix, an initial mask parameter is obtained;

[0025] Obtaining a mask position parameter based on a product of the initial mask parameter and the second position detection feature;

[0026] The mask position parameters are subjected to probability distribution processing to obtain the mask attention parameters.

[0027] In some embodiments, fusing the object contour map and the parsing map to obtain a fused data map includes:

[0028] Acquire a plurality of contour pixel values ​​in the object contour map, and acquire a plurality of resolution pixel values ​​of the resolution map;

[0029] updating the pixel points of the object contour map one by one according to the contour pixel value and the parsed pixel value of each pixel point;

[0030] The fused data graph is obtained according to the updated object contour graph.

[0031] In some embodiments, updating the pixel points of the object contour map according to the contour pixel value and the parsed pixel value of each pixel point includes:

[0032] When the contour pixel value of the pixel point is one and the parsed pixel value of the pixel point is not zero, obtaining the object part corresponding to the pixel point in the parsed image;

[0033] The part color of the object part is obtained, and the pixel points of the object contour map are colored based on the part color.

[0034] In some embodiments, inputting the fused data graph into a gait recognition network for gait recognition to obtain a gait recognition result of the target object includes:

[0035] The fused data graph is sequentially subjected to two-dimensional convolution, batch normalization and nonlinear processing to obtain fused data features;

[0036] Performing global pooling processing on the fused data features to obtain dimension-reduced fused features;

[0037] The dimension reduction fusion features are used for gait recognition to obtain the gait recognition result.

[0038] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a target object detection device, wherein the module includes:

[0039] A parsing graph acquisition module, used to acquire detection feature graphs corresponding to a plurality of consecutive detection images of the target object, and perform attention parsing on the detection feature graphs to obtain parsing graphs, wherein the parsing graphs include at least one part feature of the target object;

[0040] A data fusion module, used for obtaining an object contour map corresponding to each of the detection images, and fusing the object contour map with the analytical map to obtain a fused data map;

[0041] The gait recognition module is used to input the fusion data graph into a gait recognition network for gait recognition to obtain a gait recognition result of the target object.

[0042] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the target object detection method described in the first aspect when executing the computer program.

[0043] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium, and the storage medium stores a computer program. When the computer program is executed by a processor, it implements the target object detection method described in the first aspect above.

[0044] The target object detection method, device, electronic device and storage medium proposed in the embodiment of the present application include: first, obtaining detection feature maps corresponding to multiple continuous detection images of the target object, and performing attention analysis on the detection feature map to obtain a parsing map, the parsing map includes at least one part feature of the target object; then, obtaining an object contour map corresponding to each detection image, and fusing the object contour map and the parsing map to obtain a fused data map; finally, inputting the fused data map into a gait recognition network for gait recognition to obtain a gait recognition result of the target object. The embodiment of the present application utilizes the acquisition of a parsing map containing part features of the target object to improve the input key information for gait detection, and further fuses the parsing map and the object contour map to further restore and enhance these important shape information while retaining the part features of the target object, thereby obtaining a fused data map with more key information, so that when gait recognition is performed based on the fused data map, the accuracy of the gait recognition result is effectively improved.

[0045] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flow chart of a target object detection method provided by one embodiment of the present application.

[0047] Figure 2 yes Figure 1 Flow chart of step 101 in FIG.

[0048] Figure 3 yes Figure 2Flow chart of step 202 in FIG.

[0049] Figure 4 yes Figure 2 Flow chart of step 203 in FIG.

[0050] Figure 5 yes Figure 4 Flow chart of step 401 in FIG.

[0051] Figure 6 yes Figure 1 Flow chart of step 102 in FIG.

[0052] Figure 7 yes Figure 6 Flowchart of step 602 in FIG.

[0053] Figure 8 yes Figure 1 Flow chart of step 103 in FIG.

[0054] Fig. 9 It is a structural diagram of the Resnet9 network architecture provided in another embodiment of the present application.

[0055] Fig.10 This is a flowchart of gait recognition of a target object provided by another embodiment of the present application.

[0056] Fig.11 It is a structural schematic diagram of a target object detection device provided in yet another embodiment of the present application.

[0057] Fig.12 This is a schematic diagram of the hardware structure of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] It should be noted that although the functional modules are divided in the device schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0061] First, some nouns involved in this application are analyzed:

[0062] Transformer-based detectors are computer vision models that use the Transformer architecture to perform object detection tasks. This type of detector is able to improve detection speed and efficiency without sacrificing accuracy by combining traditional convolutional neural networks (CNNs) with self-attention mechanisms. These detectors usually adopt an encoder-decoder structure, where the encoder is responsible for extracting features from the image, while the decoder is responsible for predicting the location and category of the object. Due to its parallel processing capabilities and effective capture of long-range dependencies, Transformer-based detectors have shown excellent performance in handling multi-target detection in complex scenes.

[0063] RGB is an additive color model, commonly used to display colors on screen. "RGB" stands for red, green, and blue, and different combinations of these three colors can produce a wide range of colors. In the RGB color model, each color (red, green, and blue) has an intensity value, which is usually an integer from 0 to 255. When the intensity of red, green, and blue is set to the maximum value (255), white is obtained; when all intensity values ​​are set to 0, black is obtained. Various other colors are created by combining different intensities of red, green, and blue.

[0064] ResNet-50 (Residual Network 50 layers) is a deep learning convolutional neural network model, which was proposed by Facebook Artificial Intelligence Research Lab (FAIR) in the 2015 paper "Deep Residual Learning for Image Recognition". ResNet-50 is famous for its excellent performance in image recognition tasks, especially for its excellent results in the ImageNet Large Scale Visual Recognition Challenge (ILSVRC). The core innovation of ResNet-50 lies in its residual block design, which solves the gradient vanishing and gradient exploding problems in deep network training by introducing skip connections (also called short-circuit connections). Traditional deep neural networks become more difficult to train as the number of layers increases, and are prone to degradation problems (that is, as the network deepens, the accuracy no longer improves or even decreases). ResNet uses a residual learning framework to make it easier for information and gradients to propagate forward and backward, allowing for effective training of deeper networks.

[0065] Gait recognition is a long-range biometric recognition technology that authenticates an individual by capturing and analyzing their unique walking patterns. Compared with other biometric technologies such as face recognition and fingerprint recognition, gait recognition has some unique advantages. For example, gait data can be acquired from a long distance without approaching the collection device, and data can be collected discretely without attracting attention. With these advantages, gait recognition has gradually been widely used in security, identity authentication, social security and other fields.

[0066] In the related art, gait recognition detection for an object usually uses the object's contour or skeleton information as a model input for gait recognition. However, since the contour or skeleton information contains limited key information, the detection results obtained by this gait recognition method have low accuracy.

[0067] In order to improve the detection accuracy of gait recognition of the target object, the embodiment of the present application utilizes the acquisition of a parsing graph containing part features of the target object to improve the input key information used for gait detection, and further fuses the parsing graph and the object contour graph to further restore and enhance these important shape information while retaining the part features of the target object, thereby obtaining a fused data graph with more key information, so that when gait recognition is performed based on the fused data graph, the accuracy of the gait recognition results can be effectively improved.

[0068] The target object detection method, device, electronic device and storage medium provided in the embodiments of the present application will be further described below. The target object detection method provided in the embodiments of the present application is applied to a server, a processor or an intelligent terminal.

[0069] First, the target object detection method in the embodiment of the present application is described in detail. Figure 1 , is an optional flow chart of the target object detection method provided in an embodiment of the present application, Figure 1 The method in the embodiment may include but is not limited to steps 101 to 103. Figure 1 The order of step 101 to step 103 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0070] Step 101: Obtain detection feature maps corresponding to multiple consecutive detection images of the target object, and perform attention analysis on the detection feature maps to obtain analysis maps.

[0071] The following is a detailed description of step 101.

[0072] In some embodiments, in response to a gait recognition request for a detection image sequence consisting of a plurality of continuous detection images containing a target object, the detection image sequence is input into an encoder for feature extraction. The extracted features of different scales are then input into a backbone network (usually using a ResNet-50 network) for feature fusion as shown in the following formula (1).

[0073] F fused = fuse({Backbone i (I)}) (1)

[0074] Where I∈R B×H×W×3 Represents the input RGB detection image sequence, Backbone i is the operation of the i-th layer backbone network, and fuse is the fusion operation of multiple layers. In addition, B represents the number of data samples, H represents the height of the image (usually in pixels), W represents the width of the image (usually in pixels), and 3 refers to the number of channels of the image.

[0075] Next, the fused features are upsampled to restore a higher resolution detection feature map as shown in the following formula (2).

[0076] F decoded =BilinearUpsample(F fused ) (2)

[0077] Among them, BilinearUpsample is bilinear upsampling, which is a commonly used image processing technique in deep learning. It is mainly used to expand the image size while maintaining the image quality. In convolutional neural networks (CNN), especially in tasks that require image reconstruction (such as semantic segmentation, image restoration, etc.), the upsampling step is very important because it can help restore detail information.

[0078] In this embodiment, a pre-trained model based on M2FP is used. The pre-trained model is based on the Mask2Former architecture and has been improved to adapt to human body analysis. It can segment multiple features from the input RGB image, including body parts such as head and arms, and objects such as watches and backpacks. However, for gait recognition tasks, accessories such as watches and backpacks should be regarded as noise information.

[0079] Next, in order to obtain accurate gait recognition results for the target object, it is necessary to first perform attention analysis on the detection feature map to obtain a parsed map containing the key information of the target object (i.e., at least one part feature of the target object, such as the head, body, arms, legs, and feet of the human body), as described below.

[0080] Reference Figure 2 , performing attention parsing on the detected feature map to obtain a parsing map, including the following steps 201 to 203.

[0081] Step 201: Flatten the detection feature map to obtain a flattened detection feature.

[0082] Step 202: Add position coding to the flattened detection feature to obtain a position detection feature.

[0083] Steps 201 to 202 are described in detail below.

[0084] In some embodiments, after obtaining the detection feature map F corresponding to the detection image sequence decoded After that, first detect the feature map F decoded After flattening, the flattened detection feature is obtained as shown in the following formula (3).

[0085] F flattened =Flatten(F decoded ) (3)

[0086] Among them, Flatten is a flattening operation, which is mainly used to convert multi-dimensional data into one-dimensional data. And,

[0087] ”'

[0088] F flattened ∈R B×(H×W)×C , C is the feature dimension parameter.

[0089] Next, the flattened detection feature F flattened Position encoding is added to obtain position detection features, thereby improving the position information of the features, so as to facilitate the subsequent fusion of the parsing map and the object contour map, as described below.

[0090] Reference Figure 3 , adding position coding to the flattened detection feature to obtain the position detection feature, including the following steps 301 to 304.

[0091] Step 301: Obtain position parameters, index parameters, and feature dimension parameters.

[0092] Step 302: Based on the ratio of the index parameter to the feature dimension parameter and based on the preset parameters, exponential processing is performed to obtain the feature parameters.

[0093] Step 303: Obtain the position coding feature based on the sine value corresponding to the ratio of the position parameter to the feature parameter.

[0094] Step 304: Perform superposition processing based on the position coding feature and the flattened detection feature to obtain the position detection feature.

[0095] Steps 301 to 304 are described in detail below.

[0096] In some embodiments, based on the set position parameters (including x and y), the index parameter i of the channel, the feature dimension parameter C, based on the ratio C of the index parameter i to the feature dimension parameter, and based on the preset parameter (i.e., 10000), exponential processing is performed to obtain the feature parameter 10000. 2i / C ; Further, based on the position parameters (including x and y) and the characteristic parameters 10000 2i / C The sine value corresponding to the ratio of is obtained, and the position encoding feature is shown in the following formula (4).

[0097]

[0098] Next, based on the position encoding feature P pos (x,y) and the flattened detection feature F flattened Perform superposition processing to obtain the position detection feature F pos As shown in the following formula (5).

[0099] F pos =F flattened +P pos (5)

[0100] Through the above steps 301 to 304, the position coding features obtained by sinusoidal processing of the index parameters, feature dimension parameters, and preset parameters are used to perform position coding superposition processing on the flattened detection features to obtain position detection features that effectively improve the position information, thereby facilitating the subsequent fusion of the analytical map and the object contour map.

[0101] Step 203: Perform attention analysis on the position detection features to obtain a parsing graph.

[0102] The following is a detailed description of step 203.

[0103] In some embodiments, when flattening the detection feature F flattened Perform position encoding to obtain position detection feature F pos Afterwards, the unknown detection features will be further analyzed by attention to obtain a parsed graph containing key information of the target object (i.e., at least one part feature of the target object, such as the head, body, arms, legs, and feet of the human body), as described below.

[0104] Reference Figure 4 , performing attention analysis on the position detection features to obtain a parsing graph, including the following steps 401 to 403.

[0105] Step 401: Obtain a query vector and a mask matrix, and perform probability distribution processing based on the query vector, position detection features, and the mask matrix to obtain mask attention parameters.

[0106] Step 401 is described in detail below.

[0107] In some embodiments, based on the query vector Q∈R M×C and mask matrix M, further combined with the query vector Q∈R M×C , position detection feature F pos And the mask matrix M is processed by probability distribution to obtain the attention parameters required for attention analysis in attention processing. M×C represents different parsing targets (including background query, component query and human query), and in this embodiment, more attention is paid to the query of the part features of the target object (such as head, body, arms, legs and feet), so as to improve the key information for gait recognition of the target object. The following will further describe how to obtain the mask attention parameters.

[0108] Reference Figure 5 , probability distribution processing is performed based on the query vector, position detection features and mask matrix to obtain mask attention parameters, including the following steps 501 to 504.

[0109] Step 501: Obtain the feature vector dimension of the position detection feature.

[0110] Step 502: Detect the product of the feature and the query vector based on the first position, divide it by the root of the feature vector dimension, and add the mask matrix to obtain initial mask parameters.

[0111] Step 503: Obtain mask position parameters based on the product of the initial mask parameters and the second position detection feature.

[0112] Step 504: Perform probability distribution processing on the mask position parameters to obtain mask attention parameters.

[0113] Steps 501 to 504 are described in detail below.

[0114] In some embodiments, the feature F is detected based on the determined position pos The feature vector dimension d k , and the position detection feature F pos The first position detection feature K and the second position detection feature V in the first position detection feature K are further divided by the square root of the feature vector dimension based on the transpose of the first position detection feature K and the product of the query vector Q. Add the mask matrix M to get the initial mask parameters Next, based on the product of the initial mask parameter and the second position detection feature V, the mask position parameter is obtained Finally, the mask position parameters are probability-distributed to obtain the mask attention parameters as shown in the following formula (6).

[0115]

[0116] Step 402: parse the mask attention parameters to obtain the position mask and category prediction corresponding to each position detection feature.

[0117] Step 403: Select a position mask whose category prediction matches the part of the target object from the multiple position masks as the target position mask, and generate a parsing graph corresponding to the target position mask based on matrix multiplication.

[0118] Steps 402 to 403 are described in detail below.

[0119] In some embodiments, the Transformer Decoder is then used to fuse and interact the input features with the query vector through a series of attention mechanisms to parse the mask attention parameters to obtain the position mask and category prediction corresponding to each position detection feature, and select the position mask whose category prediction matches the part of the target object (i.e., the head, body, arms, legs, and feet for human gait recognition) from multiple position masks as the target position mask. Finally, a parsing graph corresponding to the target position mask, i.e., the semantic segmentation result, is generated based on matrix multiplication.

[0120] Through the above steps 401 to 403, and steps 501 to 504, the mask attention parameters obtained by calculating the position detection features, the query vector and the mask matrix as shown in formula (6) can be used to more accurately obtain the position masks and category predictions corresponding to different position detection features, and select the position mask matching the part of the target object from multiple position masks to generate a parsing graph, so that only the parsing features corresponding to the part of the target object related to gait recognition are retained to exclude other interference information irrelevant to gait recognition, thereby further improving the accuracy and reliability of subsequent gait recognition.

[0121] Step 102: Obtain an object contour map corresponding to each detection image, and fuse the object contour map and the analysis map to obtain a fused data map.

[0122] The following is a detailed description of step 102.

[0123] In some embodiments, the analytical graph obtained in step 101 focuses on segmenting various parts of the target object, but in the generation process, information such as the object's contour, edge and shape may be lost. However, the unique body shape and posture of the target object are key information in gait recognition. In this embodiment, by combining the analytical graph with the object contour graph, not only can the characteristics of various parts of the human body be retained, but also these important shape information can be restored and enhanced, so that the overall gait characteristics of the target object can be more accurately captured in the subsequent gait recognition process.

[0124] Therefore, in this embodiment, an object contour map corresponding to the target object in the detection image is first obtained, and then the object contour map and the analysis map are fused to obtain a fused data map for subsequent improvement of gait recognition accuracy, as described below.

[0125] Reference Figure 6 , the object contour map and the analytical map are fused to obtain a fused data map, including the following steps 601 to 603.

[0126] Step 601: Acquire multiple contour pixel values ​​in the object contour map, and acquire multiple parsed pixel values ​​of the parsed map.

[0127] Step 602: updating the pixel points of the object contour map one by one according to the contour pixel value and the parsed pixel value of each pixel point.

[0128] Steps 601 to 602 are described in detail below.

[0129] In some embodiments, based on the obtained object contour map and analysis map, the contour pixel value W(x,y) corresponding to each pixel point in the object contour map (i.e., coordinates (x,y)) is first obtained, and the analysis pixel value P(x,y) corresponding to each pixel point in the analysis map (i.e., coordinates (x,y)) is obtained.

[0130] It can be understood that in the object outline image, except for the body part of the target object which is white, the rest of the background parts are black.

[0131] Then, the pixel points of the object contour map are updated one by one according to the contour pixel value W(x,y) and the parsed pixel value P(x,y) of each pixel point (ie, coordinates (x,y)), as described below.

[0132] Reference Figure 7 , according to the contour pixel value and the parsed pixel value of each pixel, the pixel points of the object contour map are updated, including the following steps 701 to 702.

[0133] Step 701: when the contour pixel value of a pixel point is one and the parsed pixel value of the pixel point is not zero, obtain the object part corresponding to the pixel point in the parsed image.

[0134] Step 702: Obtain the part color of the object part, and color the pixel points of the object contour map based on the part color.

[0135] Steps 701 to 702 are described in detail below.

[0136] In some embodiments, for each pixel point (i.e., coordinates (x, y)), when the contour pixel value of the pixel point is one, i.e., W(x, y) = 1, and the resolution pixel value of the pixel point is not zero, i.e., P(x, y) ≠ 0, the object part k corresponding to the pixel point in the resolution map is obtained, and the part color C of the object part is further determined. k , and then the pixel points of the object contour map are colored based on the part color, as shown in the following formula (7).

[0137]

[0138] Among them, F(x,y) represents the pixel value of the final fused image at (x,y).

[0139] Step 603: Obtain the fused data graph according to the updated object contour graph.

[0140] Step 603 is described in detail below.

[0141] In some embodiments, after traversing all the pixel points and coloring the object contour map according to the matching results, the updated object contour map is used as the fused data map.

[0142] Through the above steps 601 to 603, and steps 701 to 702, the pixel values ​​corresponding to each pixel point in the object contour map and the analytical map are used for matching, and the object contour map is colored with the corresponding part color according to the matching result to complete the fusion of the object contour map and the analytical map, so as to retain the characteristics of various parts of the human body, and restore and enhance these important shape information, so that the overall gait characteristics of the target object can be more accurately captured in the subsequent gait recognition process, so as to obtain a more accurate gait recognition result of the target object.

[0143] Step 103: Input the fused data graph into a gait recognition network to perform gait recognition, and obtain a gait recognition result of the target object.

[0144] The following is a detailed description of step 103.

[0145] In some embodiments, the gait recognition network uses a Gaitbase network to train and recognize the fused images. Gaitbase is a neural network structure designed specifically for gait recognition tasks. It has powerful spatiotemporal feature extraction capabilities by introducing multi-level convolutional neural networks and adaptive feature aggregation modules.

[0146] In order to further improve the accuracy of gait recognition of target objects in practical applications, the gait recognition network needs to be trained.

[0147] In this embodiment, the generated fusion data graph is used as the input of the Gaitbase network for feature learning during the training process. Since these fusion data graphs combine the advantages of contour graphs and analytical graphs, the Gaitbase network can better capture the motion characteristics of each part of the object and avoid interference from external objects. At the same time, the multi-level feature extraction mechanism of the Gaitbase network makes it more robust when processing complex scenes and different perspectives.

[0148] Based on the gait recognition network, after obtaining the fused data graph of the target object in actual application, the fused data graph is further input into the gait recognition network for gait recognition to obtain the gait recognition result of the target object, as described in detail below.

[0149] Reference Figure 8 , inputting the fused data graph into the gait recognition network for gait recognition, and obtaining the gait recognition result of the target object, including the following steps 801 to 803.

[0150] Step 801: The fused data graph is sequentially subjected to two-dimensional convolution, batch normalization and nonlinear processing to obtain fused data features.

[0151] Step 802: Perform global pooling processing on the fused data features to obtain dimension-reduced fused features.

[0152] Step 803: Perform gait recognition on the dimension reduction fusion features to obtain a gait recognition result.

[0153] Steps 801 to 803 are described in detail below.

[0154] In some embodiments, the main line of the gait recognition network adopts the Resnet9 network architecture, referring to Fig. 9 , is a schematic diagram of the structure of a Resnet9 network architecture provided in an embodiment of the present application. Fig. 9 As shown in , the Resnet9 network architecture adopted in this embodiment consists of multiple convolutional layers and residual blocks, and each residual block contains two convolutional layers and a skip connection to prevent the gradient disappearance problem.

[0155] After the fused data graph is input into the gait recognition network, the fused data graph is sequentially subjected to two-dimensional convolution, batch normalization and nonlinear processing to obtain fused data features, and then the fused data features are globally pooled (i.e., first, the time dimension is integrated into the global features of the sequence through time series pooling, and then the feature graph is divided into multiple horizontal areas for pooling to capture local features in the image) to obtain reduced-dimensional fused features, and finally, gait recognition is performed on the reduced-dimensional fused features through a fully connected layer to obtain gait recognition results (i.e., classification and feature mapping).

[0156] Through the above steps 801 to 803, the Gaitbase network with powerful spatiotemporal feature extraction capabilities and the Resnet9 network architecture with powerful feature extraction capabilities while maintaining model efficiency are used to perform gait recognition on the fused data graph obtained from the contour graph and the analytical graph, so as to obtain a more accurate gait recognition result of the target object, thereby greatly improving the accuracy of target object detection.

[0157] Reference Fig.10 , is a flow chart of gait recognition of a target object provided in an embodiment of the present application. Fig.10 As shown in , the gait recognition of the target object is divided into three stages. In the first stage, the detection image is transformed into a human body analysis image through the pre-trained Mask2Former module to exclude noise information. In the second stage, the human body analysis image generated in the first stage and the binary contour map are fused. The two complement each other and can simultaneously retain the detailed features and global shape information of the human body, thereby improving the accuracy and robustness of gait recognition. In the third stage, the spatiotemporal features in the fused image are extracted and input into the gait recognition network to accurately identify the individual's gait pattern and perform identity authentication.

[0158] The target object detection method, device, electronic device and storage medium proposed in the embodiments of the present application include: first, obtaining detection feature maps corresponding to multiple continuous detection images of the target object, flattening the detection feature maps to obtain flattened detection features, obtaining position parameters, index parameters, and feature dimension parameters, performing exponential processing based on the ratio of the index parameters to the feature dimension parameters and based on preset parameters to obtain feature parameters, obtaining position coding features based on the sine value corresponding to the ratio of the position parameters to the feature parameters, performing superposition processing based on the position coding features and the flattened detection features to obtain position detection features, obtaining feature vector dimensions of the position detection features, obtaining initial mask parameters based on the product of the first position detection feature and the query vector, and dividing by the root of the feature vector dimension, and adding a mask matrix, obtaining mask position parameters based on the product of the initial mask parameters and the second position detection feature, performing probability distribution processing on the mask position parameters to obtain mask attention parameters, and parsing the mask attention parameters to obtain each The method comprises the following steps: obtaining a position mask and a category prediction corresponding to each position detection feature, selecting a position mask whose category prediction matches the part of the target object from multiple position masks as the target position mask, and generating a parsing graph corresponding to the target position mask based on matrix multiplication, wherein the parsing graph includes at least one part feature of the target object; then, obtaining an object contour map corresponding to each detection image, obtaining multiple contour pixel values ​​in the object contour map, and obtaining multiple parsing pixel values ​​of the parsing graph, and obtaining the object part corresponding to the pixel point in the parsing graph one by one when the contour pixel value of the pixel point is one and the parsing pixel value of the pixel point is not zero, obtaining the part color of the object part, and coloring the pixel points of the object contour map based on the part color, and obtaining a fused data map according to the updated object contour map; finally, performing two-dimensional convolution, batch normalization and nonlinearization processing on the fused data map in sequence to obtain a fused data feature, performing global pooling processing on the fused data feature to obtain a reduced-dimensional fused feature, performing gait recognition on the reduced-dimensional fused feature, and obtaining a gait recognition result of the target object.

[0159] The embodiment of the present application utilizes the mask attention parameters obtained by calculating the position detection features, the query vector and the mask matrix as shown in formula (6), so that the position masks and category predictions corresponding to different position detection features can be obtained more accurately, and the position mask matching the part of the target object is selected from multiple position masks to generate a parsing graph, so that only the parsing features corresponding to the part of the target object related to gait recognition are retained to exclude other interference information irrelevant to gait recognition, thereby further improving the accuracy and reliability of subsequent gait recognition; and, the pixel values ​​corresponding to each pixel point in the object contour map and the parsing graph are matched, and the corresponding part color is matched to the object contour map according to the matching result. The lines are colored to complete the fusion of the object contour map and the analytical map, so that the characteristics of various parts of the human body can be retained, and these important shape information can be restored and enhanced, so that the overall gait characteristics of the target object can be more accurately captured in the subsequent gait recognition process, so as to obtain a more accurate gait recognition result of the target object; and, the Gaitbase network with powerful spatiotemporal feature extraction capabilities and the Resnet9 network architecture with powerful feature extraction capabilities while maintaining model efficiency are used to perform gait recognition on the fused data map obtained from the contour map and the analytical map, so as to obtain a more accurate gait recognition result of the target object, thereby greatly improving the accuracy of target object detection.

[0160] The present application also provides a target object detection device, which can implement the above target object detection method. Fig.11 , the device 1100 comprises:

[0161] A parsing graph acquisition module 1110 is used to acquire detection feature graphs corresponding to a plurality of consecutive detection images of a target object, and perform attention parsing on the detection feature graph to obtain a parsing graph, wherein the parsing graph includes at least one part feature of the target object;

[0162] The data fusion module 1120 is used to obtain an object contour map corresponding to each detection image, and fuse the object contour map with the analytical map to obtain a fused data map;

[0163] The gait recognition module 1130 is used to input the fused data graph into the gait recognition network to perform gait recognition and obtain the gait recognition result of the target object.

[0164] In some embodiments, the parsing graph acquisition module 1110 is further used to:

[0165] Flatten the detection feature map to obtain a flattened detection feature;

[0166] Add position encoding to the flattened detection feature to obtain a position detection feature;

[0167] The position detection features are subjected to attention parsing to obtain a parsing graph.

[0168] In some embodiments, the parsing graph acquisition module 1110 is further used to:

[0169] Get position parameters, index parameters, and feature dimension parameters;

[0170] Based on the ratio of the index parameter to the feature dimension parameter, exponential processing is performed based on the preset parameters to obtain the feature parameter;

[0171] Based on the sine value corresponding to the ratio of the position parameter and the feature parameter, the position coding feature is obtained;

[0172] The position detection feature is obtained by superimposing the position coding feature and the flattened detection feature.

[0173] In some embodiments, the parsing graph acquisition module 1110 is further used to:

[0174] Obtain the query vector and mask matrix, and perform probability distribution processing based on the query vector, position detection features and mask matrix to obtain mask attention parameters;

[0175] The mask attention parameters are parsed to obtain the position mask and category prediction corresponding to each position detection feature;

[0176] A position mask whose category prediction matches the part of the target object is selected from multiple position masks as the target position mask, and a parsing graph corresponding to the target position mask is generated based on matrix multiplication.

[0177] In some embodiments, the parsing graph acquisition module 1110 is further used to:

[0178] Get the feature vector dimension of the position detection feature;

[0179] The product of the feature and the query vector is detected based on the first position, and divided by the root of the feature vector dimension, and then added to the mask matrix to obtain the initial mask parameters;

[0180] Obtaining mask position parameters based on the product of the initial mask parameters and the second position detection feature;

[0181] The mask position parameters are processed with probability distribution to obtain the mask attention parameters.

[0182] In some embodiments, the data fusion module 1120 is further configured to:

[0183] Obtaining a plurality of contour pixel values ​​in the object contour map, and obtaining a plurality of resolution pixel values ​​of the resolution map;

[0184] The pixel points of the object contour map are updated one by one according to the contour pixel value and the parsed pixel value of each pixel point;

[0185] A fused data map is obtained according to the updated object contour map.

[0186] In some embodiments, the data fusion module 1120 is further configured to:

[0187] When the contour pixel value of the pixel point is one and the resolution pixel value of the pixel point is not zero, the object part corresponding to the pixel point in the resolution image is obtained;

[0188] The part color of the object part is obtained, and the pixel points of the object contour map are colored based on the part color.

[0189] In some embodiments, the gait recognition module 1130 is further configured to:

[0190] The fused data graph is sequentially subjected to two-dimensional convolution, batch normalization, and nonlinear processing to obtain fused data features;

[0191] Perform global pooling on the fused data features to obtain dimension-reduced fused features;

[0192] The dimension reduction fusion features are used for gait recognition to obtain the gait recognition result.

[0193] In the above embodiments, the description of each embodiment has its own emphasis. For the part not described in detail in a certain embodiment, the specific implementation of the target object detection device is basically the same as the specific implementation of the above target object detection method, and will not be repeated here.

[0194] In the embodiment of the present application, the target object detection device uses the mask attention parameters obtained by calculating the position detection features, the query vector and the mask matrix as shown in formula (6), so as to more accurately obtain the position masks and category predictions corresponding to different position detection features, and select the position mask matching the part of the target object from multiple position masks to generate a parsing diagram, so that only the parsing features corresponding to the part of the target object related to gait recognition are retained to exclude other interference information irrelevant to gait recognition, thereby further improving the accuracy and reliability of subsequent gait recognition; and, the pixel values ​​corresponding to each pixel point in the object contour map and the parsing diagram are matched, and the corresponding part color is matched according to the matching result. The contour map is colored to complete the fusion of the object contour map and the analytical map, so that the characteristics of various parts of the human body can be retained, and these important shape information can be restored and enhanced, so that the overall gait characteristics of the target object can be captured more accurately in the subsequent gait recognition process, so as to obtain a more accurate gait recognition result of the target object; and, the Gaitbase network with powerful spatiotemporal feature extraction capabilities and the Resnet9 network architecture with powerful feature extraction capabilities while maintaining model efficiency are used to perform gait recognition on the fused data map obtained by the contour map and the analytical map, so as to obtain a more accurate gait recognition result of the target object, thereby greatly improving the accuracy of target object detection.

[0195] The present application also provides an electronic device, including:

[0196] at least one memory;

[0197] at least one processor;

[0198] at least one program;

[0199] The program is stored in the memory, and the processor executes the at least one program to implement the target object detection method implemented in the present application. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.

[0200] See also Fig.12 , Fig.12 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0201] The processor 1201 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0202] The memory 1202 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device or RAM (Random Access Memory). The memory 1202 can store operating systems and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1202, and the processor 1201 is used to call and execute the target object detection method of the embodiment of the present application;

[0203] Input / output interface 1203, used to implement information input and output;

[0204] The communication interface 1204 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);

[0205] A bus 1205 that transmits information between various components of the device (e.g., the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204);

[0206] The processor 1201 , the memory 1202 , the input / output interface 1203 and the communication interface 1204 are connected to each other in communication within the device via the bus 1205 .

[0207] An embodiment of the present application further provides a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the above-mentioned target object detection method is implemented.

[0208] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0209] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0210] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0211] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0212] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0213] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0214] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0215] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0216] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0217] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0218] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0219] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A target object detection method, characterized in that: include: Acquire detection feature maps corresponding to a plurality of consecutive detection images of the target object, respectively, and perform attention analysis on the detection feature maps to obtain analysis maps, wherein the analysis maps include at least one part feature of the target object; Acquire an object contour map corresponding to each of the detection images, and fuse the object contour map with the analytical map to obtain a fused data map; The fused data graph is input into a gait recognition network for gait recognition to obtain a gait recognition result of the target object.

2. The target object detection method according to claim 1, characterized in that: The step of performing attention parsing on the detection feature graph to obtain a parsed graph includes: Flattening the detection feature map to obtain a flattened detection feature; Adding a position code to the flattened detection feature to obtain a position detection feature; The position detection features are subjected to attention analysis to obtain the analysis graph.

3. The target object detection method according to claim 2, characterized in that: The adding position coding to the flattened detection feature to obtain the position detection feature includes: Get position parameters, index parameters, and feature dimension parameters; Based on the ratio of the index parameter to the feature dimension parameter, and based on the preset parameter, exponential processing is performed to obtain the feature parameter; Obtaining a position coding feature based on a sine value corresponding to a ratio of the position parameter to the feature parameter; The position detection feature is obtained by performing a superposition process based on the position coding feature and the flattening detection feature.

4. The target object detection method according to claim 2, characterized in that: The step of performing attention parsing on the position detection feature to obtain the parsing graph includes: Obtaining a query vector and a mask matrix, and performing probability distribution processing based on the query vector, the position detection feature and the mask matrix to obtain a mask attention parameter; Parsing the mask attention parameters to obtain a position mask and a category prediction corresponding to each position detection feature; A position mask whose category prediction matches the part of the target object is selected from the plurality of position masks as a target position mask, and the parsing graph corresponding to the target position mask is generated based on matrix multiplication.

5. The target object detection method according to claim 4, characterized in that: The position detection feature includes a first position detection feature and a second position detection feature, and the probability distribution processing is performed based on the query vector, the position detection feature and the mask matrix to obtain the mask attention parameter, including: Obtaining a feature vector dimension of the position detection feature; Based on the product of the first position detection feature and the query vector, divided by the square root of the feature vector dimension, and then added to the mask matrix, an initial mask parameter is obtained; Obtaining a mask position parameter based on a product of the initial mask parameter and the second position detection feature; The mask position parameters are subjected to probability distribution processing to obtain the mask attention parameters.

6. The target object detection method according to claim 1, characterized in that: The step of fusing the object contour map and the analytical map to obtain a fused data map includes: Acquire a plurality of contour pixel values ​​in the object contour map, and acquire a plurality of resolution pixel values ​​of the resolution map; updating the pixel points of the object contour map one by one according to the contour pixel value and the parsed pixel value of each pixel point; The fused data graph is obtained according to the updated object contour graph.

7. The target object detection method according to claim 6, characterized in that: The updating of the pixel points of the object contour map according to the contour pixel value and the parsed pixel value of each pixel point comprises: When the contour pixel value of the pixel point is one and the parsed pixel value of the pixel point is not zero, obtaining the object part corresponding to the pixel point in the parsed image; The part color of the object part is obtained, and the pixel points of the object contour map are colored based on the part color.

8. The target object detection method according to claim 1, characterized in that: The step of inputting the fused data graph into a gait recognition network for gait recognition to obtain a gait recognition result of the target object includes: The fused data graph is sequentially subjected to two-dimensional convolution, batch normalization and nonlinear processing to obtain fused data features; Performing global pooling processing on the fused data features to obtain dimension-reduced fused features; The dimension reduction fusion features are used for gait recognition to obtain the gait recognition result.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the target object detection method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target object detection method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Gait recognition method and device based on body type transformation and storage medium

    CN115240269A

  • Deep learning based robot target recognition and motion detection method, storage medium and apparatus

    US11763485B1

  • End-to-end multimodal gait recognition method based on deep learning

    US20220343686A1