A method, apparatus, electronic device, and storage medium for target object recognition
By acquiring and processing the segmentation mask of historical video images, and using the transformed network layer to generate high-quality masks and gait features, the problem of incomplete gait recognition in complex street scenes is solved, and the recognition accuracy and robustness of target objects are improved.
Patent Information
- Application Number
- CN202510346036.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-24
AI Technical Summary
In the existing gait recognition method, it is difficult to extract complete gait features in complex street scenarios, resulting in a significant decrease in recognition effect.
By obtaining the video images and segmentation masks of the historical target object, the mask feature queue is determined, and the transformed network layer and gait recognition network are used to generate high-quality masks and timing gait features, and the mask loss of the target object is completed, and the recognition accuracy is improved.
In complex street scenarios, the recognition accuracy and robustness of target objects are significantly improved, and the problem of incomplete gait characteristics is solved.
Smart Images

Figure CN119851356B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and particularly to a method, apparatus, electronic device, and storage medium for target object recognition. Background Art
[0002] With the acceleration of the urbanization process, the urban population density is increasing continuously, and urban management and public safety issues are becoming increasingly prominent, requiring further empowerment of the security monitoring industry. Among them, how to quickly and accurately identify and locate key pedestrians in a complex street environment has become an important issue in the fields of urban management and public safety. Traditional monitoring means mainly rely on methods such as person tracking and face recognition, but these methods have many limitations in practical applications.
[0003] For example, in common face recognition, due to the limited viewing angle and clarity of the monitoring camera, especially in long-distance and low-light environments, the accuracy of face recognition drops significantly, and face recognition is difficult to perform in non-cooperative scenarios and is difficult to recognize when pedestrians are unconscious or uncooperative. Secondly, the target tracking method relying only on motion trajectories and appearance features is difficult to maintain effective recognition of specific key persons in complex scenarios with many occlusions or multiple pedestrians.
[0004] Therefore, gait, as a unique and difficult-to-disguise biometric feature, still has a high recognition rate under long-distance, low-resolution, and low-light conditions, and has gradually become a research hotspot. In recent years, with the rapid development of computer vision and deep learning technologies, significant progress has been made in pedestrian recognition technology based on gait recognition. Gait is a biometric feature that presents an individual's unique walking pattern. Compared with other biometric recognition methods such as face, fingerprint, and iris, gait is difficult to disguise and can be easily captured non-invasively at a certain distance. Based on the above advantages, gait recognition has high feasibility in security applications. However, most existing gait recognition methods are based on ideal laboratory environments and are difficult to be directly applied to complex street scenes with serious occlusions and interference from multiple pedestrians. Especially in actual street monitoring scenarios, the body postures and contours of pedestrians are easily occluded (such as partially blocked by other pedestrians, trees, vehicles, etc.), or some parts blend in with the background, resulting in incomplete pedestrian masks being extracted. The missing masks lead to the inability to form a complete gait sequence of pedestrians at consecutive moments, further affecting the extraction of temporal features of gait by the gait recognition model, thus causing a significant decline in the recognition effect.
[0005] Currently, for the problem of incomplete gait features extracted when performing gait recognition on target objects in related technologies, no effective solution has been proposed. Summary of the Invention
[0006] Embodiments of the present application provide a method, an apparatus, an electronic device, and a storage medium for target object recognition, so as to at least solve the problem of incomplete gait features extracted in the related art when performing gait recognition on a target object.
[0007] In a first aspect, embodiments of the present application provide a method for target object recognition.
[0008] In some embodiments, the method for target object recognition includes the following steps:
[0009] Obtain historical video images including a historical target object, determine a first segmentation mask corresponding to the historical target object, and determine historical target object features according to the historical video images and the first segmentation mask;
[0010] Determine a mask feature queue corresponding to the historical target object according to the historical target object features and a first transformation network layer;
[0011] Obtain a video image to be recognized including a target object to be recognized, determine a second segmentation mask corresponding to the target object to be recognized, and determine target object features to be recognized according to the video image to be recognized and the second segmentation mask, where the object identifier of the historical target object is the same as the object identifier of the target object to be recognized;
[0012] Determine a high-quality mask corresponding to the target object to be recognized according to the mask feature queue, the target object features to be recognized, and a second transformation network layer;
[0013] Determine temporal gait features corresponding to the target object to be recognized according to the high-quality mask and a gait recognition network, and determine a key target object according to the temporal gait features.
[0014] In some embodiments, the first transformation network layer includes a mask key transformation network layer and a mask value transformation network layer. The determining of the mask feature queue corresponding to the historical target object according to the historical target object features and the first transformation network layer includes:
[0015] Input the historical target object features into the mask key transformation network layer to determine mask key features;
[0016] Input the historical target object features into the mask value transformation network layer to determine mask value features;
[0017] Determine the mask feature queue corresponding to the historical target object according to the mask key features and the mask value features.
[0018] In some of these embodiments, the second transformation network layer includes a query-key transformation network layer and a query-value transformation network layer. Determining the high-quality mask corresponding to the target object to be recognized based on the mask feature queue, the target object feature to be recognized, and the second transformation network layer includes:
[0019] Inputting the target object feature to be recognized into the query-key transformation network layer to determine query-key features;
[0020] Inputting the target object feature to be recognized into the query-value transformation network layer to determine query-value features;
[0021] Determining the high-quality mask based on the mask key features, the mask value features, the query key features, and the query value features.
[0022] In some of these embodiments, determining the high-quality mask based on the mask key features, the mask value features, the query key features, and the query value features includes:
[0023] Determining an attention weight matrix based on the query key features and the mask key features;
[0024] Determining the high-quality mask based on the attention weight matrix, the query value features, and the mask value features.
[0025] In some of these embodiments, the method further includes:
[0026] Obtaining a video image to be tracked. When the number of target objects included in the video image to be tracked is less than the target objects to be recognized, determining the missing target objects in the video image to be tracked as disappearing target objects, obtaining a video image to be restored including the target objects to be restored, and determining the gait feature corresponding to the target object to be restored as the gait feature to be restored;
[0027] Determining the gait feature corresponding to the disappearing target object as the disappearing target gait feature. When the similarity between the gait feature to be restored and the disappearing target gait feature is greater than or equal to the gait feature similarity threshold, determining the object identifier of the target object to be restored as the same as the object identifier of the disappearing target object.
[0028] In some of these embodiments, there are multiple target objects to be recognized. After determining the temporal gait feature corresponding to the target object to be recognized based on the high-quality mask and the gait recognition network, and determining the key target object based on the temporal gait feature, it further includes:
[0029] Obtain the object inference duration corresponding to the key target object, and determine the non-key target objects included in the multiple to-be-identified target objects when the object inference duration is greater than the inference duration threshold;
[0030] Determine the target objects other than the non-key target objects in the to-be-identified video image as the current to-be-identified target objects.
[0031] In some embodiments, after determining the temporal gait feature corresponding to the to-be-identified target object according to the high-quality mask and the gait recognition network, and determining the key target object according to the temporal gait feature, the following steps are further included:
[0032] Obtain the key target gait feature in the target database, and determine the to-be-identified target object as a suspected key target object when the similarity between the temporal gait feature and the key target gait feature is greater than the feature similarity threshold;
[0033] Add the high-quality mask corresponding to the suspected key target object to the mask feature queue to determine the current mask feature queue.
[0034] In a second aspect, an embodiment of the present application provides an apparatus for identifying a target object.
[0035] In some embodiments, the apparatus for identifying a target object includes a historical feature determination module, a mask queue determination module, a to-be-identified feature determination module, a high-quality mask determination module, and a key target determination module:
[0036] The historical feature determination module is configured to obtain a historical video image including a historical target object, determine a first segmentation mask corresponding to the historical target object, and determine a historical target object feature according to the historical video image and the first segmentation mask;
[0037] The mask queue determination module is configured to determine a mask feature queue corresponding to the historical target object according to the historical target object feature and a first transformation network layer;
[0038] The to-be-identified feature determination module is configured to obtain a to-be-identified video image including a to-be-identified target object, determine a second segmentation mask corresponding to the to-be-identified target object, and determine a to-be-identified target object feature according to the to-be-identified video image and the second segmentation mask, where the object identifier of the historical target object is the same as the object identifier of the to-be-identified target object;
[0039] The high-quality mask determination module is configured to determine a high-quality mask corresponding to the to-be-identified target object according to the mask feature queue, the to-be-identified target object feature, and a second transformation network layer;
[0040] The key target determination module is configured to determine the temporal gait features corresponding to the target object to be recognized according to the high-quality mask and the gait recognition network, and determine the key target object according to the temporal gait features.
[0041] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the target object recognition method described in the first aspect above is implemented.
[0042] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the program is executed by a processor, the target object recognition method described in the first aspect above is implemented.
[0043] Compared with the related art, the target object recognition method, device, electronic device, and storage medium provided by the embodiments of the present application obtain historical video images containing historical target objects, determine the first segmentation mask corresponding to the historical target objects, and determine the historical target object features according to the historical video images and the first segmentation mask; determine the mask feature queue corresponding to the historical target objects according to the historical target object features and the first transformation network layer; obtain the video image to be recognized containing the target object to be recognized, determine the second segmentation mask corresponding to the target object to be recognized, and determine the target object features to be recognized according to the video image to be recognized and the second segmentation mask, where the object identifier of the historical target object is the same as the object identifier of the target object to be recognized; determine the high-quality mask corresponding to the target object to be recognized according to the mask feature queue, the target object features to be recognized, and the second transformation network layer; determine the temporal gait features corresponding to the target object to be recognized according to the high-quality mask and the gait recognition network, and determine the key target object according to the temporal gait features, solving the problem of incomplete gait features extracted in the related art when performing gait recognition on target objects, and improving the accuracy of target object recognition.
[0044] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0046] Figure 1 is a hardware structure block diagram of a terminal for the target object recognition method according to an embodiment of the present application;
[0047] Figure 2 is a flowchart of a target object recognition method according to an embodiment of the present application;
[0048] Figure 3 is a flowchart of another target object recognition method according to an embodiment of the present application;
[0049] Figure 4 is a structural block diagram of a target object recognition device according to an embodiment of the present application. Detailed implementation manners
[0050] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be described and explained below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without making creative efforts fall within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and time-consuming, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes made on the basis of the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.
[0051] Referring to "embodiment" in the present application means that a specific feature, structure or characteristic described in conjunction with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0052] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one kind", "the" and the like involved in this application do not indicate a limitation in quantity and may represent a singular or plural number. The terms "include", "comprise", "have" and any variations thereof involved in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The words such as "connect", "be connected", "couple" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0053] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. Taking running on a terminal as an example, Figure 1 is the hardware structure block diagram of the terminal of the target object recognition method in the embodiment of the present invention. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in the figure is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those shown in Figure 1 the figure, or have a different configuration from that shown in Figure 1 the figure.
[0054] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the target object recognition method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0055] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0056] This embodiment provides a target object recognition method. Figure 2 is a flowchart of the target object recognition method according to the embodiments of the present application, as Figure 2 shown, and the process includes the following steps:
[0057] Step S201, obtain a historical video image containing a historical target object, determine a first segmentation mask corresponding to the historical target object, and determine historical target object features according to the historical video image and the first segmentation mask.
[0058] The video image in the embodiments of the present application may be video data containing a street scene, collected by surveillance cameras distributed in each block, and transmitted to the cloud server in real time through a network for processing. Specifically, after extracting features from each collected video picture using the algorithm in the detection and segmentation network, the detection result is output, where the pedestrian ID detected in the video image is denoted as , the corresponding detection box is denoted as , and the segmented mask result is denoted as .
[0059] Further, a tracking algorithm is used to maintain the IDs of the detection boxes in the above sequence of pictures based on information such as distance and feature similarity, and a historical video image queue is constructed for each ID to store the sequence of historical video images corresponding to the ID, and the pedestrians therein are determined as historical target objects. The historical video image queue corresponding to the historical target object with ID n is denoted as , and the first segmentation mask corresponding to the historical target object is determined. The first segmentation mask queue corresponding to the historical target object with ID n is denoted as . Where m is the length of the historical video image queue, and can be optionally 5. In addition, a detection box queue can be constructed for each ID to store the sequence of the most recent detection boxes. The detection box queue corresponding to the historical target object with ID n is denoted as .
[0060] The historical video images and the first segmentation masks in the historical video image queue corresponding to each historical target object are respectively input into the feature encoding network of the pedestrian mask restoration module. After splicing and then scaling, the historical target object features corresponding to the historical target object can be obtained.
[0061] Step S202: Determine the mask feature queue corresponding to the historical target object according to the historical target object features and the first transformation network layer.
[0062] The feature encoding network of the pedestrian mask restoration module includes a first transformation network layer. The historical target object features are input into the first transformation network layer to determine the mask feature queue corresponding to the historical target object. Wherein, the length of the mask feature queue is the same as that of the historical video image queue.
[0063] Step S203: Obtain the video image to be recognized containing the target object to be recognized, determine the second segmentation mask corresponding to the target object to be recognized, and determine the target object features to be recognized according to the video image to be recognized and the second segmentation mask, where the object identifier of the historical target object is the same as that of the target object to be recognized.
[0064] Obtain the video image to be recognized containing the target object to be recognized, where the object identifier of the target object to be recognized is the same as that of the historical target object. Determine the second segmentation mask corresponding to the target object to be recognized, and input the video image to be recognized and the second segmentation mask into the feature decoding network of the pedestrian mask restoration module. After splicing and then scaling, the target object features to be recognized corresponding to the target object to be recognized can be obtained.
[0065] Step S204: Determine the high-quality mask corresponding to the target object to be recognized according to the mask feature queue, the target object features to be recognized, and the second transformation network layer.
[0066] The feature decoding network of the pedestrian mask restoration module includes a second transformation network layer. The mask feature queue and the feature of the target object to be recognized are input into the second transformation network layer to determine a high-quality mask corresponding to the target object to be recognized.
[0067] Step S205: Determine the temporal gait feature corresponding to the target object to be recognized according to the high-quality mask and the gait recognition network, and determine the key target object according to the temporal gait feature.
[0068] According to the high-quality masks corresponding to multiple frame video images including the target object to be recognized, a high-quality mask sequence is formed. When the length of the high-quality mask sequence of the target object to be recognized reaches the requirement of the number of walking frames required for gait recognition, the high-quality mask sequence is input into the gait recognition network to determine the temporal gait feature corresponding to the target object to be recognized.
[0069] Match and compare the temporal gait feature with the key human gait feature library, calculate the similarity corresponding to the temporal gait feature with the Euclidean distance, and for the temporal gait feature whose similarity reaches the preset gait similarity threshold, determine that the target object to be recognized corresponding to it is the key target object.
[0070] Through the above steps, the embodiment of the present application obtains historical video images including historical target objects, determines the first segmentation mask corresponding to the historical target objects, and determines the historical target object features according to the historical video images and the first segmentation mask, so as to obtain the temporal features based on the historical target object features. According to the historical target object features and the first transformation network layer, a mask feature queue including temporal features corresponding to the historical target object is determined. This mask feature queue can perform missing complementation of the locally missing mask for the video image to be detected. Therefore, further, obtain the video image to be recognized including the target object to be recognized, determine the second segmentation mask corresponding to the target object to be recognized, and determine the target object to be recognized features according to the video image to be recognized and the second segmentation mask. According to the mask feature queue, the target object to be recognized features and the second transformation network layer, determine a more complete high-quality mask corresponding to the target object to be recognized, so as to generate a more complete foreground contour of the target object. According to the high-quality mask and the gait recognition network, input the restored high-quality pedestrian walking mask sequence into the gait recognition network to extract gait features, so as to determine the temporal gait feature corresponding to the target object to be recognized, and determine the key target object according to the temporal gait feature, that is, match and compare with the target database including key target gait features to find the key target object, thus solving the problem that the extracted gait features are incomplete when performing gait recognition on the target object in the related art, and significantly improving the accuracy and robustness of gait recognition and target object recognition.
[0071] In some of these embodiments, the first transformation network layer includes a masked key transformation network layer and a masked value transformation network layer, and step S202 includes:
[0072] Step S2021: Input the historical target object features into the masked key transformation network layer to determine the masked key features.
[0073] Step S2022: Input the historical target object features into the masked value transformation network layer to determine the masked value features.
[0074] In the embodiments of the present application, the first transformation network layer includes a masked key transformation network layer and a masked value transformation network layer. Specifically, the masked key transformation network layer can be a four-layer transformation network layer, sequentially including Layer Normalization (LM), Self-Attention, Layer Normalization and Multilayer Perceptron (MLP); the masked value transformation network layer can be a four-layer transformation network layer with a network structure similar to that of the masked key transformation network layer, but its parameter weights are different from those of the masked key transformation network layer.
[0075] Input the historical target object features into the masked key transformation network layer and the masked value transformation network layer respectively to extract deep features, and perform feature residual transfer after the Self-Attention and the Multilayer Perceptron, so as to determine the masked key features and the masked value features .
[0076] Step S2023: Determine the masked feature queue corresponding to the historical target object according to the masked key features and the masked value features.
[0077] Store in the masked feature queue of the pedestrian corresponding to its tracking, that is, determine the masked feature queue corresponding to the historical target object.
[0078] Through the above steps, the embodiments of the present application provide a specific method for determining the masked feature queue based on the masked key transformation network layer and the masked value transformation network layer, with high feasibility and high masked extraction accuracy.
[0079] In some of these embodiments, the second transformation network layer includes a query key transformation network layer and a query value transformation network layer, and step S204 includes:
[0080] Step S2041: Input the features of the target object to be recognized into the query key transformation network layer to determine the query key features.
[0081] Step S2042: Input the features of the target object to be recognized into the query value transformation network layer to determine the query value features.
[0082] In the embodiments of the present application, the second transformation network layer includes a query-key transformation network layer and a query-value transformation network layer. Specifically, the query-key transformation network layer can be a four-layer transformation network layer, sequentially including Layer Normalization (LN), Self-Attention, Layer Normalization and Multilayer Perceptron (MLP); the query-value transformation network layer can be a four-layer transformation network layer with a network structure similar to that of the query-key transformation network layer, but its parameter weights are different from those of the query-key transformation network layer.
[0083] Input the features of the target object to be recognized into the query-key transformation network layer and the query-value transformation network layer respectively to extract deep features, so as to determine the query-key features and the masked-value features .
[0084] Step S2043: Determine a high-quality mask according to the masked-key features, masked-value features, query-key features and query-value features.
[0085] Furthermore, determine the high-quality mask corresponding to the target object to be recognized according to the masked-key features, masked-value features, query-key features and query-value features.
[0086] Through the above steps, the embodiments of the present application provide a specific method for determining query-key features and query-value features based on the query-key transformation network layer and the query-value transformation network layer, and further determine the high-quality mask corresponding to the target object to be recognized, improving the accuracy of the obtained high-quality mask and having high feasibility.
[0087] In some of the embodiments, step S2043 includes:
[0088] Step S2143: Determine an attention weight matrix according to the query-key features and the masked-key features.
[0089] Specifically, in the embodiments of the present application, perform a dot product operation on the transpose of the query-key features and the masked-value features , input the obtained score matrix into the softmax function, calculate the similarity between the query-key features and the first segmentation mask, and obtain the attention weight matrix for determining where to fuse the value features from the mask feature queue, where Ck is the number of feature dimensions of the query-key features:
[0090]
[0091] Step S2243: Determine a high-quality mask according to the attention weight matrix, the query-value features and the masked-value features.
[0092] In the embodiment of the present application, the attention weight matrix is multiplied by to obtain the result of weighted summation, which represents the association information between the second segmentation mask and the masks in the mask feature queue:
[0093]
[0094] And the association information is added to the current query value feature in a splicing manner to determine the feature to be repaired. The feature to be repaired is input into the segmentation head network in the aforementioned detection segmentation network to obtain a repaired high-quality mask .
[0095] Through the above steps, the embodiment of the present application introduces an attention mechanism to extract and fuse features, so as to obtain a high-quality mask to implement subsequent gait and target recognition, and further improve the accuracy of target recognition.
[0096] This embodiment also provides a target object recognition method. Figure 3 is a flowchart of another target object recognition method according to the embodiment of the present application. As Figure 3 shown, on the basis of steps S201 to S205, this process further includes the following steps:
[0097] Step S301, obtain the video image to be tracked. When the number of target objects included in the video image to be tracked is less than the number of target objects to be recognized, determine the target objects missing in the video image to be tracked as disappearing target objects, obtain the video image to be restored including the target objects to be restored, and determine the gait feature corresponding to the target objects to be restored as the gait feature to be restored.
[0098] In the embodiment of the present application, when the tracking of the target object in the video frame is interrupted due to occlusion or multi-target intersection, the historical gait feature is called, and the historical gait feature is used as a supplement to the feature set required for tracking to maintain the target object ID. Obtain the video image to be tracked, and the video image to be tracked can be an image obtained after the video image to be recognized. The number of target objects included in the video image to be tracked is less than the number of target objects to be recognized in the video image to be recognized, that is, some target objects are not recognized. At this time, determine the target objects missing in the video image to be tracked due to not being recognized as disappearing target objects. Specifically, the disappearing target objects can be moved to the recently disappeared pedestrian queue, and the disappearing target object ID set can be denoted as . Further, continue to obtain the video image after the video image to be tracked. When a new target object appears in the obtained image compared with its previous frame, determine the obtained frame image as the video image to be restored, determine the newly appeared target object in it as the target object to be restored, and determine the gait feature corresponding to the target object to be restored as the gait feature to be restored.
[0099] Step S302: Determine the gait feature corresponding to the disappearing target object as the disappearing target gait feature. When the similarity between the gait feature to be restored and the disappearing target gait feature is greater than or equal to the gait feature similarity threshold, determine the object identifier of the target object to be restored as the same as that of the disappearing target object.
[0100] Determine the gait feature corresponding to the disappearing target object as the disappearing target gait feature. When a new target object to be restored appears, compare its corresponding gait feature to be restored with the gait feature of the disappearing target object that is the closest in the recently disappeared pedestrian queue. Restore the object ID of the target object whose similarity reaches the gait feature similarity threshold, that is, determine the object identifier of the newly appeared target object to be restored as the same as that of the previous disappearing target object. In addition, each time a new target object to be restored appears and is restored and matched with the disappearing target object, if there is a disappearing target object in the recently disappeared pedestrian queue that fails to match the gait feature to be restored, record the number of unmatched times of this disappearing target object. When the number of unmatched times is greater than the preset unmatched threshold, delete this disappearing target object from the recently disappeared pedestrian queue and do not perform the next match.
[0101] Through the above steps, in the embodiment of the present application, when the target object in the picture causes tracking interruption due to occlusion or multi-pedestrian intersection, the historical gait feature is called, and the gait feature is used as a supplement to the feature set required for tracking to maintain the target object ID. When a new target object appears, compare its gait feature with the gait features in the recently disappeared pedestrian queue, and restore the object ID of the target object whose similarity reaches the threshold, so as to alleviate the problem of tracking failure caused by occlusion and intersection and improve the integrity of the target object information.
[0102] In some embodiments, the target object recognition method includes multiple target objects to be recognized. After step S205, it includes:
[0103] Step S206: Obtain the object inference duration corresponding to the determined key target object. When the object inference duration is greater than the inference duration threshold, determine the non-key target objects included in the multiple target objects to be recognized.
[0104] In the embodiment of the present application, the time required from obtaining the video picture to determining the key target object can be determined as the object inference duration. If the object inference duration is greater than the preset inference duration threshold, sort the temporal gait features of the target objects to be recognized according to the similarity with the gait features in the database, and sequentially determine the target objects to be recognized as non-key target objects in ascending order of similarity until the object inference duration is not greater than the preset inference duration threshold. Specifically, when the object inference duration is greater than the preset inference duration threshold, add the non-key target objects to the non-key person queue, and the set of object IDs of the target objects in the non-key person queue is denoted as ; when the inference duration of the object is not greater than the preset inference duration threshold, then pop the non-critical target objects in descending order of gait similarity.
[0105] Step S207: Determine the target objects other than the non-critical target objects in the video image to be recognized as the current target objects to be recognized.
[0106] Determine the target objects other than the non-critical target objects in the video image to be recognized as the current target objects to be recognized. In this way, when obtaining the video image to be recognized containing the target objects to be recognized again, the target objects to be recognized therein will not include the non-critical target objects, that is, the non-critical target objects will no longer perform occluded pedestrian mask restoration and gait recognition, so that the recognition frequency can be reduced to reduce the object inference duration of target object recognition.
[0107] Through the above steps, when there are too many target objects in the picture or the transmission speed is unstable, resulting in gait recognition delay, the embodiments of the present application downsample and recognize non-critical target objects. During the delay, only detect and track the target objects in the non-critical person queue, without performing segmentation and gait recognition, reducing the queue length and recognition frequency of target object gait recognition, improving the recognition efficiency in complex street scenes, and saving resources.
[0108] In some embodiments, after step S205, it further includes:
[0109] Step S208: Obtain the gait features of the key targets in the target database. When the similarity between the temporal gait features and the gait features of the key targets is greater than the feature similarity threshold, determine the target objects to be recognized as suspected key target objects.
[0110] Specifically, obtain the gait features of the key targets in the target database, determine the target objects whose similarity between the temporal gait features and the gait features of the key targets is greater than the feature similarity threshold as suspected key target objects, and add their IDs to the set of suspected key persons, denoted as , when the similarity between the temporal gait features and the gait features of the suspected key target objects is less than the feature similarity threshold, delete it from the set.
[0111] Step S209: Add the high-quality masks corresponding to the suspected key target objects to the mask feature queue to determine the current mask feature queue.
[0112] In the embodiments of the present application, the high-quality mask corresponding to the suspected key target object is added to the mask feature queue to form a new mask feature queue, which is determined as the current mask feature queue and re-input to the mask recovery module to obtain the corresponding high-quality mask. Specifically, the sliding window sequence recognition method can be used for the target objects in the set of suspected key target objects (half of the pictures are repeated between adjacent mask sequences input to the gait recognition network). That is, when the similarity between the temporal gait features output by the gait recognition and the key target gait features is higher than the threshold, the latter half of the mask picture sequence is extracted from the high-quality mask picture queue of the suspected key target object input to the gait recognition network, and is added to the mask feature queue of the suspected key target object ID again for overclocking recognition to accelerate the gait recognition frequency of the suspected key target object.
[0113] In addition, in the embodiments of the present application, multiple key target objects can be determined, and the similarities between the multiple key target objects and the key target gait features in the target database are sorted, and the information of the first n (n can be preset as needed) target objects is continuously maintained for manual inspection to determine whether they are key persons; when the number of times a target object is continuously recognized as a key target object is greater than the preset threshold, a notification "key target object has been found" is issued.
[0114] Through the above steps, when there is a target object in the picture whose similarity with the key target features in the database is higher than the threshold, the suspected key target object is overclocked and recognized, and the sliding window sequence recognition method (there is a certain sequence overlap between windows) is used for overclocking recognition of the target objects in the possible key person queue, so as to adjust the target object recognition frequency, improve the recognition speed of key target objects and improve the efficiency in complex street scenes.
[0115] It should be noted that the steps shown in the above process or the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0116] This embodiment also provides a target object recognition device, which is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0117] Figure 4 is a structural block diagram of the target object recognition device according to the embodiments of the present application, as Figure 4As shown in the figure, the device includes a historical feature determination module 10, a mask queue determination module 20, a feature to be recognized determination module 30, a high-quality mask determination module 40, and a key target determination module 50:
[0118] The historical feature determination module 10 is configured to obtain a historical video image including a historical target object, determine a first segmentation mask corresponding to the historical target object, and determine historical target object features based on the historical video image and the first segmentation mask;
[0119] The mask queue determination module 20 is configured to determine a mask feature queue corresponding to the historical target object according to the historical target object features and the first transformation network layer;
[0120] The feature to be recognized determination module 30 is configured to obtain a video image to be recognized including a target object to be recognized, determine a second segmentation mask corresponding to the target object to be recognized, and determine target object features to be recognized based on the video image to be recognized and the second segmentation mask, wherein the object identifier of the historical target object is the same as the object identifier of the target object to be recognized;
[0121] The high-quality mask determination module 40 is configured to determine a high-quality mask corresponding to the target object to be recognized according to the mask feature queue, the target object features to be recognized, and the second transformation network layer;
[0122] The key target determination module 50 is configured to determine temporal gait features corresponding to the target object to be recognized according to the high-quality mask and the gait recognition network, and determine the key target object according to the temporal gait features.
[0123] It should be noted that the above-mentioned modules can be functional modules or program modules, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned modules can be located in the same processor; or the above-mentioned modules can also be located in different processors in any combination form.
[0124] This embodiment also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0125] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0126] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:
[0127] Obtain a historical video image containing a historical target object, determine a first segmentation mask corresponding to the historical target object, and determine historical target object features based on the historical video image and the first segmentation mask;
[0128] Determine a mask feature queue corresponding to the historical target object according to the historical target object features and the first transformation network layer;
[0129] Obtain a video image to be recognized containing the target object to be recognized, determine a second segmentation mask corresponding to the target object to be recognized, and determine target object features to be recognized based on the video image to be recognized and the second segmentation mask, where the object identifier of the historical target object is the same as the object identifier of the target object to be recognized;
[0130] Determine a high-quality mask corresponding to the target object to be recognized according to the mask feature queue, the target object features to be recognized, and the second transformation network layer;
[0131] Determine the temporal gait features corresponding to the target object to be recognized according to the high-quality mask and the gait recognition network, and determine the key target object according to the temporal gait features.
[0132] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated here.
[0133] In addition, in combination with the target object recognition method in the above embodiments, an embodiment of the present application can be implemented by providing a storage medium. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the target object recognition methods in the above embodiments is implemented.
[0134] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope described in this specification.
[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.
[0136] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for identifying a target object, characterized in that It includes the following steps: Obtain a historical video image containing a historical target object, determine a first segmentation mask corresponding to the historical target object, and determine historical target object features based on the historical video image and the first segmentation mask; Determine a first transformation network layer including a mask key transformation network layer and a mask value transformation network layer, input the historical target object features into the mask key transformation network layer to determine mask key features, and input the historical target object features into the mask value transformation network layer to determine mask value features; Determine a mask feature queue corresponding to the historical target object according to the mask key features and the mask value features; Obtain a video image to be recognized containing a target object to be recognized, determine a second segmentation mask corresponding to the target object to be recognized, and determine target object features to be recognized based on the video image to be recognized and the second segmentation mask, wherein the object identifier of the historical target object is the same as the object identifier of the target object to be recognized; Determine a second transformation network layer including a query key transformation network layer and a query value transformation network layer, input the target object features to be recognized into the query key transformation network layer to determine query key features, and input the target object features to be recognized into the query value transformation network layer to determine query value features; Determine a high-quality mask corresponding to the target object to be recognized according to the mask key features, the mask value features, the query key features and the query value features; Determine the temporal gait features corresponding to the target object to be recognized according to the high-quality mask and the gait recognition network, and determine the key target object according to the temporal gait features.
2. The target object recognition method according to claim 1, wherein, The determining the high-quality mask according to the mask key features, the mask value features, the query key features and the query value features includes: Determine an attention weight matrix according to the query key features and the mask key features; Determine the high-quality mask according to the attention weight matrix, the query value features and the mask value features.
3. The target object recognition method according to claim 1 or 2, characterized in that, The method further includes: Obtain a video image to be tracked. When the number of target objects included in the video image to be tracked is less than that of the target objects to be recognized, determine the missing target objects in the video image to be tracked as disappearing target objects, obtain a video image to be restored containing target objects to be restored, and determine the gait features corresponding to the target objects to be restored as gait features to be restored; Determine the gait features corresponding to the disappearing target objects as disappearing target gait features. When the similarity between the gait features to be restored and the disappearing target gait features is greater than or equal to the gait feature similarity threshold, determine the object identifier of the target object to be restored as the same as the object identifier of the disappearing target object.
4. The target object recognition method according to claim 3, characterized in that There are multiple target objects to be recognized. After determining the temporal gait features corresponding to the target objects to be recognized according to the high-quality mask and the gait recognition network, and determining the key target object according to the temporal gait features, it further includes: Obtain the object inference duration corresponding to the key target object, and determine the non-key target objects included in the multiple to-be-identified target objects when the object inference duration is greater than the inference duration threshold; Determine the target objects other than the non-key target objects in the to-be-identified video image as the current to-be-identified target objects.
5. The target object recognition method according to claim 4, wherein After determining the temporal gait features corresponding to the to-be-identified target objects according to the high-quality mask and the gait recognition network, and determining the key target objects according to the temporal gait features, it further includes: Obtain the key target gait features in the target database, and determine that the to-be-identified target object is a suspected key target object when the similarity between the temporal gait features and the key target gait features is greater than the feature similarity threshold; Add the high-quality mask corresponding to the suspected key target object to the mask feature queue to determine the current mask feature queue.
6. An object recognition device, characterized in that, It includes a historical feature determination module, a mask feature determination module, a mask queue determination module, a to-be-identified feature determination module, a query feature determination module, a high-quality mask determination module, and a key target determination module: The historical feature determination module is used to obtain a historical video image including historical target objects, determine the first segmentation mask corresponding to the historical target objects, and determine historical target object features according to the historical video image and the first segmentation mask; The mask feature determination module is used to determine a first transformation network layer including a mask key transformation network layer and a mask value transformation network layer, input the historical target object features into the mask key transformation network layer to determine mask key features, and input the historical target object features into the mask value transformation network layer to determine mask value features; The mask queue determination module is used to determine the mask feature queue corresponding to the historical target objects according to the mask key features and the mask value features; The to-be-identified feature determination module is used to obtain a to-be-identified video image including to-be-identified target objects, determine the second segmentation mask corresponding to the to-be-identified target objects, and determine to-be-identified target object features according to the to-be-identified video image and the second segmentation mask, where the object identifier of the historical target object is the same as the object identifier of the to-be-identified target object; The query feature determination module is used to determine a second transformation network layer including a query key transformation network layer and a query value transformation network layer, input the to-be-identified target object features into the query key transformation network layer to determine query key features, and input the to-be-identified target object features into the query value transformation network layer to determine query value features; The high-quality mask determination module is used to determine the high-quality mask corresponding to the to-be-identified target objects according to the mask key features, the mask value features, the query key features, and the query value features; The key target determination module is used to determine the temporal gait features corresponding to the to-be-identified target objects according to the high-quality mask and the gait recognition network, and determine the key target objects according to the temporal gait features.
7. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the target object recognition method according to any one of claims 1 to 5.
8. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the target object recognition method according to any one of claims 1 to 5 when running.
Citation Information
Patent Citations
Gait recognition method, gait recognition model training method and device, terminal and storage medium
CN115761879A