Deformable group identification method, device, electronic device and storage medium

By extracting and clustering human key points on traffic scene videos, combined with deformable characteristics judgment, the problems of insufficient group modeling research and low recognition accuracy in the existing technology are solved, and more efficient deformable group recognition is achieved.

CN114529866BActive Publication Date: 2025-05-06YANSHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210032653.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2025-05-06
Estimated Expiration
2042-01-12

AI Technical Summary

Technical Problem

The prior art has the problem of less research on group modeling in pedestrian detection and recognition and semantic understanding, especially under the characteristics of pedestrian deformation, uncertain number and congestion in traffic scenarios, the group recognition accuracy is low.

Method used

By obtaining the video frame to be identified, a single frame image data is extracted, a preset human key point extraction model is input to obtain human key point information, clustering people based on the head key point information, and determining whether there are deformable characteristics in the group, thereby identifying deformable groups.

Benefits of technology

It improves the identification accuracy of deformable groups, can effectively deal with the deformable and congestion problems of pedestrians in traffic scenarios, and provides more accurate population detection and analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529866B_ABST
    Figure CN114529866B_ABST
Patent Text Reader

Abstract

The present application provides a deformable group recognition method, device, electronic device and storage medium, the method comprising: obtaining a video to be recognized, and performing frame extraction on the video to be recognized to obtain single-frame image data; the video to be recognized includes pedestrians; the single-frame image data is input into a preset human key point extraction model to obtain human key point information of all pedestrians; the human key point information includes head key point information; according to the head key point information, the pedestrians are clustered and divided into at least one group; it is determined whether there is a deformable feature in each group, and the group corresponding to the deformable feature is a deformable group. The scheme has high deformable group recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a deformable group recognition method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of artificial intelligence, information transmission and other technologies, intelligent technologies such as autonomous driving, assisted driving, and intelligent transportation have become hot topics in current research and application. Autonomous driving is an important technology to solve traffic congestion in the future, and can greatly improve production efficiency and traffic efficiency. Autonomous driving research mainly includes technologies such as environmental perception, planning and decision-making, and vehicle automatic control. Among them, environmental perception technology mainly detects and identifies targets in traffic scenes, and performs semantic analysis and understanding of traffic scenes to achieve information interaction between intelligent driving vehicles and the external traffic environment. It is the basis for decision-making planning and automatic control of autonomous driving vehicles, and is of great significance for ensuring the safe travel of traffic participants in the traffic environment and improving the safety of intelligent driving vehicles.

[0003] Pedestrians in mixed traffic conditions are the most vulnerable targets among road users. Pedestrian safety is the focus and difficulty of environmental perception research for autonomous vehicles. Groups are a common phenomenon among pedestrians in congested traffic and play an important role in influencing crowd behavior.

[0004] At present, a lot of research has been carried out in the field of pedestrian detection and recognition and semantic understanding, but there are few studies on group modeling. Deformable group modeling is the basis of deformable group congestion detection, deformable group occlusion tracking and trajectory prediction. Due to the characteristics of pedestrians in traffic scenes such as deformability, uncertain number and congestion occlusion, deformable group modeling is challenging.

[0005] Most recent group recognition methods are based on deep learning, which mainly uses a deep convolutional neural network model with an attention mechanism to extract features and send them to a fully connected layer for group detection. However, due to the deformability of pedestrians, congestion and occlusion within the group, the detection accuracy is low. Summary of the invention

[0006] The purpose of the embodiments of this specification is to provide a deformable group recognition method, device, electronic device and storage medium.

[0007] To solve the above technical problems, the embodiments of the present application are implemented in the following ways:

[0008] In a first aspect, the present application provides a deformable group recognition method, the method comprising:

[0009] Obtain a video to be identified, and perform frame extraction on the video to be identified to obtain single-frame image data; the video to be identified includes pedestrians;

[0010] Input the single-frame image data into the preset human key point extraction model to obtain the human key point information of all pedestrians; the human key point information includes the head key point information;

[0011] According to the key point information of the head, the pedestrians are clustered and divided into at least one group;

[0012] Determine whether there is a deformable feature in each group, and the group corresponding to the deformable feature is a deformable group. The determination of whether there is a deformable feature in each group includes: determining the number of key points in the same part of all pedestrians based on the human body key point information of all pedestrians; determining the mean and standard deviation based on the number of key points in the same part of all pedestrians; determining whether there is occlusion inside the group based on the mean and standard deviation, and if there is occlusion, determining that there is a deformable feature in the group.

[0013] In one of the embodiments, the head key point information includes head coordinates;

[0014] According to the key point information of the head, the pedestrians are clustered and divided into at least one group, including:

[0015] According to the head coordinates, the attraction and repulsion potentials between pedestrians are calculated;

[0016] By comparing the size of the attraction potential and the repulsion potential, the pedestrians are clustered. Through repeated iterations, the iteration is terminated until there are no remaining unclassified pedestrians, and the pedestrians are divided into at least one group.

[0017] In one embodiment, the attraction potential and repulsion potential between pedestrians are calculated based on the head coordinates, including:

[0018] Select any head coordinate from all head key point information as the first head coordinate;

[0019] The attraction potential and repulsion potential of the first head coordinate and the second head coordinate are calculated respectively; the second head coordinate is any head coordinate among all the head coordinates except the first head coordinate.

[0020] In one embodiment, clustering pedestrians is performed by comparing the magnitudes of the attraction potential and the repulsion potential, including:

[0021] If the attraction potential between the first head coordinate and the second head coordinate is greater than or equal to the repulsion potential, the first head coordinate and the second head coordinate are of the same category, and the pedestrian corresponding to the first head coordinate and the pedestrian corresponding to the second head coordinate are divided into one group;

[0022] If the attraction potential between the first head coordinate and the second head coordinate is smaller than the repulsion potential, the first head coordinate and the second head coordinate are not of the same type, and the first head coordinate and the second head coordinate do not belong to the same group.

[0023] In one of the embodiments, the method further includes: if there are missed detections in the pedestrian's human body key point information, supplementing the missed human body key point information and setting the missed key point coordinates to null values.

[0024] In one embodiment, the key points of the human body include right ankle, right knee, right hip, left hip, left knee, left ankle, pelvis, chest, upper neck, head, right wrist, right elbow, right shoulder, left shoulder, left elbow, and left wrist.

[0025] In a second aspect, the present application provides a deformable group recognition device, the device comprising:

[0026] The acquisition module is used to acquire the video to be identified and perform frame extraction on the video to be identified to obtain single-frame image data; the video to be identified includes pedestrians;

[0027] A determination module is used to input the single-frame image data into a preset human key point extraction model to obtain the human key point information of all pedestrians; the human key point information includes the head key point information;

[0028] A clustering module is used to cluster pedestrians according to the key point information of the head, and divide the pedestrians into at least one group;

[0029] The processing module is used to determine whether there are deformable features in each group. The group corresponding to the deformable features is a deformable group. The determination of whether there are deformable features in each group includes: determining the number of key points in the same part of all pedestrians based on the human body key point information of all pedestrians; determining the mean and standard deviation based on the number of key points in the same part of all pedestrians; and determining whether there is occlusion inside the group based on the mean and standard deviation. If there is occlusion, it is determined that there are deformable features in the group.

[0030] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the deformable group recognition method according to the first aspect is implemented.

[0031] In a fourth aspect, the present application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the deformable group recognition method as in the first aspect.

[0032] It can be seen from the technical solution provided in the above embodiments of this specification that:

[0033] By inputting the single-frame image data extracted from the video frame to be identified into the preset human key point extraction model, the human key point information is obtained, and then the pedestrians are divided into at least one group according to the head key point information in the human key point information. Finally, by determining whether there are deformable features in each group, the deformable group is identified with high recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0035] Figure 1 A schematic diagram of the process of the deformable group recognition method provided in this application;

[0036] Figure 2 A flow chart for generating a deformable group detection dataset provided by this application;

[0037] Figure 3 Schematic diagram of key points of the human body provided for this application;

[0038] Figure 4 A schematic diagram of the structure of the deformable group recognition device provided in this application;

[0039] Figure 5 A schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION

[0040] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0041] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0042] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments of the present application description without departing from the scope or spirit of the present application. Other embodiments derived from the present application description will be apparent to those skilled in the art. The present application description and examples are exemplary only.

[0043] The words “include,” “including,” “have,” “contain,” etc. used in this document are open-ended terms, meaning including but not limited to.

[0044] Unless otherwise specified, "parts" in this application are all calculated by mass.

[0045] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0046] Reference Figure 1 , which shows a flow chart applicable to the deformable group recognition method provided in the embodiment of the present application. It can be understood that the deformable group recognition method provided in the embodiment of the present application can be applied to deformable group detection in complex traffic scenes where people and vehicles are mixed, and can also be applied to scenes such as video surveillance of crowds in complex indoor and outdoor scenes.

[0047] like Figure 1 As shown, the deformable group recognition method may include:

[0048] S110, obtaining a video to be identified, and performing frame extraction on the video to be identified to obtain single-frame image data; the video to be identified includes pedestrians.

[0049] The video to be identified can be a video of a complex traffic scene with mixed pedestrians and vehicles, or a surveillance video of a crowd in a complex indoor and outdoor scene, etc., which varies with different usage scenarios and is not limited here. The video to be identified includes multiple pedestrians, and the characteristics of the multiple pedestrians (such as height, body shape, gender, etc.) can be the same or different, which is not limited here. The video to be identified can be acquired in real time through an acquisition device, or collected from the Internet, stored locally, or acquired from a public data set, etc., which is not limited here.

[0050] The video to be identified may be cut into 5-second segments at a frame rate of 25 frames per second to obtain single-frame image data.

[0051] S120, inputting the single-frame image data into a preset human key point extraction model to obtain human key point information of all pedestrians; the human key point information includes head key point information.

[0052] The preset human key point extraction model is a pre-trained network model. The training of the preset human key point extraction model can be carried out through the following steps:

[0053] 1) Establish a deformable group detection dataset

[0054] Specifically, a deformable crowd detection dataset can be established by using public datasets (such as the open source human posture dataset MPII), online collection, and local collection. The dataset can be divided into a training set and a test set at a ratio of 4:1. To filter public datasets, you can select public datasets with dense pedestrians; to collect videos online, you can enter crowd detection keywords on the website and download related videos. For data diversity, replace keywords and repeat the search multiple times; to collect videos locally, you can collect traffic scene videos with dense pedestrians.

[0055] like Figure 2 As shown, for public datasets, public dataset parsing and then data label extraction can be performed. For videos collected on the Internet, download and save mixed traffic videos (and repeat this step by modifying the search keywords), then delete videos with sparse pedestrians and video clips containing noise or blurred images, and perform video frame extraction on the remaining videos. The videos can be cut into 5-second segments at a frame rate of 25 frames per second, and then the coordinates and categories of key points of the human body in each frame of the video can be annotated according to the data annotation format of the public dataset MPII. For locally collected videos, video frame extraction is performed, and the videos can also be cut into 5-second segments at a frame rate of 25 frames per second, and then the individual pedestrians in each frame of the video can be annotated according to the data annotation format of the public dataset MPII, and a deformable group key point detection dataset with 16 key points is constructed. As shown in FIG. Figure 3 As shown, the numbers and names of the 16 key points of the human body are 0-right ankle joint, 1-right knee joint, 2-right hip joint, 3-left hip joint, 4-left knee joint, 5-left ankle joint, 6-pelvis, 7-chest, 8-upper neck, 9-head, 10-right wrist joint, 11-right elbow joint, 12-right shoulder joint, 13-left shoulder joint, 14-left elbow joint, and 15-left wrist joint.

[0056] 2) Train the preset human key point extraction model

[0057] Specifically, the training set images are input into the human body key point extraction network for training to obtain a preset human body key point extraction model.

[0058] It can be understood that before the training set images are input into the human key point extraction model, the images can be preprocessed, wherein the preprocessing may include: normalizing the image size, adjusting the size to 256×256, rotating the image by 90°, 180°, 270°, blurring the image, and adding noise (Gaussian noise, salt and pepper noise) to the image, and other operations.

[0059] Then the preprocessed images of the training set are input into the human key point extraction model. First, the backbone network is used to extract features to obtain feature maps. Then, the feature maps are sent to the confidence and association network branches to obtain the confidence and association of the key points. After obtaining these two pieces of information, the matching algorithm in graph theory is used to obtain local associations. Finally, the joint points of the same person are connected to obtain the key point information of the individual pedestrian.

[0060] The single-frame image data not obtained in S110 is input into the trained preset human key point extraction model, so that the human key point information of all pedestrians in each frame of the image can be obtained, and all the extracted human key point information can be recorded as a set to be identified.

[0061] S130: Cluster the pedestrians according to the key point information of the heads, and divide the pedestrians into at least one group.

[0062] Specifically, the pedestrians are clustered. Since the key points of the head can represent the pedestrians and are less likely to be blocked, the key points of the head are selected to represent the pedestrians. First, the result of the human body key point extraction in step S120 is used to find the coordinates of the key points of the head of each pedestrian. Then, the artificial potential field method is used to cluster the key points of the head to achieve the clustering of pedestrians. The clustering result is the crowd detection result.

[0063] Optionally, step S130 may include:

[0064] According to the head coordinates, the attraction and repulsion potentials between pedestrians are calculated, which may include:

[0065] Select any head coordinate from all head key point information as the first head coordinate;

[0066] The attraction potential and repulsion potential of the first head coordinate and the second head coordinate are calculated respectively; the second head coordinate is any head coordinate among all the head coordinates except the first head coordinate.

[0067] By comparing the size of the attraction potential and the repulsion potential, the pedestrians are clustered, and the iteration is repeated until there are no pedestrians left unclassified, and the iteration is terminated to divide the pedestrians into at least one group. Among them, if the attraction potential of the first head coordinate and the second head coordinate is greater than or equal to the repulsion potential, the first head coordinate and the second head coordinate are of the same category, and the pedestrian corresponding to the first head coordinate and the pedestrian corresponding to the second head coordinate are divided into one group; if the attraction potential of the first head coordinate and the second head coordinate is less than the repulsion potential, the first head coordinate and the second head coordinate are not of the same category, and the first head coordinate and the second head coordinate do not belong to the same group.

[0068] Exemplarily, step S130 may include:

[0069] a) The key points of the head can represent individual pedestrians. First, the key point information of the head is extracted from the key points. From the key point information of the head, any head coordinate is selected and the coordinate of the point is recorded as (x1, y1);

[0070] b) Assume that the head coordinates of the selected head key point information are the initial point (x1, y1) and the head coordinates belong to cluster1 (the first cluster);

[0071] c) Use artificial potential field to cluster individual pedestrians.

[0072] The artificial potential field is used to calculate the initial point (x1, y1) and all other head coordinates (x i ,y i ) and compare the attractive and repulsive potentials. If the attractive potential is greater than or equal to the repulsive potential, then the point (x i ,y i ) and the initial point (x1, y1) belong to the same category, both are classified as cluster1. If the attractive potential is less than the repulsive potential, then point (x i ,y i ) does not belong to the same category as the initial point (x1, y1). i ,y i ) is identified as a remaining point and enters the set to be identified;

[0073] The attractive potential calculation formula is:

[0074]

[0075] Where ξ represents the gravitational gain, (x i ,y i ) is the coordinate of the target point (i.e., the second head coordinate), i takes values ​​of 0, 1, 2, …, n, ρ[(x1, y1), (x i ,y i )] The current point (x1, y1) and the target point (x i ,y i ) between the two sides;

[0076] The repulsive potential calculation formula is:

[0077]

[0078] Where η represents the repulsive force gain, (x i ,y i ) is the coordinate of the target point, i takes values ​​of 0, 1, 2, ..., n. ρ0 represents the distance threshold of the key point, and points greater than this distance will not produce repulsive influence;

[0079] d) For the remaining points that are not classified in the identification set, reselect the initial point, repeat the calculation steps of attraction and repulsion potential, cluster the identification geometry until there are no unclustered head key points, cluster the head key points, and generate m clusters, recorded as cluster1…cluster m .

[0080] Due to occlusion, deformation and other problems in the crowd, the deep learning model for extracting key points of the human body may miss key points of the human body.

[0081] Therefore, in one embodiment, the deformable group recognition further includes: if there are missed detections in the human body key point information of pedestrians, supplementing the missed human body key point information and setting the missed key point coordinates to null values.

[0082] Specifically, the key point information of the pedestrians obtained in S120 is checked to find pedestrians with less than 16 key points, and the key point information that is missed is supplemented with null values ​​(null, null). The supplemented group is represented as C m [(x i ,y i )] wi×16 , where C m Belong to cluster m The pedestrian set, (x i ,y i ) is the key point coordinate, w i is the pedestrian ID, w i =1,2,3…,n.

[0083] S140: Determine whether there is a deformable feature in each group, and the group corresponding to the deformable feature is the deformable group.

[0084] Wherein, judging whether there is a deformable feature in each group may include:

[0085] According to the key point information of all pedestrians, determine the number of key points of the same part of all pedestrians;

[0086] Determine the mean and standard deviation based on the number of key points in the same part of all pedestrians;

[0087] Based on the mean and standard deviation, it is determined whether there is occlusion inside the group. If there is occlusion, it is determined that there are deformable features in the group.

[0088] For example, statistics C m The number of key points of the same part of all individuals in the same part, and the statistical curve C drawn by part q ;

[0089] According to C qFind the local maximum and use it to calculate the mean E q , standard deviation δ q ;

[0090] Determine whether there is occlusion inside the group: m The internal statistical values ​​of all parts are judged, and the formula is:

[0091] N jq -(E q +αδ q )≥T

[0092] Among them, N jq represents the number of key points of the qth part of the jth group, T represents the threshold of the deformation feature, α is a constant, the values ​​of α and T are determined by experiments, j takes values ​​of 0, 1, 2, ..., m; q takes values ​​of 0, 1, 2, ..., 15;

[0093] Count the number of features that meet the deformable condition, recorded as num q ;

[0094] The formula is used to determine whether the group is a deformable group:

[0095]

[0096] Determine whether the group is a deformable group, satisfying If the condition is met, the group has a deformable feature and is judged to be a deformable group in mixed traffic flow;

[0097] Outputs a deformable group collection.

[0098] The embodiment of the present application combines crowd detection with human body posture estimation, realizes human body key point detection through deep learning, avoids manual feature extraction, and reduces the complexity of the algorithm.

[0099] The embodiment of the present application utilizes the theory of artificial potential field to achieve the clustering of pedestrians based on the key points of human body posture, and can accurately divide the multi-pedestrian scene into several pedestrian groups, thus achieving accurate detection of the crowd in mixed traffic scenes. The designed deformable feature detection method can determine whether the group has deformable features and whether there is occlusion problem inside the group by judging the number of feature points.

[0100] The embodiments of the present application can effectively identify people in a video and determine deformable people, thereby preventing the occurrence of dangerous events and having broad application value.

[0101] Reference Figure 4 , which shows a schematic structural diagram of a deformable group recognition device described according to an embodiment of the present application.

[0102] like Figure 4 As shown, the deformable group recognition device 400 may include:

[0103] The acquisition module 410 is used to acquire the video to be identified and perform frame extraction on the video to be identified to obtain single-frame image data; the video to be identified includes pedestrians;

[0104] The determination module 420 is used to input the single-frame image data into a preset human key point extraction model to obtain the human key point information of all pedestrians; the human key point information includes the head key point information;

[0105] A clustering module 430, configured to cluster pedestrians according to the key point information of the head, and divide the pedestrians into at least one group;

[0106] The processing module 440 is used to determine whether there are deformable features in each group. The group corresponding to the deformable features is a deformable group. The determination of whether there are deformable features in each group includes: determining the number of key points in the same part of all pedestrians based on the human body key point information of all pedestrians; determining the mean and standard deviation based on the number of key points in the same part of all pedestrians; determining whether there is occlusion inside the group based on the mean and standard deviation, and if there is occlusion, determining that there are deformable features in the group.

[0107] Optionally, the head key point information includes head coordinates;

[0108] The clustering module 430 is further configured to:

[0109] According to the head coordinates, the attraction and repulsion potentials between pedestrians are calculated;

[0110] By comparing the size of the attraction potential and the repulsion potential, the pedestrians are clustered. Through repeated iterations, the iteration is terminated until there are no remaining unclassified pedestrians, and the pedestrians are divided into at least one group.

[0111] Optionally, the clustering module 430 is further configured to:

[0112] Select any head coordinate from all head key point information as the first head coordinate;

[0113] The attraction potential and repulsion potential of the first head coordinate and the second head coordinate are calculated respectively; the second head coordinate is any head coordinate among all the head coordinates except the first head coordinate.

[0114] Optionally, the clustering module 430 is further configured to:

[0115] If the attraction potential between the first head coordinate and the second head coordinate is greater than or equal to the repulsion potential, the first head coordinate and the second head coordinate are of the same category, and the pedestrian corresponding to the first head coordinate and the pedestrian corresponding to the second head coordinate are divided into one group;

[0116] If the attraction potential between the first head coordinate and the second head coordinate is smaller than the repulsion potential, the first head coordinate and the second head coordinate are not of the same type, and the first head coordinate and the second head coordinate do not belong to the same group.

[0117] The deformable group recognition device 400 may further include: a supplementing module, which is used to supplement the missed human key point information if there is any missed human key point information of pedestrians, and set the missed key point coordinates to null values.

[0118] Optionally, human body key points include right ankle, right knee, right hip, left hip, left knee, left ankle, pelvis, chest, upper neck, head, right wrist, right elbow, right shoulder, left shoulder, left elbow, and left wrist.

[0119] The present embodiment provides a deformable group recognition device that can execute the above method embodiments. Its implementation principle and technical effects are similar and will not be described in detail here.

[0120] Figure 5 FIG. 1 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 5 As shown, a structural schematic diagram of an electronic device 300 suitable for implementing an embodiment of the present application is shown.

[0121] like Figure 5 As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage part 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0122] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.

[0123] In particular, according to the embodiments of the present disclosure, the above reference Figure 1 The described process may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code for executing the above-described deformable group recognition method. In such an embodiment, the computer program may be downloaded and installed from a network via the communication portion 309, and / or installed from a removable medium 311.

[0124] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of the code, and the aforementioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0125] The units or modules involved in the embodiments described in the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor. The names of these units or modules do not constitute limitations on the units or modules themselves in certain circumstances.

[0126] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop, a mobile phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0127] As another aspect, the present application further provides a storage medium, which may be the storage medium included in the aforementioned device in the above embodiment; or may be a storage medium that exists independently and is not assembled into the device. The storage medium stores one or more programs, and the aforementioned programs are used by one or more processors to execute the deformable group recognition method described in the present application.

[0128] Storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0129] It should be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0130] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

Claims

1. A deformable group recognition method, characterized in that: The method comprises: Acquire a video to be identified, and perform frame extraction on the video to be identified to obtain single-frame image data; the video to be identified includes pedestrians; Inputting the single-frame image data into a preset human key point extraction model to obtain human key point information of all the pedestrians; the human key point information includes head key point information; According to the head key point information, the pedestrians are clustered and divided into at least one group; Determine whether there is a deformable feature in each of the groups, and the group corresponding to the deformable feature is a deformable group. The determining whether there is a deformable feature in each of the groups includes: determining the number of key points in the same part of all pedestrians based on the key point information of the human body of all pedestrians; determining the mean and standard deviation based on the number of key points in the same part of all pedestrians; determining whether there is occlusion inside the group based on the mean and standard deviation, and if there is occlusion, determining that there is a deformable feature in the group.

2. The method according to claim 1, characterized in that The head key point information includes head coordinates; The step of clustering the pedestrians according to the head key point information to divide the pedestrians into at least one group includes: Calculating the attraction potential and repulsion potential between the pedestrians according to the head coordinates; The pedestrians are clustered by comparing the attraction potential and the repulsion potential, and the iteration is terminated until there are no remaining unclassified pedestrians, and the pedestrians are divided into at least one group.

3. The method according to claim 2, characterized in that The step of calculating the attraction potential and repulsion potential between the pedestrians according to the head coordinates includes: Select any head coordinate from all the head key point information as the first head coordinate; The attraction potential and repulsion potential of the first head coordinate and the second head coordinate are calculated respectively; the second head coordinate is any head coordinate among all the head coordinates except the first head coordinate.

4. The method according to claim 3, characterized in that The clustering of the pedestrians by comparing the magnitudes of the attraction potential and the repulsion potential comprises: If the attraction potential between the first head coordinate and the second head coordinate is greater than or equal to the repulsion potential, the first head coordinate and the second head coordinate are of the same category, and the pedestrian corresponding to the first head coordinate and the pedestrian corresponding to the second head coordinate are divided into one group; If the attraction potential between the first head coordinate and the second head coordinate is smaller than the repulsion potential, the first head coordinate and the second head coordinate are not of the same type, and the first head coordinate and the second head coordinate do not belong to the same group.

5. The method according to claim 1, characterized in that The method further includes: if there are missed detections in the pedestrian's human body key point information, supplementing the missed human body key point information and setting the missed key point coordinates to null values.

6. The method according to claim 1, characterized in that The key points of the human body include right ankle joint, right knee joint, right hip joint, left hip joint, left knee joint, left ankle joint, pelvis, chest, upper neck, head, right wrist joint, right elbow joint, right shoulder joint, left shoulder joint, left elbow joint, and left wrist joint.

7. A deformable group recognition device, characterized in that: The device comprises: An acquisition module is used to acquire a video to be identified and perform frame extraction on the video to be identified to obtain single-frame image data; the video to be identified includes pedestrians; A determination module, used for inputting the single frame image data into a preset human key point extraction model to obtain human key point information of all the pedestrians; the human key point information includes head key point information; A clustering module, used for clustering the pedestrians according to the head key point information, and dividing the pedestrians into at least one group; The processing module is used to determine whether there is a deformable feature in each of the groups. The group corresponding to the deformable feature is a deformable group. The determination of whether there is a deformable feature in each of the groups includes: determining the number of key points in the same part of all pedestrians based on the human body key point information of all pedestrians; determining the mean and standard deviation based on the number of key points in the same part of all pedestrians; and determining whether there is occlusion inside the group based on the mean and standard deviation. If there is occlusion, it is determined that there is a deformable feature in the group.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the deformable group recognition method as described in any one of claims 1 to 6 is implemented.

9. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the deformable group recognition method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Artificial potential field obstacle avoidance method for omnidirectional-wheel mobile robot

    CN108614561A

  • Method and device for generating shielded face image

    CN111667403A