A method and device for crowd positioning, an electronic device, and a storage medium
By locating the key points of the human body and filtering the target neighborhood of the population images, the problem of low accuracy of population positioning in dense scenarios is solved, and high-precision population positioning and behavior analysis is achieved.
Patent Information
- Application Number
- CN202210146593.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-02-17
AI Technical Summary
The prior art has low accuracy in population positioning, especially in dense scenarios, and it is difficult to effectively distinguish different pedestrians, resulting in a high false detection rate.
By locating the human body key points on the population image, an initial positioning map is obtained, and the target neighborhood is determined based on the position of the initial human body key points, filtering to obtain the target positioning map. This method uses target neighborhoods of different sizes to reduce the probability of false detection and improve positioning accuracy.
The accuracy of population positioning in dense scenarios is achieved, the probability of false detection of multiple initial human key points corresponding to the same human body is reduced, and the accuracy of population counting and behavior analysis is improved.
Smart Images

Figure CN114550086B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and apparatus for crowd localization, an electronic device, and a storage medium. Background Art
[0002] With the growth of the population and the acceleration of the urbanization process, the behavior of a large number of people gathering is increasing and the scale is getting larger. Crowd analysis is of great significance for public safety and urban planning. Common crowd analysis tasks include crowd counting, group behavior analysis, crowd localization, etc. Among them, crowd localization is the basis for other crowd analysis tasks. Crowd localization refers to estimating the positions of head key points included in an image or video through a computer vision algorithm, and determining the coordinates of the head key points included in the image or video, so as to provide a data basis for subsequent crowd analysis tasks such as crowd counting and group behavior analysis. The accuracy of crowd localization directly affects the accuracy of crowd counting and the results of crowd behavior analysis. Therefore, there is an urgent need for a crowd localization method with a relatively high accuracy. Summary of the Invention
[0003] The present disclosure provides a technical solution for a method and apparatus for crowd localization, an electronic device, and a storage medium.
[0004] According to one aspect of the present disclosure, there is provided a method for crowd localization, including: performing human key point localization on a crowd image to obtain an initial localization map corresponding to the crowd image, where the initial localization map is used to indicate the positions of initial human key points included in the crowd image; determining a target neighborhood corresponding to the initial human key points based on the positions of the initial human key points in the crowd image; and filtering the initial localization map based on the target neighborhood corresponding to the initial human key points to obtain a target localization map, where the target localization map is used to indicate the positions of target human key points included in the crowd image.
[0005] In a possible implementation manner, the determining a target neighborhood corresponding to the initial human key points based on the positions of the initial human key points in the crowd image includes: for any one of the initial human key points, determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales at different positions in the crowd image.
[0006] In a possible implementation manner, the initial human key points are initial head key points.
[0007] In a possible implementation, determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image and a preset perspective mapping relationship includes: determining the target image scale corresponding to the position of the initial human head key point in the crowd image based on the preset perspective mapping relationship; determining the height of the head frame corresponding to the initial human head key point based on the target image scale; and determining the target neighborhood corresponding to the initial human head key point based on the height of the head frame corresponding to the initial human head key point.
[0008] In a possible implementation, determining the target neighborhood corresponding to the initial human head key point based on the height of the head frame corresponding to the initial human head key point includes: when the height of the head frame is greater than a preset head frame height threshold, determining the target neighborhood based on a first neighborhood radius; or when the height of the head frame is less than or equal to the preset head frame height threshold, determining the target neighborhood based on a second neighborhood radius, where the first neighborhood radius is greater than the second neighborhood radius.
[0009] In a possible implementation, performing human key point localization on a crowd image to obtain an initial localization map corresponding to the crowd image includes: performing human key point localization on the crowd image to determine a predicted localization map corresponding to the crowd image, where the predicted localization map is used to indicate the prediction confidence of each pixel point in the crowd image being a human key point; and performing image processing on the predicted localization map based on a preset confidence threshold to obtain the initial localization map.
[0010] In a possible implementation, filtering the initial localization map based on the target neighborhood corresponding to the initial human key point to obtain a target localization map includes: for any initial human head key point i, determining whether there is at least one other initial human head key point in the target neighborhood corresponding to the initial human head key point i; when there is at least one other initial human head point j in the target neighborhood, determining the prediction confidence corresponding to the initial human head point i and the prediction confidence corresponding to the at least one other initial human head point j based on the predicted localization map; and determining the target human head key point in the target neighborhood based on the initial human head key point i and the initial human head key point with the maximum prediction confidence among the at least one other initial human head key points j.
[0011] In a possible implementation, the method further includes: determining the position of the target human foot key point included in the crowd image based on the target localization map and the preset perspective mapping relationship.
[0012] In a possible implementation, determining the positions of the target human foot key points included in the crowd image based on the target positioning map and the preset perspective mapping relationship includes: for any one of the target human head key points, determining the first image coordinates of the target human head key point in the crowd image according to the target positioning map; based on the preset perspective mapping relationship, performing coordinate transformation on the first image coordinates to obtain the second image coordinates of the target human foot key point corresponding to the target human head key point in the crowd image.
[0013] In a possible implementation, performing coordinate transformation on the first image coordinates based on the preset perspective mapping relationship to obtain the second image coordinates of the target human foot key point corresponding to the target human head key point in the crowd image includes: based on the preset perspective mapping relationship, determining the target image scale corresponding to the target human head key point; based on the target image scale, determining the image distance between the target human head key point and the target human foot key point; according to the first image coordinates and the image distance, determining the second image coordinates of the target human foot key point in the crowd image.
[0014] In a possible implementation, the method further includes: obtaining a plurality of labeled human body frames obtained by performing human body frame labeling on pedestrians at different positions in the crowd image; based on the plurality of labeled human body frames, determining the preset perspective mapping relationship.
[0015] In a possible implementation, determining the preset perspective mapping relationship based on the plurality of labeled human body frames includes: for any one of the labeled human body frames, determining the reference image scale corresponding to the reference human key point in the labeled human body frame; according to the third image coordinates of the reference human key point in each labeled human body frame and the reference image scale corresponding to the reference human key point in each labeled human body frame, fitting to obtain the preset perspective mapping relationship.
[0016] According to an aspect of the present disclosure, there is provided a crowd positioning device, including: a human key point positioning module, configured to perform human key point positioning on a crowd image to obtain an initial positioning map corresponding to the crowd image, where the initial positioning map is used to indicate the positions of the initial human key points included in the crowd image; a target neighborhood determination module, configured to determine a target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image; a filtering module, configured to filter the initial positioning map based on the target neighborhood corresponding to the initial human key point to obtain a target positioning map, where the target positioning map is used to indicate the positions of the target human key points included in the crowd image.
[0017] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.
[0018] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0019] In an embodiment of the present disclosure, for a crowd image, human key point localization can be performed to obtain an initial localization map that end-to-end indicates the positions of the initial human key points included in the crowd image. Based on the positions of the initial human key points in the crowd image, target neighborhoods corresponding to different initial human key points are determined. Furthermore, the initial localization map is filtered by using the target neighborhoods of different sizes that match different initial human key points, so as to reduce the false detection probability that multiple initial human key points correspond to the same human body, and obtain a target localization map with a higher accuracy rate.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Other features and aspects of the present disclosure will become clear according to the following detailed description of exemplary embodiments with reference to the accompanying drawings. Description of the Drawings
[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure.
[0022] Figure 1 A flowchart showing a crowd localization method according to an embodiment of the present disclosure;
[0023] Figure 2 A schematic diagram showing a crowd image and its corresponding preset perspective mapping relationship according to an embodiment of the present disclosure;
[0024] Figure 3 A block diagram showing a crowd localization device according to an embodiment of the present disclosure;
[0025] Figure 4 A block diagram showing an electronic device according to an embodiment of the present disclosure;
[0026] Figure 5 A block diagram showing another electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0027] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0028] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.
[0029] As used herein, the term "and / or" merely describes an association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0030] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0031] Fast crowd analysis is of great significance for public safety and urban planning. Common crowd analysis tasks include crowd counting, group behavior analysis, crowd positioning, etc. Among them, crowd positioning is the basis for other crowd analysis tasks. The accuracy of crowd positioning directly affects the accuracy of crowd counting and the results of crowd behavior analysis. Specifically, crowd positioning is a technology for estimating the positions of all human body key points in a video or image in a monitoring scenario through computer vision algorithms. Finally, the coordinates of all human body key points in the picture can be obtained, providing basic data for subsequent analysis tasks such as crowd counting and group behavior analysis.
[0032] Common population positioning methods rely on object detection algorithms. The population positioning task is transformed into a human head object detection task, and finally the center of the human head detection box is used as the positioning result of the human head key point. Population positioning algorithms based on human head object detection: On the one hand, the training of a human head object detection model based on a convolutional neural network relies on a large amount of labeled data. For the population positioning task, a large number of human head box annotations are required. For very dense population scenes, the cost of annotating human head boxes is relatively high. On the other hand, since the human heads in the distance are relatively small in the crowded scene, the annotated human head boxes are not accurate, which will affect the training of the human head object detection model. In addition, most of the high-performance human head object detection frameworks in the related technologies are two-stage detections, that is, first, the image is feature-extracted through a pre-trained feature extraction network, and then the extracted features are sent to a Region of Interest (ROI) network to obtain candidate regions where human heads may exist. Then, ROI pooling / ROI alignment is performed on the candidate regions to map the two-dimensional features into a fixed-length feature vector. Finally, the feature vector is sent to two neural networks for classification and position regression respectively to obtain the detection result. The above detection process is relatively complex and is obviously not practical for population analysis with high real-time requirements.
[0033] In addition to the population positioning method based on human head object detection, there is currently also a population positioning algorithm that directly generates a target positioning map. This method avoids the limitations of the algorithm based on human head object detection: The original population image is input, and through the convolutional neural network, a target positioning map of the same size as the original population image is directly output end-to-end. In the target positioning map, the target human head key point is represented by 1, otherwise it is 0. Summing the target positioning map can obtain the population counting index.
[0034] Then, the convolutional neural network can only output the initial positioning map, and image post-processing needs to be performed on the initial positioning map to obtain the final target positioning map. Filtering (for example, non-maximum suppression processing NMS) operations need to be performed during image post-processing. In the related technologies, a neighborhood with a fixed neighborhood radius is used for filtering. In this way, it is easy to cause misdetection of multiple human head key points of the same human body head in places where the human head is relatively large, or missed detection in places where the human head is relatively small.
[0035] Embodiments of the present disclosure provide a method for crowd positioning, which can be applied to crowd positioning in dense scenarios. By performing human key point positioning on the crowd images collected in dense scenarios, an initial positioning map for indicating the positions of the initial human key points included in the crowd images can be obtained end-to-end. Then, based on the positions of the initial human key points in the crowd images, different target neighborhoods corresponding to different initial human key points are determined. Furthermore, multiple initial human key points in the initial positioning map are filtered using target neighborhoods of different sizes to reduce the false detection probability that the same human corresponds to multiple initial human key points, and a target positioning map with higher accuracy is obtained.
[0036] The crowd positioning method provided by the embodiments of the present disclosure will be described in detail below.
[0037] Figure 1 The flowchart of a crowd positioning method according to an embodiment of the present disclosure is shown. This crowd positioning method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. This crowd positioning method can be implemented by a processor calling computer-readable instructions stored in a memory. Alternatively, the crowd positioning method can be executed by a server. As Figure 1 shown, this crowd positioning method may include:
[0038] In step S11, human key point positioning is performed on the crowd image to obtain an initial positioning map corresponding to the crowd image, where the initial positioning map is used to indicate the positions of the initial human key points included in the crowd image.
[0039] The crowd image here is an image containing a dense crowd, which can be obtained by an image acquisition device after image acquisition of a dense crowd within a certain spatial range, or a key image frame containing a dense crowd obtained from a video, or obtained by other means. The present disclosure does not make specific limitations on this.
[0040] By performing human key point positioning on the crowd image, an initial positioning map for indicating the positions of the initial human key points included in the crowd image is obtained end-to-end. For example, by performing head key point positioning on the crowd image, an initial positioning map for indicating the positions of the initial head key points included in the crowd image is obtained end-to-end. The specific process of head key point positioning will be described in detail later in combination with possible implementation manners of the present disclosure, and will not be elaborated here.
[0041] In step S12, based on the positions of the initial human key points in the crowd image, the target neighborhoods corresponding to the initial human key points are determined.
[0042] For the initial human key points included in the crowd image indicated by the initial positioning map, based on the positions of the initial human key points in the crowd image, the size of the target neighborhood where the initial human key points need to be processed subsequently can be determined. The process of determining the target neighborhood corresponding to the initial human key points based on the positions of the initial human key points in the crowd image will be described in detail later in combination with possible implementation manners of the present disclosure, and will not be elaborated here.
[0043] In step S13, the initial positioning map is filtered based on the target neighborhoods corresponding to the initial human key points to obtain a target positioning map, where the target positioning map is used to indicate the positions of the target human key points included in the crowd image.
[0044] The initial positioning map is filtered according to different target neighborhoods corresponding to different initial human key points to obtain a target positioning map with a relatively high accuracy. The process of filtering the initial positioning map based on the target neighborhoods corresponding to the initial human key points will be described in detail later in combination with possible implementation manners of the present disclosure, and will not be elaborated here.
[0045] In the embodiment of the present disclosure, for human key point positioning of the crowd image, an initial positioning map for indicating the positions of the initial human key points included in the crowd image can be obtained end-to-end. Based on the positions of the initial human key points in the crowd image, the target neighborhoods corresponding to different initial human key points are determined. Then, the initial positioning map is filtered by using different-sized target neighborhoods matched with different initial human key points to reduce the false detection probability that the same human corresponds to multiple initial human key points, and a target positioning map with a relatively high accuracy is obtained.
[0046] In a possible implementation manner, determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image includes: for any one initial human key point, determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales at different positions in the crowd image.
[0047] The image scale can be the number of pixel rows required to represent a unit height in the real world at a certain position in the crowd image. The unit height can be flexibly set according to actual situations. For example, the unit height can be 1 meter, and the present disclosure does not make specific limitations thereto.
[0048] In the crowd image, the image scale corresponding to the pedestrians in the foreground is large, while the image scale corresponding to the pedestrians in the background is small. For example, for two pedestrians A and B with the same height of 1.7 meters, the number of pixel rows required for pedestrian A in the foreground of the crowd image is p1, and the number of pixel rows required for pedestrian B in the background of the crowd image is p2, where p1 > p2.
[0049] Here, "far" means that the distance between the real pedestrian corresponding to the pedestrian in the crowd image and the image acquisition device for collecting the crowd image is far, and "near" means that the distance between the real pedestrian corresponding to the pedestrian in the crowd image and the image acquisition device for collecting the crowd image is near.
[0050] The preset perspective mapping relationship can indicate the image scales corresponding to different positions in the crowd image. The process of determining the preset perspective mapping relationship will be described in detail later in combination with possible implementation manners of the present disclosure, and will not be elaborated here.
[0051] After determining the preset perspective mapping relationship corresponding to the crowd image, for any initial human key point indicated by the initial positioning map, based on the position of the initial human key point in the crowd image and the preset perspective mapping relationship, the image scale corresponding to the initial human key point can be determined, and then based on the image scale corresponding to the initial human key point, the target neighborhood corresponding to the initial human key point can be determined.
[0052] Based on the position of the initial human key point in the crowd image, in addition to determining the image scale corresponding to the initial human key point based on the preset perspective mapping relationship to determine the target neighborhood corresponding to the initial human key point, the target neighborhood corresponding to the initial human head key point can also be directly determined based on the position of the initial human key point in the crowd image and the preset neighborhood radius corresponding to different positions in the preset crowd image. Other ways to determine the target area can also be adopted, and the embodiments of the present disclosure do not make specific limitations thereto.
[0053] In a possible implementation manner, the crowd positioning method further includes: obtaining a plurality of labeled human frames obtained by performing human frame annotation on pedestrians at different positions in the crowd image; and determining the preset perspective mapping relationship based on the plurality of labeled human frames.
[0054] By selecting pedestrians at different positions of far, middle, and near in the crowd image for human frame annotation, a plurality of labeled human frames in the crowd image can be obtained. Based on the proportional relationship between the height of the labeled human frame and the actual height of the pedestrian, the image scale corresponding to a limited position (the position of the labeled human frame) in the crowd image can be determined. Furthermore, based on the image scale corresponding to the limited position, further fitting can be performed to effectively obtain the image scale corresponding to each position in the crowd image, that is, the preset perspective mapping relationship is obtained.
[0055] Figure 2A schematic diagram showing a crowd image and its corresponding preset perspective mapping relationship according to an embodiment of the present disclosure. As Figure 2 shown, pedestrians at different positions of far, middle, and near are selected in the crowd image for human body box annotation, and four annotated human body boxes A, B, C, and D at different positions in the crowd image are obtained. Furthermore, based on the four annotated human body boxes A, B, C, and D, fitting is performed to effectively obtain the preset perspective mapping relationship corresponding to the crowd image.
[0056] In a possible implementation manner, determining the preset perspective mapping relationship based on multiple annotated human body boxes includes: for any one annotated human body box, determining the reference image scale corresponding to the reference human key points in the annotated human body box; and fitting to obtain the preset perspective mapping relationship according to the third image coordinates of the reference human key points in each annotated human body box and the reference image scale corresponding to the reference human key points in each annotated human body box.
[0057] Along the column direction of the crowd image, the image scales at different positions change linearly. Therefore, after determining the image scales corresponding to the reference human key points at finite positions in the crowd image according to the annotated human body boxes, linear function fitting can be used to effectively obtain the image scale corresponding to each position in the crowd image, that is, the preset perspective mapping relationship corresponding to the crowd image is obtained.
[0058] Since pedestrians stand vertically, taking the human foot key point as the reference human key point, the height of the annotated human body box can be regarded as the height of the pedestrian in the crowd image. The height of the annotated human body box can be represented by the number of pixel rows occupied by the annotated human body box. For example, if the annotated human body box occupies 17 pixel rows in the crowd image, the height of the annotated human body box is 17. Assuming that the real height of the pedestrian corresponding to the annotated human body box is 1.7 meters, the position of the reference human foot key point in the annotated human body box can be determined, indicating that 17 pixel rows are required for 1.7 meters in the real world. Assuming that the unit height is 1m, therefore, at the position of the reference human foot key point in the annotated human body box, 10 pixel rows are required for 1 meter in the real world, that is, the reference image scale corresponding to the reference human foot key point in the annotated human body box is 10. The real height of the pedestrian corresponding to the annotated human body box can be appropriately selected according to the actual situation, and the present disclosure does not make specific limitations on this.
[0059] The reference human foot key point in the annotated human body box can be the midpoint of the bottom edge of the annotated human body box, or other pixel points in the annotated human body box. The present disclosure does not make specific limitations on this.
[0060] Still taking the above Figure 2For example, after obtaining four labeled human body bounding boxes A, B, C, and D in the crowd image, the reference image scale corresponding to the reference human foot key points in each labeled human body bounding box is determined in the above manner. Furthermore, according to the third image coordinates of the reference human foot key points in the four labeled human body bounding boxes and their corresponding reference image scales, a linear function fitting is performed to obtain a linear mapping function p = a * y + b.
[0061] Among them, the image coordinates refer to the position coordinates in the pixel coordinate system of the crowd image. For example, taking the upper left corner of the crowd image as the coordinate origin (0, 0), the direction parallel to the row direction of the image is the x-axis direction, and the direction parallel to the column direction of the image is the y-axis direction to construct the pixel coordinate system of the crowd image. The units of the abscissa and ordinate of the image coordinates are both pixel points. For example, if the image coordinates of the reference human foot key point are (10, 15), it indicates that the reference human foot key point is the pixel point located at the 10th row and 15th column in the crowd image.
[0062] The linear mapping function p = a * y + b is the functional representation form of the preset perspective mapping relationship corresponding to the crowd image. Among them, a and b are parameters obtained by linear function fitting, y is the ordinate of the image coordinates at different positions in the crowd image, and p is the image scale corresponding to this position. Using the linear mapping function p = a * y + b, the image scale corresponding to each position in the crowd image can be determined.
[0063] Before or after determining the preset perspective mapping relationship corresponding to the crowd image, human key point localization is performed on the crowd image to obtain the initial localization map corresponding to the crowd image.
[0064] In a possible implementation manner, the initial human key point is the initial human head key point.
[0065] Since in a dense crowd image, the bodies of different pedestrians are severely occluded, therefore, determining the human head key point as the human key point can effectively distinguish different pedestrians and improve the crowd localization accuracy.
[0066] The human head key point can be the center point of the human head, or other preset key points of the human head. The present disclosure does not make specific limitations on this.
[0067] In a possible implementation manner, performing human key point localization on the crowd image to obtain the initial localization map corresponding to the crowd image includes: performing human head key point localization on the crowd image to determine the predicted localization map corresponding to the crowd image, where the predicted localization map is used to indicate the predicted confidence of each pixel point in the crowd image being a human head key point; based on a preset confidence threshold, performing image processing on the predicted localization map to obtain the initial localization map.
[0068] Perform head key point localization on the crowd image, and end-to-end determine the prediction confidence of each pixel point in the crowd image as a head key point, and then perform threshold segmentation on the prediction localization map through a preset confidence threshold to determine the positions of the initial head key points included in the crowd image.
[0069] In one example, a trained head key point localization neural network can be used to perform head key point localization on the crowd image. Specifically, input the crowd image into the trained head key point localization neural network, and after the localization of the head key point localization neural network, directly output the prediction localization map. The specific network structure of the trained head key point localization neural network and the training process can adopt the network structure and training process in the related technology, and the present disclosure does not make specific limitations thereon.
[0070] In one example, the pixel value of each pixel point in the prediction localization map represents the prediction confidence of that pixel point, that is, the probability that the pixel point is a head key point. Perform a sigmoid operation on the prediction localization map so that the pixel value of each pixel point in the prediction localization map is between 0 and 1. For example, if the pixel value of a certain pixel point in the prediction localization map is 0.7, it means that the probability that the pixel point is a head key point is 0.7.
[0071] Since the prediction localization map is only used to indicate the prediction confidence of each pixel point in the crowd image as a head key point, therefore, by presetting the confidence threshold and performing threshold segmentation on the prediction localization map, an initial localization map for indicating the positions of the initial head key points included in the crowd image can be effectively obtained. The specific value of the preset confidence threshold can be flexibly set according to the actual situation, and the present disclosure does not make specific limitations thereon.
[0072] Compare the pixel value of each pixel point in the prediction localization map with the preset confidence threshold. When the pixel value of a certain pixel point in the prediction localization map is greater than or equal to the preset confidence threshold, determine the pixel value of the pixel point with the same relative position in the initial localization map as 1; when the pixel value of a certain pixel point in the prediction localization map is less than the preset confidence threshold, determine the pixel value of the pixel point with the same relative position in the initial localization map as 0.
[0073] The initial localization map and the crowd image have the same size. The positions of the pixel points with a pixel value of 1 in the initial localization map are used to indicate the positions of the initial head key points included in the crowd image. For example, when the pixel value of the pixel point with the image coordinates (x, y) in the initial localization map is 1, it can be determined that the pixel point with the image coordinates (x, y) in the crowd image is an initial head key point; when the pixel value of the pixel point with the image coordinates (x, y) in the initial localization map is 0, it can be determined that the pixel point with the image coordinates (x, y) in the crowd image is a part other than the initial head key point.
[0074] After determining the positions of the initial head key points included in the crowd image, the target neighborhood corresponding to each initial head key point can be determined based on a preset perspective mapping relationship.
[0075] In a possible implementation manner, determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image and the preset perspective mapping relationship includes: determining the target image scale corresponding to the initial head key point in the crowd image based on the preset perspective mapping relationship; determining the head box height corresponding to the initial head key point based on the target image scale; and determining the target neighborhood corresponding to the initial head key point based on the head box height corresponding to the initial head key point.
[0076] Based on the preset perspective mapping relationship, the target image scale corresponding to the initial head key point can be quickly determined, and then the head box heights corresponding to the initial head key points at different positions can be determined, so that the target neighborhood matching the head box height can be further determined.
[0077] For example, for an initial head key point i in the initial positioning map, the image coordinates of the initial head key point i in the crowd image are (h x , h y ), then according to the preset perspective mapping relationship (linear mapping function p = a*y + b) corresponding to the crowd image, the target scale corresponding to the initial head key point i can be determined as p i = a*h y + b. Assuming that the true head box height of the pedestrians in the crowd image is 0.4 meters * 0.4 meters, the head box height corresponding to the initial head key point i in the crowd image is s i = 0.4*p i . According to the head box height s i = 0.4*p i corresponding to the initial head key point, the target neighborhood with a matching size is determined.
[0078] In a possible implementation manner, determining the target neighborhood corresponding to the initial head key point based on the head box height corresponding to the initial head key point includes: in the case where the head box height is greater than the preset head box height threshold, determining the target neighborhood based on the first neighborhood radius; or, in the case where the head box height is less than or equal to the preset head box height threshold, determining the target neighborhood based on the second neighborhood radius, where the first neighborhood radius is greater than the second neighborhood radius.
[0079] When the height of the head frame is greater than the preset head frame height threshold, it can be determined that the size of the head frame is relatively large. Therefore, a relatively large first neighborhood radius is used for subsequent filtering processing. When the height of the head frame is less than or equal to the preset head frame height threshold, it can be determined that the size of the head frame is relatively small. Therefore, a relatively small second neighborhood radius is used for subsequent filtering processing. By flexibly determining the neighborhood radius, the accuracy of the filtering operation can be improved. The specific values of the preset head frame height threshold, the first neighborhood radius, and the second neighborhood radius can be flexibly set according to the actual situation, and the present disclosure does not make specific limitations on this.
[0080] In one example, the head frame height threshold is 32. For a certain initial head key point i, when the height s of its corresponding head frame i > 32, its target neighborhood is determined based on the first neighborhood radius 2; when the height s of its corresponding head frame i <= 32, its target neighborhood is determined based on the second neighborhood radius 1.
[0081] When the first preset neighborhood radius is 2, the target neighborhood corresponding to the initial head key point i includes the pixel points whose pixel distance from the initial head key point i does not exceed 2 pixel points. When the second preset neighborhood radius is 1, the target neighborhood corresponding to the initial head key point i includes the pixel points whose pixel distance from the initial head key point i does not exceed 1 pixel point.
[0082] After determining the target neighborhood corresponding to each initial head key point in the initial positioning map based on the above method, the initial positioning map is filtered using the target neighborhood corresponding to each initial head key point to obtain a target positioning map with higher accuracy.
[0083] In a possible implementation manner, filtering the initial positioning map based on the target neighborhood corresponding to the initial human key point to obtain a target positioning map includes: for any initial head key point i, determining whether there is at least one other initial head key point in the target neighborhood corresponding to the initial head key point i; when there is at least one other initial head point j in the target neighborhood, determining the predicted confidence corresponding to the initial head point i and the predicted confidence corresponding to at least one other initial head point j according to the predicted positioning map; and determining the initial head key point with the maximum predicted confidence among the initial head key point i and at least one other initial head key point j as the target head key point in the target neighborhood.
[0084] To avoid the problem of false detection of multiple initial head key points corresponding to the same human head, the initial positioning map is further filtered using the target neighborhood of the initial head key point to obtain a target positioning map with higher accuracy.
[0085] For any initial human head key point i in the initial positioning map, after determining its corresponding target neighborhood according to the above method, check whether there are other initial human head key points in its target neighborhood. If there is another initial human head key point j with image coordinates (x j , y j ), then according to the predicted positioning map, determine the predicted confidence corresponding to the initial human head key point i and the predicted confidence corresponding to the initial human head key point j. When the predicted confidence of the initial human head key point i is greater than the predicted confidence of the initial human head key point j, keep the pixel value of the pixel point with image coordinates (x i , y i ) as 1, and update the pixel value of the pixel point with image coordinates (x j , y j ) to 0, that is, filter out the initial human head key point j in the initial positioning map. And so on, traverse each initial human head key point in the initial positioning map to obtain the final target positioning map.
[0086] In a possible implementation, the crowd positioning method further includes: determining the positions of the target human foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship.
[0087] When analyzing the positions of pedestrians in a monitoring system, generally, the positions of human feet need to be known. For example, in the case of traffic flow statistics, lines are generally drawn in advance according to the actual set positions, and then whether a line is crossed is judged based on the position relationship of the human feet in the front and rear frames. Therefore, in the crowd positioning task, how to accurately locate the positions of human feet is crucial.
[0088] After determining the target positioning map based on the foregoing method, the positions of the target human head key points included in the crowd image can be determined. Furthermore, based on the target positioning map and the preset perspective mapping relationship, the positions of the target human foot key points included in the crowd image can be further quickly determined.
[0089] In a possible implementation, determining the positions of the target human foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship includes: for any target human head key point, determining the first image coordinates of the target human head key point in the crowd image according to the target positioning map; based on the preset perspective mapping relationship, performing coordinate transformation on the first image coordinates to obtain the second image coordinates of the corresponding target human foot key point of the target human head key point in the crowd image.
[0090] Since the preset perspective mapping relationship can indicate the image scales at different positions in the crowd image, based on the preset perspective mapping relationship, the first image coordinates of the target human head key point can be subjected to coordinate transformation to obtain the second image coordinates of the corresponding target human foot key point of the target human head key point in the crowd image.
[0091] For example, for a target head key point in the target positioning map, the first image coordinate of the target head key point in the crowd image is (h x , h y ), and the first image coordinate is known. According to the preset perspective mapping relationship corresponding to the crowd image, the first image coordinate (h x , h y ) can be coordinate-transformed to obtain the second image coordinate of the target foot key point corresponding to the target head key point, which is (f x , f y ).
[0092] In a possible implementation manner, based on the preset perspective mapping relationship, coordinate-transforming the first image coordinate to obtain the second image coordinate of the target foot key point corresponding to the target head key point in the crowd image includes: determining the target image scale corresponding to the target head key point based on the preset perspective mapping relationship; determining the image distance between the target head key point and the target foot key point based on the target image scale; and determining the second image coordinate of the target foot key point in the crowd image according to the first image coordinate and the image distance.
[0093] Since the pedestrians are standing vertically, it can be determined that f x = h x . According to the preset perspective mapping relationship (linear mapping function p = a*y + b), it can be determined that the target scale corresponding to the target foot key point is a*f x + b. Assuming that in the real world, the distance from the center of the head to the feet is 1.5 meters, that is, the real distance between the target head key point and the target foot key point is 1.5 meters, then the image distance between the target foot key point and the target head key point in the crowd image is 1.5*(a*f x + b), that is, h x = f y - 1.5*(a*f x + b), and further it can be determined that f y = (h y + 1.5*b) / (1 - 1.5*a). Therefore, the second image coordinate of the target foot key point in the crowd image is (h x , (h y + 1.5*b) / (1 - 1.5*a)).
[0094] By traversing each target human head key point in the target positioning map, the position of the target human foot key point corresponding to each target human head key point in the crowd image can be determined. Furthermore, based on the positions of the target human foot key points in adjacent front and rear frame crowd images, crowd behavior analysis such as passenger flow statistics and behavior trajectory analysis can be performed. The present disclosure does not make specific limitations on this.
[0095] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0096] In addition, the present disclosure also provides a crowd positioning device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any crowd positioning method provided by the present disclosure. The corresponding technical solutions and descriptions can be referred to the corresponding records in the method part and will not be elaborated further.
[0097] Figure 3 The block diagram of a crowd positioning device according to an embodiment of the present disclosure is shown. As Figure 3 shown, the device 30 includes:
[0098] A human key point positioning module 31, configured to perform human key point positioning on a crowd image to obtain an initial positioning map corresponding to the crowd image, where the initial positioning map is used to indicate the positions of the initial human key points included in the crowd image;
[0099] A target neighborhood determination module 32, configured to determine a target neighborhood corresponding to an initial human key point based on the position of the initial human key point in the crowd image;
[0100] A filtering module 33, configured to filter the initial positioning map based on the target neighborhood corresponding to the initial human key point to obtain a target positioning map, where the target positioning map is used to indicate the positions of the target human key points included in the crowd image.
[0101] In a possible implementation manner, the target neighborhood determination module 32 is specifically configured to:
[0102] For any initial human key point, determine a target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales at different positions in the crowd image.
[0103] In a possible implementation manner, the initial human key point is an initial human head key point.
[0104] In a possible implementation, the target neighborhood determination module 32 includes:
[0105] The first determination sub-module is configured to determine the target image scale corresponding to the position of the initial human head key point in the crowd image based on a preset perspective mapping relationship;
[0106] The second determination sub-module is configured to determine the height of the human head frame corresponding to the initial human head key point based on the target image scale;
[0107] The third determination sub-module is configured to determine the target neighborhood corresponding to the initial human head key point based on the height of the human head frame corresponding to the initial human head key point.
[0108] In a possible implementation, the third determination sub-module is specifically configured to:
[0109] When the height of the human head frame is greater than a preset human head frame height threshold, determine the target neighborhood based on the first neighborhood radius; or,
[0110] When the height of the human head frame is less than or equal to the preset human head frame height threshold, determine the target neighborhood based on the second neighborhood radius, where the first neighborhood radius is greater than the second neighborhood radius.
[0111] In a possible implementation, the human key point localization module 31 includes:
[0112] The human key point localization sub-module is configured to perform human key point localization on the crowd image to determine a predicted localization map corresponding to the crowd image, where the predicted localization map is used to indicate the predicted confidence of each pixel point in the crowd image being a human key point;
[0113] The fourth determination sub-module is configured to perform image processing on the predicted localization map based on a preset confidence threshold to obtain an initial localization map.
[0114] In a possible implementation, the filtering module 33 is specifically configured to:
[0115] For any initial human head key point i, determine whether there is at least one other initial human head key point in the target neighborhood corresponding to the initial human head key point i;
[0116] When there is at least one other initial human head point j in the target neighborhood, determine the predicted confidence corresponding to the initial human head point i and the predicted confidence corresponding to at least one other initial human head point j based on the predicted localization map;
[0117] Based on the initial human head key point i and the initial human head key point with the maximum predicted confidence among at least one other initial human head key points j, determine the target human head key point in the target neighborhood.
[0118] In a possible implementation, the device 30 further includes:
[0119] A human foot key point determination module, configured to determine the positions of target human foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship.
[0120] In a possible implementation, the human foot key point determination module includes:
[0121] A fifth determination sub-module, configured to, for any target human head key point, determine the first image coordinates of the target human head key point in the crowd image according to the target positioning map;
[0122] A sixth determination sub-module, configured to perform coordinate conversion on the first image coordinates based on the preset perspective mapping relationship to obtain the second image coordinates of the target human foot key point corresponding to the target human head key point in the crowd image.
[0123] In a possible implementation, the sixth determination sub-module is specifically configured to:
[0124] Determine the target image scale corresponding to the target human head key point based on the preset perspective mapping relationship;
[0125] Determine the image distance between the target human head key point and the target human foot key point based on the target image scale;
[0126] Determine the second image coordinates of the target human foot key point in the crowd image according to the first image coordinates and the image distance.
[0127] In a possible implementation, the device 30 further includes:
[0128] An acquisition module, configured to acquire a plurality of labeled human body frames obtained by performing human body frame labeling on pedestrians at different positions in the crowd image;
[0129] A perspective mapping relationship determination module, configured to determine the preset perspective mapping relationship based on the plurality of labeled human body frames.
[0130] In a possible implementation, the perspective mapping relationship determination module is specifically configured to:
[0131] For any one of the labeled human body frames, determine the reference image scale corresponding to the reference human key point in the labeled human body frame;
[0132] Fit to obtain the preset perspective mapping relationship according to the third image coordinates of the reference human key points in each labeled human body frame and the reference image scales corresponding to the reference human key points in each labeled human body frame.
[0133] This method has a specific technical association with the internal structure of a computer system and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing the data storage volume, reducing the data transmission volume, increasing the hardware processing speed, etc.), thereby obtaining the technical effect of improving the internal performance of the computer system in line with natural laws.
[0134] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0135] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0136] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.
[0137] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.
[0138] The electronic device can be provided as a terminal, a server or other forms of devices.
[0139] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Referring to Figure 4 , the electronic device 800 can be a terminal device such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.
[0140] Referring to Figure 4 , the electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0141] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0142] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.
[0143] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0144] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0145] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0146] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0147] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a complementary metal oxide semiconductor (CMOS) or a charge coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0148] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as Wi-Fi, 2G, 3G, 4G, Long Term Evolution (LTE) of Universal Mobile Telecommunications Technology, 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0149] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0150] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 804 including computer program instructions, and the computer program instructions can be executed by a processor 820 of the electronic device 800 to complete the above method.
[0151] The present disclosure relates to the field of augmented reality. By acquiring image information of a target object in a real environment, relevant features, states, and attributes of the target object are detected or recognized by means of various vision-related algorithms, so as to obtain an AR effect combining virtual and real that matches a specific application. Exemplarily, the target object may involve a face, limbs, gestures, actions, etc. related to the human body, or identification marks, markers related to objects, or sand tables, display areas, or display items related to a venue or place. The vision-related algorithms may involve visual positioning, SLAM, three-dimensional reconstruction, image registration, background segmentation, key point extraction and tracking of an object, pose or depth detection of an object, etc. The specific application may not only involve interaction scenarios such as guiding, navigation, explanation, reconstruction, virtual effect overlay display, etc. related to a real scene or item, but also involve special effect processing related to people, such as makeup beautification, limb beautification, special effect display, virtual model display, etc. The relevant features, states, and attributes of the target object can be detected or recognized by means of a convolutional neural network. The convolutional neural network is a network model obtained by training a model based on a deep learning framework.
[0152] Figure 5 A block diagram of another electronic device according to an embodiment of the present disclosure is shown. Referring to Figure 5 , the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 5 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0153] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ), or the like.
[0154] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0155] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0156] The computer-readable storage medium may be a tangible device that can retain and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0157] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0158] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0159] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0160] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data processing apparatus, an apparatus is created that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions that implement various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0161] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0162] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0163] The computer program product can be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0164] The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments. The similarities or likenesses between them can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0165] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0166] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0167] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for crowd positioning, characterized in that, Including: Performing human key point localization on a crowd image to obtain an initial localization map corresponding to the crowd image, where the initial localization map is used to indicate the positions of initial human key points included in the crowd image; Determining a target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image; Filtering the initial localization map based on the target neighborhood corresponding to the initial human key point to obtain a target localization map, where the target localization map is used to indicate the positions of target human key points included in the crowd image; Wherein, the determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image includes: For any one of the initial human key points, determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales at different positions in the crowd image.
2. The method according to claim 1, wherein The initial human key point is an initial human head key point.
3. The method according to claim 2, characterized in that, The determining the target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image and a preset perspective mapping relationship includes: Determining a target image scale corresponding to the position of the initial human head key point in the crowd image based on the preset perspective mapping relationship; Determining the height of a human head frame corresponding to the initial human head key point based on the target image scale; Determining the target neighborhood corresponding to the initial human head key point based on the height of the human head frame corresponding to the initial human head key point.
4. The method according to claim 3, wherein The determining the target neighborhood corresponding to the initial human head key point based on the height of the human head frame corresponding to the initial human head key point includes: In the case where the height of the human head frame is greater than a preset human head frame height threshold, determining the target neighborhood based on a first neighborhood radius; or, In the case where the height of the human head frame is less than or equal to the preset human head frame height threshold, determining the target neighborhood based on a second neighborhood radius, where the first neighborhood radius is greater than the second neighborhood radius.
5. The method according to any one of claims 2 to 4, characterized in that, The performing human key point localization on a crowd image to obtain an initial localization map corresponding to the crowd image includes: Performing human key point localization on the crowd image to determine a predicted localization map corresponding to the crowd image, where the predicted localization map is used to indicate the prediction confidence of each pixel point in the crowd image being a human key point; Performing image processing on the predicted localization map based on a preset confidence threshold to obtain the initial localization map.
6. The method according to claim 5, characterized in that, The filtering the initial localization map based on the target neighborhood corresponding to the initial human key point to obtain a target localization map includes: For any initial human head key point i, determining whether there is at least one other initial human head key point in the target neighborhood corresponding to the initial human head key point i; In the case that there is at least one other initial human head point j in the target neighborhood, based on the predicted positioning map, determine the predicted confidence corresponding to the initial human head point i and the predicted confidence corresponding to the at least one other initial human head point j; Based on the initial human key point i and the initial human key point with the maximum predicted confidence among the at least one other initial human key points j, determine the target human key point in the target neighborhood.
7. The method according to claim 6, characterized in that The method further includes: Based on the target positioning map and the preset perspective mapping relationship, determine the position of the target human foot key point included in the crowd image.
8. The method according to claim 7, wherein The determining the position of the target human foot key point included in the crowd image based on the target positioning map and the preset perspective mapping relationship includes: For any one of the target human head key points, according to the target positioning map, determine the first image coordinate of the target human head key point in the crowd image; Based on the preset perspective mapping relationship, perform coordinate transformation on the first image coordinate to obtain the second image coordinate of the target human foot key point corresponding to the target human head key point in the crowd image.
9. The method according to claim 8, wherein The performing coordinate transformation on the first image coordinate based on the preset perspective mapping relationship to obtain the second image coordinate of the target human foot key point corresponding to the target human head key point in the crowd image includes: Based on the preset perspective mapping relationship, determine the target image scale corresponding to the target human head key point; Based on the target image scale, determine the image distance between the target human head key point and the target human foot key point; According to the first image coordinate and the image distance, determine the second image coordinate of the target human foot key point in the crowd image.
10. The method according to any one of claims 2 to 4, characterized in that The method further includes: Obtain a plurality of labeled human body frames obtained by performing human body frame labeling on pedestrians at different positions in the crowd image; Based on the plurality of labeled human body frames, determine the preset perspective mapping relationship.
11. The method according to claim 10, wherein The determining the preset perspective mapping relationship based on the plurality of labeled human body frames includes: For any one of the labeled human body frames, determine the reference image scale corresponding to the reference human key point in the labeled human body frame; According to the third image coordinate of the reference human key point in each labeled human body frame and the reference image scale corresponding to the reference human key point in each labeled human body frame, fit to obtain the preset perspective mapping relationship.
12. A crowd positioning device, characterized in that, Includes: A human key point positioning module, configured to perform human key point positioning on a crowd image to obtain an initial positioning map corresponding to the crowd image, where the initial positioning map is used to indicate the positions of the initial human key points included in the crowd image; A target neighborhood determination module, configured to determine a target neighborhood corresponding to the initial human key point based on the position of the initial human key point in the crowd image; A filtering module, configured to filter the initial positioning map based on the target neighborhood corresponding to the initial human key point to obtain a target positioning map, where the target positioning map is used to indicate the positions of the target human key points included in the crowd image; Among them, the target neighborhood determination module is specifically configured to: For any one of the initial human key points, based on the position of the initial human key point in the crowd image and a preset perspective mapping relationship, determine the target neighborhood corresponding to the initial human key point, where the preset perspective mapping relationship is used to indicate the image scales at different positions in the crowd image.
13. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 11.
14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Target object positioning method and device, electronic equipment and readable storage medium
CN112802108A