A method and device for crowd positioning, an electronic device, and a storage medium

By transforming the relationship between the key points of the head and perspective mapping of the crowd images, the problem of inaccurate positioning of the key points of the human foot in dense scenes is solved, and high-precision population positioning and behavior analysis is achieved.

CN114550085BActive Publication Date: 2025-07-18SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210146591.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-07-18
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

In dense scenarios, it is difficult for the prior art to accurately locate key points of people's feet in crowd images, resulting in poor pedestrian detection results and prone to missed detection, affecting the accuracy of crowd positioning.

Method used

By positioning the key points of the head of the crowd image, using the preset perspective mapping relationship, the position of the key points of the foot in the crowd image is determined, and the coordinate conversion is used to combine the image scale and actual height to accurately locate the key points of the foot.

Benefits of technology

It realizes the high-precision positioning of key points in crowd images in dense scenes, improves the accuracy of crowd positioning, and supports subsequent passenger flow statistics and behavioral analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550085B_ABST
    Figure CN114550085B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a crowd positioning method, an apparatus, an electronic device, and a storage medium. The method includes: performing head key point positioning on a crowd image to obtain a target positioning map corresponding to the crowd image, where the target positioning map is used to indicate the positions of target head key points included in the crowd image; and determining the positions of target foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales corresponding to different positions in the crowd image. Embodiments of the present disclosure can effectively obtain the positions of target foot key points with relatively high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method and device for crowd positioning, an electronic device, and a storage medium. Background Art

[0002] With the growth of the population and the acceleration of the urbanization process, the behavior of a large number of people gathering is increasing and the scale is getting larger. Crowd analysis is of great significance for public safety and urban planning. Common crowd analysis tasks include crowd counting, group behavior analysis, crowd positioning, etc. Among them, crowd positioning is the basis for other crowd analysis tasks. When analyzing the position of pedestrians in a monitoring system, generally, the position of the human feet needs to be known. For example, in the case of traffic flow statistics, lines are usually drawn in advance, and then whether a line is crossed is judged according to the position relationship of the human feet in the front and rear frames. Therefore, in the crowd positioning task, it is crucial to accurately locate the position of the human feet. Summary of the Invention

[0003] The present disclosure proposes a technical solution for a method and device for crowd positioning, an electronic device, and a storage medium.

[0004] According to an aspect of the present disclosure, there is provided a method for crowd positioning, including: performing head key point positioning on a crowd image to obtain a target positioning map corresponding to the crowd image, where the target positioning map is used to indicate the positions of target head key points included in the crowd image; and determining the positions of target human foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales corresponding to different positions in the crowd image.

[0005] In a possible implementation manner, the determining the positions of target human foot key points included in the crowd image based on the target positioning map and the preset perspective mapping relationship includes: for any one of the target head key points, determining a first image coordinate of the target head key point in the crowd image according to the target positioning map; and performing coordinate transformation on the first image coordinate based on the preset perspective mapping relationship to obtain a second image coordinate of the target human foot key point corresponding to the target head key point in the crowd image.

[0006] In a possible implementation manner, the coordinate transformation of the first image coordinate based on the preset perspective mapping relationship to obtain the second image coordinate of the target human foot key point corresponding to the target human head key point in the crowd image includes: determining the target image scale corresponding to the target human foot key point based on the preset perspective mapping relationship; determining the image distance between the target human head key point and the target human foot key point based on the target image scale; and determining the second image coordinate of the target human foot key point in the crowd image according to the first image coordinate and the image distance.

[0007] In a possible implementation manner, the head key point positioning of the crowd image to obtain the target positioning map corresponding to the crowd image includes: performing head key point positioning on the crowd image to determine the predicted positioning map corresponding to the crowd image, where the predicted positioning map is used to indicate the prediction confidence of each pixel point in the crowd image being a head key point; performing image processing on the predicted positioning map based on a preset confidence threshold to obtain an initial positioning map, where the initial positioning map is used to indicate the positions of the initial head key points included in the crowd image; determining the target neighborhood corresponding to each initial head key point in the initial positioning map; and performing filtering processing on the target neighborhood corresponding to each initial head key point based on the predicted positioning map to obtain the target positioning map.

[0008] In a possible implementation manner, the determining the target neighborhood corresponding to each initial head key point in the initial positioning map includes: determining the target neighborhood corresponding to each initial head key point according to a first preset neighborhood radius.

[0009] In a possible implementation manner, the determining the target neighborhood corresponding to each initial head key point in the initial positioning map includes: for any one of the initial head key points, determining the target neighborhood corresponding to the initial head key point based on the position of the initial head key point in the crowd image and the preset perspective mapping relationship.

[0010] In a possible implementation manner, the determining the target neighborhood corresponding to the initial head key point based on the position of the initial head key point in the crowd image and the preset perspective mapping relationship includes: determining the target image scale corresponding to the position of the initial head key point in the crowd image based on the preset perspective mapping relationship; determining the head frame height corresponding to the initial head key point based on the target image scale; and determining the target neighborhood corresponding to the initial head key point based on the head frame height corresponding to the initial head key point.

[0011] In a possible implementation manner, determining the target neighborhood corresponding to the initial human head key point based on the height of the human head bounding box corresponding to the initial human head key point includes: when the height of the human head bounding box is greater than a preset human head bounding box height threshold, determining the target neighborhood corresponding to the initial human head key point based on a second preset neighborhood radius; or, when the height of the human head bounding box is less than or equal to the preset human head bounding box height threshold, determining the target neighborhood corresponding to the initial human head key point based on a third neighborhood radius, where the second neighborhood radius is greater than the third neighborhood radius.

[0012] In a possible implementation manner, filtering the target neighborhood corresponding to each initial human head key point based on the predicted positioning map to obtain the target positioning map includes: for any initial human head key point i, determining whether there is at least one other initial human head center point in the target neighborhood corresponding to the initial human head key point i; when there is at least one other initial human head point j in the target neighborhood corresponding to the initial human head key point i, determining the predicted confidence corresponding to the initial human head point i and the predicted confidence corresponding to the at least one other initial human head point j based on the predicted positioning map; determining the target human head key point in the target neighborhood corresponding to the initial human head key point i based on the initial human head key point i and the initial human head key point with the highest predicted confidence among the at least one other initial human head key points j.

[0013] In a possible implementation manner, the method further includes: obtaining a plurality of labeled human body bounding boxes obtained by performing human body bounding box annotation on pedestrians at different positions in the crowd image; determining the preset perspective mapping relationship based on the plurality of labeled human body bounding boxes.

[0014] In a possible implementation manner, determining the preset perspective mapping relationship based on the plurality of labeled human body bounding boxes includes: for any one of the labeled human body bounding boxes, determining the reference image scale corresponding to the reference human foot key point in the labeled human body bounding box; fitting the preset perspective mapping relationship according to the third image coordinate of the reference human foot key point in each labeled human body bounding box and the reference image scale corresponding to the reference human foot key point in each labeled human body bounding box.

[0015] According to one aspect of the present disclosure, there is provided a crowd positioning device, including: a head key point positioning module, configured to perform head key point positioning on a crowd image to obtain a target positioning map corresponding to the crowd image, where the target positioning map is used to indicate the positions of target head key points included in the crowd image; a human foot key point determination module, configured to determine the positions of target human foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales corresponding to different positions in the crowd image.

[0016] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.

[0017] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0018] In an embodiment of the present disclosure, when performing head key point positioning on a crowd image, a target positioning map for indicating the positions of target head key points included in the crowd image can be obtained end-to-end. Since the preset perspective mapping relationship can be used to indicate the image scales corresponding to different positions in the crowd image, therefore, based on the target head key points and the preset perspective mapping relationship, the positions of target human foot key points with relatively high accuracy can be effectively obtained.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Other features and aspects of the present disclosure will become clear according to the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings herein are incorporated into the specification and form a part of the specification, which illustrate embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure.

[0021] Figure 1 A flowchart showing a crowd positioning method according to an embodiment of the present disclosure;

[0022] Figure 2 A schematic diagram showing a crowd image and its corresponding preset perspective mapping relationship according to an embodiment of the present disclosure;

[0023] Figure 3 A block diagram showing a crowd positioning device according to an embodiment of the present disclosure;

[0024] Figure 4A block diagram of an electronic device according to an embodiment of the present disclosure is shown;

[0025] Figure 5 A block diagram of another electronic device according to an embodiment of the present disclosure is shown. Detailed implementation manners

[0026] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0027] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior or better than other embodiments.

[0028] The term "and / or" herein merely describes an association relationship of associated objects and means that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent any one or more elements selected from the set composed of A, B, and C.

[0029] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0030] When analyzing the position of a pedestrian in a monitoring system, generally, the position of the human foot needs to be known. For example, in the case of traffic flow statistics, generally, a line is drawn in advance according to the actually set position, and then it is determined whether the line is crossed according to the position relationship of the human foot in the front and back frames. Therefore, in the crowd positioning task, how to accurately locate the position of the human foot is crucial. In the related art, a human body box is used for pedestrian detection and tracking, and the midpoint of the bottom edge of the detected human body box is determined as the position of the human foot key point. However, in a dense scene, the pedestrians are severely occluded, resulting in a poor effect of using the human body box for pedestrian detection, and there are likely to be many missed detections, thus affecting the functional accuracy.

[0031] Embodiments of the present disclosure provide a crowd positioning method, which can be applied to crowd positioning in dense scenarios. By performing head key point positioning on the crowd image collected in the dense scenario, a target positioning map for indicating the positions of the target head key points included in the crowd image can be obtained end-to-end. Since the preset perspective mapping relationship can be used to indicate the image scales corresponding to different positions in the crowd image, based on the target head key points and the preset perspective mapping relationship, the positions of the target foot key points with relatively high accuracy can be effectively obtained.

[0032] The crowd positioning method provided by the embodiments of the present disclosure will be described in detail below.

[0033] Figure 1 The flowchart of a crowd positioning method according to an embodiment of the present disclosure is shown. This crowd positioning method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. This crowd positioning method can be implemented by a processor calling computer-readable instructions stored in a memory. Alternatively, the crowd positioning method can be executed by a server. As Figure 1 shown, this crowd positioning method may include:

[0034] In step S11, perform head key point positioning on the crowd image to obtain a target positioning map corresponding to the crowd image, where the target positioning map is used to indicate the positions of the target head key points included in the crowd image.

[0035] The crowd image here is an image containing a dense crowd. It can be obtained by an image acquisition device after image acquisition of a dense crowd within a certain spatial range, or a key image frame containing a dense crowd obtained from a video, or obtained by other means. The present disclosure does not make specific limitations on this.

[0036] Perform head key point positioning on the crowd image to obtain end-to-end a target positioning map for indicating the positions of the target head key points included in the crowd image. The specific process of head key point positioning will be described in detail later in combination with possible implementation manners of the present disclosure, and will not be elaborated here.

[0037] The head key point can be the center point of the human head, or other preset key points of the human head. The present disclosure does not make specific limitations on this.

[0038] In step S12, based on the target positioning map and the preset perspective mapping relationship, determine the positions of the target human foot key points included in the crowd image, where the preset perspective mapping relationship is used to indicate the image scales corresponding to different positions in the crowd image.

[0039] The image scale here is used to represent the proportional relationship between the actual size and the image size in the crowd image. For example, the image scale can be the number of pixel rows required at a certain position in the crowd image to represent a unit height in the real world. The unit height can be flexibly set according to the actual situation. For example, the unit height can be 1 meter, and the present disclosure does not make specific limitations on this.

[0040] For the pedestrians closer in the crowd image, the corresponding image scale is larger, and for the pedestrians farther away, the corresponding image scale is smaller. For example, for two pedestrians A and B with the same height of 1.7 meters, the number of pixel rows required for pedestrian A closer in the crowd image is p1, and the number of pixel rows required for pedestrian B farther away in the crowd image is p2, where p1 > p2.

[0041] Here, "far" means that the distance between the real pedestrian corresponding to the pedestrian in the crowd image and the image acquisition device for acquiring the crowd image is far, and "near" means that the distance between the real pedestrian corresponding to the pedestrian in the crowd image and the image acquisition device for acquiring the crowd image is near.

[0042] The preset perspective mapping relationship can indicate the image scales corresponding to different positions in the crowd image. The process of determining the preset perspective mapping relationship will be described in detail later in combination with possible implementation manners of the present disclosure, and will not be elaborated here.

[0043] For the target human head key points included in the crowd image indicated by the target positioning map, based on the preset perspective mapping relationship, the image scale corresponding to the target human head key points in the crowd image can be determined. Then, based on the image scale corresponding to the target human head key points in the crowd image and in combination with the real height of the pedestrian corresponding to the target human head key points, the positions of the target human foot key points corresponding to the pedestrian in the crowd image can be effectively determined. The process of determining the positions of the target human foot key points included in the crowd image based on the target positioning map and the preset perspective mapping relationship will be described in detail later in combination with possible implementation manners of the present disclosure, and will not be elaborated here.

[0044] In the embodiments of the present disclosure, for the crowd image, human head key point positioning can be performed end-to-end to obtain a target positioning map for indicating the positions of the target human head key points included in the crowd image. Since the preset perspective mapping relationship can be used to indicate the image scales corresponding to different positions in the crowd image, based on the target human head key points and the preset perspective mapping relationship, the positions of the target human foot key points with relatively high accuracy can be effectively obtained.

[0045] In a possible implementation, the crowd positioning method further includes: obtaining a plurality of labeled human body frames obtained by performing human body frame annotation on pedestrians at different positions in the crowd image; and determining a preset perspective mapping relationship based on the plurality of labeled human body frames.

[0046] By selecting pedestrians at different positions (far, middle, and near) in the crowd image for human body frame annotation, a plurality of labeled human body frames in the crowd image can be obtained. Based on the proportional relationship between the height of the labeled human body frame and the actual height of the pedestrian, the image scale corresponding to a limited position (the position of the labeled human body frame) in the crowd image can be determined. Furthermore, based on the image scale corresponding to the limited position, a fitting can be further performed to effectively obtain the image scale corresponding to each position in the crowd image, that is, the preset perspective mapping relationship is obtained.

[0047] Figure 2 A schematic diagram showing a crowd image and its corresponding preset perspective mapping relationship according to an embodiment of the present disclosure is shown. As Figure 2 As shown, by selecting pedestrians at different positions (far, middle, and near) in the crowd image for human body frame annotation, four labeled human body frames A, B, C, and D at different positions in the crowd image are obtained. Furthermore, based on the four labeled human body frames A, B, C, and D, a fitting is performed to effectively obtain the preset perspective mapping relationship corresponding to the crowd image.

[0048] In a possible implementation, determining a preset perspective mapping relationship based on a plurality of labeled human body frames includes: for any one of the labeled human body frames, determining a reference image scale corresponding to a reference human foot key point in the labeled human body frame; and fitting to obtain the preset perspective mapping relationship according to the third image coordinate of the reference human foot key point in each labeled human body frame and the reference image scale corresponding to the reference human foot key point in each labeled human body frame.

[0049] Along the column direction of the crowd image, the image scales at different positions change linearly. Therefore, after determining the image scale corresponding to the reference human foot key point at a limited position in the crowd image according to the labeled human body frame, a linear function fitting can be used to effectively obtain the image scale corresponding to each position in the crowd image, that is, the preset perspective mapping relationship corresponding to the crowd image is obtained.

[0050] Since pedestrians stand vertically, the height of the labeled human body box can be regarded as the height of the pedestrian in the crowd image. The height of the labeled human body box can be represented by the number of pixel rows occupied by the labeled human body box. For example, if the labeled human body box occupies 17 pixel rows in the crowd image, the height of the labeled human body box is 17. Assuming that the true height of the pedestrian corresponding to the labeled human body box is 1.7 meters, the position of the reference human foot key point in the labeled human body box can be determined. It takes 17 pixel rows to represent 1.7 meters in the real world. Assuming the unit height is 1m, therefore, it takes 10 pixel rows to represent 1 meter in the real world at the position of the reference human foot key point in the labeled human body box, that is, the reference image scale corresponding to the reference human foot key point in the labeled human body box is 10. The true height of the pedestrian corresponding to the labeled human body box can be appropriately selected according to the actual situation, and the present disclosure does not make specific limitations on this.

[0051] The reference human foot key point in the labeled human body box can be the midpoint of the bottom edge of the labeled human body box, or other pixel points in the labeled human body box. The present disclosure does not make specific limitations on this.

[0052] Still taking the above Figure 2 as an example, after obtaining four labeled human body boxes A, B, C, and D in the crowd image, the reference image scale corresponding to the reference human foot key point in each labeled human body box is determined in the above manner. Furthermore, a linear function fitting is performed according to the third image coordinates of the reference human foot key points in the four labeled human body boxes and their corresponding reference image scales, and a linear mapping function p = a*y + b is obtained.

[0053] Among them, the image coordinates refer to the position coordinates in the pixel coordinate system of the crowd image. For example, taking the upper left corner of the crowd image as the coordinate origin (0, 0), the direction parallel to the row direction of the image is the x-axis direction, and the direction parallel to the column direction of the image is the y-axis direction to construct the pixel coordinate system of the crowd image. The units of the abscissa and ordinate of the image coordinates are both pixel points. For example, if the third image coordinate of the reference human foot key point is (10, 15), it indicates that the reference human foot key point is the pixel point located at the 10th row and 15th column in the crowd image.

[0054] The linear mapping function p = a*y + b is the functional representation form of the preset perspective mapping relationship corresponding to the crowd image. Among them, a and b are parameters obtained by linear function fitting, y is the ordinate of the image coordinates at different positions in the crowd image, and p is the image scale corresponding to this position. Using the linear mapping function p = a*y + b, the image scale corresponding to each position in the crowd image can be determined.

[0055] Before or after determining the preset perspective mapping relationship corresponding to the crowd image, head key point positioning is performed on the crowd image to obtain the target positioning map corresponding to the crowd image.

[0056] In a possible implementation, head key point localization is performed on a crowd image to obtain a target localization map corresponding to the crowd image, including: performing head key point localization on the crowd image to determine a predicted localization map corresponding to the crowd image, where the predicted localization map is used to indicate the predicted confidence of each pixel point in the crowd image being a head key point; based on a preset confidence threshold, performing image processing on the predicted localization map to obtain an initial localization map, where the initial localization map is used to indicate the positions of the initial head key points included in the crowd image; determining the target neighborhood corresponding to each initial head key point in the initial localization map; and based on the predicted localization map, performing filtering processing on the target neighborhood corresponding to each initial head key point to obtain the target localization map.

[0057] Perform head key point localization on the crowd image, end-to-end determine the predicted confidence of each pixel point in the crowd image being a head key point, and then perform threshold segmentation on the predicted localization map through the preset confidence threshold to determine the initial localization map used to indicate the positions of the initial head key points included in the crowd image, and then determine the target neighborhood of each initial head key point in the initial localization map, so that based on the predicted localization map, further filtering processing is performed on the target neighborhood corresponding to the initial head key point to obtain a target localization map with higher accuracy to accurately indicate the positions of the target head key points included in the crowd image.

[0058] In one example, a trained head key point localization neural network can be used to perform head key point localization on the crowd image. Specifically, input the crowd image into the trained head key point localization neural network, and after the localization of the head key point localization neural network, directly output the predicted localization map. The specific network structure of the trained head key point localization neural network and the training process can adopt the network structure and training process in related technologies, and the present disclosure does not make specific limitations thereon.

[0059] In one example, the pixel value of each pixel point in the predicted localization map represents the predicted confidence of this pixel point, that is, the probability that this pixel point is a head key point. Perform a sigmoid operation on the predicted localization map to make the pixel value of each pixel point in the predicted localization map between 0 and 1. For example, if the pixel value of a certain pixel point in the predicted localization map is 0.7, it means that the probability that this pixel point is a head key point is 0.7.

[0060] Since the predicted localization map is only used to indicate the predicted confidence of each pixel point in the crowd image being a head key point, therefore, by presetting the confidence threshold and performing threshold segmentation on the predicted localization map, an initial localization map used to indicate the positions of the initial head key points included in the crowd image can be effectively obtained. The specific value of the preset confidence threshold can be flexibly set according to the actual situation, and the present disclosure does not make specific limitations thereon.

[0061] Compare the pixel values of each pixel point in the predicted localization map with a preset confidence threshold. When the pixel value of a certain pixel point in the predicted localization map is greater than or equal to the preset confidence threshold, determine the pixel value of the corresponding pixel point in the initial localization map as 1; when the pixel value of a certain pixel point in the predicted localization map is less than the preset confidence threshold, determine the pixel value of the corresponding pixel point in the initial localization map as 0.

[0062] The initial localization map and the crowd image have the same size. The positions of the pixel points with pixel value 1 in the initial localization map are used to indicate the positions of the initial head key points included in the crowd image. For example, when the pixel value of the pixel point with image coordinates (x, y) in the initial localization map is 1, it can be determined that the pixel point with image coordinates (x, y) in the crowd image is an initial head key point; when the pixel value of the pixel point with image coordinates (x, y) in the initial localization map is 0, it can be determined that the pixel point with image coordinates (x, y) in the crowd image is a part other than the initial head key point.

[0063] To avoid the problem of false detection that one human head corresponds to multiple initial head key points, further determine the target neighborhood corresponding to each initial head key point in the initial localization map, and filter the target neighborhood corresponding to each initial head key point to obtain a target localization map with higher accuracy. In the target localization map, one human head corresponds to one target head key point.

[0064] In a possible implementation manner, determining the target neighborhood corresponding to each initial head key point in the initial localization map includes: determining the target neighborhood corresponding to each initial head key point according to a first preset neighborhood radius.

[0065] By presetting a fixed first preset neighborhood radius, the target neighborhood corresponding to each initial head key point can be quickly determined. The specific value of the first neighborhood radius can be flexibly set according to the actual situation, and the present disclosure does not make specific limitations on this.

[0066] For example, if the first preset neighborhood radius is 2, then for any initial head key point i, the target neighborhood corresponding to the initial head key point i includes the pixel points whose pixel distance from the initial head key point i does not exceed 2 pixel points.

[0067] In a possible implementation manner, determining the target neighborhood corresponding to each initial head key point in the initial localization map includes: for any initial head key point, determining the target neighborhood corresponding to the initial head key point based on the position of the initial head key point in the crowd image and a preset perspective mapping relationship.

[0068] After determining the positions of the initial head key points included in the crowd image based on the initial positioning map, the target neighborhood matching each initial head key point can be determined based on a preset perspective mapping relationship.

[0069] In a possible implementation manner, determining the target neighborhood corresponding to the initial head key point based on the position of the initial head key point in the crowd image and the preset perspective mapping relationship includes: determining the target image scale corresponding to the position of the initial head key point in the crowd image based on the preset perspective mapping relationship; determining the head box height corresponding to the initial head key point based on the target image scale; and determining the target neighborhood corresponding to the initial head key point based on the head box height corresponding to the initial head key point.

[0070] Based on the preset perspective mapping relationship, the target image scale corresponding to the position of the initial head key point in the crowd image can be quickly determined, and then the head box height corresponding to the initial head key point can be determined based on the target image scale, so that the target neighborhood matching it can be further determined according to the head box height.

[0071] For example, for a certain initial head key point i in the initial positioning map, the image coordinates of the initial head key point i in the crowd image are (h x , h y ), then according to the preset perspective mapping relationship (linear mapping function p = a*y + b) corresponding to the crowd image, the target scale corresponding to the initial head key point i can be determined as p i = a*h y + b. Assuming that the true head box height of the pedestrians in the crowd image is 0.4 meters * 0.4 meters, then the head box height corresponding to the initial head key point i in the crowd image is s i = 0.4*p i . According to the head box height s i = 0.4*p i corresponding to the initial head key point, the target neighborhood with a matching size is determined.

[0072] In a possible implementation manner, determining the target neighborhood corresponding to the initial head key point based on the head box height corresponding to the initial head key point includes: when the head box height is greater than the preset head box height threshold, determining the target neighborhood corresponding to the initial head key point based on the second preset neighborhood radius; or, when the head box height is less than or equal to the preset head box height threshold, determining the target neighborhood corresponding to the initial head key point based on the third neighborhood radius, where the second neighborhood radius is greater than the third neighborhood radius.

[0073] When the height of the head box is greater than the preset head box height threshold, it can be determined that the size of the head box is relatively large. Therefore, a larger second neighborhood radius is used for subsequent filtering processing; when the height of the head box is less than or equal to the preset head box height threshold, it can be determined that the size of the head box is relatively small. Therefore, a smaller third neighborhood radius is used for subsequent filtering processing. By flexibly determining the neighborhood radius, the accuracy of the filtering operation can be improved. The specific values of the preset head box height threshold, the second neighborhood radius, and the third neighborhood radius can be flexibly set according to the actual situation, and the present disclosure does not make specific limitations on this.

[0074] In one example, the head box height threshold is 32. For a certain initial head key point i, when the height s of its corresponding head box i > 32, its target neighborhood is determined based on the second neighborhood radius 2; when the height s of its corresponding head box i <= 32, its target neighborhood is determined based on the third neighborhood radius 1.

[0075] When the second preset neighborhood radius is 2, the target neighborhood corresponding to the initial head key point i includes pixel points whose pixel distance from the initial head key point i does not exceed 2 pixel points. When the third preset neighborhood radius is 1, the target neighborhood corresponding to the initial head key point i includes pixel points whose pixel distance from the initial head key point i does not exceed 1 pixel point.

[0076] After determining the target neighborhood corresponding to each initial head key point in the initial positioning map based on the above method, the initial positioning map is filtered using the target neighborhood corresponding to each initial head key point to obtain a target positioning map with a higher accuracy.

[0077] In a possible implementation manner, based on the predicted positioning map, filtering processing is performed on the target neighborhood corresponding to each initial head key point to obtain a target positioning map, including: for any initial head key point i, determining whether there is at least one other initial head key point in the target neighborhood corresponding to the initial head key point i; when there is at least one other initial head point j in the target neighborhood corresponding to the initial head key point i, based on the predicted positioning map, determining the predicted confidence corresponding to the initial head point i, and the predicted confidence corresponding to at least one other initial head point j; based on the initial head key point i and the initial head key point with the largest predicted confidence among at least one other initial head key points j, determining the target head key point in the target neighborhood corresponding to the initial head key point i.

[0078] For the initial positioning map, the image coordinates are (x i , y i) The initial head key point i, checks whether there are other initial head key points within its target neighborhood. If there is another initial head key point j with image coordinates (x j , y j ), then according to the predicted localization map, determine the predicted confidence corresponding to the initial head key point i, and the predicted confidence corresponding to the initial head point j. When the predicted confidence of the initial head key point i is greater than the predicted confidence of the initial head key point j, keep the pixel value of the pixel point with image coordinates (x i , y i ) as 1, and update the pixel value of the pixel point with image coordinates (x j , y j ) to 0, that is, filter out the initial head key point j in the initial localization map. And so on, traverse each initial head key point in the initial localization map to obtain the final target localization map.

[0079] After determining the target localization map corresponding to the crowd image and the preset perspective mapping relationship based on the above method, the positions of the target foot key points included in the crowd image can be determined based on the target localization map and the preset perspective mapping relationship.

[0080] In a possible implementation, determining the positions of the target foot key points included in the crowd image based on the target localization map and the preset perspective mapping relationship includes: for any one target head key point, determine the first image coordinates of the target head key point in the crowd image according to the target localization map; based on the preset perspective mapping relationship, perform coordinate transformation on the first image coordinates to obtain the second image coordinates of the target foot key point corresponding to the target head key point in the crowd image.

[0081] Since the preset perspective mapping relationship can indicate the image scales at different positions in the crowd image, therefore, based on the preset perspective mapping relationship, the first image coordinates of the target head key point can be coordinate-transformed to obtain the second image coordinates of the target foot key point corresponding to the target head key point in the crowd image.

[0082] For example, for a certain target head key point in the target localization map, the first image coordinates of the target head key point in the crowd image are (h x , h y ), and the first image coordinates are known. According to the preset perspective mapping relationship corresponding to the crowd image, the first image coordinates (h x , h y ) can be coordinate-transformed to obtain the second image coordinates of the target foot key point corresponding to the target head key point as (f x , f y ).

[0083] In a possible implementation, based on a preset perspective mapping relationship, coordinate transformation is performed on the first image coordinates to obtain the second image coordinates of the target human foot key point corresponding to the target human head key point in the crowd image, including: determining the target image scale corresponding to the target human foot key point based on the preset perspective mapping relationship; determining the image distance between the target human head key point and the target human foot key point based on the target image scale; and determining the second image coordinates of the target human foot key point in the crowd image according to the first image coordinates and the image distance.

[0084] Since pedestrians stand vertically, it is possible to determine that f x = h x . According to the preset perspective mapping relationship (linear mapping function p = a*y + b), it can be determined that the target scale corresponding to the target human foot key point is a*f x + b. Assuming that in the real world, the distance from the center of the human head to the human foot is 1.5 meters, that is, the real distance between the target human head key point and the target human foot key point is 1.5 meters, then the image distance between the target human foot key point and the target human head key point in the crowd image is 1.5*(a*f x + b), that is, h x = f y - 1.5*(a*f x + b), and then it can be determined that f y = (h y + 1.5*b) / (1 - 1.5*a). Therefore, the second image coordinates of the target human foot key point in the crowd image are (h x , (h y + 1.5*b) / (1 - 1.5*a)).

[0085] By traversing each target human head key point in the target positioning map, the position of the target human foot key point corresponding to each target human head key point in the crowd image can be determined. Furthermore, based on the positions of the target human foot key points in adjacent front and back frame crowd images, crowd behavior analysis such as passenger flow statistics and behavior trajectory analysis can be performed. The present disclosure does not make specific limitations in this regard.

[0086] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0087] In addition, the present disclosure also provides a crowd positioning device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any crowd positioning method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be elaborated further.

[0088] Figure 3 A block diagram of a crowd positioning device according to an embodiment of the present disclosure is shown. As Figure 3 shown, the device 30 includes:

[0089] A head key point positioning module 31, configured to perform head key point positioning on a crowd image to obtain a target positioning map corresponding to the crowd image, where the target positioning map is used to indicate the positions of target head key points included in the crowd image;

[0090] A human foot key point determination module 32, configured to determine the positions of target human foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales corresponding to different positions in the crowd image.

[0091] In a possible implementation manner, the human foot key point determination module 32 includes:

[0092] A first determination sub-module, configured to, for any one target head key point, determine a first image coordinate of the target head key point in the crowd image according to the target positioning map;

[0093] A second determination sub-module, configured to perform coordinate conversion on the first image coordinate based on the preset perspective mapping relationship to obtain a second image coordinate of the target human foot key point corresponding to the target head key point in the crowd image.

[0094] In a possible implementation manner, the second determination sub-module is specifically configured to:

[0095] Determine a target image scale corresponding to the target human foot key point based on the preset perspective mapping relationship;

[0096] Determine an image distance between the target head key point and the target human foot key point based on the target image scale;

[0097] Determine the second image coordinate of the target human foot key point in the crowd image according to the first image coordinate and the image distance.

[0098] In a possible implementation manner, the head key point positioning module 31 includes:

[0099] A third determination sub-module, configured to perform head key point positioning on the crowd image to determine a predicted positioning map corresponding to the crowd image, where the predicted positioning map is used to indicate the predicted confidence levels of each pixel point in the crowd image being head key points;

[0100] The fourth determination sub-module is configured to perform image processing on the predicted positioning map based on a preset confidence threshold to obtain an initial positioning map, where the initial positioning map is used to indicate the positions of the initial head key points included in the crowd image;

[0101] The fifth determination sub-module is configured to determine a target neighborhood corresponding to each initial head key point in the initial positioning map;

[0102] The sixth determination sub-module is configured to perform filtering processing on the target neighborhood corresponding to each initial head key point based on the predicted positioning map to obtain a target positioning map.

[0103] In a possible implementation manner, the fifth determination sub-module is specifically configured to:

[0104] Determine the target neighborhood corresponding to each initial head key point according to a first preset neighborhood radius.

[0105] In a possible implementation manner, the fifth determination sub-module includes:

[0106] The first determination unit is configured to, for any initial head key point, determine the target neighborhood corresponding to the initial head key point based on the position of the initial head key point in the crowd image and a preset perspective mapping relationship.

[0107] In a possible implementation manner, the first determination unit includes:

[0108] The first determination subunit is configured to determine the target image scale corresponding to the position of the initial head key point in the crowd image based on the preset perspective mapping relationship;

[0109] The second determination subunit is configured to determine the head box height corresponding to the initial head key point based on the target image scale;

[0110] The third determination subunit is configured to determine the target neighborhood corresponding to the initial head key point based on the head box height corresponding to the initial head key point.

[0111] In a possible implementation manner, the third determination subunit is specifically configured to:

[0112] In the case where the head box height is greater than a preset head box height threshold, determine the target neighborhood corresponding to the initial head key point according to a second preset neighborhood radius; or,

[0113] In the case where the head box height is less than or equal to the preset head box height threshold, determine the target neighborhood corresponding to the initial head key point according to a third neighborhood radius, where the second neighborhood radius is greater than the third neighborhood radius.

[0114] In a possible implementation manner, the sixth determination sub-module is specifically configured to:

[0115] For any initial human head key point i, determine whether there is at least one other initial human head center point in the target neighborhood corresponding to the initial human head key point i;

[0116] When there is at least one other initial human head point j in the target neighborhood corresponding to the initial human head key point i, based on the prediction positioning map, determine the prediction confidence corresponding to the initial human head point i and the prediction confidence corresponding to at least one other initial human head point j;

[0117] Based on the initial human head key point with the maximum prediction confidence among the initial human head key point i and at least one other initial human head key point j, determine the target human head key point in the target neighborhood corresponding to the initial human head key point i.

[0118] In a possible implementation manner, the apparatus 30 further includes:

[0119] An acquisition module, configured to acquire a plurality of labeled human body frames obtained by performing human body frame labeling on pedestrians at different positions in the crowd image;

[0120] A perspective mapping relationship determination module, configured to determine a preset perspective mapping relationship based on the plurality of labeled human body frames.

[0121] In a possible implementation manner, the perspective mapping relationship determination module is specifically configured to:

[0122] For any one of the labeled human body frames, determine the reference image scale corresponding to the reference human foot key point in the labeled human body frame;

[0123] According to the third image coordinate of the reference human foot key point in each labeled human body frame and the reference image scale corresponding to the reference human foot key point in each labeled human body frame, fit to obtain a preset perspective mapping relationship.

[0124] In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be described in detail here.

[0125] The embodiments of the present disclosure further propose a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0126] The embodiments of the present disclosure further propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.

[0127] Embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0128] The electronic device may be provided as a terminal, a server, or other forms of devices.

[0129] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Referring to Figure 4 , the electronic device 800 may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, and other terminal devices.

[0130] Referring to Figure 4 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0131] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0132] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0133] The power supply component 806 provides power for various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0134] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0135] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0136] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0137] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a complementary metal oxide semiconductor (CMOS) or charge coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0138] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as Wi-Fi, 2G, 3G, 4G, Long Term Evolution (LTE) of Universal Mobile Telecommunications Technology, 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0139] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0140] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 804 including computer program instructions, and the computer program instructions can be executed by a processor 820 of the electronic device 800 to complete the above-described method.

[0141] The present disclosure relates to the field of augmented reality. By acquiring the image information of a target object in the real environment, relevant features, states, and attributes of the target object are detected or recognized through various vision-related algorithms, thereby obtaining an AR effect that combines virtual and real and matches specific applications. Exemplarily, the target object may involve the face, limbs, gestures, actions, etc. related to the human body, or identification markers, landmarks related to objects, or sand tables, display areas, or display items related to venues or places. The vision-related algorithms may involve visual positioning, SLAM, three-dimensional reconstruction, image registration, background segmentation, key point extraction and tracking of objects, pose or depth detection of objects, etc. Specific applications can not only involve interactive scenarios such as navigation, guidance, explanation, reconstruction, virtual effect overlay display related to real scenes or items, but also involve special effect processing related to people, such as makeup beautification, limb beautification, special effect display, virtual model display, etc. The relevant features, states, and attributes of the target object can be detected or recognized through a convolutional neural network. The convolutional neural network is a network model obtained by training the model based on a deep learning framework.

[0142] Figure 5 FIG. shows a block diagram of another electronic device according to an embodiment of the present disclosure. Referring to Figure 5 , the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 5 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0143] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM) or the like.

[0144] In an exemplary embodiment, a non - volatile computer - readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the above - mentioned computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above - mentioned method.

[0145] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer - readable storage medium having thereon computer - readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0146] A computer - readable storage medium may be a tangible device that can retain and store instructions for use by an instruction - execution device. A computer - readable storage medium may be, for example (but not limited to), an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non - exhaustive list) of the computer - readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read - only memory (ROM), an erasable programmable read - only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read - only memory (CD - ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer - readable storage medium used herein is not construed as being an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0147] The computer - readable program instructions described herein can be downloaded from the computer - readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer - readable program instructions from the network and forwards the computer - readable program instructions for storage in the computer - readable storage medium in each computing / processing device.

[0148] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0149] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0150] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0151] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0152] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0153] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0154] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated herein.

[0155] Those skilled in the art can understand that in the above methods of the specific implementation manners, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0156] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0157] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for crowd positioning, characterized in that Including: Performing head key point localization on a crowd image to obtain a target localization map corresponding to the crowd image, where the target localization map is used to indicate the positions of target head key points included in the crowd image; Based on the target localization map and a preset perspective mapping relationship, determining the positions of target foot key points included in the crowd image, where the preset perspective mapping relationship is used to indicate the image scales corresponding to different positions in the crowd image; Wherein, the method further includes: Obtaining a plurality of labeled human body frames obtained by performing human body frame annotation on pedestrians at different positions in the crowd image; Based on the plurality of labeled human body frames, determining the preset perspective mapping relationship.

2. The method according to claim 1, characterized in that, The determining the positions of target foot key points included in the crowd image based on the target localization map and the preset perspective mapping relationship includes: For any one of the target head key points, determining a first image coordinate of the target head key point in the crowd image according to the target localization map; Based on the preset perspective mapping relationship, performing coordinate transformation on the first image coordinate to obtain a second image coordinate of the target foot key point corresponding to the target head key point in the crowd image.

3. The method according to claim 2, wherein The performing coordinate transformation on the first image coordinate based on the preset perspective mapping relationship to obtain a second image coordinate of the target foot key point corresponding to the target head key point in the crowd image includes: Based on the preset perspective mapping relationship, determining a target image scale corresponding to the target foot key point; Based on the target image scale, determining an image distance between the target head key point and the target foot key point; According to the first image coordinate and the image distance, determining the second image coordinate of the target foot key point in the crowd image.

4. The method according to any one of claims 1 to 3, characterized in that The performing head key point localization on a crowd image to obtain a target localization map corresponding to the crowd image includes: Performing head key point localization on the crowd image to determine a predicted localization map corresponding to the crowd image, where the predicted localization map is used to indicate the prediction confidence of each pixel point in the crowd image being a head key point; Based on a preset confidence threshold, performing image processing on the predicted localization map to obtain an initial localization map, where the initial localization map is used to indicate the positions of initial head key points included in the crowd image; Determining a target neighborhood corresponding to each initial head key point in the initial localization map; Based on the predicted localization map, performing filtering processing on the target neighborhood corresponding to each initial head key point to obtain the target localization map.

5. The method according to claim 4, wherein The determining a target neighborhood corresponding to each initial head key point in the initial localization map includes: According to a first preset neighborhood radius, determining a target neighborhood corresponding to each initial head key point.

6. The method according to claim 4, characterized in that, The determining a target neighborhood corresponding to each initial head key point in the initial localization map includes: For any one of the initial head key points, based on the position of the initial head key point in the crowd image and the preset perspective mapping relationship, determine the target neighborhood corresponding to the initial head key point.

7. The method according to claim 6, wherein The determining the target neighborhood corresponding to the initial head key point based on the position of the initial head key point in the crowd image and the preset perspective mapping relationship includes: Based on the preset perspective mapping relationship, determine the target image scale corresponding to the position of the initial head key point in the crowd image; Based on the target image scale, determine the head box height corresponding to the initial head key point; Based on the head box height corresponding to the initial head key point, determine the target neighborhood corresponding to the initial head key point.

8. The method according to claim 7, wherein The determining the target neighborhood corresponding to the initial head key point based on the head box height corresponding to the initial head key point includes: When the head box height is greater than the preset head box height threshold, determine the target neighborhood corresponding to the initial head key point based on the second preset neighborhood radius; or, When the head box height is less than or equal to the preset head box height threshold, determine the target neighborhood corresponding to the initial head key point based on the third neighborhood radius, where the second neighborhood radius is greater than the third neighborhood radius.

9. The method according to claim 4, wherein The filtering process of the target neighborhood corresponding to each initial head key point based on the predicted positioning map to obtain the target positioning map includes: For any initial head key point i, determine whether there is at least one other initial head center point in the target neighborhood corresponding to the initial head key point i; When there is at least one other initial head point j in the target neighborhood corresponding to the initial head key point i, based on the predicted positioning map, determine the predicted confidence corresponding to the initial head point i and the predicted confidence corresponding to the at least one other initial head point j; Based on the initial head key point i and the initial head key point with the highest predicted confidence among the at least one other initial head key points j, determine the target head key point in the target neighborhood corresponding to the initial head key point i.

10. The method according to claim 1, characterized in that, The determining the preset perspective mapping relationship based on the multiple labeled human body boxes includes: For any one of the labeled human body boxes, determine the reference image scale corresponding to the reference human foot key point in the labeled human body box; According to the third image coordinates of the reference human foot key point in each labeled human body box and the reference image scale corresponding to the reference human foot key point in each labeled human body box, fit to obtain the preset perspective mapping relationship.

11. A crowd positioning device, characterized in that, Includes: A head key point positioning module, configured to perform head key point positioning on a crowd image to obtain a target positioning map corresponding to the crowd image, where the target positioning map is used to indicate the positions of the target head key points included in the crowd image; A human foot key point determination module, configured to determine the positions of target human foot key points included in the crowd image based on the target positioning map and a preset perspective mapping relationship, where the preset perspective mapping relationship is used to indicate the image scales corresponding to different positions in the crowd image; Wherein, the device further includes: An acquisition module, configured to acquire a plurality of labeled human body frames obtained by performing human body frame labeling on pedestrians at different positions in the crowd image; A perspective mapping relationship determination module, configured to determine the preset perspective mapping relationship based on the plurality of labeled human body frames.

12. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Regional population retention detection method and device and storage medium

    CN107481260A

  • Method and apparatus for detecting persons, and non-transitory computer-readable recording medium

    US20160335490A1