A crowd counting method and device, electronic equipment and storage medium
By combining head landmark localization and crowd density detection, and selecting an appropriate density distribution map for crowd counting, the accuracy problem of crowd counting in different scenarios is solved, achieving high-precision crowd statistics in both sparse and dense scenarios.
Patent Information
- Application Number
- CN202210306681.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2042-03-25
AI Technical Summary
Existing technologies are not accurate enough in crowd counting in various scenarios, especially in sparsely populated and densely populated scenarios where misidentification occurs.
A method combining head key point localization and crowd density detection is used to obtain a first crowd density distribution map and a second crowd density distribution map, respectively. The appropriate distribution map is selected according to the number of people threshold for crowd counting. The number of people is calculated by combining Gaussian kernel rendering and density distribution map weighting.
It improves the accuracy of crowd counting and is suitable for various scenarios, including sparsely populated and densely populated areas, accurately counting the number of people in the entire area and the region of interest.
Smart Images

Figure CN114663837B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a crowd counting method and device, electronic equipment and storage medium. BACKGROUND
[0002] The crowd counting technology refers to a technology of evaluating real-time number of people, distribution of people, density of crowd and the like in a video picture through a computer vision algorithm. In the related technology, a regression manner is usually used to predict a crowd density distribution map, and the number of people in the whole picture is estimated while the position distribution of the crowd is estimated. The accuracy of this manner is relatively high in a crowded scene, but misrecognition is prone to occur in a sparse scene. Therefore, how to improve the accuracy of crowd counting in various scenes becomes a problem to be solved at present. SUMMARY
[0003] The present disclosure provides a technical solution of a crowd counting method and device, electronic equipment and storage medium.
[0004] According to an aspect of the present disclosure, a crowd counting method is provided, comprising:
[0005] obtaining a crowd image;
[0006] obtaining a first number of people corresponding to the crowd image and a first crowd density distribution map corresponding to the crowd image based on head key point positioning on the crowd image;
[0007] obtaining a second crowd density distribution map corresponding to the crowd image based on crowd density detection on the crowd image;
[0008] selecting a target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map based on the first number of people and a first preset number of people threshold;
[0009] determining a crowd counting result of the crowd image based on the target crowd density distribution map.
[0010] The crowd counting method provided by the present disclosure can be applied to crowd counting in a sparse scene and crowd counting in a crowded scene. In the present disclosure, a first crowd density distribution map and a second crowd density distribution map are obtained based on head key point positioning and crowd density detection, and crowd counting is performed in combination with the first crowd density distribution map and the second crowd density distribution map, thereby improving the accuracy of crowd counting in various scenes.
[0011] In a possible implementation, the target crowd density distribution corresponding to the crowd image is selected from the first crowd density distribution and the second crowd density distribution based on the first number and a first preset number threshold, and the method comprises the following steps.
[0012] In a case where the first number is less than the first preset number threshold, the first crowd density distribution is determined as the target crowd density distribution; or
[0013] In a case where the first number is greater than or equal to the first preset number threshold, the second crowd density distribution is determined as the target crowd density distribution.
[0014] In the embodiments of the present disclosure, in a sparse personnel scene, crowd counting is performed based on head key point positioning, and in a dense crowd scene, crowd counting is performed based on crowd density detection, thereby improving the accuracy of crowd counting.
[0015] In a possible implementation, the crowd counting result of the crowd image is determined based on the target crowd density distribution, and the method comprises the following steps.
[0016] The total number of people in the crowd image is determined based on the target crowd density distribution; and / or
[0017] The number of people in a region of interest in the crowd image is determined based on the target crowd density distribution and the region of interest.
[0018] In the embodiments of the present disclosure, the global number of people can be counted, and the number of people in a partial region can also be counted, thereby improving flexibility, refining the counting result, and further improving the accuracy of crowd counting.
[0019] In a possible implementation, the total number of people in the crowd image is determined, and the method comprises the following steps.
[0020] The total number of people in the crowd image is obtained based on the weighting of the density values of the pixel points in the target crowd density distribution; and / or
[0021] The number of people in the region of interest is determined, and the method comprises the following steps.
[0022] The number of people in the region of interest is obtained based on the weighting of the density values of the pixel points corresponding to the region of interest in the target crowd density distribution.
[0023] In a possible implementation, the method further comprises the following steps.
[0024] A second number of people corresponding to a region of interest in the crowd image is determined based on head key point positioning performed on the crowd image.
[0025] In a case where the second number of people is less than a second preset number threshold, the second number of people is determined as the number of people in the region of interest.
[0026] In a case where the second number of people is greater than or equal to the second preset number threshold, the number of people in the region of interest is obtained based on a weighted sum of density values corresponding to pixel points in the second crowd density distribution map corresponding to the region of interest.
[0027] In a possible implementation, the obtaining, based on head key point positioning performed on the crowd image, of the first number of people corresponding to the crowd image and the first crowd density distribution map corresponding to the crowd image comprises:
[0028] performing head key point positioning on the crowd image to obtain a target positioning map corresponding to the crowd image, wherein the target positioning map is used to indicate positions of target head key points included in the crowd image;
[0029] determining the first number of people based on the target positioning map;
[0030] determining the first crowd density distribution map corresponding to the crowd image based on the target positioning map.
[0031] In a possible implementation, the performing head key point positioning on the crowd image to obtain a target positioning map corresponding to the crowd image comprises:
[0032] performing head key point positioning on the crowd image to obtain a predicted positioning map corresponding to the crowd image, wherein the predicted positioning map is used to indicate a predicted confidence of each pixel point in the crowd image being a head key point;
[0033] performing image processing on the predicted positioning map based on a preset confidence threshold to obtain an initial positioning map, wherein the initial positioning map is used to indicate positions of initial head key points included in the crowd image;
[0034] determining a target neighborhood corresponding to each initial head key point in the initial positioning map;
[0035] performing filtering processing on the target neighborhood corresponding to each initial head key point based on the predicted positioning map to obtain the target positioning map corresponding to the crowd image.
[0036] In a possible implementation, the determining a target neighborhood corresponding to each initial head key point in the initial positioning map comprises:
[0037] According to a preset neighborhood radius, a target neighborhood corresponding to each of the initial head key points is determined, wherein the preset neighborhood radius is determined based on a position of the initial head key point in the crowd image and a preset perspective relationship corresponding to the crowd image, and the preset perspective mapping relationship corresponding to the crowd image is used to indicate image scales corresponding to different positions in the crowd image.
[0038] In a possible implementation, the filtering processing on the target neighborhood corresponding to each of the initial head key points based on the predicted positioning map to obtain the target positioning map comprises:
[0039] For any one of the initial head key points, it is determined whether there is at least one other initial head key point in the target neighborhood corresponding to the initial head key point i.
[0040] In the case that there is at least one other initial head key point j in the target neighborhood corresponding to the initial head key point i, the predicted confidence corresponding to the initial head key point i and the predicted confidence corresponding to the at least one other initial head key point j are determined based on the predicted positioning map.
[0041] Based on the initial head key point i and the initial head key point with the maximum predicted confidence among the at least one other initial head key point j, a target head key point in the target neighborhood corresponding to the initial head key point i is determined.
[0042] In a possible implementation, the first crowd density distribution map corresponding to the crowd image is obtained based on the target positioning map, comprising:
[0043] According to a position of each target head key point indicated by the target positioning map, a Gaussian kernel is used to render each target head key point to obtain the first crowd density distribution map.
[0044] According to an aspect of the present disclosure, a crowd counting device is provided, comprising:
[0045] An acquisition module is configured to acquire a crowd image.
[0046] A positioning module is configured to obtain a first number of people corresponding to the crowd image and a first crowd density distribution map corresponding to the crowd image based on head key point positioning performed on the crowd image.
[0047] A detection module is configured to obtain a second crowd density distribution map corresponding to the crowd image based on crowd density detection performed on the crowd image.
[0048] The first selection module is configured to select a target crowd density distribution graph corresponding to the crowd image from the first crowd density distribution graph and the second crowd density distribution graph based on the first number of people and a first preset number of people threshold.
[0049] The first determination module is configured to determine a crowd counting result of the crowd image based on the target crowd density distribution graph.
[0050] According to an aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the above method.
[0051] According to an aspect of the present disclosure, a computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the above method.
[0052] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure. Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0053] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the technical solutions of the present disclosure together with the specification.
[0054] Figure 1 A flow chart of a crowd counting method according to an embodiment of the present disclosure is shown;
[0055] Figure 2 An exemplary schematic diagram of a target positioning map in an embodiment of the present disclosure is shown;
[0056] Figure 3 An exemplary schematic diagram of a first crowd density distribution graph in an embodiment of the present disclosure is shown;
[0057] Figure 4 A schematic diagram of a crowd image and its corresponding preset perspective mapping relationship according to an embodiment of the present disclosure is shown;
[0058] Figure 5 A block diagram of a crowd counting device according to an embodiment of the present disclosure is shown;
[0059] Figure 6 A block diagram of an electronic device 800 according to an embodiment of the present disclosure is shown;
[0060] Figure 7 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0061] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numbers in different drawings represent the same or similar elements. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.
[0062] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0063] The term "and / or", used herein only means an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0064] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, elements and circuits that are well known to those skilled in the art are not described in detail in order to highlight the main idea of the present disclosure.
[0065] For a sparse personnel scene, the positioning result output by the crowd positioning model has high accuracy, so using the number of positioning results as the result of crowd counting will be more accurate. For a dense crowd scene, there is serious occlusion between people, and the size of people changes dramatically from near to far, so it is difficult to detect each individual by detection. The accuracy of the result of the crowd positioning model is poor.
[0066] For a dense crowd scene, the accuracy of the crowd density distribution map predicted by the dense crowd counting model is high, so estimating the number of people based on the crowd density distribution map will be more accurate. For a sparse personnel scene, the dense crowd counting model is prone to misidentification, and the number of errors is large.
[0067] The embodiment of the present disclosure provides a crowd counting method, which can be applied to crowd counting in a sparse personnel scene and crowd counting in a dense crowd scene. In the embodiment of the present disclosure, based on head key point positioning and crowd density detection, a first crowd density distribution map and a second crowd density distribution map are obtained respectively, and crowd counting is performed by combining the first crowd density distribution map and the second crowd density distribution map, thereby improving the accuracy of crowd counting in various scenes.
[0068] Figure 1 A flowchart of a crowd counting method according to an embodiment of the present disclosure is shown. The crowd counting method can be performed by an electronic device such as a terminal device or a server, and the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The crowd positioning method can be implemented by calling computer readable instructions stored in a memory by a processor. Alternatively, the crowd positioning method can be performed by a server. As shown in the figure, the crowd counting method can include: Figure 1
[0069] In step S11, a crowd image is acquired.
[0070] The crowd image can be an image containing sparse personnel or an image containing dense crowds. It can be an image obtained by image acquisition equipment on personnel in a certain space range, an image frame obtained from a video, or obtained by other means, which is not limited in the present disclosure.
[0071] In step S12, based on head key point positioning on the crowd image, a first number of people corresponding to the crowd image and a first crowd density distribution map corresponding to the crowd image are obtained.
[0072] The first number of people can be used to represent the number of people obtained based on head key point positioning on the crowd image. The first crowd density distribution map can be used to represent the crowd density distribution map obtained based on head key point positioning on the crowd image.
[0073] In one possible implementation, step S12 can include: performing head key point positioning on the crowd image to obtain a target positioning map corresponding to the crowd image, wherein the target positioning map is used to indicate the position of a target head key point included in the crowd image; determining the first number of people based on the target positioning map; determining the first crowd density distribution map corresponding to the crowd image based on the target positioning map.
[0074] The head key point positioning on the crowd image is performed to obtain a target positioning map indicating the position of a target head key point included in the crowd image. The specific process of head key point positioning will be described in detail in the following possible implementation of the present disclosure, which will not be described here. The head key point can be the center point of the head or other preset key points of the head, which is not limited in the embodiments of the present disclosure.
[0075] Since each target head key point in the target positioning map represents a person, the number of target head key points indicated in the target positioning map can be determined as the first number of people. For example, if the target positioning map indicates the positions of 10 target head key points, the first number of people can be determined as 10; if the target positioning map indicates the positions of 100 target head key points, the first number of people can be determined as 100.
[0076] In a possible implementation, determining the first crowd density distribution map corresponding to the crowd image based on the target positioning map can include: using a Gaussian kernel to render each target head key point according to the position of each target head key point indicated in the target positioning map, to obtain the first crowd density distribution map.
[0077] Figure 2 An example of a target positioning map in an embodiment of the present disclosure is shown. It is assumed that the size of the crowd image is 10*10, and the size of the target positioning map corresponding to the crowd image is also 10*10. As shown in FIG. 3, the upper left corner is taken as the coordinate origin, and the positions of three target head key points are indicated in the target positioning map, as shown by the gray squares in FIG. 3. Figure 2 Figure 2 The coordinates of the three positions are (2, 2), (6, 3), and (4, 7) respectively.
[0078] Figure 3 An example of a first crowd density distribution map in an embodiment of the present disclosure is shown. It is assumed that the size of each head frame is 3 pixels*3 pixels, and the Gaussian kernel is used to render each target head key point as shown in FIG. 4. Figure 2 After the Gaussian kernel is used to render each target head key point, the value of each pixel in the head frame where each target head key point is located can be obtained, representing the probability that each pixel point is a person. As shown in FIG. 5, for any head frame, the sum of the values of the pixel points in the head frame is 1. Figure 3
[0079] The process of using the Gaussian kernel to render each target head key point to obtain the first crowd density distribution map can refer to related technologies, which will not be described here.
[0080] In step S13, a second crowd density distribution map corresponding to the crowd image is obtained based on the crowd density detection on the crowd image.
[0081] The second crowd density distribution map can be used to represent a crowd density distribution map obtained based on crowd density detection on the crowd image. In an example, the crowd density detection on the crowd image can be performed based on a direct regression manner. For example, the crowd image is input into a trained deep convolutional neural network, and a crowd density distribution map of the crowd image (i.e., the second crowd density distribution map) can be output. In another example, the crowd density detection on the crowd image can be performed based on a density map regression manner. Details of the crowd density detection on the crowd image based on the density map regression manner will be described below in combination with possible implementation manners of the present disclosure, and thus will not be described herein.
[0082] In step S14, based on the first number and a first preset number threshold, a target crowd density distribution map corresponding to the crowd image is selected from the first crowd density distribution map and the second crowd density distribution map.
[0083] The first preset number threshold can be used to determine whether the scene corresponding to the crowd image is a sparse personnel scene or a crowd dense scene. The first preset number threshold can be set as needed, for example, the first preset number threshold can be set to 50 or 100, etc. In a possible implementation manner, the first preset number threshold can be set according to the size of the spatial range corresponding to the crowd image. In the case that the spatial range corresponding to the crowd image is large, the first preset number threshold is large; in the case that the spatial range corresponding to the crowd image is small, the first preset number threshold is small.
[0084] The target crowd density distribution map can be used to represent a crowd density distribution map used for determining a crowd counting result of the crowd image subsequently. In the embodiment of the present disclosure, based on the size relationship between the first number and the first preset number threshold, it is determined whether the scene corresponding to the crowd image is a sparse personnel scene or a crowd dense scene, and then it is determined whether the first crowd density distribution map is determined as the target crowd density distribution map or the second crowd density distribution map is determined as the target crowd density distribution map.
[0085] In a possible implementation manner, step S14 can include: in the case that the first number is less than the first preset number threshold, the first crowd density distribution map is determined as the target crowd density distribution map.
[0086] When the first number is less than the first preset number threshold, it indicates that the scene corresponding to the crowd image is a sparse personnel scene. Considering that the accuracy of the head key point positioning result is high and the accuracy of the crowd density detection result is low in the sparse personnel scene, in this case, the first crowd density distribution map obtained based on the head key point positioning on the crowd image can be determined as the target crowd density distribution map, so as to improve the accuracy of the crowd counting result.
[0087] In a possible implementation, step S14 can include: determining the second crowd density distribution map as the target crowd density distribution map in a case where the first number is greater than or equal to the first preset number threshold.
[0088] When the first number is greater than or equal to the first preset number threshold, it indicates that the scene corresponding to the crowd image is a crowd dense scene. Considering that the accuracy of the head key point positioning result is low and the accuracy of the crowd density detection result is high in the crowd dense scene, the second crowd density distribution map obtained based on the crowd density detection on the crowd image can be determined as the target crowd density distribution map in this case, so as to improve the accuracy of the crowd counting result.
[0089] It should be noted that after step S12 is performed, the first number and the first preset number threshold can be compared first. In a case where the first number is less than the first preset number threshold, the first crowd density distribution map is directly determined as the target crowd density distribution map. In a case where the first number is greater than or equal to the first number, steps S13 and S14 are performed to determine the target crowd density distribution map.
[0090] In step S15, a crowd counting result of the crowd image is determined based on the target crowd density distribution map.
[0091] In the target crowd density distribution map, the value of each pixel point represents the probability that the pixel point is a person, and the sum of the values of the pixel points in each head frame is 1. Therefore, based on the target crowd density distribution map, the crowd counting result of the crowd image can be determined.
[0092] The counting result of the crowd image includes but is not limited to the total number of people in the crowd image, the number of people in a region of interest in the crowd image, and the like. The region of interest in the crowd image can be set as needed. In an example, a queuing area can be determined as the region of interest, or a dining area can be determined as the region of interest. In this way, the number of people in the queuing area or the number of people in the dining area can be obtained by combining the region of interest labeled in the crowd image and the number of people in the target crowd density distribution map. It can be understood that the queuing area and the dining area are only examples of the region of interest, and the region of interest can also be other regions. The region of interest in the crowd image can be one or more, for example, the crowd image can include a queuing area and a dining area at the same time, and the number of people in each region of interest in the crowd image can be determined respectively.
[0093] In a possible implementation, step S15 can include determining the total number of people in the crowd image based on the target crowd density distribution map. Determining the total number of people in the crowd image can include obtaining the total number of people in the crowd image based on weighting of the density values corresponding to the pixel points in the target crowd density distribution map. The weight used in weighting can be set as needed, and is not limited in the embodiments of the present disclosure.
[0094] In a possible implementation, step S15 can include determining the number of people in the region of interest based on the target crowd density distribution map and the region of interest in the crowd image. Determining the number of people in the region of interest includes obtaining the number of people in the region of interest based on weighting of the density values corresponding to the pixel points in the target crowd density distribution map that correspond to the region of interest. The weight used in weighting can be set as needed, and is not limited in the embodiments of the present disclosure.
[0095] Since the target crowd density distribution map of the crowd image has the same size as the crowd image, the pixels are one-to-one corresponding between the two, and the value of each pixel in the target crowd density distribution map represents the probability that the corresponding pixel in the crowd image is a person. Therefore, according to the position of the region of interest of the crowd image, the region of interest in the target crowd density distribution map can be determined, and the number of people in the region of interest can be determined.
[0096] In the embodiments of the present disclosure, based on head key point positioning and crowd density detection, a first crowd density distribution map and a second crowd density distribution map are obtained respectively, and crowd counting is performed in combination with the first crowd density distribution map and the second crowd density distribution map, thereby improving the accuracy of crowd counting in various scenarios.
[0097] In a possible implementation, the crowd counting method can further include: in response to the region of interest being preset in the crowd image by calibration or the like, determining a second number of people corresponding to the region of interest in the crowd image based on head key point positioning performed on the crowd image; based on the second number of people and a second preset number threshold, selecting the target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map; and step S15 of determining the crowd counting result of the region of interest in the crowd image based on the target crowd density distribution map can further include: determining the number of people in the region of interest based on the target crowd density distribution map and the region of interest in the crowd image.
[0098] The user can calibrate a region of interest in the crowd image as needed, for example, a queuing area or a dining area can be calibrated as a region of interest. The user can calibrate one or more regions of interest in the crowd image, and the number of people in each region of interest can be determined in the embodiments of the present application.
[0099] Based on the head key point positioning of the crowd image, the position of each head key point in the crowd image can be obtained. In combination with the position of the region of interest in the crowd image, it can be determined which head key points are located within the region of interest and which head key points are located outside the region of interest, and then the number of people in the region of interest in the crowd image, i.e., the second number of people mentioned above, can be obtained.
[0100] The second preset number threshold can be used to determine whether the scene corresponding to the region of interest in the crowd image is a sparse personnel scene or a crowd dense scene, for example, whether the scene corresponding to the queuing area is a sparse personnel scene or a crowd dense scene, or whether the scene corresponding to the dining area is a sparse personnel scene or a crowd dense scene. In the case where there are multiple regions of interest in the crowd image, the second number of people corresponding to each region of interest in the crowd image can be determined respectively, and the second number of people corresponding to each region of interest can be compared with the second preset number threshold to determine whether the scene corresponding to each region of interest is a sparse personnel scene or a crowd dense scene.
[0101] The second preset number threshold can be set as needed, for example, the second preset number threshold can be set to 30 or 50, etc. In a possible implementation manner, the second preset number threshold can be set according to the size of the spatial range corresponding to the region of interest. In the case where the spatial range corresponding to the region of interest is large, the second preset number threshold is large; in the case where the spatial range corresponding to the region of interest is small, the second preset number threshold is small. It can be understood that the second preset number threshold is smaller than the first preset number threshold.
[0102] In the case where the second number of people is less than the second preset number threshold, it indicates that the scene corresponding to the region of interest in the crowd image is a sparse personnel scene, and therefore the first crowd density distribution map can be determined as the target crowd density distribution map corresponding to the crowd image. In the case where the second number of people is greater than or equal to the second preset number threshold, it indicates that the scene corresponding to the region of interest in the crowd image is a crowd dense scene, and therefore the second crowd density distribution map can be determined as the target crowd density distribution map corresponding to the crowd image. It can be understood that because the second number of people corresponding to different regions of interest is different, the scene corresponding to some regions of interest can be a sparse personnel scene and the scene corresponding to some regions of interest can be a crowd dense scene, and therefore the target crowd density distribution map corresponding to the crowd image selected can be different when different regions of interest are processed.
[0103] Since the target crowd density distribution map is selected for the region of interest, when determining the crowd counting result of the region of interest in the crowd image based on the target crowd density distribution map, the number of people in the region of interest is actually determined. In a possible implementation, the position of the region of interest in the target crowd density distribution map can be determined according to the position of the region of interest in the crowd image, the pixel points belonging to the region of interest in the target crowd density distribution map are found out, the density values corresponding to the pixel points belonging to the region of interest in the target crowd density distribution map are weighted, and the number of people in the region of interest can be obtained. The weight used when weighting can be set according to requirements, and the embodiments of the present application do not limit this.
[0104] It should be noted that in the case where the second number is less than the second preset number threshold, the second number can be directly determined as the number of people in the region of interest, and the steps of selecting the target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map based on the second number and the second preset number threshold, and determining the number of people in the region of interest based on the target crowd density distribution map and the region of interest in the crowd image can be omitted.
[0105] In the embodiments of the present application, the number of people in the region of interest can be accurately determined.
[0106] The specific process of head key point positioning is described in detail below.
[0107] In a possible implementation, the head key point positioning of the crowd image to obtain the target positioning map corresponding to the crowd image can include: performing head key point positioning on the crowd image to determine a predicted positioning map corresponding to the crowd image, wherein the predicted positioning map is used to indicate the predicted confidence of each pixel point in the crowd image being a head key point; performing image processing on the predicted positioning map based on a preset confidence threshold to obtain an initial positioning map, wherein the initial positioning map is used to indicate the position of an initial head key point included in the crowd image; determining a target neighborhood corresponding to each initial head key point in the initial positioning map; performing filtering processing on the target neighborhood corresponding to each initial head key point based on the predicted positioning map to obtain the target positioning map corresponding to the crowd image.
[0108] The head key point positioning of the crowd image is performed, the prediction confidence of each pixel point in the crowd image being a head key point is determined in an end-to-end manner, and then the prediction positioning map is threshold segmented by using a preset confidence threshold to determine an initial positioning map used for indicating the positions of the initial head key points included in the crowd image, and then a target neighborhood of each initial head key point in the initial positioning map is determined, so that the target neighborhood corresponding to the initial head key point is further filtered based on the prediction positioning map to obtain a target positioning map with high precision, which accurately indicates the positions of the target head key points included in the crowd image.
[0109] In an example, the trained head key point positioning neural network can be used to perform head key point positioning on the crowd image. Specifically, the crowd image is input into the trained head key point positioning neural network, and the prediction positioning map is directly output after positioning by the head key point positioning neural network. The specific network structure of the trained head key point positioning neural network and the training process can adopt the network structure and the training process in related technologies, and the present disclosure does not make specific limitations thereto.
[0110] In an example, the pixel value of each pixel point in the prediction positioning map represents the prediction confidence of the pixel point, i.e., the probability of the pixel point being a head key point. The sigmoid operation is performed on the prediction positioning map so that the pixel value of each pixel point in the prediction positioning map is between 0 and 1. For example, the pixel value of a certain pixel point in the prediction positioning map is 0.7, which means that the probability of the pixel point being a head key point is 0.7.
[0111] Since the prediction positioning map is only used to indicate the prediction confidence of each pixel point in the crowd image being a head key point, the initial positioning map used for indicating the positions of the initial head key points included in the crowd image can be effectively obtained by threshold segmenting the prediction positioning map by using the preset confidence threshold. The specific value of the preset confidence threshold can be flexibly set according to actual conditions, and the present disclosure does not make specific limitations thereto.
[0112] The pixel value of each pixel point in the prediction positioning map is compared with the preset confidence threshold. In the case that the pixel value of a certain pixel point in the prediction positioning map is greater than or equal to the preset confidence threshold, the pixel value of the pixel point at the corresponding position in the initial positioning map is determined as 1; in the case that the pixel value of a certain pixel point in the prediction positioning map is less than the preset confidence threshold, the pixel value of the pixel point at the corresponding position in the initial positioning map is determined as 0.
[0113] The initial positioning map and the crowd image have the same size, and the position of a pixel point with a pixel value of 1 in the initial positioning map is used to indicate the position of an initial head key point included in the crowd image. For example, in a case where the pixel value of a pixel point with an image coordinate of (x, y) in the initial positioning map is 1, it can be determined that the pixel point with the image coordinate of (x, y) in the crowd image is an initial head key point; in a case where the pixel value of the pixel point with the image coordinate of (x, y) in the initial positioning map is 0, it can be determined that the pixel point with the image coordinate of (x, y) in the crowd image is a part other than the initial head key point.
[0114] In order to avoid the problem of false detection of multiple initial head key points corresponding to the head of the same human body, the target neighborhood corresponding to each initial head key point in the initial positioning map is further determined, and the target neighborhood corresponding to each initial head key point is filtered to obtain a target positioning map with higher precision, in which the head of the same human head corresponds to one target head key point.
[0115] In a possible implementation, determining the target neighborhood corresponding to each initial head key point in the initial positioning map includes: determining the target neighborhood corresponding to each initial head key point according to a preset neighborhood radius.
[0116] In an example, the preset neighborhood radius can be fixed, and in this case, the preset neighborhood radius can be referred to as a first preset neighborhood radius. In another example, the preset neighborhood radius can be determined according to the position of the initial head key point in the crowd image and a preset perspective relationship corresponding to the crowd image, and in this case, according to the size of the head box height, a second preset neighborhood radius or a third preset neighborhood radius can be selected to determine the target neighborhood. The process of determining the target neighborhood corresponding to each initial head key point according to the first preset neighborhood radius and the process of determining the target neighborhood corresponding to each initial head key point based on the position of the initial head key point in the crowd image and the preset perspective mapping relationship corresponding to the crowd image (corresponding to the second preset neighborhood radius and the third preset neighborhood radius) are described below.
[0117] In a possible implementation, determining the target neighborhood corresponding to each initial head key point in the initial positioning map includes: determining the target neighborhood corresponding to each initial head key point according to a first preset neighborhood radius.
[0118] By using the pre-set fixed first preset neighborhood radius, the target neighborhood corresponding to each initial head key point can be quickly determined. The specific value of the first neighborhood radius can be flexibly set according to actual conditions, and the present disclosure does not make a specific limitation in this regard.
[0119] For example, the first preset neighborhood radius is 2, and for any initial head key point i, the target neighborhood corresponding to the initial head key point i includes pixel points with a pixel distance of no more than 2 pixels from the initial head key point i.
[0120] In a possible implementation, the target neighborhood corresponding to each initial head key point in the initial positioning map is determined, including: for any initial head key point, based on the position of the initial head key point in the crowd image and the preset perspective mapping relationship corresponding to the crowd image, the target neighborhood corresponding to the initial head key point is determined.
[0121] The preset perspective mapping relationship is used to indicate the image scale corresponding to different positions in the crowd image. Because the installation angles of different image acquisition devices are different, the image scales corresponding to the crowd images acquired by different image acquisition devices are different. In the embodiment of the present disclosure, it is necessary to determine the preset perspective mapping relationship corresponding to the crowd image acquired by each image acquisition device. The process of determining the preset perspective mapping relationship corresponding to the crowd image is described below.
[0122] In a possible implementation, a plurality of labeled human body boxes of pedestrians in different positions in the crowd image can be obtained by labeling the human body boxes of the pedestrians in different positions in the crowd image; and based on the plurality of labeled human body boxes, the preset perspective mapping relationship corresponding to the crowd image is determined.
[0123] By labeling the human body boxes of the pedestrians in different positions (far, middle, and near) in the crowd image, a plurality of labeled human body boxes in the crowd image can be obtained. Based on the proportional relationship between the height of the labeled human body box and the actual height of the pedestrian, the image scale corresponding to the limited position (the position of the labeled human body box) in the crowd image can be determined. Further, based on the image scale corresponding to the limited position, the image scale corresponding to each position in the crowd image is effectively fitted, that is, the preset perspective mapping relationship corresponding to the crowd image is obtained.
[0124] Figure 4 A schematic diagram of a crowd image and its corresponding preset perspective mapping relationship according to an embodiment of the present disclosure is shown. As shown in Figure 2 Four labeled human body boxes A, B, C, and D in different positions in the crowd image are obtained by labeling the human body boxes of the pedestrians in different positions (far, middle, and near) in the crowd image. Further, based on the four labeled human body boxes A, B, C, and D, the preset perspective mapping relationship corresponding to the crowd image is effectively fitted.
[0125] In a possible implementation, the preset perspective mapping relationship corresponding to the crowd image is determined based on the plurality of labeled human body boxes, including: for any one labeled human body box, determining a reference image scale corresponding to a reference human body key point in the labeled human body box; and fitting the preset perspective mapping relationship corresponding to the crowd image according to the third image coordinates of the reference human body key point in each labeled human body box and the reference image scale corresponding to the reference human body key point in each labeled human body box.
[0126] The image scales at different positions are linearly changed along the column direction of the crowd image, and therefore, after the image scales corresponding to the reference human body key points at limited positions in the crowd image are determined according to the labeled human body boxes, the image scale corresponding to each position in the crowd image can be effectively obtained by linear function fitting, that is, the preset perspective mapping relationship corresponding to the crowd image is obtained.
[0127] Since the pedestrian is vertically standing, the height of the labeled human body box can be regarded as the height of the pedestrian in the crowd image by taking the human foot key point as the reference human body key point. The height of the labeled human body box can be represented by the number of pixel rows occupied by the labeled human body box. For example, if the labeled human body box occupies 17 pixel rows in the crowd image, the height of the labeled human body box is 17. Assuming that the real height of the pedestrian corresponding to the labeled human body box is 1.7 meters, the position of the reference human foot key point in the labeled human body box can be determined, indicating that 1.7 meters in the real world requires 17 pixel rows. Assuming that the unit height is 1 m, therefore, the position of the reference human foot key point in the labeled human body box indicates that 1 meter in the real world requires 10 pixel rows, that is, the reference image scale corresponding to the reference human foot key point in the labeled human body box is 10. The real height of the pedestrian corresponding to the labeled human body box can be appropriately selected according to actual conditions, and the disclosure does not make a specific limitation in this regard.
[0128] The reference human foot key point in the labeled human body box can be the midpoint of the bottom edge of the labeled human body box, and can also be other pixel points in the labeled human body box, and the disclosure does not make a specific limitation in this regard.
[0129] Still taking the above example, after four labeled human body boxes A, B, C and D are labeled in the crowd image, the reference image scale corresponding to the reference human foot key point in each labeled human body box is determined in the above manner. Further, linear function fitting is performed according to the third image coordinates of the reference human foot key points in the four labeled human body boxes and the reference image scales corresponding thereto, to obtain a linear mapping function p=a*y+b. Figure 4
[0130] The image coordinates refer to position coordinates in a pixel coordinate system of the crowd image. For example, a pixel coordinate system of the crowd image is constructed with the upper left corner of the crowd image as the coordinate origin (0, 0), the direction parallel to the row direction of the image as the direction of the x-axis, and the direction parallel to the column direction of the image as the direction of the y-axis. The horizontal and vertical coordinates of the image coordinates are both in units of pixels. For example, if the image coordinates of the reference human foot key point are (10, 15), it indicates that the reference human foot key point is a pixel point located at the 10th row and the 15th column in the crowd image.
[0131] The linear mapping function p = a * y + b is a function representation of the preset perspective mapping relationship corresponding to the crowd image. Wherein, a and b are parameters obtained by linear function fitting, y is the vertical coordinate of the image coordinates of different positions in the crowd image, and p is the image scale corresponding to the position. By using the linear mapping function p = a * y + b, the image scale corresponding to each position in the crowd image can be determined.
[0132] Before or after determining the preset perspective mapping relationship corresponding to the crowd image, the human head key point positioning is performed on the crowd image to obtain an initial positioning map corresponding to the crowd image.
[0133] After determining the positions of the initial human head key points included in the crowd image based on the initial positioning map, the target neighborhood matched with each initial human head key point can be determined based on the preset perspective mapping relationship.
[0134] In one possible implementation, based on the position of the initial human head key point in the crowd image and the preset perspective mapping relationship, the target neighborhood corresponding to the initial human head key point is determined, including: based on the preset perspective mapping relationship, determining the target image scale corresponding to the position of the initial human head key point in the crowd image; based on the target image scale, determining the head frame height corresponding to the initial human head key point; based on the head frame height corresponding to the initial human head key point, determining the target neighborhood corresponding to the initial human head key point.
[0135] Based on the preset perspective mapping relationship, the target image scale corresponding to the position of the initial human head key point in the crowd image can be quickly determined, and then the head frame height corresponding to the initial human head key point is determined based on the target image scale, so that the target neighborhood matched with the initial human head key point can be further determined according to the head frame height.
[0136] For example, for an initial head key point i in the initial positioning map, the image coordinates of the initial head key point i in the crowd image are (hx, hy), and according to the preset perspective mapping relationship (linear mapping function p = a * y + b) corresponding to the crowd image, the target scale corresponding to the initial head key point i can be determined as pi = a * hy + b. Assuming that the real head frame height corresponding to the pedestrian in the crowd image is 0.4 meters * 0.4 meters, the head frame height corresponding to the initial head key point i in the crowd image is si = 0.4 * pi. According to the head frame height corresponding to the initial head key point i, that is, si = 0.4 * pi, the target neighborhood matching the size is determined.
[0137] In a possible implementation, based on the head frame height corresponding to the initial head key point, the target neighborhood corresponding to the initial head key point is determined, including: in a case where the head frame height is greater than a preset head frame height threshold, determining the target neighborhood corresponding to the initial head key point based on a second preset neighborhood radius; or, in a case where the head frame height is less than or equal to the preset head frame height threshold, determining the target neighborhood corresponding to the initial head key point based on a third preset neighborhood radius, wherein the second preset neighborhood radius is greater than the third preset neighborhood radius.
[0138] In a case where the head frame height is greater than the preset head frame height threshold, it can be determined that the head frame is relatively large in size, and therefore a larger second preset neighborhood radius is used for subsequent filtering processing. In a case where the head frame height is less than or equal to the preset head frame height threshold, it can be determined that the head frame is relatively small in size, and therefore a smaller third preset neighborhood radius is used for subsequent filtering processing. By flexibly determining the neighborhood radius, the accuracy of the filtering operation can be improved. The specific values of the preset head frame height threshold, the second preset neighborhood radius, and the third preset neighborhood radius can be flexibly set according to actual conditions, and the present disclosure does not make specific limitations thereto.
[0139] In an example, the head frame height threshold is 32, and for an initial head key point i, in a case where the head frame height si corresponding to the initial head key point i is greater than 32, the target neighborhood of the initial head key point i is determined based on a second preset neighborhood radius 2; and in a case where the head frame height si corresponding to the initial head key point i is less than or equal to 32, the target neighborhood of the initial head key point i is determined based on a third preset neighborhood radius 1.
[0140] In a case where the second preset neighborhood radius is 2, the target neighborhood corresponding to the initial head key point i includes pixel points with a pixel distance of no more than 2 pixel points from the initial head key point i. In a case where the third preset neighborhood radius is 1, the target neighborhood corresponding to the initial head key point i includes pixel points with a pixel distance of no more than 1 pixel point from the initial head key point i.
[0141] After determining the target neighborhood corresponding to each initial head key point in the initial positioning map based on the above manner, the initial positioning map is filtered by using the target neighborhood corresponding to each initial head key point to obtain a target positioning map with higher accuracy.
[0142] In a possible implementation, the target neighborhood corresponding to each initial head key point is filtered based on the predicted positioning map to obtain the target positioning map, including: determining, for any initial head key point i, whether there is at least one other initial head key point in the target neighborhood corresponding to the initial head key point i; in the case that there is at least one other initial head key point j in the target neighborhood corresponding to the initial head key point i, determining, based on the predicted positioning map, a predicted confidence of the initial head key point i and a predicted confidence of the at least one other initial head key point j; and determining, based on the initial head key point with the maximum predicted confidence from the initial head key point i and the at least one other initial head key point j, a target head key point in the target neighborhood corresponding to the initial head key point i.
[0143] For an initial head key point i with an image coordinate (xi, yi) in the initial positioning map, it is detected whether there is another initial head key point in the target neighborhood of the initial head key point i. If there is another initial head key point j with an image coordinate (xj, yj), then according to the predicted positioning map, a predicted confidence of the initial head key point i and a predicted confidence of the initial head key point j are determined. In the case that the predicted confidence of the initial head key point i is greater than the predicted confidence of the initial head key point j, the pixel value of the pixel point with the image coordinate (xi, yi) is kept as 1, and the pixel value of the pixel point with the image coordinate (xj, yj) is updated as 0, that is, the initial head key point j in the initial positioning map is filtered out. In this way, each initial head key point in the initial positioning map is traversed to obtain the final target positioning map.
[0144] The specific process of the crowd density detection is described below.
[0145] In a possible implementation, the method can further include: performing structure similarity perception on the crowd image to obtain a second crowd density distribution map corresponding to the crowd image. Specifically, the crowd image can be input into a trained crowd statistical model to obtain the second crowd density distribution map corresponding to the crowd image. The crowd statistical model is obtained by training a neural network based on a preset loss function. In an optional embodiment, the neural network can be a convolutional neural network (CNN), such as a multi-column convolutional network (MCNN) or a deep dilated convolutional network (CSRNet).
[0146] In a possible implementation, the method further includes: acquiring a training image and a corresponding crowd distribution density image of the training image respectively, wherein the crowd density distribution image includes a real density map and a predicted density map; obtaining a head block mean square error loss function, a structural similarity loss function and a background square difference loss function respectively according to the training image and the crowd distribution density image; and obtaining a preset loss function according to the head block mean square error loss function, the structural similarity loss function and the background square difference loss function.
[0147] In one example, a preset region of a crowd image can be acquired, wherein the preset region includes a head block; a real density map and a predicted density map corresponding to each preset region are determined; a head block mean square error loss function, a structural similarity loss function and a background square difference loss function are determined according to the real density map and the predicted density map corresponding to each preset region, and a preset loss function can be obtained.
[0148] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without deviating from the principle logic. Limited by the length, the present disclosure will not be repeated. Those skilled in the art can understand that the specific execution order of each step in the above-mentioned method should be determined according to its function and possible internal logic.
[0149] In addition, the present disclosure also provides a crowd counting device, an electronic device, a computer readable storage medium and a program, all of which can be used to implement any one of the crowd counting methods provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding description in the method part, and will not be repeated.
[0150] Figure 5 A block diagram of a crowd counting device according to an embodiment of the present disclosure is shown. As shown in Figure 5 The device 50 includes:
[0151] An acquisition module 51 is configured to acquire a crowd image.
[0152] A positioning module 52 is configured to obtain a first number of people corresponding to the crowd image and a first crowd density distribution map corresponding to the crowd image based on head key point positioning performed on the crowd image.
[0153] A detection module 53 is configured to obtain a second crowd density distribution map corresponding to the crowd image based on crowd density detection performed on the crowd image.
[0154] A first selection module 54 is configured to select a target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map based on the first number of people and a first preset number of people threshold.
[0155] The first determining module 55 is configured to determine a crowd counting result of the crowd image based on the target crowd density distribution map.
[0156] In a possible implementation, the first selecting module is further configured to:
[0157] In a case where the first number of people is less than the first preset number threshold, the first crowd density distribution map is determined as the target crowd density distribution map; or
[0158] In a case where the first number of people is greater than or equal to the first preset number threshold, the second crowd density distribution map is determined as the target crowd density distribution map.
[0159] In a possible implementation, the first determining module is further configured to:
[0160] determine a total number of people in the crowd image based on the target crowd density distribution map; and / or
[0161] determine a number of people in a region of interest in the crowd image based on the target crowd density distribution map and the region of interest.
[0162] In a possible implementation,
[0163] The determination of the total number of people in the crowd image includes:
[0164] obtaining the total number of people in the crowd image based on weighting of density values corresponding to respective pixels in the target crowd density distribution map; and / or
[0165] The determination of the number of people in the region of interest includes:
[0166] obtaining the number of people in the region of interest based on weighting of density values corresponding to respective pixels in the target crowd density distribution map and corresponding to the region of interest.
[0167] In a possible implementation, the apparatus further includes:
[0168] The second determining module is configured to, in response to the region of interest being marked in the crowd image, determine a second number of people corresponding to the region of interest in the crowd image based on head key point positioning performed on the crowd image.
[0169] The second selecting module is configured to select, based on the second number of people and a second preset number threshold, a target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map.
[0170] The first determining module is further configured to determine a number of people in a region of interest in the crowd image based on the target crowd density distribution map and the region of interest.
[0171] In a possible implementation, the positioning module is further configured to:
[0172] perform head key point positioning on the crowd image to obtain a target positioning map corresponding to the crowd image, wherein the target positioning map is used to indicate positions of target head key points included in the crowd image;
[0173] determine the first number of people based on the target positioning map.
[0174] determine a first crowd density distribution map corresponding to the crowd image based on the target positioning map.
[0175] In a possible implementation, the head key point positioning on the crowd image to obtain the target positioning map corresponding to the crowd image comprises:
[0176] perform head key point positioning on the crowd image to obtain a predicted positioning map corresponding to the crowd image, wherein the predicted positioning map is used to indicate a predicted confidence of each pixel point in the crowd image being a head key point;
[0177] perform image processing on the predicted positioning map based on a preset confidence threshold to obtain an initial positioning map, wherein the initial positioning map is used to indicate positions of initial head key points included in the crowd image;
[0178] determine a target neighborhood corresponding to each initial head key point in the initial positioning map;
[0179] perform filtering processing on the target neighborhood corresponding to each initial head key point based on the predicted positioning map to obtain the target positioning map corresponding to the crowd image.
[0180] In a possible implementation, the determination of the target neighborhood corresponding to each initial head key point in the initial positioning map comprises:
[0181] determine the target neighborhood corresponding to each initial head key point according to a preset neighborhood radius, wherein the preset neighborhood radius is determined based on a position of the initial head key point in the crowd image and a preset perspective relationship corresponding to the crowd image, and the preset perspective mapping relationship corresponding to the crowd image is used to indicate image scales corresponding to different positions in the crowd image.
[0182] In a possible implementation, the filtering processing on the target neighborhood corresponding to each of the initial head key points based on the predicted positioning map to obtain the target positioning map comprises:
[0183] For any one of the initial head key points, it is determined whether there is at least one other initial head key point in the target neighborhood corresponding to the initial head key point i;
[0184] In the case that there is at least one other initial head key point j in the target neighborhood corresponding to the initial head key point i, it is determined, based on the predicted positioning map, a predicted confidence corresponding to the initial head key point i and a predicted confidence corresponding to the at least one other initial head key point j;
[0185] Based on the initial head key point i and the initial head key point with the maximum predicted confidence in the at least one other initial head key point j, a target head key point in the target neighborhood corresponding to the initial head key point i is determined.
[0186] In a possible implementation, the obtaining of the first crowd density distribution map corresponding to the crowd image based on the target positioning map comprises:
[0187] According to the position of each target head key point indicated by the target positioning map, a Gaussian kernel is used to render each target head key point to obtain the first crowd density distribution map.
[0188] The method has specific technical association with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing data storage amount, reducing data transmission amount, improving hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system in accordance with the natural law.
[0189] In the embodiments of the present disclosure, based on the head key point positioning and the crowd density detection, the first crowd density distribution map and the second crowd density distribution map are obtained respectively, and the crowd counting is performed by combining the first crowd density distribution map and the second crowd density distribution map, so that the accuracy of the crowd counting is improved in various scenes.
[0190] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or contains modules which can be used to execute the methods described in the above method embodiments, and the specific implementation can be referred to the description of the above method embodiments. For brevity, it will not be described here.
[0191] The embodiments of the present disclosure further provide a computer readable storage medium having stored thereon computer program instructions, which when executed by a processor, implement the method described above. The computer readable storage medium can be a volatile or non-volatile computer readable storage medium.
[0192] The embodiments of the present disclosure further provide an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored by the memory to execute the method described above.
[0193] The embodiments of the present disclosure further provide a computer program product, comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in a processor of an electronic device, the processor in the electronic device executes the method described above.
[0194] The electronic device can be provided as a terminal, a server or other forms of devices.
[0195] Figure 6 A block diagram of an electronic device 800 according to an embodiment of the present disclosure is shown. The electronic device 800 can be, for example, a terminal device such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.
[0196] Referring to Figure 6 The electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0197] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with displaying, making phone calls, data communications, camera operations and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the methods described above. In addition, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0198] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or nonvolatile memory, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disc, or optical disc.
[0199] The power supply component 806 supplies power for various components of the electronic device 800. The power supply component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0200] The multimedia component 808 includes a screen providing an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a back camera. The front camera and / or the back camera can receive external multimedia data when the electronic device 800 is in an operation mode, such as a photographing mode or a video mode. Each of the front camera and the back camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0201] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.
[0202] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, etc. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0203] The sensor component 814 includes one or more sensors for providing status assessments for various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration / g-force and temperature of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 814 can also include a light sensor, such as a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensor, utilized in an imaging application. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0204] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as a Wireless Fidelity (Wi-Fi), a second-generation (2G) mobile communication technology, a third-generation (3G) mobile communication technology, a fourth-generation (4G) mobile communication technology, a long-term evolution (LTE) of a Universal Mobile Telecommunication System (UMTS), a fifth-generation (5G) mobile communication technology, or a combination thereof. In an example embodiment, the communication component 816 receives a broadcast signal or broadcast related information from an external broadcasting management system through a broadcast channel. In an example embodiment, the communication component 816 can further include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technology.
[0205] In an example embodiment, the electronic device 800 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements to perform the above-described methods.
[0206] In an example embodiment, a non-transitory computer-readable storage medium, such as the memory 804 including computer program instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to perform the above-described methods.
[0207] Figure 7 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to FIG. 19, the electronic device 1900 includes a processing component 1922, further including one or more processors, and a memory resource represented by a memory 1932, for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method. Figure 7
[0208] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Microsoft Windows Server operating system (Windows Server TM ), Apple's graphical user interface-based operating system (Mac OSX TM ), multi-user multi-process computer operating system (Unix TM ), free and open source Unix-like operating system (Linux TM ), open source Unix-like operating system (FreeBSD TM ) or the like.
[0209] In an exemplary embodiment, a non-transitory computer readable storage medium, such as the memory 1932 including computer program instructions executable by the processing component 1922 of the electronic device 1900 to perform the above-described method is also provided.
[0210] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0211] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0212] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0213] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0214] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0215] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0216] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0217] The flow diagrams and the block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0218] The computer program product can be embodied in a tangible medium of
[0219] The above description of the various embodiments is intended to be illustrative in all aspects, rather than being restrictive. Those skilled in the art can refer to the description of the various embodiments to make modifications and / or improvements.
[0220] Those skilled in the art can understand that, in the above-described method of the specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible inherent logic.
[0221] Having described above several embodiments of the disclosure, any modifications and variations that fall within the scope of the described embodiments are also intended to be within the scope of the disclosure. As will be apparent to those skilled in the art, some modifications and variations to the embodiments described above can be practiced while staying within the scope and spirit of the described embodiments. The foregoing description of the described embodiments has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the described embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. It is intended that the disclosed embodiments be limited only by the claims.
Claims
1. A method of crowd counting, characterized by, The method comprises: obtaining a crowd image; based on the head key point positioning of the crowd image, obtaining the first number corresponding to the crowd image and the first crowd density distribution map corresponding to the crowd image; based on the crowd density detection of the crowd image, obtaining the second crowd density distribution map corresponding to the crowd image; based on the first number and the first preset number threshold, selecting the target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map; based on the target crowd density distribution map, determining the crowd counting result of the crowd image; based on the head key point positioning of the crowd image, obtaining the first number corresponding to the crowd image and the first crowd density distribution map corresponding to the crowd image, comprising: determining the target neighborhood corresponding to each initial head key point according to the preset neighborhood radius, wherein the preset neighborhood radius is determined based on the position of the initial head key point in the crowd image and the preset perspective relationship corresponding to the crowd image, and the preset perspective mapping relationship corresponding to the crowd image is used to indicate the image scale corresponding to different positions in the crowd image.
2. The method of claim 1, wherein, based on the first number and the first preset number threshold, selecting the target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map, comprising: in the case that the first number is less than the first preset number threshold, determining the first crowd density distribution map as the target crowd density distribution map; or in the case that the first number is greater than or equal to the first preset number threshold, determining the second crowd density distribution map as the target crowd density distribution map.
3. The method according to claim 1 or 2, characterized in that, based on the target crowd density distribution map, determining the crowd counting result of the crowd image, comprising: based on the target crowd density distribution map, determining the total number of people in the crowd image; and / or based on the target crowd density distribution map and the region of interest in the crowd image, determining the number of people in the region of interest.
4. The method of claim 3, wherein: determining the total number of people in the crowd image comprises: based on the weighting of the density values of the pixel points in the target crowd density distribution map, obtaining the total number of people in the crowd image; and / or determining the number of people in the region of interest comprises: based on the weighting of the density values of the pixel points corresponding to the region of interest in the target crowd density distribution map, obtaining the number of people in the region of interest.
5. The method of claim 1, wherein, The method further comprises: in response to the presence of the region of interest in the crowd image, based on the head key point positioning of the crowd image, determining the second number corresponding to the region of interest in the crowd image; based on the second number and the second preset number threshold, selecting the target crowd density distribution map corresponding to the crowd image from the first crowd density distribution map and the second crowd density distribution map; The people counting result of the region of interest of the crowd image is determined based on the target crowd density distribution map, and the people counting result of the region of interest of the crowd image includes: Based on the target crowd density distribution map and the region of interest in the crowd image, the number of people in the region of interest is determined.
6. The method according to any one of claims 1 or 2, characterized in that, The first number of people corresponding to the crowd image and the first crowd density distribution map corresponding to the crowd image are obtained based on the head key point positioning of the crowd image, and the head key point positioning of the crowd image includes: The target positioning map corresponding to the crowd image is obtained by performing head key point positioning on the crowd image, wherein the target positioning map is used to indicate the position of the target head key point included in the crowd image; The first number of people is determined based on the target positioning map; The first crowd density distribution map corresponding to the crowd image is determined based on the target positioning map.
7. The method of claim 6, wherein, The target positioning map corresponding to the crowd image is obtained by performing head key point positioning on the crowd image, and the head key point positioning on the crowd image includes: The prediction positioning map corresponding to the crowd image is obtained by performing head key point positioning on the crowd image, wherein the prediction positioning map is used to indicate the prediction confidence of each pixel point being a head key point in the crowd image; The initial positioning map is obtained by performing image processing on the prediction positioning map based on a preset confidence threshold, wherein the initial positioning map is used to indicate the position of the initial head key point included in the crowd image; The target neighborhood corresponding to each initial head key point in the initial positioning map is determined; The target positioning map corresponding to the crowd image is obtained by filtering the target neighborhood corresponding to each initial head key point based on the prediction positioning map.
8. The method of claim 7, wherein, The target positioning map is obtained by filtering the target neighborhood corresponding to each initial head key point based on the prediction positioning map, and the filtering includes: For any one of the initial head key points, it is determined whether there is at least one other initial head key point in the target neighborhood corresponding to the initial head key point i; In the case that there is at least one other initial head key point j in the target neighborhood corresponding to the initial head key point i, the prediction confidence of the initial head key point i and the prediction confidence of the at least one other initial head key point j are determined based on the prediction positioning map; The target head key point in the target neighborhood corresponding to the initial head key point i is determined based on the initial head key point with the maximum prediction confidence among the initial head key point i and the at least one other initial head key point j.
9. The method of claim 7, wherein, The first crowd density distribution map corresponding to the crowd image is obtained based on the target positioning map, and the first crowd density distribution map is obtained by rendering each target head key point using a Gaussian kernel according to the position of each target head key point indicated by the target positioning map. The device includes:
10. A people counting device, characterized by An acquisition module is configured to acquire a crowd image; A positioning module is configured to obtain a first number of people corresponding to the crowd image and a first crowd density distribution map corresponding to the crowd image based on head key point positioning of the crowd image. The detection module is configured to obtain a second crowd density distribution corresponding to the crowd image based on crowd density detection on the crowd image. The first selection module is configured to select a target crowd density distribution corresponding to the crowd image from the first crowd density distribution and the second crowd density distribution based on the first number of people and a first preset number of people threshold. The first determination module is configured to determine a crowd counting result of the crowd image based on the target crowd density distribution. The positioning module is further configured to: determine a target neighborhood corresponding to each initial head key point according to a preset neighborhood radius, wherein the preset neighborhood radius is determined based on a position of the initial head key point in the crowd image and a preset perspective relationship corresponding to the crowd image, and the preset perspective mapping relationship corresponding to the crowd image is used to indicate image scales corresponding to different positions in the crowd image.
11. An electronic device, comprising: comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the method of any one of claims 1 to 9.
12. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by the processor, implement the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Crowd density statistical method and device, electronic equipment and storage medium
CN109858424A
Object counting method and device, electronic equipment and storage medium
CN111652107A