Crowd situation acquisition method, device, storage medium and system
By determining the blind spot and non-blind spot ranges in the video acquisition area, using space-time perception information and simulation deduction, the blind spot population situation is indirectly acquired, and the problems of increased costs and unsatisfactory results caused by hardware addition in the existing technology are solved, and the perception of the population situation in the whole region is realized.
Patent Information
- Application Number
- CN202111357329.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-11-16
AI Technical Summary
In the prior art, the method of reducing the blind spots of video acquisition by adding video acquisition hardware leads to increased costs but the results are not ideal, and the whole-domain perception of crowd situation cannot be achieved.
By determining the blind spot and non-blind spot ranges in the target video acquisition area, the space-time perception information of the blind spot range is obtained, and simulation deduction combined with the crowd situation in the non-blind spot range can indirectly perceive the crowd situation in the blind spot range, and the crowd situation in the whole region can be obtained.
The entire domain perception of crowd situations can be achieved without adding hardware, reducing costs and improving perception effects.
Smart Images

Figure CN114283373B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence simulation, and specifically to a method, device, storage medium and system for acquiring crowd situation. Background Art
[0002] For precise control of large crowds, global video capture and perception prediction of crowd movements are crucial. However, in practice, crowds are often obscured by objects such as trees and buildings, resulting in blind spots in the field of view of video capture and perception equipment, making global perception of crowd movements difficult. To address this issue, those skilled in the art have been exploring various methods to address blind spots in video capture.
[0003] A common approach to reducing blind spots in video capture is to install more hardware devices in the video capture area. However, this approach suffers from the following drawbacks: more hardware devices lead to higher costs, and adding hardware devices still cannot completely resolve the issue of being unable to fully perceive crowd situations due to blind spots in video capture.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, storage medium and system for acquiring crowd situation, so as to at least solve the technical problem in the prior art of achieving full-area perception of crowd situation by reducing blind spots in video acquisition by adding video acquisition hardware, resulting in increased costs but unsatisfactory results.
[0006] According to one aspect of an embodiment of the present application, a crowd situation acquisition method is provided, including: determining a first area range and a second area range within a target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; acquiring spatiotemporal perception information associated with the first area range; predicting a first crowd situation within the first area range based on the spatiotemporal perception information; and acquiring a target crowd situation within the target video acquisition area range by using the first crowd situation and the second crowd situation within the second area range.
[0007] According to another aspect of an embodiment of the present application, a crowd situation acquisition method is also provided, including: receiving a video capture image from a client, wherein the content displayed by the video capture image includes: a target video capture area range; based on the video capture image, determining a first area range and a second area range within the target video capture area range, obtaining spatiotemporal perception information associated with the first area range, predicting a first crowd situation in the first area range based on the spatiotemporal perception information, and obtaining a target crowd situation in the target video capture area range using the first crowd situation and the second crowd situation in the second area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; and feeding back the target crowd situation to the client.
[0008] According to another aspect of an embodiment of the present application, a crowd situation acquisition device is also provided, including: a determination module, used to determine a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; a first acquisition module, used to acquire spatiotemporal perception information associated with the first area range; a prediction module, used to predict a first crowd situation in the first area range based on the spatiotemporal perception information; a second acquisition module, used to use the first crowd situation and the second crowd situation in the second area range to acquire the target crowd situation in the target video acquisition area range.
[0009] According to another aspect of an embodiment of the present application, a storage medium is further provided, wherein the storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute any one of the above-mentioned crowd situation acquisition methods.
[0010] According to another aspect of an embodiment of the present application, a crowd situation acquisition system is also provided, including: a processor; and a memory, connected to the above-mentioned processor, for providing the above-mentioned processor with instructions for processing the following processing steps: determining a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; obtaining spatiotemporal perception information associated with the first area range; predicting a first crowd situation in the first area range based on the spatiotemporal perception information; and using the first crowd situation and the second crowd situation in the second area range to obtain the target crowd situation in the target video acquisition area range.
[0011] In an embodiment of the present application, a first area range and a second area range are determined within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; by obtaining spatiotemporal perception information associated with the first area range, and predicting the first crowd situation in the first area range based on the spatiotemporal perception information, and utilizing the first crowd situation and the second crowd situation in the second area range, the target crowd situation in the target video acquisition area range is obtained.
[0012] It is easy to notice that through the embodiments of the present application, based on the interaction patterns between people in the blind spot and people in the non-blind spot, it is possible to indirectly perceive the situation of people in the blind spot by using simulation deduction. By combining the situation of people in the blind spot with the situation of people in the non-blind spot, the situation of people in the entire area within the video capture range can be obtained.
[0013] Therefore, the embodiment of the present application achieves the purpose of predicting the crowd situation in the video capture blind spot through the simulation estimation algorithm and then obtaining the global crowd situation, thereby achieving the technical effect of being able to perform global perception of the crowd situation without adding video capture hardware, and thus solving the technical problem in the existing technology of reducing the video capture blind spot by adding video capture hardware to achieve global perception of the crowd situation, resulting in increased costs but unsatisfactory results. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0015] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a crowd situation acquisition method is shown;
[0016] Figure 2 is a flow chart of a method for acquiring crowd situation according to an embodiment of the present application;
[0017] Figure 3 is a flowchart of an optional crowd situation acquisition method according to an embodiment of the present application;
[0018] Figure 4 is a flow chart of another method for acquiring crowd situation according to an embodiment of the present application;
[0019] Figure 5 This is a schematic diagram of acquiring crowd situation on a cloud server according to an embodiment of the present application;
[0020] Figure 6 is a structural diagram of a crowd situation acquisition device according to an embodiment of the present application;
[0021] Figure 7 It is a structural block diagram of another computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0024] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0025] Crowd situation: This refers to the state of each pedestrian in a crowd, as well as the overall state of the crowd. Common crowd situation parameters include crowd density and the number of pedestrians in the crowd. Situational awareness: This refers to the perception and understanding of various elements or objects in a dynamic environment in a specific time and space, as well as the prediction of future states.
[0026] Example 1
[0027] According to an embodiment of the present application, an embodiment of a crowd situation acquisition method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0028] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing a crowd situation acquisition method is shown in FIG. Figure 1As shown, the computer terminal 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0029] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the crowd situation acquisition method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned crowd situation acquisition method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0031] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0032] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0033] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the aforementioned computer device (or mobile device).
[0034] Under the above operating environment, this application provides Figure 2 A method for acquiring crowd situation is shown in FIG. Figure 2 is a flow chart of a method for acquiring crowd situation according to an embodiment of the present application. Figure 2 As shown, the crowd situation acquisition method includes:
[0035] Step S202: determining a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range;
[0036] Step S204: obtaining spatiotemporal perception information associated with the first area range;
[0037] Step S206, predicting a first crowd situation in a first area based on the spatiotemporal perception information;
[0038] Step S208 , using the first crowd situation and the second crowd situation in the second area, obtain the target crowd situation in the target video acquisition area.
[0039] It is easy to notice that through the embodiments of the present application, based on the interaction patterns between people in the blind spot and people in the non-blind spot, it is possible to indirectly perceive the situation of people in the blind spot by using simulation deduction. By combining the situation of people in the blind spot with the situation of people in the non-blind spot, the situation of people in the entire area within the video capture range can be obtained.
[0040] Therefore, the embodiment of the present application achieves the purpose of predicting the crowd situation in the video capture blind spot through the simulation estimation algorithm and then obtaining the global crowd situation, thereby achieving the technical effect of being able to perform global perception of the crowd situation without adding video capture hardware, and thus solving the technical problem in the existing technology of reducing the video capture blind spot by adding video capture hardware to achieve global perception of the crowd situation, resulting in increased costs but unsatisfactory results.
[0041] Optionally, the crowd situation acquisition method provided in this application can be applied to, but is not limited to, scenarios such as crowd video monitoring and precise crowd flow control. By adopting the crowd situation acquisition method in the embodiments of this application, it is possible to obtain the global crowd situation within the video capture area through simulation estimation using the video captured by the original equipment without adding video capture hardware.
[0042] Optionally, the blind spot range may be a range within the video capture area where the hardware device cannot capture video of the crowd due to obstacles blocking the crowd. The blind spot range may also be a range within the video capture area that cannot be covered by existing hardware equipment. The non-blind spot range is a range within the video capture area where the hardware device can always capture video of the crowd normally.
[0043] Optionally, the data used to determine the first area range and the second area range within the target video acquisition area range may be a video captured by a hardware device within the target video acquisition area range.
[0044] Optionally, the crowd situation may include: the crowd density of the crowd; the number of pedestrians in the crowd; the position and movement status of each of the multiple pedestrians in the crowd, etc. In practical applications, by studying the crowd situation, it is possible to achieve safety monitoring and precise control of the crowd.
[0045] Figure 3 is a flow chart of an optional crowd situation acquisition method according to an embodiment of the present application; Figure 3As shown, during the process of global crowd situation perception for video acquisition area Area0, due to the presence of blind area Area1 within video acquisition area Area0, the crowd situation within Area1 cannot be directly perceived. In this case, according to the embodiment of the present application, the global crowd situation of video acquisition area Area0 can be indirectly perceived through the interaction between the crowd in blind area Area1 and the crowd in non-blind area Area2.
[0046] Specifically, it is still Figure 3 As shown, the indirect global crowd situation awareness for the video acquisition area Area0 includes the following method steps:
[0047] Step S302, determining the blind area range, dividing the space grid and time segment;
[0048] Step S304: Based on the historical positions of people in the non-blind zone over the past period of time, execute Algorithm 1 to estimate their initial movement direction;
[0049] Step S306: Using each person's initial trajectory, loop through Algorithm 2 to gradually form each person's position trajectory.
[0050] Step S308: Using the pedestrian position prediction results, obtain the movement trajectories of pedestrians entering the blind spot from the non-blind spot and pedestrians leaving the blind spot. Calculate the pedestrian density and number of pedestrians in the blind spot.
[0051] Step S310 , using the density and number of people in the blind spot and the directly perceived density and number of people in the non-blind spot, obtain the density and number of people in the entire area.
[0052] The above method steps implement the following indirect perception process: using the crowds entering and exiting the blind area Area1 from the non-blind area Area2 in the past period of time T, the crowd status in the blind area Area1 at the current moment is estimated; combined with the crowd status in the non-blind area Area2, the crowd perception of the entire area Area0 is obtained.
[0053] In an optional embodiment, in step S202, determining a first area range and a second area range within the target video acquisition area range includes the following method steps:
[0054] Step S221, dividing the target video acquisition area into a plurality of unit areas;
[0055] Step S222, dividing the first preset time range into multiple time segments;
[0056] Step S223, in each of the multiple time segments, counting the number of people in each of the multiple unit areas in turn to obtain a statistical result;
[0057] Step S224: determining the first area range and the second area range based on the statistical results.
[0058] The multiple unit areas may be multiple areas of identical size and shape, used to divide the target wired control area. The multiple time segments may be multiple time intervals of identical length, used to divide the first preset time range. The first preset time range may be a manually set time range used to determine the first and second area ranges, and the first preset time range may be selected from the time period during which video capture is performed within the target video capture area.
[0059] For each of the multiple time segments, the number of people in each of the multiple unit areas is counted within the time segment. The number of people in the multiple unit areas within the multiple time segments is obtained as a statistical result. Based on the statistical result, the first area range and the second area range can be determined. The first area range is a blind area range, and the second area range is a non-blind area range.
[0060] In an optional embodiment, in step S224, determining the first area range and the second area range based on the statistical results includes the following method steps:
[0061] Step S2241, determining a first unit area and a second unit area from the plurality of unit areas based on the statistical results, wherein the first unit area has a continuously empty crowd count in some or all of the plurality of time segments, and the second unit area is the remaining unit areas in the plurality of unit areas except the first unit area;
[0062] Step S2242: determine a first area range using the first unit area, and determine a second area range using the second unit area.
[0063] Based on the statistical results, i.e., the number of people in the multiple unit areas during the multiple time segments, the multiple unit areas can be divided into first unit areas and second unit areas. The division is based on the following: if no people are counted in a unit area during some or all of the multiple time segments, the unit area belongs to the first unit area; otherwise, the unit area belongs to the second unit area.
[0064] According to the above division criteria, the first unit area obtained by the division is determined as the first area range, that is, the blind area range; the second unit area obtained by the division is determined as the second area range, that is, the non-blind area range.
[0065] Still as Figure 3 shown, taking the example of performing global population situation perception on the video acquisition area Area0, in step S302, to determine the blind area range Area1, dividing the spatial grid and time segments, including the following method steps: Divide the video acquisition area Area0 into L×W squares according to a fixed side length, where L is the number of squares divided along the length direction of the video acquisition area Area0, and W is the number of squares divided along the width direction of the video acquisition area Area0; preset a time range T1 for determining the blind area and non-blind area ranges, and divide this time range T1 into n1 time segments Δt1; use the camera to collect the coordinate position information of the people in the video acquisition area Area0 within each time segment Δt1 of the n1 time segments in the above time range T1; count the number of pedestrians in each square within each time segment Δt1 of the n time segments in the above time range T1, and record this statistical result as R; determine the blind area range Area1 within the above video acquisition area Area0 according to this statistical result R.
[0066] Optionally, the criterion for determining whether each square in the L×W squares within the video acquisition area Area0 belongs to the blind area range Area1 can be: If in the statistical result R, the number of pedestrians in a square is zero within consecutive n0 (n * <n0<n1) time segments Δt1, then this square belongs to the blind area range Area1; otherwise, this square belongs to the non-blind area range Area2. Among them, preset another time range TA (TA<T1) as the minimum continuous time range for determining whether a square belongs to the blind area range. Here TA = n * ×Δt1, and take this minimum continuous time range TA as 1 hour, that is, at least within 1 hour, the number of pedestrians in this square continues to be zero to determine that this square belongs to the blind area range.
[0067] Optionally, the criterion for determining whether each square in the L×W squares within the video acquisition area Area0 belongs to the blind area range Area1 can also be: If in the statistical result R, the number of pedestrians in a square is zero within all n1 time segments Δt1, then this square belongs to the blind area range Area1; otherwise, this square belongs to the non-blind area range Area2. Among them, since the above T1>TA, therefore, n1>n * .
[0068] In an optional embodiment, in step S204, obtain the spatio-temporal perception information associated with the first area range, including the following method steps:
[0069] Step S241, estimating the initial movement direction of the crowd within the second area within a second preset time range, wherein the second preset time range is a preset historical time range;
[0070] Step S242 , using the initial movement direction to predict the movement trajectory of the crowd within a third preset time range, to obtain spatiotemporal perception information, wherein the third preset time range is a preset future time range, and the movement trajectory is a future movement trajectory.
[0071] The second preset time range may be a manually set time range for determining the spatiotemporal perception information. The second preset time range may be selected from the historical period of video capture within the target video capture area. Based on the statistical results, the initial movement direction of the crowd within the second preset time range can be estimated. The second area is a non-blind zone, and the initial movement direction is the possible movement direction of the crowd at the current moment.
[0072] Based on the initial movement direction, the movement direction of the group within the third preset time range can be predicted, thereby obtaining the spatiotemporal perception information. The third preset time range can be a manually set time range for determining the spatiotemporal perception information, and the third preset time range can be selected from a future time range during which video is captured within the target video capture area. The movement trajectory represents the possible movement trajectory of the group within the future time range.
[0073] In an optional embodiment, in step S241, estimating the initial movement direction of the crowd located in the second area within the second preset time range includes the following method steps:
[0074] Step S2411: Analyze the multiple first positions using a first neural network model to determine a second position, wherein the first neural network model is obtained through machine learning training using multiple sets of data, each set of data including: pedestrian positions and crowd flow data at multiple sampling moments, the multiple first positions being multiple consecutive historical positions of each pedestrian in the crowd within the second area at multiple historical moments, and the second positions being the position of each pedestrian in the crowd within the second area at a moment next to the multiple historical moments;
[0075] Step S2412: Determine an initial movement direction using the third position and the second position, wherein the third position is the first position corresponding to a previous moment adjacent to the second position.
[0076] The second preset time range, i.e., the preset historical time range, includes multiple consecutive historical moments. The crowd within the second area includes multiple pedestrians, i.e., multiple pedestrians within the non-blind spot. The multiple first positions are the multiple consecutive historical positions of each of the multiple pedestrians at the multiple consecutive historical moments; the multiple second positions are the position of each of the multiple pedestrians at the next moment after the multiple consecutive historical moments.
[0077] The first neural network model is a neural network model trained through machine learning using multiple sets of data, each of which includes pedestrian location and crowd flow data at multiple sampling times. The first neural network model can obtain the second location by analyzing the multiple first locations.
[0078] The third position is the first position corresponding to the previous moment adjacent to the second position. Based on the third position and the second position, the initial movement direction can be determined. The initial movement direction is the possible movement direction of the people within the second area at the current moment.
[0079] Still like Figure 3 As shown, taking the global crowd situation awareness of the video acquisition area Area0 as an example, in step S304, according to the historical position PT of the crowd in the non-blind area Area2 within the past time range T, Algorithm 1 is executed to estimate the initial movement direction of the crowd, wherein the time range T includes the -Tth to -1th moments, and the crowd includes Nt pedestrians, that is, the historical position PT includes the position coordinates of the above Nt pedestrians from the -Tth to -1th moments in the past.
[0080] Optionally, the input of the above algorithm 1 can be the historical position PT of the crowd in the non-blind area Area2 within the past period of time T = {p (-T) ,…,P (t) ,…,p (-1) |-T≤t≤-1}; where the pedestrian position data at the tth moment is The position of each pedestrian represents the coordinates of the square where the pedestrian is located, recorded as The value range of x is 1~L, and the value range of y is 1~W; the output of the algorithm 1 can be the pedestrian position at the current moment (the 0th moment), that is,
[0081] Optionally, the model of the above algorithm 1 can be a convolutional neural network (CNN), denoted as CNN1. The input layer size of CNN1 is T×L×W, and the output layer size is L×W. CNN1 includes the following six-layer structure:
[0082] 1) The first two-dimensional convolution layer in the spatial direction, which includes 4 convolution kernels, where the size of each convolution kernel is The output size is T×4×8×8;
[0083] 2) The second 2D convolutional layer in the spatial direction, which includes 16 convolution kernels, where each convolution kernel has a size of 8×8 and an output size of T×16;
[0084] 3) A one-dimensional convolution layer in the time direction, which includes 4 convolution kernels, where the size of each convolution kernel is T and the output size is 4×16, that is, the output is a 64-dimensional vector;
[0085] 4) The first two-dimensional deconvolution layer in the spatial direction, the output size of this two-dimensional deconvolution layer is
[0086] 5) A second 2D deconvolution layer in the spatial direction, where the output size of the 2D deconvolution layer is 4×L×W;
[0087] 6) A one-dimensional convolutional layer in the channel dimension, where the output size of the one-dimensional convolutional layer is L x W.
[0088] It should be noted that the activation function used between two adjacent layers in the six-layer structure of CNN1 can be a Rectified Linear Unit (ReLU) function in the non-negative range. In addition, Batch Normalization is required before the activation function.
[0089] Optionally, the training process of the above CNN1 is as follows:
[0090] 1) Collect training data;
[0091] In the non-blind area Area2, the pedestrian position data and pedestrian flow data are collected at each moment in the past time range T (from the -Tth moment to the -1th moment). For example, for the tth moment, the above training data includes: pedestrian position data Among them, the position of each pedestrian Represents the coordinates of the square where the pedestrian is located (1≤x≤L, 1≤y≤W); the real pedestrian flow data is represented by the pedestrian flow vector Among them, L represents the length of the pedestrian flow vector, which is also the number of line segments for calculating pedestrian flow. The two endpoints of the lth line segment are and The flow vector of people at the lth position It represents the number of people passing through the lth line segment at time t (i.e., flow rate).
[0092] 2) Introducing constraints;
[0093] During the training process of Algorithm 1, three constraints are introduced, namely, the following three assumptions are made in the calculation: there is an upper limit on the pedestrian movement speed; the cost of changing the pedestrian density distribution is as low as possible; and the pedestrian flow caused by the trajectory of pedestrians from the current location to the predicted location must be as close as possible to the actual pedestrian flow.
[0094] 3) Determine the objective function;
[0095] After introducing the above three constraints, the objective function used to train Algorithm 1 is determined as shown in the following formula (1):
[0096]
[0097] In formula (1), min is the minimum function, max is the maximum function; v is the upper limit of the average speed of the person assumed in the constraint condition; Δt is the duration of each moment from the -Tth to the -1th moment in the time range T.
[0098] In formula (1), represents the position of the nth person at the current moment (i.e., the 0th moment) predicted by the model, where θ is the weight parameter of CNN1; Indicates the position prediction result of the crowd in Area2 at the current moment (i.e., the 0th moment) The predicted crowd density at the current moment (i.e., the 0th moment) obtained by statistics; D (0) Indicates the actual crowd density at the current moment (i.e., the 0th moment); Represents the actual crowd flow at the current moment (i.e., the 0th moment); N -1 is the number of pedestrians at the previous moment of the current moment (i.e., moment 0).
[0099] In formula (1), function f represents the distance between two pedestrian density distributions, which may be the Wasserstein distance, and is used to represent the minimum cost of the above-mentioned change in pedestrian density distribution.
[0100] In formula (1), the function It is used to determine whether the straight line path of the nth pedestrian from the position at the -1th moment to the predicted current moment (i.e., the 0th moment) intersects with the above-mentioned lth line segment. If so, the function The value of is 1; otherwise, the function The value of is 0. Specifically, the function The value rule of is shown in the following formula (2):
[0101]
[0102] In formula (2), the function g(p1,p2,q1,q2) indicates whether points q1 and q2 are on the same side of the line p1p2. If they are on the same side, the value of the function g(p1,p2,q1,q2) is 0; otherwise, the value of the function g(p1,p2,q1,q2) is 1.
[0103] 4) Optimize the objective function.
[0104] By optimizing the above objective function using the stochastic gradient descent method, the target weights of CNN1 can be obtained.
[0105] Optionally, the model testing process of the above algorithm 1 can be: using the target weight of CNN1 obtained by the above optimization objective function, each test data can be tested.
[0106] Specifically, based on the historical positions PT of the crowd in the non-blind area Area2 over the past T time periods (from -Tth to -1th), the current position of each pedestrian in the crowd can be estimated. The initial movement direction of each pedestrian is then obtained by subtracting their positions at -1th and the current time period (i.e., 0th).
[0107] In an optional embodiment, in step S242, the movement trajectory of the crowd within the third preset time range is predicted using the initial movement direction to obtain spatiotemporal perception information, including the following method steps:
[0108] Step S2421: Analyze the plurality of fourth positions using a second neural network model to determine a fifth position, wherein the second neural network model is trained using multiple sets of data through machine learning, each set of the multiple sets of data including: pedestrian positions and crowd flow data at multiple sampling moments before a preset moment, the plurality of fourth positions being a plurality of consecutive historical positions corresponding to multiple historical moments before a target moment for each pedestrian in the crowd within the second area, and the fifth position being the position of each pedestrian in the crowd within the second area at a next moment after the target moment;
[0109] Step S2422: Use multiple fifth positions obtained by gradual deduction through the second neural network model to obtain spatiotemporal perception information.
[0110] The crowd within the second area includes multiple pedestrians, i.e., multiple pedestrians within the non-blind area. The preset moment is the current moment of video capture, and the target moment is the current moment when the position of each of the multiple pedestrians is predicted to be at the next moment. The target moment includes multiple historical moments, which may be the historical moments of video capture or multiple moments at which the position of each of the multiple pedestrians has been predicted. The multiple fourth positions are the multiple consecutive historical positions of each of the multiple pedestrians at the multiple historical moments; the fifth position is the position of each of the multiple pedestrians at the next moment after the target moment.
[0111] The second neural network model is a neural network model trained through machine learning using multiple additional sets of data, each of which includes pedestrian location and crowd flow data at multiple sampling times before the predetermined time. The second neural network model can derive the fifth position by analyzing the multiple fourth positions.
[0112] By using the second neural network model multiple times, multiple fifth positions can be gradually deduced, thereby obtaining the spatiotemporal perception information. When the second neural network model is used multiple times for analysis, the next moment after the target moment corresponding to each analysis is used as the target moment for the next analysis. In other words, the target moment for each analysis is the next moment after the target moment corresponding to the previous analysis.
[0113] Still like Figure 3 As shown, taking the global crowd situation awareness of the video acquisition area Area0 as an example, in step S306, the initial trajectory of the crowd in the non-blind area Area2 is used to loop through Algorithm 2 to gradually form the position trajectory PX of each pedestrian in the non-blind area Area2. That is, starting from the t-th moment (t≥1), Algorithm 2 can predict the position of each pedestrian at the next moment (i.e., the t+1-th moment) based on the pedestrian's historical position and a part of the predicted future position of the pedestrian. Repeated use of Algorithm 2 can form the position of each pedestrian in the crowd in the non-blind area Area2 at multiple moments in the future, that is, predict the future movement trajectory of the crowd.
[0114] Optionally, the model of the above algorithm 2 can be a convolutional neural network CNN2, where the input layer size of CNN2 is T×L×W and the output layer size is L×W.
[0115] Optionally, the input of the above CNN2 can be the pedestrian position of the past T moments before the t+1 moment (i.e., the t-T+1 moment to the t-th moment), where the pedestrian position data at the s-th moment is The position of the nth pedestrian Represents the coordinates of the square where the pedestrian is located (1≤x≤L, 1≤y≤W); the output of CNN2 can be the pedestrian position at the t+1th moment
[0116]
[0117] It should be noted that each time CNN2 is used, the pedestrian position output by CNN2 is used as the pedestrian position input for the next use of CNN2, thus implementing a cyclic use of Algorithm 2. Specifically, the pedestrian position output at time t+1 when CNN2 is used in one use is used as the pedestrian position input at time t when CNN2 is used in the next use.
[0118] Similar to CNN1, CNN2 also includes the aforementioned six-layer structure, but the two correspond to different network parameters and are trained according to their respective network parameters. Optionally, the training process of the above CNN2 is as follows:
[0119] 1) Collect training data;
[0120] In the non-blind area Area2, the pedestrian position data and pedestrian flow data within T moments before the t+1 moment (i.e., the t-T+1 moment to the t-th moment) are collected. For example, for the r-th moment (t-T+1≤r≤t), the above training data includes: the pedestrian position data at the r-th moment Among them, the position of each pedestrian Represents the coordinates of the square where the pedestrian is located (1≤x≤L,1≤y≤W); the real pedestrian flow data at the r+1th moment is expressed as a pedestrian flow vector Among them, L represents the length of the pedestrian flow vector, which is also the number of line segments for calculating pedestrian flow. The two endpoints of the lth line segment are and The flow vector of people at the lth position It represents the number of people passing through the lth line segment at the r+1th moment (i.e., flow rate).
[0121] 2) Introducing constraints;
[0122] During the training of CNN2, four constraints are introduced, namely, the following four assumptions are made in the calculation: there is an upper limit on the speed of pedestrian movement; the cost of changing the pedestrian density distribution is as low as possible; the pedestrian flow rate caused by the trajectory of pedestrians from the current location to the predicted location should be as close as possible to the actual pedestrian flow rate; and the direction of the pedestrian movement trajectory should remain stable.
[0123] 3) Determine the objective function;
[0124] After introducing the above four constraints, the objective function for training CNN2 is determined as shown in the following formula (3):
[0125]
[0126] In formula (3), min is the minimum function, max is the maximum function; v is the upper limit of the average speed of the person assumed in the constraint condition; Δt is the duration of each moment from the t-T+1th moment to the tth moment.
[0127] In formula (3), represents the position of the nth person at time t+1 predicted by the model, where θ′ is the weight parameter of CNN2; Represents the predicted result of the position of the crowd in Area2 at the t+1th moment The predicted crowd density at time t+1 obtained by statistics; D (t+1) The actual crowd density at time t+1; Represents the real crowd flow at the t+1th moment.
[0128] In formula (3), It is the angular deviation of the pedestrian's movement direction between the previous and next steps, which is used to indicate the stability of the pedestrian's trajectory direction. The specific calculation method is shown in the following formula (4):
[0129]
[0130] 4) Optimize the objective function.
[0131] By using the stochastic gradient descent method, the above objective function can be optimized to obtain the target weight of the CNN2.
[0132] Optionally, the CNN2 test process can be as follows: using the target weight of the CNN2 obtained from the optimization objective function, each test data can be tested to obtain the pedestrian position deduced in a single step. Algorithm 2 can be used repeatedly to gradually deduce the position of each pedestrian at multiple future moments.
[0133] In an optional embodiment, in step S206, predicting a first crowd situation in a first area based on spatiotemporal perception information includes the following method steps:
[0134] Step S261: Acquire a first portion of the crowd and a second portion of the crowd based on the spatiotemporal perception information, wherein the first portion of the crowd is the crowd entering the first area from the second area, and the second portion of the crowd is the crowd entering the second area from the first area;
[0135] Step S262: predicting the situation of the first group of people using the first group of people and the second group of people.
[0136] The first group of people mentioned above are those who enter the first area from the second area, that is, those who enter the blind area from the non-blind area, causing pedestrians to disappear from the video capture field of view. The second group of people mentioned above are those who enter the second area from the first area, that is, those who enter the non-blind area from the blind area, causing pedestrians to appear in the video capture field of view. The first group of people and the second group of people can be obtained based on the above-mentioned spatiotemporal perception information.
[0137] By using the first part of the population and the second part of the population, a first population situation can be predicted. The first population situation may be the population density, the number of people, etc. within the first area.
[0138] Still like Figure 3 As shown, taking the global crowd situation awareness of the video acquisition area Area0 as an example, in step S308, at a certain moment, the predicted result of the pedestrian position in the non-blind area Area2 is used to obtain the number Min of pedestrians whose movement trajectory enters the blind area Area1 from the non-blind area Area2, and the number Mout of pedestrians who leave the blind area Area1; using the number Min and the number Mout, the pedestrian density D1 in the blind area range Area1 at that moment can be obtained; by subtracting the number Min from the number Mout, the number C1 of pedestrians in the blind area range Area1 at that moment can be calculated.
[0139] According to step S208, the first crowd situation within the first area and the second crowd situation within the second area are combined to obtain the target crowd situation within the target video acquisition area. The second crowd situation may be the crowd density, the number of people, etc. within the second area. The target crowd situation may be the density and number of people within the entire target video acquisition area, i.e., the first area and the second area.
[0140] Still like Figure 3 As shown, taking the global crowd situation awareness of the video acquisition area Area0 as an example, in step S310, at a certain moment, the pedestrian density D1 and the number of pedestrians C1 in the blind area range Area1 at that moment are used, combined with the crowd density D2 and the number of pedestrians C2 in the non-blind area range Area2 directly perceived by the hardware device at that moment, the global crowd density D0 and the global crowd number C0 in the entire area Area0 at that moment can be obtained.
[0141] One embodiment of the present application further provides a method for acquiring crowd situation, which is run on a cloud server. Figure 4 is a flow chart of another optional crowd situation acquisition method according to an embodiment of the present application, such as Figure 4 As shown, the crowd situation acquisition method includes:
[0142] Step S402: receiving a video capture image from a client, wherein the video capture image displays content including: a target video capture area range;
[0143] Step S404: Based on the video capture image, a first area range and a second area range are determined within the target video capture area range, spatiotemporal perception information associated with the first area range is obtained, a first crowd situation within the first area range is predicted based on the spatiotemporal perception information, and a target crowd situation within the target video capture area range is obtained using the first crowd situation and the second crowd situation within the second area range, wherein the first area range is a blind area range and the second area range is a non-blind area range;
[0144] Step S406: Feedback the target population situation to the client.
[0145] Optionally, Figure 5 is a schematic diagram of acquiring crowd situation on a cloud server according to an embodiment of the present application, such as Figure 5 As shown, the client uploads the video capture image to the cloud server, wherein the content displayed in the video capture image includes: the target video capture area range; the cloud server determines the first area range and the second area range within the target video capture area range based on the video capture image, obtains the spatiotemporal perception information associated with the first area range, predicts the first crowd situation in the first area range based on the spatiotemporal perception information, and obtains the target crowd situation in the target video capture area range using the first crowd situation and the second crowd situation in the second area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; the cloud server feeds back the above-mentioned target crowd situation as a crowd situation acquisition result to the above-mentioned client, and the final crowd situation acquisition result will be provided to the user through the client.
[0146] It should be noted that the above-mentioned crowd situation acquisition method provided in the embodiment of the present application can be, but is not limited to, applicable to actual application scenarios such as precise passenger flow control. Through the interaction between the SaaS server and the client, the target crowd situation is obtained by predicting the crowd situation within the blind spot based on the spatiotemporal perception information and combining it with the crowd situation within the non-blind spot, and the returned target crowd situation is provided to the user through the client.
[0147] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0149] Example 2
[0150] According to an embodiment of the present application, a device embodiment for implementing the above-mentioned crowd situation acquisition method is also provided. Figure 6 is a structural diagram of a crowd situation acquisition device according to an embodiment of the present application, such as Figure 6 As shown, the device includes: a determination module 601, a first acquisition module 602, a prediction module 603, and a second acquisition module 604, wherein:
[0151] The determination module 601 is used to determine the first area range and the second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; the first acquisition module 602 is used to obtain the spatiotemporal perception information associated with the first area range; the prediction module 603 is used to predict the first crowd situation in the first area range based on the spatiotemporal perception information; the second acquisition module 604 is used to use the first crowd situation and the second crowd situation in the second area range to obtain the target crowd situation in the target video acquisition area range.
[0152] Optionally, the determination module 601 is further used to: divide the target video acquisition area range to obtain multiple unit areas; divide the first preset time range to obtain multiple time segments; in each of the multiple time segments, count the number of people in each of the multiple unit areas in turn to obtain statistical results; and determine the first area range and the second area range based on the statistical results.
[0153] Optionally, the determination module 601 is also used to: determine a first unit area and a second unit area from multiple unit areas based on statistical results, respectively, wherein the number of people in the first unit area is continuously empty in some or all of the continuous time segments of the multiple time segments, and the second unit area is the remaining unit areas in the multiple unit areas except the first unit area; use the first unit area to determine the range of the first area, and use the second unit area to determine the range of the second area.
[0154] Optionally, the first acquisition module 602 is also used to: estimate the initial movement direction of the crowd located in the second area within a second preset time range, wherein the second preset time range is a preset historical time range; use the initial movement direction to predict the movement trajectory of the crowd within a third preset time range to obtain spatiotemporal perception information, wherein the third preset time range is a preset future time range, and the movement trajectory is a future movement trajectory.
[0155] Optionally, the first acquisition module 602 is also used to: analyze multiple first positions using a first neural network model to determine a second position, wherein the first neural network model is obtained through machine learning training using multiple sets of data, and each set of data in the multiple sets of data includes: pedestrian positions and crowd flow data at multiple sampling moments, the multiple first positions are multiple continuous historical positions of each pedestrian in the crowd located within the second area at multiple historical moments, and the second position is the position of each pedestrian in the crowd located within the second area at the next moment after the multiple historical moments; the third position and the second position are used to determine the initial movement direction, wherein the third position is the first position corresponding to the previous moment adjacent to the second position.
[0156] Optionally, the first acquisition module 602 is further used to: analyze multiple fourth positions using the second neural network model to determine a fifth position, wherein the second neural network model is obtained through machine learning training using multiple sets of data, and each set of data includes: pedestrian positions and crowd flow data at multiple sampling moments before a preset moment, the multiple fourth positions are multiple continuous historical positions corresponding to multiple historical moments before the target moment for each pedestrian in the crowd within the second area, and the fifth position is the position of each pedestrian in the crowd within the second area at the next moment after the target moment; the multiple fifth positions obtained by stepwise deduction through the second neural network model are used to obtain spatiotemporal perception information.
[0157] Optionally, the prediction module 603 is also used to: obtain a first part of the population and a second part of the population based on spatiotemporal perception information, wherein the first part of the population is a population entering the first area range from the second area range, and the second part of the population is a population entering the second area range from the first area range; and use the first part of the population and the second part of the population to predict the situation of the first population.
[0158] It should be noted that the determination module 601, the first acquisition module 602, the prediction module 603, and the second acquisition module 604 correspond to steps S202 to S208 in Example 1. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0159] In an embodiment of the present application, a first area range and a second area range are determined within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; by obtaining spatiotemporal perception information associated with the first area range, and predicting the first crowd situation in the first area range based on the spatiotemporal perception information, and utilizing the first crowd situation and the second crowd situation in the second area range, the target crowd situation in the target video acquisition area range is obtained.
[0160] It is easy to notice that through the embodiments of the present application, based on the interaction patterns between people in the blind spot and people in the non-blind spot, it is possible to indirectly perceive the situation of people in the blind spot by using simulation deduction. By combining the situation of people in the blind spot with the situation of people in the non-blind spot, the situation of people in the entire area within the video capture range can be obtained.
[0161] Therefore, the embodiment of the present application achieves the purpose of predicting the crowd situation in the video capture blind spot through the simulation estimation algorithm and then obtaining the global crowd situation, thereby achieving the technical effect of being able to perform global perception of the crowd situation without adding video capture hardware, and thus solving the technical problem in the existing technology of reducing the video capture blind spot by adding video capture hardware to achieve global perception of the crowd situation, resulting in increased costs but unsatisfactory results.
[0162] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.
[0163] Example 3
[0164] According to an embodiment of the present application, an embodiment of an electronic device is also provided. The electronic device can be any computing device in a computing device group. The electronic device includes: a processor and a memory, wherein:
[0165] The memory is connected to the above-mentioned processor and is used to provide the above-mentioned processor with instructions for processing the following processing steps: determining a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; obtaining spatiotemporal perception information associated with the first area range; predicting a first crowd situation in the first area range based on the spatiotemporal perception information; and using the first crowd situation and the second crowd situation in the second area range to obtain a target crowd situation in the target video acquisition area range.
[0166] In an embodiment of the present application, a first area range and a second area range are determined within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; by obtaining the spatiotemporal perception information associated with the first area range, and predicting the first crowd situation of the first area range based on the spatiotemporal perception information, the first crowd situation and the second crowd situation of the second area range are used to obtain the target crowd situation of the target video acquisition area range. It is easy to notice that through the embodiment of the present application, according to the interaction rules between the crowd in the blind area and the crowd in the non-blind area, the situation of the crowd in the blind area can be indirectly perceived by means of simulation deduction. The blind area range is combined with the crowd situation in the non-blind area to obtain the global crowd situation within the video acquisition range.
[0167] Therefore, the embodiment of the present application achieves the purpose of predicting the crowd situation in the video capture blind spot through the simulation estimation algorithm and then obtaining the global crowd situation, thereby achieving the technical effect of being able to perform global perception of the crowd situation without adding video capture hardware, and thus solving the technical problem in the existing technology of reducing the video capture blind spot by adding video capture hardware to achieve global perception of the crowd situation, resulting in increased costs but unsatisfactory results.
[0168] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.
[0169] Example 4
[0170] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0171] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0172] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the crowd situation acquisition method: determining a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; obtaining spatiotemporal perception information associated with the first area range; predicting a first crowd situation in the first area range based on the spatiotemporal perception information; and using the first crowd situation and the second crowd situation in the second area range to obtain the target crowd situation in the target video acquisition area range.
[0173] Optionally, Figure 7 is a structural block diagram of another computer terminal according to an embodiment of the present application, such as Figure 7 As shown, the computer terminal may include: one or more (only one is shown in the figure) processors 122 , a memory 124 , and a peripheral interface 126 .
[0174] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the crowd situation acquisition method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned crowd situation acquisition method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, corporate intranet, local area network, mobile communication network and combinations thereof.
[0175] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: determine the first area range and the second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; obtain the spatiotemporal perception information associated with the first area range; predict the first crowd situation in the first area range based on the spatiotemporal perception information; use the first crowd situation and the second crowd situation in the second area range to obtain the target crowd situation in the target video acquisition area range.
[0176] Optionally, the processor may also execute the program code of the following steps: dividing the target video acquisition area to obtain a plurality of unit areas; dividing the first preset time range to obtain a plurality of time segments; in each of the plurality of time segments, counting the number of people in each of the plurality of unit areas in turn to obtain statistical results; and determining the first area range and the second area range based on the statistical results.
[0177] Optionally, the processor may also execute the program code of the following steps: determining a first unit area and a second unit area respectively from a plurality of unit areas based on statistical results, wherein the number of people in the first unit area is continuously empty in part or all of the continuous time segments of the plurality of time segments, and the second unit area is the remaining unit areas in the plurality of unit areas except the first unit area; determining the range of the first area using the first unit area, and determining the range of the second area using the second unit area.
[0178] Optionally, the processor may also execute the program code of the following steps: estimating the initial movement direction of a crowd located within a second area within a second preset time range, wherein the second preset time range is a preset historical time range; using the initial movement direction to predict the movement trajectory of the crowd within a third preset time range to obtain spatiotemporal perception information, wherein the third preset time range is a preset future time range, and the movement trajectory is a future movement trajectory.
[0179] Optionally, the processor may also execute the program code of the following steps: using a first neural network model to analyze multiple first positions and determine a second position, wherein the first neural network model is obtained through machine learning training using multiple sets of data, and each set of data in the multiple sets of data includes: pedestrian positions and crowd flow data at multiple sampling moments, the multiple first positions are multiple continuous historical positions of each pedestrian in the crowd located within the second area at multiple historical moments, and the second position is the position of each pedestrian in the crowd located within the second area at the next moment after the multiple historical moments; using the third position and the second position to determine the initial movement direction, wherein the third position is the first position corresponding to the previous moment adjacent to the second position.
[0180] Optionally, the processor may also execute the program code of the following steps: using the second neural network model to analyze multiple fourth positions and determine the fifth position, wherein the second neural network model is obtained through machine learning training using multiple sets of data, and each set of data includes: pedestrian positions and crowd flow data at multiple sampling moments before a preset moment, the multiple fourth positions are multiple continuous historical positions corresponding to multiple historical moments before the target moment for each pedestrian in the crowd within the second area, and the fifth position is the position of each pedestrian in the crowd within the second area at the next moment after the target moment; the multiple fifth positions obtained by stepwise deduction through the second neural network model are used to obtain spatiotemporal perception information.
[0181] Optionally, the processor may also execute the program code of the following steps: obtaining a first part of the population and a second part of the population based on spatiotemporal perception information, wherein the first part of the population is a population entering the first area from the second area, and the second part of the population is a population entering the second area from the first area; and using the first part of the population and the second part of the population to predict the situation of the first population.
[0182] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receiving a video capture image from the client, wherein the content displayed by the video capture image includes: the target video capture area range; based on the video capture image, determining the first area range and the second area range within the target video capture area range, obtaining the spatiotemporal perception information associated with the first area range, predicting the first crowd situation in the first area range based on the spatiotemporal perception information, and obtaining the target crowd situation in the target video capture area range using the first crowd situation and the second crowd situation in the second area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; and feeding back the target crowd situation to the client.
[0183] In an embodiment of the present application, a first area range and a second area range are determined within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; by obtaining spatiotemporal perception information associated with the first area range, and predicting the first crowd situation in the first area range based on the spatiotemporal perception information, and utilizing the first crowd situation and the second crowd situation in the second area range, the target crowd situation in the target video acquisition area range is obtained.
[0184] It is easy to notice that through the embodiments of the present application, based on the interaction patterns between people in the blind spot and people in the non-blind spot, it is possible to indirectly perceive the situation of people in the blind spot by using simulation deduction. By combining the situation of people in the blind spot with the situation of people in the non-blind spot, the situation of people in the entire area within the video capture range can be obtained.
[0185] Therefore, the embodiment of the present application achieves the purpose of predicting the crowd situation in the video capture blind spot through the simulation estimation algorithm and then obtaining the global crowd situation, thereby achieving the technical effect of being able to perform global perception of the crowd situation without adding video capture hardware, and thus solving the technical problem in the existing technology of reducing the video capture blind spot by adding video capture hardware to achieve global perception of the crowd situation, resulting in increased costs but unsatisfactory results.
[0186] It can be understood by those skilled in the art that Figure 7The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 7 It does not limit the structure of the above electronic device. For example, the computer terminal may also include Figure 7 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 7 Different configurations shown.
[0187] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0188] According to an embodiment of the present application, an embodiment of a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the crowd situation acquisition method provided in the above embodiment 1.
[0189] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0190] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: determining a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; obtaining spatiotemporal perception information associated with the first area range; predicting a first crowd situation in the first area range based on the spatiotemporal perception information; and obtaining a target crowd situation in the target video acquisition area range using the first crowd situation and the second crowd situation in the second area range.
[0191] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: dividing the target video acquisition area range to obtain multiple unit areas; dividing the first preset time range to obtain multiple time segments; in each of the multiple time segments, counting the number of people in each of the multiple unit areas in turn to obtain statistical results; and determining the first area range and the second area range based on the statistical results.
[0192] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: determining a first unit area and a second unit area from multiple unit areas based on statistical results, respectively, wherein the number of people in the first unit area is continuously empty in some or all of the continuous time segments of multiple time segments, and the second unit area is the remaining unit areas of the multiple unit areas except the first unit area; determining the range of the first area using the first unit area, and determining the range of the second area using the second unit area.
[0193] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: estimating the initial movement direction of a crowd located in a second area within a second preset time range, wherein the second preset time range is a preset historical time range; using the initial movement direction to predict the movement trajectory of the crowd within a third preset time range to obtain spatiotemporal perception information, wherein the third preset time range is a preset future time range, and the movement trajectory is a future movement trajectory.
[0194] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: using a first neural network model to analyze multiple first positions to determine a second position, wherein the first neural network model is obtained through machine learning training using multiple sets of data, and each set of data in the multiple sets of data includes: pedestrian positions and crowd flow data at multiple sampling moments, the multiple first positions are multiple continuous historical positions of each pedestrian in the crowd located within the second area at multiple historical moments, and the second position is the position of each pedestrian in the crowd located within the second area at the next moment after the multiple historical moments; using the third position and the second position to determine the initial movement direction, wherein the third position is the first position corresponding to the previous moment adjacent to the second position.
[0195] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: using a second neural network model to analyze multiple fourth positions to determine a fifth position, wherein the second neural network model is obtained through machine learning training using multiple sets of data, and each set of data in the multiple sets of data includes: pedestrian positions and crowd flow data at multiple sampling moments before a preset moment, the multiple fourth positions are multiple continuous historical positions corresponding to multiple historical moments before the target moment for each pedestrian in the crowd within the second area, and the fifth position is the position of each pedestrian in the crowd within the second area at the next moment after the target moment; the multiple fifth positions obtained by stepwise deduction through the second neural network model are used to obtain spatiotemporal perception information.
[0196] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining a first part of the population and a second part of the population based on spatiotemporal perception information, wherein the first part of the population is a population entering the first area range from the second area range, and the second part of the population is a population entering the second area range from the first area range; and using the first part of the population and the second part of the population to predict the situation of the first population.
[0197] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: receiving a video capture image from a client, wherein the content displayed by the video capture image includes: the target video capture area range; based on the video capture image, determining a first area range and a second area range within the target video capture area range, obtaining spatiotemporal perception information associated with the first area range, predicting a first crowd situation in the first area range based on the spatiotemporal perception information, and obtaining a target crowd situation in the target video capture area range using the first crowd situation and the second crowd situation in the second area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; and feeding back the target crowd situation to the client.
[0198] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0199] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0200] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0201] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0202] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0203] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0204] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for acquiring crowd situation, characterized in that: include: Determine a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range, and the second area range is a non-blind area range; Acquiring spatiotemporal perception information associated with the first area; Acquire a first part of the crowd and a second part of the crowd based on the spatiotemporal perception information, wherein the first part of the crowd is a crowd entering the first area from the second area, and the second part of the crowd is a crowd entering the second area from the first area; Predicting the situation of the first group of people in the first area using the first group of people and the second group of people; The target crowd situation in the target video acquisition area is obtained by utilizing the first crowd situation and the second crowd situation in the second area.
2. The crowd situation acquisition method according to claim 1, characterized in that: Determining the first area range and the second area range within the target video acquisition area range includes: Dividing the target video acquisition area into a plurality of unit areas; Dividing the first preset time range to obtain multiple time segments; In each of the multiple time segments, counting the number of people in each of the multiple unit areas in turn to obtain a statistical result; The first area range and the second area range are determined based on the statistical results.
3. The crowd situation acquisition method according to claim 2, characterized in that: Determining the first area range and the second area range based on the statistical result includes: Determine a first unit area and a second unit area from the multiple unit areas based on the statistical result, wherein the number of people in the first unit area is continuously empty in some or all consecutive time segments of the multiple time segments, and the second unit area is the remaining unit areas of the multiple unit areas except the first unit area; The first area range is determined using the first unit area, and the second area range is determined using the second unit area.
4. The method for acquiring crowd situation according to claim 1, characterized in that: Acquiring the spatiotemporal perception information associated with the first area range includes: estimating an initial movement direction of a crowd within the second area within a second preset time range, wherein the second preset time range is a preset historical time range; The movement trajectory of the crowd within a third preset time range is predicted using the initial movement direction to obtain the spatiotemporal perception information, wherein the third preset time range is a preset future time range and the movement trajectory is a future movement trajectory.
5. The method for acquiring crowd situation according to claim 4, characterized in that: Estimating the initial movement direction of the crowd located in the second area within the second preset time range includes: Analyzing the plurality of first positions using a first neural network model to determine a second position, wherein the first neural network model is obtained through machine learning training using multiple sets of data, each set of the multiple sets of data including: pedestrian positions and crowd flow data at multiple sampling moments, the plurality of first positions being a plurality of consecutive historical positions corresponding to multiple historical moments of each pedestrian in a crowd located within the second area, and the second position being a position of each pedestrian in the crowd located within the second area at a moment immediately following the multiple historical moments; The initial movement direction is determined using a third position and the second position, wherein the third position is the first position corresponding to a previous moment adjacent to the second position.
6. The method for acquiring crowd situation according to claim 4, characterized in that: Predicting the movement trajectory of the crowd within the third preset time range using the initial movement direction to obtain the spatiotemporal perception information includes: Analyzing the plurality of fourth positions using a second neural network model to determine a fifth position, wherein the second neural network model is obtained through machine learning training using multiple sets of data, each set of the multiple sets of data including: pedestrian positions and crowd flow data at multiple sampling moments before a preset moment, the plurality of fourth positions being a plurality of consecutive historical positions corresponding to multiple historical moments before a target moment for each pedestrian in the crowd within the second area, and the fifth position being the position of each pedestrian in the crowd within the second area at a next moment after the target moment; The spatiotemporal perception information is obtained by using the multiple fifth positions gradually deduced through the second neural network model.
7. A method for acquiring crowd situation, characterized in that: include: Receiving a video capture image from a client, wherein the content displayed in the video capture image includes: a target video capture area range; Based on the video capture image, a first area range and a second area range are determined within the target video capture area range, spatiotemporal perception information associated with the first area range is obtained, and a first part of the population and a second part of the population are obtained based on the spatiotemporal perception information, wherein the first part of the population is a population entering the first area range from the second area range, and the second part of the population is a population entering the second area range from the first area range; the first part of the population and the second part of the population are used to predict the first crowd situation in the first area range, and the first crowd situation and the second crowd situation in the second area range are used to obtain the target crowd situation in the target video capture area range, wherein the first area range is a blind area range, and the second area range is a non-blind area range; Feedback the target population situation to the client.
8. A crowd situation acquisition device, characterized in that: include: A determination module, configured to determine a first area range and a second area range within a target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; A first acquisition module is used to acquire spatiotemporal perception information associated with the first area; a prediction module, configured to obtain a first portion of the crowd and a second portion of the crowd based on the spatiotemporal perception information, wherein the first portion of the crowd is a crowd entering the first area from the second area, and the second portion of the crowd is a crowd entering the second area from the first area; and predict the situation of the first crowd in the first area using the first portion of the crowd and the second portion of the crowd; The second acquisition module is used to acquire the target crowd situation within the target video acquisition area by using the first crowd situation and the second crowd situation within the second area.
9. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the crowd situation acquisition method according to any one of claims 1 to 7.
10. A crowd situation acquisition system, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Step 1: determining a first area range and a second area range within the target video acquisition area range, wherein the first area range is a blind area range and the second area range is a non-blind area range; Step 2: Acquire spatiotemporal perception information associated with the first area; Step 3: Acquire the first part of the population and the second part of the population based on the spatiotemporal perception information, wherein: The first part of the crowd is the crowd entering the first area from the second area, and the second part of the crowd is the crowd entering the second area from the first area; the first part of the crowd and the second part of the crowd are used to predict the situation of the first crowd in the first area; Step 4: Using the first crowd situation and the second crowd situation in the second area, obtain the target crowd situation in the target video acquisition area.
Citation Information
Patent Citations
Monitoring blind area crowd state deduction method based on Bayesian network
CN103530601A
Crowd trajectory prediction method based on multi-precision interaction
CN113362367A