Monitoring method and system for spatiotemporal contact behavior sequences of people in urban public spaces

By constructing a three-dimensional map model of urban public spaces and identifying pedestrian information sequences, the problem of insufficient adaptability of existing technologies to three-dimensional spatial environments is solved, and high-precision and efficient pedestrian contact behavior detection and backtracking are achieved.

CN116912877BActive Publication Date: 2025-09-19SHENZHEN UNIV
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202310500567.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-06
Publication Date
2025-09-19
Estimated Expiration
2043-05-06

Smart Images

  • Figure CN116912877B_ABST
    Figure CN116912877B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for monitoring the spatiotemporal contact behavior sequences of people in urban public spaces. The method includes constructing a three-dimensional map model based on point cloud data of urban public spaces, identifying each video frame image in a plurality of video sequences to form a pedestrian information sequence; mapping the position information in the pedestrian information to the three-dimensional map model to form a plurality of pedestrian trajectories; and extracting the spatiotemporal contact behavior of people in the plurality of pedestrian trajectories based on a preset social distance. The present application obtains a pedestrian information sequence by extracting the video sequence, then projects the pedestrian information sequence to a three-dimensional map model to form a three-dimensional pedestrian trajectory, and finally forms a spatiotemporal contact behavior sequence of people based on the pedestrian trajectories. In this way, pedestrian contact behavior is detected from the spatiotemporal dimension, which not only improves the accuracy of pedestrian contact behavior recognition, but also can well adapt to the current rapidly developing three-dimensional urban space environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image analysis and target positioning, and in particular to a method and system for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces. Background Art

[0002] As surveillance cameras increase in density and coverage across urban public spaces, the volume of unstructured video data is surging. Traditional video analysis methods rely primarily on manual, continuous observation of surveillance footage, requiring significant time, human resources, and data storage space. This approach is unsuitable for addressing the challenges of investigating major infectious diseases, public safety emergencies, and criminal investigations.

[0003] Automated monitoring of the spatiotemporal contact behaviors of people in urban public spaces is therefore crucial for addressing major infectious diseases, public safety emergencies, and criminal investigations. Intelligent video analysis using technologies like artificial intelligence and neural networks is particularly crucial for improving the efficiency and accuracy of analyzing massive amounts of video. However, existing automated monitoring technologies primarily focus on visually detecting pedestrians' two-dimensional coordinates and lack the ability to retrieve the spatiotemporal information about direct or indirect contact behaviors, making them difficult to adapt to the rapidly evolving three-dimensional spatial environment of cities.

[0004] Therefore existing technology still needs to be improved and improved. Summary of the Invention

[0005] The technical problem to be solved by this application is to provide a method and system for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces in response to the shortcomings of existing technologies.

[0006] In order to solve the above technical problems, the first aspect of the embodiments of the present application provides a method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, the method comprising:

[0007] Acquiring point cloud data of an urban public space, and determining a three-dimensional map model corresponding to the urban public space based on the point cloud data;

[0008] Acquire multiple video sequences of the urban public space, and identify pedestrian information in each video frame image in each video sequence to form a pedestrian information sequence, wherein the pedestrian information includes identity information and location information;

[0009] Mapping the position information of each pedestrian information to a three-dimensional map model to obtain the three-dimensional position information corresponding to each position information, and sorting the three-dimensional position information corresponding to the same identity information in chronological order to obtain a number of pedestrian trajectories;

[0010] The spatiotemporal contact behaviors of several pedestrian trajectories are extracted based on the preset social distance, and a spatiotemporal contact behavior sequence of the crowd is formed according to the extracted spatiotemporal contact behaviors of the crowd.

[0011] The method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, wherein determining the three-dimensional map model corresponding to the urban public space based on the point cloud data specifically includes:

[0012] Constructing a three-dimensional map coordinate system corresponding to the urban public space, and determining a first conversion model from the point cloud coordinate system corresponding to the point cloud data to the three-dimensional map coordinate system;

[0013] For each point cloud data point in the point cloud data, determining a neighborhood of the point cloud data point, determining a projected tangent plane between the point cloud data point and each neighboring point in the neighborhood based on a first transformation model, and projecting each neighboring point onto the projected tangent plane;

[0014] The projected tangent plane corresponding to each point cloud data trajectory point is triangulated to obtain a three-dimensional map model.

[0015] The method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, wherein, for each point cloud data point in the point cloud data, determining the neighborhood of the point cloud data point specifically comprises:

[0016] downsampling the point cloud data using a voxel filtering algorithm to obtain target point cloud data;

[0017] For each point cloud data point in the target point cloud data, a KD tree nearest neighbor search method is used to determine the neighborhood corresponding to the point cloud data trajectory point, wherein the neighborhood includes a preset number of neighborhood points.

[0018] The method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, wherein the mapping of the position information in each pedestrian information to a three-dimensional map model to obtain the three-dimensional position information corresponding to each position information specifically includes:

[0019] Obtaining a second conversion model between a pixel coordinate system corresponding to each video sequence and a three-dimensional map coordinate system corresponding to the three-dimensional map model;

[0020] For each pedestrian information, converting the position information in the pedestrian information into a three-dimensional map coordinate system based on a second conversion model corresponding to the pixel coordinate system of the pedestrian information to obtain the sole position information corresponding to the position information;

[0021] The height information corresponding to the pedestrian information is determined, and three-dimensional coordinate information composed of the sole position information and the height information is used as the three-dimensional position information corresponding to the position information.

[0022] The method for monitoring the spatiotemporal contact behavior sequence of crowds in urban public spaces, wherein the spatiotemporal contact behavior of crowds based on the extraction of a plurality of pedestrian trajectories based on a preset social distance specifically includes:

[0023] Calculate the discrete Fréchet distance between two pedestrian trajectories in a number of pedestrian trajectories to obtain the social distance between two pedestrians;

[0024] For each pedestrian, a target social distance greater than a preset social distance is selected from all social distances corresponding to the pedestrian, and the contact pedestrian, contact time, and contact position corresponding to the pedestrian are determined based on the trajectory point pairs corresponding to the obtained target social distance;

[0025] For each contact pedestrian of the pedestrian, determining a trajectory line connecting any first trajectory point in the pedestrian trajectory of the pedestrian and any second trajectory point in the pedestrian trajectory of the contact pedestrian, and detecting whether the trajectory line intersects a triangular facet in the three-dimensional map model to determine obstacle information between the pedestrian and the contact pedestrian;

[0026] The contact pedestrian, contact time, contact position, social distance and obstacle information corresponding to the pedestrian are used as a crowd spatiotemporal contact behavior to obtain the crowd spatiotemporal contact behaviors corresponding to several pedestrian trajectories.

[0027] The method for monitoring the spatiotemporal contact behavior sequence of crowds in urban public spaces, wherein the spatiotemporal contact behavior of crowds includes contact with pedestrians, contact time, contact location, social distance, and obstacle information; and forming the spatiotemporal contact behavior sequence of crowds based on the extracted spatiotemporal contact behavior of crowds specifically includes:

[0028] The pedestrian corresponding to each pedestrian trajectory is used as a node, wherein the node stores identity information, pedestrian trajectory and timestamp information;

[0029] For each pedestrian, a connection edge is constructed between the pedestrian and the pedestrian contacting the pedestrian to form a crowd spatiotemporal contact behavior sequence, wherein the crowd spatiotemporal contact behavior is stored on the connection edge.

[0030] The method for monitoring the spatiotemporal contact behavior sequence of a crowd in an urban public space, wherein, after forming the spatiotemporal contact behavior sequence of a crowd based on the extracted spatiotemporal contact behavior of the crowd, the method further comprises:

[0031] Determine backtracking constraints;

[0032] Based on the backtracking constraint condition, the temporal and spatial contact behavior sequence of the crowd is queried by a breadth-first algorithm to obtain the backtracking pedestrian group corresponding to the target pedestrian.

[0033] A second aspect of the present application provides a system for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, the system comprising:

[0034] an acquisition module, configured to acquire point cloud data of an urban public space and determine a three-dimensional map model corresponding to the urban public space based on the point cloud data;

[0035] an identification module, configured to obtain a plurality of video sequences of the urban public space and identify pedestrian information in each video frame image in each video sequence to form a pedestrian information sequence, wherein the pedestrian information sequence includes a plurality of pedestrian information, each pedestrian information including identity information and location information;

[0036] A mapping module is used to map the position information in each pedestrian information to a three-dimensional map model to obtain the three-dimensional position information corresponding to each position information, and to sort the three-dimensional position information corresponding to the same identity information in chronological order to obtain a number of pedestrian trajectories;

[0037] The extraction module is used to extract the spatiotemporal contact behaviors of a number of pedestrian trajectories based on a preset social distance, and form a spatiotemporal contact behavior sequence of a crowd according to the extracted spatiotemporal contact behaviors of the crowd.

[0038] A third aspect of an embodiment of the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces as described above.

[0039] A fourth aspect of an embodiment of the present application provides a terminal device, comprising: a processor, a memory, and a communication bus; the memory stores a computer-readable program that can be executed by the processor;

[0040] The communication bus realizes the connection and communication between the processor and the memory;

[0041] When the processor executes the computer-readable program, the processor implements the steps of any of the above-described methods for monitoring the spatiotemporal contact behavior sequences of people in urban public spaces.

[0042] Beneficial effects: Compared with the prior art, the present application provides a method and system for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, the method comprising obtaining point cloud data of an urban public space, and determining a three-dimensional map model corresponding to the urban public space based on the point cloud data; obtaining several video sequences of the urban public space, and identifying pedestrian information of each video frame image in each video sequence to form a pedestrian information sequence; mapping the position information in each pedestrian information to a three-dimensional map model to obtain three-dimensional position information corresponding to each position information, and sorting the three-dimensional position information corresponding to the same identity information in chronological order to obtain several pedestrian trajectories; extracting the spatiotemporal contact behavior of the crowd of several pedestrian trajectories based on a preset social distance, and forming a spatiotemporal contact behavior sequence of the crowd according to the extracted spatiotemporal contact behavior of the crowd. This application extracts several video sequences to obtain pedestrian information sequences, then projects the pedestrian information sequences onto a three-dimensional map model to form three-dimensional pedestrian trajectories, and finally forms a crowd spatiotemporal contact behavior sequence based on the three-dimensional pedestrian trajectories. In this way, pedestrian contact behavior is detected from the spatiotemporal dimension, which can not only improve the accuracy of pedestrian contact behavior recognition, but also adapt well to the rapidly developing three-dimensional urban spatial environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without inventive work.

[0044] Figure 1 This is a flow chart of the method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces provided in this application.

[0045] Figure 2 This is an example diagram of the method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces provided in this application.

[0046] Figure 3 for Figure 2 Schematic diagram of the spatiotemporal contact behavior sequence of a crowd in .

[0047] Figure 4 for Figure 2 Schematic diagram of the backtrace results in .

[0048] Figure 5 is the local projection map of the neighborhood points in the neighborhood.

[0049] Figure 6 Flowchart for greedy projection triangulation.

[0050] Figure 7Schematic diagram of pedestrian height calculation.

[0051] Figure 8 This is a schematic diagram of the structural principles of the monitoring system for the spatiotemporal contact behavior sequence of people in urban public spaces provided in this application.

[0052] Figure 9 This is a schematic diagram of the structure of the terminal device provided in this application. DETAILED DESCRIPTION

[0053] This application provides a method and system for monitoring the temporal and spatial interaction behavior sequences of people in urban public spaces. To clarify and clarify the purpose, technical solutions, and effects of this application, the following further describes this application in detail with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are intended only to illustrate this application and are not intended to limit it.

[0054] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0055] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0056] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.

[0057] Research has found that with the increasing density and coverage of surveillance cameras in urban public spaces, the amount of unstructured video data has surged. Traditional video analysis methods rely primarily on continuous manual observation of surveillance video, which consumes significant time, human resources, and data storage space, making them unsuitable for applications such as investigating major infectious diseases, public safety emergencies, and criminal investigations. Therefore, automated monitoring of the spatiotemporal contact behaviors of people in urban public spaces is crucial for these areas. The application of technologies such as artificial intelligence and neural networks for intelligent video analysis is particularly important for improving the efficiency and accuracy of analyzing massive amounts of video.

[0058] Existing automated monitoring technologies primarily focus on visual detection of pedestrian coordinates in two dimensions, such as pedestrian target recognition, posture detection, and social distance detection. Social distance detection, a crucial component of pedestrian contact behavior, includes methods such as motion vector analysis for safe distance risk assessment and deep CNN-based social distance monitoring models based on bird's-eye view analysis. However, existing research primarily calculates social distance by directly calculating the pixel distance between pedestrians, failing to detect three-dimensional social distance and hindering its application in complex three-dimensional urban environments. When utilizing multiple surveillance cameras in public areas for coordinated monitoring, it is necessary to consider the three-dimensional structure of the scene, both indoors and outdoors, as well as the spatial transformation relationship between the three-dimensional model and the multi-pixel coordinate system. In indoor scenes, building structure and camera parameters significantly impact the detection of crowd contact behavior. Existing technologies lack research on the spatial transformation model between the three-dimensional model and the multi-pixel coordinate system, limiting them to two-dimensional social distance detection. This can lead to misjudgments of pedestrian contacts obstructed by building structures and lacks the ability to detect and track contact behavior in spatiotemporal space.

[0059] In order to solve the above problems, in an embodiment of the present application, point cloud data of an urban public space is obtained, and a three-dimensional map model corresponding to the urban public space is determined based on the point cloud data; several video sequences of the urban public space are obtained, and pedestrian information of each video frame image in each video sequence is identified to form a pedestrian information sequence; the position information in each pedestrian information is mapped to the three-dimensional map model to obtain the three-dimensional position information corresponding to each position information, and the three-dimensional position information corresponding to the same identity information is sorted in chronological order to obtain several pedestrian trajectories; the spatiotemporal contact behavior of the crowd of the several pedestrian trajectories is extracted based on the preset social distance, and a crowd spatiotemporal contact behavior sequence is formed based on the extracted crowd spatiotemporal contact behavior. The present application obtains a pedestrian information sequence by extracting several video sequences, and then projects the pedestrian information sequence to a three-dimensional map model to form a three-dimensional pedestrian trajectory, and finally forms a crowd spatiotemporal contact behavior sequence based on the three-dimensional pedestrian trajectory. In this way, pedestrian contact behavior is detected from the spatiotemporal dimension, which can not only improve the accuracy of pedestrian contact behavior recognition, but also adapt well to the current rapidly developing three-dimensional urban space environment. At the same time, after obtaining the crowd's spatiotemporal contact behavior sequence, the crowd's spatiotemporal contact behavior sequence can be used to trace the contact behavior between pedestrians, which can effectively reduce the cost of video analysis and improve the efficiency and accuracy of video analysis.

[0060] The application content will be further explained below through description of embodiments in conjunction with the accompanying drawings.

[0061] This embodiment provides a method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces. Figures 1-4 As shown, the method includes:

[0062] S10: Acquire point cloud data of an urban public space, and determine a three-dimensional map model corresponding to the urban public space based on the point cloud data.

[0063] Specifically, the point cloud data is obtained by collecting point clouds of urban public spaces, for example, by using radar lasers or drones, etc. The urban public spaces may be large urban public spaces with dense traffic, such as stations, hospitals, and airports.

[0064] 3D map models are used to reflect the spatial information of urban public spaces. The coordinate system of the point cloud data is the coordinate system of the acquisition device corresponding to the point cloud data, and the 3D map coordinate system corresponding to the 3D map model must match the urban public space. Therefore, when determining a 3D map model based on point cloud data, a 3D map coordinate system for the urban public space can be first constructed. Then, a first transformation model can be constructed between the point cloud coordinate system corresponding to the point cloud data and the 3D map coordinate system. Based on the first transformation model, each point in the point cloud data is transformed into the 3D map coordinate system to obtain the 3D map model.

[0065] In one implementation, determining the three-dimensional map model corresponding to the urban public space based on the point cloud data specifically includes:

[0066] S11, constructing a three-dimensional map coordinate system corresponding to the urban public space, and determining a first conversion model from the point cloud coordinate system corresponding to the point cloud data to the three-dimensional map coordinate system;

[0067] S12. For each point cloud data point in the point cloud data, determining a neighborhood of the point cloud data point;

[0068] S13, determining a projection tangent plane between the point cloud data point and each neighborhood point in the neighborhood based on the first conversion model, and projecting each neighborhood point onto the projection tangent plane;

[0069] S14. Triangulate the projected tangent plane corresponding to each point cloud data trajectory point to obtain a three-dimensional map model.

[0070] Specifically, the three-dimensional map coordinate system is a three-dimensional coordinate system of an urban public space, wherein the coordinate origin of the three-dimensional map coordinate system is a scene feature point in the urban public space, and the plane where the coordinate axis is located can coincide with or be parallel to the rigid structure of the urban public space. In this embodiment, the process of constructing the three-dimensional map coordinate system can be: according to preset conditions, at least four scene feature points are selected as reference points in the urban public space, one of the selected reference points is used as the coordinate origin, and the plane where the coordinate axis is located is made to coincide with or be parallel to the rigid structure of the urban public space to establish the three-dimensional map coordinate system, wherein the preset conditions can be that the selected reference points meet the difference requirements and repeatability requirements, and the difference requirement refers to that the reference point is a significant point in the urban public space, such as a corner point, an edge point, etc. The repeatability requirement refers to that the same reference point appears repeatedly in different perspectives and is rotationally, photometrically, and scale invariant.

[0071] Furthermore, after determining the three-dimensional map coordinate system, the position coordinates of each reference point that is not used as the coordinate origin are determined, and then a first conversion model for converting the point cloud coordinate system to the three-dimensional map coordinate system is determined based on the position coordinates of each reference point and the position coordinates of the reference point cloud data points corresponding to each reference trajectory point in the point cloud data, wherein the first conversion model includes rotation parameters and translation parameters, the rotation parameters are used to reflect the posture conversion relationship from the point cloud coordinate system to the three-dimensional map coordinate system, and the translation parameters are used to reflect the distance from the origin of the three-dimensional map coordinate system to the origin of the point cloud coordinate system on the X, Y, and Z axes.

[0072] In this embodiment, the first conversion model can be expressed as:

[0073] ;

[0074] ;

[0075] ;

[0076] ;

[0077] ;

[0078] Among them, the scaling factor Used to describe the scale conversion relationship between the point cloud coordinate system and the three-dimensional map coordinate system; translation matrix Used to describe the distance between the origin of the 3D map coordinate system and the origin of the point cloud coordinate system on the X, Y, and Z axes; rotation matrix Used to describe the posture transformation relationship from the 3D map coordinate system to the point cloud coordinate system. The rotation value of the 3D map coordinate system transformed to the point cloud coordinate system with the x-axis as the rotation axis. The rotation value of the 3D map coordinate system transformed to the point cloud coordinate system with the y-axis as the rotation axis. The rotation value of the 3D map coordinate system transformed to the point cloud coordinate system with the z-axis as the rotation axis.

[0079] In one implementation, for each point cloud data point in the point cloud data, determining the neighborhood of the point cloud data point specifically includes:

[0080] downsampling the point cloud data using a voxel filtering algorithm to obtain target point cloud data;

[0081] For each point cloud data point in the target point cloud data, a KD tree nearest neighbor search method is used to determine the neighborhood corresponding to the point cloud data trajectory point, wherein the neighborhood includes a preset number of neighborhood points.

[0082] Specifically, because point cloud data contains a large number of point cloud data points, using all of them to construct a 3D map model would be time-consuming. Therefore, after determining the first conversion model, the point cloud data can be filtered and used to construct the 3D map model, thereby increasing the speed of 3D map model construction. In this embodiment, a pixel filtering algorithm is used to downsample the point cloud data, significantly reducing the number of point clouds while preserving the scene's structural features. This improves the speed of 3D map model construction while ensuring the accuracy of the constructed 3D map model.

[0083] After obtaining the target point cloud data, a KD tree-based nearest neighbor search algorithm can be used to search the neighborhood of each point cloud data point. The neighborhood of each point cloud data point includes a preset number of neighboring points. In other words, the neighborhood of each point cloud data point includes the same number of neighboring points. For example, a KD tree-based nearest neighbor search algorithm can be used to determine a neighborhood containing k nearest neighbor points.

[0084] Furthermore, after obtaining the neighborhood of each point cloud data point, the point cloud data point and the neighboring points in its neighborhood are converted into a three-dimensional map coordinate system through the first transformation model, and then a two-dimensional tangent plane between the point cloud data point and the neighboring points in its neighborhood is determined, and the two-dimensional tangent plane is used as a projection plane, and finally each neighboring point is projected onto the projected tangent plane. The process of determining the projected tangent plane can be: since the three-dimensional neighborhood points in the neighborhood of the point cloud data point are in the three-dimensional map model and the point cloud data point is in the neighborhood, Corresponding 3D data points There is a one-to-one correspondence between the projection tangent planes, so the PCA normal estimation method of the greedy projection algorithm can be used to calculate the approximate normal vector of each 3D data point, and the projection tangent plane of the 3D data point can be solved by the plane equation. For example, assuming the 3D data point Normal vector ,point Pass The projected tangent plane can be expressed as:

[0085] (1).

[0086] like Figure 5 As shown in the figure, after obtaining the projection tangent plane, the projection matrix method can be used to project the point projection value in the space onto the tangent plane, and the three-dimensional point on the projection tangent plane can be obtained through a series of operations such as translation and rotation. The projection on , where the projection matrix The formula is:

[0087] (2)

[0088] in, is the translation transformation matrix, To surround Axis rotation Spend, To surround Axis rotation Degrees, in the following form:

[0089] ; ; .

[0090] Based on formulas (1) and (2), any point On its projected tangent plane Projection on:

[0091] .

[0092] Furthermore, after projecting the neighborhood points of each point cloud data point onto the projection tangent plane, each projection tangent plane is triangulated. The triangulation can be performed using a greedy algorithm, and the local optimal point is selected as the extension point each time. The extended point is then mapped back to space through the projection relationship, and this process is repeated to obtain a three-dimensional map model.

[0093] In one implementation, Figure 6 As shown, the process of triangulation by greedy algorithm can be:

[0094] a1. Select any point in the point cloud data As the initial growth point, use the KD tree algorithm to search for its neighboring points and find the distance within its neighborhood. The nearest point And use straight lines to connect them to get line segments , and use it as a growing side of the triangle;

[0095] a2. Calculate the distance edge The nearest point , then construct the first triangle patch with spatial position information ;

[0096] a3. Based on step a2, continue to find the distance triangle The nearest point on any side forms a new spatial triangle;

[0097] a4. Repeat step a2 until all points are selected and a complete topological structure is constructed.

[0098] S20: Acquire several video sequences of the urban public space, and identify pedestrian information in each video frame image in each video sequence to form a pedestrian information sequence.

[0099] Specifically, the multiple video sequences are all captured by image acquisition devices installed in urban public spaces. For example, the urban public space is equipped with multiple surveillance cameras, and the multiple surveillance cameras correspond one-to-one to the multiple video sequences. Each video sequence is formed by video frames captured by its corresponding surveillance camera during a preset time period, and each surveillance camera corresponds to the same preset time period. In other words, the collection time period corresponding to each video sequence in the multiple video sequences is the same. For example, the collection time period corresponding to the multiple video sequences is the video sequence collected between 10:00 and 11:00 on May 5, 2023.

[0100] The pedestrian information includes identity information and location information, wherein the identity information is the identity characteristics of the pedestrian, and the location information is used to reflect the location of the pedestrian in the public space of the city. Each video frame image in each video sequence can include multiple pedestrians, so the pedestrian information corresponding to each video frame image can be obtained by performing pedestrian detection on each video frame image in each video sequence. After the pedestrian information corresponding to each video frame image is obtained, the sequence consisting of all the obtained pedestrian information is used to form a pedestrian information sequence. In addition, since different video frames in the same video sequence or video frames in different video sequences can carry the same pedestrian, and video frames in different video sequences can contain the same pedestrian at the same acquisition time, in order to facilitate the subsequent determination of pedestrian trajectories, when determining pedestrian information, the timestamp corresponding to the video frame image can be used as the timestamp of the pedestrian information determined based on the video frame image, so that the pedestrian information can be merged and sorted based on the timestamp to form a pedestrian trajectory.

[0101] In one implementation, the pedestrian information can be obtained by detection based on a deep learning pedestrian detection model, wherein the detection process based on the deep learning pedestrian detection model can be: first use the FastReID model to extract the identity feature information of the pedestrian and store it in the database, use the YOLOv5 model to perform pedestrian target detection on the video sequence, combine the DeepSORT algorithm to realize pedestrian tracking processing, and replace the representation extraction model in DeepSORT with the FastReID feature extraction model, match the same pedestrian captured by different cameras or at different times, complete the re-identification of the pedestrian, and then detect the pedestrian and obtain the position information of the pedestrian in the video frame image.

[0102] S30 , mapping the position information in each pedestrian information to a three-dimensional map model to obtain three-dimensional position information corresponding to each position information, and sorting the three-dimensional position information corresponding to the same identity information in chronological order to obtain a number of pedestrian trajectories.

[0103] Specifically, the three-dimensional position information refers to the position of the pedestrian in the three-dimensional map model, where the three-dimensional position information includes top-of-head position information, which in turn includes sole-of-foot position information and the pedestrian's height information. It is understood that, based on a second conversion model between the pixel coordinate system corresponding to the position information and the three-dimensional map coordinate system, the position information can be converted to the three-dimensional map coordinate system to obtain sole-of-foot position information. The height information is then determined based on the pedestrian's height, and finally, the top-of-head position information is determined based on the height information and sole-of-foot position information.

[0104] In one implementation, mapping the position information in each pedestrian information to the three-dimensional map model to obtain the three-dimensional position information corresponding to each position information specifically includes:

[0105] S31, obtaining a second conversion model between a pixel coordinate system corresponding to each video sequence and a three-dimensional map coordinate system corresponding to the three-dimensional map model;

[0106] S32. For each piece of pedestrian information, convert the position information in the pedestrian information into a three-dimensional map coordinate system based on a second conversion model corresponding to the pixel coordinate system of the pedestrian information to obtain the sole position information corresponding to the position information;

[0107] S33. Determine the height information corresponding to the pedestrian information, and use the three-dimensional coordinate information composed of the sole position information and the height information as the three-dimensional position information corresponding to the position information.

[0108] Specifically, for the image acquisition device corresponding to each video sequence in a plurality of video sequences, a second conversion model can be established between the pixel coordinate system corresponding to each image acquisition device and the three-dimensional map coordinate system. The second conversion model can be used to convert the coordinate points in the pixel coordinate system to the three-dimensional map coordinate system. In this embodiment, the second conversion model can be determined by first obtaining the intrinsic camera parameters of the acquisition device, where the intrinsic camera parameters may include the camera focal length. , the physical size of each pixel of the image and , the image origin coordinates of the camera optical center and . Secondly, in the 3D map coordinate system At least four reference calibration points are selected, the camera extrinsic parameters are calculated based on the selected reference calibration points, and the pose information of the acquisition device in the three-dimensional map model space is determined based on the camera extrinsic parameters, wherein the pose information includes the rotation matrix and translation vectors Finally, the second transformation model from the pixel coordinate system to the 3D map coordinate system is established by combining the camera intrinsic parameters and posture information.

[0109] In this embodiment, the second conversion model can be expressed as:

[0110] ;

[0111] ;

[0112] ;

[0113] ;

[0114] Among them, the transformation matrix in and Implement rigid body transformations, including scaling, rotation, and translation, Implement perspective transformation, Achieve full scale transformation.

[0115] The position information may include the sole position and the top position of the head. After obtaining the second conversion model corresponding to each video sequence, the sole position can be converted into the The second conversion model corresponding to the input pedestrian information can obtain the foot position information Then, if Figure 7 As shown, the height from the acquisition device to the ground and the horizontal distance from the camera to the pedestrian Calculate the straight-line distance from the camera to the pedestrian , the distance Camera focal length The actual height of the pedestrian is obtained by multiplying the ratio of the height difference of the pedestrian in the image frame. , considering the height of the floor , add up to get the height of the pedestrian's head in the 3D map model , combined with the plantar position information The three-dimensional coordinate information of the pedestrian's head can be obtained .in, , and The calculation formulas are: ; ; ;in, Indicates the pixel coordinate in the y direction of the top of the head. Indicates the pixel coordinate in the y direction of the sole position.

[0116] In one implementation, to improve the integrity of pedestrian trajectories, after determining the pedestrian trajectory based on pedestrian information, an interpolation method can be used to weight the pedestrian coordinate information of the time period of adjacent timestamps detected and allocated to the time period, thereby filling in the missing pedestrian coordinate information and obtaining the complete pedestrian trajectory. The identity information and the trajectory information of the pedestrian trajectory are stored in the database. The calculation formula of the interpolated position information is as follows:

[0117] ;

[0118] ;

[0119] in, is the time interval between video frames, is the location of the pedestrian at the last timestamp, is the position of the next timestamp, It is Pedestrian coordinates to be completed, is the time point when the pedestrian was last detected. It's the next time point.

[0120] S40: extracting the spatiotemporal contact behaviors of a number of pedestrian trajectories based on a preset social distance, and forming a spatiotemporal contact behavior sequence of a crowd according to the extracted spatiotemporal contact behaviors of the crowd.

[0121] Specifically, the spatiotemporal contact behavior of a crowd includes at least the contact with pedestrians, contact time, contact location and social distance. Based on the spatiotemporal contact behavior of a crowd, the time, location, person and contact distance of the contact event with the pedestrian can be determined. In actual applications, there are many buildings in urban public spaces. Although the social distance between two pedestrians is very small, there may be obstacles between the two pedestrians. Therefore, when determining the spatiotemporal contact behavior of a crowd, it is also possible to detect whether there is an obstacle between the two pedestrians, that is, the spatiotemporal contact behavior of a crowd can also include obstacle information. Accordingly, the spatiotemporal contact behavior of a crowd can be expressed as ,in, Indicates the identity information of the pedestrian in contact. Indicates contact time, Indicates the contact position, Indicates social distance, Indicates obstacle information.

[0122] In one implementation, the extraction of a plurality of pedestrian trajectories based on a preset social distance specifically includes:

[0123] S41. Calculate the discrete Fréchet distance between two pedestrian trajectories in a plurality of pedestrian trajectories to obtain the social distance between the two pedestrians;

[0124] S42: For each pedestrian, select a target social distance greater than a preset social distance from all social distances corresponding to the pedestrian, and determine the pedestrian contact, contact time, and contact location corresponding to the pedestrian based on the obtained trajectory point pairs corresponding to the target social distance;

[0125] S43. For each contact pedestrian, determine a trajectory line connecting any first trajectory point in the pedestrian's trajectory and any second trajectory point in the contact pedestrian's trajectory, and detect whether the trajectory line intersects a triangular facet in the three-dimensional map model to determine obstacle information between the pedestrian and the contact pedestrian.

[0126] S44. The contact pedestrian, contact time, contact position, social distance, and obstacle information corresponding to the pedestrian are used as a crowd spatiotemporal contact behavior to obtain crowd spatiotemporal contact behaviors corresponding to several pedestrian trajectories.

[0127] Specifically, in step S41, each pedestrian trajectory in the plurality of pedestrian trajectories corresponds to a pedestrian. The discrete Fréchet distance between two pedestrian trajectories in the plurality of pedestrian trajectories can reflect the social distance between the two pedestrians corresponding to the two pedestrian trajectories. Thus, by calculating the discrete Fréchet distance between each pair of pedestrian trajectories in the plurality of pedestrian trajectories, the social distance between each pair of pedestrians in the plurality of pedestrian trajectories can be obtained. In this embodiment, the discrete Fréchet distance determination process can be as follows: first, the first pedestrian trajectory and the second pedestrian trajectory in the two pedestrian trajectories are sequentially combined into a trajectory point pair sequence, wherein the trajectory point pair sequence includes a plurality of trajectory point pair groups, each of which corresponds one-to-one to a trajectory point a included in the first pedestrian trajectory. Each trajectory point pair group is formed based on its corresponding trajectory point a and all trajectory points b in the second pedestrian trajectory. Each trajectory point pair group includes a plurality of trajectory point pairs, each of which corresponds one-to-one to a trajectory point b included in the second pedestrian trajectory. Each trajectory point pair includes the trajectory point a corresponding to its trajectory point pair group and the trajectory point b corresponding to the trajectory point pair. Next, the Euclidean distances between the two trajectory points in each trajectory point pair are calculated to obtain a distance subsequence. The average of the Euclidean distances in the distance subsequence is used as the distance corresponding to the trajectory point pair, forming a distance sequence. Finally, the distance subsequence corresponding to the maximum distance in the distance sequence is selected, and the minimum distance in the selected distance subsequence is used as the discrete Fréchet distance. That is, the minimum distance in the selected distance subsequence is used as the social distance between the pedestrian corresponding to the first pedestrian trajectory and the pedestrian corresponding to the second pedestrian trajectory.

[0128] For example, suppose two of the pedestrian trajectories are and , then the trajectory point sequence = = ; Calculate each trajectory point pair group in the trajectory point pair sequence Each trajectory point pair Euclidean distance ,in, , Pedestrian trajectories The corresponding pedestrian trajectory point coordinates, human trajectory The corresponding pedestrian's trajectory point coordinates are based on The Euclidean distance of each trajectory point pair in determines its corresponding distance sequence , and then calculate the average value of each distance subsequence to get the distance sequence , then in the distance sequence Select the distance subsequence with the largest distance , and finally calculate the subsequence The minimum Euclidean distance in is taken as the social distance between two pedestrians, where the distance can be expressed as , Pedestrian trajectory The corresponding pedestrians, Pedestrian trajectory The corresponding pedestrians.

[0129] Furthermore, in step S42, after obtaining the social distance between two pedestrians, each pedestrian will obtain the social distance between the pedestrian and all pedestrians except himself. Based on this, for each pedestrian, it is recorded as an initial pedestrian; the social distance between the initial pedestrian and other pedestrians is compared with the preset social distance respectively to obtain a target social distance that is less than the preset social distance, wherein the target social distance can be 0, 1, or multiple. Here, obtaining the target social distance is used as an example for explanation. After obtaining the target social distance, the target pedestrian corresponding to the target social distance can be determined, and then the target pedestrian is regarded as a pedestrian in spatiotemporal contact with the initial pedestrian, the identity information of the target pedestrian is regarded as the contact pedestrian, the time pair consisting of the timestamp of the trajectory point in the initial pedestrian corresponding to the target social distance and the timestamp in the pedestrian trajectory of the target pedestrian is regarded as the contact time, the coordinate information of the trajectory point in the initial pedestrian corresponding to the target social distance is regarded as the contact position, or the coordinate information pair consisting of the coordinate information of the trajectory point in the initial pedestrian and the coordinate information in the pedestrian trajectory of the target pedestrian is regarded as the contact position, etc.

[0130] Furthermore, in step S43, after obtaining the contact pedestrians corresponding to each pedestrian, in order to accurately determine the temporal and spatial contact patterns of the crowd, the presence of obstacles between the pedestrian and the contact pedestrian is detected to obtain obstacle information between the pedestrian and the contact pedestrian. In this embodiment, this obstacle information can be determined based on whether a target trajectory point pair exists in the pedestrian's trajectory and the contact pedestrian's trajectory, where the line connecting the target trajectory point pair contacts at least one triangular facet in the three-dimensional map model. If the target trajectory point pair exists, it is determined that an obstacle exists between the pedestrian and the contact pedestrian; if the target trajectory point pair does not exist, it is determined that no obstacle exists between the pedestrian and the contact pedestrian.

[0131] In one implementation, the process of determining the target trajectory point pair may be:

[0132] First, assume that the vertex coordinates of each triangle are , , , find the normal vector of the plane where the triangle is located , the formula is as follows:

[0133] ;

[0134] ;

[0135] .

[0136] Next, construct the plane equation, the formula is as follows:

[0137] ;

[0138] Substitute the coordinates of the first and second trajectory points into the plane equation to obtain the result 、 ,in, and The expressions are:

[0139] ;

[0140] .

[0141] Finally, calculate The result, if If the result is positive, it is considered that there is no obstacle between the two pedestrians; if If the result is negative, the two pedestrians are on the opposite sides of the plane. The intersection of the line connecting the two pedestrians and the plane is , and then determine whether the intersection point falls within the patch range. If so, it is considered that there is an obstacle between pedestrians. The following is the intersection calculation formula:

[0142] ;

[0143] ;

[0144] .

[0145] Furthermore, in step S44, after obtaining the contact pedestrian, contact time, contact location, social distance, and obstacle information, the contact pedestrian, contact time, contact location, social distance, and obstacle information can be used as a pedestrian's crowd spatiotemporal contact behavior. Thus, through the above method, all crowd spatiotemporal contact behaviors corresponding to each pedestrian can be obtained, that is, all crowd spatiotemporal contact behaviors between the obtained pedestrian trajectories.

[0146] After obtaining all the spatiotemporal contact behaviors of the crowd, a spatiotemporal contact behavior sequence of the crowd can be formed based on all the spatiotemporal contact behaviors of the crowd, wherein the spatiotemporal contact behavior sequence of the crowd can reflect the spatiotemporal contact behavior of the crowd between any two pedestrians, that is, the spatiotemporal contact pedestrians with each pedestrian can be determined through the spatiotemporal contact behavior of the crowd, as well as the time, place and contact form of the contact behavior, wherein the contact form includes social distance and whether there are obstacles.

[0147] In one implementation, the crowd spatiotemporal contact behavior includes contact with pedestrians, contact time, contact location, social distance, and obstacle information; and forming a crowd spatiotemporal contact behavior sequence based on the extracted crowd spatiotemporal contact behavior specifically includes:

[0148] The pedestrian corresponding to each pedestrian trajectory is used as a node, wherein the node stores identity information, pedestrian trajectory and timestamp information;

[0149] For each pedestrian, a connection edge is constructed between the pedestrian and the pedestrian contacting the pedestrian to form a crowd spatiotemporal contact behavior sequence, wherein the crowd spatiotemporal contact behavior is stored on the connection edge.

[0150] Specifically, the spatiotemporal contact behavior sequence of a crowd is expressed in the form of graph data, where the nodes in the graph data are pedestrians. Specifically, a node is created for each pedestrian, storing the identity information, pedestrian trajectory, and timestamp information corresponding to the node. Then, based on the spatiotemporal contact behavior of the crowd, the nodes corresponding to two pedestrians in spatiotemporal contact are connected with edges, and the spatiotemporal contact behavior between the two pedestrians is stored on the connecting edges, where the contact time, contact location, social distance, and obstacle information in the spatiotemporal contact behavior of the crowd are stored. The spatiotemporal contact behavior sequence constructed in this embodiment is mainly applied in fields such as major infectious diseases, sudden public safety incidents, and criminal case investigation. The spatiotemporal contact behavior sequence of the crowd can be used to quickly determine the people in the spatiotemporal contact of the query, thereby effectively reducing the cost of video analysis and improving the efficiency and accuracy of video analysis.

[0151] In one implementation, after generating a crowd spatiotemporal contact behavior sequence, the crowd spatiotemporal contact behavior sequence can be backtracked to identify crowds that meet the query criteria. Based on this, after forming a crowd spatiotemporal contact behavior sequence based on the extracted crowd spatiotemporal contact behaviors, the method further includes:

[0152] Determine backtracking constraints;

[0153] Based on the backtracking constraint condition, the temporal and spatial contact behavior sequence of the crowd is queried by a breadth-first algorithm to obtain the backtracking pedestrian group corresponding to the target pedestrian.

[0154] Specifically, the backtracking constraint conditions are pre-set. In this embodiment, since the spatiotemporal contact behavior includes direct contact behavior and indirect contact behavior, direct contact behavior refers to the behavior of pedestrians violating social distance at the same time when the pedestrian trajectories intersect and there are no obstacles blocking them, that is, the pedestrian trajectories have intersections in the three dimensions of X, Y, and Z and the time dimension T is included in [0, t1]; indirect contact behavior means that the pedestrian trajectories have intersections in the three dimensions of X, Y, and Z, and the time dimension T is included in [t1, t2]. Both t1 and t2 are preset time thresholds, and t1 is less than t2. Therefore, the backtracking constraint condition can be that there is an intersection of trajectories in the same space and time, that is, the backtracking constraint condition limits the position coordinates and time conditions to ensure that the X, Y, and Z dimensions are the same, and the time dimension T is contained in [0, t1]. It can also be that there is an intersection of the same spatial trajectories at different times. The backtracking constraint condition limits the position coordinate information, and the time condition needs to be set to a threshold, that is, to ensure that the three dimensions X, Y, and Z are the same, the interval of T is set to [0, t], and the time dimension T is contained in [t1, t2].

[0155] After determining the backtracking constraints, a backtracking model can be determined based on the backtracking constraints. First, the target pedestrian and a preset number of groups of people are determined, which are represented as level-1, level-2, level-3, ..., level-k respectively. Then, the backtracking model is used to backtrack the spatiotemporal contact behavior sequence of the crowd. When the spatiotemporal sequence is traversed or the number of people of the corresponding group category centered on the query target pedestrian is found, the characteristic screenshots and identity information of these people are output. In this embodiment, the backtracking model can use a breadth-first algorithm for backtracking. The backtracking process can be to first confirm the starting node and put the starting node at the end of the queue; then, starting from the starting node, visit all its neighboring nodes in sequence. If the neighboring node has not been visited, it is added to the queue; then the next node is taken from the queue and the above operation is repeated until the queue is empty or the target node is found. The pseudo code of the backtracking process can be:

[0156] Def BFS(node,value):

[0157] nextlist=node.all neighbors

[0158] for ele in nextlist:

[0159] if not ele.ischecked:

[0160] ele.ischecked=True

[0161] if ele.value==value:

[0162] return ele

[0163] else:

[0164] Append all neighbors of ele to nextlist

[0165] nextlist.pop() #Pop the current ele from nextlist

[0166] In summary, this embodiment provides a method for monitoring the spatiotemporal contact behavior sequence of crowds in urban public spaces, the method comprising obtaining point cloud data of an urban public space and determining a three-dimensional map model corresponding to the urban public space based on the point cloud data; obtaining a plurality of video sequences of the urban public space and identifying pedestrian information of each video frame image in each video sequence to form a pedestrian information sequence; mapping the position information in each pedestrian information to a three-dimensional map model to obtain three-dimensional position information corresponding to each position information, and sorting the three-dimensional position information corresponding to the same identity information in chronological order to obtain a plurality of pedestrian trajectories; extracting the spatiotemporal contact behavior of crowds from a plurality of pedestrian trajectories based on a preset social distance, and forming a spatiotemporal contact behavior sequence of crowds based on the extracted spatiotemporal contact behavior of crowds. This application obtains a pedestrian information sequence by extracting a plurality of video sequences, then projects the pedestrian information sequence to a three-dimensional map model to form a three-dimensional pedestrian trajectory, and finally forms a spatiotemporal contact behavior sequence of crowds based on the three-dimensional pedestrian trajectory. In this way, pedestrian contact behavior is detected from the spatiotemporal dimension, which not only improves the accuracy of pedestrian contact behavior recognition, but also can well adapt to the current rapidly developing three-dimensional urban space environment.

[0167] Based on the above-mentioned method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, this embodiment provides a system for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, such as Figure 8 As shown, the system includes:

[0168] An acquisition module 100 is configured to acquire point cloud data of an urban public space and determine a three-dimensional map model corresponding to the urban public space based on the point cloud data;

[0169] an identification module 200 for acquiring a plurality of video sequences of the urban public space and identifying pedestrian information in each video frame image in each video sequence to form a pedestrian information sequence, wherein the pedestrian information sequence includes a plurality of pedestrian information, each pedestrian information including identity information and location information;

[0170] A mapping module 300 is configured to map the position information in each pedestrian information to a three-dimensional map model to obtain three-dimensional position information corresponding to each position information, and to sort the three-dimensional position information corresponding to the same identity information in chronological order to obtain a plurality of pedestrian trajectories;

[0171] The extraction module 400 is used to extract the spatiotemporal contact behaviors of a number of pedestrian trajectories based on a preset social distance, and form a spatiotemporal contact behavior sequence of a crowd according to the extracted spatiotemporal contact behaviors of the crowd.

[0172] Based on the above-mentioned method for monitoring the spatiotemporal contact behavior sequences of people in urban public spaces, this embodiment provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in the method for monitoring the spatiotemporal contact behavior sequences of people in urban public spaces as described in the above-mentioned embodiment.

[0173] Based on the above-mentioned monitoring method for the spatiotemporal contact behavior sequence of people in urban public spaces, this application also provides a terminal device, such as Figure 9 As shown, it includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via bus 24. The display screen 21 is configured to display a preset user guidance interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logic instructions in the memory 22 to execute the method described in the above embodiment.

[0174] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0175] The memory 22, as a computer-readable storage medium, can be configured to store software programs or computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes the software programs, instructions, or modules stored in the memory 22 to perform functional applications and data processing, thereby implementing the methods in the above embodiments.

[0176] The memory 22 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal device. In addition, the memory 22 may include high-speed random access memory and non-volatile memory. For example, various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, may also be transient storage media.

[0177] In addition, the specific process of loading and executing the multiple instructions in the storage medium and the processor in the terminal device has been described in detail in the above method and will not be described here one by one.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces, characterized by: The method comprises: Acquiring point cloud data of an urban public space, and determining a three-dimensional map model corresponding to the urban public space based on the point cloud data; Acquire multiple video sequences of the urban public space, and identify pedestrian information in each video frame image in each video sequence to form a pedestrian information sequence, wherein the pedestrian information includes identity information and location information; Mapping the position information of each pedestrian information to a three-dimensional map model to obtain the three-dimensional position information corresponding to each position information, and sorting the three-dimensional position information corresponding to the same identity information in chronological order to obtain a number of pedestrian trajectories; Extracting the spatiotemporal contact behaviors of several pedestrian trajectories based on a preset social distance, and forming a spatiotemporal contact behavior sequence based on the extracted spatiotemporal contact behaviors of the crowd; The step of mapping the position information in each pedestrian information to the three-dimensional map model to obtain the three-dimensional position information corresponding to each position information specifically includes: Obtaining a second conversion model between a pixel coordinate system corresponding to each video sequence and a three-dimensional map coordinate system corresponding to the three-dimensional map model; For each pedestrian information, converting the position information in the pedestrian information into a three-dimensional map coordinate system based on a second conversion model corresponding to the pixel coordinate system of the pedestrian information to obtain the sole position information corresponding to the position information; determining height information corresponding to the pedestrian information, and using three-dimensional coordinate information formed by the sole position information and the height information as three-dimensional position information corresponding to the position information; The extraction of a plurality of pedestrian trajectories based on a preset social distance specifically includes: Calculate the discrete Fréchet distance between two pedestrian trajectories in a number of pedestrian trajectories to obtain the social distance between two pedestrians; For each pedestrian, a target social distance greater than a preset social distance is selected from all social distances corresponding to the pedestrian, and the contact pedestrian, contact time, and contact position corresponding to the pedestrian are determined based on the trajectory point pairs corresponding to the obtained target social distance; For each contact pedestrian of the pedestrian, determining a trajectory line connecting any first trajectory point in the pedestrian trajectory of the pedestrian and any second trajectory point in the pedestrian trajectory of the contact pedestrian, and detecting whether the trajectory line intersects a triangular facet in the three-dimensional map model to determine obstacle information between the pedestrian and the contact pedestrian; The contact pedestrian, contact time, contact position, social distance and obstacle information corresponding to the pedestrian are used as a crowd spatiotemporal contact behavior to obtain the crowd spatiotemporal contact behaviors corresponding to several pedestrian trajectories.

2. The method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces according to claim 1 is characterized in that: Determining the three-dimensional map model corresponding to the urban public space based on the point cloud data specifically includes: Constructing a three-dimensional map coordinate system corresponding to the urban public space, and determining a first conversion model from the point cloud coordinate system corresponding to the point cloud data to the three-dimensional map coordinate system; For each point cloud data point in the point cloud data, determining a neighborhood of the point cloud data point, determining a projected tangent plane between the point cloud data point and each neighboring point in the neighborhood based on a first transformation model, and projecting each neighboring point onto the projected tangent plane; The projected tangent plane corresponding to each point cloud data trajectory point is triangulated to obtain a three-dimensional map model.

3. The method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces according to claim 2 is characterized in that: For each point cloud data point in the point cloud data, determining the neighborhood of the point cloud data point specifically includes: downsampling the point cloud data using a voxel filtering algorithm to obtain target point cloud data; For each point cloud data point in the target point cloud data, a KD tree nearest neighbor search method is used to determine the neighborhood corresponding to the point cloud data trajectory point, wherein the neighborhood includes a preset number of neighborhood points.

4. The method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces according to claim 1 is characterized in that: The spatiotemporal contact behavior of the crowd includes contacted pedestrians, contact time, contact location, social distance and obstacle information; The forming of a crowd spatiotemporal contact behavior sequence based on the extracted crowd spatiotemporal contact behavior specifically includes: The pedestrian corresponding to each pedestrian trajectory is used as a node, wherein the node stores identity information, pedestrian trajectory and timestamp information; For each pedestrian, a connection edge is constructed between the pedestrian and the pedestrian contacting the pedestrian to form a crowd spatiotemporal contact behavior sequence, wherein the crowd spatiotemporal contact behavior is stored on the connection edge.

5. The method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces according to claim 1 is characterized in that: After forming a crowd space-time contact behavior sequence based on the extracted crowd space-time contact behavior, the method further includes: Determine backtracking constraints; Based on the backtracking constraint condition, the temporal and spatial contact behavior sequence of the crowd is queried by a breadth-first algorithm to obtain the backtracking pedestrian group corresponding to the target pedestrian.

6. A monitoring system for the temporal and spatial contact behavior sequence of people in urban public spaces, characterized by: The system comprises: an acquisition module, configured to acquire point cloud data of an urban public space and determine a three-dimensional map model corresponding to the urban public space based on the point cloud data; an identification module, configured to obtain a plurality of video sequences of the urban public space and identify pedestrian information in each video frame image in each video sequence to form a pedestrian information sequence, wherein the pedestrian information sequence includes a plurality of pedestrian information, each pedestrian information including identity information and location information; A mapping module is used to map the position information in each pedestrian information to a three-dimensional map model to obtain the three-dimensional position information corresponding to each position information, and to sort the three-dimensional position information corresponding to the same identity information in chronological order to obtain a number of pedestrian trajectories; An extraction module is used to extract the spatiotemporal contact behaviors of a number of pedestrian trajectories based on a preset social distance, and form a spatiotemporal contact behavior sequence of a crowd based on the extracted spatiotemporal contact behaviors of the crowd; The step of mapping the position information in each pedestrian information to the three-dimensional map model to obtain the three-dimensional position information corresponding to each position information specifically includes: Obtaining a second conversion model between a pixel coordinate system corresponding to each video sequence and a three-dimensional map coordinate system corresponding to the three-dimensional map model; For each pedestrian information, converting the position information in the pedestrian information into a three-dimensional map coordinate system based on a second conversion model corresponding to the pixel coordinate system of the pedestrian information to obtain the sole position information corresponding to the position information; determining height information corresponding to the pedestrian information, and using three-dimensional coordinate information formed by the sole position information and the height information as three-dimensional position information corresponding to the position information; The extraction of a plurality of pedestrian trajectories based on a preset social distance specifically includes: Calculate the discrete Fréchet distance between two pedestrian trajectories in a number of pedestrian trajectories to obtain the social distance between two pedestrians; For each pedestrian, a target social distance greater than a preset social distance is selected from all social distances corresponding to the pedestrian, and the contact pedestrian, contact time, and contact position corresponding to the pedestrian are determined based on the trajectory point pairs corresponding to the obtained target social distance; For each contact pedestrian of the pedestrian, determining a trajectory line connecting any first trajectory point in the pedestrian trajectory of the pedestrian and any second trajectory point in the pedestrian trajectory of the contact pedestrian, and detecting whether the trajectory line intersects a triangular facet in the three-dimensional map model to determine obstacle information between the pedestrian and the contact pedestrian; The contact pedestrian, contact time, contact position, social distance and obstacle information corresponding to the pedestrian are used as a crowd spatiotemporal contact behavior to obtain the crowd spatiotemporal contact behaviors corresponding to several pedestrian trajectories.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces as described in any one of claims 1 to 5.

8. A terminal device, characterized in that: include: processor, memory, and communication bus; The memory stores a computer-readable program executable by the processor; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, the steps of the method for monitoring the spatiotemporal contact behavior sequence of people in urban public spaces as described in any one of claims 1 to 5 are implemented.