Video monitoring method and system, electronic equipment and storage medium
By constructing a three-dimensional digital twin scene and node relationship matrix, combined with graph neural network and dynamic field of view splicing algorithm, the problem that supervisors in large video surveillance environments is difficult to monitor globally in real time, and accurate video monitoring and risk management of the target area is achieved, and the accuracy and timeliness of monitoring are improved.
Patent Information
- Application Number
- CN202510223890.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In large-scale video surveillance environments, it is difficult for supervisors to conduct global real-time monitoring of multiple storyboard images, which affects the accuracy and timeliness of video surveillance.
By constructing a three-dimensional digital twin scene and node relationship matrix, using graph neural network and dynamic vision stitching algorithm, the real-time location and action path of the target object are determined, and dynamic heat maps are generated through multi-camera vision fusion to display the flow density. Combining bone key point detection, interaction behavior analysis and Bayesian network probability reasoning framework, risk levels of action trajectories are evaluated and disposal suggestions are generated.
Comprehensive and accurate video monitoring and risk management of the target area are achieved, the accuracy and timeliness of monitoring are improved, the fatigue and error of manual monitoring are reduced, and the automation and intelligent decision-making support capabilities of the monitoring system are enhanced.
Smart Images

Figure CN120220047A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video monitoring, and specifically relates to a method, system, electronic device and storage medium for video monitoring. Background Art
[0002] In modern times when technology is developing very rapidly, surveillance cameras are deployed in many locations, such as streets, intersections, stairways, lakesides and other positions where accidents may occur. The cameras set up can effectively provide effective video evidence for security personnel or police officers, and the real-time surveillance videos can better help the staff monitor the target positions to respond promptly to possible dangerous behaviors. However, in a large-scale video monitoring environment, since the number of cameras is usually very large, supervisors need to watch multiple sub-shot images at the same time, and it is impossible to conduct global real-time monitoring of the large scene. Moreover, due to the influence of many factors such as attention and visual blind spots in the work, the effectiveness of its monitoring remains to be discussed.
[0003] Therefore, how to ensure the accuracy and timeliness of video monitoring in a large-scale video monitoring environment is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0004] This application provides a method, system, electronic device and storage medium for video monitoring, which realizes comprehensive and accurate video monitoring and risk management of the target area.
[0005] In the first aspect of this application, a method for video monitoring is provided, which is applied to a video monitoring platform. The method includes: Construct a three-dimensional digital twin scene according to the spatial coordinates of each camera in the target area. Using a graph neural network, each camera is used as a first node, and a node relationship matrix is constructed according to the spatial relationship between the cameras; According to the three-dimensional digital twin scene and the node relationship matrix, use the dynamic field-of-view stitching algorithm to determine the real-time position and action path of the target object, and generate a dynamic heat map by fusing the fields of view of multiple cameras to display the pedestrian flow density of the target area; Determine the behavior pattern of the target object through skeleton key point detection and interaction behavior analysis, and predict the action trajectory of the target object according to the behavior pattern combined with the attention mechanism; Evaluate the initial risk level of the action trajectory according to the preset risk propagation model and the Bayesian network probability inference framework, and adjust the initial risk level in combination with the pedestrian flow density to obtain the final risk level, and generate a disposal suggestion according to the final risk level.
[0006] Optionally, constructing a three-dimensional digital twin scene based on the spatial coordinates of each camera in the target area and using a graph neural network with each camera as a node to construct a node relationship matrix according to the spatial relationship between the cameras includes: Using a 3D reconstruction network to convert the two-dimensional image information of the target images collected by each camera into three-dimensional spatial information to form a three-dimensional space, identifying the objects in the target images through semantic segmentation, and mapping the objects into the three-dimensional space; Establishing a first edge between the first nodes according to the spatial relationship between the cameras, constructing an adjacency matrix of the graph through the first edge, and the elements in the adjacency matrix represent the connection weights or distances between the nodes.
[0007] Optionally, determining the real-time position and action path of the target object using a dynamic field-of-view stitching algorithm based on the three-dimensional digital twin scene and the node relationship matrix includes: Constructing a field-of-view graph of the cameras based on the three-dimensional digital twin scene and the node relationship matrix, where the field of view of each camera is a region and the overlapping part between the fields of view is a shared region; Determining the position information of the target object in different fields of view through the field-of-view graph, and calculating the action path based on the position changes of the target object in different fields of view in combination with the timestamp information; Smoothing the action path of the target object using a tracking algorithm.
[0008] Optionally, generating a dynamic heat map using multi-camera field-of-view fusion to display the crowd density in the target area includes: Fusing the fields of view of multiple cameras through the three-dimensional digital twin scene and the node relationship matrix to generate a unified view covering the target area; In the unified view, counting the number of people appearing at the target pixel points or regional units within a preset time interval to generate crowd density data; Generating a dynamic heat map based on the crowd density data, where the color or brightness of the heat map represents the level of crowd density, and the crowd density distribution in the target area is displayed in real time.
[0009] Optionally, determining the behavior pattern of the target object through skeletal key point detection and interaction behavior analysis includes: Using a skeletal key point detection algorithm to extract the skeletal key point information of the target object, where the skeletal key point information includes joint positions and pose angles; According to the skeletal key point information, identifying the interaction behavior of the target object through a pre-trained interaction behavior analysis model, where the interaction behavior includes conversation, handshake, or pushing; Determine the behavior pattern of the target object according to the type and duration of the interaction behavior, where the behavior pattern includes normal behavior, abnormal behavior, or potential dangerous behavior.
[0010] Optionally, the predicting the action trajectory of the target object according to the behavior pattern in combination with the attention mechanism includes: Construct a behavior feature vector of the target object according to the behavior pattern, where the behavior feature vector includes skeletal key point information, interaction behavior information, and environmental information; Input the behavior feature vector into a prediction model based on Transformer, and introduce an attention mechanism in the prediction model to perform weighted processing on the behavior features at different time steps; Predict the action trajectory of the target object through the prediction model, where the action trajectory includes continuing to move forward, staying, or changing direction.
[0011] Optionally, the evaluating the initial risk level of the action trajectory according to a preset risk propagation model and a Bayesian network probability inference framework includes: Use the preset risk propagation model to analyze the risk propagation path of the action trajectory of the target object in the target area and identify potential risk nodes; Construct a Bayesian network including the behavior of the target object, environmental factors, and historical events according to the potential risk nodes and the risk propagation path, where the second nodes of the Bayesian network represent random variables, and the second edges represent the conditional dependence relationship between variables; Use a conditional probability table, in combination with the random variables and the conditional dependence relationship in the Bayesian network, to calculate the initial risk level of the action trajectory of the target object.
[0012] In a second aspect of the present application, a video monitoring system is provided, including a construction module, a splicing module, a prediction module, and a risk module, where: The construction module is configured to construct a three-dimensional digital twin scene according to the spatial coordinates of each camera in the target area, and use a graph neural network to take each camera as a node and construct a node relationship matrix according to the spatial relationship between the cameras; The splicing module is configured to use a dynamic field-of-view splicing algorithm according to the three-dimensional digital twin scene and the node relationship matrix to determine the real-time position and action path of the target object, and generate a dynamic heat map by fusing multi-camera fields of view to display the crowd density of the target area; The prediction module is configured to determine the behavior pattern of the target object through skeletal key point detection and interaction behavior analysis, and predict the action trajectory of the target object according to the behavior pattern in combination with the attention mechanism; A risk module configured to evaluate an initial risk level of the action trajectory according to a preset risk propagation model and a Bayesian network probability inference framework, adjust the initial risk level in combination with the crowd density to obtain a final risk level, and generate a disposal suggestion according to the final risk level.
[0013] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, and both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of the above.
[0014] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions that, when executed, execute the method described in any one of the above.
[0015] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By constructing a three-dimensional digital twin scene based on the spatial coordinates of each camera in the target area, using a graph neural network with each camera as a node, and constructing a node relationship matrix according to the spatial relationship between cameras, accurate modeling of the monitoring area is achieved. This enables the system to accurately determine the real-time position and action path of the target object in three-dimensional space; using the dynamic field-of-view stitching algorithm, combining the three-dimensional digital twin scene and the node relationship matrix, the fields of view of multiple cameras can be stitched in real time, eliminating physical occlusion and achieving seamless tracking of the target object. This ensures that the action path of the target object is completely recorded within the monitoring area.
[0016] 2. By generating a dynamic heat map through multi-camera field-of-view fusion, the system can display the crowd density of the target area in real time. This provides intuitive visual information for monitoring personnel, helping them quickly identify and respond to crowded areas, and preventing potential safety hazards; 3. Through skeleton key point detection and interactive behavior analysis, the system can determine the behavior pattern of the target object. Combining the attention mechanism, the system can predict the future action trajectory of the target object. This provides a scientific basis for taking preventive measures in advance and helps to detect and handle potential abnormal behaviors in a timely manner.
[0017] 4. Using a preset risk propagation model and a Bayesian network probability inference framework, the system can evaluate the initial risk level of the target object's action trajectory, and dynamically adjust it in combination with the crowd density to obtain the final risk level. This enables the system to accurately evaluate potential risks based on real-time data and historical information, generate corresponding disposal suggestions, and help monitoring personnel quickly respond to and handle emergencies; 5. The entire monitoring process is highly automated. From the positioning of the target object, behavior analysis to risk assessment and the generation of disposal suggestions, no manual intervention is required. This greatly improves the monitoring efficiency and reduces the fatigue and errors of manual monitoring. The system realizes intelligent decision-making support through machine learning and deep learning technologies. Monitoring personnel can quickly take appropriate measures according to the disposal suggestions generated by the system, improving the emergency response ability. 6. Through the three-dimensional digital twin scenario, monitoring personnel can intuitively view the layout of the monitoring area and the action paths of the target objects. This makes the monitoring more intuitive and easy to understand, helping to make decisions quickly. The dynamic heat map shows the crowd density in real time, helping monitoring personnel quickly identify high-risk areas. This visualization tool improves the efficiency and accuracy of monitoring. Description of the Drawings
[0018] Figure 1 is a schematic flowchart of the method for video monitoring disclosed in the embodiments of the present application; Figure 2 is a schematic block diagram of the system for video monitoring disclosed in the embodiments of the present application; Figure 3 is a schematic structural diagram of an electronic device disclosed in the embodiments of the present application.
[0019] Description of the Reference Numerals: 201, construction module; 202, splicing module; 203, prediction module; 204, risk module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Embodiments
[0020] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0021] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for instance" is intended to present relevant concepts in a specific manner.
[0022] In the description of the embodiments of the present application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0023] This embodiment discloses a method for video monitoring, which is applied to a video monitoring platform. Figure 1 It is a schematic flowchart of the method for video monitoring disclosed in the embodiments of the present application, as Figure 1 shown. The method includes the following steps: S101. Construct a three-dimensional digital twin scene based on the spatial coordinates of each camera in the target area. Using a graph neural network, each camera is used as a first node, and a node relationship matrix is constructed according to the spatial relationship between the cameras; S102. According to the three-dimensional digital twin scene and the node relationship matrix, use a dynamic field-of-view stitching algorithm to determine the real-time position and action path of the target object, and generate a dynamic heat map by fusing the fields of view of multiple cameras to display the pedestrian flow density in the target area; S103. Determine the behavior pattern of the target object through skeleton key point detection and interaction behavior analysis, and predict the action trajectory of the target object according to the behavior pattern combined with the attention mechanism; S104. Evaluate the initial risk level of the action trajectory according to a preset risk propagation model and a Bayesian network probability inference framework, and adjust the initial risk level in combination with the pedestrian flow density to obtain the final risk level, and generate a disposal suggestion according to the final risk level.
[0024] Based on the spatial coordinates of each camera within the target area, a 3D reconstruction network is used to convert the two-dimensional image information collected by the cameras into three-dimensional spatial information, thereby constructing an accurate three-dimensional digital twin scene. Through semantic segmentation technology, objects in the target image (such as buildings, roads, pedestrians, etc.) are identified and mapped into the three-dimensional space to form a complete three-dimensional digital twin scene. A graph neural network is adopted, with each camera regarded as a first node. Edges between nodes are established according to the spatial relationships between cameras (such as distance, angle, field of view overlap, etc.), and the adjacency matrix of the graph is constructed through these edges. The elements in the adjacency matrix represent the connection weights or distances between nodes, thus constructing a node relationship matrix reflecting the spatial topological relationship of the cameras. Based on the constructed three-dimensional digital twin scene and the node relationship matrix, a dynamic field of view stitching algorithm is used to fuse the fields of view of multiple cameras, eliminate physical occlusions, and determine the position and movement path of the target object in real time. Through technologies such as image registration and image fusion, the images within the overlapping fields of view are aligned and fused to generate a seamless stitched image, thereby accurately capturing the movement trajectory of the target object. Using the multi-camera field of view fusion technology, the number of people appearing in the pixel points or regional units of the target area within a preset time interval is counted to generate crowd density data. Based on these data, a dynamic heat map is generated to visually display the crowd density distribution of the target area through color or brightness changes, and the aggregation and evacuation states of the crowd are reflected in real time. Through the skeleton key point detection algorithm, the skeleton key point information of the target object (including joint positions and pose angles) is extracted, and combined with the interactive behavior analysis model, the interactive behaviors between the target objects (such as talking, shaking hands, pushing, etc.) are identified. According to the type and duration of the interactive behavior, the behavior pattern of the target object is determined and classified as normal behavior, abnormal behavior, or potential dangerous behavior. Based on the determined behavior pattern, a behavior feature vector of the target object is constructed, which includes skeleton key point information, interactive behavior information, and environmental information. The behavior feature vector is input into a prediction model based on Transformer, and an attention mechanism is introduced into the model to weight the behavior features at different time steps, thereby predicting the future movement trajectory of the target object, such as continuing to move forward, staying, or changing direction. Using a preset risk propagation model, the possible risk propagation paths caused by the movement trajectory of the target object within the monitoring area are analyzed, and potential risk nodes and influence ranges are identified. Combining with the Bayesian network probability inference framework, a probability graph model containing the behavior of the target object, environmental factors, and historical events is constructed. Each node is modeled through a conditional probability table (CPT), and the initial risk level corresponding to the movement trajectory of the target object is calculated under given conditions. Combining with the crowd density data displayed by the dynamic heat map, the initial risk level is adjusted to obtain the final risk level.Generate corresponding disposal suggestions according to the final risk level to guide the monitoring personnel to take appropriate countermeasures, such as emergency evacuation, resource allocation, or safety reminders, etc., so as to effectively respond to and prevent potential security threats.
[0025] Optionally, constructing a three-dimensional digital twin scene based on the spatial coordinates of each camera in the target area, and using a graph neural network, taking each camera as a node, and constructing a node relationship matrix according to the spatial relationship between the cameras includes: Using a 3D reconstruction network to convert the two-dimensional image information of the target image collected by each camera into three-dimensional space information to form a three-dimensional space, identifying the objects in the target image through semantic segmentation, and mapping the objects into the three-dimensional space; Establishing a first edge between the first nodes according to the spatial relationship between the cameras, and constructing an adjacency matrix of the graph through the first edge, where the elements in the adjacency matrix represent the connection weights or distances between the nodes.
[0026] The 3D reconstruction network is a deep learning model that can convert two-dimensional image information into three-dimensional spatial information. It infers the three-dimensional shape and position of an object by analyzing feature points, textures, and geometric structures in the image. For example, in the shopping mall surveillance scenario, the 3D reconstruction network can convert objects such as pedestrians, shelves, and aisles in the two-dimensional images captured by each camera into three-dimensional models, forming a complete three-dimensional digital twin scenario. Semantic segmentation technology is used to identify objects in the target image. It divides the image into multiple parts and assigns a semantic label to each part, such as "pedestrian", "vehicle", "building", etc. Through semantic segmentation, objects in the image can be accurately identified and mapped into the three-dimensional space. Mapping an object from a two-dimensional image to the three-dimensional space requires coordinate transformation. The position and pose of each object in the three-dimensional space are calculated based on its projection in the two-dimensional image and the internal and external parameters of the camera. Using the principle of geometric projection and camera calibration technology, the two-dimensional coordinates of the object can be converted into three-dimensional coordinates, thus constructing a three-dimensional scene containing the position and shape of the object. Each camera serves as a node in the graph and establishes edges with other nodes according to their spatial relationships (such as distance, angle, field of view overlap, etc.). For example, if the fields of view of two cameras overlap, an edge is established between them, indicating a spatial connection between them. The adjacency matrix is a two-dimensional array used to represent the connection relationships between nodes in the graph. The size of the matrix is N×N, where N (a positive integer) is the number of nodes. The elements in the matrix represent the connection weights or distances between nodes. For example, if there is an edge between two nodes, the corresponding matrix element is a non-zero value, representing their connection weight or distance. The connection weights can be calculated based on factors such as the spatial distance between cameras and the degree of field of view overlap. The graph neural network (GNN) is a neural network specifically designed to process graph-structured data. In the video monitoring system, the GNN can utilize the information in the node relationship matrix to model and analyze camera nodes. Through the GNN, complex relationships between camera nodes can be mined, the monitoring layout can be optimized, and the efficiency and accuracy of video monitoring can be improved. For example, the GNN can identify the fields of view of the cameras that are most important for monitoring key areas, thus making more reasonable decisions in resource allocation and data processing.
[0027] By adopting a 3D reconstruction network, the two-dimensional image information collected by the camera is converted into three-dimensional space information, and an accurate three-dimensional scene model can be formed. This method not only improves the realism of the scene, but also can effectively capture the objects and structures in the space, providing a reliable basis for subsequent monitoring and analysis. Taking each camera as a node and constructing a node relationship matrix according to the spatial relationship between the cameras can effectively represent the topological structure between the cameras. This structured representation enables the system to better understand the mutual relationship between the cameras, thereby optimizing the monitoring strategy and resource allocation. By establishing an adjacency matrix, the system can quantify the connection weights or distances between nodes. This quantification method provides an important reference for subsequent dynamic field-of-view stitching and target tracking, enabling the monitoring system to achieve more efficient target detection and behavior analysis in complex environments.
[0028] Optionally, the step of determining the real-time position and action path of the target object by using the dynamic field-of-view stitching algorithm based on the three-dimensional digital twin scene and the node relationship matrix includes: Constructing a field-of-view graph of the cameras based on the three-dimensional digital twin scene and the node relationship matrix, where the field of view of each camera is a region, and the overlapping part between the fields of view is a shared region; Determining the position information of the target object in different fields of view through the field-of-view graph, and calculating the action path based on the position changes of the target object in different fields of view and combining the timestamp information; Smoothing the action path of the target object by using a tracking algorithm.
[0029] Suppose that in a large shopping mall, multiple cameras are installed to monitor the activities of people in the mall. The spatial coordinates of each camera are known and are modeled through a three-dimensional digital twin scenario and a node relationship matrix. There are three cameras (A, B, and C) in the mall, and their fields of view partially overlap, covering the main passages and some store areas in the mall. Definition of the field of view area: The field of view of camera A covers the mall entrance and part of the passage; the field of view of camera B covers the passage and some store areas, overlapping with the field of view of camera A in the passage part; the field of view of camera C covers the store area and another passage, overlapping with the field of view of camera B in the store area. Based on the three-dimensional digital twin scenario and the node relationship matrix, a field of view graph of the cameras is constructed. The field of view of each camera is an area, and the overlapping part between the fields of view is the shared area. In the field of view graph, the fields of view of cameras A, B, and C are respectively represented as areas A, B, and C, and the shared areas are A∩B and B∩C. Suppose a target object (such as a customer) enters from the mall entrance and first appears in the field of view of camera A. Through the target detection algorithm, the position of the target object in the field of view of camera A is determined. The target object moves along the passage and enters the shared area of cameras A and B. Through the field of view graph, the system can identify the position information of the target object in the fields of view of the two cameras. Combining the timestamp information, the system records the position changes of the target object in the fields of view of cameras A and B and preliminarily calculates its movement path. The target object continues to move and enters the shared area of cameras B and C. The system updates the position information of the target object again and, combining the timestamp information, further calculates its movement path. Since the positions of the target object in the fields of view of different cameras may have noise and errors, the system uses the Kalman filter algorithm to smooth the movement path of the target object. The Kalman filter continuously corrects the position estimate of the target object through two steps: state prediction and measurement update, reducing the influence of noise and errors. After being processed by the Kalman filter, the movement path of the target object is smoother and more continuous, and can more accurately reflect its movement trajectory in the mall. The monitoring personnel can better track the movement of the target object through the smoothed movement path and timely discover and handle potential abnormal behaviors.
[0030] Using the camera spatial coordinates in the 3D digital twin scenario and the camera connection relationships in the node relationship matrix, construct the field of view graph of the cameras. The field of view of each camera is defined as a region, and the overlapping part between the fields of view is defined as the shared region. The field of view graph can visually display the monitoring ranges of multiple cameras and their overlapping situations, providing a basis for subsequent determination of the target object's position and calculation of the action path. Through the field of view graph, the position relationship of the target object in the fields of view of different cameras and the collaborative monitoring situation between the cameras can be clearly seen. Using the field of view graph, determine the position information of the target object in the fields of view of different cameras. When the target object moves from the field of view of one camera to that of another camera, through the shared region in the field of view graph, seamless tracking of the target object can be achieved. According to the position changes of the target object in different fields of view and combined with the timestamp information, calculate its action path. Through the timestamp information, the position of the target object at different time points can be accurately determined, thereby generating its action trajectory. This method can reflect the movement situation of the target object in real time, providing a basis for subsequent risk assessment and disposal suggestions. Using tracking algorithms, such as Kalman filtering, etc., smooth the action path of the target object. The tracking algorithm can predict the current position and future position of the target object based on its historical position information and motion state, thereby reducing the noise and error of position estimation. Through the smoothing process of the tracking algorithm, the action path of the target object is smoother and more continuous, improving the accuracy and stability of the real-time position and action path. This is crucial for accurately evaluating the behavior pattern of the target object and predicting its future action trajectory, and helps to detect and handle potential security threats in a timely manner.
[0031] Optionally, the generating a dynamic heat map by fusing the fields of view of multiple cameras to display the crowd density in the target area includes: Fuse the fields of view of multiple cameras through the 3D digital twin scenario and the node relationship matrix to generate a unified view covering the target area; In the unified view, count the number of people appearing at the target pixel points or regional units within a preset time interval to generate crowd density data; Generate a dynamic heat map based on the crowd density data, where the color or brightness of the heat map represents the level of crowd density, and display the crowd density distribution in the target area in real time.
[0032] Through the three-dimensional digital twin scenario and the node relationship matrix, the fields of view of multiple cameras are fused to generate a unified view covering the target area. Specifically, using the spatial information in the three-dimensional digital twin scenario and the camera position relationship in the node relationship matrix, the fields of view of multiple cameras are aligned and fused. For example, in a surveillance system of a large shopping mall, by fusing the fields of view of multiple cameras, a unified view covering the entire mall can be generated, eliminating the limitations of the fields of view of individual cameras and achieving comprehensive surveillance of the entire mall. In the unified view, the number of people appearing at the target pixel points or regional units within a preset time interval is counted to generate crowd density data. Specifically, by counting the number of people at each pixel point or regional unit in the unified view, crowd density data can be obtained. For example, in the unified view of the shopping mall, the crowd density data can be generated by counting the number of people in each regional unit, so as to understand the crowd distribution in the mall. According to the crowd density data, a dynamic heat map is generated, where the color or brightness of the heat map represents the level of crowd density, and the crowd density distribution of the target area is displayed in real time. Specifically, using the crowd density data, a dynamic heat map can be generated, and the level of crowd density is visually shown through the change of color or brightness. For example, in the surveillance system of the shopping mall, the crowd density distribution in the mall can be displayed in real time through the dynamic heat map, helping the management personnel to timely understand the crowd changes in the mall and take corresponding management measures.
[0033] Through the three-dimensional digital twin scenario and the node relationship matrix, the fields of view of multiple cameras are fused to generate a unified view covering the target area. This fusion method can eliminate the limitations of the fields of view of individual cameras, provide a comprehensive and continuous surveillance perspective, and enable the surveillance personnel to more intuitively observe the overall situation of the target area. In the unified view, the number of people appearing at the target pixel points or regional units within a preset time interval is counted to generate crowd density data. In this way, the crowd distribution in the target area can be accurately captured, providing data support for subsequent risk assessment and disposal suggestions. According to the crowd density data, a dynamic heat map is generated, where the color or brightness of the heat map represents the level of crowd density, and the crowd density distribution of the target area is displayed in real time. This visualization method enables the surveillance personnel to quickly identify high-density crowd areas, timely discover potential safety hazards, and improve the surveillance efficiency and safety.
[0034] Optionally, the determination of the behavior pattern of the target object by skeletal key point detection and interaction behavior analysis includes: Using a skeletal key point detection algorithm, extract the skeletal key point information of the target object, where the skeletal key point information includes joint point positions and pose angles; Based on the skeletal key point information, identify the interaction behavior of the target object through a pre-trained interaction behavior analysis model, where the interaction behavior includes conversation, handshake or pushing and shoving; Determine the behavior pattern of the target object according to the type and duration of the interaction behavior, where the behavior pattern includes normal behavior, abnormal behavior or potentially dangerous behavior.
[0035] Adopt advanced skeletal key point detection algorithms such as OpenPose. These algorithms can detect the key points of the human body in real time, including joint positions and pose angles. The OpenPose algorithm can handle multi-person pose estimation through a deep learning-based model, using Part Affinity Fields (PAFs) to associate different key points, thereby achieving accurate key point detection. The algorithm output includes the coordinate information of each key part of the human body, such as the head, shoulders, elbows, wrists, hips, knees and ankles. These coordinate information are used to determine the posture and movement of the human body. For example, by the change in the position of the key points, it can be judged whether the human body is walking, standing or sitting, etc. Use pre-trained deep learning models, such as models based on convolutional neural networks (CNNs) or graph neural networks (GNNs), to analyze the extracted skeletal key point information. These models can identify various interaction behaviors, including conversation, handshake or pushing and shoving, etc. For example, by analyzing the relative positions and movement trajectories between key points, the model can judge whether two people are shaking hands or pushing and shoving. The model can identify the characteristics of different behaviors by learning a large amount of labeled data. For example, conversation behavior usually involves the orientation of the head and body as well as hand movements, while pushing and shoving behavior involves more violent physical contact and movements. Classify the behavior pattern of the target object as normal behavior, abnormal behavior or potentially dangerous behavior according to the identified interaction behavior type and duration. For example, a short handshake may be regarded as normal behavior, while a long-term pushing and shoving or conflict may be regarded as abnormal or potentially dangerous behavior. By analyzing the duration and frequency of the behavior, the system can further judge the nature of the behavior. For example, a person staying in a certain area for a long time may be regarded as abnormal behavior, while passing by quickly may be normal behavior. For example, normal behavior: A person walks normally in a shopping mall and talks with others. Key point detection shows that their posture and movements conform to the characteristics of normal walking and conversation, and the system identifies their behavior pattern as normal behavior. Abnormal behavior: A person suddenly runs quickly in a shopping mall. Key point detection shows that their speed and movement amplitude are abnormal, and the system identifies their behavior pattern as abnormal behavior. Potentially dangerous behavior: Two people push and shove in a shopping mall. Key point detection shows that their physical contact and movements are violent, and the system identifies their behavior pattern as potentially dangerous behavior.
[0036] Using a skeletal key point detection algorithm, the skeletal key point information of the target object can be accurately extracted, including the joint point positions and pose angles. This information provides the basic data for subsequent behavior analysis, enabling the system to accurately capture the action details and pose changes of the target object. Based on the extracted skeletal key point information, through a pre-trained interaction behavior analysis model, the system can identify the interaction behaviors between target objects, such as conversations, handshakes, or shoves. This recognition ability enables the system not only to monitor individual behaviors but also to understand the interaction relationships between target objects, providing support for more complex behavior pattern analysis. According to the type and duration of the identified interaction behaviors, the system can determine the behavior pattern of the target object and classify it as normal behavior, abnormal behavior, or potentially dangerous behavior. This classification method enables the system to quickly identify and respond to potential security threats, improving the efficiency and security of monitoring. Through the above steps, the system can more accurately understand the behaviors of target objects, timely detect abnormal behaviors or potentially dangerous behaviors, thereby improving the accuracy and efficiency of monitoring. This is of great significance for ensuring the safety of the target area, especially in crowded public places such as shopping malls and stations. The determined behavior pattern can be an important input for an intelligent decision-making and early warning system, helping monitoring personnel to quickly react and take appropriate measures. For example, when a potentially dangerous behavior is detected, the system can automatically trigger an alarm to notify the monitoring personnel to intervene, thus effectively preventing and reducing the occurrence of safety accidents.
[0037] Optionally, the predicting the action trajectory of the target object according to the behavior pattern in combination with the attention mechanism includes: Constructing a behavior feature vector of the target object according to the behavior pattern, where the behavior feature vector includes skeletal key point information, interaction behavior information, and environmental information; Inputting the behavior feature vector into a prediction model based on Transformer, and introducing an attention mechanism in the prediction model to perform weighted processing on the behavior features at different time steps; Predicting the action trajectory of the target object through the prediction model, where the action trajectory includes continuing to move forward, staying, or changing direction.
[0038] The construction of the behavioral feature vector is achieved by integrating the skeletal key point information, interaction behavior information, and environmental information of the target object. The skeletal key point information includes the positions and pose angles of the joints, and these data can accurately describe the body posture and action details of the target object. The interaction behavior information covers the interaction behaviors between the target object and other objects, such as conversations, handshakes, or shoves, etc., and these behaviors are recognized by a pre-trained interaction behavior analysis model. The environmental information includes the environmental characteristics where the target object is located, such as lighting, weather, crowd density, etc., and this information has an important impact on the behavior pattern of the target object. The above multi-dimensional information is fused to form a comprehensive behavioral feature vector. This vector not only contains the body posture and action information of the target object, but also contains its interaction information with the environment and other objects, providing a rich feature basis for subsequent action trajectory prediction. The constructed behavioral feature vector is input into a prediction model based on Transformer. The Transformer model is famous for its self-attention mechanism and can effectively handle the long-range dependencies in sequence data. In the prediction model, the self-attention mechanism weights the behavioral features at different time steps, enabling the model to pay more attention to those time steps that are more crucial for predicting the action trajectory. Through the encoder and decoder structures of the Transformer model, the model can learn the temporal information and spatial information in the behavioral feature vector. The encoder is responsible for encoding the input behavioral feature vector and extracting high-level feature representations; the decoder then generates the action trajectory prediction of the target object based on the output of the encoder. The predicted action trajectory includes the possible action directions of the target object, such as continuing to move forward, staying, or changing direction, etc. The action trajectory predicted through the above process can provide real-time monitoring and early warning functions for the video monitoring system. The system can identify possible abnormal behaviors or potential dangerous behaviors of the target object in advance according to the predicted action trajectory, and thus issue an early warning in a timely manner to take corresponding measures. The prediction results can also provide decision-making support for the monitoring personnel. For example, in a crowded public place, if it is predicted that the target object may change direction or stay, the monitoring personnel can adjust the monitoring resources in advance and focus on relevant areas to ensure public safety.
[0039] Construct a behavior feature vector of the target object according to the behavior pattern, which includes skeletal key point information, interaction behavior information, and environmental information. Such a multi-dimensional feature vector can comprehensively describe the behavior state and environmental background of the target object, providing rich feature information for subsequent action trajectory prediction. Input the behavior feature vector into a prediction model based on Transformer, and introduce an attention mechanism into the model. The attention mechanism can weight the behavior features at different time steps, enabling the model to pay more attention to key behavior features and time steps, thereby improving the accuracy and robustness of the prediction. Through the prediction model, the action trajectory of the target object can be predicted, including continuing to move forward, staying, or changing direction, etc. This prediction ability enables the monitoring system to identify potential behaviors of the target object in advance, providing a basis for taking timely preventive measures, thereby improving the initiative and security of monitoring.
[0040] Optionally, the evaluating the initial risk level of the action trajectory according to the preset risk propagation model and the Bayesian network probability inference framework includes: Use the preset risk propagation model to analyze the risk propagation path of the action trajectory of the target object in the target area and identify potential risk nodes; Construct a Bayesian network including the behavior, environmental factors, and historical events of the target object according to the potential risk nodes and the risk propagation path. The second nodes of the Bayesian network represent random variables, and the second edges represent the conditional dependence relationships between the variables; Use the conditional probability table, combined with the random variables and the conditional dependence relationships in the Bayesian network, to calculate the initial risk level of the action trajectory of the target object.
[0041] Through a preset risk propagation model, analyze the possible risk propagation paths that the action trajectory of the target object may trigger within the target area. This step aims to identify potential risk nodes that may exist in the action trajectory. These nodes may be key positions or areas in the action trajectory, with a relatively high probability of risk propagation. During the process of analyzing the risk propagation paths, the system will identify potential risk nodes. These nodes may be key positions or areas in the action trajectory, with a relatively high probability of risk propagation. For example, within a monitored area, certain areas may be regarded as high-risk nodes due to dense crowds or the presence of dangerous goods. Based on the identified potential risk nodes and risk propagation paths, construct a Bayesian network that includes the behavior of the target object, environmental factors, and historical events. In this network, nodes represent random variables, and edges represent the conditional dependence relationships between variables. The nodes in the Bayesian network represent random variables, such as the behavior pattern of the target object, environmental factors (such as weather, lighting, crowd density, etc.), and historical events (such as similar events that occurred in this area in the past). Edges represent the conditional dependence relationships between these variables, that is, how the state of one variable affects the state of another variable. In the constructed Bayesian network, use the conditional probability table (CPT) to model each node. The conditional probability table details the probabilities of the current node taking each possible value given the state of the parent node. Through the conditional probability table, combined with the random variables and conditional dependence relationships in the Bayesian network, calculate the initial risk level of the target object's action trajectory. This step utilizes the probability inference ability of the Bayesian network to infer the initial risk level of the action trajectory based on the known variable states (such as the behavior pattern of the target object, environmental factors, etc.). After obtaining the initial risk level, combine it with the crowd density data displayed by the dynamic heat map to adjust the initial risk level. Areas with a relatively high crowd density may increase the risk level because a high-density crowd may lead to an accelerated risk propagation speed or an expanded influence range. Based on the adjusted risk level, determine the final risk level of the target object's action trajectory. This step ensures the accuracy and real-time nature of the risk assessment, providing a scientific basis for subsequent disposal suggestions. According to the final risk level, generate corresponding disposal suggestions. These suggestions may include emergency evacuation, resource allocation, safety reminders, etc., to help monitoring personnel promptly take appropriate countermeasures to effectively respond to and prevent potential security threats.
[0042] Analyzing the risk propagation path of the action trajectory of the target object in the target area using a preset risk propagation model can identify potential risk nodes. This helps to discover in advance the key positions that may trigger risks and provides a basis for subsequent risk assessment and disposal. According to the identified potential risk nodes and risk propagation paths, a Bayesian network containing the behavior of the target object, environmental factors, and historical events is constructed. This network represents random variables through second nodes and conditional dependencies between variables through second edges, enabling more accurate modeling and analysis of the complex relationships between risk factors. Using the conditional probability table, combined with the random variables and conditional dependencies in the Bayesian network, the initial risk level of the action trajectory of the target object is calculated. This provides a quantitative basis for subsequent comprehensive risk assessment by combining other factors such as the crowd density, making the risk assessment more scientific and accurate.
[0043] This embodiment also discloses a video monitoring system. Figure 2 It is a schematic diagram of the modules of the video monitoring system disclosed in the embodiments of the present application, as Figure 2 shown. The system includes a construction module, a stitching module, a prediction module, and a risk module, where: The construction module is configured to construct a three-dimensional digital twin scene according to the spatial coordinates of each camera in the target area. Using a graph neural network, each camera is regarded as a node, and a node relationship matrix is constructed according to the spatial relationship between the cameras; The stitching module is configured to determine the real-time position and action path of the target object using the dynamic field-of-view stitching algorithm according to the three-dimensional digital twin scene and the node relationship matrix, and generate a dynamic heat map using multi-camera field-of-view fusion to display the crowd density in the target area; The prediction module is configured to determine the behavior pattern of the target object through skeleton key point detection and interaction behavior analysis, and predict the action trajectory of the target object according to the behavior pattern combined with the attention mechanism; The risk module is configured to evaluate the initial risk level of the action trajectory according to the preset risk propagation model and the Bayesian network probability inference framework, and adjust the initial risk level in combination with the crowd density to obtain the final risk level, and generate a disposal suggestion according to the final risk level.
[0044] Optionally, the construction module is configured to: Use a 3D reconstruction network to convert the two-dimensional image information of the target image collected by each camera into three-dimensional spatial information to form a three-dimensional space, identify the objects in the target image through semantic segmentation, and map the objects into the three-dimensional space; Establish the first edges between the first nodes according to the spatial relationship between the cameras, and construct the adjacency matrix of the graph through the first edges. The elements in the adjacency matrix represent the connection weights or distances between the nodes.
[0045] Optionally, the splicing module is configured to: Construct the field of view graph of the cameras based on the three-dimensional digital twin scene and the node relationship matrix, where the field of view of each camera is a region, and the overlapping part between the fields of view is the shared region; Determine the position information of the target object in different fields of view through the field of view graph, and calculate the action path according to the position change of the target object in different fields of view, combined with the timestamp information. Use the tracking algorithm to smooth the action path of the target object.
[0046] Optionally, the splicing module is configured to: Fuse the fields of view of multiple cameras through the three-dimensional digital twin scene and the node relationship matrix to generate a unified view covering the target area; In the unified view, count the number of people appearing at the target pixel points or regional units within a preset time interval to generate the pedestrian flow density data; Generate a dynamic heat map according to the pedestrian flow density data, where the color or brightness of the heat map represents the level of pedestrian flow density, and the pedestrian flow density distribution of the target area is displayed in real time.
[0047] Optionally, the prediction module is configured to: Use the skeleton key point detection algorithm to extract the skeleton key point information of the target object, and the skeleton key point information includes the joint point position and the pose angle; According to the skeleton key point information, identify the interaction behavior of the target object through a pre-trained interaction behavior analysis model, and the interaction behavior includes conversation, handshake or pushing; Determine the behavior pattern of the target object according to the type and duration of the interaction behavior, and the behavior pattern includes normal behavior, abnormal behavior or potential dangerous behavior.
[0048] Optionally, the prediction module is configured to: Construct the behavior feature vector of the target object according to the behavior pattern, and the behavior feature vector includes skeleton key point information, interaction behavior information and environmental information; Input the behavior feature vector into the prediction model based on Transformer, and introduce the attention mechanism in the prediction model to weight the behavior features at different time steps; Predict the action trajectory of the target object through the prediction model, where the action trajectory includes moving forward, staying, or changing direction.
[0049] Optionally, the risk module is configured to: Use the preset risk propagation model to analyze the risk propagation path of the action trajectory of the target object within the target area, and identify potential risk nodes; Construct a Bayesian network including the behavior, environmental factors, and historical events of the target object according to the potential risk nodes and the risk propagation path. The second nodes of the Bayesian network represent random variables, and the second edges represent the conditional dependence relationships between the variables; Use the conditional probability table, combined with the random variables and the conditional dependence relationships in the Bayesian network, to calculate the initial risk level of the action trajectory of the target object.
[0050] It should be noted that when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0051] This embodiment also discloses an electronic device. Refer to Figure 3 , the electronic device may include: at least one processor, at least one communication bus, a user interface, a network interface, and at least one memory.
[0052] Among them, the communication bus is used to realize the connection and communication between these components.
[0053] Among them, the user interface may include a display screen (Display) and a camera (Camera). Optionally, the user interface may also include a standard wired interface and a wireless interface.
[0054] Among them, the network interface may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0055] Among them, the processor may include one or more processing cores. The processor uses various interfaces and circuits to connect various parts within the entire server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling the data stored in the memory, it performs various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately through a single chip.
[0056] Among them, the memory may include random access memory (RAM) and may also include read-only memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory may also be at least one storage device located far from the aforementioned processor. As Figure 3 shown, in a memory as a computer storage medium, there may be included an operating system, a network communication module, a user interface module, and an application program for the method of video monitoring.
[0057] In Figure 3In the electronic device shown, the user interface is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor can be used to call the application program stored in the memory that monitors video. When executed by one or more processors, the electronic device is caused to execute the method of one or more of the above embodiments.
[0058] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0059] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0060] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0061] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0062] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0063] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0064] The foregoing are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other implementation manners of the present disclosure after considering the disclosure of the specification. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A video monitoring method, characterized in that: Applied to a video monitoring platform, the method comprises: Build a three-dimensional digital twin scene based on the spatial coordinates of each camera in the target area, use a graph neural network, take each camera as the first node, and build a node relationship matrix based on the spatial relationship between cameras; A dynamic view stitching algorithm is used according to the three-dimensional digital twin scene and the node relationship matrix to determine the real-time position and action path of the target object, and a dynamic heat map is generated by multi-camera view fusion to display the crowd density of the target area; Determine the behavior pattern of the target object through skeleton key point detection and interactive behavior analysis, and predict the action trajectory of the target object based on the behavior pattern combined with the attention mechanism; The initial risk level of the action trajectory is evaluated according to a preset risk propagation model and a Bayesian network probabilistic reasoning framework, and the initial risk level is adjusted in combination with the crowd density to obtain a final risk level, and a disposal recommendation is generated according to the final risk level.
2. The video monitoring method according to claim 1, characterized in that: The method of constructing a three-dimensional digital twin scene according to the spatial coordinates of each camera in the target area, using a graph neural network, taking each camera as a node, and constructing a node relationship matrix according to the spatial relationship between the cameras includes: A 3D reconstruction network is used to convert the two-dimensional image information of the target image collected by each camera into three-dimensional space information to form a three-dimensional space, objects in the target image are identified through semantic segmentation, and the objects are mapped into the three-dimensional space; A first edge between first nodes is established according to a spatial relationship between cameras, and an adjacency matrix of a graph is constructed through the first edge, wherein elements in the adjacency matrix represent connection weights or distances between nodes.
3. The video monitoring method according to claim 1, characterized in that: The method of using a dynamic view stitching algorithm according to the three-dimensional digital twin scene and the node relationship matrix to determine the real-time position and action path of the target object includes: Constructing a camera field of view graph based on the three-dimensional digital twin scene and the node relationship matrix, wherein the field of view of each camera is an area, and the overlapping parts between the fields of view are shared areas; Determine the position information of the target object in different views through the view map, and calculate the action path according to the position change of the target object in different views in combination with the timestamp information; The movement path of the target object is smoothed using a tracking algorithm.
4. The video monitoring method according to claim 1, characterized in that: The method of generating a dynamic heat map by fusion of multiple camera fields of view to display the crowd density of the target area includes: The fields of view of multiple cameras are fused through the three-dimensional digital twin scene and the node relationship matrix to generate a unified view covering the target area; In the unified view, the number of people appearing in the target pixel point or area unit within a preset time interval is counted to generate crowd density data; A dynamic heat map is generated based on the crowd density data, wherein the color or brightness of the heat map indicates the crowd density, and the crowd density distribution of the target area is displayed in real time.
5. The video monitoring method according to claim 1, characterized in that: Determining the behavior pattern of the target object by skeleton key point detection and interactive behavior analysis includes: Utilizing a skeleton key point detection algorithm, extracting skeleton key point information of the target object, wherein the skeleton key point information includes joint point positions and posture angles; According to the skeleton key point information, identifying the target object's interactive behavior through a pre-trained interactive behavior analysis model, wherein the interactive behavior includes talking, shaking hands, or pushing; The behavior pattern of the target object is determined according to the type and duration of the interactive behavior, where the behavior pattern includes normal behavior, abnormal behavior or potentially dangerous behavior.
6. The video monitoring method according to claim 5, characterized in that: Predicting the target object's action trajectory according to the behavior pattern in combination with the attention mechanism includes: Constructing a behavior feature vector of the target object according to the behavior pattern, wherein the behavior feature vector includes skeleton key point information, interactive behavior information and environment information; Inputting the behavior feature vector into a Transformer-based prediction model, introducing an attention mechanism into the prediction model to perform weighted processing on the behavior features at different time steps; The prediction model is used to predict the movement trajectory of the target object, where the movement trajectory includes continuing to move forward, stopping, or changing direction.
7. The video monitoring method according to claim 1, characterized in that: The initial risk level of the action trajectory is evaluated according to the preset risk propagation model and the Bayesian network probability reasoning framework, including: Analyze the risk propagation path of the target object's action trajectory in the target area using the preset risk propagation model to identify potential risk nodes; Constructing a Bayesian network including the behavior, environmental factors and historical events of the target object according to the potential risk nodes and the risk propagation path, wherein the second node of the Bayesian network represents a random variable, and the second edge represents a conditional dependency relationship between the variables; The initial risk level of the target object's action trajectory is calculated by using a conditional probability table in combination with the random variables and the conditional dependency in the Bayesian network.
8. A video monitoring system, characterized in that: It includes building module, splicing module, prediction module and risk module, among which: A construction module is configured to construct a three-dimensional digital twin scene according to the spatial coordinates of each camera in the target area, using a graph neural network, taking each camera as a node, and constructing a node relationship matrix according to the spatial relationship between the cameras; A splicing module is configured to use a dynamic view splicing algorithm according to the three-dimensional digital twin scene and the node relationship matrix to determine the real-time position and action path of the target object, and use multi-camera view fusion to generate a dynamic heat map to display the crowd density of the target area; A prediction module, configured to determine the behavior pattern of the target object through skeleton key point detection and interactive behavior analysis, and predict the action trajectory of the target object according to the behavior pattern in combination with an attention mechanism; The risk module is configured to evaluate the initial risk level of the action trajectory according to a preset risk propagation model and a Bayesian network probabilistic reasoning framework, adjust the initial risk level in combination with the crowd density to obtain a final risk level, and generate a disposal recommendation based on the final risk level.
9. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Cited By
Regional risk early warning method, electronic equipment, storage medium and program product
CN120494537A
Limited space safety monitoring method, device and equipment and readable storage medium
CN120747877A
Limited space safety monitoring method, device and equipment and readable storage medium
CN120747877B
Monitoring equipment layout method and device based on digital twinning, equipment and medium
CN121173935A
Tower perimeter security and protection method and system based on multi-view fusion
CN121397187A