Complex scene real-time human body behavior recognition method and system

By using multi-camera collaboration and three-dimensional modeling technology in the video surveillance system, the problems of human behavior recognition and multi-view data fusion in complex scenarios are solved, and high-precision and high-efficiency monitoring effects are achieved, and the level of safety management in public places is improved.

CN120088862APending Publication Date: 2025-06-03WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221472.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

It is difficult for existing video surveillance systems to achieve real-time and accurate human behavior recognition and multi-view data fusion in complex scenarios, resulting in false alarms, missed alarms and inefficiency.

Method used

The real-time human behavior recognition method for complex scenes based on multi-camera collaboration and three-dimensional modeling is adopted. Through deep learning models and depth estimation technology, accurate modeling and behavior analysis of human body's three-dimensional postures is realized, and data processing is optimized through multi-view fusion technology.

Benefits of technology

It improves the intelligence level of the monitoring system and the accuracy and efficiency of scene monitoring, and can achieve real-time and accurate human behavior recognition and multi-view data fusion in complex scenarios, enhancing the safety management capabilities of public places.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088862A_ABST
    Figure CN120088862A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time human body behavior recognition method and system in a complex scene, and belongs to the field of data image processing. The method comprises the following steps: deploying a camera network in a target scene and synchronizing a timestamp, and obtaining internal and external parameters by using a Zhang Zhengyou calibration method; camera video streams are collected in real time and preprocessed; capturing video frames by using a deep learning model to carry out three-dimensional human body modeling, extracting human body three-dimensional posture information of each frame for carrying out real-time three-dimensional modeling on a human body, and correcting the depth of human body three-dimensional modeling by combining internal reference of a camera and using information in a real-time depth map; based on depth estimation and multi-camera fusion, human body three-dimensional reconstruction and unified coordinate mapping are realized; and behavior classification and data association are carried out on the reconstructed three-dimensional human body model, and human body behavior types are marked. According to the invention, building environment and character activity information can be efficiently integrated, real-time performance and accuracy are achieved, and abnormal behaviors are detected and early warned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital twins and intelligent monitoring, and particularly to a method and system for real-time human behavior recognition in complex scenarios based on multi-camera collaboration and 3D modeling. Background Art

[0002] With the rapid development of urbanization, efficiently and intelligently managing complex public places such as high-speed railway stations, airports, and shopping malls has become a key issue in modern urban management. Especially in scenarios with frequent personnel flow and complex environments, traditional monitoring means can no longer meet the requirements in terms of real-time performance, accuracy, and security.

[0003] Most current video monitoring systems rely on two-dimensional monitoring and manual intervention, lacking the ability to deeply analyze dynamic scenarios and recognize behaviors, resulting in problems such as false alarms, missed alarms, and low efficiency in personnel monitoring of multi-view and large-scale areas. Therefore, it is particularly important to develop a real-time human behavior recognition system based on 3D modeling and multi-camera collaboration.

[0004] In recent years, the continuous progress of digital twin technology, 3D modeling, and deep learning technology has made human behavior recognition and multi-camera data fusion gradually become an important research direction in the field of intelligent monitoring. By combining sensor data, computer vision, and artificial intelligence technologies, not only can the intelligent level of the monitoring system be improved, but also high-precision 3D data can be obtained in real time, providing a more comprehensive and accurate analysis of personnel activities. Although these technologies have made certain progress in many fields, they still face many challenges in real-time human behavior recognition and 3D modeling in complex scenarios, especially in large-scale and multi-view dynamic scenarios. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to propose a method for real-time human behavior recognition in complex scenarios based on multi-camera collaboration and 3D modeling in view of the limitations in the prior art. This method realizes precise real-time monitoring and behavior analysis of personnel in complex scenarios by integrating 3D modeling, deep learning, and multi-camera data fusion technologies. This technical solution not only improves the intelligent level of the monitoring system but also provides effective support for the safety management of public places such as high-speed railway stations and airports, having broad application prospects and important social value.

[0006] The technical solution adopted by the present invention to solve its technical problems is: a method for real-time human behavior recognition in complex scenarios, including the following steps:

[0007] Step 1, deploy a camera network in the target scenario and synchronize timestamps, and use the Zhang Zhengyou calibration method to obtain internal and external parameters;

[0008] Step 2: Collect the camera video stream in real time and perform preprocessing;

[0009] Step 3: Use a deep learning model to capture video frames for 3D human body modeling, extract the 3D human body pose information of each frame for real-time 3D human body modeling, and combine the internal parameters of the camera to correct the depth of the 3D human body modeling using the information in the real-time depth map;

[0010] Step 4: Based on depth estimation and multi-camera fusion, achieve 3D human body reconstruction and unified coordinate mapping;

[0011] Step 5: Classify the behavior of the reconstructed 3D human body model and associate the data, and label the human body behavior type.

[0012] Furthermore, in Step 1, the Network Time Protocol is adopted to synchronize the camera clock with the unified time source and perform time synchronization processing on the camera data;

[0013] The preprocessing in Step 2 includes compressing and encoding the original video data, dynamically adjusting the compression bit rate according to the network bandwidth, decoding the received video stream to restore the original image frames, and performing image enhancement.

[0014] Furthermore, in Step 3, the implementation method of capturing video frames for 3D human body modeling is as follows:

[0015] The first step: Determine the unified frame rate and calculate the inter-frame interval time:

[0016]

[0017] where f is the unified frame rate and ΔT represents the time interval between each frame;

[0018] Use the central server to sort and align the multi-camera frames received according to the time stamps, and the frame set {F 1 ,F 2 ,…,F n} satisfies:

[0019] ||T frame (F i )-T frame (F j )||≤Δt (2)

[0020] where F i represents the frame data captured by the i-th camera, T frame (F i ) represents the time stamp of the frame F i , and Δt refers to the maximum allowable time stamp difference, and the time stamp difference between any two frames cannot exceed Δt;

[0021] Preprocess the captured frame data, specifically including: resolution processing and pixel value normalization processing to meet the input specifications of the deep learning model. The pixel value normalization processing formula is as follows:

[0022]

[0023] where NormalizedValue represents the normalized pixel value, OriginalValue represents the original pixel value, μ is the mean of the image pixels, and σ is the standard deviation of the image pixels;

[0024] In the second step, input the image frame into the ROMP model for 3D human body modeling, and ROMP will return the 3D key points of the human skeleton:

[0025] P rel =[X rel ,Y rel ,Z rel (4)

[0026] where P rel represents the relative coordinates of the model, and X rel ,Y rel ,Z rel represent the relative positions of the point in the horizontal direction, vertical direction, and vertical space respectively.

[0027] Furthermore, in step three, the specific method for obtaining the real-time depth map is as follows:

[0028] Step1, for an ordinary camera, the method for obtaining the depth map is as follows:

[0029] In step Step11, preprocess the image: divide it into k×k small grid blocks and adjust the resolution using bilinear interpolation;

[0030] In step Step12, convert the channel order of the input image from the original format of (H, W, C) to the format of (C, H, W), where H represents the number of pixels of the image in the vertical direction, W represents the number of pixels of the image in the horizontal direction, and C represents the number of color channels of the image; convert the image data type to the float32 type and expand the batch dimension in the 0th dimension to make it meet the model input format (B, C, H, W), where B is the batch size;

[0031] In step Step13, input the preprocessed small grid blocks into the MiDaS lightweight model and output the corresponding real-time depth map, stretch it back to the original resolution and splice the k×k depth map grids.

[0032] Step2, for an RGB-D camera, directly obtain the real-time depth data and perform invalid value filling and denoising processing on the obtained depth map.

[0033] Further, in step three, the specific implementation method for correcting the depth of the human body three-dimensional modeling by using the information in the real-time depth map is as follows:

[0034] Calibrate the relative coordinates through the depth map in combination with the camera internal parameters. The formula is:

[0035] P u =[X u , Y u , X u (5)

[0036]

[0037] where D(u, v) is the depth value of the pixel point (u, v) in the depth map, (X u , Y u , Z u ) represents the three-dimensional coordinates of the depth image pixel point, and K is the camera internal parameter matrix, which is specifically as follows:

[0038]

[0039] where f x , f y are the focal lengths in the horizontal and vertical directions respectively, and c x , c y are the principal point coordinates of the camera.

[0040] Further, in order to facilitate mapping to a unified platform in step four, the parts covered by multiple camera views are subjected to multi-view fusion and optimization, that is, the human bodies in the multi-view coverage area are corresponded and preferentially retained; the determination of the same human body adopts a human joint matching mechanism. For two views a and b with an overlapping area, if the Euclidean distance of the pelvic joint in the human body coordinate sequence in a or b is not greater than the threshold, it is determined as the same person; if there are multiple coordinate sequences in b that satisfy the condition for a certain joint coordinate sequence in a, then the weighted Euclidean distances of multiple key nodes in the above human body coordinate sequences are calculated respectively, and the one with the smallest weight is determined as the same human body. For the multiple absolute three-dimensional coordinates of different views determined to be the same human body, it is considered that the less the key nodes are occluded, the more accurate the coordinate sequence is. Therefore, the model with less weighted occlusion of the key nodes is selected and retained, and the rest are discarded, so that each human body in the platform corresponds to only one human body three-dimensional coordinate sequence.

[0041] Further, the specific implementation method of step four is as follows:

[0042] In the first step, convert the relative three-dimensional coordinates to absolute three-dimensional coordinates by using the camera external parameters. The formula is as follows:

[0043] P world =[X world , Yworld ,Z world (8)

[0044]

[0045] Among them, (X u ,Y u ,Z u ) are the relative three-dimensional coordinates after deep calibration, R is the camera rotation matrix, t is the camera translation vector, and (X world ,Y world ,Z world ) are the absolute three-dimensional coordinates in the converted world coordinate system;

[0046] Second, perform multi-view fusion. The specific method is as follows:

[0047] (1), The human joint feature matching mechanism determines the same human body under multiple views. The method is as follows:

[0048] (11), First, perform preliminary matching through the pelvic joint. For the pelvic joint p b =(x b ,y b ,z b ), assuming that the coordinates obtained under different views are p b1 and p b2 respectively, calculate the Euclidean distance of the pelvic joint between the two views as d b :

[0049]

[0050] Set the human pelvic Euclidean distance determination threshold ∈. If d b ≤∈, it is determined that it is the same human body under the two views;

[0051] (12), If there are multiple candidate matches, use the weighted average of key joint points for fine matching. Assume that the key node set is {K 1 ,K 2 ,...,K 5}, corresponding to the three-dimensional coordinates of the pelvis, chest, head, left shoulder, and right shoulder. Among them, the Euclidean distance between each key node K k and other candidate views under the determination view are d 1 ,d 2 ,d 3 ,d 4 ,d 5 respectively. Calculate the Euclidean distance of the weighted average of human key nodes, and select the match corresponding to the smallest Euclidean distance as the most accurate key node matching result;

[0052] (2) For multiple absolute three-dimensional coordinates of different perspectives of a unified human body, select and retain the model with less weighted occlusion of key nodes, and discard the rest. Assume the key node set is {K 1 , K 2 ,..., K 5}. For each key node K k , if the joint is occluded, O k = 1; if not occluded, O k = 0; Calculate the occlusion degree through weighted average, and select the three-dimensional coordinate model with the smallest occlusion degree as the final three-dimensional human body model.

[0053] Furthermore, the Euclidean distance weighted average formula for human body key nodes is:

[0054]

[0055] The calculation formula for the occlusion degree is as follows:

[0056]

[0057] Furthermore, the specific implementation method of step five is as follows:

[0058] First step, based on the three-dimensional coordinate sequence, use the long short-term memory network model to extract the dynamic behavior characteristics of the human body, construct a time series data stream to capture the temporal changes of human behavior;

[0059] Second step, adopt a lightweight behavior classification model to classify the time series data of the human body joint positions;

[0060] Third step, real-time detect and update the behavior category of each human body, and associate it with the reconstructed three-dimensional human body model.

[0061] The present invention also provides a complex scene real-time human behavior recognition system, including:

[0062] A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the complex scene real-time human behavior recognition method as described in the above technical solution.

[0063] The beneficial effects produced by the present invention are:

[0064] 1. Through precise camera network deployment and timestamp synchronization, it ensures the precise alignment of data obtained from multiple perspectives, improving the accuracy and efficiency of scene monitoring; through multi-camera fusion technology and human joint feature matching mechanism, it realizes the precise fusion and optimization of human body information under different perspectives, improving the visibility and integrity under multi-perspective monitoring;

[0065] 2. By using deep learning models (such as ROMP and MiDaS models), accurate and real-time three-dimensional human pose information can be extracted from video data, and depth parameters can be corrected by combining depth estimation, improving the modeling accuracy of human poses.

[0066] 3. By adopting LSTM and lightweight behavior classification models, the dynamic behaviors of people can be recognized and classified in real time, and abnormal behaviors can be effectively detected, enhancing the intelligent monitoring ability of the system.

[0067] In summary, the method of the present invention is simple and practical. By using deep learning models for human pose recognition and behavior classification, this method can effectively process activity data in different environments, has strong adaptability, can perform real-time monitoring and display of dynamic scenes from different perspectives in a multi-camera environment, and has high reliability, practicability and feasibility, and is suitable for real-time monitoring and visualization display in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0069] Figure 1 is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0071] As Figure 1 shown, a real-time human behavior recognition method for complex scenes based on multi-camera collaboration and three-dimensional modeling provided by an embodiment of the present invention specifically includes the following steps:

[0072] Step 1. For a specific application scenario, in order to facilitate the visualization of the monitoring platform, determine the scope and accuracy requirements of the target scene, and select an appropriate data acquisition method (such as laser scanning, drone shooting, photogrammetry, etc.). Preprocess and fuse the collected point cloud data and image data, remove noise and calibration errors, and use a triangular mesh reconstruction algorithm to convert it into a three-dimensional model. In order to enhance the authenticity of the scene, map the texture information of the scene to the surface of the three-dimensional model, and select a suitable three-dimensional engine to import the processed three-dimensional model.

[0073] Step 2: According to the geometric characteristics, monitoring requirements, and environmental conditions of the target scenario, optimize the deployment locations of the cameras to ensure field of view coverage and avoid visual occlusion. To facilitate data transmission, ensure that all cameras are in the same network environment. To ensure timestamp synchronization, adopt the Network Time Protocol to synchronize the camera clocks with a unified time source and perform time synchronization processing on the camera data. The conversion between the image coordinate system, camera coordinate system, and world coordinate system requires the internal and external parameters and scaling factors of the camera. Therefore, use the Zhang Zhengyou calibration method to calibrate each camera separately to facilitate subsequent coordinate system conversion and processing.

[0074] Step 3: The videos captured in real-time by each camera are transmitted to the central server through the network. On the central server, the original video data is compressed and encoded, and the compression bitrate is dynamically adjusted according to the network bandwidth. The server decodes the received video stream, restores the original image frames, and performs preprocessing (such as image enhancement, etc.) to ensure the efficiency and accuracy of subsequent processing. Among them, the resolution adjustment is performed in the subsequent steps.

[0075] Step 4: Human pose modeling and depth estimation. Use a deep learning model (such as the ROMP model) to perform human pose modeling on each video frame and extract the three-dimensional human pose information of each frame to facilitate real-time three-dimensional modeling of the human body. Given that the three-dimensional human pose information extracted by the deep learning model may not be accurate enough for depth information, combined with the internal parameters of the camera, use the information in the depth map to correct the depth of the three-dimensional human modeling to improve the accuracy of the human modeling relative to the three-dimensional coordinates. There are the following two ideas for obtaining the depth map:

[0076] 1. If an ordinary camera is used, the MiDaS model (Ranftl et al., 2022) can be used for monocular depth estimation to generate depth maps of the human body and the scene. The input of the MiDaS model is an image with a resolution of 384×384. To improve the accuracy of the depth map, the image is divided into 4×4 small grids, the resolution is adjusted by bilinear interpolation, and then pixel normalization is performed. Finally, the depth map is stretched to the original resolution, and the small grids are merged to restore the complete depth information;

[0077] 2. If an RGB-D camera is used, the depth map can be directly obtained. In this case, only invalid value filling and denoising processing are required for the obtained depth map.

[0078] Step 5: Use the external parameters of the camera to convert the relative three-dimensional coordinates of the human body calibrated by depth in the camera coordinate system into absolute three-dimensional coordinates in the world coordinate system. To facilitate mapping to a unified platform, the overlapping parts covered by multiple camera views are subjected to multi-view fusion and optimization, that is, the human bodies in the multi-view coverage area are corresponding and preferentially retained. The determination of the same human body adopts a human joint matching mechanism. For two views a and b with an overlapping area, if the Euclidean distance of the pelvis joint in the human body coordinate sequence in a or b is not greater than the threshold, it is determined to be the same person; if for a certain joint coordinate sequence in a, there are multiple coordinate sequences in b that meet the condition, then the weighted Euclidean distance of multiple key nodes in the above human body coordinate sequences is calculated respectively, and the one with the smallest weight is determined to be the same human body. For the multiple absolute three-dimensional coordinates of different views determined to be the same human body, it is considered that the less the key nodes are blocked, the more accurate the coordinate sequence is. Therefore, the model with less weighted occlusion of key nodes is selected and retained, and the rest are discarded, so that each human body in the platform corresponds to only one human body three-dimensional coordinate sequence.

[0079] Step 6: To further enhance the real-time monitoring function of the platform, a human behavior recognition and classification module is introduced. Based on the three-dimensional coordinate sequence, a long short-term memory network (LSTM) is used to extract the dynamic behavior features of the human body and classify them into basic behaviors (such as walking, running, standing, etc.) or abnormal behaviors (such as falling, running fast, etc.).

[0080] Step 7: Use a lightweight rendering engine (such as Unity3D) to optimize the rendering pipeline, and give priority to rendering key human body information to ensure the real-time performance of the platform. Highlight processing and classification annotation display are used for crowded areas and abnormal behavior areas, so that users can pay attention to abnormal situations in time.

[0081] In Step 2, the specific processes of camera network deployment, timestamp synchronization, and internal and external parameter acquisition are as follows:

[0082] Step 201: Optimize the camera position configuration and unify the camera network environment.

[0083] Step 202: Adopt the Network Time Protocol (NTP) to record the timestamp of each frame of image and synchronize it.

[0084] Step 203: Calibrate the camera using the Zhang Zhengyou calibration method to obtain the internal and external parameters of the camera, namely the reference point and the scaling factor.

[0085] In Step 4, the specific process of using a deep learning model to capture video frames for three-dimensional human body modeling and correcting and calibrating the relative three-dimensional coordinates in combination with depth estimation is as follows:

[0086] Step 401: Capture video frames for three-dimensional human body modeling according to the following method:

[0087] Step 4011: Determine the unified frame rate and calculate the inter-frame interval time:

[0088]

[0089] Among them, f is the unified frame rate, and ΔT represents the time interval between each frame.

[0090] Use the central server to sort and align the received multi-camera frames according to timestamps. The frame set {F 1 ,F 2 ,…,F n} satisfies:

[0091] ||T frame (F i )-T frane (F j )||≤Δt (2)

[0092] Among them, F i represents the frame data captured by the i-th camera, and T frame (F i ) represents the timestamp of frame F i . Δt refers to the maximum allowable timestamp difference, that is, the timestamp difference between any two frames cannot exceed Δt;

[0093] Preprocess the captured frame data, specifically including: resolution processing and pixel value normalization processing to meet the input specifications of the deep learning model. The pixel value normalization processing formula is as follows:

[0094]

[0095] Among them, NormalizedValue represents the normalized pixel value, OriginalValue represents the original pixel value, μ is the mean of the image pixels, and σ is the standard deviation of the image pixels.

[0096] Step 4012: Input the image frame into the ROMP model (Regression of Multiple 3D people) for 3D human body modeling. ROMP will return the 3D key points of the human skeleton:

[0097]

[0098] Among them, P rel represents the relative coordinates of the model, and X rel ,Y rel ,Z rel respectively represent the relative positions of the point in the horizontal direction, vertical direction, and vertical space.

[0099] Step 402: The specific method for obtaining the real-time depth map is as follows:

[0100] Step 4021. For an ordinary camera, the method for obtaining a depth map is as follows:

[0101] Step 40211. Preprocess the image: Divide it into 4×4 small grid blocks, and use bilinear interpolation to adjust the resolution to 384×384.

[0102] Step 40212. Convert the channel order of the input image from the original (H,W,C) format to the (C,H,W) format, where H represents the number of pixels in the vertical direction of the image, W represents the number of pixels in the horizontal direction of the image, and C represents the number of color channels of the image; convert the image data type to the float32 type, and expand the batch dimension in the 0th dimension to make it meet the model input format (B,C,H,W), where B is the batch size.

[0103] Step 40213. Input the preprocessed small grid blocks into the MiDaS lightweight model and output the corresponding real-time depth map, stretch and return to the original resolution, and splice the 4×4 depth map grids.

[0104] Step 4022. For an RGB-D camera, directly obtain real-time depth data, and perform invalid value filling and denoising processing on the obtained depth map.

[0105] Step 403. Calibrate the relative coordinates through the depth map in combination with the camera internal parameters. The formula is:

[0106] P u =[X u ,Y u ,Z u (5)

[0107]

[0108] where D(u,v) is the depth value of the pixel point (u,v) in the depth map, and (X u ,Y u ,Z u ) represents the three-dimensional coordinates of the depth image pixel point, that is, the relative three-dimensional coordinates after depth calibration. K is the camera internal parameter matrix, specifically as follows:

[0109]

[0110] where f x , f y are the focal lengths in the horizontal and vertical directions respectively, and c x , c y are the principal point coordinates of the camera.

[0111] In step 5, after converting the relative three-dimensional coordinates of the human body into absolute three-dimensional coordinates, the specific process of three-dimensional multi-view fusion and mapping of the human body is as follows:

[0112] Step 501: Use the extrinsic parameters of the camera to convert the relative three-dimensional coordinates into absolute three-dimensional coordinates. The formula is as follows:

[0113] P world =[X world ,Y world ,Z world (8)

[0114]

[0115] where R is the camera rotation matrix, t is the camera translation vector, and (X world ,Y world ,Z world ) are the absolute three-dimensional coordinates in the world coordinate system after conversion.

[0116] Step 502: Perform multi-view fusion. The specific method is as follows:

[0117] Step 5021: The human joint feature matching mechanism determines the same human body under multiple views. The method is as follows:

[0118] Step 50211: First, perform a preliminary match through the pelvic joint. For the pelvic joint p b =(x b ,y b ,z b ), assuming the coordinates obtained under different views are p b1 and p b2 respectively, calculate the Euclidean distance between the pelvic joints in the two views as d b :

[0119]

[0120] Set the human pelvic Euclidean distance determination threshold ∈. If d b ≤∈, it is determined that it is the same human body under the two views;

[0121] Step 50212: If there are multiple candidate matches, use the weighted average of the key joint points for fine matching. Assume the key node set is {K 1 ,K 2 ,...,K 5}, corresponding to the three-dimensional coordinates of the pelvis, chest, head, left shoulder, and right shoulder. Among them, the Euclidean distance between each key node K k and other candidate views in the determination view is d 1 ,d 2 ,d 3 ,d4 , d 5 , the weighted average formula of the Euclidean distance of the key body nodes is as follows:

[0122]

[0123] According to the Euclidean distance D after weighted average avg , select the match corresponding to the minimum Euclidean distance as the most accurate key node matching result.

[0124] Step 5022: For multiple absolute three-dimensional coordinates of different perspectives of a determined unified human body, select and retain the model with less weighted occlusion of key nodes, and discard the rest. Assume the key node set is {K 1 , K 2 ,..., K 5}, corresponding to the three-dimensional coordinates of the pelvis, chest, head, left shoulder, and right shoulder. For each key node K k , if the joint is occluded, O k = 1; if not occluded, O k = 0; the weighted average formula is:

[0125]

[0126] where w represents the occlusion degree; select the three-dimensional coordinate model with the least occlusion as the final model.

[0127] Step 5023: Map all human bodies to the three-dimensional platform according to their absolute three-dimensional coordinates.

[0128] In step 6, the specific process of identifying and classifying human behaviors is as follows:

[0129] Step 601: Based on the three-dimensional joint positions of each human body at different time points, use models such as the long short-term memory network (LSTM) to extract the dynamic behavior characteristics of the human body, construct a time-series data stream to capture the time-series changes of human behaviors, and ensure the continuity and dynamic recognition of behaviors.

[0130] Step 602: Use a lightweight behavior classification model (such as the MobileNet-LSTM model) to classify the time-series data of human joint positions.

[0131] Step 603: Real-time detect and update the behavior category of each human body, associate it with its three-dimensional model, and synchronize the behavior changes to the digital twin platform.

[0132] In step 7, the specific process of rendering and displaying is as follows:

[0133] Step 701: Use the lightweight rendering engine Unity3D to optimize the rendering pipeline and give priority to rendering key human body information.

[0134] Step 702: Distinguish different people through color mapping or labels, and use highlighting processing and classification annotation display for crowded areas and abnormal behavior areas.

[0135] As can be seen from the above embodiments, through the application of digital twin technology, the present invention can achieve real-time monitoring and display of activities in complex scenarios. By adopting the three-dimensional building modeling and point cloud data fusion technology, the present invention can construct a high-precision virtual scene model and effectively render it through a three-dimensional engine (such as Unity3D), ensuring the authenticity and interactivity of the scene display; it can seamlessly integrate data streams, video streams and behavior analysis structures into the digital twin platform for display, meeting the real-time monitoring and data display requirements in various complex scenarios.

[0136] On the other hand, the embodiment of the present invention also provides a real-time human behavior recognition system for complex scenarios, including:

[0137] A processor and a memory, where the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the real-time human behavior recognition method for complex scenarios as described in the above technical solution.

[0138] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A real-time human behavior recognition method for complex scenes, characterized by: The steps include: Step 1: deploy a camera network in the target scene and synchronize timestamps, and use Zhang Zhengyou’s calibration method to obtain internal and external parameters; Step 2: real-time acquisition of camera video stream and preprocessing; Step 3: Use the deep learning model to capture video frames for 3D human body modeling, extract the 3D posture information of the human body in each frame, and use it to perform real-time 3D modeling of the human body. Combined with the internal parameters of the camera, the depth of the 3D modeling of the human body is corrected using the information in the real-time depth map; Step 4: Based on depth estimation and multi-camera fusion, realize 3D reconstruction of the human body and unified coordinate mapping; Step 5: classify the behavior and associate the data of the reconstructed 3D human body model, and mark the human behavior type.

2. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: In step 1, the network time protocol is used to synchronize the camera clock with a unified time source, and time synchronization processing is performed on the camera data; The preprocessing in step 2 includes compressing and encoding the original video data, dynamically adjusting the compression bit rate according to the network bandwidth, decoding the received video stream, restoring the original image frame, and performing image enhancement.

3. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: In step 3, the implementation method of grabbing video frames for 3D human body modeling is as follows: The first step is to determine the uniform frame rate and calculate the interval time between frames: Where f is the uniform frame rate, ΔT represents the time interval between each frame; Use a central server to sort and align the received multi-camera frames according to timestamps. The frame set {F1, F2, …, F n }satisfy: ||T frame (F i )-T frame (F j )||≤Δt (2) Among them, F i Represents the frame data captured by the i-th camera, T frame (F i ) represents frame F i The timestamp of the frame, Δt refers to the maximum allowed timestamp difference, and the timestamp difference between any two frames cannot exceed Δt; The captured frame data is preprocessed, including resolution processing and pixel value normalization processing, to meet the input specifications of the deep learning model. The pixel value normalization processing formula is as follows: Among them, NormalizedValue represents the normalized pixel value, OriginalValue represents the original pixel value, μ is the mean value of the image pixels, and σ is the standard deviation of the image pixels; The second step is to input the image frame into the ROMP model for 3D human body modeling. ROMP will return the 3D key points of the human skeleton: P rel =[X rel ,Y rel ,Z rel ] (4) Among them, P rel Indicates the relative coordinates of the model, X rel ,Y rel ,Z rel They respectively represent the relative position of a point in the horizontal direction, vertical direction, and vertical space.

4. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: In step 3, the specific method for obtaining the real-time depth map is: Step 1: For ordinary cameras, the method to obtain the depth map is: Step 11, preprocess the image: divide it into k×k small grid blocks, and adjust the resolution using bilinear interpolation; Step 12: Convert the channel order of the input image from the original format (H, W, C) to the format (C, H, W), where H represents the number of pixels of the image in the vertical direction, W represents the number of pixels of the image in the horizontal direction, and C represents the number of color channels of the image; convert the image data type to float32 type, and expand the batch dimension in the 0th dimension to satisfy the model input format (B, C, H, W), where B is the batch size; In step 13, the preprocessed small grid blocks are input into the MiDaS lightweight model and the corresponding real-time depth map is output, which is stretched back to the original resolution and the k×k depth map grids are spliced. Step 2: For RGB-D cameras, directly obtain real-time depth data, and perform invalid value filling and denoising on the obtained depth map.

5. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: In step 3, the specific implementation method of using the information in the real-time depth map to correct the depth of the three-dimensional modeling of the human body is as follows: Combined with the camera internal parameters, the relative coordinates are calibrated through the depth map. The formula is: P u =[X u ,Y u ,Z u ] (5) Where D(u,v) is the depth value of the pixel (u,v) in the depth map, (X u ,Y u ,Z u ) represents the three-dimensional coordinates of the depth map pixel, and K is the camera intrinsic parameter matrix, which is as follows: Among them, f x , f y For the horizontal and vertical focal lengths, c x , c y is the principal point coordinate of the camera.

6. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: In step 4, in order to facilitate mapping to a unified platform, the parts covered by multiple camera perspectives are fused and optimized from multiple perspectives, that is, the human body in the area covered by multiple perspectives is matched and preferentially retained; The determination of the same human body adopts the human joint matching mechanism. For the two perspectives a and b with overlapping areas, if the Euclidean distance of the pelvic joint in a human coordinate sequence in a and b is not greater than the threshold, it is determined to be the same person; if for a joint coordinate sequence in a, there are multiple coordinate sequences in b that meet the condition, then the weighted Euclidean distances of multiple key nodes in the above human coordinate sequence are calculated respectively, and the one with the smallest weight is determined to be the same person. For multiple absolute three-dimensional coordinates of different perspectives determined to be the same human body, it is determined that the less occlusion of the key nodes, the more accurate the coordinate sequence, so the model with less weighted occlusion of the key nodes is selected to be retained, and the rest is discarded, so that each human body in the platform corresponds to only one human three-dimensional coordinate sequence.

7. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: The specific implementation of step 4 is as follows: The first step is to use the camera external parameters to convert the relative three-dimensional coordinates into absolute three-dimensional coordinates. The formula is as follows: P world =[X world ,Y world ,Z world ] (8) Among them, (X u ,Y u ,Z u ) is the relative 3D coordinate after depth calibration, R is the camera rotation matrix, t is the camera translation vector, (X world ,Y world ,Z world ) is the absolute three-dimensional coordinate in the world coordinate system after conversion; The second step is to perform multi-view fusion. The specific method is as follows: (1) The human joint feature matching mechanism determines the same human body from multiple perspectives. The method is as follows: (11), firstly, the pelvic joint is preliminarily matched, and for the pelvic joint p b =(x b ,y b ,z b ), assuming that the coordinates obtained at different viewing angles are p b1 and p b2 , calculate the Euclidean distance of the pelvic joint under two viewing angles as d b : Set the human pelvis Euclidean distance judgment threshold ∈, if d b ≤∈, it is determined to be the same person from two perspectives; (12), if there are multiple candidate matches, the weighted average of the key joint points is used for fine matching. Assume that the key node set is {K1, K2, ..., K5}, corresponding to the three-dimensional coordinates of the pelvis, chest, head, left shoulder, and right shoulder, where each key node K k The Euclidean distances of the other candidate viewpoints under the judgment viewpoint are d1, d2, d3, d4, and d5 respectively. The weighted average Euclidean distance of the key nodes of the human body is calculated, and the match corresponding to the minimum Euclidean distance is selected as the most accurate key node matching result; (2) For multiple absolute 3D coordinates of different viewpoints of a unified human body, the model with less weighted occlusion of key nodes is selected and the rest are discarded. Assuming that the key node set is {K1, K2, ..., K5}, for each key node K k , if the joint is occluded, O k =1; if not blocked, O k =0; the occlusion degree is obtained by weighted average, and the three-dimensional coordinate model with the smallest occlusion degree is selected as the final three-dimensional human body model.

8. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: The weighted average formula of the Euclidean distance of key nodes of the human body is: The calculation formula for the occlusion degree is as follows:

9. The complex scene real-time human behavior recognition method according to claim 1, characterized in that: The specific implementation of step five is as follows: The first step is to extract the dynamic behavior characteristics of the human body based on the three-dimensional coordinate sequence using the long short-term memory network model and construct a time series data stream to capture the temporal changes of human behavior. In the second step, a lightweight behavior classification model is used to classify the time series data of human joint positions; The third step is to detect and update the behavior category of each human body in real time and associate it with the reconstructed 3D human body model.

10. Complex scene real-time human behavior recognition system, characterized by: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the complex scene real-time human behavior recognition method as described in any one of claims 1 to 9.