A security scene abnormal behavior recognition system based on multi-source video space-time alignment

CN122657197BActive Publication Date: 2026-09-22ZHUHAI QUANBAO NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611149292.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-22
Estimated Expiration
2046-07-31

AI Technical Summary

Technical Problem

[0002]在医院等复杂医疗环境中,人员流动频繁、区域功能多样化,现有视频监控系统通常依赖单一摄像头或分散部署的摄像头网络进行监控,但在多楼层、多通道、多功能区域的场景中,单摄像头视角存在明显局限性,难以实现对医护人员、患者及陪护人员的跨区域连续追踪

Benefits of technology

本发明提供一种基于多源视频时空对齐的安全场景异常行为识别系统。本发明通过分区域多站点激光扫描获取医院各楼层结构及功能空间点云,并进行统计滤波、体素下采样和全局拼接,构建带有墙体、地面及固定设备语义标注的静态三维背景模型及统一空间坐标体系,为多摄像头视频数据的空间映射和动态目标定位提供精确基准。利用多路视频的帧级时间对齐、极线校正与立体匹配,结合重叠视野区域的跨视角融合,生成全局点云纹理信息并提取动态目标,实现跨摄像头连续的三维动态融合模型。同时,本发明通过动态物体的三维空间位置与外观特征完成跨帧身份关联和跨摄像头轨迹合并,并对稀疏或遮挡点云进行结构插值补全,保证人体骨架及刚体目标结构完整。结合麦克风阵列采集的环境声音和同步视频数据提取多模态行为特征,并通过行为-空间与行为-角色一致性映射进行判定,实现异常事件的统计分析与告警生成。在三维空间模型中高亮渲染并动态提示异常区域,并可通过镜头交互选择最优视频源叠加显示实时图像,嵌入三维空间进行动态渲染更新,从而全面提升医院安全场景下的异常行为发现、追踪和响应能力,确保监控的实时性、准确性及全面覆盖。通过本发明提供的一种基于多源视频时空对齐的安全场景异常行为识别系统,能够实现对医院全院范围内人员行为的精确监控和智能化管理,显著提升异常行为识别的准确性和及时性,保障医院运行环境的安全与高效。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122657197B_ABST
    Figure CN122657197B_ABST
Patent Text Reader

Abstract

The application discloses a security scene abnormal behavior recognition system based on multi-source video space-time alignment, comprising: a static background modeling module for constructing a static three-dimensional background model with semantic annotation and a three-dimensional space coordinate system; a dynamic three-dimensional reconstruction module for generating three-dimensional reconstruction data across cameras; a cross-view tracking completion module for cross-frame identity association and cross-camera trajectory merging; a multi-modal behavior recognition module for determining personnel behavior classification results; a medical semantic anomaly detection module for counting semantic anomaly candidate segments and generating abnormal alarm information; and a real scene interactive rendering module for integrating and embedding dynamic objects and event annotation data into a three-dimensional space model for real-time rendering update. The application can realize accurate monitoring and intelligent management of personnel behavior in the whole hospital, significantly improve the accuracy and timeliness of abnormal behavior recognition, and ensure the safety and efficiency of the hospital operation environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and image processing technology, and in particular to a security scene abnormal behavior recognition system based on multi-source video spatiotemporal alignment. Background Technology

[0002] In complex medical environments such as hospitals, where personnel flow is frequent and area functions are diverse, existing video surveillance systems typically rely on single cameras or distributed camera networks for monitoring. However, in scenarios with multiple floors, multiple channels, and multiple functions, the perspective of a single camera has significant limitations, making it difficult to achieve continuous cross-area tracking of medical staff, patients, and caregivers. Single-camera monitoring is easily affected by factors such as limited perspective, obstruction, changes in lighting, and insufficient coverage, leading to loss of target identity, broken trajectory, and inaccurate dynamic behavior recognition. Furthermore, traditional systems often rely solely on video image information, lacking the ability to comprehensively analyze multimodal features such as sound, facial expressions, and body movements, making it difficult to make real-time judgments of abnormal behaviors in complex hospital scenarios. Abnormal events in hospital environments are diverse, including patient falls, sudden calls for help, and doctor-patient conflicts. These events often occur on different floors, in different rooms, or in corridors, making it difficult for a single camera to capture complete information simultaneously, resulting in insufficient sensitivity and accuracy in anomaly recognition. Meanwhile, due to the lack of a unified data structure and a hospital-wide dynamic integration mechanism, monitoring personnel often cannot quickly locate the spatial position of an event or associate it with a specific target when an anomaly is detected, increasing the difficulty of security management and response time. While existing multi-camera systems can provide wider coverage, they still face technical bottlenecks in cross-view target continuous tracking, 3D spatial coordinate mapping, and real-time visualization of dynamic targets, especially in maintaining target identity continuity, ensuring the integrity of posture information, and the application of multimodal information fusion. Therefore, how to achieve continuous cross-view tracking by multiple cameras throughout the hospital, combine multimodal information such as sound, facial expressions, and limb movements for abnormal behavior identification, and achieve real-time interactive visualization and unified data integration through 3D spatial models has become a crucial technical issue for improving the efficiency and accuracy of hospital security monitoring. Summary of the Invention

[0003] This invention provides a security scene abnormal behavior recognition system based on multi-source video spatiotemporal alignment, mainly comprising: The static background modeling module is used to acquire raw point cloud data by multi-site laser scanning in different regions, perform statistical filtering, voxel downsampling and ICP global stitching, and construct a static 3D background model with semantic annotation and a 3D spatial coordinate system. The dynamic 3D reconstruction module is used to obtain global point cloud texture information and extract dynamic targets by stereo matching and depth back projection of multiple video frames based on the spatial pose parameters of the camera and the 3D spatial coordinate system, and to generate cross-camera 3D reconstruction data. The cross-view tracking and completion module is used to perform cross-frame identity association and cross-camera trajectory merging based on the spatial position and appearance features of the 3D model of dynamic objects, and to perform structural interpolation completion for sparse or occluded parts of the point cloud. The multimodal behavior recognition module is used to extract audio spectrum features, facial expression features, and human skeleton limb movement trajectories from environmental sound signals and synchronous video data collected by microphone arrays pre-deployed in the hospital setting, and to determine the classification results of personnel behavior. The medical semantic anomaly detection module is used to statistically analyze semantic anomaly candidate segments within a continuous time window and generate anomaly alarm information based on the semantic labels of medical functional areas in the 3D spatial model, the target behavior categories of personnel, and the relationship between the participating target roles, through behavior-space and behavior-role consistency mapping. The real-scene interactive rendering module is used to select the optimal video source and capture the real-time video image area by the camera's viewpoint position when the camera zooms in, and integrate the dynamic object and event annotation data into the 3D spatial model for real-time rendering and updates.

[0004] Furthermore, the static background modeling module is used to acquire raw point cloud data through multi-site laser scanning in different regions, perform statistical filtering, voxel downsampling, and ICP global stitching to construct a static 3D background model with semantic annotation and a 3D spatial coordinate system, including: Based on the physical location information of multiple cameras deployed within the hospital setting and the deployment parameters of the laser point cloud scanning equipment, a multi-site, area-based laser scanning method was used to acquire raw point cloud data of the hospital's floor structure, passageways, and functional spaces. Statistical filtering was applied to the acquired point cloud data to remove outlier noise points, and a voxel mesh downsampling method was used to unify point density, resulting in standardized point cloud data. The point clouds from different scanning sites were globally stitched together using the ICP registration algorithm to obtain an overall spatial point cloud model. Based on this overall spatial point cloud model, an initial 3D structural model was constructed using surface reconstruction and triangulation methods. A semantic segmentation model was then used to classify and label the contours of walls, floors, and fixed equipment, resulting in a static 3D background model with semantic information. Based on the calibration image data acquired during the camera calibration process, the Zhang Zhengyou calibration method was used to calculate the intrinsic and extrinsic parameter matrices, constructing a camera projection matrix. The transformation relationship between the camera coordinate system and the 3D model coordinate system was determined, and the poses of each camera were uniformly mapped to the 3D spatial coordinate system to obtain 3D spatial reference data.

[0005] Furthermore, the dynamic 3D reconstruction module is used to obtain global point cloud texture information based on the camera's spatial pose parameters and the 3D spatial coordinate system through stereo matching and depth back projection of multiple video frames, and to extract dynamic targets to generate cross-camera 3D reconstruction data, including: Based on the camera's spatial pose parameters and the 3D spatial coordinate system, frame-level timestamp alignment and epipolar correction are performed on multiple video streams to obtain corrected video images. A semi-global stereo matching algorithm is used to perform stereo matching calculations on the disparity map, and the disparity values ​​are converted into depth values ​​through triangulation to generate an initial dense depth map. Based on the camera's extra-field depth map, image pixels are mapped to 3D space through back projection to obtain 3D point cloud data corresponding to the video frames. Using a field-of-view overlap coordinate merging unit, the field-of-view overlap area is determined according to the field of view of each camera. Image sub-regions in the overlap area and single-view data in the non-overlapping area are extracted and merged across the field of view spatial coordinates to obtain global spatial point cloud texture information. The process involves comparing global spatial point cloud texture information with a static 3D background model to extract dynamic target point cloud regions with positional deviations exceeding a preset threshold. Voxel filtering, mesh reconstruction, and skeleton extraction are then performed on the dynamic point clouds to obtain a 3D model of the dynamic object. This 3D model is mapped to a 3D spatial coordinate system to generate a 3D dynamic fusion model. Based on the spatial overlap between multiple cameras, consistency optimization and boundary smoothing algorithms are used to adjust the fusion results, resulting in a continuous 3D spatial model across cameras. By establishing a mapping relationship between video content and 3D space, dynamic video information is continuously mapped into 3D space to obtain global 3D reconstruction data containing dynamic objects.

[0006] It also includes a field-of-view overlap coordinate merging unit, used to determine the field-of-view overlap area based on the field of view of each camera, extract the image sub-region of the overlap area and the single-view data of the non-overlapping area, and perform cross-view spatial coordinate unified merging to obtain global spatial point cloud texture information, specifically including: The overlapping area between adjacent cameras is calculated based on the field of view of each camera. Image sub-regions belonging only to the overlapping area are extracted as input for cross-view fusion, while non-overlapping areas are directly retained as single-view data. For image frames in the overlapping area, feature points are extracted and descriptors are calculated using SIFT or ORB algorithms. A set of overlapping area frames with temporal consistency is obtained through spatiotemporal alignment. The RANSAC algorithm is used to remove feature point matching outliers, obtaining high-confidence cross-view feature correspondences within the overlapping area. Through epipolar geometric constraints and projection consistency verification algorithms, different projection positions of the same target within the overlapping area are jointly corrected to minimize reprojection errors. The corrected overlapping area results are merged with the single-view data of the non-overlapping area using unified spatial coordinates to form cross-view video frames, which are then mapped into three-dimensional space. High-resolution voxel mapping is used in active target areas, while voxel precision is reduced in inactive areas to control the computational load, resulting in global spatial point cloud texture information.

[0007] Furthermore, the cross-view tracking and completion module is used to perform cross-frame identity association and cross-camera trajectory merging based on the spatial position and appearance features of the dynamic object's 3D model, and to perform structural interpolation completion for sparse or occluded parts of the point cloud, including: Based on the 3D model of the dynamic object and its spatial coordinates, a temporary tracking identifier is assigned to each dynamic object; the appearance feature vector of each dynamic object in the video frame is extracted and compared with the feature library of historical frames. Based on the comparison results, the cross-frame identity association is determined, and the temporary identifier is mapped to a persistent identity label; in response to the identity identifier break caused by the movement of the dynamic object between different camera fields of view, the position, velocity vector and appearance features of the object before it disappears in the previous camera are recorded and spatiotemporally matched with the position, velocity vector and appearance features of the newly appearing target in the next camera. If the match is successful, they are merged into a cross-camera continuous trajectory with the same identity. To address the incomplete structure extraction caused by sparse point clouds or occlusion in 3D models of dynamic objects, key structural points of the corresponding targets in video frames are detected, and prediction and smoothing are performed through temporal filtering. The coordinates of the predicted key structural points are back-projected into 3D space, and interpolation is performed to complete the missing parts or parts with confidence scores below the threshold in the original structure extraction results. Specifically, for human targets, the positions of human joints are completed to form complete skeleton data, and for non-human rigid body targets, the key points of the structural contour and the positions of the bounding box vertices are completed. Based on the completed data, the category label, identity label, continuous 3D trajectory, and corresponding posture or motion state sequence of each dynamic object are obtained.

[0008] Furthermore, the multimodal behavior recognition module is used to extract audio spectrum features, facial expression features, and human skeletal limb movement trajectories from environmental sound signals and synchronous video data collected by a microphone array pre-deployed in the hospital setting, and to determine the personnel behavior classification results, including: A microphone array pre-deployed within the hospital setting collects ambient sound signals. Short-time Fourier transform is used to extract audio spectral features, and a sound source calculation method based on time-of-arrival (TOA) is employed to determine the spatial direction and distance of the sound, obtaining spatial location data. This spatial location data is mapped to a 3D spatial model using spatial coordinate transformation and matched with target trajectory data to determine the correlation between sound and people. Based on synchronously acquired video data, a convolutional neural network is used to extract and classify facial region images, obtaining facial expression features. A 3D convolutional neural network is used to detect the coordinates of key human points and connect them to form a skeleton sequence, obtaining limb movement trajectory features. Sound features, facial expression features, and limb movement trajectory features are stored in a personnel behavior monitoring database. Through this database, historical data and behavior classification labels for personnel sound features, facial expression features, and limb movement trajectory features are obtained. A random forest algorithm is used to train the model, constructing a personnel behavior classification model. Based on real-time acquired personnel sound features, facial expression features, and limb movement trajectory features, the personnel behavior classification model is used to determine personnel behavior category labels.

[0009] Furthermore, the medical semantic anomaly detection module is used to statistically analyze candidate semantic anomalies within a continuous time window based on the semantic labels of medical functional areas in the three-dimensional spatial model, the categories of personnel target behaviors, and the relationships of participating target roles, through behavior-space and behavior-role consistency mapping, and to generate anomaly alarm information, including: Based on the hospital's functional area information marked in the 3D spatial model, a spatial semantic label is assigned to each spatial coordinate point, including emergency room, intensive care unit, general ward, corridor, infusion area, and operating area. Based on the spatial semantic label and behavior category label corresponding to the current personnel's spatial coordinates, a behavior-space consistency mapping table is queried to determine the judgment status of the target personnel's behavior under the current spatial semantics. The status includes normal, abnormal, or requiring review. If the result is abnormal, it is judged as spatial semantic inconsistency. Based on the role relationship category between the currently participating targets, a behavior-role consistency mapping table is queried to determine the judgment status of the target personnel's behavior under the current role relationship. Role relationships include medical staff-patient, medical staff-caregiver, and patient-patient. If the result is abnormal, it is determined to be an inconsistency in role relationship; if the behavior, spatial semantics, and role relationship are all normal, it is determined to be normal behavior; if any one of the behavior, spatial semantics, or role relationship is abnormal, the corresponding time window is marked as a semantic abnormality candidate segment; statistical analysis is performed on the semantic abnormality candidate segments within the continuous time window. If the number of consecutive occurrences of the candidate segment within the preset time range exceeds the preset threshold, it is determined to be an abnormal event. The abnormal event is mapped to the corresponding area in the 3D spatial model, and an abnormal alarm information containing event type, occurrence timestamp, spatial coordinates, target identifier, and behavior category label is generated and pushed to the monitoring terminal. The abnormal area is highlighted and dynamically flashed in the 3D spatial model.

[0010] Furthermore, the real-scene interactive rendering module is used to filter the optimal video source and extract a real-time video image region based on the camera's viewpoint position when the camera zooms in, and to integrate and embed dynamic object and event annotation data into a three-dimensional spatial model for real-time rendering updates, including: Based on a static 3D background model and the spatial pose parameters of each camera, a bidirectional mapping relationship is established between the spatial coordinates of the 3D model and the pixel coordinates of the camera images. When a user zooms in on the camera in the 3D model interface, the 3D spatial coordinates corresponding to the current center point of the field of view are calculated based on the current viewpoint position, line of sight, and field of view parameters. Candidate cameras that can cover these spatial coordinates are then selected. Based on the angle between the line of sight and the camera's optical axis, and the distance between the camera and the target point, the optimal video source is selected. The pixel coordinate region of this spatial coordinate point in the original video frame is calculated through back projection. The corresponding image region is extracted from the real-time video stream and superimposed on the 3D model interface. If the camera zooms out, the detail preview mode is automatically exited. All category labels, identity labels, continuous trajectories, posture sequences, sound association results, and event annotation data of all dynamic objects are integrated through a unified data structure. The integrated results are embedded into the 3D spatial model for real-time rendering and dynamic updates.

[0011] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention provides a security scene abnormal behavior recognition system based on multi-source video spatiotemporal alignment. The invention acquires point clouds of the structural and functional spaces of each floor of a hospital through multi-site laser scanning in different regions. Statistical filtering, voxel downsampling, and global stitching are then performed to construct a static 3D background model with semantic annotations for walls, floors, and fixed equipment, along with a unified spatial coordinate system. This provides a precise benchmark for spatial mapping and dynamic target localization of multi-camera video data. Utilizing frame-level temporal alignment, epipolar correction, and stereo matching of multiple video streams, combined with cross-view fusion of overlapping visual areas, global point cloud texture information is generated, and dynamic targets are extracted, achieving a continuous 3D dynamic fusion model across cameras. Simultaneously, the invention completes cross-frame identity association and cross-camera trajectory merging through the 3D spatial position and appearance features of dynamic objects, and performs structural interpolation to complete sparse or occluded point clouds, ensuring the integrity of the human skeleton and rigid body target structures. Multimodal behavioral features are extracted by combining environmental sound data collected by a microphone array and synchronous video data, and judgment is made through behavior-space and behavior-role consistency mapping, enabling statistical analysis and alarm generation of abnormal events. This invention highlights and dynamically highlights abnormal areas in a 3D spatial model, and allows for interactive selection of the optimal video source to overlay and display real-time images. It embeds these images into the 3D space for dynamic rendering and updates, thereby comprehensively enhancing the ability to detect, track, and respond to abnormal behavior in hospital security scenarios, ensuring real-time monitoring, accuracy, and comprehensive coverage. The security scenario abnormal behavior recognition system based on multi-source video spatiotemporal alignment provided by this invention enables precise monitoring and intelligent management of personnel behavior throughout the hospital, significantly improving the accuracy and timeliness of abnormal behavior recognition, and ensuring the safety and efficiency of the hospital's operating environment. Attached Figure Description

[0012] Figure 1 This is a flowchart of a security scene abnormal behavior recognition system based on multi-source video spatiotemporal alignment according to the present invention; Figure 2 This is a schematic diagram of a security scene abnormal behavior recognition system based on multi-source video spatiotemporal alignment according to the present invention. Detailed Implementation

[0013] To make the objectives, technical strategies, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The user information, such as user tags, obtained was with the consent of the users and complies with relevant laws and policies.

[0014] like Figures 1-2 This embodiment discloses a security scene abnormal behavior recognition system based on multi-source video spatiotemporal alignment, which may specifically include: Step S101, the static background modeling module, is used to acquire raw point cloud data by using multi-site laser scanning in different regions, perform statistical filtering, voxel downsampling and ICP global stitching, and construct a static 3D background model with semantic annotation and a 3D spatial coordinate system.

[0015] Based on the physical location information of multiple cameras deployed within the hospital setting and the deployment parameters of the laser point cloud scanning equipment, a multi-site, area-based laser scanning method was used to acquire raw point cloud data of the hospital's floor structure, corridor areas, and functional spaces. Statistical filtering was applied to the acquired point cloud data to remove outlier noise points, and a voxel mesh downsampling method was used to unify point density, resulting in standardized point cloud data. The point clouds from different scanning sites were globally stitched together using the ICP registration algorithm to obtain an overall spatial point cloud model. Based on this overall spatial point cloud model, an initial 3D structural model was constructed using surface reconstruction and triangulation methods. A semantic segmentation model was then used to classify and label the contours of walls, floors, and fixed equipment, obtaining a static 3D background model with semantic information. Based on the calibration image data acquired during the camera calibration process, the Zhang Zhengyou calibration method was used to calculate intrinsic and extrinsic parameter matrices, constructing a camera projection matrix. The transformation relationship between the camera coordinate system and the 3D model coordinate system was determined, and the poses of each camera were uniformly mapped to the 3D spatial coordinate system to obtain 3D spatial reference data.

[0016] For example, a six-story general hospital was modeled in 3D. Each floor has an area of ​​approximately 2,500 square meters, with a total building area of ​​approximately 15,000 square meters. Eighteen laser scanning stations were deployed within the hospital, each covering an area of ​​approximately 800-1000 square meters, with regional scanning to ensure full coverage. At each station, the laser scanner acquired approximately 6 million raw point cloud data points per second. After statistical filtering, approximately 3% of outliers were removed. Then, voxel mesh downsampling was used to standardize the point cloud data to 5 points per cubic centimeter. Subsequently, the point clouds from the 18 stations were globally stitched together using the ICP registration algorithm. The initial average registration error was approximately 0.08 meters, exceeding the preset threshold of 0.05 meters. Therefore, approximately 500 mismatched points were identified, and the transformation matrix was recalculated for iterative optimization. After four iterations, the average registration error finally converged to 0.048 meters. Simultaneously, the relative poses of each station were recorded for subsequent multi-camera dynamic video fusion and spatial reference verification. Based on the global point cloud model, surface reconstruction and triangulation were performed on each floor's corridor, ward, operating room, and nurse station to construct an initial 3D structural model. A semantic segmentation model was then used to classify and label walls, floors, fixed equipment outlines, beds, and medical equipment, achieving a semantic representation of the static 3D background model. During the camera calibration phase, images of the calibration board were acquired, and the intrinsic and extrinsic parameter matrices for each camera were calculated using Zhang Zhengyou's method. A camera projection matrix was constructed, and the coordinates of each camera were uniformly mapped to a 3D spatial coordinate system. This yielded the spatial coordinates of a camera installed on the third-floor corridor as (15.2 meters, 8.5 meters, 3.2 meters), with a pitch angle of -10° and an azimuth angle of 90°. This 3D spatial reference data was used for subsequent hospital-wide video data mapping and 3D dynamic fusion, ensuring spatial consistency and dynamic scene visualization under multi-camera coverage.

[0017] Step S102, the dynamic 3D reconstruction module is used to obtain global point cloud texture information and extract dynamic targets by stereo matching and depth back projection of multiple video frames based on the camera spatial pose parameters and 3D spatial coordinate system, and generate cross-camera 3D reconstruction data.

[0018] Based on the camera's spatial pose parameters and the 3D spatial coordinate system, a time synchronization mechanism is used to perform frame-level timestamp alignment on multiple video streams to obtain a synchronized video frame sequence. Epipolar correction is then applied to the video frames using multi-view geometric constraints to obtain the corrected video image. Based on the camera's spatial pose parameters and the 3D spatial coordinate system, a semi-global stereo matching algorithm is used to perform stereo matching calculations on the disparity map. The disparity values ​​are converted into depth values ​​using triangulation principles to generate an initial dense depth map. Based on the camera's extra-field depth map, image pixels are mapped to the 3D spatial coordinate system using a back-projection method to obtain the 3D point cloud data corresponding to the video frames. Using a field-of-view overlap coordinate merging unit, the overlapping areas are determined based on the field of view of each camera. Image sub-regions in the overlapping areas and single-view data from non-overlapping areas are extracted and merged across viewpoints to obtain global spatial point cloud texture information. Based on a static 3D background model, the global spatial point cloud texture information is compared with the static background model. Dynamic target point cloud regions with positional deviations exceeding a preset threshold are extracted. Voxel filtering, mesh reconstruction, and skeleton extraction are then performed on the dynamic point clouds to obtain a 3D model of the dynamic object. Based on video frames and 3D spatial reference data, the 3D models of dynamic objects are mapped onto a 3D spatial coordinate system to generate a 3D dynamic fusion model of the hospital scene. According to the spatial overlap areas between multiple cameras, consistency optimization and boundary smoothing algorithms are used to adjust the fusion results, obtaining a continuous 3D spatial model across cameras. Based on the fused global 3D reconstruction model, by establishing a mapping relationship between video content and 3D space, dynamic video information (excluding static modeling) is continuously mapped into 3D space, obtaining global 3D reconstruction data containing dynamic objects.

[0019] For example, 64 cameras were deployed throughout the hospital, approximately 10 to 12 per floor, covering wards, corridors, emergency areas, operating rooms, nurses' stations, and public areas. Each camera had a field of view of approximately 90 degrees, with about 40% overlap between adjacent cameras. Multiple video streams underwent frame-level timestamp alignment at a frame rate of 30 frames per second, with a corrected resolution of 1920×1080 pixels. Multi-view geometric correction was performed using epipolar constraints. For each video frame, a semi-global stereo matching algorithm was used to generate a disparity map, containing approximately 2.5 million pixels per frame. Triangulation was used to convert the disparity into depth values, generating an initial dense depth map with an average depth error of approximately 1.8 cm per pixel. Based on backprojection from the camera's external depth map, each frame generated a 3D point cloud containing approximately 2.2 million points. For overlapping areas, sub-region images were extracted and merged with single-view data from non-overlapping areas to obtain global spatial point cloud texture information. This information was then compared with a static 3D background model, and point cloud regions with positional deviations exceeding 5 cm were extracted as dynamic targets. Voxel filtering is applied to the dynamic point cloud, downsampling the number of points per cubic centimeter from 25 to 5. Mesh reconstruction and skeleton extraction are then performed to generate a 3D model of the dynamic object. For example, when a nurse pushes a medicine cart through a corridor, its dynamic model occupies a space of approximately 2 meters × 0.8 meters × 1.5 meters in 3D space. The dynamic object model is mapped to a 3D coordinate system, and consistency optimization and boundary smoothing algorithms are applied in overlapping areas across cameras to control the positional deviation of the same target under different camera views to within 1 centimeter, thereby generating a 3D dynamic fusion model covering the entire hospital. Based on this, a mapping relationship between video frames and 3D space is established to achieve real-time updates of all dynamic information except for the static background, ultimately obtaining global 3D dynamic reconstruction data including medical staff, patients, and mobile devices.

[0020] The field-of-view overlap coordinate merging unit is used to determine the field-of-view overlap area based on the field of view range of each camera, extract the image sub-region of the overlap area and the single-view data of the non-overlapping area, and perform cross-view spatial coordinate merging to obtain global spatial point cloud texture information.

[0021] The overlapping area between adjacent cameras is calculated based on the field of view of each camera. Image sub-regions belonging only to the overlapping area are extracted as input for cross-view fusion, while non-overlapping areas are directly retained as single-view data. For image frames in the overlapping area, feature points are extracted and descriptors are calculated using SIFT or ORB algorithms. A set of overlapping area frames with temporal consistency is obtained through spatiotemporal alignment. The RANSAC algorithm is used to remove feature point matching outliers, obtaining high-confidence cross-view feature correspondences within the overlapping area. Through epipolar geometric constraints and projection consistency verification algorithms, different projection positions of the same target within the overlapping area are jointly corrected to minimize reprojection errors. The corrected overlapping area results are merged with the single-view data of the non-overlapping areas using unified spatial coordinates to form cross-view video frames, which are then mapped into 3D space. High-resolution voxel mapping is used in active target areas, while voxel precision is reduced in inactive areas to control the computational load, resulting in global spatial point cloud texture information.

[0022] For example, based on the field of view of each camera, the overlapping area image sub-region is first extracted as the cross-view fusion input, while the remaining non-overlapping areas are directly retained as single-view data. For each frame of the overlapping area, the SIFT algorithm is used to extract approximately 2500 feature points and calculate descriptors. A time-consistent frame set is obtained through timestamp alignment. The RANSAC algorithm is used to remove approximately 5% of outlier matching points, retaining high-confidence cross-view feature correspondences. On this basis, the projection position of the same target in the overlapping area is jointly corrected. If, in the overlapping camera area of ​​a six-story corridor in a hospital, the initial projection position of a patient has a pixel deviation of approximately 5 cm measured on camera A and camera B, the shortest distance between the patient's projection point on the image plane of camera A and the corresponding epipolar line of camera B is calculated using epipolar geometric constraints. This distance is used to adjust the projection position of camera B. Combined with projection consistency verification, it is checked whether the reprojection error of the adjusted projection point on the two cameras is less than a preset threshold of 1 cm. After three rounds of iterative optimization, the final deviation of the patient's position mapped in three-dimensional space is controlled within 1 cm. After correcting overlapping areas, the data is merged with single-view data from non-overlapping areas using unified spatial coordinates to form a complete cross-view video frame, which is then mapped to 3D space. In dynamic, active areas such as corridors and emergency room passages, high-resolution voxel mapping is used, retaining approximately 8 points per cubic centimeter, while in wards and inactive areas, this is reduced to 2 points per cubic centimeter to control computational load. The final result is a global spatial point cloud texture information covering the entire hospital, with each floor containing approximately 20 to 25 million 3D point clouds.

[0023] Step S103, cross-view tracking and completion module, is used to perform cross-frame identity association and cross-camera trajectory merging based on the spatial position and appearance features of the dynamic object 3D model, and to perform structural interpolation completion for sparse or occluded parts of the point cloud.

[0024] Based on the 3D model of a dynamic object and its position coordinates in 3D space, a temporary tracking identifier is assigned to each dynamic object. The appearance feature vector of each dynamic object in the corresponding region of a video frame is extracted, and the similarity of the extracted appearance feature vector in the current frame with the appearance feature vectors in a database of historical frames is compared. Based on the comparison results, cross-frame identity association is determined, and a mapping from temporary identifiers to persistent identity tags is performed. To address the issue of identity fragmentation caused by viewpoint switching when a dynamic object moves between different camera views, the last position, velocity vector, and appearance feature vector of the dynamic object before disappearing in the previous camera are recorded. These are then spatiotemporally matched with the position, velocity vector, and appearance feature vector of a newly appearing target in the view of the next camera. If a match is successful, the tracks are merged into a continuous cross-camera trajectory with the same identity. To address the incomplete structure extraction in 3D models of dynamic objects due to sparse point clouds or occlusion, key structural points of the corresponding targets in video frames are detected. Temporal filtering is used to predict and smooth the positions of these key structural points, and the predicted coordinates are back-projected into 3D space. For missing parts or parts with confidence scores below a preset threshold in the original structure extraction results, interpolation is performed to complete the data. For targets identified as human bodies, the completion targets are the joint positions, forming complete skeleton data. For targets identified as non-human rigid bodies, the completion targets are the structural contour key points and bounding box vertex positions. Based on the complete skeleton data of human body targets and the structural contour key points and bounding box vertex positions of non-human rigid body targets, the 3D model of the dynamic object is completed, resulting in a category label, identity label, continuous 3D trajectory, and corresponding posture or motion state sequence for each dynamic object.

[0025] For example, when performing dynamic 3D tracking throughout the hospital, the system assigns a temporary tracking identifier to each moving target. For instance, patient A and nurse B moving in the corridor are assigned the identifiers 1001 and 1002, respectively. The system extracts the appearance feature vector of each target from video frames. For example, patient A's color histogram, clothing texture, and body contour features in a particular frame form a feature vector of approximately 512 dimensions. This vector is then compared with the feature vectors in a database from historical frames. If the similarity exceeds 0.85, the temporary identifier is mapped to a persistent identity label. During the process of Patient A moving from the third-floor emergency corridor to the fourth-floor ward, the identification marker was interrupted due to the camera's perspective shift. The coordinates (15.2m, 8.5m, 3.2m), velocity vector (1.2m / s, 0m / s, 0m / s), and appearance feature vector of Patient A before disappearing on the third floor were recorded. These were then compared with the initial position (15.3m, 8.6m, 6.2m), velocity vector (1.1m / s, 0.1m / s, 0m / s), and appearance feature vector of a newly appearing target on the fourth floor for spatiotemporal continuity matching. Successful matching resulted in the trajectories being merged to form a continuous trajectory across cameras. For the 3D model of the dynamic object, due to occlusion or sparse point clouds in the emergency and surgical areas, key structural points of Patient A were detected as missing. Temporal filtering was used to predict the positions of key points on the arms and legs, and the predicted coordinates were back-projected into 3D space. Points with a confidence level below 0.6 in the original model were interpolated to complete the data, ultimately yielding the complete skeleton data. Similarly, for nurse B pushing the medicine cart, the key points of its structural outline and the positions of the bounding box vertices are completed to form a complete rigid body model. After the above processing, each dynamic object has a clear category label throughout the hospital, including patient or medical staff, identity label, continuous three-dimensional trajectory, and precise posture or motion state sequence, enabling continuous tracking and visualization of dynamic behaviors across floors and areas in three-dimensional space.

[0026] Step S104: The multimodal behavior recognition module is used to extract audio spectrum features, facial expression features, and human skeleton limb movement trajectories from environmental sound signals and synchronous video data collected by a microphone array pre-deployed in the hospital setting, and determine the personnel behavior classification results.

[0027] A microphone array pre-deployed within the hospital setting collects ambient sound signals. Short-time Fourier transform is used to extract audio spectral features, and a sound source calculation method based on time-of-arrival (TOA) is employed to determine the spatial direction and distance of the sound, obtaining spatial location data. This spatial location data is mapped to a 3D spatial model using spatial coordinate transformation and matched with target trajectory data to determine the correlation between sound and people. Based on synchronously acquired video data, a convolutional neural network is used to extract and classify facial region images, obtaining facial expression features. A 3D convolutional neural network is used to detect the coordinates of key points on the human body and connect them to form a skeleton sequence, obtaining limb movement trajectory features. Sound features, facial expression features, and limb movement trajectory features are stored in a personnel behavior monitoring database. Through this database, historical data and behavior classification labels for personnel sound features, facial expression features, and limb movement trajectory features are obtained. A random forest algorithm is used to train the model, constructing a personnel behavior classification model. Based on real-time acquired personnel sound features, facial expression features, and limb movement trajectory features, the personnel behavior classification model is used to determine personnel behavior category labels.

[0028] For example, a hospital deployed a 48-microphone array to collect ambient sound signals in real time, with an audio sampling rate of 44.1 kHz. Spectral features were extracted within a 20-millisecond window using short-time Fourier transform, and the spatial direction and distance of the sound source were calculated using the time-of-arrival method. For instance, if a patient shouted in the fourth-floor infusion area, the system measured the sound source location as (12.5 meters, 8.7 meters, 3.0 meters). This spatial location of the sound was mapped to a 3D model and matched with the patient's trajectory detected in the corresponding video frame, establishing a correlation between the sound and the person. In the synchronously acquired video data, facial features were extracted using a convolutional neural network to obtain a 256-dimensional facial expression feature vector. Simultaneously, a 3D convolutional neural network was used to detect key points on the human body and connect them to form a skeleton sequence, extracting limb movement trajectory features. 30 frames of dynamic skeleton data were generated per second and stored in the personnel behavior monitoring database. By utilizing a personnel behavior monitoring database, approximately 500,000 historical behavior samples collected from across the hospital over the past three months were obtained. These samples included voice characteristics, facial expression characteristics, and limb movement trajectories, and were labeled with behavior categories such as normal walking, falling and calling for help, and disputes. A random forest algorithm was used to train the model, with the number of trees in the forest set to 500, the maximum depth to 20, the minimum number of sample splits to 10, and the maximum number of features to be the square root of the total number of features. The trained personnel behavior classification model can analyze voice characteristics, facial expression characteristics, and limb movement trajectories collected in real time across the hospital, and output behavior category labels in real time every second. It can accurately identify abnormal behaviors such as falls, calls for help, or abnormal disputes, and combine spatial semantic labels to determine the location where the behavior occurred, thereby obtaining real-time three-dimensional behavior monitoring data covering the entire hospital.

[0029] Step S105, Medical semantic anomaly detection module, is used to statistically analyze semantic anomaly candidate segments within a continuous time window and generate anomaly alarm information based on the semantic labels of medical functional areas in the three-dimensional spatial model, the target behavior categories of personnel, and the relationship of participating target roles, through behavior-space and behavior-role consistency mapping.

[0030] Based on the hospital's functional area information already labeled in the 3D spatial model, each spatial coordinate point is assigned a corresponding spatial semantic label. These labels include emergency room, intensive care unit, general ward, corridor, infusion area, and operating area. Based on the spatial semantic label corresponding to the current personnel target's spatial coordinates and the currently identified behavior category label, a preset behavior-spatial consistency mapping table is queried to determine the judgment status of the personnel target's behavior category under the current spatial semantic label. The preset behavior-spatial consistency mapping table records the judgment status of each behavior category under each spatial semantic label, with judgment statuses including normal, abnormal, or requiring review. If the query result is abnormal, it is judged as spatial semantic inconsistency. Based on the role relationship categories between the currently participating targets and the preset behavior-role consistency mapping table, the judgment status of the current behavior category under the current role relationship is queried. Role relationship categories include medical staff-patient, medical staff-caregiver, and patient-patient. The preset behavior-role consistency mapping table records the judgment status of each behavior category under each role relationship category, with judgment statuses including normal, abnormal, or requiring review. If the query result is abnormal, it is determined to be due to inconsistent role relationships. If the behavior category, spatial semantics, and role relationships are all normal, it is determined to be normal behavior. If any of the behavior category, spatial semantics, or role relationships is abnormal, the corresponding time window is marked as a semantic anomaly candidate segment. Statistical analysis is performed based on the frequency of occurrence of semantic anomaly candidate segments within consecutive time windows. If the number of consecutive occurrences of a candidate segment within a preset time range exceeds a preset threshold, it is determined to be an abnormal event. The abnormal event is mapped to the corresponding area in the 3D spatial model according to its spatial location, involved targets, and behavior type, and an anomaly alarm message containing the event type, occurrence timestamp, spatial coordinates, involved target identifiers, and behavior category labels is generated. The anomaly alarm message is pushed to the monitoring terminal, and the corresponding area of ​​the abnormal event is highlighted and dynamically flashed in the 3D spatial model.

[0031] For example, the 3D spatial model precisely labels key areas on each floor. Spatial semantic labels are assigned to the emergency room, intensive care unit, general wards, corridors, infusion area, and operating area, with a semantic resolution of 0.1 meters for each spatial coordinate point to accurately map dynamic behavior. In the third-floor infusion area, a patient A's location coordinates are (12.5 meters, 8.7 meters, 3.0 meters), and their behavior is identified as "intense shouting." By querying the preset behavior-spatial consistency mapping table, this behavior is determined to be abnormal in the infusion area, indicating spatial semantic inconsistency. Simultaneously, the patient's role relationship with medical staff B is medical staff-patient, which is also determined to be abnormal through the behavior-role consistency mapping table, indicating role inconsistency. Therefore, this time window is marked as a semantically abnormal candidate segment. Within a consecutive 15-second time window, this candidate segment appeared 6 times, exceeding the preset threshold of 3 times. The system determined this event to be an abnormal event and generated an abnormal alarm message containing the event type "abnormal shouting," the occurrence timestamp "May 6, 2026, 14:23:15," spatial coordinates (12.5 meters, 8.7 meters, 3.0 meters), the target identifiers A and B, and the behavior category label "intense shouting." Simultaneously, the corresponding area in the infusion area was highlighted and dynamically flashed in the 3D spatial model. Meanwhile, in the corridor on the same floor, a medical staff member C pushed a medicine cart at coordinates (15.2 meters, 5.0 meters, 3.2 meters). The behavior was identified as normal movement and pushing, the spatial semantic label was "corridor," and the behavior-space consistency mapping table determined it to be normal. The role relationship was medical staff-patient, which was also determined to be normal through the behavior-role consistency mapping table. Therefore, the system determined this behavior to be normal, did not generate an alarm, and maintained a static display in the 3D spatial model. In addition, in a general ward on the fourth floor, a patient D engaged in a "pulling" behavior while interacting with another patient E. The spatial coordinates were (18.3 meters, 6.5 meters, 6.0 meters), and the spatial semantic label was "general ward." This behavior was judged as normal in the spatial mapping table. However, the role relationship between patient D and patient E was patient-patient, which was judged as abnormal through the behavior-role consistency mapping table. Therefore, it was judged as a role relationship inconsistency. This time window was marked as a semantic anomaly candidate segment. After statistical analysis of continuous time windows, if the number of consecutive occurrences exceeded the preset threshold of 3 times, an abnormal event was generated and an alarm was pushed. At the same time, the corresponding area of ​​the fourth-floor ward was highlighted and dynamically flashed in the 3D spatial model.

[0032] Step S106, Real-scene interactive rendering module, is used to select the optimal video source and capture the real-time video image area by the viewpoint position of the camera when zooming in, and integrate the dynamic object and event annotation data into the three-dimensional space model for real-time rendering and updating.

[0033] Based on a static 3D background model and the spatial pose parameters of each camera, a bidirectional mapping relationship is established between the spatial coordinates of the 3D model and the pixel coordinates of the camera images. When a user zooms in on the 3D model interface, the 3D spatial coordinates corresponding to the current center point of the field of view are calculated based on the current viewpoint position, viewing direction, and field of view parameters. Candidate cameras that can cover these spatial coordinates are then selected. Based on the angle between the viewing direction and the camera's optical axis, and the distance between the camera and the target point, the optimal video source is selected. The pixel coordinate region of this spatial coordinate point in the original video frame is calculated through back projection, and the corresponding image region is extracted from the real-time video stream and superimposed on the 3D model interface. If the camera zooms out, the detail preview mode is automatically exited. All category labels, identity labels, continuous trajectories, posture sequences, sound association results, and event annotation data of all dynamic objects are integrated through a unified data structure. The integrated results are embedded into the 3D spatial model for real-time rendering and dynamic updates.

[0034] For example, the system calibrated all 64 cameras in the hospital. After calculating the intrinsic and extrinsic parameter matrices of each camera using Zhang Zhengyou's calibration method, the system established a correspondence between the coordinates of any point in the 3D spatial model and the pixel coordinates of the camera. For instance, the 3D coordinates of point P in the third-floor corridor are (15.2 meters, 6.8 meters, 3.0 meters). The system can project this point onto the camera image plane using the projection matrix of camera A, with corresponding pixel coordinates of (950, 540). At the same time, the system can also back-project the pixel (950, 540) in the camera image back into the 3D space to reconstruct its 3D coordinates as (15.2 meters, 6.8 meters, 3.0 meters), with an error of less than 1 centimeter. Through similar calibration and mapping of all cameras in the hospital to the 3D model, a bidirectional mapping between any point in the 3D space and the pixel coordinates of the camera is achieved. When monitoring personnel zoom in on the 3D model interface of the third-floor emergency corridor, the system calculates the 3D spatial coordinates (15.0m, 6.5m, 3.0m) corresponding to the center point of the field of view based on the current viewpoint position (14.8m, 6.2m, 3.2m), the direction of the line of sight, and the field of view angle parameters. It then selects three candidate cameras covering these spatial coordinates and further evaluates the optical axis angle between the candidate cameras and the direction of the line of sight, as well as the distance from the camera to the target point. The camera with the smallest angle to the line of sight and a distance of approximately 5 meters is selected as the optimal video source. The system calculates the pixel area of ​​this 3D spatial coordinate point in the original video frame of the selected camera through back projection, extracts an image area of ​​approximately 400×400 pixels, and overlays it onto the 3D model interface to achieve a preview of local details. If the personnel zoom out more than 10 meters, the system automatically exits the detail preview mode and restores the image to a global view. At the same time, all dynamic object category labels, identity labels, continuous 3D trajectories, posture sequences, sound association results, and abnormal event annotation data are integrated through a unified data structure and embedded into a 3D spatial model. This enables real-time rendering and dynamic updates at 30 frames per second in the interface, allowing monitoring personnel to intuitively view the dynamic behavior and event status of medical staff, patients, and equipment throughout the hospital.

[0035] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the concept of this application. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A security scene abnormal behavior recognition system based on multi-source video spatiotemporal alignment, characterized in that, The system includes: The static background modeling module is used to acquire raw point cloud data by multi-site laser scanning in different regions, perform statistical filtering, voxel downsampling and ICP global stitching, and construct a static 3D background model with semantic annotation and a 3D spatial coordinate system. The dynamic 3D reconstruction module is used to obtain global point cloud texture information and extract dynamic targets by stereo matching and depth back projection of multiple video frames based on the spatial pose parameters of the camera and the 3D spatial coordinate system, and to generate cross-camera 3D reconstruction data. The cross-view tracking and completion module is used to perform cross-frame identity association and cross-camera trajectory merging based on the spatial position and appearance features of the 3D model of dynamic objects, and to perform structural interpolation completion for sparse or occluded parts of the point cloud. The multimodal behavior recognition module is used to extract audio spectrum features, facial expression features, and human skeleton limb movement trajectories from environmental sound signals and synchronous video data collected by microphone arrays pre-deployed in the hospital setting, and to determine the classification results of personnel behavior. The medical semantic anomaly detection module is used to statistically analyze semantic anomaly candidate segments within a continuous time window and generate anomaly alarm information based on the semantic labels of medical functional areas in the 3D spatial model, the target behavior categories of personnel, and the relationship between the participating target roles, through behavior-space and behavior-role consistency mapping. The real-scene interactive rendering module is used to filter the optimal video source and capture the real-time video image area by the viewpoint position when the camera zooms in, and integrate and embed dynamic object and event annotation data into the three-dimensional space model for real-time rendering and updating. The medical semantic anomaly detection module is used to statistically analyze candidate semantic anomalies within a continuous time window based on the semantic labels of medical functional areas in a three-dimensional spatial model, the categories of personnel target behaviors, and the relationships between participating target roles. This is achieved through behavior-space and behavior-role consistency mapping, and the module generates anomaly alarm information, including: Based on the hospital's functional area information marked in the 3D spatial model, a spatial semantic label is assigned to each spatial coordinate point, including emergency room, intensive care unit, general ward, corridor, infusion area, and operating area. Based on the spatial semantic label and behavior category label corresponding to the current personnel's spatial coordinates, a behavior-space consistency mapping table is queried to determine the judgment status of the target personnel's behavior under the current spatial semantics. The status includes normal, abnormal, or requiring review. If the result is abnormal, it is judged as spatial semantic inconsistency. Based on the role relationship category between the currently participating targets, a behavior-role consistency mapping table is queried to determine the judgment status of the target personnel's behavior under the current role relationship. Role relationships include medical staff-patient, medical staff-caregiver, and patient-patient. If the result is abnormal, it is determined to be an inconsistency in role relationship; if the behavior, spatial semantics, and role relationship are all normal, it is determined to be normal behavior; if any one of the behavior, spatial semantics, or role relationship is abnormal, the corresponding time window is marked as a semantic abnormality candidate segment; statistical analysis is performed on the semantic abnormality candidate segments within the continuous time window. If the number of consecutive occurrences of the candidate segment within the preset time range exceeds the preset threshold, it is determined to be an abnormal event. The abnormal event is mapped to the corresponding area in the 3D spatial model, and an abnormal alarm information containing event type, occurrence timestamp, spatial coordinates, target identifier, and behavior category label is generated and pushed to the monitoring terminal. The abnormal area is highlighted and dynamically flashed in the 3D spatial model.

2. The system according to claim 1, wherein, The static background modeling module is used to acquire raw point cloud data through multi-site laser scanning in different regions, perform statistical filtering, voxel downsampling, and ICP global stitching to construct a static 3D background model with semantic annotation and a 3D spatial coordinate system, including: Based on the physical location information of multiple cameras deployed within the hospital setting and the deployment parameters of the laser point cloud scanning equipment, a multi-site, area-based laser scanning method was used to acquire raw point cloud data of the hospital's floor structure, passageways, and functional spaces. Statistical filtering was applied to the acquired point cloud data to remove outlier noise points, and a voxel mesh downsampling method was used to unify point density, resulting in standardized point cloud data. The point clouds from different scanning sites were globally stitched together using the ICP registration algorithm to obtain an overall spatial point cloud model. Based on this overall spatial point cloud model, an initial 3D structural model was constructed using surface reconstruction and triangulation methods. A semantic segmentation model was then used to classify and label the contours of walls, floors, and fixed equipment, resulting in a static 3D background model with semantic information. Based on the calibration image data acquired during the camera calibration process, the Zhang Zhengyou calibration method was used to calculate the intrinsic and extrinsic parameter matrices, constructing a camera projection matrix. The transformation relationship between the camera coordinate system and the 3D model coordinate system was determined, and the poses of each camera were uniformly mapped to the 3D spatial coordinate system to obtain 3D spatial reference data.

3. The system according to claim 1, wherein, The dynamic 3D reconstruction module is used to acquire global point cloud texture information based on the camera's spatial pose parameters and the 3D spatial coordinate system through stereo matching and depth back projection of multiple video frames, and to extract dynamic targets to generate cross-camera 3D reconstruction data, including: Based on the camera's spatial pose parameters and the 3D spatial coordinate system, frame-level timestamp alignment and epipolar correction are performed on multiple video streams to obtain corrected video images. A semi-global stereo matching algorithm is used to perform stereo matching calculations on the disparity map, and the disparity values ​​are converted into depth values ​​through triangulation to generate an initial dense depth map. Based on the camera's extra-field depth map, image pixels are mapped to 3D space through back projection to obtain 3D point cloud data corresponding to the video frames. Using a field-of-view overlap coordinate merging unit, the field-of-view overlap area is determined according to the field of view of each camera. Image sub-regions in the overlap area and single-view data in the non-overlapping area are extracted and merged across the field of view spatial coordinates to obtain global spatial point cloud texture information. The process involves comparing global spatial point cloud texture information with a static 3D background model to extract dynamic target point cloud regions with positional deviations exceeding a preset threshold. Voxel filtering, mesh reconstruction, and skeleton extraction are then performed on the dynamic point clouds to obtain a 3D model of the dynamic object. This 3D model is mapped to a 3D spatial coordinate system to generate a 3D dynamic fusion model. Based on the spatial overlap between multiple cameras, consistency optimization and boundary smoothing algorithms are used to adjust the fusion results, resulting in a continuous 3D spatial model across cameras. By establishing a mapping relationship between video content and 3D space, dynamic video information is continuously mapped into 3D space to obtain global 3D reconstruction data containing dynamic objects.

4. The system according to claim 3, wherein, The aforementioned field-of-view overlap coordinate merging unit determines the field-of-view overlap region based on the field of view of each camera, extracts image sub-regions from the overlap region and single-view data from the non-overlapping region, and performs unified merging of cross-view spatial coordinates to obtain global spatial point cloud texture information, including: The overlapping area between adjacent cameras is calculated based on the field of view of each camera. Image sub-regions belonging only to the overlapping area are extracted as input for cross-view fusion, while non-overlapping areas are directly retained as single-view data. For image frames in the overlapping area, feature points are extracted and descriptors are calculated using SIFT or ORB algorithms. A set of overlapping area frames with temporal consistency is obtained through spatiotemporal alignment. The RANSAC algorithm is used to remove feature point matching outliers, obtaining high-confidence cross-view feature correspondences within the overlapping area. Through epipolar geometric constraints and projection consistency verification algorithms, different projection positions of the same target within the overlapping area are jointly corrected to minimize reprojection errors. The corrected overlapping area results are merged with the single-view data of the non-overlapping area using unified spatial coordinates to form cross-view video frames, which are then mapped into three-dimensional space. High-resolution voxel mapping is used in active target areas, while voxel precision is reduced in inactive areas to control the computational load, resulting in global spatial point cloud texture information.

5. The system according to claim 1, wherein, The cross-view tracking and completion module is used to perform cross-frame identity association and cross-camera trajectory merging based on the spatial position and appearance features of the dynamic object's 3D model, and to perform structural interpolation completion for sparse or occluded parts of the point cloud, including: Based on the 3D model of the dynamic object and its spatial coordinates, a temporary tracking identifier is assigned to each dynamic object; the appearance feature vector of each dynamic object in the video frame is extracted and compared with the feature library of historical frames. Based on the comparison results, the cross-frame identity association is determined, and the temporary identifier is mapped to a persistent identity label; in response to the identity identifier break caused by the movement of the dynamic object between different camera fields of view, the position, velocity vector and appearance features of the object before it disappears in the previous camera are recorded and spatiotemporally matched with the position, velocity vector and appearance features of the newly appearing target in the next camera. If the match is successful, they are merged into a cross-camera continuous trajectory with the same identity. To address the incomplete structure extraction caused by sparse point clouds or occlusion in 3D models of dynamic objects, key structural points of the corresponding targets in video frames are detected, and prediction and smoothing are performed through temporal filtering. The coordinates of the predicted key structural points are back-projected into 3D space, and interpolation is performed to complete the missing parts or parts with confidence scores below the threshold in the original structure extraction results. Specifically, for human targets, the positions of human joints are completed to form complete skeleton data, and for non-human rigid body targets, the key points of the structural contour and the positions of the bounding box vertices are completed. Based on the completed data, the category label, identity label, continuous 3D trajectory, and corresponding posture or motion state sequence of each dynamic object are obtained.

6. The system according to claim 1, wherein, The multimodal behavior recognition module is used to extract audio spectrum features, facial expression features, and human skeletal limb movement trajectories from environmental sound signals and synchronous video data collected by a microphone array pre-deployed in the hospital setting, and to determine the personnel behavior classification results, including: A microphone array pre-deployed within the hospital setting collects ambient sound signals. Short-time Fourier transform is used to extract audio spectral features, and a sound source calculation method based on time-of-arrival (TOA) is employed to determine the spatial direction and distance of the sound, obtaining spatial location data. This spatial location data is mapped to a 3D spatial model using spatial coordinate transformation and matched with target trajectory data to determine the correlation between sound and people. Based on synchronously acquired video data, a convolutional neural network is used to extract and classify facial region images, obtaining facial expression features. A 3D convolutional neural network is used to detect the coordinates of key human points and connect them to form a skeleton sequence, obtaining limb movement trajectory features. Sound features, facial expression features, and limb movement trajectory features are stored in a personnel behavior monitoring database. Through this database, historical data and behavior classification labels for personnel sound features, facial expression features, and limb movement trajectory features are obtained. A random forest algorithm is used to train the model, constructing a personnel behavior classification model. Based on real-time acquired personnel sound features, facial expression features, and limb movement trajectory features, the personnel behavior classification model is used to determine personnel behavior category labels.

7. The system according to claim 1, wherein, The real-scene interactive rendering module is used to filter the optimal video source and extract a real-time video image region based on the camera's viewpoint position when the camera zooms in. It then integrates and embeds the dynamic object and event annotation data into a 3D spatial model for real-time rendering updates, including: Based on the static 3D background model and the spatial pose parameters of each camera, a two-way mapping relationship is established between the spatial coordinates of the 3D model and the pixel coordinates of the camera image. When a user zooms in on the camera in the 3D model interface, the system calculates the 3D spatial coordinates corresponding to the current center point of the field of view based on the current viewpoint position, line of sight, and field of view parameters, and filters candidate cameras that can cover these spatial coordinates. Based on the angle between the line of sight and the camera's optical axis, and the distance between the camera and the target point, the system selects the optimal video source and calculates the pixel coordinate region of the 3D spatial coordinates in the original video frame through back projection. The corresponding image region is then extracted from the real-time video stream and superimposed on the 3D model interface. If the camera zooms out, the system automatically exits the detail preview mode. All category labels, identity labels, continuous trajectories, posture sequences, sound association results, and event annotation data of all dynamic objects are integrated through a unified data structure, and the integrated results are embedded into the 3D spatial model for real-time rendering and dynamic updates.

Citation Information

Patent Citations

  • Dormitory personnel abnormal behavior early warning method and system based on millimeter wave radar

    CN122090601A

  • Scene activity analysis using statistical and semantic features learnt from object trajectory data

    US20120170802A1