Railway station three-dimensional video generation method and system based on adaptive perspective

By deploying multiple sensors and depth sensors in railway passenger stations, and combining building information models and comprehensive scoring functions, adaptive 3D videos are generated, solving the problems of single viewpoint and occlusion in traditional monitoring systems, and realizing efficient and comprehensive 3D video monitoring.

CN120751105BActive Publication Date: 2026-03-03BEIJING GUOTIE HUACHEN COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511251162.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-03-03
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Traditional railway station monitoring systems mainly use two-dimensional cameras, which have a single perspective and cannot meet the requirements of full coverage and real-time performance in complex scenes. They also suffer from occlusion and blind spots. Existing three-dimensional reconstruction methods lack real-time performance and completeness in dynamic crowd scenes.

Method used

Multiple cameras and depth sensors are deployed in railway passenger stations. Video streams and point cloud data are collected through a network clock synchronization mechanism to generate a real-time sensor dataset with a unified timestamp. A static geometric model is constructed by combining it with building information modeling. The best virtual viewpoint is selected through a comprehensive scoring function for 3D rendering to generate an adaptive 3D video.

Benefits of technology

It achieves more comprehensive and three-dimensional scene perception, reduces occlusion, improves the quality and comprehensiveness of video surveillance, avoids the occlusion and blind spot problems in traditional methods, and provides more efficient 3D video rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751105B_ABST
    Figure CN120751105B_ABST
Patent Text Reader

Abstract

The application discloses a railway station three-dimensional video generation method and system based on adaptive view angle, and relates to the technical field of passenger station video monitoring, wherein the method comprises the following steps: arranging multiple cameras and depth sensors in a railway station, and synchronously collecting video streams and point cloud data by using a network clock synchronization mechanism to generate a real-time sensing data set with a unified timestamp; constructing a static geometric model based on the building information model of the passenger station, and fusing the static geometric model with the real-time data set to generate a dynamic scene point cloud; for multiple candidate virtual view angles, calculating view angle scores by comprehensively evaluating the coverage, occlusion rate, view angle transformation cost and key area weight of the virtual view angles; selecting the best virtual view angle with the highest score, and generating corresponding three-dimensional video frames by three-dimensional rendering of the dynamic scene point cloud based on the virtual view angle. Through the comprehensive evaluation mechanism of the virtual view angle, adaptive monitoring coverage is realized, and the monitoring precision and view angle flexibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of passenger station video surveillance technology, and in particular to a method and system for generating 3D video of railway passenger stations based on adaptive perspective. Background Technology

[0002] Railway passenger stations are key hubs in modern urban public transportation systems. Because waiting, boarding, alighting, and commercial / leisure activities are concentrated in a limited space, passenger density is high and mobility is strong, posing enormous pressure on safety management and operational scheduling. Video surveillance, as a primary safety technology, has been widely deployed in areas such as platforms, waiting rooms, entrances / exits, and passageways, using closed-circuit television systems to monitor passenger behavior, equipment operation, and emergencies in real time.

[0003] Traditional railway station monitoring systems primarily use two-dimensional cameras, with security personnel observing the footage from each camera on multiple screens in a monitoring room. In this model, the cameras are fixed in location, offering a single perspective and lacking depth information, making it difficult to accurately analyze complex scenarios such as obstructions, crowding, and lingering crowds. Human surveillance also suffers from fatigue and missed detections. In real-world environments, personnel evacuation guidance, abnormal behavior identification, and emergency response require real-time and comprehensive information, which two-dimensional video based on a fixed perspective cannot effectively represent the spatial relationships of events, especially in multi-level station buildings and elevated transportation spaces. While multi-camera systems expand coverage, frequent image switching and the lack of a global three-dimensional perspective still fail to meet the goal of "full coverage with no blind spots" in monitoring.

[0004] In recent years, with the development of artificial intelligence and computer vision technologies, some research has begun to explore the deployment of smart cameras in railway stations to automatically identify anomalies through behavioral analysis algorithms. However, most of these technologies are based on monocular or binocular perspectives, emphasizing the detection of pedestrians, objects, or abnormal events in a single image, and their performance depends on target resolution and the degree of occlusion. Ozer et al., in describing their multi-camera intelligent monitoring system, pointed out that traditional target detection and tracking algorithms are typically suitable for low-occlusion, high-resolution scenes, while complex backgrounds and lighting changes can lead to failures in foreground extraction and tracking. Their system requires hardware to process large amounts of data and cope with changes in lighting flicker, shadows, and rapid exposure adjustments. This indicates that traditional two-dimensional monitoring has technical bottlenecks in areas such as lighting, occlusion, and multi-view calibration.

[0005] On the other hand, multi-view 3D reconstruction technology has been widely studied in the field of computer vision. Reconstructing a scene using multiple cameras or a moving camera can yield 3D point clouds and depth information. Typical multi-baseline stereo algorithms require selecting reference viewpoints and calculating depth by mapping pixel correspondences between different viewpoints. Princeton University's multi-view reconstruction tutorial points out that multi-baseline stereo is prone to occlusion problems when the number of cameras increases or the baseline is large, and occlusion contributes error values ​​to the cost function. Volumetric methods such as voxel coloring can explicitly consider occlusion to some extent, but determining which viewpoints each voxel is visible from and arranging the scanning order remain challenges. Traditional 3D reconstruction methods are typically geared towards static scenes and struggle to meet the real-time and completeness requirements of dynamic crowd scenes in railway stations. Summary of the Invention

[0006] This application provides a method, system, storage medium, computer program product, and electronic device for generating 3D videos of railway passenger stations based on adaptive perspective, which at least solves the problems of fixed perspective, limited coverage, occlusion, and insufficient real-time performance in current related technologies for 2D monitoring systems.

[0007] In a first aspect, embodiments of this application provide a method for generating 3D video of a railway station based on an adaptive perspective. The method includes: deploying multiple cameras and multiple depth sensors at the railway station; synchronously collecting time-series video streams and point cloud data from each sensor through a network clock synchronization mechanism; generating a real-time sensor dataset with a unified timestamp; constructing a static geometric model reflecting the layout of railway station facilities based on the building information model of the railway station; and fusing the static geometric model with the real-time sensor dataset to generate a corresponding dynamic scene point cloud; and for each candidate virtual perspective in the candidate virtual perspective set, statistically analyzing the coverage, occlusion rate, and perspective transformation cost corresponding to the candidate virtual perspective. The system calculates the corresponding viewpoint score by assigning weights to key areas and calling a comprehensive scoring function. The virtual viewpoint is the viewpoint of the virtual camera. The system selects the best virtual observation viewpoint with the highest corresponding viewpoint score and performs 3D rendering on the dynamic scene point cloud based on the best virtual observation viewpoint to generate the corresponding 3D video frame. The coverage represents the proportion of key targets visible under the current candidate virtual viewpoint, the occlusion rate represents the proportion of the area of ​​occluded targets under the current candidate virtual viewpoint to the total area, the viewpoint transformation cost represents the transformation cost required to move from the viewpoint of the previous 3D video frame to the current candidate virtual viewpoint, and the key area weight represents the importance of the key areas covered by the viewpoint of the current candidate frame.

[0008] Secondly, embodiments of this application provide a 3D video generation system for railway passenger stations based on adaptive perspective. The system includes: a heterogeneous sensor acquisition unit, used to deploy multiple cameras and depth sensors at the railway passenger station, synchronously acquire time-series video streams and point cloud data from each sensor through a network clock synchronization mechanism, and generate a real-time sensor dataset with a unified timestamp; a dynamic point cloud generation unit, used to construct a static geometric model reflecting the layout of railway passenger station facilities based on the building information model of the railway passenger station, and fuse the static geometric model with the real-time sensor dataset to generate a corresponding dynamic scene point cloud; and a virtual perspective scoring unit, used to calculate the coverage corresponding to each candidate virtual perspective in the candidate virtual perspective set. The system calculates the viewpoint score by considering factors such as occlusion rate, viewpoint transformation cost, and key area weight, and calls a comprehensive scoring function. The virtual viewpoint is the viewpoint of the virtual camera. A 3D video rendering unit is used to select the best virtual viewing viewpoint with the highest corresponding viewpoint score and perform 3D rendering on the dynamic scene point cloud based on the best virtual viewing viewpoint to generate the corresponding 3D video frame. The coverage represents the proportion of key targets visible under the current candidate virtual viewpoint, the occlusion rate represents the proportion of the area of ​​occluded targets to the total area under the current candidate virtual viewpoint, the viewpoint transformation cost represents the transformation cost required to move from the viewpoint of the previous 3D video frame to the current candidate virtual viewpoint, and the key area weight represents the importance of the key areas covered by the viewpoint of the current candidate frame.

[0009] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the adaptive perspective-based three-dimensional video generation method for railway passenger stations according to any embodiment of this application.

[0010] Fourthly, embodiments of this application provide a storage medium storing a computer program thereon, characterized in that, when the program is executed by a processor, it implements the steps of the method for generating three-dimensional video of a railway passenger station based on an adaptive perspective according to any embodiment of this application.

[0011] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the adaptive perspective-based three-dimensional video generation method for railway passenger stations according to any embodiment of this application.

[0012] The method and system for generating 3D videos of railway passenger stations based on adaptive perspective provided in this application can achieve at least the following technical effects:

[0013] (1) By deploying multiple cameras and depth sensors at railway passenger stations and synchronously collecting video streams and point cloud data through a network clock synchronization mechanism, the information collected by different sensors can be combined to generate a real-time sensing dataset with a unified timestamp, providing a more comprehensive and three-dimensional scene perception capability than traditional two-dimensional monitoring systems. In addition, the generation of dynamic scene point clouds and the fusion of static geometric models enable the monitoring system to capture the spatial information of key targets in real time.

[0014] (2) By evaluating the candidate virtual viewpoint set, including a comprehensive analysis of dimensions such as viewpoint coverage, occlusion rate, viewpoint change cost, and key area weight, the best virtual observation viewpoint with the highest viewpoint score is selected. Thus, multi-dimensional optimization is used to improve the accuracy of viewpoint selection, reduce occlusion to a greater extent, and more intelligently adapt to the complex changes in the current scene, thereby ensuring the quality and comprehensiveness of video surveillance.

[0015] (3) By performing 3D rendering of dynamic scene point clouds based on the optimal virtual viewing angle, clearer and more realistic 3D surveillance video frames can be generated. Thus, selecting the optimal viewing angle and smoothly transitioning avoids the image stuttering and discontinuity caused by frequent switching of camera views in traditional methods. At the same time, the consideration of the cost of viewing angle transformation makes 3D video rendering more efficient, avoids the waste of excessive computing resources, and improves rendering efficiency.

[0016] This technical solution introduces a comprehensive scoring mechanism based on virtual perspectives, which intelligently selects the optimal virtual 3D video rendering viewpoint, achieving adaptive monitoring coverage and avoiding the occlusion and blind spot problems found in traditional monitoring. Furthermore, the spatial understanding based on Building Information Modeling and the real-time sensor data synchronization mechanism further ensure the system's efficiency and real-time performance, providing more accurate and comprehensive technical support for the safety monitoring and operational scheduling of railway passenger stations. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating an example of a method for generating 3D video of a railway passenger station based on an adaptive perspective, according to an embodiment of this application, is shown.

[0019] Figure 2 A flowchart illustrating an example of generating dynamic scene point clouds according to an embodiment of this application is shown.

[0020] Figure 3 A flowchart illustrating an example of adaptive occlusion compensation depth estimation according to an embodiment of this application is shown.

[0021] Figure 4 A simulation diagram illustrating an example of comparing the detection accuracy of the baseline algorithm and the adaptive viewpoint algorithm over 10 time frames is shown.

[0022] Figure 5 A simulation diagram illustrating an example of comparing the occlusion rate of the baseline algorithm and the adaptive viewpoint algorithm over 10 time frames is shown.

[0023] Figure 6 A structural block diagram of an example of a railway passenger station 3D video generation system based on an adaptive perspective, according to an embodiment of this application, is shown. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] It should be noted that some experts and scholars have proposed some novel directions for video surveillance of railway passenger transport.

[0026] Specifically, in "A Train Station Surveillance System: Challenges and Solutions," Ozer et al. introduced a smart camera system for train stations. The system employs a multi-layered "smart camera-server" architecture, using smart cameras to detect passenger gestures in real time and send alerts to the monitoring center. The paper points out that the basic principle of multi-camera monitoring systems is that dedicated security personnel view multiple screens in the monitoring room; however, personnel attention is limited, making it difficult to guarantee system effectiveness. To improve efficiency, smart cameras need to automatically identify anomalies and push relevant images to operators. However, traditional target detection and tracking algorithms are typically suitable for low-occlusion and high-resolution scenes, but not for station environments with complex backgrounds and unstable lighting. The paper also summarizes the main challenges encountered during the deployment of the smart camera system: 1) Hardware-wise, it requires processing massive amounts of video data for real-time computation; 2) Environmentally, station lighting suffers from flicker, shadows, and reflections, causing frequent changes in camera parameters and exposure; 3) Calibration-wise, the complex structures of platforms, stairs, and ramps make camera geometric calibration difficult; 4) Tracking-wise, multiple people obstructing the view and changes in posture cause target tracking to be easily lost, requiring specific graph theory algorithms to solve the problem. Therefore, two-dimensional multi-camera systems struggle to fully describe complex three-dimensional spaces.

[0027] In traditional 3D reconstruction technology, multiple views Figure 3 3D reconstruction reconstructs scene depth by combining image information from multiple cameras. Multi-baseline stereo algorithms calculate depth by selecting a reference view and matching pixels across views. Princeton University's multi-view reconstruction textbook points out that occlusion problems increase significantly with the number of cameras or the baseline distance, introducing incorrect matches into the error function and affecting reconstruction accuracy. Volumetric methods such as voxel coloring discretize the scene into a voxel mesh and determine whether each voxel belongs to an object by judging whether its color consistency is satisfied in different images. However, this method requires pre-determining the visibility order of voxels: all voxels that occlude a particular voxel must be processed before calculating its color consistency; otherwise, incorrect judgments will occur. Furthermore, voxel coloring relies on good lighting and texture; if the scene lacks texture or the lighting changes drastically, "bumps" or loss of detail can easily occur. These traditional 3D reconstruction methods are mainly used for static or slowly changing scenes, lacking support for real-time dynamic scenes, and are difficult to apply directly to environments like railway stations where there are dense crowds and frequent occlusion.

[0028] The paper "Recent Advances in Video Analytics for Rail Network Surveillance for Security, Trespass and Suicide Prevention - A Survey" points out that with the development of machine learning, many closed-circuit television (CCTV) monitoring systems integrate multiple sensors to automatically detect static objects and abnormal events for managing public spaces. These methods involve object detection, behavior recognition, and anomaly detection, but the review points out that railway monitoring still faces significant challenges. Existing algorithms rely on large amounts of training data, and their performance fluctuates with changes in lighting and environment. Furthermore, the paper emphasizes that future railway monitoring needs to integrate different types of sensors (electro-optical, thermal infrared, acoustic, etc.) and computing architectures to form an integrated monitoring and decision-making system. This also means that a single two-dimensional vision algorithm cannot meet the needs of safety management; cross-domain information fusion and three-dimensional scene understanding are required.

[0029] It should be understood that the above description of the relevant technologies is intended only to help the public better understand the inventive spirit and motivation of this application, and is not intended to limit this application. Furthermore, the technical solutions described in the above-mentioned relevant technologies are not prior art, and may also be undisclosed technical solutions, such as those under research or in the laboratory stage.

[0030] The technical solutions in this application, including the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information, comply with relevant laws and regulations and do not violate public order and good morals.

[0031] Digital twins are a relatively new concept that has emerged in recent years, referring to the construction of a digital copy of a physical system using sensor data and models. For railway passenger stations, a digital twin needs to include various information such as station structure, equipment status, and passenger flow dynamics. However, existing digital twins are mostly based on Building Information Modeling (BIM) to construct static scenes, unable to update the dynamic status of personnel and equipment in real time; simultaneously, the modeling process requires a large amount of manual data collection and point cloud post-processing, lacking the ability to quickly generate dynamic 3D videos. Furthermore, 3D scanning equipment such as LiDAR is costly in railway scenarios, with limited deployment range, making comprehensive coverage difficult. Therefore, integrating digital twins with multi-view dynamic video to generate real-time 3D monitoring images is a current research gap.

[0032] Figure 1 A flowchart illustrating an example of a method for generating 3D video of a railway passenger station based on an adaptive perspective, according to an embodiment of this application, is shown.

[0033] Regarding the executing entity of the method in this application embodiment, it can be any controller or processor with computing or processing capabilities, specifically a railway passenger transport intelligent monitoring platform. By integrating multi-sensor data, constructing a digital scene model, and employing a dynamic perspective optimization algorithm, it generates real-time three-dimensional video of the station space, providing intuitive and effective support for security command, passenger flow analysis, and emergency response.

[0034] In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of both, and the type of terminal or electronic device can be diverse, such as mobile phones, tablets, or desktop computers, etc.

[0035] like Figure 1 As shown, in step S110, multiple cameras and multiple depth sensors are deployed at the railway passenger station. The time-series video streams and point cloud data of each sensor are collected synchronously through a network clock synchronization mechanism, and a real-time sensor dataset with a unified timestamp is generated.

[0036] In some implementations, multiple heterogeneous sensing nodes are deployed in different key areas of the railway station. Each node contains multiple high-resolution cameras and depth sensors. The cameras are installed according to the principle of "full coverage with no blind spots," with viewing angles designed based on the building structure and passenger flow paths. Depth sensors are mainly deployed in areas with high passenger flow or severe obstruction to provide accurate depth information. During data acquisition, each camera generates a continuous sequence of images, and the depth sensor outputs a point cloud sequence. To ensure data fusion quality, the sensors need to be calibrated and synchronized.

[0037] In sensor calibration and spatiotemporal synchronization, the first step is to calibrate the intrinsic and extrinsic parameters of the cameras. In intrinsic parameter calibration, the focal length, principal point, and distortion coefficients of each camera are calibrated to obtain the intrinsic parameter matrix. In extrinsic parameter calibration, significant feature points (such as corners and pillars) are selected based on the station's static geometric model to calibrate the cameras' extrinsic parameters, determine their pose, and thus establish the transformation relationship between the world coordinate system and the camera coordinate system.

[0038] Furthermore, to achieve time synchronization, the system employs a unified network clock, aligning all image and point cloud data through timestamps. During acquisition, if the timestamp deviation of a particular sensor exceeds a threshold, its data can be adjusted using interpolation methods. Thus, the unified timestamps of the time-series dataset enable precise matching of data from different sensors, avoiding visual information misalignment or spatiotemporal inconsistencies caused by data deviations.

[0039] In step S120, a static geometric model reflecting the layout of railway passenger station facilities is constructed based on the building information model of the railway passenger station, and the static geometric model is fused with the real-time sensor dataset to generate a corresponding dynamic scene point cloud.

[0040] It should be noted that the station's Building Information Modeling (BIM) provides the geometric information of each structural component. The model is exported using computer-aided design software and converted into a mesh suitable for 3D rendering. A depth sensor is used to scan the station environment when there are no passengers to obtain point clouds. The scanned point cloud and BIM model are registered using the Iterative Closest Point (ICP) algorithm, and the model coordinate system is optimized. The resulting static geometric model is denoted as... .

[0041] During implementation, the BIM model first provides a precise geometric framework, including the 3D coordinates, dimensions, and relative positions of each facility. Based on this, synchronously acquired real-time sensor data (such as video streams and point cloud data) is fused. Point cloud data, typically acquired by depth sensors, contains spatial location and depth information about each point in the scene. By interfacing with the static geometric model, the real-time acquired data can be combined with the existing spatial structure to generate a dynamic scene point cloud, reflecting the current facility layout and personnel activities within the station. The real-time acquired point cloud... With static geometric model Fusion to construct a dynamic scene point cloud of the complete scene In addition, a spatial octree is used to manage the scene in layers, which significantly improves the efficiency of subsequent rendering and viewpoint evaluation.

[0042] In step S130, for each candidate virtual view in the candidate virtual view set, the coverage, occlusion rate, view transformation cost and key area weight corresponding to the candidate virtual view are calculated, and the comprehensive scoring function is called to calculate the corresponding view score.

[0043] Here, the virtual perspective refers to the perspective of a virtual camera. The virtual perspective defines the viewing angle and field of view, and it selects the optimal viewing angle by simulating the camera's position and orientation in three-dimensional space. When generating images or videos, the virtual camera's perspective determines the content of the rendered image.

[0044] Specifically, coverage represents the proportion of key targets visible in the current candidate virtual viewpoint, occlusion rate represents the proportion of the area of ​​occluded targets in the current candidate virtual viewpoint to the total area, viewpoint transformation cost represents the transformation cost required to move from the viewpoint of the previous 3D video frame to the current candidate virtual viewpoint, and key region weight represents the importance of key regions covered by the viewpoint of the current candidate frame.

[0045] Based on the evaluation of the four factors mentioned above, the system calculates the total score for each candidate virtual viewpoint using a comprehensive scoring function. For example, the scores of each factor can be weighted and fused to obtain the final viewpoint score. Thus, through a multi-dimensional virtual viewpoint evaluation mechanism, the optimal observation angle can be accurately selected, ensuring coverage of key areas and maximizing the effectiveness and stability of the viewpoint. This allows each frame of video to clearly display the main dynamics of the scene and avoids the problems caused by frequent viewpoint switching in traditional monitoring systems.

[0046] In step S140, the best virtual viewing angle with the highest corresponding viewpoint score is selected, and the dynamic scene point cloud is rendered in three dimensions based on the best virtual viewing angle to generate the corresponding three-dimensional video frame.

[0047] Here, through the aforementioned virtual perspective evaluation, the system selects the best virtual perspective with the highest score. Then, based on this perspective, the dynamic scene point cloud is rendered in 3D to generate the final 3D video frame. The rendering process combines the aforementioned static geometric model and real-time sensor data, taking into account factors such as lighting, texture, and occlusion in the scene to generate a highly realistic 3D video.

[0048] The 3D rendering process can be implemented using various depth map rendering techniques in computer graphics. For example, firstly, based on the selected optimal virtual viewpoint, a corresponding viewpoint matrix is ​​determined, and a viewpoint transformation is performed to map point cloud data onto the image plane. Then, depth information in the scene is processed using ray tracing or rasterization algorithms to generate clear 3D video frames. Because the cost of viewpoint transformation is comprehensively considered during the optimal viewpoint calculation, the rendering process smoothly transitions with the rendering results of the previous frame, avoiding visual stuttering caused by frequent viewpoint switching. This allows for rapid response and efficient generation of video frames in dynamic scenes.

[0049] Figure 2 A flowchart illustrating an example of generating dynamic scene point clouds according to an embodiment of this application is shown.

[0050] like Figure 2 As shown, in step S210, the real-time acquired depth data and the dynamic targets in the video stream are fused according to the real-time sensing dataset to generate a dynamic scene depth map.

[0051] Here, information captured from different types of sensors (such as depth sensors and cameras) is processed synchronously to generate a depth map of a dynamic scene, thereby reflecting the depth information of objects, people, and spatial structures in the scene.

[0052] Specifically, depth information (i.e., the distance from each pixel to the sensor) is provided by a depth sensor. Depth data is typically stored in the form of a point cloud, where each point contains spatial coordinates and a corresponding depth value. Furthermore, dynamic targets in the video stream can usually be identified using computer vision algorithms (such as deep learning-based object detection models). Dynamic targets in the video stream are extracted, labeled, or segmented to form a two-dimensional target region corresponding to the depth data.

[0053] In data fusion, depth data and dynamic target detection results from the video stream need to be fused through spatial alignment and temporal synchronization. Specifically, the data acquired by the depth sensor is mapped to each frame of the video stream, mapping depth information to the image data of the dynamic target. For example, clock synchronization technology ensures that the depth data and video frames are time-consistent, and camera calibration technology performs geometric calibration between cameras, ensuring that the depth image and video image are aligned in the same space.

[0054] In some implementations, the depth value of an image pixel can be calculated using feature matching and disparity calculation.

[0055] Specifically, select a reference camera. In its image Detecting feature points In other cameras Image In the process, corresponding points are searched through depth-guided block matching. Calculate parallax The baseline length obtained through calibration and focal length Calculate feature points based on the disparity formula. depth :

[0056] Equation (1)

[0057] in, To prevent division by zero of constants.

[0058] This generates a dynamic scene depth map containing scene depth information. The depth values ​​of both the background and dynamic targets are accurately captured, and the depth information of the dynamic targets can be effectively updated to reflect the movement and changes of the targets.

[0059] In step S220, occluded pixel regions in the dynamic scene depth map are identified, and the dynamic scene depth map is compensated based on the depth data of the occluded pixel regions to generate a dynamic scene point cloud.

[0060] It should be noted that in real-world scenarios, targets and obstacles may obstruct part of the line of sight or affect the acquisition of depth data. Therefore, the handling of occluded areas is crucial to ensuring the accuracy of 3D reconstruction.

[0061] Specifically, occlusion problems typically occur when a dynamic target or object obstructs other objects in the view, resulting in incomplete depth information acquisition. In this process, the system analyzes pixel data in the depth map of the dynamic scene to identify the occluded areas, using methods such as foreground / background separation and depth outlier detection.

[0062] Once the occluded areas are identified, the missing depth data can be filled in using compensation methods. These compensation strategies can be diverse, such as interpolation based on surrounding pixels, depth data inference, or physical model-based compensation. For example, in depth data inference, the missing areas are filled in by calculating the depth information before and after the target.

[0063] By compensating for occlusion areas, the system eliminates the problem of missing depth data caused by dynamic target occlusion or scene complexity, enabling it to recover and improve missing data in dynamic scene depth maps. Ultimately, the compensated depth map is converted into a 3D point cloud, accurately recording the position, depth, and color information of each point in 3D space. Thus, the generated dynamic scene point cloud not only reflects the motion trajectory of dynamic targets but also accurately reproduces the physical structure and spatial relationships of the scene.

[0064] Regarding the implementation details of step S220, in some examples of embodiments of this application, considering the dynamic and occluded characteristics of the passenger station environment, this paper proposes an Adaptive Occlusion-Compensated Depth Estimation (AOCDE) algorithm, which uses multi-view images and depth sensor data to jointly estimate the depth.

[0065] Figure 3 A flowchart illustrating an example of adaptive occlusion compensation depth estimation according to an embodiment of this application is shown.

[0066] like Figure 3 As shown, in step S310, the image gradient consistency and bidirectional matching are used to detect whether a pixel is occluded. When at least one first pixel does not satisfy bidirectional matching or color consistency in multiple views, each first pixel is marked as an occluded pixel region.

[0067] Specifically, image gradient consistency is achieved by comparing a reference image ( ) and source image ( Occlusion is detected by using the gradient of the corresponding pixel in the image.

[0068] Equation (2)

[0069] In the formula, It is a reference image. medium pixel gradient at, It is the source image medium pixel gradient at, This represents the gradient consistency threshold. When the image gradient is inconsistent, it indicates that the pixel may be occluded.

[0070] Bidirectional matching detects occlusion by comparing the matching of the same pixel in the reference image and the source image.

[0071] Equation (3)

[0072] Equation (4)

[0073] In the formula, It is a reference image. Pixels With source image Corresponding pixel in the middle Difference measures (such as color difference, pixel value difference); It is the source image Pixels Compared with reference image The corresponding pixel in Difference measurement.

[0074] Bidirectional matching is used to detect mismatches between the reference and source images. If the matching error in a certain direction exceeds a certain threshold, such as... If the pixel is occluded, then the pixel is marked as an occluded area.

[0075] In step S320, it is detected whether the occluded pixel area is located within the depth sensor coverage area.

[0076] The coverage area of ​​a depth sensor refers to the spatial range within which the depth sensor can directly sense and provide effective depth data. In practical applications, depth sensors (such as LiDAR or infrared depth cameras) typically have a certain field of view and measurement distance; areas outside this range cannot acquire effective depth information. Therefore, for pixels marked as occluded, it is necessary to confirm whether they are within the effective coverage area of ​​the depth sensor to determine whether depth sensor data can be used for compensation.

[0077] In step S331, if the occluded pixel area is located within the coverage area of ​​the depth sensor, the point cloud data provided by the depth sensor is used to calculate the depth value corresponding to the occluded pixel area using the nearest neighbor interpolation algorithm.

[0078] In some implementations, the image coordinates (e.g., the pixel's position in the image) of the marked occluded pixel region are converted into three-dimensional spatial coordinates based on the region's location. Next, assuming the depth point cloud data is sparsely distributed in space, the nearest neighbor interpolation algorithm is used to find the point closest to the occluded pixel's location in the depth point cloud data, based on the occluded pixel's neighborhood position, thereby obtaining the corresponding depth value. This effectively supplements the depth information of the occluded region.

[0079] In step S333, if the occluded pixel area is not located within the depth sensor coverage area, a Kalman filter-based depth estimation algorithm is used to predict the depth value of the occluded pixel area.

[0080] Kalman filtering is an effective recursive filtering method widely used for state estimation of dynamic systems. It can estimate missing depth information based on existing depth data and the motion trajectory of a dynamic target. Specifically, the motion trajectory of the target is first estimated using dynamic target detection and tracking algorithms in the video stream (such as target detection based on optical flow or deep learning). The motion trajectory provides the target's positional changes over time, helping the Kalman filter to estimate the depth value. Then, the Kalman filter predicts the missing depth value by considering the target's dynamic model (e.g., uniform linear motion or acceleration model) and measurement noise (from sensor and image processing errors).

[0081] Equation (5)

[0082] In the formula, For a moment Depth estimation, For a moment The depth value, It is the depth mean of the spatial neighborhood of the occluded pixel region. This represents the Kalman filter gain.

[0083] Here, by combining the depth estimate from the previous time step with the mean depth of the surrounding spatial region, Kalman gain is utilized. The two values ​​are weighted and fused to obtain a more accurate depth estimate. Kalman gain. Its function is to balance the influence between historical depth values ​​and the spatial depth mean, ensuring that changes in dynamic scenarios are taken into account when making predictions.

[0084] By combining the depth information from the previous moment with the depth data of the local spatial region, the weighting mechanism of Kalman filtering is used to improve the accuracy and robustness of depth estimation. It can accurately estimate the depth value of the occluded area in the depth sensor blind zone, ensuring the integrity of depth data in dynamic scenes.

[0085] In step S340, the predicted depth map of the occluded pixel region is fused with the dynamic scene depth map, and the fused depth map is converted into a 3D point cloud to obtain the corresponding dynamic scene point cloud.

[0086] Here, depth information of occluded pixel regions predicted by depth sensors or Kalman filtering algorithms is fused with a dynamic scene depth map. The estimated depth map is transformed into a 3D point cloud and mapped to the world coordinate system using the camera extrinsic matrix. Point clouds from different cameras and depth sensors are fused, and noise is removed using weighted averaging and voxel filtering. Motion compensation is applied to consecutive time frames to ensure the spatiotemporal continuity of the point cloud.

[0087] Through this fusion process, the depth information of occluded areas is effectively supplemented into the original depth map, resulting in a more complete and accurate scene representation. Next, the fused depth map is converted into a 3D point cloud, where each point represents a spatial location and its corresponding depth value within the scene. Therefore, the resulting dynamic scene point cloud is not only more complete and accurate but also better reflects the spatial structure and dynamic changes of the actual scene, providing a high-quality 3D model.

[0088] Regarding the implementation details of viewpoint selection in step S130, in some examples of embodiments of this application, the Adaptive Viewpoint Optimization (AVO) algorithm can be used to dynamically plan the optimal viewpoint sequence in real time by analyzing scene distribution and event attention.

[0089] It should be noted that in multi-view 3D reconstruction technology, choosing a suitable viewing angle is crucial to determining the final video quality. Traditional methods typically use a fixed reference viewpoint or a preset path, but in dynamic environments, this can easily lead to certain areas being occluded for extended periods or the image becoming discontinuous.

[0090] The specific implementation details of the AVO algorithm mainly include the construction of the candidate view set, the design of view evaluation index, and the calculation of view scoring function.

[0091] Regarding the construction of the candidate viewpoint set, the viewpoint position parameters of the virtual camera in 3D space are set. , This indicates the location information of the virtual camera. This represents the coordinates of the virtual camera in three-dimensional space; and sets the viewing angle using Euler angles. definition, These are yaw angle, pitch angle, and roll angle, respectively.

[0092] Specifically, the spatial position and Euler angles of the virtual camera can be set according to the space of the static set model to ensure that the virtual camera can cover all key areas in the station (such as platforms, waiting rooms, passages, etc.), while taking into account the maximum coverage of the viewing angle and the layout of the passages in the station.

[0093] Furthermore, the monitoring area is extracted from the static geometric model and discretized to generate a preset number of candidate virtual viewpoints covering the monitoring area; among which, the first... Each candidate virtual perspective is represented as .

[0094] To ensure comprehensive monitoring of key areas of the station, the candidate viewpoint set is generated by discretizing the station's architectural spatial layout. Multiple candidate virtual viewpoints are generated through reasonable spatial division, and each viewpoint is evaluated. Finally, the optimal viewpoint is selected for dynamic 3D monitoring video frame rendering.

[0095] Here, the monitoring area mainly refers to the feasible area in the railway passenger station, such as the waiting room and platform. Based on the spatial geometric model of the station, the accessible area of ​​the station is discretized and sampled. Multiple view positions are selected and combined with the Euler angle range to generate candidate viewpoints, ensuring that the core area of ​​the station can be effectively covered, especially the densely populated areas and key equipment areas.

[0096] To quantify the merits of each candidate viewpoint, four types of viewpoint evaluation metrics were defined: coverage, occlusion rate, viewpoint change cost, and key area weight. Then, a viewpoint scoring function was used to quantitatively calculate the viewpoint score for each candidate viewpoint.

[0097] Specifically, the comprehensive scoring function is as follows:

[0098] Equation (6)

[0099] In the formula, Indicates the current candidate virtual perspective Rating from the perspective of Indicates the current candidate virtual perspective coverage Indicates the current candidate virtual perspective occlusion rate, Indicates the current candidate virtual perspective The cost of changing perspectives Indicates the current candidate virtual perspective The weight of key regions; These are weighting coefficients, used to adjust the impact of different scoring items on the final perspective score.

[0100] In equation (6), For adjustable parameters, satisfying At each time frame, the system calculates a scoring function for all candidate viewpoints and selects the viewpoint with the highest score. The virtual camera position for the current frame.

[0101] Equation (7)

[0102] In the formula, Indicates the total number of targets in the scene; For the goal The visible area represents the target's position within the viewing angle. The area observed below; For the goal The total area.

[0103] In equation (7), coverage represents the proportion of the target area that can be observed from the current viewpoint. It reflects the effectiveness of the viewpoint; the higher the coverage, the more comprehensive the viewpoint's coverage of the scene, ensuring monitoring coverage of important areas.

[0104] Equation (8)

[0105] In the formula, Indicate target Indicates perspective The area below that is obscured.

[0106] In Equation (8), the occlusion rate is used to evaluate the occlusion of the target from the current viewpoint. This scoring item is used to judge the effectiveness of the viewpoint; the lower the occlusion rate, the clearer the target can be observed from the viewpoint. By calculating the occlusion rate, information loss caused by occlusion is reduced, providing clearer scene monitoring data.

[0107] Equation (9)

[0108] In the formula, This indicates the perspective from the previous 3D video frame. Indicates perspective The corresponding virtual camera position, Indicates perspective The corresponding virtual camera position, This represents the weighting coefficient used to adjust the pitch angle change during viewpoint transformation. Indicates perspective Euler angles, Indicates perspective Euler angles, Indicates the maximum range of change for the virtual camera. This indicates the maximum range of variation of Euler angles.

[0109] In equation (9), the viewpoint transformation cost is used to evaluate the cost of switching from the previous viewpoint to the current viewpoint, mainly considering changes in position and Euler angles (pitch, yaw, and roll). The smaller the viewpoint transformation cost, the smoother the viewpoint switching, and the higher the system's response speed. By calculating the viewpoint transformation cost, the system can evaluate the smoothness and cost of viewpoint switching, improve the system's real-time response capability, and ensure the smoothness of monitoring.

[0110] Equation (10)

[0111] In the formula, Indicates perspective The weight of key areas is as follows: Indicates perspective Field of view, Represent each key area in the scene The key areas are preset according to the task requirements; As an indicator function, when the key area The value is 1 when it is within the field of view of the current viewpoint, and 0 otherwise; Key areas The regional weight indicates the importance of that region; This represents the sum of the regional weights of all key areas in the scene.

[0112] In equation (10), the key area weight is used to evaluate the importance of key areas from a specific perspective, such as prioritizing areas with high passenger flow or potential safety hazards within the station. Specifically, based on task requirements, key areas within the station, such as entrances / exits, ticket gates, and escalators, are pre-defined, and the weight value for each key area is determined. Configure settings based on the importance of the region.

[0113] In some examples of embodiments of this application, the regional weight of key areas can also be dynamically adjusted based on real-time monitoring results.

[0114] Specifically, based on dynamic scene point cloud identification, passenger flow density and abnormal behavior events in key areas are identified. Abnormal behavior events include any one of the following: falling events, going against the flow of traffic, or crossing the warning zone.

[0115] Here, sub-point cloud analysis is performed on each corresponding key area in the dynamic scene point cloud to calculate the passenger flow density index.

[0116] Equation (11)

[0117] In the formula, Key areas Passenger flow density index Key areas The number of passenger flow point clouds in the data. Key areas Total point cloud volume covering the area.

[0118] Anomaly detection can identify dynamic objects in videos using various object detection algorithms (such as YOLO) and use depth information to determine whether abnormal behavior (such as falling, going against the flow, or crossing a restricted area) has occurred. These events will be marked as abnormal behavior and trigger further responses.

[0119] Furthermore, when the passenger flow density in the first key area exceeds the preset passenger flow density threshold and / or abnormal behavior events are detected, the regional weight of the first key area is increased.

[0120] Here, by dynamically increasing or decreasing the monitoring weight of various areas of the station based on real-time passenger flow density and abnormal behavior event detection results, the system can optimize monitoring strategies according to the real-time situation within the station. For example, when the passenger flow density of a certain area exceeds a threshold, or when an abnormal behavior event occurs, the monitoring weight of that area is automatically increased, ensuring that key areas receive priority attention throughout the monitoring system. This improves the response speed to abnormal events and ensures that dynamic high-risk areas are under effective monitoring.

[0121] Regarding the details of the rendering of the 3D video frame in step S140, in some examples of the embodiments of this application, the dynamic scene point cloud is cropped based on the best virtual viewing angle, the geometry outside the virtual camera's field of view is removed, and a depth test is performed using a Z-buffer (depth buffer) to ensure the visibility of pixels.

[0122] When generating 3D video frames, the dynamic scene point cloud first needs to be cropped using the optimal virtual viewing angle. The purpose of cropping is to remove geometry outside the virtual camera's field of view, reducing the rendering burden and ensuring that only the portion visible to the camera is rendered, thus improving rendering efficiency. To ensure that the depth information of each pixel accurately reflects the position of objects in the scene, depth testing is performed using Z-buffer technology to determine the visibility of each pixel, and the depth values ​​are compared to determine whether the pixel needs to be rendered.

[0123] Video frames captured by multiple cameras are projected onto the surface of a dynamic scene point cloud to form texture maps. The ambient lighting distribution is estimated using infrared images from a depth sensor. Combined with local lighting calculations based on the Phong (Lambert) model of a railway passenger station, the brightness of the 3D video frames is fitted.

[0124] Here, texture mapping is used to support more realistic 3D video rendering. Specifically, by projecting video frames captured by multiple cameras onto the surface of a dynamic scene point cloud, corresponding texture maps are generated, which can map detailed information of the station scene (such as texture and color) onto the 3D model. Next, the ambient lighting distribution is estimated using infrared images from a depth sensor, and lighting calculations are performed on each pixel using the Phong lighting model to ensure that the generated 3D video frames have realistic brightness and lighting effects.

[0125] Specifically, the ambient light intensity is estimated using infrared images from a depth sensor. This provides background information for lighting calculations. Local lighting calculations are performed on each pixel using the Phong model, which calculates the brightness of each pixel by considering the combination of ambient light, diffuse light, and specular light.

[0126] Equation (12)

[0127] In the formula, The direction of the light source, For the surface normal vector, For reflection vector, The optimal virtual observation view vector. Indicates ambient light intensity. Indicates the ambient light reflectance. Indicates the diffuse reflectance coefficient. Indicates the specular reflection coefficient. Indicates the light intensity of the light source. The specular index is used to control the degree of scattering in specular reflection.

[0128] Based on texture mapping and lighting calculations, the generated 3D video frames not only possess high-quality details but also realistically reproduce lighting and environmental changes, greatly enhancing the realism of the visual effects. By combining depth sensor data and the Phong lighting model, it can accurately reflect brightness and lighting changes in the scene, improving the viewing experience of the video frames.

[0129] In this embodiment, during the rendering process, a Z-buffer is used for depth testing to ensure accurate scene display. Infrared images from a depth sensor are used to estimate the ambient lighting distribution, and the Phong lighting model is used to calculate local lighting effects, ensuring the rendered result has realistic lighting and reflection effects. Furthermore, the continuously rendered video sequence is compressed by a video encoder and transmitted to the monitoring center. The encoder can dynamically adjust the bitrate and resolution based on network bandwidth. To ensure system real-time performance, GPU-based parallel rendering and encoding can be employed.

[0130] To verify the effectiveness of the proposed adaptive perspective 3D video generation method, a comparative experiment was designed. Since real-world station environments are complex and difficult to replicate, this paper employs a combination of virtual scene simulation and real-world data to construct a 3D simulated station including platforms, waiting rooms, and entrances / exits. Pedestrian models are randomly generated within the scene to simulate passenger behaviors such as entering / exiting the station, waiting for trains, and boarding. Obstacles (such as pillars and billboards) and lighting variations are added to certain areas.

[0131] 1. Experimental Setup

[0132] Baseline Algorithm: Traditional multi-view 3D reconstruction and fixed-view rendering. A fixed reference viewpoint and preset camera path are used, and the position and pose of the virtual camera are not adjusted throughout the experiment. Depth estimation uses the standard multi-baseline stereo method without occlusion compensation.

[0133] Adaptive Viewpoint Algorithm: The proposed adaptive viewpoint optimization algorithm is employed. The candidate viewpoint set covers multiple high- and low-angle viewpoints above the platform and waiting room. Depth estimation uses the AOCDE algorithm, supplemented by data from a depth sensor. The parameters of the viewpoint scoring function are set as follows: The weight of key areas is dynamically adjusted based on passenger flow density.

[0134] Evaluation Indicators: To comprehensively evaluate the results, three indicators are set:

[0135] Detection accuracy: The accuracy of pedestrian detection based on 3D video, used to evaluate the algorithm's ability to identify targets.

[0136] Occlusion rate: In the rendering results, the proportion of the occluded target area to the total target area is used to measure the effect of viewpoint selection on occlusion.

[0137] Rendering latency: The delay from data acquisition to 3D screen display, used to verify the system's real-time performance.

[0138] 2. Experimental Results

[0139] Figure 4 The diagram shows a simulation of the detection accuracy of a baseline algorithm versus an adaptive viewpoint algorithm over 10 time frames.

[0140] like Figure 4 As shown, the detection accuracy of the adaptive view algorithm is consistently higher than that of the baseline algorithm and maintains a steady increase over time, indicating that the adaptive view can avoid occlusion in dynamic scenes and improve the completeness of pedestrian detection.

[0141] Figure 5 A simulation diagram illustrating an example of comparing the occlusion rate of the baseline algorithm and the adaptive viewpoint algorithm over 10 time frames is shown.

[0142] like Figure 5 As shown, the occlusion rate of the baseline algorithm decreases slowly over time, but remains between 0.23 and 0.30; the occlusion rate of the adaptive viewpoint algorithm is significantly lower and stabilizes after the 7th frame. This indicates that dynamically selecting the viewpoint can significantly reduce occlusion and improve the integrity of the observation.

[0143] Furthermore, experiments showed that the average rendering latency of the adaptive viewpoint algorithm was 150 ms, while the baseline algorithm's latency was 120 ms. Although the latency was slightly higher, it still met the real-time monitoring requirements (<200 ms). Overall, the adaptive viewpoint algorithm significantly outperformed the baseline algorithm in terms of detection accuracy and occlusion rate, and its real-time performance was acceptable.

[0144] 3. Discussion

[0145] Experimental results show that the proposed adaptive viewpoint-based 3D video generation method can effectively improve monitoring quality by adjusting the position and pose of the virtual camera used for 3D video rendering in real time. Its main advantages include:

[0146] Strong occlusion handling capability: By comprehensively considering coverage and occlusion rate through a viewpoint scoring function, the observation angle is automatically adjusted in occluded areas, significantly reducing the occlusion rate. Compared with the traditional fixed viewpoint, the occlusion rate decreased by approximately 40% in the experiment.

[0147] High target detection accuracy: Without changing the hardware configuration, the algorithm significantly improves the target detection accuracy by optimizing the observation angle and depth estimation, especially in detecting occluded pedestrians in crowded areas.

[0148] Flexibility and scalability: The weights in the perspective scoring function can be adjusted according to business needs. When unexpected events occur, the system can quickly focus on key areas, enabling timely attention to abnormal scenarios. The candidate perspective set can also be expanded or reduced in real time based on station structure and network bandwidth, adapting to station deployments of different sizes.

[0149] Real-time performance is acceptable: Although the adaptive viewpoint algorithm requires additional viewpoint calculations and depth compensation, resulting in a slight increase in rendering latency, it remains within an acceptable range under real-time monitoring. Latency can be further reduced in the future through hardware acceleration or algorithm optimization.

[0150] This paper addresses the problems of single viewpoint, severe occlusion, and lack of 3D information in traditional video surveillance at railway passenger stations, and proposes a 3D video generation method based on adaptive viewpoint. Through multi-sensor acquisition, dynamic depth estimation, viewpoint optimization, and real-time rendering, the system can generate continuous and complete 3D monitoring images in complex station environments. Compared with traditional multi-camera fixed-viewpoint methods, this method significantly improves target detection accuracy and occlusion handling capabilities. Experimental results show that the detection accuracy is improved by an average of approximately 15%, the occlusion rate is reduced by nearly 40%, and real-time performance is maintained.

[0151] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0152] Figure 6 A structural block diagram of an example of a railway passenger station 3D video generation system based on an adaptive perspective, according to an embodiment of this application, is shown.

[0153] like Figure 6 As shown, the railway passenger station 3D video generation system 600 based on adaptive perspective includes a heterogeneous sensing acquisition unit 610, a dynamic point cloud generation unit 620, a virtual perspective scoring unit 630, and a 3D video rendering unit 640.

[0154] The heterogeneous sensing acquisition unit 610 is used to deploy multiple cameras and multiple depth sensors in railway passenger stations. It synchronously acquires the time-series video streams and point cloud data of each sensor through a network clock synchronization mechanism and generates a real-time sensing dataset with a unified timestamp.

[0155] The dynamic point cloud generation unit 620 is used to construct a static geometric model reflecting the layout of railway passenger station facilities based on the building information model of the railway passenger station, and to fuse the static geometric model with the real-time sensing dataset to generate a corresponding dynamic scene point cloud.

[0156] The virtual viewpoint scoring unit 630 is used to calculate the coverage, occlusion rate, viewpoint transformation cost and key area weight of each candidate virtual viewpoint in the candidate virtual viewpoint set, and call the comprehensive scoring function to calculate the corresponding viewpoint score; the virtual viewpoint is the viewpoint of the virtual camera.

[0157] The 3D video rendering unit 640 is used to select the best virtual viewing angle with the highest corresponding viewpoint score, and to perform 3D rendering of the dynamic scene point cloud based on the best virtual viewing angle to generate corresponding 3D video frames.

[0158] Wherein, the coverage represents the proportion of key targets visible in the current candidate virtual viewpoint, the occlusion rate represents the proportion of the area of ​​occluded targets in the current candidate virtual viewpoint to the total area, the viewpoint transformation cost represents the transformation cost required to move from the viewpoint of the previous 3D video frame to the current candidate virtual viewpoint, and the key area weight represents the importance of the key areas covered by the viewpoint of the current candidate frame.

[0159] In some embodiments, this application provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions. The execution instructions can be read and executed by electronic devices (including but not limited to computers, servers, or network devices) to perform the steps of any of the above-described methods for generating 3D videos of railway passenger stations based on adaptive perspective.

[0160] In some embodiments, this application also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of any of the above-described methods for generating 3D videos of railway passenger stations based on adaptive perspective.

[0161] In some embodiments, this application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of a method for generating a 3D video of a railway passenger station based on an adaptive perspective.

[0162] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0163] The electronic devices in this application can exist in various forms, including but not limited to: mobile communication devices, ultra-mobile personal computer devices, portable entertainment devices, or other airborne electronic devices with data interaction functions.

[0164] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0165] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A railway station three-dimensional video generation method based on adaptive view angle, characterized in that, The method comprises: arranging multiple cameras and multiple depth sensors at a railway station, synchronously collecting time-series video streams and point cloud data of each sensor through a network clock synchronization mechanism, and generating a real-time sensor data set with a unified timestamp; constructing a static geometric model reflecting the layout of the railway station facilities based on a building information model of the railway station, and fusing the static geometric model with the real-time sensor data set to generate a corresponding dynamic scene point cloud; for each candidate virtual perspective in a candidate virtual perspective set, calculating the coverage, occlusion rate, perspective transformation cost and key area weight corresponding to the candidate virtual perspective, and calling a comprehensive scoring function to calculate a corresponding perspective score; the virtual perspective is the perspective of a virtual camera; selecting the best virtual observation perspective corresponding to the highest perspective score, and performing three-dimensional rendering on the dynamic scene point cloud based on the best virtual observation perspective to generate a corresponding three-dimensional video frame; wherein the coverage represents the proportion of key targets visible under the current candidate virtual perspective, the occlusion rate represents the proportion of the area of the target occluded under the current candidate virtual perspective to the total area, the perspective transformation cost represents the conversion cost required to move from the perspective of the previous three-dimensional video frame to the current candidate virtual perspective, and the key area weight represents the importance of the key area covered by the current candidate frame perspective.

2. The method of claim 1, wherein, The fusion of the static geometric model and the real-time sensor data set to generate a corresponding dynamic scene point cloud comprises: According to the real-time sensor data set, the dynamic targets in the real-time collected depth data and video stream are fused to generate a dynamic scene depth map; identify the occluded pixel area in the dynamic scene depth map, and compensate the dynamic scene depth map based on the depth data of the occluded pixel area, thereby generating a dynamic scene point cloud.

3. The method of claim 2, wherein, The identification of the occluded pixel area in the dynamic scene depth map and the compensation of the dynamic scene depth map based on the depth data of the occluded pixel area to generate a dynamic scene point cloud comprises: use image gradient consistency and bidirectional matching to detect whether a pixel is occluded, and mark each first pixel as an occluded pixel area when at least one first pixel does not satisfy bidirectional matching or color consistency in multiple perspectives; if the occluded pixel area is located in the depth sensor coverage area, use the point cloud data provided by the depth sensor to calculate the depth value corresponding to the occluded pixel area through the nearest neighbor interpolation algorithm; if the occluded pixel area is not located in the depth sensor coverage area, use a Kalman filter-based depth estimation algorithm to predict the depth value of the occluded pixel area: , wherein is the depth estimate at time is the depth value at time is the average depth value based on the spatial neighborhood of the occluded pixel region, is the Kalman filter gain;​​ fuse the predicted depth map of the occluded pixel area with the dynamic scene depth map, and convert the fused depth map into a three-dimensional point cloud to obtain a corresponding dynamic scene point cloud.

4. The method of claim 1, wherein, The determination operation for the candidate virtual perspective set comprises: Setting a viewing angle position parameter of the virtual camera in a three-dimensional space , Position information of the virtual camera, Coordinate of the virtual camera in a three-dimensional space; and setting a viewing angle orientation defined by Euler angles , Yaw angle, pitch angle and roll angle, respectively; The monitoring area is extracted from the static geometric model, and discretization processing is performed on the monitoring area to generate a preset number of candidate virtual view angles covering the monitoring area; wherein a first candidate virtual view angle is represented as . .

5. The method of claim 1, wherein, The comprehensive scoring function is: , In the formula, Indicates the current candidate virtual perspective Rating from the perspective of Indicates the current candidate virtual perspective coverage Indicates the current candidate virtual perspective occlusion rate, Indicates the current candidate virtual perspective The cost of changing perspectives Indicates the current candidate virtual perspective The weight of key regions; These are weighting coefficients, used to adjust the impact of different scoring items on the final perspective score; , In the formula, Indicates the total number of targets in the scene; For the goal The visible area represents the target's position within the viewing angle. The area observed below; For the goal The total area; , In the formula, represents the target represents the occluded area at the viewing angle under the viewing angle; , wherein represents the view angle from the previous three-dimensional video frame, represents the view angle corresponding virtual camera position, represents the view angle corresponding virtual camera position, represents a weight coefficient for adjusting the change of the pitch angle in the view angle transformation process, represents the view angle Euler angle, represents the view angle Euler angle, represents the maximum change range of the virtual camera, represents the maximum change range of the Euler angle; , In the formula, represents the weight of the key area under the view angle , represents the field of view of the view angle , represents the weight of each key area in the scene , the key area is preset according to the task requirement; is an indication function, the value is 1 when the key area is in the field of view of the current view angle, otherwise 0; is the area weight of the key area , indicating the importance of the area; represents the sum of the area weights of all key areas in the scene.

6. The method of claim 5, wherein, It also includes: identify passenger flow density and abnormal behavior events of each of the key areas based on the dynamic scene point cloud; the abnormal behavior events include any one of the following: a fall event, a reverse event, or a barrier crossing event; when it is monitored that the passenger flow density of a first key area exceeds a preset passenger flow density threshold and / or there is an abnormal behavior event, increase the area weight of the first key area.

7. The method of claim 1, wherein, the three-dimensional rendering of the dynamic scene point cloud based on the optimal virtual observation view angle to generate a corresponding three-dimensional video frame, comprising: cutting the dynamic scene point cloud based on the optimal virtual observation view angle, removing geometric bodies outside the virtual camera field of view, and using Z-buffer for depth test to ensure the visibility of pixels; projecting the video frames captured by the multiple cameras onto the surface of the dynamic scene point cloud to form a texture map, and estimating the ambient light distribution using the infrared image of the depth sensor, combined with local light calculation based on the railway station Phong model, to realize the brightness fitting of the three-dimensional video frame: , wherein, is the light source direction, is the surface normal, is the reflection vector, is the optimal virtual viewing angle vector, represents the ambient light intensity, represents the ambient light reflection coefficient, represents the diffuse reflection coefficient, represents the specular reflection coefficient, represents the light intensity of the light source, is the highlight exponent, which controls the degree of scattering of the specular reflection.

8. An adaptive view-based railway station 3D video generation system, characterized in that, the system comprises: a heterogeneous sensor acquisition unit for arranging multiple cameras and multiple depth sensors at a railway station, synchronously acquiring time sequence video streams and point cloud data of each sensor through a network clock synchronization mechanism, and generating a real-time sensor data set with a unified timestamp; a dynamic point cloud generation unit for constructing a static geometric model reflecting the layout of railway station facilities based on a building information model of the railway station, and fusing the static geometric model with the real-time sensor data set to generate a corresponding dynamic scene point cloud; a virtual view scoring unit for, for each candidate virtual view in a candidate virtual view set, counting the coverage, occlusion rate, view transformation cost and key area weight corresponding to the candidate virtual view, and calling a comprehensive scoring function to calculate a corresponding view score; the virtual view is the view of a virtual camera; a three-dimensional video rendering unit for selecting an optimal virtual observation view angle with the highest corresponding view score, and three-dimensionally rendering the dynamic scene point cloud based on the optimal virtual observation view angle to generate a corresponding three-dimensional video frame; wherein the coverage represents the proportion of key targets visible under the current candidate virtual view, the occlusion rate represents the proportion of the area of the target occluded under the current candidate virtual view to the total area, the view transformation cost represents the conversion cost required to move from the view of the previous three-dimensional video frame to the current candidate virtual view, and the key area weight represents the importance of the key area covered by the current candidate frame view.

Citation Information

Patent Citations

  • Method and system for extracting visual image information of fully mechanized coal mining face of underground coal mine

    CN118609062A

  • Excavator environment virtual view angle display method and device

    CN119251415A