A method for simulating and constructing subway scene image samples based on virtual reality technology

The subway scene image is constructed through virtual reality technology, and the problem of insufficient sample data in the subway station scene is solved, high-quality RGBD data samples are generated, and the accuracy and robustness of the pedestrian object detection algorithm is improved.

CN118097564BActive Publication Date: 2025-07-18NANJING SAC RAIL TRAFFIC ENG CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410473111.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-07-18
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

The lack of sufficient real sample data sets in the subway station scenario limits the performance and generalization capabilities of the pedestrian object detection algorithm.

Method used

Virtual reality technology is used to construct subway scene images, simulate the appearance, posture and movement of pedestrians through scene modeling, generate high-quality RGBD data samples, and use binocular parallax algorithm to restore the depth map and automatically mark the object position.

Benefits of technology

It provides rich data samples, improves the accuracy and robustness of the object detection algorithm in subway station scenarios, meets the needs of algorithm training, and improves data generation efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118097564B_ABST
    Figure CN118097564B_ABST
Patent Text Reader

Abstract

The present invention designs a method for simulating and constructing subway scene image samples based on virtual reality technology. This method uses virtual reality technology to create synthetic subway station scene images, including simulating various changes in the appearance, posture, actions, etc. of pedestrians. In this way, a large number of diverse and complex image samples can be generated for training and testing pedestrian object detection algorithms. The advantage of this method is that it overcomes the problem of limited real datasets, provides richer data samples, and helps to improve the accuracy and robustness of object detection algorithms in subway station scenes. Through virtual reality technology, researchers can precisely control all aspects of the scene, thus better meeting the needs of algorithm training. This has a positive promoting effect on strengthening human body detection technology in fields such as subway station security monitoring systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In view of the situation that the development of the pedestrian object detection algorithm is restricted by the lack of subway station sample data sets, the present invention proposes a method for simulating and constructing image samples for pedestrian object detection in subway stations. Background Art

[0002] The pedestrian object detection algorithm in the subway station scenario has important significance, which is mainly reflected in the following aspects: The pedestrian object detection algorithm can be used to monitor the safety status of the subway station. By detecting and tracking pedestrians in real time, the system can timely discover abnormal behaviors, emergencies or unusual crowd gatherings, improving the safety of the subway station; There are usually a large number of people in the subway station, especially during peak hours. The pedestrian object detection algorithm can help manage the flow of people, analyze the degree of congestion, and give early warnings of possible congestion situations to optimize the station layout and the guidance of the flow of people; In a busy subway station, people often lose their items. The pedestrian object detection algorithm can assist the security monitoring system to quickly locate and track the lost items, improving the efficiency of finding lost items; The pedestrian object detection algorithm is an important part of the intelligent transportation system. By real-time monitoring of pedestrians in the subway station scenario, the traffic signals can be adjusted more intelligently, the traffic efficiency can be improved, and the congestion can be reduced. Generally speaking, the pedestrian object detection algorithm in the subway station scenario is of great significance for improving the safety, efficiency and service level of the station, and helps to build a more intelligent, safe and convenient urban transportation system.

[0003] When it comes to the pedestrian object detection algorithm in the subway station scenario, a major challenge is the lack of sufficient real sample data sets. Since the subway station belongs to a closed environment, collecting real data may be restricted by privacy and security aspects. Therefore, the lack of sample data may limit the performance and generalization ability of the object detection algorithm.

[0004] In recent years, some progress has been made in the research on the method of simulating and constructing subway scene image samples in the field of computer vision. The following is the development status of some common methods for simulating and constructing subway scene image samples: (1) Data augmentation technology: A common method is to use data augmentation technology to generate more samples by performing operations such as rotation, flipping, scaling, and brightness adjustment on existing real data. This helps to increase the diversity of the dataset, but for some complex scenes and objects, it may still not provide enough samples; (2) Image synthesis technology: Some studies adopt image synthesis technology to synthesize elements in the real scene (such as pedestrians, backgrounds, etc.) into the virtual environment. This can create synthetic images with diversity for training models. However, the authenticity of the synthetic images and the degree of matching with the actual scene are challenges that need to be considered; (3) Deep generative models: Some studies use deep generative models (such as generative adversarial networks, GANs) to generate realistic subway scene images. This method can learn the feature distribution in the scene, thus generating image samples with a high sense of realism; (4) Cross-domain transfer learning: Some studies adopt cross-domain transfer learning in the simulation of subway scene image samples, migrating the model trained in other scenes to the subway scene to improve the utilization efficiency of sample data.

[0005] The above traditional methods for obtaining image training samples also have the following problems: (1) Data augmentation technology generates more samples by transforming real data, but its limitation is that it cannot capture new features in the real scene. For complex scenes and objects, simple transformations may not provide enough sample diversity, resulting in a decline in the performance of the model when dealing with new situations; (2) Image synthesis technology can create synthetic images with diversity, but authenticity and the degree of matching with the actual scene are key issues. Synthetic images may not be able to fully simulate the lighting, shadows, and details in the real world, thus affecting the generalization ability of the model in the real scene; (3) Deep generative models such as generative adversarial networks (GANs) can generate realistic images, but their training may be relatively complex and require a large amount of computing resources. In addition, mode collapse or blurring may occur during the training process of the model, resulting in distorted or unrealistic images being generated; (4) Cross-domain transfer learning can improve the utilization efficiency of sample data, but it may also be limited by the differences between different scenes. The model trained in other scenes does not always effectively adapt to the subway scene, especially when the scene differences are large, the effect of transfer learning may be limited. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the object of the present invention is to design a method for simulating and constructing subway scene image samples based on virtual reality technology, which uses virtual reality technology to create synthetic subway station scene images, including simulating various changes in the appearance, posture, actions, etc. of pedestrians. In this way, a large number of diverse and complex image samples can be generated for training and testing pedestrian object detection algorithms.

[0007] To achieve the above object, the technical solution adopted by the present invention is: a method for simulating and constructing subway scene image samples based on virtual reality technology, which is based on a scene modeling module, a simulation module, and an image training sample generation module. The scene modeling module is used to simulate the real subway station site; the simulation module is used to obtain a top view of pedestrians passing through the turnstile in the form of an image; the image training sample generation module is used to automatically label the positions of pedestrians and objects in the obtained image training samples. Specifically, it includes the following steps:

[0008] Step 1: Establish a virtual subway station model and model pedestrians and their personal belongings.

[0009] Step 2: Establish an automatic simulation model for pedestrians passing through the turnstile. First, construct a blueprint for the process of pedestrians passing through the turnstile. When it comes to pedestrian animations, add pedestrian models to the virtual scene and create corresponding animations for them.

[0010] Step 3: Generate image training samples. First, automatically capture the top view of pedestrians passing through the turnstile in the virtual environment through a virtual camera to generate monocular and binocular RGB images, then use the binocular disparity algorithm to restore the depth map, and automatically annotate the RGBD data through a script.

[0011] Furthermore, the subway station scene modeling adopts the method of segmentation and recombination to model the basic structure and interactive elements in the subway station. The modeling process not only needs to pay attention to the appearance shape, but also needs to add textures and materials to the models. The basic structure includes platforms, turnstiles, indoor ceilings, floors, walls, seats, billboards, lamps, and signs; the interactive elements include passengers, subway vehicles, and ticket vending machines.

[0012] Furthermore, when modeling pedestrians, various actions shown by pedestrians when passing through the turnstile need to be defined, including but not limited to walking and staying. By using 3D modeling software to apply skeletal animations or keyframe animations to pedestrian models, the real motion state can be restored.

[0013] Furthermore, in Step 2, virtual reality technology is used to implement the simulation of the behavior of pedestrians passing through the turnstile. According to the contact between the human bounding box of pedestrians when passing through the turnstile and the bounding boxes in front of and behind the turnstile as the trigger condition, the opening and closing of the turnstile are automatically controlled.

[0014] Further, step 3 is specifically as follows: First, set a virtual camera on top of the turnstile in the virtual engine, automatically acquire image samples of pedestrians passing through the turnstile, generate data with monocular and binocular RGB images, and use the binocular disparity algorithm to restore the depth map corresponding to the RGB; finally, automatically label the positions of people and objects in the restored depth map samples to achieve the automatic acquisition of pedestrian turnstile image samples and the automatic labeling of the positions of people and objects.

[0015] Compared with the prior art, the beneficial effects of this technical solution are as follows:

[0016] 1. The present invention overcomes the problem of limited real datasets, provides richer data samples, and helps improve the accuracy and robustness of object detection algorithms in subway station scenarios.

[0017] 2. Through virtual reality technology, researchers of the present invention can precisely control various aspects of the scenario, thus better meeting the needs of algorithm training, which has a positive promoting effect on strengthening human detection technology in fields such as subway station security monitoring systems.

[0018] 3. The present invention can obtain high-quality RGBD data with object position detection information, and the entire process is automatic, capable of generating the required training samples while saving manpower and material resources to meet the needs of algorithm training and semantic annotation information.

[0019] 4. This automated method not only improves efficiency but also ensures the quality and consistency of the generated data, providing a reliable basis for the development of simulation software and pedestrian behavior analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic flowchart of a method for simulating and constructing subway scene image samples based on virtual reality technology in this embodiment.

[0021] Figure 2 It is a schematic flowchart of the subway station scene modeling process in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following further describes the specific embodiments of the method for generating and constructing image samples of the present invention with reference to the drawings.

[0023] Combined with Figure 1 , the method for simulating and constructing subway scene image samples based on virtual reality technology in this embodiment includes: simulating the real subway station site, simulating the behavior of pedestrians passing through the turnstile, and automatically labeling the positions of pedestrians and objects in the acquired image training samples.

[0024] The structure of the subway station is relatively complex. In this embodiment, a segmentation and recombination modeling method is adopted to more accurately restore the real scene. The subway station scene modeling is a complex process, usually including multiple stages.

[0025] First of all, starting from the conceptual design, it is necessary to determine the overall layout, style and functions of the subway station, including the positions and designs of multiple elements such as platforms, turnstiles, indoor ceilings, floor passages, stairs, elevators, billboards, etc. Next, use professional 3D modeling software such as Blender, 3ds Max or Maya to start creating the basic subway station model. This includes basic structures such as platforms, columns, ceilings, floors, etc. Subsequently, gradually add details such as seats, billboards, lamps, signs, subway vehicles, etc. to ensure that the model is close to the details of the real subway station. To enhance the sense of reality, it is necessary to add textures and materials to the model to ensure that parts such as walls, floors, and ceilings look realistic. Then, design appropriate lighting effects to simulate the actual lighting conditions of the subway station and perform rendering to generate the final image or animation.

[0026] During the modeling process, it is necessary to continuously optimize and adjust to improve performance and ensure smooth operation of the scene. After the modeling is completed, the model can be imported into the game engine Unreal Engine of the target platform and necessary adjustments and optimizations are carried out. Finally, tests are conducted to verify the performance of the scene on the target platform and further adjustments and optimizations are made as needed. For scenes that require interactivity, interactive elements such as passengers, trains, and ticket vending machines can be considered to make the scene more vivid and create a realistic and satisfactory virtual subway station environment.

[0027] The core task of the pedestrian passing through the turnstile sample simulation generation module is to obtain the top view of the pedestrian passing through the turnstile in the form of an image and label the position information of people and objects. To achieve this goal, the module is divided into two parts: the simulation module of the pedestrian automatically passing through the turnstile and the automatic generation module of the image training sample.

[0028] When it comes to pedestrian animation, first, pedestrian models need to be added to the virtual scene and appropriate animations need to be created for them. This includes defining actions such as pedestrians passing through turnstiles, walking, standing, etc. Using 3D modeling software, skeletal animations or keyframe animations can be applied to the pedestrian models to simulate real movements. These animations need to consider the natural behaviors of pedestrians, such as normal gait, actions of passing through turnstiles, etc. Through ingenious animation design, pedestrians in the virtual scene can look more real and vivid, thus improving the fidelity of the generated images. In the pedestrian turnstile simulation module, the behaviors of pedestrians passing through turnstiles in the subway environment are simulated, including two situations: entering the turnstile opening and leaving the turnstile opening. For the behavior of entering the turnstile opening, bounding boxes are set around the pedestrians and the turnstile gates, and the contact between the bounding boxes is used to trigger the turnstile gates to open in the opposite direction. For the behavior of leaving the turnstile opening, the separation of the bounding boxes is also used as the trigger condition to control the closing of the turnstile gates.

[0029] The image training sample generation module is divided into two main steps. First, the top views of pedestrians passing through turnstiles are automatically captured in the virtual environment by a virtual camera to generate monocular and binocular RGB images. In this process, the monocular RGB images are realistically simulated, while the binocular RGB images are used to restore the depth map through the binocular disparity algorithm to form RGBD data. Second, the system automatically annotates the generated RGBD data. By an algorithm that only renders the target object in the virtual engine and extracts the position information through the depth map, the annotation of the target object is achieved. This entire process utilizes the virtual environment and the virtual camera to generate image data with depth information by simulating the real scene, providing a large-scale and detailed annotated data set for training computer vision models, which is particularly suitable for the training of depth perception tasks.

[0030] Through a method for constructing subway scene image sample simulation based on virtual reality technology provided by this embodiment, high-quality RGBD data with object position detection information can be obtained. The entire process is automated, and the required training samples can be generated while saving manpower and material resources to meet the needs of algorithm training and semantic annotation information. This automated method not only improves efficiency but also ensures the quality and consistency of the generated data, providing a reliable basis for the development of simulation software and pedestrian behavior analysis.

[0031] The above is only the preferred embodiment of the present invention and does not constitute a limitation on the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for simulating and constructing subway scene image samples based on virtual reality technology, characterized in that, This method is implemented by a subway scene image sample simulation construction system based on virtual reality technology. The system includes a scene modeling module, a simulation module, and an image training sample generation module. The scene modeling module is used to simulate the real subway station site; the simulation module is used to obtain a top view of pedestrians passing through the turnstile in the form of an image; the image training sample generation module is used to automatically label the positions of pedestrians and objects in the obtained image training samples. Specifically, it includes the following steps: Step 1: Establish a virtual subway station model and model pedestrians and their personal belongings. Step 2: Establish an automatic simulation model for pedestrians passing through the turnstile. First, construct a blueprint for the process of pedestrians passing through the turnstile. When it comes to pedestrian animations, add pedestrian models in the virtual scene and create corresponding animations for them. In the automatic simulation module for pedestrians passing through the turnstile, simulate the behavior of pedestrians passing through the turnstile in the subway station environment, including two situations: entering the turnstile and leaving the turnstile. For the behavior of entering the turnstile, set bounding boxes around the pedestrian and the turnstile door, and trigger the turnstile door to open in the opposite direction through the contact between the bounding boxes. For the behavior of leaving the turnstile, also use the separation of the bounding boxes as the trigger condition to control the closing of the turnstile door. Step 3: Generate image training samples: First, automatically capture the top view of pedestrians passing through the turnstile in the virtual environment through a virtual camera to generate monocular and binocular RGB images. In this process, the monocular RGB images are realistically simulated, and the binocular RGB images are used to restore the depth map through the binocular disparity algorithm to form RGBD data. Secondly, the system automatically labels the generated RGBD data by only rendering the target object in the virtual engine and using an algorithm to extract position information through the depth map to achieve the annotation of the target object.

2. The method for simulating and constructing subway scene image samples based on virtual reality technology according to claim 1, wherein: The virtual subway station scene modeling uses the method of segmentation and recombination to model the basic structure and interactive elements in the subway station. The modeling process not only needs to focus on the appearance shape but also needs to add textures and materials to the model.

3. The method for simulating and constructing subway scene image samples based on virtual reality technology according to claim 1, wherein: When modeling pedestrians, various actions shown by pedestrians when passing through the turnstile need to be defined, including but not limited to walking and staying. By using 3D modeling software to apply skeletal animation or keyframe animation to the pedestrian model, the real motion state is restored.

4. The method for simulating and constructing subway scene image samples based on virtual reality technology according to claim 2, wherein: The basic structure includes platforms, turnstiles, indoor ceilings, floors, walls, seats, billboards, lamps, and signs; the interactive elements include passengers, subway vehicles, and ticket vending machines.

5. The method for simulating and constructing subway scene image samples based on virtual reality technology according to claim 1, characterized in that: In Step 2, virtual reality technology is used to implement the simulation of the behavior of pedestrians passing through the turnstile. According to the contact between the human bounding box of pedestrians when passing through the turnstile and the bounding boxes in front of and behind the turnstile as the trigger condition, the opening and closing of the turnstile are automatically controlled.

6. The method for simulating and constructing subway scene image samples based on virtual reality technology according to claim 1, wherein Specifically, Step 3 is as follows: First, set a virtual camera on top of the turnstile in the virtual engine to automatically obtain image samples of pedestrians passing through the turnstile, generate data with monocular and binocular RGB images, and use the binocular disparity algorithm to restore the depth map corresponding to the RGB. Finally, automatically label the positions of the people and objects in the depth map samples according to the restored depth map samples to achieve the automatic acquisition of image samples of pedestrians passing through the turnstile and the automatic annotation of the positions of people and objects.

Citation Information

Patent Citations

  • Virtuality and reality combined subway turnstile system transit logic verification system and method

    CN106803296A

  • Mars scene binocular data set generation method based on virtual reality

    CN115035247A