Scene modeling method and electronic device

By combining global and target area image acquisition equipment, and through time-segmented dynamic acquisition and path planning, the problems of incomplete data and low accuracy in 3D reconstruction of accident sites have been solved, achieving efficient and accurate 3D model reconstruction and supporting rapid rescue and analysis.

CN121074253BActive Publication Date: 2026-04-28BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2025-08-22
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies for 3D reconstruction of accident sites suffer from low data accuracy and incomplete data collection, resulting in missing details and blind spots in the reconstruction model, which affects the accuracy of rescue measures and the analysis of accident causes.

Method used

The system uses a first image acquisition device to acquire global images. After identifying the target area, a second image acquisition device enters the target area to acquire images. The system combines global and target area video information to construct a 3D model. The system uses multiple video units of predetermined time lengths for 3D modeling and optimizes data acquisition through a time-segmented dynamic acquisition mechanism and path planning algorithm.

Benefits of technology

It improves the efficiency and completeness of data acquisition for 3D reconstruction of accident scenes, ensures the integrity of global information and local details of the model, enhances modeling precision and accuracy, and supports rapid and accurate rescue and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074253B_ABST
    Figure CN121074253B_ABST
Patent Text Reader

Abstract

The application provides a scene modeling method and an electronic device, wherein the scene modeling method comprises: collecting a global image of a target scene by using a first image collection device to obtain global video information; identifying a target region in the target scene based on the global video information; collecting an image in the target region by using a second image collection device to enter the target region to obtain target region video information; and constructing a three-dimensional model of the target scene based on the global video information and the target region video information. The scene modeling method and the electronic device provided by the application can effectively improve the data collection efficiency, completeness of the target scene and the modeling precision of the three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D reconstruction technology, and in particular to a scene modeling method and electronic device. Background Technology

[0002] Following sudden accidents such as explosions and fires, the reconstruction of the accident scene is crucial for on-site rescue and accident cause analysis. Three-dimensional reconstruction of the accident scene requires the acquisition of data from the actual accident site using data collection equipment. However, existing data collection methods suffer from low data accuracy and incomplete data collection, resulting in low precision in the reconstructed accident scene model, difficulty in including detailed information, blind spots, or missing details. This can easily lead to rescue and emergency measures that are not in line with the current situation, threatening the personal safety of rescue personnel and adversely affecting accident cause analysis. Summary of the Invention

[0003] In view of this, the purpose of this application is to propose a scene modeling method and an electronic device to solve the above-mentioned technical problems.

[0004] To achieve the above objectives, this application provides a scene modeling method, including:

[0005] The first image acquisition device is used to acquire global images of the target scene to obtain global video information;

[0006] Identify the target region in the target scene based on the global video information;

[0007] The second image acquisition device is used to enter the target area and acquire images within the target area to obtain video information of the target area;

[0008] A three-dimensional model of the target scene is constructed based on the global video information and the target area video information.

[0009] Optionally, the step of constructing a 3D model of the target scene based on the global video information and the target region video information includes:

[0010] The global video information and the target area video information are respectively divided into multiple video units of predetermined time length;

[0011] Perform 3D modeling operations on each video unit to obtain the 3D model unit corresponding to each video unit;

[0012] All the three-dimensional model units are fused together to obtain the three-dimensional model of the target scene.

[0013] Optionally, fusing all the three-dimensional model units to obtain a three-dimensional model of the target scene includes:

[0014] Multiple three-dimensional model units corresponding to the first image acquisition device are fused according to the time information of the global video information to obtain the first model corresponding to the first image acquisition device.

[0015] The multiple three-dimensional model units corresponding to the second image acquisition device are fused according to the time information of the video information of the target area to obtain the first model corresponding to the second image acquisition device.

[0016] By fusing all the first models, a three-dimensional model of the target scene is obtained.

[0017] Optionally, performing the 3D modeling operation for each video unit includes:

[0018] The video unit is converted into a series of consecutive image frames to obtain multiple corresponding image frames;

[0019] Extract the sparse point cloud of the target scene from each image frame to obtain the point cloud data of the target scene;

[0020] The point cloud data of the target scene is fitted with a Gaussian distribution and then rendered.

[0021] Optionally, the step of using a second image acquisition device to enter the target area and acquire images within the target area to obtain video information of the target area includes:

[0022] Control the second image acquisition device to enter the target area and acquire three-dimensional point cloud data of surrounding objects in real time;

[0023] An environmental map is constructed in real time based on the 3D point cloud data of the surrounding objects;

[0024] Based on the environmental map, a path planning algorithm is used to plan the acquisition path in real time, and the second image acquisition device is used to acquire images within the target area based on the acquisition path to obtain video information of the target area.

[0025] Optionally, it also includes:

[0026] Identify the target object in the target area based on the three-dimensional point cloud data of the surrounding objects;

[0027] The location information of the target data collection object is obtained based on the environmental map;

[0028] The process involves using a path planning algorithm based on the environmental map to plan the acquisition path in real time, and then using the second image acquisition device to acquire images within the target area based on the acquisition path to obtain video information of the target area, including:

[0029] Based on the environmental map and the location information of the target data collection object, a path planning algorithm is used to plan the data collection path in real time.

[0030] The second image acquisition device is used to acquire global images of the target area and omnidirectional images of the target object along the acquisition path to obtain video information of the target area.

[0031] Optionally, it also includes:

[0032] Real-time acquisition of environmental information of the second image acquisition device;

[0033] Based on the environmental information, risk areas are identified;

[0034] The location information of the risk area is obtained based on the environmental map;

[0035] The acquisition path of the second image acquisition device is adjusted in real time based on the location information of the risk area using a dynamic obstacle avoidance algorithm, so that the second image acquisition device can avoid the risk area.

[0036] Optionally, the method further includes:

[0037] The three-dimensional model of the target scene is converted into an interactive virtual image based on virtual reality or augmented reality technology.

[0038] Optionally, the first image acquisition device is an image acquisition drone with image acquisition function;

[0039] The step of acquiring global images of the target scene using the first image acquisition device to obtain global video information includes:

[0040] Obtain the location information of the target scene;

[0041] Multiple data collection routes are planned based on the location information of the target scene;

[0042] The aforementioned acquisition drone is used to collect global images of the target scene along multiple acquisition routes to obtain global video information.

[0043] Based on the same inventive concept, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0044] As described above, this application provides a scene modeling method and electronic device for 3D reconstruction of an accident scene. First, a first image acquisition device acquires global images of the target scene to obtain the overall situation of the accident scene. Then, based on the global video information, the target area requiring key data acquisition is identified, and a second image acquisition device enters the target area to acquire images within that area. The acquired data has richer dimensions, providing data support for the accuracy of subsequent modeling. Finally, a 3D model of the target scene is constructed based on the global video information and the target area video information, realizing the reconstruction of the accident scene. In this process, considering the complex nature of the accident scene environment, different image acquisition devices are used to acquire images and videos of the target scene, resulting in more complete and richer data, and higher data acquisition efficiency. Simultaneously, the second image acquisition device can enter the target area to acquire data from complex and dangerous areas, effectively acquiring detailed information and improving data completeness, obtaining more realistic information about the accident scene. The resulting 3D model of the target scene possesses global information without lacking local details. Furthermore, directly acquiring relevant images of the target scene yields video information containing more and more accurate accident scene details, further improving the accuracy and precision of the modeling. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of a scene modeling method according to an embodiment of this application;

[0047] Figure 2 This is a schematic diagram of a scene modeling device according to an embodiment of this application;

[0048] Figure 3 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0050] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0051] When sudden accidents such as explosions and fires occur, the environment at the accident site is extremely complex and dangerous. Failure to promptly grasp the situation at the scene will adversely affect on-site rescue, accident handling, and accident cause analysis. Three-dimensional reconstruction of the accident scene refers to constructing a corresponding three-dimensional model based on the actual data of the accident site. Relevant personnel can then comprehensively understand the true situation of the accident site based on the reconstructed 3D model, which is of great significance for on-site rescue and accident cause analysis.

[0052] 3D reconstruction of an accident scene first requires collecting data from the actual accident site using acquisition equipment. Then, 3D modeling technology is used to construct a 3D model of the accident scene based on the collected data, reconstructing the situation. Rescue personnel, accident analysts, and other relevant personnel can quickly grasp the true situation of the accident scene based on the reconstructed 3D model, facilitating correct rescue decisions and accident cause analysis. Current methods for collecting accident scene information mostly rely on laser scanning, photogrammetry, or structured light scanning. These methods have significant limitations in adaptability and efficiency in complex scenarios. Laser scanning equipment is often large and expensive, has stringent operating environment requirements, and is difficult to deploy flexibly in high-risk or complex terrains. Furthermore, its acquisition accuracy is poor, making it difficult to capture detailed information such as debris and building damage at the accident scene. Photogrammetry technology mainly refers to taking photographs of the accident scene to achieve subsequent modeling. While this technology is relatively flexible, it has high requirements for lighting conditions and shooting angles, and the data processing process is complex, making it difficult to meet the rapid response needs of accident scenes. Structured light scanning technology projects specific grating patterns (such as sinusoidal stripes or square grids) onto the surface of the object being scanned using a projector. When the object's surface has irregularities, the originally regular grating pattern deforms, thereby acquiring the object's three-dimensional information. Structured light scanning technology offers high precision, but it relies on specific equipment, has high equipment requirements, and is greatly affected by ambient light. The complex conditions at accident scenes such as explosions and fires make it difficult to meet the application conditions of structured light scanning technology. The shortcomings of existing information acquisition methods result in limited information collected from accident scenes, leading to insufficient data and resulting in low precision, lack of detailed information, blind spots, or missing details in the reconstructed accident scene model. When the reconstructed model is inaccurate, relevant personnel cannot make accurate judgments, which can easily lead to rescue and emergency measures that are not in line with the actual situation, thus threatening the personal safety of rescue personnel.

[0053] Meanwhile, existing information collection mostly relies on fixed or semi-fixed collection equipment, which lacks active mobility and cannot achieve multi-angle, full-coverage data collection. This further affects the accuracy of the reconstructed accident scene model, further increases blind spots or missing details, and makes it difficult to truly restore the full picture of the accident scene.

[0054] Furthermore, for some areas where data collection is more difficult, such as the core area of ​​an accident, areas with severe building damage, and more dangerous areas, existing technologies cannot collect the corresponding data in a timely manner. Usually, it is necessary to wait until the accident site is cleared before some information can be obtained. Insufficient and incomplete data acquisition will further affect the model reconstruction efficiency and accuracy, thereby adversely affecting the efficient deployment of rescue work.

[0055] It is evident that low accuracy and efficiency in reconstructing accident scenes directly increase the difficulty of on-site rescue and accident analysis. Therefore, achieving rapid and accurate modeling of accident scenes has become a crucial issue in current research and practice.

[0056] In view of this, this application provides a scene modeling method applicable to the 3D reconstruction of accident scenes such as fires and explosions, which can effectively improve the data acquisition efficiency, completeness, and accuracy of the 3D model of the target scene. Figure 1 As shown, the method includes:

[0057] S101. Use the first image acquisition device to acquire global images of the target scene to obtain global video information;

[0058] Specifically, the target scenario can be determined based on the central area and impact range of an accident scene such as a fire or explosion. Typically, accidents like fires and explosions have a certain impact on surrounding areas and buildings, especially explosions, where the shockwave can cause severe damage to nearby structures. When assessing an accident scene or conducting emergency rescue operations, it is necessary to understand not only the situation in the central area but also the scope of the accident's impact. Therefore, determining the target scenario based on the central area and impact range of the accident scene ensures that the reconstructed accident scene model contains information about the central area and its impact range, allowing relevant personnel to better understand the actual situation at the accident scene. The area within a certain distance from the center of the accident site can be defined as the affected zone. The size of the affected zone can be determined based on the actual situation on site. For example, if the affected zone is large, the area farther from the center can be defined as the affected zone, such as 0-50m, 100m, 150m, or 200m from the center. If the affected zone is small, the area closer to the center can be defined as the affected zone, such as 0-20m, 30m, 40m, or 45m from the center. There are no specific restrictions. The center of the accident site is determined based on the center point of the accident site. Specifically, the area 0-25m, 30m, or 50m from the center point is the center. The center point is the location of the accident, such as the ignition point or the explosion point. The size of the center and affected zones can be further determined by considering the type of accident, the location of the accident, and other specific factors. There are no specific restrictions.

[0059] S102. Identify the target region in the target scene based on the global video information;

[0060] Specifically, the target area is the area where information is difficult to collect and has a significant impact on the analysis of the accident scene. These are areas that require priority collection, such as areas with concentrated accident debris, the interior of buildings, and the central area of ​​the accident scene. When global video information is collected, a pre-trained semantic segmentation model is used to segment the information contained in the global video information, automatically identifying the target areas that need to be collected, thus providing the prerequisite for the subsequent entry of a second image acquisition device.

[0061] S103. Use the second image acquisition device to enter the target area and acquire images within the target area to obtain video information of the target area;

[0062] Specifically, the second image acquisition device is a mobile unmanned device equipped with an image acquisition device. Unmanned devices include quadruped robots (such as robotic dogs) and drones. A quadruped robot is a biomimetic robot that mimics the movement of quadrupedal animals. A robotic dog is a typical quadruped robot, possessing flexible movement capabilities and excellent adaptability to complex terrain, and can be used in this application to collect information about the target area. The quadruped robot can be designed to be relatively small, enabling it to enter some internal areas of the accident site, such as inside buildings. With its superior obstacle-crossing ability and stability, it can adapt to the complex environment and terrain of the accident site, deeply penetrating dangerous areas to perform comprehensive scanning of structural damage, debris distribution, and secondary disaster risk points, achieving efficient data collection of the target area and providing a foundation for the accurate construction of the subsequent accident site model. When the unmanned device is a drone, the situation in the target area is more complex; therefore, a drone with high flight stability, high accuracy, and small size can be selected. The image acquisition device is a device with image acquisition capabilities, such as a monocular camera. Other devices with image acquisition capabilities can also be used in this application; no specific limitations are imposed.

[0063] S104. Based on the global video information and the target area video information, a three-dimensional model of the target scene is constructed.

[0064] Specifically, the global video information contains overall information about the entire accident scene, while the target area video information contains information about important areas of the accident scene. This ensures that the constructed 3D model of the target scene (i.e., the 3D reconstruction model of the accident scene) has global information while not lacking local details.

[0065] Based on S101 to S104 of this application, a three-dimensional reconstruction of an accident scene is performed. First, a first image acquisition device is used to acquire global images of the target scene to obtain the overall situation of the accident scene. Then, based on the global video information, the target area that needs to be focused on is identified, and a second image acquisition device is used to enter the target area to acquire images within the target area. The acquired data has richer dimensions, providing data support for the accuracy of subsequent modeling. Finally, a three-dimensional model of the target scene is constructed based on the global video information and the target area video information, realizing the reconstruction of the accident scene. In this process, considering the complex characteristics of the accident scene environment, different image acquisition devices are used to acquire images and videos of the target scene, resulting in more complete and richer data, and higher data acquisition efficiency. At the same time, the second image acquisition device can enter the target area to acquire data in complex and dangerous areas, effectively acquiring detailed information and improving data completeness, obtaining more real information about the accident scene. The three-dimensional model of the target scene constructed thus has global information while not lacking local details. In addition, directly acquiring relevant images of the target scene, the obtained video information contains more and more accurate accident scene details, further improving the accuracy and precision of modeling.

[0066] In some embodiments, constructing a 3D model of the target scene based on the global video information and the target region video information includes:

[0067] The global video information and the target area video information are respectively divided into multiple video units of predetermined time length;

[0068] Perform 3D modeling operations on each video unit to obtain the 3D model unit corresponding to each video unit;

[0069] All the three-dimensional model units are fused together to obtain the three-dimensional model of the target scene.

[0070] When conducting emergency rescue and handling operations at accident sites, the efficiency of model reconstruction is crucial. Accident sites are complex environments, and the collected global and target area video information contains a large amount of data. Directly using this data to construct a 3D model would impact modeling efficiency due to the large volume of data to process simultaneously, thus affecting the efficiency of 3D reconstruction and consequently hindering rescue and emergency handling efforts. Furthermore, the complexity of accident sites—for example, the possibility of secondary explosions—can lead to quality issues such as blurry data segments. If these data segments are directly used to construct the 3D model, the quality of these data segments will affect the overall result, impacting the overall modeling effectiveness and accuracy.

[0071] In this application, global video information and target area video information are divided into multiple video units of predetermined time lengths. 3D modeling is performed independently for each video unit, requiring less data processing time and effectively ensuring efficient 3D modeling. Furthermore, quality issues in any video unit do not affect the overall modeling quality, while ensuring accurate processing of data from each video unit in subsequent 3D modeling processes, thus improving modeling precision. Because 3D modeling can be performed independently for each video unit after dividing it into predetermined time lengths, data acquisition and modeling can be performed simultaneously. That is, while acquiring global images of the target scene or images within the target area, the acquired images are divided into video units, and corresponding 3D model units are constructed from these video units. When all global images are acquired to obtain global video information, or all images within the target area are acquired to obtain target area video information, the corresponding 3D model unit for each video unit is also completed. At this point, all 3D model units are directly fused to obtain the 3D model of the target scene, effectively improving the efficiency of 3D model construction and enhancing the real-time performance of 3D reconstruction of the accident scene. After obtaining the 3D model unit corresponding to each video unit, a quality assessment is performed on each 3D model unit. The main assessment indicators include geometric accuracy and texture clarity. 3D model units that do not meet the quality standards are discarded, and an automatic re-acquisition mechanism is triggered. This means that based on the acquisition information of the video unit corresponding to the 3D model unit that does not meet the quality standards, the corresponding information is re-acquired using the corresponding image acquisition device (i.e., the first image acquisition device or the second image acquisition device), and the corresponding 3D model unit is reconstructed based on the re-acquisition information, further ensuring the modeling accuracy of the subsequent target scene's 3D model. Specifically, when the video unit corresponding to the 3D model unit that does not meet the quality standards is acquired by the first image acquisition device, it is re-acquired using the first image acquisition device; similarly, when the video unit corresponding to the 3D model unit that does not meet the quality standards is acquired by the second image acquisition device, it is re-acquired using the second image acquisition device. The predetermined time length can be 1–10 minutes, specifically 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes, or other time lengths are not limited.

[0072] Multiple video units of predetermined durations are obtained from the same video. Adjacent video units may partially overlap, allowing for better fusion of subsequent 3D model units. No specific restrictions are imposed. For example, when the corresponding video length in the global video information is 4 minutes, the first, second, and third video units correspond to minutes 0-2, 1-3, and 2-4 of the global video information, respectively; when the corresponding video length in the target region video information is 5 minutes, the first, second, third, and fourth video units correspond to minutes 0-2, 1-3, 2-4, and 3-5 of the target region video information, respectively.

[0073] In some embodiments, fusing all the three-dimensional model units to obtain a three-dimensional model of the target scene includes:

[0074] Multiple three-dimensional model units corresponding to the first image acquisition device are fused according to the time information of the global video information to obtain the first model corresponding to the first image acquisition device.

[0075] The multiple three-dimensional model units corresponding to the second image acquisition device are fused according to the time information of the video information of the target area to obtain the first model corresponding to the second image acquisition device.

[0076] By fusing all the first models, a three-dimensional model of the target scene is obtained.

[0077] Both the first image acquisition device and the second image acquisition device have unique device identification information. When multiple first image acquisition devices or multiple second image acquisition devices are used for information acquisition, each first image acquisition device and each second image acquisition device has unique device identification information.

[0078] Both the global video information and the target area video information contain device identification information to identify the corresponding image acquisition device. During the fusion process of 3D model units, all 3D model units with the same device identification information (i.e., the same first image acquisition device or the same second image acquisition device) are fused first; the fusion order of 3D model units with the same device identification information is determined according to the time information of their corresponding video units in the global video information or the target area video information.

[0079] For example, when the corresponding video length in the global video information is 9 minutes, with a predetermined time length of 3 minutes, the first video unit corresponds to the content from 0 to 3 minutes of the global video information, the second video unit corresponds to the content from 3 to 6 minutes of the global video information, and the third video unit corresponds to the content from 6 to 9 minutes of the global video information. Three-dimensional modeling operations are performed on the first, second, and third video units respectively to obtain corresponding three-dimensional model units. During the fusion of the three-dimensional model units, the three-dimensional model units are fused according to the time information of the global video information, that is, the model fusion is performed in the order of the three-dimensional model units corresponding to the first video unit, the second video unit, and the third video unit, thereby obtaining the first model corresponding to the first image acquisition device, effectively ensuring the accuracy and consistency of the three-dimensional model unit fusion.

[0080] For example, when the video length corresponding to the target area video information is 6 minutes, with a predetermined time length of 2 minutes, the first video unit corresponds to the content from 0 to 2 minutes of the target area video information, the second video unit corresponds to the content from 2 to 4 minutes of the target area video information, and the third video unit corresponds to the content from 4 to 6 minutes of the target area video information. Three-dimensional modeling operations are performed on the first, second, and third video units respectively to obtain corresponding three-dimensional model units. During the fusion of the three-dimensional model units, the three-dimensional model units are fused according to the time information of the target area video information, that is, the model fusion is performed in the order of the three-dimensional model units corresponding to the first video unit, the second video unit, and the third video unit, thereby obtaining the first model corresponding to the second image acquisition device, effectively ensuring the accuracy and consistency of the three-dimensional model unit fusion.

[0081] After obtaining the first model corresponding to each device identification information, that is, after obtaining the first model corresponding to each second image acquisition device or the first model corresponding to each first image acquisition device, all the first models are then fused to obtain the three-dimensional model of the target scene.

[0082] In the fusion process of multiple 3D model units, point cloud data from each 3D model unit is used to perform initial alignment using a feature-point-based coarse registration technique. The Iterative Closest Point (ICP) algorithm is then employed to further optimize the registration accuracy, ensuring that the acquired data can be precisely aligned within the same coordinate system. Finally, 3D Gaussian Splatting is used for final fusion to achieve seamless scene reconstruction. In the fusion process of multiple first models, point cloud data from each first model is used to perform initial alignment using a feature-point-based coarse registration technique. Then, the ICP algorithm is used to further optimize the registration accuracy for finer alignment, improving the fusion accuracy of the first models. Finally, 3D Gaussian Splatting is used for final fusion to achieve seamless scene reconstruction.

[0083] This application employs a time-segmented dynamic acquisition mechanism to achieve global image and target area image acquisition. Specifically, multiple first image acquisition devices and multiple second image acquisition devices are used to implement the time-segmented dynamic acquisition mechanism. Based on the location information of the target scene, the target scene is divided into multiple acquisition areas. For each first image acquisition device, a corresponding acquisition area, acquisition time, and acquisition order are assigned. Each first image acquisition device acquires image segments according to the acquisition area, acquisition time, and acquisition order. The image segments acquired by each first image acquisition device contain its corresponding device identification information, acquisition order information, and acquisition area identification information. The image data acquired by all first image acquisition devices constitutes the global video information. After obtaining the global video information, the corresponding target area is determined, and based on the target area, multiple second image acquisition devices are assigned corresponding acquisition areas, acquisition times, and acquisition orders. Each second image acquisition device acquires images within the target area according to the acquisition area, acquisition time, and acquisition order. The image segments acquired by each second image acquisition device contain its corresponding device identification information, acquisition order information, and acquisition area identification information. The image data acquired by all second image acquisition devices constitutes the target area video information. By using a time-segmented dynamic acquisition mechanism, data acquisition efficiency and accuracy can be effectively improved. By using the equipment identification information, acquisition sequence information, and acquisition area identification information marked on each image segment, it can be effectively ensured that the data collected by different devices can be accurately matched and integrated to generate a complete and accurate three-dimensional model of the accident scene.

[0084] In some embodiments, performing the 3D modeling operation for each video unit includes:

[0085] The video unit is converted into a series of consecutive image frames to obtain multiple corresponding image frames;

[0086] Extract the sparse point cloud of the target scene from each image frame to obtain the point cloud data of the target scene;

[0087] The point cloud data of the target scene is fitted with a Gaussian distribution and then rendered.

[0088] When performing 3D modeling operations on each video unit, each video unit is first serialized, converting it into consecutive image frames. This provides a foundation for subsequent point cloud data extraction, ensuring that the temporal and spatial information of each frame is accurately preserved. Then, based on structured light technology or Structure from Motion (SFM) technology, sparse point clouds related to the target scene are extracted from each image frame to obtain the point cloud data of the target scene. The extraction of point cloud data can accurately capture the geometric information of the scene, providing rich spatial data support for 3D reconstruction. Finally, the point cloud data of the target scene is fitted with a Gaussian distribution. After fitting, 3D Gaussian sputtering technology is used for efficient rendering, effectively preserving the details of the target scene while avoiding distortion during rendering, thus improving the rendering effect and the expressiveness of the data.

[0089] In some embodiments, the step of using a second image acquisition device to enter the target area and acquire images within the target area to obtain video information of the target area includes:

[0090] Control the second image acquisition device to enter the target area and acquire three-dimensional point cloud data of surrounding objects in real time;

[0091] An environmental map is constructed in real time based on the 3D point cloud data of the surrounding objects;

[0092] Based on the environmental map, a path planning algorithm is used to plan the acquisition path in real time, and the second image acquisition device is used to acquire images within the target area based on the acquisition path to obtain video information of the target area.

[0093] A 3D LiDAR can be mounted on the second image acquisition device to collect real-time 3D point cloud data of surrounding objects. Once the second image acquisition device enters the target area, it uses the mounted 3D LiDAR to scan the surrounding objects in real time, acquiring their 3D point cloud data. Then, an environmental map is constructed using this data, and a path planning algorithm is used to plan the acquisition path in real time based on the map. The second image acquisition device then acquires images within the target area based on the planned path. The path planning algorithm can be the D*Lite algorithm, a highly efficient incremental path planning algorithm suitable for robot navigation in dynamic environments. Through reverse search and incremental updates, it quickly adjusts the path when the environment changes, avoiding the high computational cost of global replanning. The environment and terrain within the target area are often complex. After entering the target area, the second image acquisition device can autonomously plan its path based on the surrounding objects and promptly change the acquisition path, effectively ensuring the smooth progress of the acquisition work, improving data acquisition efficiency, and thus enhancing the real-time performance and efficiency of 3D reconstruction of the accident scene.

[0094] In some embodiments, it also includes:

[0095] Identify the target object in the target area based on the three-dimensional point cloud data of the surrounding objects;

[0096] The location information of the target data collection object is obtained based on the environmental map;

[0097] The process involves using a path planning algorithm based on the environmental map to plan the acquisition path in real time, and then using the second image acquisition device to acquire images within the target area based on the acquisition path to obtain video information of the target area, including:

[0098] Based on the environmental map and the location information of the target data collection object, a path planning algorithm is used to plan the data collection path in real time.

[0099] The second image acquisition device is used to acquire global images of the target area and omnidirectional images of the target object along the acquisition path to obtain video information of the target area.

[0100] After acquiring the 3D point cloud data of surrounding objects, the target objects in the target area can be determined based on this data. These target objects are those that require focused acquisition, such as structural damage, accident debris, and secondary disaster risk points. These objects have a significant impact on hazard assessment, accident cause analysis, and accident reconstruction at the accident site. Therefore, after identifying the target objects, their location information is obtained from the environmental map. Then, based on the environmental map and the target object's location information, a path planning algorithm is used to plan the acquisition path in real time. Finally, a second image acquisition device is used to acquire global images of the target area and omnidirectional images of the target objects along the acquisition path. The resulting video information of the target area not only contains overall data of the target area but also detailed data of the target objects, resulting in a more accurate model and a higher degree of accident scene reconstruction.

[0101] In some embodiments, it also includes:

[0102] Real-time acquisition of environmental information of the second image acquisition device;

[0103] Based on the environmental information, risk areas are identified;

[0104] The location information of the risk area is obtained based on the environmental map;

[0105] The acquisition path of the second image acquisition device is adjusted in real time based on the location information of the risk area using a dynamic obstacle avoidance algorithm, so that the second image acquisition device can avoid the risk area.

[0106] Specifically, environmental information includes the concentration of combustible / toxic gases, information on high temperatures or hidden ignition sources, and ground structure information. The second image acquisition device can be equipped with gas sensors and thermal imaging cameras. The gas sensors collect the concentration of combustible / toxic gases (such as methane and liquefied petroleum gas) in the environment where the second image acquisition device is located. The thermal imaging camera determines the presence of high temperatures or hidden ignition sources. 3D LiDAR collects ground structure information to determine if the ground structure has been damaged. When risk points such as high concentrations of combustible / toxic gases, the presence of high temperatures or hidden ignition sources, or severe damage to the ground structure are identified, the corresponding area can be marked as a risk area. Then, the location information of the risk area is obtained based on the environmental map, and the acquisition path is adjusted in real time according to the location information of the risk area to ensure that the second image acquisition device avoids the risk area and successfully acquires images of the target area, providing a data foundation for the subsequent construction of the 3D model.

[0107] Risk areas can be further divided into high-risk and medium-risk areas. High-risk areas are those that cannot be safely bypassed. When encountering a risk area that cannot be bypassed, the area can be marked as high-risk, and data collection along that path should be abandoned to ensure equipment safety. Medium-risk areas are those with risk points but can be bypassed. By using a safe path to bypass the risk area, data collection can proceed smoothly. Different detour distances can be set according to different risk points. For example, for areas with high temperatures or hidden fire sources, a longer detour distance can be set, such as 1-2 meters away from the risk point; for areas with damaged ground structures, a smaller detour distance can be set, such as 0.5-1 meter away from the risk point. The detour distance can be determined based on the specific risk, without any specific restrictions. Furthermore, the collected data can be dynamically adjusted according to environmental conditions, such as automatically adjusting the collection time (e.g., extending or shortening the collection time) and adjusting parameters such as transmission frequency, thereby optimizing data quality and improving the flexibility and adaptability of the collection process.

[0108] In some embodiments, the method further includes:

[0109] The three-dimensional model of the target scene is converted into an interactive virtual image based on virtual reality or augmented reality technology.

[0110] In recent years, the rapid development of virtual reality (VR) and augmented reality (AR) technologies has brought profound changes to many industries, particularly demonstrating enormous potential in accident scene reconstruction and emergency rescue. By converting the 3D model of a target scene into an interactive virtual image, users wearing corresponding VR or AR devices can interact with the 3D model of the target scene. Users can flexibly observe the accident scene from multiple angles through intuitive interactive operations, including autonomous navigation, model scaling, rotation, and translation. Furthermore, multiple users can enter the same virtual scene simultaneously through different VR devices, engaging in synchronous interaction and information sharing, significantly improving the collaborative efficiency of accident investigation and analysis, and promoting multi-party participation and real-time decision-making. Specifically, the conversion of interactive virtual images is achieved through Unity 3D, which involves importing the 3D model of the target scene into Unity 3D for rendering, converting it into a virtual image that can be interacted with using VR or AR devices.

[0111] In some embodiments, the first image acquisition device is an image acquisition drone with image acquisition function;

[0112] The step of acquiring global images of the target scene using the first image acquisition device to obtain global video information includes:

[0113] Obtain the location information of the target scene;

[0114] Multiple data collection routes are planned based on the location information of the target scene;

[0115] The aforementioned acquisition drone is used to collect global images of the target scene along multiple acquisition routes to obtain global video information.

[0116] The first image acquisition device can be a drone equipped with an image acquisition device, which is any device with image acquisition capabilities, such as a monocular camera. Other devices with image acquisition capabilities are also applicable in this application, and there are no specific limitations. However, when it is necessary to use a drone to collect data on a target scene, the location information of the target scene must first be obtained. Then, multiple acquisition routes are planned based on the location information of the target scene. Each acquisition route must have at least some overlap. The drone is controlled to fly along multiple acquisition routes to collect video from multiple angles and altitudes. By leveraging the drone's advantage of high-altitude observation, comprehensive data on the target scene can be collected, providing a comprehensive accident impact assessment.

[0117] The collected global video information, target area video information, and the final constructed 3D model of the target scene are encrypted and stored in a standardized digital form, forming a complete, traceable, and tamper-proof electronic archive. This provides key evidence for subsequent legal proceedings, accident investigations, and liability determination, and provides strong support for the intelligent upgrading of urban emergency management and disaster response capabilities.

[0118] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.

[0119] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0120] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides a scene modeling device.

[0121] refer to Figure 2 The scene modeling device includes:

[0122] The first acquisition module 201 is used to acquire global images of the target scene using the first image acquisition device to obtain global video information;

[0123] The recognition module 202 is used to identify the target region in the target scene based on the global video information;

[0124] The second acquisition module 203 is used to enter the target area using the second image acquisition device to acquire images within the target area and obtain video information of the target area;

[0125] The 3D modeling module 204 is used to construct a 3D model of the target scene based on the global video information and the target area video information.

[0126] In some embodiments, the 3D modeling module 204 includes:

[0127] A segmentation unit is used to divide the global video information and the target area video information into multiple video units of predetermined time lengths, respectively.

[0128] Modeling unit, used to perform 3D modeling operations for each video unit to obtain the 3D model unit corresponding to each video unit;

[0129] A fusion unit is used to fuse all the three-dimensional model units to obtain a three-dimensional model of the target scene.

[0130] In some embodiments, the fusion unit includes:

[0131] The first fusion element is used to fuse multiple three-dimensional model units corresponding to the first image acquisition device according to the time information of the global video information to obtain the first model corresponding to the first image acquisition device.

[0132] The second fusion element is used to fuse the multiple three-dimensional model units corresponding to the second image acquisition device according to the time information of the target area video information to obtain the first model corresponding to the second image acquisition device.

[0133] The third fusion element is used to fuse all the first models to obtain a three-dimensional model of the target scene.

[0134] In some embodiments, performing the 3D modeling operation for each video unit includes:

[0135] The video unit is converted into a series of consecutive image frames to obtain multiple corresponding image frames;

[0136] Extract the sparse point cloud of the target scene from each image frame to obtain the point cloud data of the target scene;

[0137] The point cloud data of the target scene is fitted with a Gaussian distribution and then rendered.

[0138] In some embodiments, the second acquisition module 203 includes:

[0139] The point cloud real-time acquisition unit is used to control the second image acquisition device to enter the target area and acquire the three-dimensional point cloud data of the surrounding objects in real time.

[0140] A map building unit is used to build an environmental map in real time based on the three-dimensional point cloud data of the surrounding objects;

[0141] The path planning unit is used to plan the acquisition path in real time based on the environmental map using a path planning algorithm, and to acquire images of the target area based on the acquisition path using the second image acquisition device to obtain video information of the target area.

[0142] In some embodiments, the second acquisition module 203 further includes:

[0143] An object recognition unit is used to identify target objects in the target area based on the three-dimensional point cloud data of the surrounding objects.

[0144] The first acquisition unit is used to acquire the location information of the target object based on the environmental map;

[0145] The path planning unit is also used for:

[0146] Based on the environmental map and the location information of the target data collection object, a path planning algorithm is used to plan the data collection path in real time.

[0147] The second image acquisition device is used to acquire global images of the target area and omnidirectional images of the target object along the acquisition path to obtain video information of the target area.

[0148] In some embodiments, the second acquisition module 203 further includes:

[0149] An environmental information acquisition unit is used to acquire environmental information of the second image acquisition device in real time.

[0150] A risk analysis unit is used to determine risk areas based on the environmental information;

[0151] The second acquisition unit is used to acquire the location information of the risk area based on the environmental map;

[0152] The obstacle avoidance unit is used to adjust the acquisition path of the second image acquisition device in real time based on the location information of the risk area using a dynamic obstacle avoidance algorithm, so that the second image acquisition device avoids the risk area.

[0153] In some embodiments, the apparatus further includes:

[0154] The conversion module is used to convert the 3D model of the target scene into an interactive virtual image based on virtual reality or augmented reality technology.

[0155] In some embodiments, the first image acquisition device is an image acquisition drone with image acquisition function;

[0156] The first acquisition module includes:

[0157] The third acquisition unit is used to acquire the location information of the target scene;

[0158] The route planning unit is used to plan multiple data collection routes based on the location information of the target scene;

[0159] The global image acquisition unit is used to acquire global images of the target scene using the acquisition drone along multiple acquisition routes to obtain global video information.

[0160] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.

[0161] The apparatus described above is used to implement a scene modeling method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0162] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a scene modeling method as described in any of the above embodiments.

[0163] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0164] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0165] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0166] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0167] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0168] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0169] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0170] The electronic devices described above are used to implement a scene modeling method in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0171] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute a scene modeling method as described in any of the above embodiments.

[0172] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0173] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute a scene modeling method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0174] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to execute a scene modeling method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0175] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0176] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0177] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0178] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0179] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0180] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0181] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0182] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the claims of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A scene modeling method, characterized in that, include: The first image acquisition device is used to acquire global images of the target scene to obtain global video information; Identify the target region in the target scene based on the global video information; The second image acquisition device is used to enter the target area and acquire images within the target area to obtain video information of the target area; A three-dimensional model of the target scene is constructed based on the global video information and the target area video information; The step of constructing a 3D model of the target scene based on the global video information and the target region video information includes: The global video information and the target area video information are respectively divided into multiple video units of predetermined time length; Perform 3D modeling operations on each video unit to obtain the 3D model unit corresponding to each video unit; All the three-dimensional model units are fused together to obtain the three-dimensional model of the target scene.

2. The scene modeling method according to claim 1, characterized in that, The step of fusing all the three-dimensional model units to obtain the three-dimensional model of the target scene includes: Multiple three-dimensional model units corresponding to the first image acquisition device are fused according to the time information of the global video information to obtain the first model corresponding to the first image acquisition device. The multiple three-dimensional model units corresponding to the second image acquisition device are fused according to the time information of the video information of the target area to obtain the first model corresponding to the second image acquisition device. By fusing all the first models, a three-dimensional model of the target scene is obtained.

3. The scene modeling method according to claim 1, characterized in that, The step of performing 3D modeling operations for each video unit includes: The video unit is converted into a series of consecutive image frames to obtain multiple corresponding image frames; Extract the sparse point cloud of the target scene from each image frame to obtain the point cloud data of the target scene; The point cloud data of the target scene is fitted with a Gaussian distribution and then rendered.

4. The scene modeling method according to claim 1, characterized in that, The step of using a second image acquisition device to enter the target area and acquire images within the target area to obtain video information of the target area includes: Control the second image acquisition device to enter the target area and acquire three-dimensional point cloud data of surrounding objects in real time; An environmental map is constructed in real time based on the 3D point cloud data of the surrounding objects; Based on the environmental map, a path planning algorithm is used to plan the acquisition path in real time, and the second image acquisition device is used to acquire images within the target area based on the acquisition path to obtain video information of the target area.

5. The scene modeling method according to claim 4, characterized in that, Also includes: Identify the target object in the target area based on the three-dimensional point cloud data of the surrounding objects; The location information of the target data collection object is obtained based on the environmental map; The process involves using a path planning algorithm based on the environmental map to plan the acquisition path in real time, and then using the second image acquisition device to acquire images within the target area based on the acquisition path to obtain video information of the target area, including: Based on the environmental map and the location information of the target data collection object, a path planning algorithm is used to plan the data collection path in real time. The second image acquisition device is used to acquire global images of the target area and omnidirectional images of the target object along the acquisition path to obtain video information of the target area.

6. The scene modeling method according to claim 4, characterized in that, Also includes: Real-time acquisition of environmental information of the second image acquisition device; Based on the environmental information, risk areas are identified; The location information of the risk area is obtained based on the environmental map; The acquisition path of the second image acquisition device is adjusted in real time based on the location information of the risk area using a dynamic obstacle avoidance algorithm, so that the second image acquisition device can avoid the risk area.

7. The scene modeling method according to claim 1, characterized in that, The method further includes: The three-dimensional model of the target scene is converted into an interactive virtual image based on virtual reality or augmented reality technology.

8. The scene modeling method according to claim 1, characterized in that, The first image acquisition device is an image acquisition drone with image acquisition function; The step of acquiring global images of the target scene using the first image acquisition device to obtain global video information includes: Obtain the location information of the target scene; Multiple data collection routes are planned based on the location information of the target scene; The aforementioned acquisition drone is used to collect global images of the target scene along multiple acquisition routes to obtain global video information.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Traffic situation monitoring method and device based on real scene fusion, equipment and medium

    CN113033412A