Scene reconstruction method and device and medium
By removing moving objects in the two-dimensional scene image and using neural radiation fields to perform scene reconstruction, the problem of large-scale or dynamic scenes in the prior art is solved, and the realistic scene reconstruction effect is achieved.
Patent Information
- Application Number
- CN202311466517.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
The existing neural radiation field algorithm can only be used for scene reconstruction of small-scale static scenes, and cannot effectively deal with large-scale scenes or scenes containing moving objects.
By obtaining the two-dimensional scene image, removing the moving objects therein, a segmented two-dimensional image is obtained, and then the scene reconstruction is performed using the neural radiation field.
Eliminate the impact of mobile objects on scene reconstruction, ensure the effect of scene reconstruction, and realize realistic scene reconstruction, which is especially suitable for large-scale scenes such as city-level scene reconstruction.
Smart Images

Figure CN119941971A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a scene reconstruction method, device and medium. Background Art
[0002] Neural Radiance Fields (NeRF) is a three-dimensional (3D) scene reconstruction and image synthesis algorithm that combines computer vision and computer graphics. Given a set of images with camera pose information, NeRF can learn the geometric shape and lighting texture information in the scene through a multi-layer perceptron (MLP), complete the dense 3D reconstruction of the scene, and generate images from new perspectives.
[0003] In related technologies, neural radiation fields can generally only reconstruct small-scale static scenes. Summary of the invention
[0004] In order to overcome the problems existing in the related art, the present disclosure provides a scene reconstruction method, device and medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a scene reconstruction method is provided, comprising: acquiring a two-dimensional image of a scene; removing moving objects in the two-dimensional image of the scene to obtain a segmented two-dimensional image; and reconstructing the scene based on the segmented two-dimensional image using a neural radiation field.
[0006] According to a second aspect of an embodiment of the present disclosure, a scene reconstruction device is provided, comprising: an acquisition module configured to acquire a two-dimensional image of a scene; a removal module configured to remove moving objects in the two-dimensional image of the scene to obtain a segmented two-dimensional image; and a reconstruction module configured to reconstruct the scene based on the segmented two-dimensional image using a neural radiation field.
[0007] According to a third aspect of an embodiment of the present disclosure, a scene reconstruction device is provided, comprising: a processor; and a memory for storing processor executable instructions; wherein the processor is configured to execute any one of the methods described in the first aspect.
[0008] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the scene reconstruction method provided in the first aspect of the present disclosure are implemented.
[0009] By adopting the above technical solution, since the moving objects in the two-dimensional image of the scene are first removed to obtain a segmented two-dimensional image, and then the scene is reconstructed using the neural radiation field based on the segmented two-dimensional image, the influence of the moving objects on the scene reconstruction can be eliminated, the effect of the scene reconstruction is ensured, and realistic scene reconstruction can be achieved.
[0010] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0012] Figure 1 The figure is a flow chart of a scene reconstruction method according to an exemplary embodiment.
[0013] Figure 2 The present invention is a flowchart of removing moving objects in a two-dimensional image of a scene according to an embodiment of the present invention.
[0014] Figure 3 The present invention is a flowchart of a two-dimensional image rendering according to an embodiment of the present invention.
[0015] Figure 4 It is a schematic block diagram of a scene reconstruction device according to an embodiment of the present disclosure.
[0016] Figure 5 is a schematic structural diagram of a removal module according to an embodiment of the present disclosure.
[0017] Figure 6 It is a schematic diagram of the processing result of the removal module according to an embodiment of the present disclosure.
[0018] Figure 7 is a block diagram of a vehicle according to an exemplary embodiment.
[0019] Figure 8 It is a block diagram of a device for scene reconstruction according to an exemplary embodiment. DETAILED DESCRIPTION
[0020] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0021] It should be noted that all actions of acquiring signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located, and with the authorization given by the owner of the corresponding device.
[0022] NeRF works by training a neural network model that can infer the color radiation value and volume density of each point in the scene from limited observation data. That is, NeRF uses a set of rays to sample points in the scene, and then predicts the color radiation value and volume density of each point through the neural network model. By sampling and predicting multiple rays, a 3D model of the entire scene can be obtained.
[0023] The advantage of NeRF is that it can generate high-quality 3D scene models with realistic lighting effects and detailed expressions. However, in related technologies, NeRF also has some challenges, such as the high demand for processing large-scale scenes and training data. In related technologies, neural radiation fields can generally only be used for scene reconstruction of small-scale static scenes. When there are moving objects in the scene or there is a front-to-back occlusion relationship between objects in the scene, the theoretical assumptions on which the neural radiation field algorithm is based will no longer exist, and the scene reconstruction effect will be greatly reduced.
[0024] In the field of autonomous driving, the road scenes of autonomous driving usually cover a large area of the city, and there are also dynamic targets such as moving vehicles and pedestrians in the scene. When 3D scene modeling of autonomous driving road scenes is required, the general neural radiation field algorithm cannot be directly applied to the autonomous driving scene.
[0025] The traditional neural radiation field can be regarded as a 5D function, that is: the neural radiation field function represents a local scene as a function with a 5D vector as input, which includes the 3D coordinate position (x, y, z) of the spatial point and the camera viewing direction (θ, φ), θ represents the camera inclination angle, and φ represents the camera azimuth angle; the output is the color c = (r, g, b) of the 3D spatial point related to the viewing direction and the volume density σ of the 3D spatial point position.
[0026] After obtaining the function of the neural radiation field, when synthesizing an image at a certain perspective, it is completed by traversing all pixels of the image, that is, a ray is emitted from the position of the camera's optical center to the pixel, and the volume density σ of each point on the ray and the color radiation value presented at the position of each point under the ray's perspective can be queried, that is, c = (r, g, b), where the volume density σ is used to calculate the weight, and the weighted sum of the color radiation values at the point can be obtained to obtain the image pixel color. In other words, for the neural radiation field, only one table needs to be maintained, and the subscript of each element in the table is a 5D vector (x, y, z, θ, φ). The color radiation value c = (r, g, b) and the volume density σ can be obtained by looking up the table, and then the image pixel value can be obtained by using the volume rendering method. The neural radiation field uses a neural network to construct a continuously changing 5D function to perform functions similar to a 5D query table.
[0027] Since an infinite number of points can be sampled on a ray in theory, volume rendering can be completed by integration, but this method is not convenient for programming and gradient backpropagation optimization. The general approach is: first set the size of the scene, set the nearest and farthest ends of the ray to tn and tf respectively, then divide [tn, tf] evenly into N parts, and then perform uniform random sampling in each small area. Then the color of a pixel can be simplified to the form of summation:
[0028]
[0029] Among them, δ i is the distance between two neighboring sampling points, T i for:
[0030]
[0031] The neural radiation field algorithm contains two MLPs. The first MLP is used to calculate the volume density σ of each spatial point (x, y, z) and the implicit feature vector h at the corresponding position, which can be expressed as:
[0032] σ,h=MLP_1(x,y,z)
[0033] The second MLP is used to calculate the color radiation value of each spatial point (x, y, z) under the camera view (θ, φ), which can be expressed as:
[0034] c=r,g,b=MLP_2(h,θ,φ)
[0035] Traditional neural radiation fields are suitable for single objects or small-scale closed scenes. For single objects or small-scale closed scenes, the training data volume of traditional neural radiation fields is usually 50 to 200 multi-angle pictures. The data collection of single objects or small-scale scenes can generally be completed within a few minutes. Its basic assumption is that there is no problem of front and back occlusion of moving objects in the whole process.
[0036] As mentioned above, the volume density σ is related to the spatial point position (x, y, z), has nothing to do with the camera's viewing angle (θ, φ), and has nothing to do with the time when the camera takes the photo. That is, cameras with different postures and angles shooting the same physical point should obtain the same volume density σ; the same camera shooting the same physical point at the same position and with the same posture should obtain the same volume density σ at different times.
[0037] R, G, B are not only related to the spatial point (x, y, z), but also to the camera's perspective (θ, φ), but have nothing to do with the time when the camera takes the photo. That is, a camera with the same posture angle should get the same RGB color when taking photos of the same physical spatial point at different times.
[0038] When there are moving objects in the scene, due to the presence of front and back occlusions, the photos taken at different times are obviously different in the area where the moving objects are present, that is, the RGB pixel colors of the image are different. Traditional neural radiation fields cannot handle the situation where there are moving objects in the scene.
[0039] Figure 1 FIG. 1 is a flow chart of a scene reconstruction method according to an exemplary embodiment. The scene reconstruction method is applicable to the case where there are moving objects in the scene. Figure 1 As shown, the scene reconstruction method includes the following steps S11 to S13.
[0040] In step S11, a two-dimensional image of the scene is acquired.
[0041] The two-dimensional image of the scene may be acquired by photographing the scene with a camera, or the two-dimensional image of the scene may be acquired from a storage location.
[0042] In step S12, moving objects in the two-dimensional image of the scene are removed to obtain a segmented two-dimensional image.
[0043] Moving objects may include vehicles, pedestrians, etc. moving in the scene.
[0044] In addition, an object segmentation model can be used to remove moving objects in the two-dimensional image of the scene.
[0045] In step S13, the scene is reconstructed using the neural radiation field based on the segmented two-dimensional image.
[0046] Scene reconstruction based on segmented two-dimensional images means that when reconstructing the scene, only pixels in static areas (that is, areas without moving objects) are considered, and pixels in areas with moving objects are not considered, that is, the existence of moving objects is ignored.
[0047] By adopting the above technical solution, since the moving objects in the two-dimensional image of the scene are first removed to obtain a segmented two-dimensional image, and then the scene is reconstructed using the neural radiation field based on the segmented two-dimensional image, the influence of the moving objects on the scene reconstruction can be eliminated, the effect of the scene reconstruction is ensured, and realistic scene reconstruction can be achieved.
[0048] Figure 2 FIG. 1 is a flowchart of removing moving objects in a two-dimensional image of a scene according to an embodiment of the present disclosure. Figure 2 As shown, the process includes the following steps S121 and S122.
[0049] In step S121 , features in the two-dimensional image of the scene are extracted.
[0050] In some embodiments, features in the scene two-dimensional image can be extracted by performing convolution processing on the scene two-dimensional image and performing feature fusion on the features obtained by the convolution processing. Among them, any convolutional neural network structure can be used for convolution processing, for example, the RegNet 800 convolution structure can be used. In addition, any feature fusion method can be used for feature fusion, for example, a feature pyramid network structure, a bidirectional feature pyramid network structure (BiFPN), etc. can be used for feature fusion. Through convolution processing and feature fusion, target detection and segmentation can be achieved, such as detecting moving objects from the scene two-dimensional image.
[0051] In step S122, the extracted features are processed based on the attention mechanism to generate a mask of the moving object to remove the moving object.
[0052] That is, the features after feature fusion are processed based on the attention mechanism to generate a mask of the moving object, thereby removing the moving object from the scene two-dimensional image. In addition, when performing attention processing, the number of categories of moving objects in the scene two-dimensional image can be considered. For example, if the moving objects in the scene two-dimensional image include bicycles, cars, tricycles and pedestrians, then the number of categories of moving objects is 4, so that moving objects of different categories can be detected.
[0053] The attention mechanism can be a deformable cross-attention mechanism.
[0054] By adopting the above technical solution, since it is possible to extract features from the two-dimensional image of the scene, and process the extracted features based on the attention mechanism, a mask of the moving object is generated to remove the moving object, so that the outline of the moving object in the two-dimensional image of the scene can be accurately found.
[0055] In some embodiments, the two-dimensional image of the scene is a two-dimensional image of a scene whose scale is greater than a preset threshold. For example, the scene scale is not just a single object or an indoor scene, but a city-level scene. The scene modeling range of the neural radiation field in the related art is relatively small, and generally only a scene model of a scale of 10 to 100 meters can be constructed. Due to limitations of factors such as model learning ability, number of model parameters, and image quality at a longer distance, it is impossible to complete scene reconstruction for large-scale and large-range scenes (such as parking lots, large-scale scenes at the city level). In the present disclosure, in order to be able to reconstruct scenes whose scale is greater than a preset threshold, the implementation method described below is adopted.
[0056] First, the scene two-dimensional image is divided into a plurality of sub-scene two-dimensional images, and moving objects in each sub-scene two-dimensional image are removed to obtain a segmented two-dimensional image corresponding to each sub-scene two-dimensional image.
[0057] When the scene 2D image is a 2D image of a city-level scene, the scene 2D image can be divided into multiple sub-scene 2D images in the manner of block intersections. For example, each block can correspond to an area, and each area constructs a neural radiation field; different areas correspond to different neural radiation fields, and the neural radiation fields of these different areas are combined together to complete the scene modeling of the entire city.
[0058] Then, based on the segmented 2D images corresponding to each sub-scene 2D image, the sub-scenes corresponding to each sub-scene 2D image are reconstructed using the neural radiation field to obtain a 3D reconstruction of each sub-scene. After obtaining the 3D reconstruction of each sub-scene, these 3D reconstructions can be combined to obtain a 3D reconstruction of the entire scene.
[0059] For example, if a two-dimensional image of a scene is divided into M sub-scene two-dimensional images, the moving objects in each sub-scene two-dimensional image will be removed, and then the sub-scene corresponding to the first sub-scene two-dimensional image in the M sub-scene two-dimensional images will be reconstructed to obtain a first three-dimensional reconstructed scene, and the sub-scene corresponding to the second sub-scene two-dimensional image will be reconstructed to obtain a second three-dimensional reconstructed scene, until the sub-scenes corresponding to the M sub-scene two-dimensional images have completed the scene reconstruction, wherein, when reconstructing the scene, the area of the moving object in the sub-scene two-dimensional image is not considered, and only the static area in the sub-scene two-dimensional image is considered, that is, the existence of the moving object is ignored.
[0060] By adopting the above technical solution, the scene two-dimensional image is first divided into multiple sub-scene two-dimensional images and the moving objects in each sub-scene two-dimensional image are removed to obtain the segmented two-dimensional image corresponding to each sub-scene two-dimensional image, and then based on the segmented two-dimensional image corresponding to each sub-scene two-dimensional image, the sub-scenes corresponding to each sub-scene two-dimensional image are reconstructed using the neural radiation field to obtain the three-dimensional reconstructed scene of each sub-scene, so that the scene reconstruction of large-scale scenes (such as city-level scenes, city-level autonomous driving scenes, etc.) can be effectively realized. Moreover, when a certain area in a large-scale scene changes, only the image data of this area can be re-collected and the scene reconstructed without re-reconstructing the entire scene. This scene reconstruction method is suitable for neural radiation field scene reconstruction of city-level scenes, parking scenes, etc.
[0061] In some embodiments, after the scene is reconstructed, two-dimensional image rendering may be performed to obtain a two-dimensional image at a new perspective. Figure 3 FIG. 1 is a flowchart of two-dimensional image rendering according to an embodiment of the present disclosure. Figure 3 As shown, the two-dimensional image rendering process includes steps S31 to S35.
[0062] In step S31, the camera pose corresponding to the two-dimensional image to be rendered is determined.
[0063] The camera pose includes the camera's viewing angle, such as the camera's tilt and azimuth. The camera pose can also include the camera's position.
[0064] In step S32, the three-dimensional reconstructed scene involved in the camera posture is determined.
[0065] For example, a ray can be emitted from the camera under the determined camera pose, and it can be determined which three-dimensional reconstructed scenes the ray passes through. The three-dimensional reconstructed scenes passed by the ray are the three-dimensional reconstructed scenes involved in the camera pose, that is, the camera pose is within the visible range of the neural radiation field corresponding to these three-dimensional reconstructed scenes.
[0066] In step S33, the three-dimensional reconstructed scenes involved, whose distance from the position corresponding to the camera posture is greater than the preset distance, are removed. In this way, the neural radiation field that is too far away can be removed.
[0067] In step S34, each remaining 3D reconstructed scene in the involved 3D reconstructed scene is rendered respectively to obtain a rendered 2D image. The rendering may be performed in a volume rendering manner or other rendering manners.
[0068] In step S35, weighted averaging is performed on the rendered two-dimensional image to obtain a two-dimensional rendered image under the camera posture.
[0069] By adopting the above technical solution, adjacent neural radiation fields can be combined and rendered to obtain the final rendering effect, thereby efficiently obtaining a two-dimensional rendered image from a new perspective.
[0070] Figure 4 is a schematic block diagram of a scene reconstruction device according to an embodiment of the present disclosure. The scene reconstruction device is applicable to a situation where there are moving objects in the scene. Figure 4 As shown, the scene reconstruction device according to the embodiment of the present disclosure includes: an acquisition module 41, configured to acquire a two-dimensional image of the scene; a removal module 42, configured to remove moving objects in the two-dimensional image of the scene to obtain a segmented two-dimensional image; and a reconstruction module 43, configured to reconstruct the scene based on the segmented two-dimensional image using a neural radiation field.
[0071] By adopting the above technical solution, since the moving objects in the two-dimensional image of the scene are first removed to obtain a segmented two-dimensional image, and then the scene is reconstructed using the neural radiation field based on the segmented two-dimensional image, the influence of the moving objects on the scene reconstruction can be eliminated, the effect of the scene reconstruction is ensured, and realistic scene reconstruction can be achieved.
[0072] Optionally, the removal module 42 removes the moving object in the two-dimensional image of the scene, including: extracting features in the two-dimensional image of the scene; and removing the moving object by processing the features based on an attention mechanism to generate a mask of the moving object.
[0073] Optionally, the removal module 42 extracts features in the two-dimensional image of the scene, including: performing convolution processing on the two-dimensional image of the scene; and performing feature fusion on the features obtained by the convolution processing to obtain the features in the two-dimensional image of the scene.
[0074] Figure 5 FIG. 4 is a schematic structural diagram of the removal module 42 according to an embodiment of the present disclosure. Figure 5 As shown in the figure, RegNet 800 is used to perform convolution processing on the input image, BiFPN is used to perform feature fusion, deformable cross attention is used to perform attention processing on the features after feature fusion based on the cross attention mechanism, and object mask query is used to input the number of categories N of moving objects into the deformable cross attention processing. For example, if there are four categories in the scene: cars, bicycles, tricycles, and pedestrians, then N = 4. After the deformable cross attention processing, the mask of the moving object can be obtained. It should be noted that Figure 5In the figure, RegNet 800, BiFPN, and deformable cross attention are only examples, and other convolutional structures, feature fusion structures, attention mechanisms, etc. can also be used. Figure 6 FIG. 4 is a schematic diagram of the processing result of the removal module 42 according to an embodiment of the present disclosure. Figure 6 As shown, the objects indicated by numbers 1, 3, 4, 5, 7, 9, 0, etc. in the figure are moving objects.
[0075] Optionally, the scene two-dimensional image is a two-dimensional image of a scene whose scene scale is greater than a preset threshold;
[0076] The removing module 42 removes the moving objects in the scene two-dimensional image, including: dividing the scene two-dimensional image into a plurality of sub-scene two-dimensional images, and removing the moving objects in each of the sub-scene two-dimensional images to obtain a segmented two-dimensional image corresponding to each of the sub-scene two-dimensional images;
[0077] The reconstruction module 43 performs scene reconstruction based on the segmented two-dimensional image using the neural radiation field, including: based on the segmented two-dimensional image corresponding to each of the sub-scene two-dimensional images, using the neural radiation field to reconstruct the sub-scenes corresponding to each of the sub-scene two-dimensional images, to obtain a three-dimensional reconstructed scene of each of the sub-scenes.
[0078] Optionally, the scene two-dimensional image is a two-dimensional image of a scene at a city level;
[0079] The removal module 42 divides the scene two-dimensional image into a plurality of sub-scene two-dimensional images, including: dividing the scene two-dimensional image into a plurality of sub-scene two-dimensional images in the manner of street blocks and intersections.
[0080] Optionally, the scene reconstruction device according to the embodiment of the present disclosure further includes a rendering module, which is used for:
[0081] Determine the camera pose corresponding to the two-dimensional image to be rendered;
[0082] Determining the three-dimensional reconstructed scene involved in the camera posture;
[0083] Removing the three-dimensional reconstructed scenes involved, the distance between which and the position corresponding to the camera posture is greater than a preset distance;
[0084] Rendering each remaining three-dimensional reconstructed scene in the three-dimensional reconstructed scene involved respectively to obtain a rendered two-dimensional image;
[0085] The rendered two-dimensional images are weighted averaged to obtain a two-dimensional rendered image under the camera posture.
[0086] Optionally, the removal module 42 generates a mask of the moving object to remove the moving object by processing the feature based on an attention mechanism, including: generating a mask of the moving object to remove the moving object by performing deformable cross-attention processing on the feature.
[0087] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0088] The present disclosure also provides a scene reconstruction device, comprising: a processor; and a memory for storing processor executable instructions; wherein the processor is configured to execute any one of the methods described in the present disclosure.
[0089] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of any method in the present disclosure are implemented.
[0090] Figure 7 6 is a block diagram of a vehicle 600 according to an exemplary embodiment. For example, the vehicle 600 may be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0091] Reference Figure 7 , the vehicle 600 may include various subsystems, for example, an infotainment system 610, a perception system 620, a decision control system 630, a drive system 640, and a computing platform 650. The vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 may be interconnected by wire or wireless means.
[0092] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, and a navigation system, among others.
[0093] The perception system 620 may include several sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system or other positioning systems), an inertial measurement unit (IMU), a laser radar, a millimeter wave radar, an ultrasonic radar, and a camera.
[0094] The decision control system 630 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.
[0095] The drive system 640 may include components that provide powered motion for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting energy provided by the energy source into mechanical energy.
[0096] Some or all functions of the vehicle 600 are controlled by a computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652, and the processor 651 may execute instructions 653 stored in the memory 652.
[0097] The processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include a graphics processor (Graphic Process Unit, GPU), a field programmable gate array (Field Programmable Gate Array, FPGA), a system on chip (System on Chip, SOC), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC) or a combination thereof.
[0098] The memory 652 may be implemented by any type of volatile or nonvolatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0099] In addition to the instructions 653 , the memory 652 may also store data, such as road maps, route information, and data such as the location, direction, and speed of the vehicle. The data stored in the memory 652 may be used by the computing platform 650 .
[0100] In the embodiment of the present disclosure, the processor 651 may execute instruction 653 to complete all or part of the steps of the above-mentioned scene reconstruction method.
[0101] Figure 8 1 is a block diagram of a device 1900 for scene reconstruction according to an exemplary embodiment. For example, the device 1900 may be provided as a server. Figure 8, the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above-mentioned scene reconstruction method.
[0102] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958. The device 1900 may operate based on an operating system stored in the memory 1932.
[0103] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program executable by a programmable device. The computer program has a code portion for executing the above-mentioned scene reconstruction method when executed by the programmable device.
[0104] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0105] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A scene reconstruction method, characterized in that: include: Acquire a two-dimensional image of the scene; Removing moving objects from the two-dimensional image of the scene to obtain a segmented two-dimensional image; Based on the segmented two-dimensional image, a scene is reconstructed using a neural radiation field.
2. The scene reconstruction method according to claim 1, characterized in that: The removing of the moving objects in the two-dimensional image of the scene comprises: Extracting features from the two-dimensional image of the scene; The feature is processed based on an attention mechanism to generate a mask of the moving object to remove the moving object.
3. The scene reconstruction method according to claim 2, characterized in that: The extracting features from the two-dimensional image of the scene includes: Performing convolution processing on the two-dimensional image of the scene; The features obtained by the convolution processing are subjected to feature fusion to obtain the features in the two-dimensional image of the scene.
4. The scene reconstruction method according to claim 1, characterized in that: The scene two-dimensional image is a two-dimensional image of a scene whose scale is greater than a preset threshold; The removing of the moving objects in the scene two-dimensional image to obtain the segmented two-dimensional image includes: dividing the scene two-dimensional image into a plurality of sub-scene two-dimensional images, and removing the moving objects in each of the sub-scene two-dimensional images to obtain the segmented two-dimensional image corresponding to each of the sub-scene two-dimensional images; The scene reconstruction based on the segmented two-dimensional image using the neural radiation field includes: based on the segmented two-dimensional image corresponding to each of the sub-scene two-dimensional images, the scene is reconstructed by using the neural radiation field for each sub-scene corresponding to the sub-scene two-dimensional image to obtain a three-dimensional reconstructed scene of each sub-scene.
5. The scene reconstruction method according to claim 4, characterized in that: The scene two-dimensional image is a two-dimensional image of a scene at the city level; The dividing the scene two-dimensional image into a plurality of sub-scene two-dimensional images includes: dividing the scene two-dimensional image into a plurality of sub-scene two-dimensional images in the manner of block intersections.
6. The scene reconstruction method according to claim 4, characterized in that: The scene reconstruction method further comprises: after the scene reconstruction, Determine the camera pose corresponding to the two-dimensional image to be rendered; Determining the three-dimensional reconstructed scene involved in the camera posture; Removing the three-dimensional reconstructed scenes involved, the distance between which and the position corresponding to the camera posture is greater than a preset distance; Rendering each remaining three-dimensional reconstructed scene in the three-dimensional reconstructed scene involved respectively to obtain a rendered two-dimensional image; The rendered two-dimensional images are weighted averaged to obtain a two-dimensional rendered image under the camera posture.
7. The scene reconstruction method according to claim 2, characterized in that: The process of processing the feature based on an attention mechanism to generate a mask of the moving object to remove the moving object comprises: A mask of the moving object is generated by performing a deformable cross-attention process on the features to remove the moving object.
8. A scene reconstruction device, characterized in that: include: An acquisition module is configured to acquire a two-dimensional image of a scene; a removal module configured to remove moving objects in the two-dimensional image of the scene to obtain a segmented two-dimensional image; The reconstruction module is configured to reconstruct the scene based on the segmented two-dimensional image using the neural radiation field.
9. A scene reconstruction device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.