Multi-modal three-dimensional model establishment method and device, and readable medium

Through the multimodal sensor combination of IMU, monocular RGB camera and binocular event camera, the neural radiation field generation technology is used to solve the robustness of traditional three-dimensional reconstruction under low illumination and high-speed motion, and efficient and fine three-dimensional model reconstruction and moving object removal are achieved.

WO2025152900A1PCT designated stage expired Publication Date: 2025-07-24ZTE CORP

Patent Information

Application Number
PCT/CN2025/072079
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-15
Filing Date
2025-01-13
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Traditional three-dimensional reconstruction technology is not robust enough under low illumination or low texture conditions, and it is difficult to complete accurate reconstruction at high-speed motion state, especially to effectively remove moving objects in the scene.

Method used

Using a multimodal sensor combination of IMU, monocular RGB camera and binocular event camera, high frame rate and low blurred RGB images are generated through motion compensation, dynamic fuzzy correction, interpolation processing and depth estimation, and three-dimensional model reconstruction is carried out using adversarial generation neural radiation field.

Benefits of technology

High-precision three-dimensional model reconstruction is achieved under low illumination, low texture and high-speed motion conditions, effectively removing moving objects, and improving modeling efficiency and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072079_24072025_PF_FP_ABST
    Figure CN2025072079_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-modal three-dimensional model establishment method. The method comprises: acquiring IMU measurement data, a first image group collected by a binocular event camera, and a second image collected by a monocular RGB camera; calculating a second image group and a joint depth estimation image group on the basis of the IMU measurement data, the second image, and the first image group; and generating a three-dimensional model on the basis of the second image group, the joint depth estimation image group, an initial pose, and a target image. The present disclosure further provides a multi-modal three-dimensional model establishment device and a readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Multimodal three-dimensional model building method, device and readable medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application CN 202410056821.0, filed on January 15, 2024, entitled “Multimodal three-dimensional model establishment method, device and readable medium,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of computer technology, and in particular to a method, device, and readable medium for establishing a multimodal three-dimensional model. Background Art

[0004] With the continuous development of computer vision and artificial intelligence technologies, and the increasing demand for immersive visual experiences, 3D reconstruction technology has become increasingly important in the field of computer vision, and the requirements for modeling accuracy and efficiency of 3D reconstruction technology are also becoming increasingly higher. Currently, there are some 3D modeling solutions, but these 3D modeling solutions have the following problems:

[0005] 1. Insufficient robustness in low-light or low-texture conditions. Traditional reconstruction methods typically require feature matching of RGB image pairs. However, RGB cameras have a fixed dynamic range and produce poor imaging in low-light, low-texture, or scenes with large light-ratio variations. This makes it difficult to obtain accurate feature matching and, consequently, to provide an accurate 3D model.

[0006] 2. Reconstruction is difficult during high-speed motion. Traditional visual reconstruction methods rely primarily on RGB cameras, which typically shoot at a fixed and relatively low frame rate. When the camera moves at high speed, the captured RGB images inevitably produce motion blur, making accurate 3D reconstruction difficult. Summary of the Invention

[0007] The present disclosure provides a method, device, and readable medium for establishing a multimodal three-dimensional model.

[0008] An embodiment of the present disclosure provides a method for establishing a multimodal three-dimensional model, including: obtaining inertial measurement unit (IMU) measurement data, a first image group captured by a binocular event camera, and a second image captured by a monocular RGB camera; calculating a second image group and a joint depth estimation image group based on the IMU measurement data, the second image, and the first image group, wherein the second image group is obtained by performing dynamic blur correction and interpolation processing on the second image based on the smoothed first image group, the smoothed first image group is obtained by performing motion compensation processing on the first image group based on the IMU measurement data, and the joint depth estimation image group is obtained by removing moving objects from the second image and the first image group; generating a three-dimensional model based on the second image group, the joint depth estimation image group, an initial pose, and a target image, wherein the initial pose is calculated based on the IMU measurement data.

[0009] An embodiment of the present disclosure also provides a multimodal three-dimensional model building device, comprising: one or more processors; a memory on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal three-dimensional model building method according to an embodiment of the present disclosure.

[0010] An embodiment of the present disclosure further provides a computer-readable medium having a computer program stored thereon, wherein when the program is executed, the multimodal three-dimensional model building method according to the embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIG1 is a schematic diagram of a flow chart of a method for establishing a multimodal three-dimensional model according to an embodiment of the present disclosure;

[0012] FIG2 is a schematic diagram of a process of performing motion compensation processing on a first image group according to an embodiment of the present disclosure;

[0013] FIG3 is a schematic diagram of a first image group before and after motion compensation according to an embodiment of the present disclosure;

[0014] FIG4 is a schematic diagram of interpolating a second image using a first image group according to an embodiment of the present disclosure;

[0015] FIG5 is a schematic diagram of a process of performing dynamic blur correction and frame interpolation processing on a second image according to an embodiment of the present disclosure;

[0016] FIG6 is another schematic diagram of a process for performing dynamic blur correction and frame interpolation processing on a second image according to an embodiment of the present disclosure;

[0017] FIG7 is a schematic diagram of a process of performing frame interpolation processing on a second image according to an embodiment of the present disclosure;

[0018] FIG8 is a schematic diagram of a process of generating a second image group according to an extended second image group according to an embodiment of the present disclosure;

[0019] FIG9 is a schematic diagram of a process for determining a moving object mask according to an embodiment of the present disclosure;

[0020] FIG10 is a schematic diagram of a process for calculating a joint depth estimation image group according to an embodiment of the present disclosure;

[0021] FIG11 is another schematic diagram of a process for calculating a joint depth estimation image group according to an embodiment of the present disclosure;

[0022] FIG12 is a schematic diagram of generating a three-dimensional model in an adversarial generation manner according to an embodiment of the present disclosure;

[0023] FIG13 is a schematic diagram of a multimodal three-dimensional model establishment method according to a specific embodiment of the present disclosure;

[0024] FIG14 is a schematic structural diagram of a multimodal three-dimensional model building device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings, but the example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of this disclosure to those skilled in the art.

[0026] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0027] The terms used herein are used only to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements, and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof is not excluded.

[0028] The embodiments described herein may be described with reference to plan views and / or cross-sectional views, with the aid of idealized schematic diagrams of the present disclosure. Thus, the example illustrations may be modified based on manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to the embodiments shown in the accompanying drawings, but include modifications of the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the accompanying drawings are schematic in nature, and the shapes of the regions shown in the drawings illustrate specific shapes of the regions of the elements, but are not intended to be limiting.

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.

[0030] An embodiment of the present disclosure provides a method for establishing a multimodal three-dimensional model. As shown in FIG1 , the method for establishing a multimodal three-dimensional model includes the following steps S11 to S13 .

[0031] In step S11 , IMU measurement data, a first image group captured by a binocular event camera, and a second image captured by a monocular RGB camera are acquired.

[0032] The 3D model building system of the embodiment of the present disclosure includes three sensors: a monocular RGB camera, a binocular event camera, and an IMU. These sensors provide three different modal input data for the establishment of the 3D model: RGB image frames (i.e., the second image), binocular event streams (i.e., the first image group), and IMU measurement data. The binocular event camera is used to collect the binocular event stream, i.e., the first image group. The first image group includes the image E acquired by the left event camera. k-1->τ And the image E acquired by the right event camera τ->k ; Use IMU to detect the acceleration and rotation of the object; use a monocular RGB camera to collect RGB images, that is, the second image I k .

[0033] An IMU is a device that integrates multiple inertial sensors to measure an object's acceleration and angular velocity. An IMU typically consists of a three-axis accelerometer and a three-axis gyroscope. Some advanced IMUs also include a three-axis magnetometer.

[0034] An event camera, also known as an event sensor or event vision sensor, is a new type of high-speed, high-sensitivity vision sensor. Unlike RGB cameras, event cameras use an event-driven approach to record changes in light intensity, rather than continuously capturing images at a fixed frame rate. Event cameras simulate the operation of the human eye and only generate events when light intensity changes in the scene, thereby greatly reducing the demand for data transmission and processing, and having low energy consumption. When the event camera sensor senses that the change in pixel light intensity exceeds a preset threshold, an event is generated and the timestamp, pixel location, and polarity of the event (increase or decrease) are recorded. Compared to traditional cameras that capture an image at fixed intervals, event cameras can capture light intensity changes with a time resolution of microseconds, resulting in excellent performance in high-speed, high-dynamic range scenarios.

[0035] In step S12, a second image group and a joint depth estimation image group are calculated based on the IMU measurement data, the second image and the first image group, wherein the second image is subjected to dynamic blur correction and interpolation processing based on the smoothed first image group to obtain the second image group, the first image group is subjected to motion compensation processing based on the IMU measurement data to obtain the smoothed first image group, and the joint depth estimation image group is obtained by removing moving objects from the second image and the first image group.

[0036] Due to unavoidable system noise, the raw data provided by the sensor cannot be used directly for modeling and requires preprocessing. Therefore, after acquiring the IMU measurement data, the first image group, and the second image, and before building the 3D model, the acquired multimodal data needs to be preprocessed.

[0037] For the first image group, it is necessary to use the filtered IMU measurement data (i.e., acceleration and angular velocity ) is motion compensated to obtain a smooth first image group (i.e., a binocular event frame). Motion compensation technology is used to compensate for the motion of the binocular event camera and obtain a stable binocular event stream, while also making the event information captured by the binocular event camera more accurate. Motion compensation is a technology that corrects curved event streams by aligning events corresponding to the same scene edge, and is used to improve the accuracy of event-based visual mileage measurement. In the motion compensation scheme of the disclosed embodiment, IMU measurement data is used to implement rotation compensation and translation compensation to correct each original event position and improve the accuracy of event-based visual mileage calculation. In this motion compensation scheme, rotation compensation and translation compensation are implemented according to the angular velocity from the IMU sensor and the linear velocity from the back end of the system, respectively. The motion compensation scheme takes effect after initialization by estimating the speed. The motion compensation scheme is suitable for high-speed moving loads such as autonomous driving and quadrotor flight, and can accurately and reliably perform state estimation.

[0038] Within the time interval between two consecutive frames of the second image captured by the RGB camera, the particle filter data of the IMU measurement data is used to perform motion compensation on the first image group captured by the binocular event camera. Then, the optical flow map and depth map are obtained through the binocular event stream fragments to achieve motion blur removal and frame interpolation completion of the RGB image.

[0039] In the preprocessing stage, the second image and the first image group are processed to remove moving objects, that is, the moving object mask is deleted, and the depth of the first image that completes the motion blur removal, interpolation and deletion of the moving object mask is estimated, and the depth of the first image group that completes the motion compensation is calculated. The joint depth estimation image group D is obtained according to the depth estimation result and the depth calculation result.i .

[0040] In step S12, an initial pose P0 may also be calculated based on the IMU measurement data. In some embodiments, a rotation matrix and a position matrix of a target event may be calculated based on the IMU measurement data, where the target event is the event corresponding to the first image group, and the initial pose P0 may be calculated based on the rotation matrix and the position matrix of the target event.

[0041] By preprocessing the multimodal data and based on the interaction of multimodal data, it is possible to construct a high-frame-rate motion-blurred RGB image required for building a three-dimensional model and provide an initial pose P0.

[0042] In step S13, a three-dimensional model is generated according to the second image group, the joint depth estimation image group, the initial pose and the target image, wherein the initial pose is calculated according to the IMU measurement data.

[0043] The target image is a rendered RGB image. Based on the input RGB image (ie, the second image), the depth image (ie, the joint depth estimation image group D i ), the initial pose P0 and the target image, and perform multiple rounds of iterative adversarial generation to obtain the estimated pose P i , to achieve 3D reconstruction of the neural radiation field. Except for the first round of training, which requires input of an exogenous pose (i.e., the initial pose) calculated based on IMU measurement data, the poses required for subsequent iterative training are all those obtained from the previous iterative training. After multiple rounds of iterations, the final pose is obtained as the parameter of the 3D model, thereby obtaining a 3D model of the target scene.

[0044] The multimodal three-dimensional model establishment method provided by the embodiment of the present disclosure includes: obtaining IMU measurement data, a first image group captured by a binocular event camera, and a second image captured by a monocular RGB camera; calculating a second image group and a joint depth estimation image group based on the IMU measurement data, the second image, and the first image group, wherein the second image is subjected to dynamic blur correction and interpolation processing based on the smoothed first image group to obtain the second image group, the first image group is subjected to motion compensation processing based on the IMU measurement data to obtain a smoothed first image group, and the joint depth estimation image group is obtained by removing moving objects from the second image and the first image group; generating a three-dimensional model based on the second image group, the joint depth estimation image group, the initial pose, and the target image, wherein the initial pose is calculated based on the IMU measurement data. According to the embodiments of the present disclosure, monocular RGB frames, binocular event streams, and IMU measurement data are used as input data, and multimodal data is used to reconstruct a three-dimensional model. This allows for camera pose estimation and scene reconstruction under low illumination and low texture conditions, with high robustness and fine model reconstruction effects. By utilizing the high-frequency motion compensation characteristics of IMU measurement data and the high frame rate and high dynamic range characteristics of the binocular event camera, the three-dimensional model reconstruction task can be completed even in motion, thereby improving the efficiency of modeling data acquisition and the efficiency of three-dimensional model reconstruction for large-scale outdoor scenes in the overall process.

[0045] In some embodiments, after acquiring IMU measurement data (i.e., step S11), the multimodal three-dimensional model building method according to an embodiment of the present disclosure may further include: performing noise reduction and smoothing on the IMU measurement data to obtain processed IMU measurement data. Performing noise reduction and smoothing on the IMU measurement data can reduce the impact of jitter changes in the IMU measurement data on modeling accuracy.

[0046] Conventional IMU frequencies can generally reach above 300Hz, while the IMU frequency used in the embodiment of the present disclosure can be 1000Hz. IMU measurement data will produce obvious high-frequency jitter under the influence of noise, and the jitter changes of these IMU measurement data may affect the accuracy of the three-dimensional model. Therefore, it is necessary to use a filtering algorithm to reduce noise and smooth the IMU measurement data in order to obtain filtered IMU measurement data, that is, filtered acceleration. and angular velocity

[0047] According to an embodiment of the present disclosure, a particle filter can be used to preprocess IMU measurement data. A particle filter is a filter suitable for nonlinear, non-Gaussian systems. It uses a set of random sampling points (particles) to approximate the probability density function of the system, thereby estimating and predicting the system state. Particle filters are very effective for smoothing IMU measurement data in complex nonlinear systems and non-Gaussian noise.

[0048] It should be noted that the first image group can be subjected to motion compensation processing according to the processed IMU measurement data to obtain a smoothed first image group, and the initial pose P0 can be calculated according to the processed IMU measurement data.

[0049] In some embodiments, as shown in FIG2 , performing motion compensation on the first image group according to the processed IMU measurement data to obtain a smoothed first image group may include the following steps S21 to S22 .

[0050] In step S21 , a rotation matrix and a position matrix of a target event are calculated based on the processed IMU measurement data, wherein the target event is an event corresponding to the first image group.

[0051] In step S22, for each pixel of the left-eye image and the right-eye image in the first image group, the motion-compensated coordinates of the pixel are calculated based on the original coordinates of the pixel, the rotation matrix, the position matrix and the preset motion compensation hyperparameters, wherein the motion-compensated coordinates of each pixel are used to generate a smooth first image group.

[0052] Since the IMU sensor has a high frequency, at the timestamp t i and t m There is a very short time interval Δt between them, during which the motion that occurs can be regarded as uniform motion. Formula (1) defines the rotation matrix of the compensation event Formula (2) defines the position matrix of the compensation event

[0053] in, is the rotation matrix converted from Euler angles is the measurement value of the gyroscope, b g (t i ) and n g (t i) are the gyroscope bias and noise variables respectively; P represents the particle swarm optimization algorithm function; the exponential map exp represents the mapping of elements from the mathematical space called se(3) to the space called SE(3), se(3) is the Lie algebra of SE(3), SE(3) is a mathematical space that represents the set of all possible rigid body transformations in three-dimensional space. Isomorphism matrix L m From the pixel position Extension, v i Indicates that the system side is compensating timestamp t i The speed at which the

[0054] Finally, by Convert to a homogeneous matrix to obtain the compensated position.

[0055] The coordinates of each pixel after motion compensation can be expressed by the following formula (3):

[0056] The target event is the i-th event, which can be expressed as e i ={l i ,t i ,p i}, where l i ={x i ,y i} represents the target event e i The pixel coordinates, t i Indicates the timestamp, p i Indicates event polarity; the original event starts from timestamp t i to t m , the coordinates of the pixel after motion compensation are expressed as M represents the motion compensation function, and θ represents the motion compensation hyperparameter.

[0057] FIG3 is a schematic diagram of the first image group before and after motion compensation provided by an embodiment of the present disclosure. The left side of FIG3 shows the images in the original event stream captured by the binocular event camera, that is, the images in the original first image group; the right side of FIG3 shows the images in the event stream after motion compensation, that is, the images in the first image group obtained by performing motion compensation to smooth the images. By comparing area 1 and area 2, it can be seen that the images in the original event stream have the problem of image blur caused by shooting in a moving state. After motion compensation processing, the dynamic blur is removed, making the image clearer.

[0058] When an object is in high-speed motion, the binocular event camera can continuously capture binocular event stream images due to the continuous relative motion. Compared to binocular event stream images that continuously provide event information, RGB cameras, due to their fixed low frame rate, produce fewer RGB images and are prone to motion blur.

[0059] In order to solve the above problems, the embodiment of the present disclosure uses deep learning technology to perform dynamic blur correction and video frame insertion processing on RGB images based on binocular event stream images. As shown in Figure 4, at two discrete consecutive moments, i.e., the k-1th moment and the kth moment, the RGB camera can only capture two frames of RGB images, i.e., the previous frame of RGB image I k-1 and the current frame image I k , but at any time τ in the time period [k-1, k], there is a binocular event stream image output. Based on the high frame rate and high tolerance of the event camera, the RGB image can be de-motion blurred based on the continuous event information provided by the binocular event stream image to obtain a high-quality RGB image, and the RGB image can be interpolated and extended in the time period [k-1, k]. That is, according to the image E obtained by the left event camera k-1->τ And the image E acquired by the right event camera τ->k Generate and insert RGB extended image I τ , thereby enriching the contextual information of RGB images.

[0060] In some embodiments, as shown in combination with FIG. 5 and FIG. 6 , performing motion blur correction and frame interpolation processing on the second image according to the smoothed first image group to obtain the second image group includes the following steps S31 to S35 .

[0061] In step S31 , an optical flow image group is generated according to the smoothed first image group, and triangulated depth calculation is performed on the optical flow image group to obtain a first depth calculation image group.

[0062] The smoothed first image group includes the image E obtained by the left event camera after motion compensation processing k-1->τ And the image E acquired by the right event camera τ->k The smoothed first image group (E k-1->τ and E τ->k ) is input into the Multi Layer Perceptron (MLP) network and converted into an Event Spike Tensor (EST) voxel grid representation through the MLP-based kernel function to obtain a voxel grid group (V k-1->τ and V τ->k ). Voxel Grid is a spatiotemporal histogram of events, where each voxel represents a specific pixel and time interval, which can better preserve the temporal information of events.

[0063] The voxel grid group (V k-1->τ and V τ->k ) is input into the optical flow network for feature extraction, and the optical flow image group (F k-1->τ and F τ->k), and the voxel grid group (V k-1->τ and V τ->k ) to perform triangulated depth calculation and obtain the first depth calculation image group ( and ).

[0064] In step S32, the optical flow image group is used to perform interpolation processing on the second image to obtain a second image interpolation image group.

[0065] In some embodiments, the optical flow image group (F k-1->τ and F τ->k ) for the second image I k Perform interpolation processing to obtain a second image interpolation image group (ie, step S32) (I I k-1->τ and I I τ->k ) includes the following steps S321 to S323.

[0066] In step S321 , a predicted image of the second image is generated according to the optical flow image group and the second image.

[0067] The optical flow image group (F k-1->τ and F τ->k ) and the second image I acquired by the RGB camera k Combined, we can get the RGB prediction frame I at time k P k , that is, the predicted image of the second image.

[0068] In step S322, the second image is subjected to motion blur removal processing based on the predicted image and the residual network to obtain a processed second image.

[0069] The predicted image I P k and the second image I k Input the residual network and calculate the residual flow image P k , for the residual flow image P k Perform motion blur removal to remove the dynamic blur of the RGB frame image and obtain the processed second image

[0070] In step S323, the processed second image is interpolated using the optical flow image group to obtain a second image interpolation image group.

[0071] Using the optical flow image group (F k-1->τ and F τ->k ) for the processed second image Perform interpolation to complete the image and obtain the second image interpolation image group (II k-1->τ and I I τ->k ).

[0072] In step S33 , depth estimation is performed on the second interpolated image group to obtain a first depth-estimated image group.

[0073] Insert the second image into the image group (I I k-1->τ and I I τ->k ) Input the depth estimation network and calculate the first depth estimation image group ( and ).

[0074] In step S34 , an extended second image group is generated according to the first depth calculation image group, the second image interpolation image group, and the first depth estimation image group.

[0075] The first depth calculation image group ( and ), the second image interpolation image group ( and I I τ->k ) and the first depth estimation image group ( and ) are jointly aligned, and the loss is calculated by L1-Loss to obtain the extended second image group (I F k-1->τ and I F τ->k The second image group is extended to an RGB extended image group for removing motion blur, which can enrich the RGB information between timestamp k-1 and timestamp k.

[0076] In step S35, a second image group is generated based on the expanded second image group.

[0077] For the extended second group of pictures (I F k-1->τ and I F τ->k ) to perform moving object removal and background completion operations to obtain the second image group

[0078] When reconstructing 3D models of outdoor buildings, street scenes, and large scenes, there are usually moving objects in the scene, which will make the scene model reconstruction inaccurate. Traditional methods are difficult to handle this and require a lot of time and computing resources, resulting in poor final model reconstruction results. F k-1->τ and I Fτ->k ) to remove moving objects and complete background, which can solve the above problems and improve the three-dimensional model effect.

[0079] In some embodiments, as shown in FIG8 , generating the second image group according to the extended second image group (ie, step S35 ) includes the following steps S351 to S353 .

[0080] In step S351 , a moving object mask in the extended second image group is determined based on the smoothed first image group and the processed IMU measurement data.

[0081] The image of the area where the moving object mask M is located is the extended second image group (I F k-1->τ and I F τ->k ), i.e., images of moving objects. It should be noted that stationary movable objects may constitute unique scenes, such as parking lots, airports, ports, etc., so stationary movable objects do not need to be removed.

[0082] In step S352, the image of the area corresponding to the moving object mask is deleted from the expanded second image group to obtain the expanded second image group after deletion.

[0083] In step S353, the background image is filled in the area corresponding to the moving object mask in the expanded second image group after deletion to obtain a second image group.

[0084] The second group of pictures (I F k-1->τ and I F τ->k After deleting the image of the area corresponding to the moving object mask M in step S353, the background image is filled in the area. In step S353, background completion can be achieved to obtain a second image group I with the moving object removed and the background completed. i G .

[0085] In some embodiments, as shown in FIG9 , determining a moving object mask in the extended second image group (ie, step S351 ) based on the smoothed first image group and the processed IMU measurement data includes the following steps S41 to S43 .

[0086] In step S41 , semantic segmentation processing is performed on the expanded second image group to obtain a movable object mask.

[0087] After the RGB image (i.e., the second image) is subjected to motion blur correction and interpolation processing, a set of high-quality RGB extended images (i.e., the extended second image group) is obtained. The extended second image group (IF k-1->τ and I F τ->k ) to perform semantic segmentation and obtain the movable object mask M r The deep learning instance segmentation model supports 16 types of movable objects, as shown in Table 1:

[0088] Table 1

[0089] If the second group of pictures (I F k-1->τ and I F τ->k ) appears in the image of Table 1, it will be segmented into movable object masks M along the object contours. r .

[0090] In step S42 , candidate masks of moving objects in the smoothed first image group are determined based on the processed IMU measurement data.

[0091] In some embodiments, determining a candidate mask of a moving object in the smoothed first image set based on the processed IMU measurement data (i.e., step S42) includes: determining a velocity V of a stationary object in the smoothed first image set in response to a corresponding event based on the processed IMU measurement data. s and the speed V of the moving object in the smoothed first image group in response to the corresponding event d , and according to the speed V of the stationary object in the smoothed first image group responding to the corresponding event s and the speed V of the moving object in the smoothed first image group in response to the corresponding event d , determine the smoothed first image group (E k-1->τ and E τ->k ) in the moving object candidate mask M e .

[0092] When the binocular event camera itself is in motion, stationary objects in the captured image will experience relative motion at the same rate as the binocular event camera, but in the opposite direction. In this case, the moving objects in the captured image typically move at a different rate and direction than the stationary objects. The binocular event camera can record light intensity change events in real time, while the IMU can provide camera attitude information, including displacement and rotation.

[0093] The speed V of the stationary objects in the smoothed first image group responding to the corresponding event can be s and the speed V of the moving object in the smoothed first image group in response to the corresponding event d The consistency judgment result determines the smooth first image group (E k-1->τ and E τ->k) in the moving object candidate mask M e The speed V of the stationary object in the smoothed first image group in response to the corresponding event is calculated according to the following formula (4): s , calculate the speed V of the moving object in the smoothed first image group in response to the corresponding event according to the following formula (5): d : V s =-V e ±ε (4) V d ≠±V e ±γ (5)

[0094] Among them, V e is the speed of the binocular event camera, which can be calculated based on the processed IMU measurement data; ε is a constant. In the embodiment of the present disclosure, ε=0.1; γ is a constant. In the embodiment of the present disclosure, γ=0.2.

[0095] In step S43 , a moving object mask in the expanded second image group is determined based on the movable object mask and the candidate moving object masks.

[0096] In some embodiments, determining the moving object mask in the extended second image group according to the movable object mask and the candidate moving object mask (ie, step S43) includes: calculating the movable object mask M r and the candidate mask M of the moving object e When the intersection-and-union ratio is greater than a preset threshold, the movable object mask M is determined. r is the moving object mask M.

[0097] In the embodiment of the present disclosure, when the moving object candidate mask M e and movable object mask M r When the IoU (Intersection over Union) is greater than 0.85, the movable object mask M r This is the moving object mask M that needs to be removed. The movable object mask M can be calculated according to the following formula (6): r and the candidate mask M of the moving object e Intersection over Union (IoU) mask =occur(M e ,M r ) / union(M e ,M r ) (6) IoU mask >0.85

[0098] Among them, occur(M e ,M r ) is the movable object mask M r and the candidate mask M of the moving object e The intersection, union(M e ,M r ) is the movable object mask M r and the candidate mask M of the moving object e The union of .

[0099] As people's demand for immersive visual experience gradually increases, three-dimensional reconstruction technology has become increasingly important in the field of computer vision. However, under low illumination or low texture conditions, traditional three-dimensional scene reconstruction methods often cannot provide accurate reconstruction models. In addition, traditional methods find it difficult to remove moving objects in the scene, which is a huge challenge for the high-precision requirements of three-dimensional reconstruction of the scene, such as the reconstruction of three-dimensional models of large objects such as ancient buildings, urban traffic planning, or bird's-eye view scenes. To solve the above problems, according to an embodiment of the present disclosure, the second image and the first image group are processed to remove moving objects in the preprocessing stage to obtain a joint depth estimation image group.

[0100] In some embodiments, as shown in FIG10 , calculating a joint depth estimation image group according to the IMU measurement data, the second image, and the first image group (ie, step S12 ) includes the following steps S121 to S124 .

[0101] In step S121 , after the moving object mask is determined, images in the region corresponding to the moving object mask in the smoothed first image group are deleted to obtain a deleted smoothed first image group.

[0102] After determining the moving object mask M, the second image group (I F k-1->τ and I F τ->k ) and the smoothed first image group (E k-1->τ and E τ->k ) to delete the image of the area corresponding to the moving object mask M.

[0103] Step S122 , performing triangulation depth calculation on the deleted smoothed first image group to obtain a second depth calculation image group.

[0104] After removing the mask M from the binocular event image, the second depth calculation image group is obtained by triangulation calculation

[0105] In step S123 , depth estimation is performed on the deleted extended second image group to obtain a second depth-estimated image group.

[0106] The deleted extended second image group is obtained by deleting the image of the area corresponding to the moving object mask in the extended second image group. In step S123, the deleted extended second image group is input into the depth estimation network to calculate the second depth estimation image group

[0107] In step S124 , the second depth calculation image group and the second depth estimation image group are fused to obtain a joint depth estimation image group.

[0108] As shown in Figure 11, MLP is used to calculate the second depth image group and the second depth estimation image group After feature extraction and average pooling, the shared weight adaptive fusion module is input for feature fusion processing to obtain the joint depth estimation image group (D i The joint depth estimation image group (D i ):

[0109] Among them, α, β, and ρ are the hyperparameters that the adaptive fusion module needs to learn.

[0110] In some embodiments, calculating the initial pose based on the IMU measurement data may include: calculating the rotation matrix comp of the target event based on the processed IMU measurement data R and position matrix comp L , the target event is to obtain the first image group (E k-1->τ and E τ->k ) corresponding to the event; according to the rotation matrix comp of the target event R and position matrix comp L Calculate the initial pose P0.

[0111] In some embodiments, generating a three-dimensional model based on the second image group, the joint depth estimation image group, the initial pose and the target image (ie, step S13) includes: generating a three-dimensional model based on the second image group Joint depth estimation image group D i , the initial pose P0 and the target image are iteratively calculated in an adversarial generation manner to obtain the estimated pose P that meets the convergence conditions i , estimated pose P i are the parameters of the three-dimensional model.

[0112] Estimated pose P i The displacement parameter T i and the rotation parameter Q i Indicates. irepresents the three-dimensional translation vector, Q i Represents a rotation vector in quaternion form,

[0113] In some embodiments, according to the second image group Joint depth estimation image group D i , the initial pose P0 and the target image are iteratively calculated in an adversarial generation manner to obtain the estimated pose P that meets the convergence conditions i Including: for one round of iterative calculation, calculate the predicted second image group of this round according to the target image The predicted joint depth estimation image group D′ in this round i And the predicted pose P′ of this round i ; According to the second image group of this round And the second image group of this round of prediction Calculate the first adversarial residual of this round According to the joint depth estimation image group D of this round i And the predicted joint depth estimation image group D′ of this round i Calculate the second adversarial residual ΔD of this round i , according to the estimated pose P of this round i And the predicted pose P′ of this round i Calculate the third adversarial residual ΔP of this round i , where the estimated pose of the first round is calculated based on the initial pose P0; in this round, the first adversarial residual The second adversarial residual ΔD of this round i And the third adversarial residual ΔP of this round i When all convergence conditions are met, the estimated pose of this round is determined as the parameters of the three-dimensional model.

[0114] FIG12 is a schematic diagram of generating a three-dimensional model in an adversarial generation manner according to an embodiment of the present disclosure. As shown in FIG12 , the embodiment of the present disclosure is based on the neural radiation field and combines the idea of ​​adversarial generation to use the input second image group to generate a three-dimensional model. Joint depth estimation image group D i and the predicted pose P i Perform three-dimensional reconstruction iterative modeling. In each round of three-dimensional model training, the second image group is input Joint depth estimation image group D i and the predicted pose P i As training data, a 3D reconstruction model is updated, and then the updated 3D reconstruction model is reverse rendered and the corresponding predicted second image group of the target scene in the current iteration is calculated. Predict the joint depth estimation image group D′i and predicted pose P′ i According to the second image group and predict the second group of images Calculate the first adversarial residual Estimating the image set D based on the joint depth i and predict the joint depth estimation image group D′ i Calculate the second adversarial residual ΔD i , according to the estimated pose P i and predicted pose P′ i Calculate the third adversarial residual ΔP i L1-Loss is used to measure these three sets of residuals. In each round of 3D reconstruction model training, the three sets of residuals are minimized to gradually complete the update of the 3D reconstruction model.

[0115] The disclosed embodiments provide an adversarial generative 3D model reconstruction process, which only requires the input of an exogenously calculated initial pose P0 in the first round of training iterations. In subsequent iterative training, the corresponding estimated pose can be calculated based on the residual between the 2D image projection prediction of the 3D model and the real input 2D image through an adversarial generative approach, and the estimated pose is continuously improved during the iterative training.

[0116] In some embodiments, after calculating the third adversarial residual of this round based on the estimated pose of this round and the predicted pose of this round, the multimodal three-dimensional model establishment method according to the embodiment of the present disclosure may also include: when the third adversarial residual is less than the first preset threshold and the next round of iterative calculation is performed, calculating the estimated pose of the next round of iteration based on the estimated pose of this round; when the third adversarial residual is less than the second preset threshold, updating the IMU measurement data according to the estimated pose of this round, wherein the second preset threshold is less than the first preset threshold.

[0117] At the i-th iteration, the third adversarial residual ΔP i When the value is less than the first preset threshold σ1 and the next round of iterative calculation is performed, the estimated pose P of the input of this round of iteration is updated. i , and combine the IMU measurement data to predict the pose and estimate the pose P of the next iteration i+1 , which is used as the estimated pose input for the i+1th iteration.

[0118] In the third adversarial residual ΔP i If the value is less than the second preset threshold σ2, the current IMU measurement data is updated, that is, the IMU measurement data is reinitialized to reduce the pose error caused by the accumulated error of the IMU measurement data. The second preset threshold σ2 is less than the first preset threshold σ1. In the disclosed embodiment, σ1 = 2.2 and σ2 = 1.2.

[0119] The disclosed embodiment reconstructs a three-dimensional model based on the neural radiation field. During the training process of the end-to-end deep learning model, training and iteration are performed in an adversarial generation manner. The intermediate results of the three-dimensional model obtained through training can be used to correct the IMU and reduce the impact of the IMU's accumulated errors.

[0120] To clearly illustrate the solution of the disclosed embodiment, a specific example is provided below with reference to FIG13 . In this example, a 3D model is used to construct simulated 3D street scene data, digital city landscapes, and traffic road planning for autonomous driving system training. Before commencing 3D model reconstruction, a multimodal sensor assembly consisting of a monocular RGB camera, a binocular event camera, and an IMU sensor is mounted on the top of a vehicle. A large-scale 3D model of the urban traffic street scene is reconstructed as if the vehicle were driving through the city. The lenses of the monocular RGB camera and binocular event camera face the front of the vehicle and align in the same direction. The vehicle's forward direction of travel is defined as the positive z-axis, and the coordinate system of the multimodal sensor assembly conforms to the left-hand rule. Before officially commencing vehicle-based 3D model reconstruction of the urban scene, the modeling area must be determined and the vehicle's route must be planned. The route must form a closed loop in a clockwise or counterclockwise direction, i.e., it must start from and return to the origin, and the vehicle must not backtrack during travel.

[0121] Starting from the starting point of the planned path, multimodal data is collected and the 3D model is reconstructed online. Multimodal data collection is performed at any time, including RGB images, binocular event streams, and IMU observations. Multimodal data is saved at every time stamp, and online 3D model reconstruction is performed every 2 seconds.

[0122] As shown in FIG13 , the multimodal 3D model building process includes the following steps.

[0123] Step 1: Multimodal data preprocessing

[0124] (1) The particle swarm optimization algorithm is used to reduce noise and filter the IMU observation values ​​to obtain smooth IMU observation values.

[0125] (2) Since the vehicle is in motion most of the time, the binocular event camera outputs a continuous binocular event stream. The motion characteristics of the high-frequency IMU can be used to perform motion compensation on the output of the binocular event camera to further stabilize and smooth the binocular event stream.

[0126] (3) The binocular event flow finally outputs a high-frame-rate optical flow map through a deep learning algorithm. Based on the smooth binocular event flow and the optical flow map, the low-frame-rate RGB image is subjected to dynamic blur correction and high-frame-rate RGB video interpolation is completed to obtain an RGB extended image group.

[0127] Step 2: Moving object removal and background completion

[0128] A moving object is one of the 16 movable objects defined in Table 1 above and is in motion. On the one hand, the candidate mask M of the moving object is determined based on the IMU observations and the smoothed binocular event stream. e On the other hand, the 16 movable objects in the RGB image are segmented to obtain the movable object mask M r . For the candidate mask M of the moving object e and movable object mask M r The intersection-and-union ratio is calculated to obtain a moving object mask M that satisfies both the motion state and one of the above 16 types of objects.

[0129] The object corresponding to the moving object mask M is segmented, and the background corresponding to the moving object mask M is generated using a deep learning generative network to achieve the removal of movable objects and the completion of the background of the corresponding area. Finally, an RGB image group with moving objects removed and background completed is obtained. and the binocular event stream with moving objects removed.

[0130] Step 3: Depth Estimation and Pose Estimation

[0131] The second depth calculation image group can be calculated based on the binocular event stream after the moving objects are removed. According to the RGB image group of moving object removal and background completion, a second depth estimation image group can be obtained by deep learning depth estimation algorithm according to and The joint depth estimation image group D is calculated using the adaptive joint depth estimation algorithm i .

[0132] Combining the IMU observations with the particle swarm optimization algorithm not only can the motion characteristics of the IMU be used to compensate for the motion of other sensors, but also an initial pose P0 can be calculated, which is used for the next step of adversarial generative neural radiation field training.

[0133] Step 4: Training the adversarial generative neural radiation field model

[0134] The RGB image group obtained by the above steps after moving object removal and background completion Rendered RGB image and joint depth estimation image group D i, and the initial pose P0 calculated based on the IMU observation data, can be used for three-dimensional reconstruction training of the neural radiation field. When the first training iteration is performed, the initial pose P0 calculated using the IMU observation data is used. In subsequent training iterations, the previous round of training iteration provides the estimated pose P for the next round of training. i The adversarial generative neural radiation field model proposed in this system can generate a three-dimensional model and estimate the pose. It can reconstruct the three-dimensional model by estimating the pose, and calculate the pose by projecting the three-dimensional model, and perform pairwise adversarial operations to achieve the goal of simultaneously optimizing the pose and the three-dimensional model itself.

[0135] The above steps complete online 3D scene modeling and save the multimodal data for subsequent refinement. After the online 3D model is completed—that is, the conversion of multimodal street view data into a 3D model is complete—the reconstructed 3D model can be rendered into a video, allowing for secondary 3D design and viewing on-screen via a 3D viewer for user interaction and browsing.

[0136] The disclosed embodiment utilizes multimodal data to reconstruct a three-dimensional model, which has the characteristics of high robustness and fine reconstruction effect. Traditional three-dimensional model reconstruction and rendering require the use of high-end equipment and complex software, and often require professional three-dimensional modeling designers and a lot of time costs. The disclosed embodiment uses monocular RGB images, binocular event streams and IMU sensor measurements as data input, making full use of the high-frequency motion compensation characteristics of the IMU and the high frame rate and high dynamic range characteristics of the binocular event camera. It can complete the three-dimensional model reconstruction task even in motion, improve the efficiency of modeling data acquisition, and greatly improve the efficiency of three-dimensional reconstruction of large-scale outdoor scenes in the overall process.

[0137] The disclosed embodiment adopts an end-to-end deep learning algorithm for image processing, which can improve the accuracy and efficiency of three-dimensional model reconstruction. By adopting an end-to-end deep learning algorithm, in addition to the two I / O operations of input / output, the entire calculation can be performed on the image processor (Graphics Processing Unit, GPU), realizing the functions of RGB image deblurring, RGB image high frame rate, moving target removal and background completion in the image, and reconstruction of a three-dimensional model of adversarial generative neural radiation field. Compared with traditional computer vision algorithms, it can realize three-dimensional reconstruction of large-scale outdoor scenes under the condition of high-speed movement of multimodal data acquisition equipment, and compared with the pipeline processing flow of traditional deep learning, it greatly reduces I / O operations and improves the efficiency of three-dimensional reconstruction.

[0138] The disclosed embodiments utilize an adversarial generative neural radiation field model, requiring only the input of exogenously calculated multi-sensor pose information during the first round of training. In subsequent rounds of iterative training, pose estimation and 3D scene reconstruction can be achieved in an adversarial generative manner, avoiding the pose calculation process and further improving the efficiency of 3D model reconstruction. This independence from the exogenous pose calculation process enables the multimodal 3D modeling solution proposed in the disclosed embodiments to achieve online reconstruction, i.e., simultaneous data acquisition and 3D model reconstruction.

[0139] The disclosed embodiments can be applied to the reconstruction of three-dimensional models of large-scale outdoor scenes and urban scenes, and can be fixedly mounted on high-speed moving vehicles such as vehicles and drones. By utilizing multimodal data, high-frame-rate de-motion blurred RGB images are achieved, and these RGB images remove moving objects in the scene. Only one initial pose is required, and the processed RGB images can be used to perform adversarial generative neural radiation field training. In each round of training iteration, a three-dimensional model is generated and the pose information required for the next round of iteration is output, thereby realizing online training of the three-dimensional model and ultimately achieving the generation of the three-dimensional model.

[0140] The disclosed embodiments utilize a monocular RGB camera, a binocular event camera, and an IMU sensor to achieve high-precision three-dimensional model reconstruction of the target scene even in motion. They are capable of performing camera pose estimation and scene three-dimensional model reconstruction under low illumination and low-texture conditions, and can effectively remove moving objects in the target scene while achieving high model reconstruction efficiency.

[0141] As shown in Figure 14, an embodiment of the present disclosure also provides a multimodal three-dimensional model building device, including: at least one processor 1401; a memory 1402, on which at least one program is stored. When the at least one program is executed by the at least one processor 1401, the at least one processor 1401 implements the multimodal three-dimensional model building method according to each embodiment of the present disclosure.

[0142] The multimodal three-dimensional model building apparatus according to an embodiment of the present disclosure may further include at least one I / O interface 1403 connected between the processor 1401 and the memory 1402 and configured to implement information interaction between the processor 1401 and the memory 1402 .

[0143] The processor 1401 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 1402 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) 1403 is connected between the processor 1401 and the memory 1402, and can realize information interaction between the processor 1401 and the memory 1402, including but not limited to a data bus (Bus), etc.

[0144] In some embodiments, the processor 1401 , the memory 1402 , and the I / O interface 1403 are connected to each other via a bus, and further connected to other components of the computing device.

[0145] The embodiments of the present disclosure further provide a computer-readable medium having a computer program stored thereon, which, when executed, implements the multimodal three-dimensional model building method according to each embodiment of the present disclosure.

[0146] It will be appreciated by those skilled in the art that all or some of the steps in the method disclosed above, and the functional modules / units in the device can be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0147] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A method for establishing a multi-modal three-dimensional model, comprising: Obtaining inertial measurement unit (IMU) measurement data, a first image group collected by a binocular event camera, and a second image collected by a monocular RGB camera; Calculating a second image group and a joint depth estimation image group based on the IMU measurement data, the second image, and the first image group, wherein the second image group is obtained by performing dynamic blur correction processing and frame interpolation processing on the second image according to the smoothed first image group, the smoothed first image group is obtained by performing motion compensation processing on the first image group according to the IMU measurement data, and the joint depth estimation image group is obtained by performing moving object removal processing on the second image and the first image group; Generating a three-dimensional model based on the second image group, the joint depth estimation image group, an initial pose, and a target image, wherein the initial pose is calculated according to the IMU measurement data.

2. The method according to claim 1, further comprising: Performing noise reduction processing and smoothing processing on the IMU measurement data to obtain processed IMU measurement data; Performing motion compensation processing on the first image group according to the processed IMU measurement data to obtain the smoothed first image group.

3. The method according to claim 2, wherein Performing motion compensation processing on the first image group according to the processed IMU measurement data to obtain the smoothed first image group includes: Calculating a rotation matrix and a position matrix of a target event according to the processed IMU measurement data, wherein the target event is the event corresponding to obtaining the first image group; For each pixel of the left-eye image and the right-eye image in the first image group, calculating the motion-compensated coordinates of the pixel according to the original coordinates of the pixel, the rotation matrix, the position matrix, and a preset motion compensation hyperparameter, wherein the motion-compensated coordinates of the pixel are used to generate the smoothed first image group.

4. The method according to claim 2, wherein Performing dynamic blur correction processing and frame interpolation processing on the second image according to the smoothed first image group to obtain the second image group includes: Generating an optical flow image group according to the smoothed first image group, performing triangulation depth calculation on the optical flow image group to obtain a first depth calculation image group; Performing frame interpolation processing on the second image by using the optical flow image group to obtain a second image interpolation image group; Performing depth estimation on the second image interpolation image group to obtain a first depth estimation image group; Generating an extended second image group according to the first depth calculation image group, the second image interpolation image group, and the first depth estimation image group; Generating the second image group according to the extended second image group.

5. The method according to claim 4, wherein Performing frame interpolation processing on the second image by using the optical flow image group to obtain a second image interpolation image group includes: Generating a predicted image of the second image according to the optical flow image group and the second image; Performing dynamic blur removal processing on the second image according to the predicted image and a residual network to obtain a processed second image; Performing frame interpolation processing on the processed second image by using the optical flow image group to obtain the second image interpolation image group.

6. The method according to claim 4, wherein Generating the second image group according to the extended second image group includes: Determining a moving object mask in the extended second image group according to the smoothed first image group and the processed IMU measurement data; Deleting the images in the corresponding regions of the moving object mask in the extended second image group to obtain a deleted extended second image group; Filling the background image in the corresponding regions of the moving object mask in the deleted extended second image group to obtain the second image group.

7. The method according to claim 6, wherein, Determining a moving object mask in the extended second image group according to the smoothed first image group and the processed IMU measurement data includes: Performing semantic segmentation processing on the extended second image group to obtain a movable object mask; Determining a moving object candidate mask in the smoothed first image group according to the processed IMU measurement data; Determining a moving object mask in the extended second image group according to the movable object mask and the moving object candidate mask.

8. The method according to claim 7, wherein, Determining a moving object candidate mask in the smoothed first image group according to the processed IMU measurement data includes: Determining the speed of the stationary objects in the smoothed first image group in response to corresponding events and the speed of the moving objects in the smoothed first image group in response to corresponding events according to the processed IMU measurement data; Determining a moving object candidate mask in the smoothed first image group according to the speed of the stationary objects in the smoothed first image group in response to corresponding events and the speed of the moving objects in the smoothed first image group in response to corresponding events.

9. The method according to claim 7, wherein, Determining a moving object mask in the extended second image group according to the movable object mask and the moving object candidate mask includes: Calculating the intersection over union of the movable object mask and the moving object candidate mask; When the intersection over union is greater than a preset threshold, determining the movable object mask as the moving object mask in the extended second image group.

10. The method according to claim 6, wherein, Calculating a joint depth estimation image group according to the IMU measurement data, the second image and the first image group includes: After determining the moving object mask, deleting the images in the corresponding regions of the moving object mask in the smoothed first image group to obtain a deleted smoothed first image group; Performing triangulation depth calculation on the deleted smoothed first image group to obtain a second depth calculation image group; Performing depth estimation on the deleted extended second image group to obtain a second depth estimation image group; Performing fusion processing on the second depth calculation image group and the second depth estimation image group to obtain a joint depth estimation image group.

11. The method according to any one of claims 1-10, wherein, Generating a three-dimensional model according to the second image group, the joint depth estimation image group, the initial pose and the target image includes: Performing iterative calculation in an adversarial generation manner according to the second image group, the joint depth estimation image group, the initial pose and the target image to obtain an estimated pose that meets the convergence condition, where the estimated pose is a parameter of the three-dimensional model.

12. The method according to claim 11, wherein, Performing iterative calculations in an adversarial generation manner based on the second image group, the combined depth estimation image group, the initial pose, and the target image to obtain an estimated pose that satisfies the convergence condition includes: For one round of iterative calculation, calculating the predicted second image group of this round, the predicted combined depth estimation image group of this round, and the predicted pose of this round respectively according to the target image; Calculating the first adversarial residual of this round according to the second image group of this round and the predicted second image group of this round, calculating the second adversarial residual of this round according to the combined depth estimation image group of this round and the predicted combined depth estimation image group of this round, and calculating the third adversarial residual of this round according to the estimated pose of this round and the predicted pose of this round, wherein the estimated pose of the first round is calculated according to the initial pose; When the first adversarial residual of this round, the second adversarial residual of this round, and the third adversarial residual of this round all satisfy the convergence condition, determining the estimated pose of this round as the parameters of the three-dimensional model.

13. The method according to claim 12, wherein, After calculating the third adversarial residual of this round according to the estimated pose of this round and the predicted pose of this round, the method further includes: When the third adversarial residual is less than the first preset threshold and the next round of iterative calculation is performed, calculating the estimated pose of the next round of iteration according to the estimated pose of this round; When the third adversarial residual is less than the second preset threshold, updating the IMU measurement data according to the estimated pose of this round, wherein the second preset threshold is less than the first preset threshold.

14. A multi-modal three-dimensional model establishment device, comprising: One or more processors; A memory having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the multi-modal three-dimensional model establishment method according to any one of claims 1-13.

15. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed, implementing the multi-modal three-dimensional model establishment method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Three-dimensional scene reconstruction method and device and storage medium

    CN109978931A

  • Three-dimensional attitude estimation method and device based on multi-sensor fusion

    CN111354043A

  • Three-dimensional reconstruction method and device, equipment, storage medium and program product

    CN115294280A

  • Visual positioning method and device based on event camera, electronic equipment and medium

    CN117036462A

  • Method of Depth Estimation Using a Camera and Inertial Sensor

    US20180075614A1

Cited By

  • Open field crop growth monitoring method and device

    CN121415086A

  • Unmanned aerial vehicle image-assisted path tracking and obstacle avoidance method and system

    CN121523367A