Multi-modal three-dimensional model establishing method and device and readable medium
The multi-modal 3D model reconstruction method using IMU, stereo event, and RGB cameras addresses low-light and high-speed challenges by enhancing robustness and efficiency in 3D model generation.
Patent Information
- Application Number
- CN202410056821.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional three-dimensional reconstruction technology is not robust enough under low illumination or low texture conditions, and it is difficult to complete accurate reconstruction at high speed motion conditions, especially when using RGB cameras, which are prone to motion blur.
The multimodal sensor of the inertial measurement unit IMU, binocular event camera and monocular RGB camera are used to generate high frame rate and low blurred RGB images through motion compensation, dynamic fuzzy correction and interpolation processing, combined with deep learning technology, and perform moving objects removal and background completion, and use adversarial generation of neural radiation fields for three-dimensional model reconstruction.
High robustness and high precision three-dimensional model reconstruction under low illumination, low texture and high speed motion conditions are achieved, which improves modeling data acquisition efficiency and is suitable for 3D model reconstruction in large-scale outdoor scenes.
Smart Images

Figure CN120318406A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, an apparatus, and a readable medium for establishing a multi-modal three-dimensional model. Background Art
[0002] With the continuous development of computer vision technologies and artificial intelligence technologies, and the gradually increasing demand for immersive visual experiences from people, three-dimensional reconstruction technologies have become increasingly important in the field of computer vision, and the requirements for the modeling accuracy and modeling efficiency of three-dimensional reconstruction technologies are also getting higher and higher. Currently, there are already some three-dimensional model establishment solutions, but these three-dimensional model establishments have the following problems:
[0003] 1. Insufficient robustness under low illumination or low texture conditions. Traditional reconstruction methods usually require feature matching of RGB image pairs, and RGB cameras have a fixed dynamic range. They perform poorly in scenes with low illumination, low texture, or large light ratio changes, and it is difficult to obtain accurate feature matching, and thus an accurate three-dimensional model cannot be provided.
[0004] 2. It is difficult to complete reconstruction under high-speed movement. Traditional visual reconstruction methods mainly rely on RGB cameras, and RGB cameras usually shoot at a fixed frame rate and the frame rate is relatively low. When the camera is moving at high speed, the RGB images collected inevitably have motion blur, and it is difficult to complete accurate three-dimensional reconstruction. Summary of the Invention
[0005] The present disclosure provides a method, an apparatus, and a readable medium for establishing a multi-modal three-dimensional model.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for establishing a multi-modal three-dimensional model, including:
[0007] Obtaining inertial measurement unit (IMU) measurement data, a first image group collected by a binocular event camera, and a second image collected by a monocular RGB camera;
[0008] Calculating a second image group and a joint depth estimation image group according to the IMU measurement data, the second image, and the first image group; wherein, the second image group is obtained by performing dynamic blur correction processing and frame interpolation processing on the second image according to a smoothed first image group, and the smoothed first image group is obtained by performing motion compensation processing on the first image group according to the IMU measurement data; the joint depth estimation image group is obtained by performing moving object removal processing on the second image and the first image group;
[0009] Generating a three-dimensional model according to the second image group, the joint depth estimation image group, an initial pose, and a target image, where the initial pose is calculated according to the IMU measurement data.
[0010] In another aspect, an embodiment of the present disclosure further provides a multi-modal three-dimensional model building device, including: one or more processors; a memory storing one or more programs thereon; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the multi-modal three-dimensional model building method as described above; at least one I / O interface connected between the processor and the memory and configured to implement information interaction between the processor and the memory.
[0011] In another aspect, an embodiment of the present disclosure further provides a computer-readable medium storing a computer program thereon, wherein when the program is executed, the multi-modal three-dimensional model building method as described above is implemented.
[0012] The multi-modal three-dimensional model building method provided by the embodiment of the present disclosure includes: acquiring inertial measurement unit (IMU) measurement data, a first image group collected by a binocular event camera, and a second image collected by a monocular RGB camera; calculating a second image group and a joint depth estimation image group according to the IMU measurement data, the second image, and the first image group; wherein the second image group is obtained by performing dynamic blur correction processing and interpolation processing on the second image according to the smoothed first image group, and the smoothed first image group is obtained by performing motion compensation processing on the first image group according to the IMU measurement data; generating a three-dimensional model according to the second image group, the joint depth estimation image group, the initial pose, and the target image, and the initial pose is calculated according to the IMU measurement data; the joint depth estimation image group is obtained by performing moving object removal processing on the second image and the first image group; the embodiment of the present disclosure takes a monocular RGB frame, a binocular event stream, and IMU measurement data as data inputs, and uses multi-modal data for three-dimensional model reconstruction, capable of performing camera pose estimation and scene reconstruction under low illumination and low texture conditions, with high robustness and fine model reconstruction effect; by using the high-frequency motion compensation characteristic of the IMU measurement data and the high frame rate and high dynamic range characteristics of the binocular event camera, the three-dimensional model reconstruction task can also be completed in a moving situation, improving the modeling data acquisition efficiency and overall process efficiency for three-dimensional model reconstruction of large-scale outdoor scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic flowchart of the multi-modal three-dimensional model building method provided by the embodiment of the present disclosure;
[0014] Figure 2 It is a schematic flowchart of performing motion compensation processing on the first image group provided by the embodiment of the present disclosure;
[0015] Figure 3 It is a schematic diagram of the first image group before and after motion compensation provided by the embodiment of the present disclosure;
[0016] Figure 4Schematic diagram of using a first image to team up with a second image for frame interpolation provided by an embodiment of the present disclosure;
[0017] Figure 5 Flow schematic of performing dynamic blur correction and frame interpolation processing on a second image provided by an embodiment of the present disclosure Figure 1 ;
[0018] Figure 6 Flow schematic of performing dynamic blur correction and frame interpolation processing on a second image provided by an embodiment of the present disclosure Figure 2 ;
[0019] Figure 7 Flow schematic diagram of performing frame interpolation processing on a second image provided by an embodiment of the present disclosure;
[0020] Figure 8 Flow schematic diagram of generating a second image group according to an extended second image group provided by an embodiment of the present disclosure;
[0021] Figure 9 Flow schematic diagram of determining a moving object mask provided by an embodiment of the present disclosure;
[0022] Figure 10 Flow schematic of calculating a joint depth estimation image group provided by an embodiment of the present disclosure Figure 1 ;
[0023] Figure 11 Flow schematic of calculating a joint depth estimation image group provided by an embodiment of the present disclosure Figure 2 ;
[0024] Figure 12 Schematic diagram of generating a three-dimensional model in an adversarial generation manner provided by an embodiment of the present disclosure;
[0025] Figure 13 Flow schematic diagram of a method for establishing a multi-modal three-dimensional model provided by a specific example of the present disclosure;
[0026] Figure 14 Schematic diagram of the structure of a multi-modal three-dimensional model establishment device provided by an embodiment of the present disclosure. Detailed implementation manners
[0027] Hereinafter, example embodiments will be described more fully with reference to the accompanying drawings, but the example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art.
[0028] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0029] The terms used in this disclosure are only for describing specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "consisting of" are used in this specification, the specified features, wholes, steps, operations, elements, and / or components are present, but one or more other features, wholes, steps, operations, elements, components, and / or their groups are not excluded.
[0030] The embodiments described herein may be described with reference to plan views and / or cross-sectional views by means of ideal schematic diagrams of the present disclosure. Accordingly, the example illustrations may be modified according to manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to the embodiments shown in the drawings, but include modifications of configurations formed based on manufacturing processes. Therefore, the regions illustrated in the drawings have schematic properties, and the shapes of the regions shown in the drawings illustrate the specific shapes of the regions of the elements, but are not intended to be restrictive.
[0031] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.
[0032] Embodiments of the present disclosure provide a multimodal three-dimensional model building method, as Figure 1 shown, the multimodal three-dimensional model building method includes the following steps:
[0033] Step S11, obtaining IMU measurement data, a first image group collected by a binocular event camera, and a second image collected by a monocular RGB camera.
[0034] The three-dimensional model building system of the embodiments of the present disclosure consists of three major sensors: a monocular RGB camera, a binocular event camera, and an IMU (Inertial Measurement Unit), and these sensors provide three different modalities of input data for three-dimensional model building: RGB image frames (i.e., the second image), binocular event streams (i.e., the first image group), and IMU measurement data. Among them, the binocular event camera is used to collect binocular event streams, that is, the first image group, and the first image group includes an image E k-1->τ obtained by the left-eye event camera and an image E τ->k obtained by the right-eye event camera; the IMU is used to detect the acceleration and rotational motion of an object; the monocular RGB camera is used to collect RGB images, that is, the second image I k .
[0035] An IMU is a device that integrates multiple inertial sensors and is used to measure the acceleration and angular velocity of an object. An IMU usually consists of a three-axis accelerometer and a three-axis gyroscope. Some advanced IMUs also include a three-axis magnetometer.
[0036] An event camera, also known as an event sensor or event vision sensor, is a new type of high-speed and high-sensitivity vision sensor. Different from RGB cameras, event cameras record the changes in light intensity in an event-driven manner, rather than continuously capturing images at a fixed frame rate. It mimics the operation of the human eye and generates events only when there are changes in light intensity in the scene, thus greatly reducing the requirements for data transmission and processing and having a low energy consumption overhead. When the event camera sensor senses that the change in pixel light intensity exceeds a preset threshold, it generates an event that includes the timestamp of the event occurrence, the pixel position, and the polarity of the event (increase or decrease). Compared with traditional cameras that capture an image every fixed time, event cameras can capture changes in light intensity with a time resolution at the microsecond level, making them perform excellently in high-speed and high-dynamic range scenes.
[0037] Step S12: Calculate the second image group and the joint depth estimation image group based on the IMU measurement data, the second image, and the first image group; wherein, the second image group is obtained by performing dynamic blur correction processing and interpolation processing on the second image according to the smoothed first image group, and the smoothed first image group is obtained by performing motion compensation processing on the first image group according to the IMU measurement data; the joint depth estimation image group is obtained by performing moving object removal processing on the second image and the first image group.
[0038] Due to the inevitable system noise, the raw data provided by the sensor cannot be directly used for modeling and needs to be preprocessed first. Therefore, after obtaining the IMU measurement data, the first image group, and the second image and before establishing the 3D model, the obtained multi-modal data is preprocessed.
[0039] For the first image group, it is necessary to use the filtered IMU measurement data, i.e., the acceleration and the angular velocity Motion compensation is performed to obtain a smooth first image group (i.e., binocular event frames). The motion compensation technique is used to compensate for the motion of the binocular event camera and obtain a stable binocular event stream, and at the same time, it also makes the event information captured by the binocular event camera more accurate. Motion compensation is a technique for correcting a curved event stream by aligning events corresponding to the edges of the same scene, and this technique is used to improve the accuracy of event-based visual odometry. The motion compensation scheme of the present disclosure uses IMU measurement data to achieve rotation compensation and translation compensation, which are used to correct the position of each original event and improve the accuracy of event-based visual odometers. The motion compensation scheme respectively achieves rotation compensation and translation compensation according to the angular velocity from the IMU sensor and the linear velocity from the system backend. By estimating the velocity, the motion compensation scheme takes effect after initialization. The motion compensation scheme is applicable to high-speed moving vehicles such as autonomous driving and quadcopter flight, and can accurately and reliably perform state estimation.
[0040] Within the time interval between two consecutive second images collected by the RGB camera, motion compensation is performed on the first image group collected by the binocular event camera using the particle filter data of the IMU measurement data. Then, an optical flow map and a depth map are obtained through the binocular event stream segment, realizing the removal of motion blur processing and frame interpolation completion for the RGB image.
[0041] In the preprocessing stage, motion object removal processing is also respectively performed on the second image and the first image group, that is, the motion object mask is deleted. Depth estimation is performed on the first image that has completed motion blur removal processing, frame interpolation completion, and motion object mask deletion, and depth calculation is performed on the first image group that has completed motion compensation. According to the depth estimation result and the depth calculation result, a joint depth estimation image group D is obtained. i 。
[0042] In this step, the initial pose P0 can also be calculated according to the IMU measurement data. In some embodiments, the rotation matrix and position matrix of the target event can be calculated according to the IMU measurement data. The target event is the event corresponding to obtaining the first image group, and the initial pose P0 is calculated according to the rotation matrix and position matrix of the target event.
[0043] By preprocessing the multi-modal data and based on multi-modal data interaction, a high-frame-rate motion-blur-removed RGB image required for constructing a 3D model can be built and the initial pose P0 can be provided.
[0044] Step S13, generate a 3D model according to the second image group, the joint depth estimation image group, the initial pose, and the target image, where the initial pose is calculated according to the IMU measurement data.
[0045] The target image is a rendered RGB image. Based on the input RGB image (i.e., the second image), the depth image (i.e., the joint depth estimation image group Di ) By using the adversarial generation method with multiple rounds of iteration based on the initial pose P0 and the target image, the estimated pose P is obtained i to achieve the three-dimensional reconstruction of the neural radiance field. Except for the first round of training which requires the input of the exogenous pose (i.e., the initial pose) calculated based on the IMU measurement data, the poses required for subsequent iterative training are all obtained from the poses obtained in the previous iterative training. In this way, after multiple rounds of iteration, the final pose is obtained as the parameter of the three-dimensional model, thereby obtaining the three-dimensional model of the target scene.
[0046] The multi-modal three-dimensional model establishment method provided by the embodiments of the present disclosure includes: obtaining inertial measurement unit (IMU) measurement data, the first image group collected by a binocular event camera, and the second image collected by a monocular RGB camera; calculating the second image group and the joint depth estimation image group according to the IMU measurement data, the second image, and the first image group; wherein, the second image group is obtained by performing dynamic blur correction processing and interpolation processing on the second image according to the smoothed first image group, and the smoothed first image group is obtained by performing motion compensation processing on the first image group according to the IMU measurement data; generating a three-dimensional model according to the second image group, the joint depth estimation image group, the initial pose, and the target image, and the initial pose is calculated according to the IMU measurement data; the embodiments of the present disclosure use the monocular RGB frame, the binocular event stream, and the IMU measurement data as data inputs, and use multi-modal data for three-dimensional model reconstruction, which can perform camera pose estimation and scene reconstruction under low illumination and low texture conditions, with high robustness and fine model reconstruction effect; by using the high-frequency motion compensation characteristic of the IMU measurement data and the high frame rate and high dynamic range characteristics of the binocular event camera, the three-dimensional model reconstruction task can also be completed in a moving situation, improving the efficiency of modeling data acquisition and improving the three-dimensional model reconstruction efficiency of large-scale outdoor scenes in the overall process.
[0047] In some embodiments, after obtaining the IMU measurement data (i.e., step S11), the multi-modal three-dimensional model establishment method may further include the following steps: performing noise reduction processing and smoothing processing on the IMU measurement data to obtain the processed IMU measurement data. Performing noise reduction processing and smoothing processing on the IMU measurement data can reduce the influence of the jitter change of the IMU measurement data on the modeling accuracy.
[0048] The conventional IMU frequency can generally reach more than 300Hz, and the IMU frequency used in the embodiments of the present disclosure is 1000Hz. The IMU measurement data will produce obvious high-frequency jitter under the influence of noise, and these jitter changes of the IMU measurement data may affect the accuracy of the three-dimensional model. Therefore, it is necessary to use a filtering algorithm to perform noise reduction and smoothing on the IMU measurement data in order to obtain the filtered IMU measurement data, that is, the filtered acceleration and angular velocity In the embodiments of the present disclosure, a particle filter is used to preprocess IMU measurement data. The particle filter is a filter applicable to non-linear and non-Gaussian systems. It approximates the probability density function of the system by using a set of randomly sampled points (particles), thereby realizing the estimation and prediction of the system state. The particle filter is very effective for smoothing IMU measurement data in the case of complex non-linear systems and non-Gaussian noise.
[0049] It should be noted that the smoothed first image group can be obtained by performing motion compensation processing on the first image group according to the processed IMU measurement data, and the initial pose P0 is calculated according to the processed IMU measurement data.
[0050] In some embodiments, as Figure 2 shown, obtaining the smoothed first image group by performing motion compensation processing on the first image group according to the processed IMU measurement data may include the following steps:
[0051] Step S21: Calculate the rotation matrix and position matrix of the target event according to the processed IMU measurement data, where the target event is the event of obtaining the first image group.
[0052] Step S22: For each pixel of the left-eye image and the right-eye image in the first image group, calculate the coordinates of the pixel after motion compensation according to the original coordinates of the pixel, the rotation matrix, the position matrix, and a preset motion compensation hyperparameter. The coordinates of each pixel after motion compensation are used to generate the smoothed first image group.
[0053] Since the IMU sensor has a high frequency, there is a very short time interval Δt between timestamps t i and t m . During this time interval, the motion that occurs can be regarded as uniform motion. Formula (1) defines the rotation matrix of the compensation event Formula (2) defines the position matrix of the compensation event
[0054]
[0055]
[0056] Among them, is the rotation matrix converted from Euler angles is the measurement value of the gyroscope, b g (t i ) and n g (t iare the bias and noise variables of the gyroscope respectively; P represents the particle swarm optimization algorithm function; the exponential mapping exp represents mapping an element from a mathematical space named se(3) to a space named SE(3), where se(3) is the Lie algebra of SE(3), and SE(3) is a mathematical space representing the set of all possible rigid body transformations in three-dimensional space. The isomorphism matrix L m is from the pixel position extended, v i represents the velocity of the system side at the compensation timestamp t i .
[0057] Finally, the compensated position is obtained by converting to a homogeneous matrix.
[0058] The coordinates of each pixel after motion compensation can be expressed by the following formula (3):
[0059]
[0060] The target event is the i-th event, which can be expressed as e i ={l i ,t i ,p i}, where l i ={x i ,y i} represents the pixel coordinates of the target event e i , t i represents the timestamp, and p i represents the event polarity; the original event is from the timestamp t i to t m , and the coordinates of the pixel after motion compensation are expressed as M represents the motion compensation function, and θ represents the motion compensation hyperparameter.
[0061] Figure 3 is a schematic diagram of the first image group before and after motion compensation provided by the embodiment of the present disclosure. Figure 3 On the left is the image in the original event stream obtained by the binocular event camera, that is, the image in the original first image group; Figure 3 On the right is the image in the event stream after motion compensation, that is, the image in the smoothed first image group obtained by performing motion compensation processing. By comparing Region 1 and Region 2, it can be seen that the image in the original event stream has the problem of image blurring caused by shooting in a moving state. After motion compensation processing, the dynamic blur is removed, making the image clearer.
[0062] When the object is in a high-speed motion state, due to the continuous occurrence of relative motion, the binocular event camera can continuously capture binocular event stream images. Compared with the binocular event stream images that continuously provide event information, due to its fixed low frame rate, the RGB camera has fewer RGB image frames and is prone to motion blur. To solve the above problems, embodiments of the present disclosure utilize deep learning technology to perform dynamic blur correction and video frame interpolation processing on RGB images based on binocular event stream images. As Figure 4 shown, at two discrete consecutive times, namely the (k - 1)-th time and the k-th time, the RGB camera can only capture 2 frames of RGB images, that is, the previous frame of RGB image I k-1 and the current frame image I k , but at any τ time in the time period [k - 1, k], there are binocular event stream images output. Based on the high frame rate and high dynamic range of the event camera, motion blur can be removed from the RGB image based on the continuous event information provided by the binocular event stream images to obtain high-quality RGB images, and in the time period [k - 1, k], the RGB images are interpolated and extended, that is, according to the image E k-1->τ obtained by the left-eye event camera and the image E τ->k obtained by the right-eye event camera, generate and insert the RGB extended image I τ , thereby enriching the context information of the RGB images.
[0063] In some embodiments, as shown in Figure 5 and Figure 6 , performing dynamic blur correction processing and frame interpolation processing on the second image according to the smoothed first image group to obtain the second image group includes the following steps:
[0064] Step S31, generating an optical flow image group according to the smoothed first image group, and performing triangulation depth calculation on the optical flow image group to obtain a first depth calculation image group.
[0065] The smoothed first image group includes the image E k-1->τ obtained by the left-eye event camera and the image E τ->k obtained by the right-eye event camera after motion compensation processing. The smoothed first image group (E k-1->τ and E τ->k ) after motion compensation processing is input into the MLP (Multi-Layer Perceptron) network and converted into an EST (Event Spiking Tensor) voxel grid representation through a kernel function based on MLP to obtain a voxel grid group (V k-1->τ and V τ->k ). A voxel grid is a spatio-temporal histogram of events, where each voxel represents a specific pixel and time interval, indicating that it can better retain the time information of events.
[0066] The voxel grid group (Vk-1->τ and V τ->k ) The input optical flow network is used for feature extraction to obtain a group of optical flow images (F k-1->τ and F τ->k ), and triangulation depth calculation is performed on the voxel grid group (V k-1->τ and V τ->k ) to obtain a first group of depth calculation images ( and ).
[0067] Step S32: Use the group of optical flow images to perform frame interpolation on the second image to obtain a group of second image interpolated images.
[0068] In some embodiments, as shown in combination with Figure 6 and Figure 7 , the use of the group of optical flow images (F k-1->τ and F τ->k ) to perform frame interpolation on the second image I k to obtain a group of second image interpolated images (i.e., step S32) (I I k-1->τ and I I τ->k ) includes the following steps:
[0069] Step S321: Generate a predicted image of the second image according to the group of optical flow images and the second image.
[0070] Combining the group of optical flow images (F k-1->τ and F τ->k ) with the second image I k acquired by the RGB camera, the RGB predicted frame I P k at time k can be obtained, that is, the predicted image of the second image.
[0071] Step S322: Remove the dynamic blur of the second image according to the predicted image and the residual network to obtain the processed second image.
[0072] Input the predicted image I P k and the second image I k into the residual network, calculate to obtain the residual flow image P k , perform dynamic blur removal processing on the residual flow image P k to realize dynamic blur removal of the RGB frame image, and obtain the processed second image
[0073] Step S323: Use the group of optical flow images to perform frame interpolation on the processed second image to obtain a group of second image interpolated images.
[0074] Interpolate and complete the processed second image using an optical flow image group (F k-1->τ and F τ->k ) to obtain an interpolated image group of the second image (I I k-1->τ and II τ->k ).
[0075] Step S33: Perform depth estimation on the interpolated image group of the second image to obtain a first depth estimation image group.
[0076] Input the interpolated image group of the second image (I I k-1->τ I and I τ->k I ) into a depth estimation network, and calculate to obtain a first depth estimation image group ( and ).
[0077] Step S34: Generate an extended second image group based on the first depth calculation image group, the interpolated image group of the second image, and the first depth estimation image group.
[0078] Jointly align the first depth calculation image group ( and ), the interpolated image group of the second image (I I k-1->τ and I I τ->k ), and the first depth estimation image group ( and ), calculate the loss via L1-Loss, and obtain an extended second image group (I F k-1 -> τ and I F τ->k ). This extended second image group is an RGB extended image group that removes motion blur and can enrich the RGB information between timestamp k-1 and timestamp k.
[0079] Step S35: Generate a second image group based on the extended second image group.
[0080] Perform operations to remove moving objects and complete the background on the extended second image group (I F k-1->τ and I F τ->k ) to obtain a second image group
[0081] When performing 3D model reconstruction on outdoor buildings, street scenes, and large-scale scenes, there are usually moving objects in the scene. The existence of these moving objects will make the scene model reconstruction inaccurate. Traditional methods are difficult to handle and require a large amount of time and computing resources, resulting in poor final model reconstruction effects.
[0082] In the embodiments of the present disclosure, by performing moving object removal and background completion operations on the extended second image group (I F k-1->τ and I F τ->k ), the above problems can be solved and the 3D model effect can be improved.
[0083] In some embodiments, as Figure 8 shown, generating the second image group according to the extended second image group (i.e., step S35) includes the following steps:
[0084] Step S351, determining the moving object mask in the extended second image group according to the smoothed first image group and the processed IMU measurement data.
[0085] The image in the area where the moving object mask M is located is the image to be removed in the extended second image group (I F k-1->τ and I F τ->k ), that is, the image of the moving object. It should be noted that movable objects in a stationary state may constitute unique scenes, such as parking lots, airports, ports, etc. Therefore, movable objects in a stationary state do not need to be removed.
[0086] Step S352, deleting the image in the area corresponding to the moving object mask in the extended second image group to obtain the extended second image group after deletion, and filling the background image in the area corresponding to the moving object mask in the extended second image group after deletion to obtain the second image group.
[0087] After deleting the image in the area corresponding to the moving object mask M in the extended second image group (I F k-1->τ and I F τ->k ), filling the background image in this area. In this step, background completion can be achieved to obtain the second image group with the moving object removed and the background completed.
[0088] In some embodiments, as Figure 9 shown, determining the moving object mask in the extended second image group according to the smoothed first image group and the processed IMU measurement data (i.e., step S351) includes the following steps:
[0089] Step S41: Perform semantic segmentation on the extended second image group to obtain a mask for movable objects.
[0090] After the RGB image (i.e., the second image) undergoes dynamic blur correction and frame interpolation, a set of high-quality RGB extended images (i.e., the extended second image group) is obtained. Using a deep learning instance segmentation model, the extended second image group (I F k-1->τ and I F τ->k ) is semantically segmented to obtain a mask for movable objects M r . The deep learning instance segmentation model supports 16 types of movable objects, as shown in Table 1:
[0091] Table 1
[0092] Chinese Name English Name Classification ID Pedestrian Person 1 Bicycle Rider Rider 2 Sedan Car 3 Truck Truck 4 Bus Bus 5 Recreational Vehicle Caravan 6 Trailer Trailer 7 Sports Utility Vehicle SUV 8 Motorcycle Motorcycle 9 Bicycle Bicycle 10 Train Train 11 Airplane Plane 12 Ship Boat 13 Cat Cat 14 Dog Dog 15 Bird Bird 16
[0093] If the 16 types of movable objects in Table 1 appear in the images of the extended second image group (I F k-1->τ and I F τ->k ), they will be segmented along the object contours into a mask for movable objects M r .
[0094] Step S42: Determine a candidate mask for moving objects in the smoothed first image group based on the processed IMU measurement data.
[0095] In some embodiments, the determining a candidate mask for moving objects in the smoothed first image group based on the processed IMU measurement data includes the following steps: Determine the speed V s of the stationary objects in the smoothed first image group in response to corresponding events and the speed V d of the moving objects in the smoothed first image group in response to corresponding events, and based on the speed V s of the stationary objects in the smoothed first image group in response to corresponding events and the speed V d of the moving objects in the smoothed first image group in response to corresponding events, determine a candidate mask M k-1->τ and E τ->k for moving objects in the smoothed first image group (E e .
[0096] When the binocular event camera is in a moving state itself, a relative motion will be generated for the stationary objects in the captured image, and the motion rate is the same as the motion rate of the binocular event camera, and the motion speed is opposite to that of the binocular event camera. At this time, the moving objects in the captured image are usually different from the stationary objects in terms of motion rate and direction. The binocular event camera can record the light intensity change events in real time, and the IMU can provide the attitude information of the camera, including displacement and rotation.
[0097] The velocity V of the stationary objects in the smoothed first image group responding to the corresponding events can be obtained s and the velocity V of the moving objects in the smoothed first image group responding to the corresponding events d to determine the moving object candidate mask M in the smoothed first image group (E k-1->τ and E τ->k ) according to the consistency judgment result. e Among them, the velocity V of the stationary objects in the smoothed first image group responding to the corresponding events is calculated according to the following formula (4) s , and the velocity V of the moving objects in the smoothed first image group responding to the corresponding events is calculated according to the following formula (5) d :
[0098] V s =-V e ±ε (4)
[0099] V d ≠±V e ±γ (5)
[0100] Among them, V e is the velocity of the binocular event camera and can be calculated based on the processed IMU measurement data; ε is a constant, and in the embodiments of the present disclosure, ε = 0.1; γ is a constant, and in the embodiments of the present disclosure, γ = 0.2.
[0101] Step S43, determining the moving object mask in the extended second image group according to the movable object mask and the moving object candidate mask.
[0102] In some embodiments, the determining the moving object mask according to the movable object mask and the moving object candidate mask includes the following steps: calculating the intersection over union of the movable object mask M r and the moving object candidate mask M e ; in the case where the intersection over union is greater than a preset threshold, determining the movable object mask M r as the moving object mask M.
[0103] In the embodiments of the present disclosure, when the moving object candidate mask M e and the movable object mask M rWhen the IoU (Intersection over Union) is greater than 0.85, the movable object mask M r is the movable object mask M that needs to be removed. The movable object mask M can be calculated according to the following formula (6) r and the intersection over union of the moving object candidate mask M e is as follows:
[0104] IoU mask = occur(M e , M r ) / union(M e , M r ) (6)
[0105] IoU mask > 0.85
[0106] Among them, occur(M e , M r ) is the intersection of the movable object mask M r and the moving object candidate mask M e , and union(M e , M r ) is the union of the movable object mask M r and the moving object candidate mask M e .
[0107] As people's demand for immersive visual experience gradually increases, three-dimensional reconstruction technology has become increasingly important in the field of computer vision. However, under low illumination or low texture conditions, traditional three-dimensional scene reconstruction methods often cannot provide accurate reconstruction models. In addition, traditional methods are difficult to remove moving objects in the scene, which poses a huge challenge to the high-precision requirements of scene three-dimensional reconstruction, such as the three-dimensional model reconstruction of large objects or bird's-eye view scenes such as ancient buildings and urban traffic planning. The embodiments of the present disclosure solve the above problems by performing moving object removal processing on the second image and the first image group in the preprocessing stage to obtain a combined depth estimation image group.
[0108] In some embodiments, as Figure 10 shown, calculating the combined depth estimation image group according to the IMU measurement data, the second image, and the first image group (i.e., step S12) includes the following steps:
[0109] Step S121, after determining the moving object mask, delete the images in the corresponding area of the moving object mask in the smoothed first image group to obtain the smoothed first image group after deletion.
[0110] After determining the moving object mask M, respectively in the extended second image group (I Fk-1->τ and I F τ->k ) and the smooth first image group (E k-1->τ and E τ->k ) delete the images in the regions corresponding to the moving object mask M in the smooth first image group.
[0111] Step S122, perform triangulation depth calculation on the deleted smooth first image group to obtain a second depth calculation image group.
[0112] After removing the mask M in the binocular event image, a second depth calculation image group is obtained through triangulation calculation
[0113] Step S123, perform depth estimation on the deleted extended second image group to obtain a second depth estimation image group.
[0114] The deleted extended second image group is obtained by deleting the images in the regions corresponding to the moving object mask in the extended second image group. In this step, the deleted extended second image group is input into a depth estimation network to calculate and obtain a second depth estimation image group
[0115] Step S124, perform fusion processing on the second depth calculation image group and the second depth estimation image group to obtain a joint depth estimation image group.
[0116] As Figure 11 shown, use MLP to respectively perform feature extraction on the second depth calculation image group and the second depth estimation image group , and after performing average pooling respectively, send them into an adaptive fusion module with shared weights for feature fusion processing to obtain a joint depth estimation image group (D i ). The joint depth estimation image group (D i ) can be calculated according to the following formula (7):
[0117]
[0118] where α, β, ρ are hyperparameters to be learned by the adaptive fusion module.
[0119] In some embodiments, the steps of calculating the initial pose according to the IMU measurement data may include: calculating the rotation matrix comp of the target event according to the processed IMU measurement data R and the position matrix comp L , the target event is the event corresponding to obtaining the first image group (E k-1->τ and E τ->k ); according to the rotation matrix comp of the target event Rand the position matrix comp L Calculate the initial pose P0.
[0120] In some embodiments, the generating of the three-dimensional model according to the second image group, the combined depth estimation image group, the initial pose, and the target image (i.e., step S13) includes the following steps: According to the second image group the combined depth estimation image group D i , the initial pose P0, and the target image, perform iterative calculations in an adversarial generation manner to obtain an estimated pose P that satisfies the convergence condition i , where the estimated pose P i is the parameter of the three-dimensional model.
[0121] The estimated pose P i can be represented by a displacement parameter T i and a rotation parameter Q i . T i represents a three-dimensional translation vector, Q i represents a rotation vector in quaternion form,
[0122] In some embodiments, the generating of the three-dimensional model according to the second image group the combined depth estimation image group D i , the initial pose P0, and the target image, perform iterative calculations in an adversarial generation manner to obtain an estimated pose P that satisfies the convergence condition i , including the following steps: For one round of iterative calculations, calculate the predicted second image group for this round according to the target image the predicted combined depth estimation image group D′ for this round i and the predicted pose P′ for this round i ; Calculate the first adversarial residual for this round according to the second image group for this round and the predicted second image group for this round ; Calculate the second adversarial residual ΔD for this round according to the combined depth estimation image group D for this round i and the predicted combined depth estimation image group D′ for this round i ; Calculate the third adversarial residual ΔP for this round according to the estimated pose P for this round i i and the predicted pose P′ for this round i i i i ; where the estimated pose for the first round is calculated according to the initial pose P0; When the first adversarial residual for this round the second adversarial residual ΔD for this round i and the third adversarial residual ΔP for this round i all satisfy the convergence condition, determine the estimated pose for this round as the parameter of the three-dimensional model.
[0123] Figure 12 Schematic diagram of generating a 3D model in an adversarial generation manner provided by an embodiment of the present disclosure. As Figure 12 shown, the embodiment of the present disclosure is based on a neural radiance field, combines the idea of adversarial generation, and uses the input second image group to jointly estimate the depth image group D i and the predicted pose P i to perform 3D reconstruction iterative modeling. In the process of each round of 3D model training, the input second image group is combined with the depth estimation image group D i and the predicted pose P i as training data, and a 3D reconstruction modeling update will be performed once. Subsequently, the updated model of 3D reconstruction will perform inverse rendering and calculate the corresponding predicted second image group of the target scene in the current iteration predicted joint depth estimation image group D' i and the predicted pose P' i . According to the second image group and the predicted second image group , the first adversarial residual is calculated . According to the joint depth estimation image group D i and the predicted joint depth estimation image group D' i , the second adversarial residual ΔD is calculated i . According to the estimated pose P i and the predicted pose p' i , the third adversarial residual ΔP is calculated i . These three groups of residuals are measured using L1-Loss. During one round of 3D reconstruction model training, minimizing these three groups of residuals can gradually complete the update of the 3D reconstruction model.
[0124] The embodiment of the present disclosure is an adversarial generation-based 3D model reconstruction process. Only an initially calculated external pose P0 needs to be input in the first round of training iteration. In subsequent iterative training, the corresponding estimated pose can be calculated through the adversarial generation method according to the residual between the 2D image projection prediction of the 3D model and the real input 2D image, and this estimated pose is continuously improved during the iterative training.
[0125] In some embodiments, after calculating the third adversarial residual of the current round based on the estimated pose and predicted pose of the current round, the multi-modal three-dimensional model building method may further include the following steps: when the third adversarial residual is less than the first preset threshold and the next round of iterative calculation is performed, calculating the estimated pose of the next round of iteration according to the estimated pose of the current round; when the third adversarial residual is less than the second preset threshold, updating the IMU measurement data according to the estimated pose of the current round, where the second preset threshold is less than the first preset threshold.
[0126] At the i-th round of iteration, when the third adversarial residual ΔP i is less than the first preset threshold σ1 and the next round of iterative calculation is performed, the estimated pose P input for the current round of iteration will be updated i , and pose prediction will be performed in combination with the IMU measurement data to estimate the estimated pose P i+1 of the next round of iteration, and use it as the estimated pose input for the (i + 1)-th round of iteration.
[0127] When the third adversarial residual ΔP i is less than the second preset threshold σ2, the current IMU measurement data will be updated, that is, the IMU measurement data will be re-initialized to reduce the pose error caused by the cumulative error of the IMU measurement data. The second preset threshold σ2 is less than the first preset threshold σ1. In the embodiments of the present disclosure, σ1 = 2.2 and σ2 = 1.2.
[0128] The embodiments of the present disclosure perform three-dimensional model reconstruction based on the neural radiance field. During the training process of the end-to-end deep learning model, adversarial generation is used for training and iterative refinement. The intermediate results of the trained three-dimensional model can correct the IMU and reduce the influence brought by the IMU cumulative error.
[0129] To clearly illustrate the solution of the embodiments of the present disclosure, the following is combined with Figure 13 a specific example for illustration. In this example, the three-dimensional model is used to build simulation three-dimensional street scene data, digital city landscapes, traffic road planning, etc. for the training of the autonomous driving system. Before starting the three-dimensional model reconstruction, a multi-modal sensor assembly composed of a monocular RGB camera, a binocular event camera, and an IMU sensor needs to be fixed on the top of the vehicle, and in the form of the vehicle driving in the city, large-scale three-dimensional model reconstruction of the urban traffic street scene is carried out. The lenses of the monocular RGB camera and the binocular event camera face directly in front of the vehicle and have the same direction. It is stipulated that the positive direction of the vehicle's forward driving is the positive direction of the z-axis, and the coordinate system of the multi-modal sensor group conforms to the left-hand rule. Before officially starting the three-dimensional model reconstruction of the urban scene based on the vehicle, it is necessary to determine the modeling area and plan the vehicle driving route. The driving route needs to form a closed loop in the clockwise or counterclockwise direction, that is, start from the origin and finally return to the origin, and the vehicle does not go back during the driving process.
[0130] Starting from the starting point of the planned path, multi-modal data collection and online 3D model reconstruction are carried out. Multi-modal data collection is carried out at all times, and the collected multi-modal data includes: RGB images, binocular event streams, and IMU observations. The multi-modal data at each timestamp will be saved; online 3D model reconstruction is carried out every 2 seconds.
[0131] As Figure 13 shown, the process of establishing a multi-modal 3D model includes the following steps:
[0132] 1. Multi-modal data preprocessing
[0133] (1) Use the particle swarm optimization algorithm to denoise and filter the IMU observations to obtain smooth IMU observations;
[0134] (2) Since the vehicle is in a driving state most of the time, the binocular event camera outputs a continuous binocular event stream. At this time, use the motion characteristics of the high-frequency IMU to perform motion compensation on the output of the binocular event camera to further stabilize and smooth the binocular event stream;
[0135] (3) The binocular event stream finally outputs a high-frame-rate optical flow map through a deep learning algorithm. The low-frame-rate RGB images are dynamically blurred and corrected and high-frame-rate RGB video interpolation is completed according to the smooth binocular event stream and the optical flow map to obtain an RGB extended image group.
[0136] 2. Moving object removal and background completion
[0137] The moving object is one of the 16 moving objects defined in Table 1 and is in a moving state. On the one hand, determine the moving object candidate mask M according to the IMU observations and the smooth binocular event stream e ; on the other hand, perform instance segmentation on the 16 movable objects in the RGB image to obtain the movable object mask M r . Calculate the intersection over union of the moving object candidate mask M e and the movable object mask M r to obtain the moving object mask M that satisfies both the moving state and one of the above 16 objects.
[0138] Segment the object corresponding to the moving object mask M, and use the deep learning generation network to generate the background corresponding to the moving object mask M to realize the removal of the movable object and the background completion of the corresponding area. Finally, obtain the RGB image group after moving object removal and background completion and the binocular event stream from which the moving objects have been removed.
[0139] 3. Depth estimation and pose estimation
[0140] The second depth calculation image group can be calculated based on the binocular event stream after the moving object is removed. The second depth estimation image group can be obtained from the RGB image group after the moving object is removed and the background is completed through a deep learning depth estimation algorithm. According to and The joint depth estimation image group D is calculated using the adaptive joint depth estimation algorithm. i 。
[0141] The IMU observations combined with the particle swarm optimization algorithm can not only perform motion compensation on other sensors using its motion characteristics, but also calculate an initial pose P0, which is used for the adversarial generation neural radiance field training in the next step.
[0142] 4. Adversarial Generative Neural Radiance Field Model Training
[0143] Using the RGB image group after the moving object is removed and the background is completed obtained from the previous steps The rendered RGB images and the joint depth estimation image group D i 、and the initial pose P0 calculated based on the IMU observation data, the three-dimensional reconstruction training of the neural radiance field can be carried out. When the first training iteration is performed, the initial pose P0 calculated using the IMU observation data is used, while in subsequent training iterations, the previous training iteration will provide the estimated pose P for the next round of training. i 。The adversarial generative neural radiance field model proposed in this system can generate a three-dimensional model and an estimated pose, can reconstruct the three-dimensional model through the estimated pose, and at the same time calculate the pose through the projection of the three-dimensional model, and the two are in confrontation with each other to achieve the purpose of optimizing both the pose and the three-dimensional model itself.
[0144] Through the above steps, online three-dimensional scene modeling can be completed, and at the same time, multi-modal data can be saved for subsequent refined processing. When the online three-dimensional model is completed, the conversion of multi-modal street view data into a three-dimensional model is completed. The reconstructed three-dimensional model can be rendered into a video, can be used for secondary three-dimensional design, or can be viewed on the screen through a three-dimensional viewer, providing interaction and browsing for users.
[0145] The embodiments of the present disclosure use multi-modal data for 3D model reconstruction, which are characterized by high robustness and fine reconstruction effect. Traditional 3D model reconstruction and rendering require the use of high-end devices and complex software, and often require professional 3D modeling designers and a large amount of time cost. However, the embodiments of the present disclosure use monocular RGB images, binocular event streams, and IMU sensor measurements as data inputs, making full use of the high-frequency motion compensation characteristics of the IMU and the high frame rate and high dynamic range of the binocular event camera, and can also complete the 3D model reconstruction task in a moving situation, improving the efficiency of modeling data acquisition and greatly improving the 3D reconstruction efficiency of large-scale outdoor scenes in the overall process.
[0146] The embodiments of the present disclosure use an end-to-end deep learning algorithm for image processing, which can improve the accuracy and efficiency of 3D model reconstruction. By using the end-to-end deep learning algorithm, except for the two I / O operations at the input / output, the calculation can be performed entirely on the GPU (Graphics Processing Unit), realizing the functions of RGB image deblurring, high frame rate of RGB images, removal of moving objects in the image and background completion, and adversarial generative neural radiance field 3D model reconstruction. Compared with traditional computer vision algorithms, it can achieve 3D reconstruction of large-scale outdoor scenes in the case of high-speed movement of multi-modal data acquisition devices, and compared with the pipeline processing process of traditional deep learning, it greatly reduces the I / O operations and improves the efficiency of 3D reconstruction.
[0147] The embodiments of the present disclosure use an adversarial generative neural radiance field model, which only needs to input the multi-sensor pose information calculated by external sources in the first round of training. In subsequent multi-round iterative training, pose estimation and 3D scene reconstruction can be achieved in an adversarial generative manner, avoiding the pose calculation process and further improving the 3D model reconstruction efficiency. And the process that does not rely on external pose calculation enables the multi-modal 3D model establishment scheme proposed in the embodiments of the present disclosure to achieve online reconstruction, that is, data collection and 3D model reconstruction at the same time.
[0148] The embodiments of the present disclosure can be applied to 3D model reconstruction of large-scale outdoor scenes and urban scenes, and can be fixedly mounted on high-speed moving vehicles such as vehicles and drones. Using multi-modal data, high-frame-rate deblurred RGB images are realized, and these RGB images remove the moving objects in the scene. Only one initial pose is required, and adversarial generative neural radiance field training can be performed using the processed RGB images. In each round of training iteration, a 3D model is generated and the pose information required for the next round of iteration is output, realizing online training of the 3D model and finally realizing the generation of the 3D model.
[0149] Embodiments of the present disclosure utilize a monocular RGB camera, a binocular event camera, and an IMU sensor to achieve high-precision three-dimensional model reconstruction of a target scene even in a moving state. The advantages are that it can perform camera pose estimation and three-dimensional model reconstruction of the scene under low illuminance and low texture conditions, can effectively remove moving objects in the target scene, and has a relatively high model reconstruction efficiency.
[0150] As Figure 14 shown, embodiments of the present disclosure further provide a multimodal three-dimensional model building device, and the multimodal three-dimensional model building device includes:
[0151] At least one processor 1401;
[0152] A memory 1402, on which at least one program is stored. When the at least one program is executed by the at least one processor, the at least one processor implements the multimodal three-dimensional model building method as described above;
[0153] At least one I / O interface 1403, connected between the processor and the memory, and configured to implement information interaction between the processor and the memory.
[0154] Among them, the processor 1401 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory 1402 is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) 1403 is connected between the processor 1401 and the memory 1402 and can implement information interaction between the processor 901 and the memory 902, including but not limited to a data bus (Bus), etc.
[0155] In some embodiments, the processor 1401, the memory 1402, and the I / O interface 1403 are interconnected through a bus and further connected to other components of the computing device.
[0156] Embodiments of the present disclosure further provide a computer-readable medium, on which a computer program is stored. When the computer program is executed, it implements the multimodal three-dimensional model building method provided in the foregoing embodiments.
[0157] It will be appreciated by those skilled in the art that all or some of the steps in the method disclosed above, the functional modules / units in the device may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0158] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for limiting purposes. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.
Claims
1. A method for establishing a multi-modal three-dimensional model, characterized in that, Including: Obtaining inertial measurement unit (IMU) measurement data, a first image group collected by a binocular event camera, and a second image collected by a monocular RGB camera; Calculating a second image group and a joint depth estimation image group based on the IMU measurement data, the second image, and the first image group; wherein, the second image group is obtained by performing dynamic blur correction processing and interpolation processing on the second image according to the smoothed first image group, and the smoothed first image group is obtained by performing motion compensation processing on the first image group according to the IMU measurement data; the joint depth estimation image group is obtained by performing moving object removal processing on the second image and the first image group; Generating a three-dimensional model based on the second image group, the joint depth estimation image group, an initial pose, and a target image, and the initial pose is calculated based on the IMU measurement data.
2. The method according to claim 1, characterized in that, The method further includes: performing noise reduction processing and smoothing processing on the IMU measurement data to obtain processed IMU measurement data; The smoothed first image group is obtained by performing motion compensation processing on the first image group according to the processed IMU measurement data.
3. The method according to claim 2, wherein Obtaining the smoothed first image group by performing motion compensation processing on the first image group according to the processed IMU measurement data includes: Calculating a rotation matrix and a position matrix of a target event according to the processed IMU measurement data, and the target event is the event of obtaining the first image group; For each pixel of the left-eye image and the right-eye image in the first image group, according to the original coordinates of the pixel, the rotation matrix, the position matrix, and a preset motion compensation hyperparameter, calculating the coordinates of the pixel after motion compensation, and the coordinates of each pixel after motion compensation are used to generate the smoothed first image group.
4. The method according to claim 2, wherein Obtaining the second image group by performing dynamic blur correction processing and interpolation processing on the second image according to the smoothed first image group includes: Generating an optical flow image group according to the smoothed first image group, performing triangulation depth calculation on the optical flow image group to obtain a first depth calculation image group; Performing interpolation processing on the second image by using the optical flow image group to obtain a second image interpolation image group; Performing depth estimation on the second image interpolation image group to obtain a first depth estimation image group; Generating an extended second image group according to the first depth calculation image group, the second image interpolation image group, and the first depth estimation image group; Generating the second image group according to the extended second image group.
5. The method according to claim 4, characterized in that, Performing interpolation processing on the second image by using the optical flow image group to obtain a second image interpolation image group includes: Generating a predicted image of the second image according to the optical flow image group and the second image; Performing dynamic blur removal processing on the second image according to the predicted image and a residual network to obtain a processed second image; Performing interpolation processing on the processed second image by using the optical flow image group to obtain the second image interpolation image group.
6. The method according to claim 4, wherein Generating the second image group according to the extended second image group includes: Determine the moving object mask in the extended second image group according to the smoothed first image group and the processed IMU measurement data; Delete the images in the corresponding areas of the moving object mask in the extended second image group to obtain the extended second image group after deletion, and fill the background image in the areas corresponding to the moving object mask in the extended second image group after deletion to obtain the second image group.
7. The method according to claim 6, wherein The determining the moving object mask in the extended second image group according to the smoothed first image group and the processed IMU measurement data includes: Perform semantic segmentation processing on the extended second image group to obtain a movable object mask; Determine the moving object candidate mask in the smoothed first image group according to the processed IMU measurement data; Determine the moving object mask in the extended second image group according to the movable object mask and the moving object candidate mask.
8. The method according to claim 7, wherein The determining the moving object candidate mask in the smoothed first image group according to the processed IMU measurement data includes: Determine the speed of the stationary objects in the smoothed first image group responding to the corresponding events and the speed of the moving objects in the smoothed first image group responding to the corresponding events according to the processed IMU measurement data, and determine the moving object candidate mask in the smoothed first image group according to the speed of the stationary objects in the smoothed first image group responding to the corresponding events and the speed of the moving objects in the smoothed first image group responding to the corresponding events.
9. The method according to claim 7, wherein The determining the moving object mask in the extended second image group according to the movable object mask and the moving object candidate mask includes: Calculate the intersection over union (IoU) of the movable object mask and the moving object candidate mask; When the IoU is greater than a preset threshold, determine the movable object mask as the moving object mask in the extended second image group.
10. The method according to claim 6, wherein The calculating the joint depth estimation image group according to the IMU measurement data, the second image and the first image group includes: After determining the moving object mask, delete the images in the corresponding areas of the moving object mask in the smoothed first image group to obtain the smoothed first image group after deletion; Perform triangulation depth calculation on the smoothed first image group after deletion to obtain a second depth calculation image group; Perform depth estimation on the extended second image group after deletion to obtain a second depth estimation image group; Perform fusion processing on the second depth calculation image group and the second depth estimation image group to obtain a joint depth estimation image group.
11. The method according to any one of claims 1 to 10, characterized in that, The generating a three-dimensional model according to the second image group, the joint depth estimation image group, the initial pose and the target image includes: Perform iterative calculation in an adversarial generation manner according to the second image group, the joint depth estimation image group, the initial pose and the target image to obtain a predicted pose that meets the convergence condition, and the predicted pose is the parameter of the three-dimensional model.
12. The method according to claim 11, wherein Performing iterative calculations in an adversarial generation manner based on the second image group, the combined depth estimation image group, the initial pose, and the target image to obtain an estimated pose that satisfies the convergence condition, including: For one round of iterative calculation, calculating the predicted second image group, the predicted combined depth estimation image group, and the predicted pose for this round based on the target image respectively; Calculating the first adversarial residual for this round according to the second image group for this round and the predicted second image group for this round, calculating the second adversarial residual for this round according to the combined depth estimation image group for this round and the predicted combined depth estimation image group for this round, and calculating the third adversarial residual for this round according to the estimated pose for this round and the predicted pose for this round, wherein the estimated pose for the first round is calculated based on the initial pose; When the first adversarial residual for this round, the second adversarial residual for this round, and the third adversarial residual for this round all satisfy the convergence condition, determining the estimated pose for this round as the parameters of the three-dimensional model.
13. The method according to claim 12, wherein After calculating the third adversarial residual for this round according to the estimated pose for this round and the predicted pose for this round, the method further includes: When the third adversarial residual is less than the first preset threshold and the next round of iterative calculation is performed, calculating the estimated pose for the next round of iteration according to the estimated pose for this round; When the third adversarial residual is less than the second preset threshold, updating the IMU measurement data according to the estimated pose for this round, and the second preset threshold is less than the first preset threshold.
14. A multi-modal three-dimensional model establishment device, wherein, Including: One or more processors; A memory on which one or more programs are stored; When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the multi-modal three-dimensional model establishment method according to any one of claims 1-13; At least one I / O interface connected between the processor and the memory, configured to implement the information interaction between the processor and the memory.
15. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed, it implements the multi-modal three-dimensional model establishment method according to any one of claims 1-13.