Three-dimensional reconstruction method, device, equipment, computer readable medium and program product

By identifying and repairing the masked image of the dynamic object area, the artifacts and texture blur problems caused by dynamic objects in three-dimensional reconstruction are solved, and the quality and AR experience of the three-dimensional model are improved.

CN120339510APending Publication Date: 2025-07-18LINGBAN INTELLIGENT (HANGZHOU) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510407001.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

During the three-dimensional reconstruction process, dynamic objects lead to artifacts in reconstruction results, blurred textures, feature point matching errors and information loss, affecting the AR experience.

Method used

By identifying dynamic object areas to generate mask maps, the scene image is repaired and three-dimensional reconstruction is carried out to reduce the impact of dynamic objects.

Benefits of technology

Improves the quality and accuracy of the three-dimensional reconstruction model and improves the AR experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339510A_ABST
    Figure CN120339510A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a three-dimensional reconstruction method, device and equipment, a computer readable medium and a program product. A specific embodiment of the method comprises the following steps: acquiring a scene image sequence of a target scene; for each frame of scene image in the scene image sequence, the following steps are executed: identifying a dynamic object area in the scene image; generating a mask pattern corresponding to the scene image based on the identified dynamic object area; repairing the scene image based on the mask image to obtain a repaired scene image; and performing three-dimensional reconstruction based on each obtained scene image to obtain a three-dimensional model corresponding to the target scene. According to the embodiment, the influence of a dynamic object on the quality and the accuracy of the reconstructed model can be reduced, the quality and the accuracy of the reconstructed model are improved, and the AR experience is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly, to a three-dimensional reconstruction method, apparatus, device, computer-readable medium, and program product. Background Art

[0002] Augmented reality (AR) technology is being increasingly widely used in various fields. Among them, wearing AR glasses for AR navigation has shown unique advantages and potential, and three-dimensional scanning and reconstruction of spatial scenes are required.

[0003] However, in this process, since the acquisition process will scan some dynamic objects such as people, vehicles, and TV screens, the positions or shapes of the dynamic objects in the scene are constantly changing, resulting in the inability to consistently correspond to specific spatial positions during the reconstruction process, and artifacts or false geometric shapes may appear in the reconstruction result; the movement of the dynamic objects will also interfere with texture extraction, resulting in blurred or incorrect textures; the feature points on the dynamic objects may move or disappear, affecting the stability of camera pose estimation and scene point cloud construction, and the cumulative error of feature point matching may lead to the failure of the entire reconstruction process; the dynamic objects may block important static parts, causing irrecoverable information loss; for three-dimensional reconstruction based on multiple frames or multiple perspectives, the movement of the dynamic objects causes inconsistent inter-frame information, making it difficult to perform global optimization; in summary, dynamic objects will affect the quality and accuracy of the reconstruction model, thereby affecting the entire AR experience.

[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention

[0005] The content part of the present disclosure is used to briefly introduce the inventive concepts, which will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure propose a three-dimensional reconstruction method, apparatus, electronic device, computer-readable medium, and program product to solve one or more of the technical problems mentioned in the above background art section.

[0007] In a first aspect, some embodiments of the present disclosure provide a three-dimensional reconstruction method, which includes: obtaining a sequence of scene images of a target scene; for each frame of the scene image sequence, performing the following steps: identifying a dynamic object area in the scene image; generating a mask image corresponding to the scene image based on the identified dynamic object area; repairing the scene image based on the mask image to obtain a repaired scene image; and performing three-dimensional reconstruction based on the obtained scene images to obtain a three-dimensional model corresponding to the target scene.

[0008] Optionally, the generating a mask image corresponding to the scene image based on the identified dynamic object area includes: in the scene image, converting the pixel values of the pixels within the dynamic object area to a first value and converting the pixel values of the pixels outside the dynamic object area to a second value to obtain a mask image corresponding to the scene image.

[0009] Optionally, the repairing the scene image based on the mask image to obtain a repaired scene image includes: removing the dynamic object area included in the mask image from the scene image to obtain a removed image; and performing a repair process on the removed image to obtain a repaired scene image.

[0010] Optionally, the performing three-dimensional reconstruction based on the obtained scene images to obtain a three-dimensional model corresponding to the target scene includes: generating a sparse point cloud based on the obtained scene images; generating a dense point cloud based on the sparse point cloud and the scene images; and generating a three-dimensional model corresponding to the target scene based on the dense point cloud.

[0011] Optionally, the generating a sparse point cloud based on the obtained scene images includes: for every two adjacent frames of the scene images, performing feature point matching processing on the two frames of scene images to obtain a feature point matching result corresponding to the two frames of scene images; generating camera pose information corresponding to each scene image in the scene images according to the obtained feature point matching results; and generating a sparse point cloud based on the generated camera pose information and the feature point matching results.

[0012] Optionally, the generating a three-dimensional model corresponding to the target scene based on the dense point cloud includes: performing surface reconstruction on the dense point cloud to obtain an initial three-dimensional model; and performing texture mapping on the initial three-dimensional model based on the scene images to obtain a three-dimensional model corresponding to the target scene.

[0013] Second aspect, some embodiments of the present disclosure provide a 3D reconstruction device, which includes: an acquisition unit configured to acquire a sequence of scene images of a target scene; an execution unit configured to perform the following steps for each frame of the scene image sequence: identify a dynamic object area in the scene image; generate a mask image corresponding to the scene image based on the identified dynamic object area; repair the scene image based on the mask image to obtain a repaired scene image; a reconstruction unit configured to perform 3D reconstruction based on the obtained scene images to obtain a 3D model corresponding to the target scene.

[0014] Optionally, the execution unit is further configured to: in the scene image, convert the pixel values of the pixels within the dynamic object area to a first numerical value, and convert the pixel values of the pixels outside the dynamic object area to a second numerical value to obtain a mask image corresponding to the scene image.

[0015] Optionally, the execution unit is further configured to: crop the dynamic object area included in the mask image from the scene image to obtain a cropped image; perform a repair process on the cropped image to obtain a repaired scene image.

[0016] Optionally, the reconstruction unit is further configured to: generate a sparse point cloud based on the obtained scene images; generate a dense point cloud based on the sparse point cloud and the scene images; generate a 3D model corresponding to the target scene based on the dense point cloud.

[0017] Optionally, the reconstruction unit is further configured to: for every two adjacent frames of the scene images among the scene images, perform feature point matching processing on the two frames of scene images to obtain a feature point matching result corresponding to the two frames of scene images; generate camera pose information corresponding to each scene image among the scene images based on the obtained feature point matching results; generate a sparse point cloud based on the generated camera pose information and the feature point matching results.

[0018] Optionally, the reconstruction unit is further configured to: perform surface reconstruction on the dense point cloud to obtain an initial 3D model; perform texture mapping on the initial 3D model based on the scene images to obtain a 3D model corresponding to the target scene.

[0019] Third aspect, some embodiments of the present disclosure provide an electronic device, which includes: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect above.

[0020] Fourthly, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect above.

[0021] Fifthly, some embodiments of the present disclosure provide a computer program product including a computer program, and the computer program, when executed by a processor, implements the method described in any implementation manner of the first aspect above.

[0022] The above - mentioned embodiments of the present disclosure have the following beneficial effects: Through the 3D reconstruction method of some embodiments of the present disclosure, the influence of dynamic objects on the quality and accuracy of the reconstructed model can be reduced, the quality and accuracy of the reconstructed model are improved, and thus the AR experience is enhanced. Specifically, the reasons why dynamic objects affect the quality and accuracy of the reconstructed model and thus the entire AR experience are as follows: Since the acquisition process scans some dynamic objects such as people, vehicles, and TV screens, the positions or shapes of dynamic objects in the scene are constantly changing, resulting in the inability to consistently correspond to specific spatial positions during the reconstruction process, and artifacts or false geometric shapes may appear in the reconstruction results; the movement of dynamic objects also interferes with texture extraction, resulting in blurred or incorrect textures; the feature points on dynamic objects may move or disappear, affecting the stability of camera pose estimation and scene point cloud construction, and the cumulative error of feature point matching may lead to the failure of the entire reconstruction process; dynamic objects may block important static parts, causing irrecoverable information loss; for 3D reconstruction based on multiple frames or multiple perspectives, the movement of dynamic objects leads to inconsistent inter - frame information, making it difficult to perform global optimization; in summary, dynamic objects have an impact on the quality and accuracy of the reconstructed model, thus affecting the entire AR experience. Based on this, in the 3D reconstruction method of some embodiments of the present disclosure, first, a sequence of scene images of the target scene is obtained. Then, for each frame of the scene image sequence, the following steps are performed: The first step is to identify the dynamic object regions in the above - mentioned scene image. Thus, the dynamic objects in each frame of the scene image can be identified first before 3D reconstruction. The second step is to generate a mask image corresponding to the above - mentioned scene image based on the identified dynamic object regions. Thus, the mask image can be used to distinguish dynamic objects from other regions. The third step is to repair the above - mentioned scene image based on the above - mentioned mask image to obtain a repaired scene image. Thus, the dynamic objects can be removed from the scene image and repaired. Finally, based on the obtained scene images, 3D reconstruction is performed to obtain a 3D model corresponding to the above - mentioned target scene. Thus, the repaired scene images can be used for 3D reconstruction. Also, since the scene images used for 3D reconstruction are repaired and do not include dynamic objects, the influence of dynamic objects on spatial position, texture, feature point matching, information integrity, and time synchronization during the reconstruction process can be reduced, thereby reducing the influence of dynamic objects on the quality and accuracy of the reconstructed model, improving the quality and accuracy of the reconstructed model, and further enhancing the AR experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In combination with the accompanying drawings and referring to the following specific embodiments, the above - mentioned and other features, advantages, and aspects of each embodiment of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0024] Figure 1 is an architectural diagram of an exemplary system to which some embodiments of the present disclosure can be applied;

[0025] Figure 2 is a flowchart of some embodiments of a three-dimensional reconstruction method according to the present disclosure;

[0026] Figure 3 is a schematic diagram of an application scenario of a three-dimensional reconstruction method according to some embodiments of the present disclosure;

[0027] Figure 4 is a schematic diagram of another application scenario of a three-dimensional reconstruction method according to some embodiments of the present disclosure;

[0028] Figure 5 is a schematic structural diagram of some embodiments of a three-dimensional reconstruction apparatus according to the present disclosure;

[0029] Figure 6 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Embodiments

[0030] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0031] In addition, it should be noted that, for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0032] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0033] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".

[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0035] For operations such as the collection, storage, and use of the user's personal information (such as scene images) involved in the present disclosure, before performing the corresponding operations, relevant organizations or individuals fulfill obligations including conducting a personal information security impact assessment, fulfilling the obligation to inform the personal information subject, and obtaining the prior authorization and consent of the personal information subject.

[0036] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0037] Figure 1 An exemplary system architecture 100 of a 3D reconstruction method or a 3D reconstruction device to which some embodiments of the present disclosure can be applied is shown.

[0038] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0039] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shooting applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0040] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices with a display screen and including a camera, including but not limited to smart phones, tablet computers, handheld cameras, head-mounted display devices, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as, for example, multiple software or software modules for providing distributed services, or may be implemented as a single software or software module. No specific limitation is made here.

[0041] The server 105 may be a server that provides various services, such as a background server for processing the scene image sequences collected by the terminal devices 101, 102, 103. The background server may analyze and process data such as the received scene image sequences, and feedback the processing results (such as 3D models) to the head-mounted display device that needs to load the 3D model.

[0042] It should be noted that the 3D reconstruction method provided by the embodiments of the present disclosure can be executed by the terminal device 103 or the server 105. Correspondingly, the 3D reconstruction device can be set in the terminal device 103 or the server 105. No specific limitation is made here.

[0043] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.

[0044] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0045] Continuing to refer to Figure 2 , Figure 2 shows a flow 200 of some embodiments of the 3D reconstruction method according to the present disclosure. The 3D reconstruction method includes the following steps:

[0046] Step 201, obtaining a sequence of scene images of a target scene.

[0047] In some embodiments, the execution entity (such as a head-mounted display device or a server) of the 3D reconstruction method can obtain a sequence of scene images of a target scene. For example, when the execution entity is a head-mounted display device, a camera provided on the head-mounted display device can be used to directly capture consecutive scene images of the target scene as the sequence of scene images. The head-mounted display device can be, but is not limited to: AR glasses, VR glasses, MR glasses. The above camera can be an RGB camera. When the execution entity is a server, the server can obtain a sequence of scene images of the target scene from a head-mounted display device or other types of terminal devices. Other types of terminal devices can include, but are not limited to: mobile phones, tablets, and handheld cameras. The target scene can be a scene that needs to be 3D reconstructed. The scene image can be an image of the captured target scene. The sequence of scene images can be each scene image captured continuously for the target scene. When capturing the scene images, they can be captured from different angles. To improve the 3D reconstruction effect, these images can cover the entire target scene and have a certain overlap (for example, at least 60%). In addition, the camera parameters (such as focal length, aperture, etc.) during shooting can be kept as consistent as possible.

[0048] Step 202, for each frame of the scene image sequence, perform the following steps:

[0049] Step 2021: Identify the dynamic object region in the scene image.

[0050] In some embodiments, the above-mentioned execution entity may identify the dynamic object region in the above-mentioned scene image. Among them, the dynamic object region may be the image region where the dynamic object is located in the scene image. The dynamic object may include, but is not limited to, an object that moves in the target scene or an object that can move in the target scene. For example, the dynamic object may include a person, a vehicle, a TV screen. In practice, the above-mentioned execution entity may use a pre-trained semantic segmentation model to identify the dynamic object region in the above-mentioned scene image, and may also identify the dynamic object region in the above-mentioned scene image through inter-frame difference, or optical flow method or background modeling method.

[0051] Step 2022: Generate a mask image corresponding to the scene image based on the identified dynamic object region.

[0052] In some embodiments, the above-mentioned execution entity may generate a mask image corresponding to the above-mentioned scene image based on the identified dynamic object region. In practice, the above-mentioned execution entity may perform binary mask processing on the above-mentioned scene image based on the above-mentioned dynamic object region to obtain a mask image corresponding to the above-mentioned scene image.

[0053] In some optional implementation manners of some embodiments, the above-mentioned execution entity may, in the above-mentioned scene image, convert the pixel values of the pixels within the above-mentioned dynamic object region into a first numerical value, and convert the pixel values of the pixels outside the above-mentioned dynamic object region into a second numerical value to obtain a mask image corresponding to the above-mentioned scene image. Among them, the above-mentioned first numerical value may be 255. The above-mentioned second numerical value may be 0.

[0054] Step 2023: Repair the scene image based on the mask image to obtain a repaired scene image.

[0055] In some embodiments, the above-mentioned execution entity may repair the above-mentioned scene image based on the above-mentioned mask image to obtain a repaired scene image. In practice, the above-mentioned execution entity may use the above-mentioned mask image to remove the dynamic object region from the above-mentioned scene image, and then may repair the removed dynamic object region in the scene image to obtain a repaired scene image. For example, the inpaint image repair algorithm may be used to repair the scene image.

[0056] As an example, the scene image, the corresponding mask image, and the repaired scene image may be referred to Figure 3 . Figure 3 In, the scene image includes a moving object "person". In the mask image, the "person" on the workbench is marked as white. In the repaired scene image, the "person" on the workbench disappears and the removed area is repaired.

[0057] As another example, the scene image, the corresponding mask image, and the restored scene image can be referred to Figure 4 . Figure 4 In Figure 4 , the moving object "car" is included in the scene image. In the mask image, the moving object "car" is marked white. In the restored scene image, the moving object "car" disappears and the removed area is restored.

[0058] In some optional implementation manners of some embodiments, first, the above-mentioned execution subject may remove the dynamic object area included in the above-mentioned mask image from the above-mentioned scene image to obtain a removed image. In practice, the above-mentioned mask image may be subjected to binary inversion processing to invert the pixels corresponding to the dynamic object area to black to obtain an inverted image. Then, the above-mentioned inverted image may be applied to the above-mentioned scene image to obtain a removed image with the dynamic object area removed. Finally, the above-mentioned removed image is subjected to restoration processing to obtain a restored scene image. In practice, an image restoration algorithm may be used to restore the removed image.

[0059] Step 203: Based on the obtained scene images, perform 3D reconstruction to obtain a 3D model of the corresponding target scene.

[0060] In some embodiments, the above-mentioned execution subject may perform 3D reconstruction based on the obtained scene images to obtain a 3D model of the corresponding target scene. It can be understood that when there are dynamic objects in the scene images, the obtained scene images include the restored scene images. When there are no dynamic objects in the scene images, the obtained scene images include the unrepaired scene images. In practice, the above-mentioned execution subject may adopt a 3D reconstruction method based on 2D images, and use the obtained scene images to perform 3D reconstruction to obtain a 3D model of the corresponding target scene. The 3D reconstruction method based on 2D images may include but is not limited to at least one of the following: structure-based method, point cloud-based method, and deep learning-based method. The structure-based method uses a known 3D model as a reference and reconstructs the 3D model by matching feature points in the 2D image, usually requiring a large number of reference models and computing resources. The point cloud-based method converts the 2D image into a 3D point cloud and then uses a point cloud matching algorithm to reconstruct the 3D model, usually requiring 2D images from multiple perspectives to obtain sufficient information. The deep learning-based method uses a deep learning model to learn the representation of the 3D model from the 2D image, usually requiring a large amount of training data and computing resources. The 3D model of the target scene can be used to be loaded in a head-mounted display device, and virtual content can also be superimposed on the 3D model for the wearing user to view.

[0061] In some alternative implementations of some embodiments, the above-mentioned execution entity may perform 3D reconstruction based on the obtained respective scene images through the following steps to obtain a 3D model corresponding to the above-mentioned target scene:

[0062] First step, generate a sparse point cloud based on the obtained respective scene images. Among them, the points in the above-mentioned sparse point cloud may correspond to the key points identified from the scene images.

[0063] Second step, generate a dense point cloud based on the above-mentioned sparse point cloud and the above-mentioned respective scene images. The number of points included in the above-mentioned dense point cloud is greater than the number of points included in the above-mentioned sparse point cloud.

[0064] Third step, generate a 3D model corresponding to the above-mentioned target scene based on the above-mentioned dense point cloud.

[0065] In some alternative implementations of some embodiments, the above-mentioned execution entity may generate a sparse point cloud based on the obtained respective scene images through the following steps:

[0066] First step, for every two adjacent scene images among the above-mentioned respective scene images, perform feature point matching processing on the two scene images to obtain a feature point matching result corresponding to the two scene images. In practice, the above-mentioned execution entity may use a feature point matching algorithm to perform feature point matching processing on the two scene images to obtain a feature point matching result corresponding to the two scene images. The feature point matching algorithm may include but is not limited to: SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), and ORB (Oriented FAST and Rotated BRIEF). The feature point matching result may include the image coordinates of each group of feature points matched in the two scene images. Each group of feature points may represent that the point represents the same physical position under different perspectives.

[0067] Second step, generate the camera pose information corresponding to each scene image among the above-mentioned respective scene images according to the obtained respective feature point matching results. Among them, the above-mentioned camera pose information may include the position and orientation of the camera corresponding to when shooting the scene image. In practice, the SfM (Structure from Motion) method may be used to estimate the position and orientation of the camera corresponding to the scene image as the camera pose information.

[0068] Third step, generate a sparse point cloud based on the generated respective camera pose information and the above-mentioned respective feature point matching results. In practice, the above-mentioned execution entity may use the triangulation process in the SfM method to determine the three-dimensional coordinates of each group of feature points in space, and obtain each three-dimensional coordinate as the sparse point cloud. Each point in the sparse point cloud is a spatial point determined according to the matched feature points in two or more scene images and the corresponding camera pose information.

[0069] In some alternative implementations of some embodiments, the above-mentioned execution entity can generate a three-dimensional model corresponding to the target scene based on the above-mentioned dense point cloud through the following steps:

[0070] First step, perform surface reconstruction on the above-mentioned dense point cloud to obtain an initial three-dimensional model. In practice, methods based on meshes (such as Poisson reconstruction, sphere scanning, etc.), methods based on implicit surfaces (such as Marching Cubes algorithm), or Delaunay triangulation can be used to perform surface reconstruction on the above-mentioned dense point cloud to obtain an initial three-dimensional model. Thus, a continuous and smooth surface model can be extracted from the dense point cloud.

[0071] Second step, based on the above-mentioned various scene images, perform texture mapping on the above-mentioned initial three-dimensional model to obtain a three-dimensional model corresponding to the target scene. In practice, an image-based texture mapping method can be used to perform texture mapping on the above-mentioned initial three-dimensional model using the above-mentioned various scene images to obtain a three-dimensional model corresponding to the target scene. Thus, the color information of the repaired scene image can be mapped onto the reconstructed surface, making the three-dimensional model look more realistic.

[0072] In some alternative implementations of some embodiments, the above-mentioned execution entity can generate a dense point cloud based on the above-mentioned sparse point cloud and the above-mentioned various scene images through the following steps:

[0073] First step, input the above-mentioned various scene images into the feature extraction network included in the pre-trained dense point cloud generation model to obtain the image feature maps of the various scene images. Among them, the above-mentioned dense point cloud generation model can be a pre-trained neural network model that takes the various scene images and the corresponding sparse point cloud as inputs and outputs a dense point cloud. The above-mentioned dense point cloud generation model can include an image feature extraction network, a point cloud feature enhancement network, a cross-modal fusion network, and a dense prediction network. The above-mentioned image feature extraction network can include convolutional layers, pooling layers, and normalization layers. The above-mentioned convolutional layers can use pre-trained ResNet or EfficientNet to extract image features. The above-mentioned pooling layers can be used to reduce the spatial size of the feature maps while retaining key information. The above-mentioned normalization layer can be, for example, Batch Normalization, which can be used to accelerate training and improve stability. The image feature maps can represent the texture and shape information captured in the scene images. Thus, the significant features of each scene image can be extracted, providing a basis for subsequent cross-modal fusion.

[0074] Second, input the above sparse point cloud and the obtained image feature maps into the above point cloud feature enhancement network to obtain the enhanced point cloud feature representations of each point. Among them, the above point cloud feature enhancement network may include a graph convolutional layer and a fully connected layer. The graph convolutional layer can process the sparse point cloud using a graph neural network (GNN), where each node represents a point and the edges represent the proximity relationships between points. Through multi-layer graph convolutional operations, the feature representation of each point is gradually updated. The fully connected layer can fuse the information extracted from the image feature maps into the point cloud feature representation. The enhanced point cloud feature representation can be the enhanced point cloud feature representation, and each point not only contains its own geometric information but also contains relevant image features. Thus, the local geometric relationships in the point cloud can be captured through graph convolution, and the image features can be fused to enhance the feature representation of the point cloud.

[0075] Third, input the obtained enhanced point cloud feature representations and the image feature maps into the above cross-modal fusion network to obtain the fused point cloud feature representation. Among them, the above cross-modal fusion network includes an attention mechanism and a bilinear pooling layer. The attention mechanism can use a self-attention mechanism or a cross-attention mechanism to dynamically adjust the importance of different modality features, highlighting the parts that are most helpful for reconstruction. The bilinear pooling layer can capture the second-order statistical information between features by calculating the outer product between feature maps, further enhancing the effect of feature fusion. The fused point cloud feature representation can represent the fused point cloud feature representation, integrating the complementary information of the image and the point cloud. Thus, the data from the image and the point cloud can be effectively combined to provide rich context information for subsequent prediction.

[0076] Fourth, input the above fused point cloud feature representation into the above dense prediction network to obtain the initial dense point cloud. Among them, the above dense prediction network includes various deconvolution layers, skip connections, and upsampling layers. The various deconvolution layers and skip connections can form the Decoder structure of a U-Net. These layers can gradually restore the spatial resolution of the feature maps and predict new point positions based on the fused features. The upsampling layer can increase the spatial resolution of the feature map, making the finally generated point cloud have a higher density.

[0077] In the fifth step, for each point in the above-mentioned initial dense point cloud, a normal vector corresponding to the above point is generated. In practice, a preset number of nearest neighbor points to the above point can be selected from the above-mentioned initial dense point cloud. For example, a KD tree can be used to improve the search efficiency. Then, a covariance matrix can be generated based on the coordinates of the above-mentioned preset number of nearest neighbor points and the coordinates of the above point. For example, the centroid point of the coordinates of the above-mentioned preset number of nearest neighbor points and the coordinates of the above point can be determined first. Then, for each nearest neighbor point, the coordinate difference between the above nearest neighbor point and the above centroid point and the inverse matrix of the coordinate difference can be determined, and the product matrix of the coordinate difference and the inverse matrix of the coordinate difference can be determined. Finally, the ratio of the sum of the respective product matrices to the above-mentioned preset number can be determined as the covariance matrix. Here, the coordinates and coordinate differences can be represented in the form of vectors. After that, the covariance matrix can be subjected to eigenvalue decomposition to obtain each eigenvalue and each eigenvector. Finally, the eigenvector corresponding to the smallest eigenvalue can be determined as the normal vector corresponding to the above point.

[0078] In the sixth step, based on the normal vectors of each point in the above-mentioned initial dense point cloud, consistency adjustment processing is performed on the above-mentioned initial dense point cloud to obtain a first dense point cloud. In practice, for each point in the initial dense point cloud, the included angle between the above point and the normal vectors of the corresponding nearest neighbor points can be generated to obtain each normal angle. Then, in response to determining that there is a normal angle that satisfies the preset threshold condition, adjustment processing is performed on the above point. For example, the preset threshold condition can be that the normal angle is greater than the preset threshold. The preset threshold can be 30°. Specifically, position adjustment and normal adjustment can be performed on the above point.

[0079] When performing position adjustment, first, the weight coefficients corresponding to the above point and each nearest neighbor point can be generated. The weight coefficients can be generated by a preset weight function. The weight function can include a Gaussian function or a custom attenuation function. The weight function can take the coordinates of two points as independent variables and the weight coefficients of the two points as dependent variables. The closer the distance between the two points, the greater the weight coefficient, and vice versa. Then, the sum of the obtained respective weight coefficients can be determined as the weight coefficient sum. Next, for each nearest neighbor point, the product of the weight coefficient corresponding to the nearest neighbor point and the coordinates of the nearest neighbor point can be determined as the weighted nearest neighbor point. Secondly, the sum of the obtained respective weighted nearest neighbor points can be determined as the weighted nearest neighbor point sum. Finally, the ratio of the above-mentioned weighted nearest neighbor point sum to the above-mentioned weight coefficient sum can be determined as the position of the adjusted above point.

[0080] When performing normal adjustment, for each neighboring point of the above-mentioned point, the normal weight coefficient between the normal vector of the above-mentioned point and the normal vector of the above-mentioned neighboring point can be determined first, and then the product of the above-mentioned normal weight coefficient and the normal vector of the above-mentioned neighboring point can be determined as the weighted normal vector. The normal weight coefficient can be generated by a normal weight function defined based on normal similarity. The normal weight function takes the normal vectors of two points as independent variables and the normal weight coefficient as the dependent variable. The more similar the normal vectors of the two points are, the larger the normal weight coefficient is, and vice versa. Then, the modulus of the sum of the determined weighted normal vectors can be determined as the denominator value. Finally, for each neighboring point of the above-mentioned point, the ratio of the weighted normal vector corresponding to the above-mentioned neighboring point to the above-mentioned denominator value can be determined to obtain the adjusted normal vector. Then, the normal of the above-mentioned neighboring point can be adjusted according to the above-mentioned adjusted normal vector.

[0081] Step 7: Smooth the above-mentioned first dense point cloud to obtain a second dense point cloud. In practice, a smoothing algorithm can be used to smooth the above-mentioned first dense point cloud to obtain a second dense point cloud. For example, the smoothing algorithm can be the weighted average method and the smoothing filter. Thus, the initially generated dense point cloud can be further refined to have good performance in terms of details and overall coherence.

[0082] Step 9: Post-process the above-mentioned second dense point cloud to obtain a dense point cloud. In practice, the above-mentioned execution entity can perform denoising and hole filling on the above-mentioned second dense point cloud to obtain a dense point cloud. For example, a filter (such as bilateral filtering, Gaussian filtering, etc.) can be used for denoising. Interpolation methods or other surface reconstruction techniques can be used for hole filling to fill the hole areas in the point cloud.

[0083] The above steps 1-9 are an inventive point of the embodiment of the present disclosure, which solves the technical problem of "the poor quality of the generated dense point cloud". The factors that lead to the poor quality of the generated dense point cloud are often as follows: the existing methods for generating dense point clouds only use convolution operations to extract image features for subsequent dense feature generation, and the used modality is relatively single, and the geometric structure of the point cloud cannot be effectively utilized, resulting in the poor quality of the generated dense point cloud. If the above factors are solved, the effect of improving the quality of the generated dense point cloud can be achieved. To achieve this effect, the present disclosure makes full use of the complementary advantages of image and point cloud data, improves the quality of the dense point cloud through an efficient cross-modal fusion mechanism, and also introduces a geometric perception and optimization process to ensure that the newly generated points can not only accurately reflect the surface details but also maintain geometric consistency in the entire model.

[0084] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the 3D reconstruction method of some embodiments of the present disclosure, the influence of dynamic objects on the quality and accuracy of the reconstructed model can be reduced, the quality and accuracy of the reconstructed model are improved, and thus the AR experience is enhanced. Specifically, the reasons why dynamic objects affect the quality and accuracy of the reconstructed model and thus the entire AR experience are as follows: Since some dynamic objects such as people, vehicles, and TV screens are scanned during the acquisition process, the positions or shapes of the dynamic objects in the scene are constantly changing, resulting in the inability to consistently correspond to specific spatial positions during the reconstruction process, and artifacts or false geometries may appear in the reconstruction results; the movement of dynamic objects also interferes with texture extraction, resulting in blurred or incorrect textures; the feature points on dynamic objects may move or disappear, affecting the stability of camera pose estimation and scene point cloud construction, and the cumulative error of feature point matching may lead to the failure of the entire reconstruction process; dynamic objects may occlude important static parts, causing irrecoverable information loss; for 3D reconstruction based on multiple frames or multiple perspectives, the movement of dynamic objects leads to inconsistent inter-frame information, making it difficult to perform global optimization; in summary, dynamic objects affect the quality and accuracy of the reconstructed model, thereby affecting the entire AR experience. Based on this, in the 3D reconstruction method of some embodiments of the present disclosure, first, a sequence of scene images of the target scene is obtained. Then, for each frame of the scene image sequence, the following steps are performed: The first step is to identify the dynamic object regions in the above-mentioned scene image. Thus, the dynamic objects in each frame of the scene image can be identified first before 3D reconstruction. The second step is to generate a mask image corresponding to the above-mentioned scene image based on the identified dynamic object regions. Thus, the mask image can be used to distinguish dynamic objects from other regions. The third step is to repair the above-mentioned scene image based on the above-mentioned mask image to obtain a repaired scene image. Thus, the dynamic objects can be removed from the scene image and repaired. Finally, based on the obtained scene images, 3D reconstruction is performed to obtain a 3D model corresponding to the above-mentioned target scene. Thus, the repaired scene images can be used for 3D reconstruction. Also, because the scene images used for 3D reconstruction are repaired and do not include dynamic objects, the influence of dynamic objects on spatial position, texture, feature point matching, information integrity, and time synchronization during the reconstruction process can be reduced, thereby reducing the influence of dynamic objects on the quality and accuracy of the reconstructed model, improving the quality and accuracy of the reconstructed model, and further enhancing the AR experience.

[0085] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a 3D reconstruction device. These device embodiments correspond to Figure 2 the method embodiments shown, and the device can be specifically applied to various electronic devices.

[0086] As shown in Figure 5As shown, the 3D reconstruction device 500 of some embodiments includes: an acquisition unit 501, an execution unit 502, and a reconstruction unit 503. Among them, the acquisition unit 501 is configured to acquire a sequence of scene images of a target scene; the execution unit 502 is configured to perform the following steps for each frame of scene image in the above sequence of scene images: identify the dynamic object area in the above scene image; generate a mask image corresponding to the above scene image based on the identified dynamic object area; repair the above scene image based on the above mask image to obtain a repaired scene image; the reconstruction unit 503 is configured to perform 3D reconstruction based on the obtained respective scene images to obtain a 3D model corresponding to the above target scene.

[0087] Optionally, the execution unit 502 can be further configured to: in the above scene image, convert the pixel values of the pixels within the above dynamic object area to a first numerical value, and convert the pixel values of the pixels outside the above dynamic object area to a second numerical value to obtain a mask image corresponding to the above scene image.

[0088] Optionally, the execution unit 502 can be further configured to: remove the dynamic object area included in the above mask image from the above scene image to obtain a removed image; perform a repair process on the above removed image to obtain a repaired scene image.

[0089] Optionally, the reconstruction unit 503 can be further configured to: generate a sparse point cloud based on the obtained respective scene images; generate a dense point cloud based on the above sparse point cloud and the above respective scene images; generate a 3D model corresponding to the above target scene based on the above dense point cloud.

[0090] Optionally, the reconstruction unit 503 can be further configured to: for every two adjacent frames of scene images among the above respective scene images, perform feature point matching processing on the above two frames of scene images to obtain a feature point matching result corresponding to the above two frames of scene images; generate camera pose information corresponding to each scene image among the above respective scene images according to the obtained respective feature point matching results; generate a sparse point cloud based on the generated respective camera pose information and the above respective feature point matching results.

[0091] Optionally, the reconstruction unit 503 can be further configured to: perform surface reconstruction on the above dense point cloud to obtain an initial 3D model; perform texture mapping on the above initial 3D model based on the above respective scene images to obtain a 3D model corresponding to the above target scene.

[0092] It can be understood that the various units described in the device 500 and the reference Figure 2corresponds to each step in the described method. Thus, the operations, features, and beneficial effects described above for the method also apply to the apparatus 500 and the units included therein, and will not be elaborated herein.

[0093] Reference is now made to Figure 6 , which shows a schematic structural diagram of an electronic device 600 (such as Figure 1 the head-mounted display device or server in Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0094] As Figure 6 shown, the electronic device 600 may include a processing device 601 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 602 or a program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0095] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included. Figure 6 Each block shown in

[0096] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.

[0097] It should be noted that the computer-readable medium described in some embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0098] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0099] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain a sequence of scene images of a target scene; for each frame of scene image in the sequence of scene images, perform the following steps: identify a dynamic object region in the scene image; generate a mask image corresponding to the scene image based on the identified dynamic object region; repair the scene image based on the mask image to obtain a repaired scene image; and perform three-dimensional reconstruction based on the obtained respective scene images to obtain a three-dimensional model corresponding to the target scene.

[0100] Computer program code for performing the operations of some embodiments of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0102] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, an execution unit, and a reconstruction unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring the sequence of scene images of the target scene".

[0103] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0104] Some embodiments of the present disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above three-dimensional reconstruction methods.

[0105] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.

Claims

1. A three-dimensional reconstruction method, comprising: Obtaining a sequence of scene images of a target scene; For each scene image in the sequence of scene images, performing the following steps: Identifying a dynamic object region in the scene image; Generating a mask image corresponding to the scene image based on the identified dynamic object region; Repairing the scene image based on the mask image to obtain a repaired scene image; Performing three-dimensional reconstruction based on the obtained scene images to obtain a three-dimensional model corresponding to the target scene.

2. The method according to claim 1, wherein, The generating a mask image corresponding to the scene image based on the identified dynamic object region includes: In the scene image, converting the pixel values of the pixels within the dynamic object region to a first numerical value and converting the pixel values of the pixels outside the dynamic object region to a second numerical value to obtain a mask image corresponding to the scene image.

3. The method according to claim 1, wherein The repairing the scene image based on the mask image to obtain a repaired scene image includes: Cropping the dynamic object region included in the mask image from the scene image to obtain a cropped image; Performing a repair process on the cropped image to obtain a repaired scene image.

4. The method according to claim 1, wherein The performing three-dimensional reconstruction based on the obtained scene images to obtain a three-dimensional model corresponding to the target scene includes: Generating a sparse point cloud based on the obtained scene images; Generating a dense point cloud based on the sparse point cloud and the scene images; Generating a three-dimensional model corresponding to the target scene based on the dense point cloud.

5. The method according to claim 4, wherein The generating a sparse point cloud based on the obtained scene images includes: For every two adjacent scene images in the scene images, performing feature point matching processing on the two scene images to obtain a feature point matching result corresponding to the two scene images; Generating camera pose information corresponding to each scene image in the scene images according to the obtained feature point matching results; Generating a sparse point cloud based on the generated camera pose information and the feature point matching results.

6. The method according to claim 4, wherein The generating a three-dimensional model corresponding to the target scene based on the dense point cloud includes: Performing surface reconstruction on the dense point cloud to obtain an initial three-dimensional model; Performing texture mapping on the initial three-dimensional model based on the scene images to obtain a three-dimensional model corresponding to the target scene.

7. A three-dimensional reconstruction device, comprising: An acquisition unit configured to acquire a sequence of scene images of a target scene; An execution unit configured to, for each scene image in the sequence of scene images, perform the following steps: identifying a dynamic object region in the scene image; Generating a mask image corresponding to the scene image based on the identified dynamic object region; Repairing the scene image based on the mask image to obtain a repaired scene image; A reconstruction unit configured to perform three-dimensional reconstruction based on the obtained scene images to obtain a three-dimensional model corresponding to the target scene.

8. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.

10. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.