Self-updating streaming three-dimensional Gaussian reconstruction method and device, equipment and storage medium

By using a self-updating streaming 3D Gaussian reconstruction method, a global library is dynamically updated to generate a consistent 3D Gaussian sphere distribution. This solves the problem of large data volume and difficulty in flexible updating in existing technologies, and achieves efficient 3D reconstruction results.

CN121883756APending Publication Date: 2026-04-17BEIJING INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2025-12-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing 3D reconstruction technologies suffer from large data volumes and difficulty in flexible updates, resulting in low reconstruction efficiency, especially in application scenarios with frequent incremental updates.

Method used

A self-updating streaming 3D Gaussian reconstruction method is adopted. Through the principle of real-time updating and redundancy compression, a consistent 3D Gaussian sphere distribution is generated. The global library is dynamically updated using context information, and redundant Gaussian spheres are deleted to reduce storage size.

Benefits of technology

While ensuring reconstruction accuracy, the number of Gaussian spheres is reduced, reconstruction efficiency and visual realism are improved, and storage size and rendering speed are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883756A_ABST
    Figure CN121883756A_ABST
Patent Text Reader

Abstract

The invention provides a self-updating streaming three-dimensional Gaussian reconstruction method and device, equipment and a storage medium, relates to the technical field of computer vision, solves the problem of low efficiency of a three-dimensional reconstruction scheme in related technologies, and can dynamically update a global library by using context information, thereby improving the reconstruction efficiency of the three-dimensional reconstruction scheme. Therefore, more coordinated and consistent three-dimensional Gaussian ball distribution is generated, redundant structures such as fuzzy structures and floating structures are effectively inhibited, and through continuous deletion and screening of redundant Gaussian balls, the system can greatly reduce the number of Gaussian balls while ensuring the reconstruction precision, so that the storage scale is reduced, and the reconstruction efficiency is improved. And the rendering speed and the transmission efficiency can be accelerated, so that the visual trueness of a reconstruction result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a self-updating streaming 3D Gaussian reconstruction method, apparatus, device and storage medium. Background Technology

[0002] 3D scene reconstruction is a core task in computer vision and graphics, widely used in applications such as virtual reality, autonomous driving, and film and television special effects. Related technologies achieve 3D scene reconstruction through rapid 3D reconstruction techniques, employing various spatial representation methods, such as point cloud-based reconstruction, mesh patch-based reconstruction, and 3D Gaussian splashing representation. However, each method has limitations in terms of accuracy, speed, storage efficiency, and editability. For example, the data storage scale is difficult to control. In point cloud schemes, to obtain sufficient detail, the point density needs to be significantly increased, often reaching tens or even hundreds of millions of points, resulting in a rapid expansion of data volume that is difficult to compress, significantly increasing the computational cost of rendering and optimization. Similarly, in 3D Gaussian splashing schemes, the number of Gaussian distributions is difficult to constrain, exhibiting disordered growth in long-sequence reconstructions, leading to a continuous increase in overall storage overhead and computational burden.

[0003] Furthermore, the scene structure in related technologies lacks flexible update capabilities. While mesh patching can describe object surfaces, modifying or replacing local structures often requires regenerating or optimizing the entire mesh, resulting in low efficiency and unsuitability for applications requiring frequent incremental updates. 3D Gaussian representations, due to highly coupled internal parameters, make it even more difficult to locally edit or delete objects, limiting their application expansion in interactive and dynamic environments. Consequently, the 3D reconstruction solutions provided in these technologies involve large amounts of data and are difficult to update flexibly, leading to low reconstruction efficiency. Summary of the Invention

[0004] This application provides a self-updating streaming 3D Gaussian reconstruction method, apparatus, device, and storage medium, which solves the problem of low efficiency in related 3D reconstruction schemes. This scheme can generate a more coordinated and consistent 3D Gaussian sphere distribution, improve the visual realism of the reconstruction results, and reduce the number of Gaussian spheres and storage size while ensuring reconstruction accuracy, thereby improving reconstruction efficiency.

[0005] In a first aspect, this application provides a self-updating streaming 3D Gaussian reconstruction method, which includes: In response to the 3D reconstruction operation of the target object in multiple frames of an image sequence, the current frame image is acquired and the target image associated with the current frame image is determined in the multiple frames of the image sequence. The multiple frames of the image sequence are images of the target object under any shooting angle. Based on the current frame image and the target image, feature extraction is performed on the current frame image to obtain the three-dimensional feature information of the target object on the current frame image; The target Gaussian sphere is determined by the three-dimensional Gaussian sphere stored in the preset global library, and the target Gaussian sphere is projected onto the current frame image to obtain historical feature information. The target Gaussian sphere is associated with the three-dimensional feature information of the target object's display content on the current frame image. Based on 3D feature information and historical feature information, the 3D Gaussian sphere in the global library is updated in real time.

[0006] Secondly, this application also provides a self-updating streaming 3D Gaussian reconstruction device, comprising: The image selection module is configured to, in response to the three-dimensional reconstruction operation of the target object in the multi-frame images of the image sequence, acquire the current frame image and determine the target image associated with the current frame image in the multi-frame images of the image sequence, wherein the multi-frame images in the image sequence are images of the target object under any shooting angle. The feature extraction module is configured to extract features from the current frame image based on the current frame image and the target image, and obtain the three-dimensional feature information of the target object in the current frame image; The Gaussian flattening module is configured to determine the target Gaussian sphere from the three-dimensional Gaussian spheres stored in the preset global library, and project the target Gaussian sphere onto the current frame image to obtain historical feature information. The target Gaussian sphere is associated with the three-dimensional feature information of the target object's display content on the current frame image. The data update module is configured to update the 3D Gaussian sphere in the global library in real time based on 3D feature information and historical feature information.

[0007] Thirdly, this application also provides an electronic device comprising: One or more processors; A storage device is provided for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the self-updating streaming 3D Gaussian reconstruction method of this application.

[0008] Fourthly, this application also provides a storage medium for storing computer-executable instructions, which, when executed by a processor, are used to execute the self-updating streaming 3D Gaussian reconstruction method of this application.

[0009] This application's solution can dynamically update the global library by utilizing contextual information, thereby generating a more consistent 3D Gaussian sphere distribution. This effectively suppresses redundant structures such as blurring and floating objects. Furthermore, by continuously deleting and filtering redundant Gaussian spheres, the system can significantly reduce the number of Gaussian spheres while ensuring reconstruction accuracy, thereby reducing storage size and helping to accelerate rendering speed and transmission efficiency, thus improving the visual realism of the reconstruction results. Attached Figure Description

[0010] Figure 1 A schematic diagram illustrating the steps of a self-updating streaming 3D Gaussian reconstruction method provided in an embodiment of this application.

[0011] Figure 2 This is a schematic diagram of an image sequence corresponding to a target object provided in an embodiment of this application.

[0012] Figure 3 This is a schematic diagram illustrating the steps for determining a target image according to an embodiment of this application.

[0013] Figure 4 This is a schematic diagram illustrating the steps of comparing and analyzing three-dimensional feature information and historical feature information according to an embodiment of this application.

[0014] Figure 5 This is a schematic diagram of the structure of a self-updating streaming three-dimensional Gaussian reconstruction device provided in an embodiment of this application.

[0015] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0016] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of the present application and are not intended to limit the scope of the present application. Furthermore, it should be noted that, for ease of description, the accompanying drawings only show the parts relevant to the embodiments of the present application, and not all structures. Those skilled in the art, after reading this specification, should be able to deduce that any combination of technical features can constitute an optional implementation method, provided that the technical features do not contradict each other.

[0017] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. In the description of this application, "multiple" means two or more, and "several" means one or more.

[0018] Fast 3D reconstruction techniques can employ various spatial representation methods, such as point cloud-based reconstruction, mesh patch-based reconstruction, and 3D Gaussian splashing reconstruction. The point cloud approach acquires or estimates dense or sparse 3D point clouds and uses them as the geometric representation of the scene. Point clouds directly record the 3D coordinates of each point in space, thus reflecting the original measurement accuracy well, and the implementation process is relatively simple. The mesh patch approach represents the surface of 3D objects by generating discrete geometric units such as triangular meshes, and textures can be added to the mesh to enhance visual effects. This approach produces relatively smooth and renderable results when processing regular, closed object structures. The 3D Gaussian splashing approach approximates the geometric and appearance information of the scene by placing a large number of Gaussian distributions (Gaussian spheres) in 3D space. Each Gaussian distribution has parameters such as center position, scale, orientation, and color. Through rendering optimization, it can quickly generate high-quality images from new perspectives. However, each method has limitations in terms of accuracy, speed, storage efficiency, and editability.

[0019] For example, there is a drawback in controlling the scale of data storage. In point cloud solutions, to obtain sufficient detail, the density of points needs to be significantly increased, often reaching tens or even hundreds of millions of points. This leads to a rapid expansion of data volume, making compression difficult and significantly increasing the computational cost of rendering and optimization. Similarly, in 3D Gaussian splashing solutions, the number of Gaussian distributions is difficult to constrain, exhibiting disordered growth in long image sequence reconstructions, causing a continuous increase in overall storage overhead and computational burden. Furthermore, scene structures lack flexible update capabilities. While mesh patching solutions can describe object surfaces, modifying or replacing local structures often requires regenerating or optimizing the entire mesh, resulting in low efficiency and unsuitability for applications requiring frequent incremental updates. 3D Gaussian representations, due to highly coupled internal parameters, make it even more difficult to locally edit or delete objects, limiting their application expansion in interactive and dynamic environments. Therefore, it is evident that the 3D reconstruction solutions provided by related technologies involve large data volumes and are difficult to update flexibly, leading to low reconstruction efficiency.

[0020] In response to this, this application provides a self-updating streaming 3D Gaussian reconstruction method. Based on the principles of real-time updating and redundant compression, this method effectively solves the problems of slow reconstruction speed, redundant results, and difficulty in flexible updating in related technologies. Furthermore, by adopting a feedforward processing flow, it eliminates the need for extensive iterative optimization and enables real-time 3D reconstruction within a limited time. This allows for high fidelity while maintaining a compact and controllable data scale. Compared with related technologies, this method improves compression efficiency, reconstruction efficiency, stability, and editability, providing a more efficient and reliable solution for the practical application of 3D reconstruction technology.

[0021] Figure 1 This diagram illustrates the steps of a self-updating streaming 3D Gaussian reconstruction method according to an embodiment of this application. This self-updating streaming 3D Gaussian reconstruction method can be applied to electronic devices, such as computers and servers. The electronic device has a corresponding global library that records the currently generated 3D Gaussian spheres and provides corresponding add, delete, modify, and query interfaces for updating and maintaining the global library. As shown in the diagram, this method processes image sequences to perform 3D reconstruction of target objects in the images. Specific steps include S110-S140: Step S110: In response to the three-dimensional reconstruction operation of the target object in the multi-frame images of the image sequence, acquire the current frame image and determine the target image associated with the current frame image in the multi-frame images of the image sequence.

[0022] It is conceivable that multiple frames of images of the target object are needed during 3D reconstruction. The input multiple frames are input into the electronic device in the form of an image sequence. Therefore, the image sequence includes multiple frames, each representing an image of the target object from any shooting angle. Optionally, the image sequence can also be a video stream, where the multiple frames are frames from a video taken of the target object, ensuring temporal continuity and closer correlation between adjacent frames. Alternatively, the multiple frames in the image sequence can be images acquired from multiple shooting angles, such as images of the same target object taken from different shooting angles, thus obtaining multi-view images of the target object and forming an image sequence. Figure 2 As shown, Figure 2 This is a schematic diagram of an image sequence corresponding to a target object provided in an embodiment of this application. When viewing a target object (such as...) Figure 2 During the 3D reconstruction of the cuboid shown, the input image sequence includes images from multiple different shooting perspectives, such as... Figure 2 Images pic1, pic2, and pic3 in the image correspond to different shooting angles, thus showing different sides of the target object in the image (side A, side B, and side C in the image).

[0023] After acquiring the image sequence, in response to the 3D reconstruction operation of the target object in the multi-frame images of the image sequence, the currently processed image is used as the current frame image. For example, if the currently processed image is the first frame in the image sequence, then the first frame image is used as the current frame image; if the currently processed image is the second frame in the image sequence, then the second frame image is used as the current frame image. It is conceivable that during the 3D reconstruction process, each frame in the image sequence is traversed to obtain feature information associated with the target object in the images. Therefore, after determining the current frame image, a target image associated with the current frame image is selected from the image sequence. This target image contains the same image content as the current frame image, such as... Figure 2 As shown, image pic1 is the current frame image, and its image content corresponds to surface A of the target object, while the image content of image pic3 corresponds to surfaces A, B, and C of the target object. Therefore, image pic3 is used as the target image of image pic1.

[0024] Step S120: Based on the current frame image and the target image, perform feature extraction on the current frame image and obtain the three-dimensional feature information of the target object on the current frame image.

[0025] During feature extraction, for the current frame image, the electronic device combines the current frame image with the target image to extract features. It can be assumed that the current frame image and the target image are related in terms of image content. For example, the current frame image displays side A of the target object, while the target image displays side A, side B, and side C of the target object. The side A of the target object has different features in different images. By combining the target image, feature extraction is performed on the current frame image to introduce contextual information. This allows the acquisition of three-dimensional feature information of the current frame image from different shooting angles, such as three-dimensional position, shape, pose, and texture. This ensures the consistency of the generated three-dimensional feature information in space and helps to avoid problems such as discontinuity or misalignment in local reconstruction.

[0026] Step S130: Determine the target Gaussian sphere from the three-dimensional Gaussian spheres stored in the preset global library, and project the target Gaussian sphere onto the current frame image to obtain historical feature information.

[0027] A 3D Gaussian sphere corresponds to a Gaussian distribution at a single point (i.e., a Gaussian point) in 3D space. It includes the center coordinates, covariance matrix, and radiance. The center coordinates (Position) represent the location of the Gaussian point in 3D space, specifically the center point coordinates (x, y, z). The covariance matrix describes the shape and orientation of the Gaussian point in 3D space. Represented by a 3×3 matrix, the diagonal elements represent the variance of the Gaussian point along the x, y, and z axes (i.e., the "fat" or "thickness" of the Gaussian sphere), while the off-diagonal elements represent the correlation between the axes (i.e., the degree of rotation and distortion of the Gaussian point). By adjusting the covariance matrix, the Gaussian point can be "stretched" or "rotated" into an ellipsoid to fit different shaped object surfaces. Radiance (Opacity) represents the color and transparency of the Gaussian point, typically including RGB color values ​​and opacity (Alpha value). This parameter allows the Gaussian point to blend overlapping areas, achieving a more natural surface rendering effect.

[0028] In this regard, the 3D Gaussian sphere stored in the global library corresponds to the 3D Gaussian sphere of the target object. During the processing of each frame in the image sequence, the global library is continuously updated, thus recording in the global library the 3D Gaussian spheres generated from the 3D feature information extracted from the processed images. It is conceivable that if the current frame is the first frame in the image sequence, the number of 3D Gaussian spheres in the global library can be 0, and the 3D Gaussian sphere generated from the features extracted from the current frame will be added to the global library. Of course, the global library can also pre-store 3D Gaussian spheres generated when performing 3D reconstruction of the same scene. For example, if the target object in the current image sequence has already completed 3D reconstruction of a part of the target object in another image sequence, then the 3D Gaussian sphere of that part can be added to the global library. For example, the target object in the current image sequence is multiple buildings, while the target object in another image sequence is one of the multiple buildings, and the 3D reconstruction of that building has been completed, that is, the corresponding 3D Gaussian sphere has been generated. In this case, when performing 3D reconstruction of the target object in the current image sequence, the 3D Gaussian sphere generated above can be added to the global library.

[0029] Furthermore, after determining the current frame image, the target Gaussian sphere can be determined from the 3D Gaussian sphere stored in the global library. The target Gaussian sphere is associated with the 3D feature information of the target object's displayed content in the current frame image. The target Gaussian sphere is then projected onto the current frame image to obtain historical feature information. This involves projecting the 3D Gaussian sphere into the camera coordinate system corresponding to the current frame image, and the historical feature information corresponds to the features of the 3D Gaussian sphere projected onto the current frame image from a 2D perspective. It is conceivable that the 3D Gaussian sphere can provide parameter information such as center coordinates and radiation intensity. After projection, the parameter information carried by the target Gaussian sphere can be displayed in the 2D perspective corresponding to the current frame image, thus converting the information provided by the target Gaussian sphere into a 2D image. Furthermore, the target Gaussian sphere provides context alignment constraints for the current frame image.

[0030] Step S140: Based on the three-dimensional feature information and historical feature information, update the three-dimensional Gaussian sphere in the global library in real time.

[0031] Furthermore, by comparing and analyzing 3D feature information and historical feature information, the electronic device can optimize the global library based on the comparison analysis results and update the 3D Gaussian spheres in the global library. It is conceivable that under different shooting angles, some 3D Gaussian spheres may incorrectly represent the target object. For example, after projection, a color block may appear in front of the target object, forming a floating object, while this floating object does not exist in front of the target object in the current frame image. By comparing and analyzing 3D feature information and historical feature information, the aforementioned judgment of erroneous 3D Gaussian spheres can be achieved, thereby identifying and deleting these erroneous spheres to update the global library. Moreover, after processing all images in the image sequence, the global library is optimized, and the 3D Gaussian spheres within it can represent the target object well. Therefore, 3D reconstruction of the target object can be achieved by rendering and projecting all 3D Gaussian spheres in the global library.

[0032] As can be seen from the above scheme, the scheme of this application can dynamically update the global library by utilizing context information, thereby generating a more coordinated and consistent 3D Gaussian sphere distribution, effectively suppressing redundant structures such as blur and floating objects, and through continuous deletion and filtering of redundant Gaussian spheres, the system can significantly reduce the number of Gaussian spheres while ensuring reconstruction accuracy, thereby reducing the storage scale, helping to speed up rendering and transmission efficiency, and thus improving the visual realism of the reconstruction results.

[0033] In one embodiment, when selecting a target image for the current frame, the electronic device can employ different selection strategies based on the media type corresponding to the image sequence. The media type corresponding to the image sequence includes video stream image types and multi-view image types. To this end, by determining the media type corresponding to the image sequence and selecting images sequentially as the current frame image according to the order of multiple frames in the image sequence, the order of the current frame image in the image sequence can be determined. Optionally, the electronic device can determine the media type corresponding to the image sequence through type labels configured in the image sequence, such as pre-configuring corresponding type labels according to the media type before inputting the image sequence, where different type labels correspond to different media types.

[0034] Furthermore, when the media type corresponding to the image sequence is a video stream image type, several frames that are sequentially adjacent to the current frame image or at a preset distance are selected as target images. It is understandable that for a video stream image type image sequence, multiple frames are related. For example, in a video where the target object is captured from an initial shooting angle and the camera moves continuously to acquire images of the target object from different shooting angles, the image content displayed in two adjacent frames is more correlated. Therefore, selecting target images from other images that are sequentially close to the current frame image makes it easier to obtain the correlation between the images and the target object, thus facilitating the extraction of the target object's three-dimensional feature information.

[0035] To address this, this solution can select several frames that are sequentially adjacent to the current frame as target images. If the current frame is the first frame in the image sequence, then all frames adjacent to the current frame are considered target images. For example, the second frame can be a target image, and the third frame can also be a target image. The number of target images selected can be set based on actual application requirements. Optionally, in some embodiments, multiple frames before and after the current frame can be used as corresponding target images. For example, if the preceding frame is the second frame in the image sequence, then both the first and third frames are considered target images. Optionally, this solution can also select several frames that are sequentially separated from the current frame by a preset interval as target images. This preset interval represents the number of images in the sequence. For example, if the preset interval is set to 1 frame, then if the current frame is the first frame in the image sequence, the third frame can be selected as the target image.

[0036] Furthermore, when the media type corresponding to the image sequence is a multi-view image type, based on the pose information included in the images, several frames whose camera positions and the camera positions corresponding to the current frame image are located in a preset adjacent spatial region in the camera coordinate system are determined as target images. It is understandable that images record camera positions through pose information, and different camera positions correspond to different image shooting angles. Moreover, it is conceivable that the closer two cameras are, the stronger the correlation between the images they capture. Therefore, in the camera coordinate system, an adjacent spatial region can be determined by using the camera position corresponding to the current frame image as the center and a preset distance value as the radius. (Refer to...) Figure 2 The camera position corresponding to image pic1 is point P1, the camera position corresponding to image pic2 is point P2, and the camera position corresponding to image pic3 is point P3. If the distance between point P3 and point P1 is shorter than the distance between point P2 and point 1, and point P3 is located within the adjacent spatial region, then image pic3 is selected as the target image.

[0037] Therefore, this scheme selects the associated target image to process the current frame image, thereby introducing contextual information, which helps to make the generated 3D feature information spatially consistent and effectively avoids the problem of discontinuous or misaligned local reconstruction.

[0038] Optionally, the pose information includes the camera extrinsic matrix. The inverse of the camera extrinsic matrix is ​​called the c2w (camera-to-world) matrix, also known as the camera pose, which is used to transform points in the camera coordinate system to the world coordinate system. Figure 3 The figure illustrates the steps for determining a target image according to an embodiment of this application. By performing calculations based on the camera extrinsic parameter matrix, the coordinates of the camera position in the world coordinate system can be determined. This allows for the transformation of camera positions corresponding to different images into the same coordinate system to determine camera distances and to identify whether the corresponding image is the target image. The specific steps are as follows: Step S210: Based on the camera extrinsic matrix carried by each image, transform the camera position corresponding to the current frame image and the camera positions corresponding to other images in the image sequence from the camera coordinate system to the world coordinate system, and obtain the first camera coordinates corresponding to the current frame image and multiple second camera coordinates corresponding to different other images.

[0039] Step S220: Based on the coordinates of the first camera and multiple coordinates of the second camera, determine the camera distance value between each other image and the current frame image.

[0040] Step S230: If the camera distance value is less than or equal to a preset distance value, determine other images corresponding to the camera distance value as target images.

[0041] It is conceivable that the inverse of the camera extrinsic matrix can be used to transform points in the camera coordinate system to the world coordinate system. The c2w matrix can also be written as a combination of rotation matrices and translation vectors. Let R be the rotation matrix of c2w. C Camera coordinates P C World coordinates P W If the translation vector is C, then:

[0042] Furthermore, the values ​​of the c2w matrix directly describe the orientation and origin of the camera coordinate system. Specifically, the first to third columns of the rotation matrix represent the directions of the X, Y, and Z axes of the camera coordinate system in the world coordinate system, respectively; the translation vector represents the corresponding position of the camera origin in the world coordinate system. Thus, through coordinate transformation, the camera coordinates corresponding to different images can be determined, such as the first camera coordinates corresponding to the current frame image and multiple second camera coordinates corresponding to different other images.

[0043] Coordinate values ​​within the same coordinate system can be used to calculate corresponding distance values. Therefore, for each second camera coordinate, the camera distance between that coordinate and the first camera coordinate is calculated, and the calculated camera distance is compared with a preset distance value. This preset distance value is used to determine whether the camera coordinates corresponding to other images are located in the adjacent spatial region of the camera coordinates corresponding to the current frame image. Furthermore, if the camera distance value is less than or equal to the preset distance value, it can be determined that the camera coordinates corresponding to that image are located in the adjacent spatial region of the camera coordinates corresponding to the current frame image, and thus that image can be identified as the target image. Optionally, the selection of a target image can be achieved by calculating the camera distance value corresponding to the image and determining whether the image is a target image. Alternatively, after determining all camera distance values, the existence of a target image can be determined based on the camera distance values ​​and the preset distance value.

[0044] To address this, this solution selects a target image, thereby introducing a related image into the image sequence for the current frame image, thus introducing contextual information and better acquiring the three-dimensional feature information of the current frame image.

[0045] Optionally, the 3D feature information includes depth and color information. During feature extraction, this scheme combines the current frame image and the target image to introduce constraints from other images on the target object in the current frame image. This allows for the utilization of the target object's features from other perspectives to better obtain the corresponding 3D feature information of the target object in the current frame image. Specifically, a planar scanning algorithm projects the target image onto the current frame image to obtain the depth information of the target object on the current frame image. Understandably, the planar scanning method matches a reference image by projecting a set of images onto a plane and then onto a reference image. This step involves comparing the warped image with the reference image to measure their dissimilarity, such as through a matching window. If the plane is close to the true depth of a pixel in the reference image, the corresponding difference value will be low. Therefore, by testing multiple planes, the depth generated by the best matching plane for each pixel is selected, and then a depth map of the reference image is generated. Based on this, the planar scanning algorithm can determine the depth information corresponding to the target object in the current frame image.

[0046] Furthermore, based on the RGB color values ​​of pixels at the same location in different images, the color information of the target object in the current frame is determined. It's conceivable that each pixel in different images corresponds to a specific RGB color value. By determining the distribution of the same location of the target object in different images, such as by referring to... Figure 2 The current frame image is image pic1, whose image content corresponds to surface A of the target object. Image pic3, on the other hand, corresponds to surfaces A, B, and C of the target object. Image pic3, as the target image, shares surface A with image pic1. Therefore, taking surface A as the overlapping area, the RGB color values ​​of each pixel within that area can be determined in both images pic1 and pic3. Optionally, the RGB color values ​​of pixels at the same location can be determined in the current frame image by weighted averaging, thereby fusing the target images and achieving contextual fusion to more accurately represent the three-dimensional feature information of the current frame image.

[0047] In one embodiment, this solution compares and analyzes 3D feature information and historical feature information, and then updates the global library based on the comparison and analysis results, such as... Figure 4 As shown, Figure 4 This application provides a schematic diagram illustrating the steps for comparing and analyzing three-dimensional feature information and historical feature information according to an embodiment of the present application. The specific steps include: Step S310: Compare the depth information carried in the 3D feature information and the historical feature information, and update the global library if it is determined that the target Gaussian sphere renders the target object in an incorrect state.

[0048] Step S320: Compare the color information carried in the 3D feature information and the historical feature information, and update the global library if it is determined that the target Gaussian sphere renders the target object in an incorrect state.

[0049] Understandably, by comparing depth information, it's possible to determine whether the projection of the target Gaussian sphere onto the viewpoint of the current frame image contains blur, repetition, or floating objects, thus determining whether the rendering of the target object by the Gaussian sphere is erroneous. Furthermore, by comparing color information, it's possible to determine whether the projection of the target Gaussian sphere onto the viewpoint of the current frame image contains color errors or color variations, thus determining whether the rendering of the target object by the Gaussian sphere is erroneous. Therefore, through comparative analysis of different parameters, redundant or low-quality Gaussian spheres in the library are identified, such as those with blurred, repetitive, or floating objects after projection, to optimize the global library, such as by deleting redundant Gaussian spheres and adding new ones.

[0050] Optionally, during the comparison of depth information, based on the depth information carried by the 3D feature information, a first depth feature extracted from the target object in the current frame image is determined, and based on the depth information carried by the historical feature information, a second depth feature is determined by projecting the target Gaussian sphere onto the current frame image. That is, extracting the first and second depth features based on the depth information, it is conceivable that the first and second depth features include the depth at each corresponding pixel point in the image.

[0051] Furthermore, by projecting the target Gaussian sphere onto the 2D viewpoint corresponding to the current frame image, the extracted second depth feature can determine the depth of each pixel based on the same camera position as the first depth feature. Then, the first and second depth features are compared to determine the comparison result at the same position in the current frame image. Specifically, if the depth corresponding to the second depth feature is less than the depth corresponding to the first depth feature, the rendering of the target Gaussian sphere is determined to be incorrect, and the target Gaussian sphere is deleted from the global library. It can be understood that in depth comparison, if there is a case where the depth corresponding to the second depth feature is less than the depth corresponding to the first depth feature, it means that after the target Gaussian sphere is rendered, there will be corresponding floating objects obstructing the target object, thus the target Gaussian sphere can be considered to be in an incorrect state, and therefore deleted from the global library.

[0052] Furthermore, if the second depth feature is empty, the rendering of the target Gaussian sphere for the target object is determined to be in an incorrect state, and a 3D Gaussian sphere generated based on the 3D feature information is added to the global library. It is understandable that in depth comparison, if the second depth feature is empty, it means that the region is empty after the target Gaussian sphere is rendered, i.e., the region is not rendered. In this case, the target Gaussian sphere can also be considered in an incorrect state, but it is still retained, and a 3D Gaussian sphere generated based on the 3D feature information of the current frame image is added to the global library. It is conceivable that in some embodiments, if the historical feature information obtained through the target Gaussian sphere is the same as the 3D feature information corresponding to the current frame image, it can be determined that there is a redundant Gaussian sphere. Therefore, the target Gaussian sphere can be deleted, and a 3D Gaussian sphere generated based on the 3D feature information is added to the global library, thus representing the target object with the latest 3D Gaussian sphere.

[0053] Optionally, during the comparison of color information, based on the color information carried by the 3D feature information, a first color feature extracted from the target object in the current frame image is determined, and based on the color information carried by the historical feature information, a second color feature is determined by projecting the target Gaussian sphere onto the current frame image. That is, extracting the first and second color features based on color information can be understood as including the RGB color values ​​of corresponding pixels in the image. Furthermore, by projecting the target Gaussian sphere onto the 2D viewpoint corresponding to the current frame image, the extracted second color feature can determine the RGB color values ​​of corresponding pixels based on the same camera position as the first color feature. Then, the first and second color features are compared to determine the comparison result at the same position in the current frame image. Specifically, if the RGB color value corresponding to the second color feature is not equal to the RGB color value corresponding to the first color feature, the rendering of the target Gaussian sphere of the target object is determined to be in an incorrect state, and the target Gaussian sphere is deleted from the global library. It is understandable that in color comparison, if there are unequal RGB color values, it means that the rendering result of the target Gaussian sphere will be inconsistent with the current frame image. Therefore, the target Gaussian sphere can be considered to be in an incorrect state, and thus the target Gaussian sphere will be deleted from the global library.

[0054] Furthermore, if the second color feature is empty, the rendering of the target Gaussian sphere on the target object is determined to be in an incorrect state, and a 3D Gaussian sphere generated based on the 3D feature information is added to the global library. It is understandable that in color contrast, if the second color feature is empty, it means that the area is empty after the target Gaussian sphere is rendered, i.e., no rendering is performed on that area. In this case, the target Gaussian sphere can also be considered in an incorrect state, but it is still retained, and a 3D Gaussian sphere generated based on the 3D feature information of the current frame image is added to the global library.

[0055] Therefore, this scheme achieves real-time updates of the global library by comparing and analyzing the newly extracted 3D feature information with historical feature information. Through this mechanism, the reconstruction results remain compact and efficient, which not only reduces unnecessary computational overhead but also effectively reduces the burden of data storage and rendering.

[0056] Figure 5 This is a schematic diagram of the structure of a self-updating streaming 3D Gaussian reconstruction device provided in an embodiment of this application. The device is used to execute the self-updating streaming 3D Gaussian reconstruction method provided in the above embodiment, and has the corresponding functional modules and beneficial effects of executing the method. As shown in the figure, the device includes an image selection module 401, a feature extraction module 402, a Gaussian flattening module 403, and a data update module 404.

[0057] The image selection module 401 is configured to, in response to the three-dimensional reconstruction operation of the target object in the multi-frame images of the image sequence, acquire the current frame image and determine the target image associated with the current frame image in the multi-frame images of the image sequence, wherein the multi-frame images in the image sequence are images of the target object under any shooting angle. The feature extraction module 402 is configured to extract features from the current frame image based on the current frame image and the target image, and obtain the three-dimensional feature information of the target object in the current frame image; The Gaussian flattening module 403 is configured to determine the target Gaussian sphere from the three-dimensional Gaussian spheres stored in the preset global library, and project the target Gaussian sphere onto the current frame image to obtain historical feature information. The target Gaussian sphere is associated with the three-dimensional feature information of the target object's display content on the current frame image. The data update module 404 is configured to update the three-dimensional Gaussian sphere in the global library in real time based on three-dimensional feature information and historical feature information.

[0058] Based on the above embodiments, the image selection module 401 is specifically configured as follows: Determine the media type corresponding to the image sequence, and select the images as the current frame image in the order of the multiple frames in the image sequence; When the media type corresponding to the image sequence is a video stream image type, select several frames that are adjacent to the current frame image in order or at a preset distance as the target image; When the media type corresponding to the image sequence is a multi-view image type, based on the pose information included in the image, several frames of images whose camera positions are located in a preset adjacent spatial region in the camera coordinate system with the camera positions corresponding to the current frame image are determined as target images.

[0059] Based on the above embodiments, the pose information includes the camera extrinsic parameter matrix, and the image selection module 401 is further configured as follows: Based on the camera extrinsic matrix carried by each image, the camera position corresponding to the current frame image and the camera positions corresponding to other images in the image sequence are transformed from the camera coordinate system to the world coordinate system, and the first camera coordinates corresponding to the current frame image and multiple second camera coordinates corresponding to different other images are obtained. Based on the coordinates of the first camera and multiple second camera coordinates, determine the camera distance value between each other image and the current frame image; If the camera distance value is less than or equal to a preset distance value, other images corresponding to the camera distance value are selected as the target image.

[0060] Based on the above embodiments, the three-dimensional feature information includes depth information and color information, and the feature extraction module 402 is specifically configured as follows: The target image is projected onto the current frame image using a planar scanning algorithm to obtain the depth information of the target object in the current frame image. The color information of the target object in the current frame image is determined based on the RGB color values ​​of the pixels at the same location in different images.

[0061] Based on the above embodiments, the data update module 404 is specifically configured as follows: The depth information carried in the 3D feature information and the historical feature information is compared, and the global library is updated when it is determined that the target Gaussian sphere renders the target object in an incorrect state. The color information carried in the 3D feature information and historical feature information is compared, and the global library is updated if it is determined that the target Gaussian sphere renders the target object in an incorrect state.

[0062] Based on the above embodiments, the data update module 404 is further configured as follows: Based on the depth information carried by the three-dimensional feature information, the first depth feature extracted from the target object in the current frame image is determined; Based on the depth information carried by historical feature information, the second depth feature obtained by projecting the target Gaussian sphere onto the current frame image is determined; The first depth feature and the second depth feature are compared to determine the comparison result at the same location in the current frame image; If the depth corresponding to the second depth feature is less than the depth corresponding to the first depth feature, the target Gaussian sphere is determined to be rendering the target object as an error, and the target Gaussian sphere is deleted from the global library. If the second depth feature is empty, the target Gaussian sphere is determined to be rendering the target object as an error, and a three-dimensional Gaussian sphere generated based on the three-dimensional feature information is added to the global library.

[0063] Based on the above embodiments, the data update module 404 is further configured as follows: Based on the color information carried by the three-dimensional feature information, the first color feature extracted from the target object in the current frame image is determined; Based on the color information carried by historical feature information, the second color feature obtained by projecting the target Gaussian sphere onto the current frame image is determined; The first color feature and the second color feature are compared to determine the comparison result at the same position in the current frame image; If the RGB color value corresponding to the second color feature is not equal to the RGB color value corresponding to the first color feature, the target Gaussian sphere is determined to be rendering the target object as an error, and the target Gaussian sphere is deleted from the global library. If the second color feature is empty, determine that the target Gaussian sphere renders the target object in an incorrect state, and add a three-dimensional Gaussian sphere generated based on the three-dimensional feature information to the global library.

[0064] It is worth noting that in the embodiments of the above-mentioned device, the modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each module are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.

[0065] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of this application. The device is used to execute the self-updating streaming 3D Gaussian reconstruction method provided in the above embodiment, and has corresponding functional modules and beneficial effects for executing the method. As shown in the figure, the device includes a processor 501, a memory 502, an input device 503, and an output device 504. The number of processors 501 can be one or more; one processor 501 is shown as an example in the figure. The processor 501, memory 502, input device 503, and output device 504 can be connected via a bus or other means; a bus connection is shown as an example in the figure. The memory 502, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the self-updating streaming 3D Gaussian reconstruction method in the embodiments of this application. The processor 501 executes various corresponding functional applications and data processing by running the software programs, instructions, and modules stored in the memory 502, thereby realizing the above-mentioned self-updating streaming 3D Gaussian reconstruction method.

[0066] The memory 502 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data recorded or created during use. Furthermore, the memory 502 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 502 may further include memory remotely configured relative to the processor 501, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0067] The input device 503 can be used to input corresponding digital or character information to the processor 501, and to generate key signal inputs related to the user settings and function control of the device; the output device 504 can be used to send or display key signal outputs related to the user settings and function control of the device.

[0068] This application also provides a storage medium storing computer-executable instructions, which, when executed by a processor, are used to perform relevant operations in the self-updating streaming 3D Gaussian reconstruction method provided in any embodiment of this application.

[0069] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0070] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0071] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

Claims

1. A self-updating streaming 3D Gaussian reconstruction method, characterized in that, include: In response to a three-dimensional reconstruction operation of a target object in a multi-frame image of an image sequence, the current frame image is acquired and a target image associated with the current frame image is determined in the multi-frame image of the image sequence, wherein the multi-frame image of the image sequence is an image of the target object at any shooting angle. Based on the current frame image and the target image, feature extraction is performed on the current frame image to obtain the three-dimensional feature information of the target object on the current frame image; The target Gaussian sphere is determined by storing a three-dimensional Gaussian sphere in a preset global library, and the target Gaussian sphere is projected onto the current frame image to obtain historical feature information. The target Gaussian sphere is associated with the three-dimensional feature information of the target object's display content on the current frame image. Based on the three-dimensional feature information and the historical feature information, the three-dimensional Gaussian sphere in the global library is updated in real time.

2. The self-updating streaming 3D Gaussian reconstruction method according to claim 1, characterized in that, The step of responding to a 3D reconstruction operation of a target object in a multi-frame image sequence, acquiring the current frame image and determining the target image associated with the current frame image in the multi-frame image sequence, includes: Determine the media type corresponding to the image sequence, and select the images as the current frame image in the order of the multiple frames in the image sequence. When the media type corresponding to the image sequence is a video stream image type, select several frames that are adjacent in order to the current frame image or at a preset distance as the target image; When the media type corresponding to the image sequence is a multi-view image type, based on the pose information included in the image, several frames whose camera positions are located in a preset adjacent spatial region in the camera coordinate system with the camera positions corresponding to the current frame image are determined as the target image.

3. The self-updating streaming 3D Gaussian reconstruction method according to claim 2, characterized in that, The pose information includes a camera extrinsic parameter matrix. The step of determining a number of frames whose camera positions are located within a preset adjacent spatial region based on the pose information included in the image, and using these frames as the target image, includes: Based on the camera extrinsic matrix carried by each image, the camera position corresponding to the current frame image and the camera positions corresponding to other images in the image sequence are transformed from the camera coordinate system to the world coordinate system, and the first camera coordinates corresponding to the current frame image and multiple second camera coordinates corresponding to different other images are obtained. Based on the first camera coordinates and multiple second camera coordinates, determine the camera distance value between each other image and the current frame image; If the camera distance value is less than or equal to a preset distance value, other images corresponding to the camera distance value are determined as the target image.

4. The self-updating streaming 3D Gaussian reconstruction method according to claim 1, characterized in that, The three-dimensional feature information includes depth information and color information. The step of extracting features from the current frame image and obtaining the three-dimensional feature information of the target object in the current frame image based on the current frame image and the target image includes: The target image is projected onto the current frame image using a planar scanning algorithm to obtain the depth information of the target object on the current frame image. The color information of the target object in the current frame image is determined based on the RGB color values ​​of the pixels at the same location in different images.

5. The self-updating streaming 3D Gaussian reconstruction method according to claim 1, characterized in that, The step of updating the three-dimensional Gaussian sphere in the global library in real time based on the three-dimensional feature information and the historical feature information includes: The depth information carried in the three-dimensional feature information and the historical feature information is compared, and the global library is updated if it is determined that the target Gaussian sphere renders the target object in an incorrect state. The color information carried in the three-dimensional feature information and the historical feature information is compared, and the global library is updated if it is determined that the target Gaussian sphere renders the target object in an incorrect state.

6. The self-updating streaming 3D Gaussian reconstruction method according to claim 5, characterized in that, The step of comparing the depth information carried in the three-dimensional feature information and the historical feature information, and updating the global library when it is determined that the target Gaussian sphere renders the target object in an incorrect state, includes: Based on the depth information carried by the three-dimensional feature information, the first depth feature extracted from the target object in the current frame image is determined; Based on the depth information carried by the historical feature information, a second depth feature is determined by projecting the target Gaussian sphere onto the current frame image; The first depth feature and the second depth feature are compared to determine the comparison result at the same position in the current frame image; If the depth corresponding to the second depth feature is less than the depth corresponding to the first depth feature, it is determined that the target Gaussian sphere renders the target object in an incorrect state, and the target Gaussian sphere is deleted from the global library. If the second depth feature is empty, the target Gaussian sphere is determined to be rendering the target object as an error, and a three-dimensional Gaussian sphere generated based on the three-dimensional feature information is added to the global library.

7. The self-updating streaming 3D Gaussian reconstruction method according to claim 5 or 6, characterized in that, The step of comparing the color information carried in the three-dimensional feature information and the historical feature information, and updating the global library when it is determined that the target Gaussian sphere renders the target object in an incorrect state, includes: Based on the color information carried by the three-dimensional feature information, the first color feature extracted from the target object in the current frame image is determined; Based on the color information carried by the historical feature information, the second color feature obtained by projecting the target Gaussian sphere onto the current frame image is determined; The first color feature and the second color feature are compared to determine the comparison result at the same position in the current frame image; If the RGB color value corresponding to the second color feature is not equal to the RGB color value corresponding to the first color feature, it is determined that the target Gaussian sphere renders the target object in an error state, and the target Gaussian sphere is deleted from the global library. If the second color feature is empty, the target Gaussian sphere is determined to be rendering the target object in an incorrect state, and a three-dimensional Gaussian sphere generated based on the three-dimensional feature information is added to the global library.

8. A self-updating streaming 3D Gaussian reconstruction device, characterized in that, include: The image selection module is configured to, in response to a three-dimensional reconstruction operation of a target object in a multi-frame image of an image sequence, acquire the current frame image and determine the target image associated with the current frame image in the multi-frame image of the image sequence, wherein the multi-frame images in the image sequence are images of the target object at any shooting angle. The feature extraction module is configured to extract features from the current frame image based on the current frame image and the target image, and obtain the three-dimensional feature information of the target object on the current frame image; The Gaussian flattening module is configured to determine a target Gaussian sphere from a three-dimensional Gaussian sphere stored in a preset global library, and project the target Gaussian sphere onto the current frame image to obtain historical feature information. The target Gaussian sphere is associated with the three-dimensional feature information of the target object's display content on the current frame image. The data update module is configured to update the three-dimensional Gaussian sphere in the global library in real time based on the three-dimensional feature information and the historical feature information.

9. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the self-updating streaming 3D Gaussian reconstruction method as described in any one of claims 1-7.

10. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a processor, are used to perform the self-updating streaming 3D Gaussian reconstruction method as described in any one of claims 1-7.