A Spatial Intelligent 3D Video Generation Method and System with Spatial Adaptive Noise Injection
By constructing a three-dimensional Gaussian space representation and an adaptive noise injection method, the problem of uneven quality in 3D video generation under sparse view conditions is solved, and higher quality and more efficient 3D video generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING FEIDU TECH CO LTD
- Filing Date
- 2026-05-07
- Publication Date
- 2026-06-02
AI Technical Summary
Under sparse view conditions, differences in geometric reliability and appearance consistency in different spatial regions during 3D video generation lead to uneven quality of the generated results, with insufficient repair of local areas or excessive destruction of correct areas, affecting the overall quality and generation efficiency.
By constructing a three-dimensional Gaussian spatial representation, analyzing the reconstruction quality, generating a spatial confidence target map, generating an adaptive diffusion time step field based on the confidence, performing spatially non-uniform noise injection and denoising processing, and generating a target three-dimensional video.
It improves the overall generation quality and local structural accuracy of 3D video under sparse view conditions, and increases generation efficiency.
Smart Images

Figure CN122138026A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of spatial intelligence and 3D generative artificial intelligence, and in particular to a spatial intelligent 3D video generation method and system with spatial adaptive noise injection. Background Technology
[0002] With the increasing demand for 3D digital content generation, technologies such as 3D scene modeling, neural rendering, and generative diffusion models have become current research hotspots and have made significant progress.
[0003] In 3D video generation tasks under sparse view conditions, a uniform diffusion noise injection strategy is typically employed, applying noise perturbation of the same intensity to all spatial regions of the 3D scene. However, different spatial regions in 3D videos exhibit significant differences in geometric reliability and appearance consistency: some regions are fully observed from multiple perspectives, resulting in high reconstruction quality; while occluded regions and viewpoint blind spots show obvious geometric distortion or texture mismatch. Using a globally uniform noise injection method ignores these significant differences in geometric reliability and appearance consistency across different spatial regions in the 3D video, leading to spatially uneven quality in the generated results. This results in insufficient restoration of local areas or excessive destruction of correct areas, thus affecting the overall quality and generation efficiency of the 3D video.
[0004] Therefore, how to improve the overall quality of 3D video generation under sparse view conditions is a key research topic for those skilled in the art. Summary of the Invention
[0005] In a first aspect, this application provides a spatially adaptive noise-injected intelligent 3D video generation method, characterized in that the method includes: constructing a 3D Gaussian space representation based on at least two input images, and rendering a 3D video prior sequence based on the 3D Gaussian space representation and a preset camera motion trajectory; analyzing the reconstruction quality of the 3D video prior sequence in the spatial dimension, obtaining the values of at least two quality indicators corresponding to a preset sub-region, and obtaining a quality evaluation result corresponding to the 3D video prior sequence in time and space, wherein the preset sub-region refers to any sub-region obtained by regularly dividing any 2D image in the 3D video prior sequence; converting the at least two quality indicators corresponding to each preset sub-region in the quality evaluation result into a single comprehensive confidence score, and generating a spatial confidence target map, wherein... Confidence level is positively correlated with image quality. Based on the spatial confidence target map, a spatially adaptive target diffusion time step field is generated. The target diffusion time step field includes a diffusion time step corresponding to each preset sub-region. The diffusion time step is used to determine the noise perturbation amplitude applied to the corresponding preset sub-region of the 3D video prior sequence. The diffusion time step is negatively correlated with confidence level. Based on the target diffusion time step field, spatially non-uniform noise injection is performed on the 3D video prior sequence to obtain a noise-perturbed 3D video prior sequence. A local redraw mask is generated based on the target diffusion time step field to indicate the degree of participation of each preset sub-region in the denoising process. The noise-perturbed 3D video prior sequence, the local redraw mask, and the target diffusion time step field are input into a single-step diffusion generation network for single-step denoising to generate the target 3D video.
[0006] The spatial intelligent 3D video generation method with spatial adaptive noise injection provided in this application introduces a spatial confidence level reflecting image quality during the diffusion generation process. Based on the confidence level of reconstruction quality, noise injection is controlled to achieve differentiated generation control of different spatial regions in the 3D video, thereby improving the overall generation quality, local structural accuracy and generation efficiency of the 3D video under sparse view conditions.
[0007] Secondly, this application also provides a spatially adaptive noise-injected intelligent 3D video generation system, including: sparse view Figure 3A three-dimensional prior generation module is used to construct a three-dimensional Gaussian space representation based on at least two input images, and render a three-dimensional video prior sequence based on the three-dimensional Gaussian space representation and a preset camera motion trajectory. A three-dimensional video spatial quality assessment module is used to analyze the reconstruction quality of the three-dimensional video prior sequence in the spatial dimension, obtain the values of at least two quality indicators corresponding to a preset sub-region, and obtain a quality assessment result corresponding to the three-dimensional video prior sequence in time and space. The preset sub-region refers to any sub-region obtained after regularly dividing any two-dimensional image in the three-dimensional video prior sequence. A spatial confidence map construction module is used to convert the at least two quality indicators corresponding to each preset sub-region in the quality assessment result into a single comprehensive confidence score, and generate a spatial confidence target map, wherein the confidence score is positively correlated with image quality. A spatial adaptive diffusion time step field generation module is used to generate a spatial confidence target map based on the spatial confidence target map. The diagram illustrates the generation of a spatially adaptive target diffusion time step field, which includes a diffusion time step corresponding to each preset sub-region. The diffusion time step determines the noise perturbation amplitude applied to the corresponding preset sub-region of the 3D video prior sequence, wherein the diffusion time step is negatively correlated with the confidence level. A local redraw mask and spatial noise injection module is used to perform spatially non-uniform noise injection on the 3D video prior sequence based on the target diffusion time step field, obtaining a noise-perturbed 3D video prior sequence. This module also generates a local redraw mask based on the target diffusion time step field, indicating the degree of participation of each preset sub-region in the denoising process. A single-step diffusion 3D video generation module inputs the noise-perturbed 3D video prior sequence, the local redraw mask, and the target diffusion time step field into a single-step diffusion generation network for single-step denoising, generating the target 3D video.
[0008] Thirdly, this application also provides a spatially adaptive noise injection-based intelligent three-dimensional video generation apparatus, including a unit for performing any spatially adaptive noise injection-based intelligent three-dimensional video generation method in the first aspect.
[0009] Fourthly, this application also provides a computer storage medium that can store multiple instructions adapted for loading and execution by a processor of the spatial intelligent 3D video generation method with arbitrary spatial adaptive noise injection in the first aspect.
[0010] Fifthly, embodiments of this application also provide a computer program product containing instructions that, when the computer program product is run on an electronic device, cause the electronic device to execute the spatial intelligent three-dimensional video generation method with arbitrary spatial adaptive noise injection as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application also provide a chip module, including a transceiver component and a chip, wherein the chip is used to execute the spatial intelligent three-dimensional video generation method with arbitrary spatial adaptive noise injection as described in the first aspect.
[0012] In a seventh aspect, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method described in any one of the first aspects is performed when the processor executes the program.
[0013] It is understood that the spatially adaptive noise injection spatial intelligent 3D video generation system, apparatus, computer storage medium, computer program, computer program product, chip system, and electronic device provided above are all used to execute the method shown in any implementation of the first aspect of the embodiments of this application. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0014] Figure 1 This is a schematic flowchart of the spatial intelligent 3D video generation method with spatial adaptive noise injection provided in the embodiments of this application; Figure 2 This is a schematic diagram of the method for constructing 3D video priors based on input images provided in an embodiment of this application; Figure 3 This is a schematic flowchart of a method for converting quality assessment results into confidence maps, provided in an embodiment of this application. Figure 4 This is a schematic flowchart of a method for determining the target diffusion time step field based on a spatial confidence target map, provided in an embodiment of this application. Figure 5 This is a schematic diagram of the distillation optimization and model update steps provided in the embodiments of this application; Figure 6 This is a schematic diagram of the architecture of the spatial intelligent 3D video generation system with spatial adaptive noise injection provided in the embodiments of this application. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described below in conjunction with the accompanying drawings.
[0016] It should be noted that this application embodiment uses an electronic device as an example to illustrate the execution subject of the spatial intelligent 3D video generation method with spatial adaptive noise injection provided in this application. This electronic device can also be understood as the spatial intelligent 3D video generation device with spatial adaptive noise injection shown in this application embodiment. In this application embodiment, the electronic device can be a microprocessor or computer for executing program code, etc. Any electronic device that can be used to execute the method provided in this application embodiment is within the protection scope of this application embodiment, and this application does not impose any limitations. For example, the electronic device can be a desktop computer, a laptop, a mobile terminal, a 32-bit microprocessor, or a 64-bit microprocessor, etc., and this application embodiment does not limit this.
[0017] Please see Figure 1 , Figure 1 A flowchart illustrating the spatially adaptive noise injection method for generating intelligent 3D video in accordance with embodiments of this application. Figure 1 As shown, the spatially adaptive noise-injected intelligent 3D video generation method includes the following steps: S101, the electronic device constructs a three-dimensional Gaussian space representation based on at least two input images, and renders a three-dimensional video prior sequence based on the above three-dimensional Gaussian space representation and a preset camera motion trajectory.
[0018] Specifically, the electronic device receives at least two input images from the user and obtains the camera parameters corresponding to the input images, and constructs a three-dimensional Gaussian space representation based on these parameters.
[0019] The at least two input images are two-dimensional images of the same static scene captured from different viewpoints, and can originate from real camera shots, synthetic data, or rendering results from a simulated environment. The at least two input images are allowed to have one or a combination of the following: significant difference in viewpoint baseline, obvious occlusion in local areas, changes in object scale or viewpoint distortion, or inconsistent lighting or color distribution.
[0020] The camera parameters of the input image include intrinsic and extrinsic parameters. Intrinsic parameters include, but are not limited to, focal length and principal point coordinates; extrinsic parameters are the camera pose matrix (containing a rotation matrix reflecting the viewpoint direction and a translation vector reflecting the spatial position). Electronic devices can obtain the camera parameters of the input image through offline calibration or online calculation using motion reconstruction algorithms.
[0021] For some possible implementations, please refer to Figure 2 The above step S101 (the electronic device constructs a three-dimensional Gaussian space representation based on at least two input images) specifically includes: S1011, obtain sparse map and camera parameters.
[0022] That is, the electronic device acquires at least two input images (sparse graphs) and camera parameters for each input image.
[0023] S1012, Gaussian initialization.
[0024] That is, the electronic device initializes a set of three-dimensional Gaussian elements in three-dimensional space based on the above at least two input images and the camera parameters of each input image, and obtains a three-dimensional Gaussian field.
[0025] As an example, the properties of this Gaussian element can include the coordinates of its three-dimensional spatial center, anisotropic scaling parameters (covariance matrix), spatial rotation parameters, and color and transparency information. The anisotropic scaling parameters represent the radius or scaling degree of the Gaussian element along its three principal axes (X, Y, Z) in its local coordinate system, while the spatial orientation or rotation parameters represent the orientation of the Gaussian element's local coordinate system relative to the world coordinate system.
[0026] For example, an electronic device initializes a set of Gaussian primitives in 3D space based on an input image and its camera parameters. Typically, an initial sparse point cloud of the scene is first generated using a motion reconstruction structure or depth estimation method. Each 3D point and its average color in each input image are used as the initial center position and initial color of the Gaussian primitive, respectively. The spatial distribution of each Gaussian primitive is described by a 3D Gaussian function, whose covariance matrix is initialized to an isotropic small value, i.e., the scale parameter along the three principal axes is set to a preset small constant, the rotation parameter is initialized to a unit quaternion, and the opacity is initialized to an intermediate value (e.g., 0.5). Thus, the system obtains a set of initialized 3D Gaussian primitives, where each primitive contains optimizable attributes such as position, covariance matrix, color, and opacity.
[0027] S1013, multi-view consistency optimization.
[0028] In other words, the electronic device acquires projected images of a 3D Gaussian field under different input viewpoints, and then aligns these projected images with the corresponding input images under their respective viewpoints to optimize the spatial position and appearance attributes of the Gaussian primitives, thus obtaining a 3D Gaussian spatial representation. Here, the input viewpoint refers to the known camera viewpoint corresponding to the camera parameters of the input image.
[0029] That is, the three-dimensional Gaussian primitives are projected onto the image plane of the known camera viewpoint corresponding to each input image, and are aligned with the corresponding real input image (i.e., the reconstruction error between the rendered image and the real image is calculated), thereby backpropagating to optimize the spatial position and appearance attributes of the Gaussian primitives.
[0030] For example, the electronic device iteratively optimizes and adjusts the parameters of the 3D Gaussian primitives through a multi-view error joint optimization mechanism. Specifically, the electronic device first uses the current 3D Gaussian primitive set to perform a differentiable rendering process under the camera parameter conditions corresponding to each input viewpoint, thereby obtaining a synthesized image under each viewpoint. The synthesized image represents the projection result of the current 3D Gaussian field under the corresponding viewpoint. Subsequently, the electronic device compares the synthesized image with the input image of the corresponding viewpoint to calculate the reconstruction error under that viewpoint. This error can be calculated through pixel-level color difference, structural similarity measurement, or feature space difference. After obtaining the reconstruction error of each viewpoint, the electronic device summarizes the errors of all viewpoints to form an overall consistency loss function. This loss function reflects the overall degree of difference between the current 3D Gaussian representation and the real image under all observation views. When a certain Gaussian primitive produces a large error under multiple viewpoints, it indicates that the spatial position, scale, or color parameters of the Gaussian primitive may be deviated and need to be adjusted. On this basis, the attributes of the 3D Gaussian primitives are optimized through a parameter update mechanism to obtain a 3D Gaussian spatial representation.
[0031] For example, by using differentiable rendering and gradient descent optimization, the spatial location and appearance attributes of Gaussian primitives are iteratively optimized to align them with all input images. Specifically, the differences between the synthetic and real images are calculated (e.g., a weighted sum of photometric and structural similarity losses). Then, using backpropagation, the gradient value of the total loss across multiple input images with respect to the optimizable attribute parameters of each Gaussian primitive is calculated. An adaptive optimization algorithm is then used to update the attributes of all Gaussian primitives based on this gradient value. During the optimization process, operations such as cloning (splitting primitives that cover high-gradient regions of the image and are too large), pruning (removing primitives whose opacity is consistently below a threshold), and resetting (resetting primitives with anisotropic scales exceeding a threshold to their initial small scale) are periodically performed to improve the representational power of the Gaussian field. This process continues until the loss converges or a preset number of iterations is reached, resulting in a coarse 3D Gaussian scene representation that is initially aligned geometrically and visually with the sparse input images.
[0032] Understandably, the multi-view consistency optimization process can ensure high geometric accuracy in some observable areas under sparse view conditions.
[0033] In this embodiment, before acquiring the prior sequence of 3D video, the electronic device needs to acquire parameters related to the 3D video generation target. These parameters include, but are not limited to, preset camera motion trajectories (translation, rotation, or a combination of motion), the number of output video frames, and the video spatial resolution and temporal sampling density. Specifically, the electronic device can determine these parameters related to the 3D video generation target based on user requirements or the default design of the upper-layer application.
[0034] Refer again Figure 2 The electronic device renders a 3D video prior sequence based on a 3D Gaussian space representation and a preset camera motion trajectory, specifically including: S1014 performs temporal dimension expansion and prior video rendering to obtain a coarse 3D video prior sequence.
[0035] That is, based on the aforementioned three-dimensional Gaussian space representation, preset camera motion trajectory, output video frame count, and video spatial resolution and temporal sampling density, the electronic device renders the three-dimensional Gaussian field (three-dimensional Gaussian space representation) frame by frame, generating a three-dimensional video prior sequence with temporal continuity (in... Figure 2 (This is illustrated in the context of "rough 3D video priors").
[0036] In other words, based on the 3D Gaussian space representation, the 3D Gaussian field is extended in time dimension. By interpolating the camera trajectory or implicit temporal parameters, a coarse but temporally continuous 3D video sequence is rendered. This video serves as the initial prior video (coarse 3D video prior) for subsequent diffusion generation, and it may possess one or more of the following characteristics: relatively stable background structure, significant geometric or texture errors in occluded and under-observed regions, and overall temporal continuity but uneven detail quality. Structural stability refers to the high consistency between the reprojection results of a region and the input image across multiple input viewpoints, with relatively small reconstruction errors, and the projection positions satisfying geometric consistency constraints at different viewpoints. Occluded regions refer to spatial locations that are obscured by other objects in the scene from certain viewpoints, resulting in incomplete observation information. Under-observed regions refer to spatial areas lacking sufficient observation information due to a limited number of input viewpoints or insufficient viewpoint coverage. Overall temporal continuity but uneven detail quality means that while the overall scene structure changes continuously between adjacent frames, some regions may exhibit detail instability due to insufficient observation or reconstruction errors.
[0037] It should be noted that the application scenario of this application is the generation of scenarios using 3D vision and video generation technology, specifically a scenario that utilizes 3D scene representation to generate dynamic videos. The aforementioned 3D video prior sequence (the meaning of 3D in the target 3D video below is the same) is visually presented as a 2D image sequence. The prefix "3D" is used to emphasize its source and connotation, rather than its visual form. Each frame of the 2D image in the 3D video is not captured by a real-world camera, but rendered from a 3D model represented in 3D Gaussian space. Visually, it presents as a 2D image sequence, but contains dynamic content with complete 3D spatial information. Each pixel in the image corresponds to a point in 3D space.
[0038] S102, the electronic device analyzes the reconstruction quality of the above-mentioned three-dimensional video prior sequence in the spatial dimension, obtains the values of at least two quality indicators corresponding to the preset sub-region, and obtains the quality evaluation results corresponding to the above-mentioned three-dimensional video prior sequence in time and space.
[0039] In this embodiment, the electronic device analyzes the reconstruction quality of each two-dimensional image in the three-dimensional video prior sequence in the spatial dimension, based on the aforementioned three-dimensional video prior sequence and the aforementioned three-dimensional Gaussian space representation data. The obtained quality assessment result corresponds to the three-dimensional video prior sequence in time and space. It can be understood that the quality assessment result includes quality assessment data corresponding to each two-dimensional image in the three-dimensional video prior sequence, and the quality assessment data includes multiple quality indicators for a preset sub-region. The evaluation process can be implemented through a trainable lightweight network or through a regularized geometric consistency metric.
[0040] In this embodiment of the application, the preset sub-region refers to any sub-region obtained by regularly dividing any two-dimensional image in the three-dimensional video prior sequence, and the preset sub-region contains at least two pixels.
[0041] As an example, multiple quality metrics for a single preset sub-region include at least two of the following: The multi-view projection error value is determined by statistical aggregation of the individual error values of all Gaussian points projected onto the preset sub-region. The error value of a single Gaussian point is used to represent a quantitative index of the geometric and appearance consistency of the Gaussian point under different observation views in the three-dimensional Gaussian space representation. The different observation views include: the real camera view corresponding to the above-mentioned at least two input images used to construct the three-dimensional Gaussian space representation, and / or at least one adjacent virtual view in the above-mentioned preset camera motion trajectory.
[0042] As an example, for any target Gaussian point in the three-dimensional Gaussian space representation, its three-dimensional center coordinates are first obtained, and then the three-dimensional center coordinates are projected and transformed onto the two-dimensional pixel plane of each observation view using the camera intrinsic and extrinsic parameter matrices corresponding to different observation views, to obtain a set of projected pixel positions that have a corresponding relationship in spatial geometry.
[0043] It should be noted that when the above observation viewpoint is at least one virtual viewpoint corresponding to adjacent virtual rendering frames on the preset camera motion trajectory, since the camera parameters of these virtual frames are completely known and the image content has been generated by the renderer, after obtaining the above projection pixel position, this step can directly index and read the image feature information at the corresponding projection pixel coordinates in different virtual frames for difference comparison, without introducing additional feature matching or optical flow estimation steps.
[0044] Subsequently, image feature information at each of the projected pixel locations is extracted. This image feature information may include, but is not limited to, pixel color values, and / or deep semantic feature values extracted based on neural networks or pre-trained visual feature extraction networks. Based on the extracted image feature information, a difference metric between feature vectors under any two different observation perspectives is calculated. This difference metric can be calculated using methods such as mean squared error, structural similarity index, or perceptual loss function value, and is recorded as the individual multi-view projection error value corresponding to the target Gaussian point. The two different observation perspectives may include one or more of the following perspective combinations: a combination of two real input perspectives, a combination of two adjacent virtual perspectives, or a combination of a virtual perspective and a nearest real input perspective.
[0045] After obtaining the individual error values of all Gaussian points in the 3D Gaussian space representation, a rasterization rendering process is used to determine the set of all Gaussian points projected onto each preset sub-region. Statistical aggregation operations are then performed on the individual error values corresponding to these Gaussian points. These aggregation operations include, but are not limited to, arithmetic mean calculation, weighted average calculation, or maximum value filtering. The final aggregation result is the multi-view projection error value corresponding to the preset sub-region. A multi-view projection error value within a preset reasonable range indicates a high reconstruction quality for the corresponding region; values below the lower limit or above the upper limit correspond to a lower reconstruction quality evaluation result. For example, the preset reasonable range is greater than or equal to 0.005 and less than or equal to 0.050, or other suitable values depending on the application scenario requirements; this paper does not impose any limitations on this.
[0046] The local Gaussian element density value is a quantitative indicator that is determined statistically based on the number of all Gaussian points projected onto a preset sub-region. It is used to represent the richness and redundancy of the geometric structure in the local space corresponding to the preset sub-region in the three-dimensional Gaussian space representation.
[0047] As an example, for any preset sub-region on a 2D image frame, the set of Gaussian points in the 3D Gaussian space representation corresponding to the 2D pixels in the preset sub-region is statistically analyzed. Alternatively, during the process of projecting 3D Gaussian points onto the pixel region using a rasterizer, the set of Gaussian points whose color contribution weight to any pixel in the sub-region exceeds a preset threshold, or whose coverage area of the projected 2D covariance matrix intersects with the sub-region, is recorded. The number of Gaussian points in this set is counted, and the count result or the value normalized based on the area of the sub-region is determined as the local Gaussian pixel density value corresponding to the preset sub-region. A density value within a reasonable range indicates good quality. If the density value is abnormally low, it indicates that the 3D surface corresponding to the preset sub-region lacks sufficient Gaussian description, and the rendering result is prone to holes or blurriness. If the density value is abnormally high and exceeds the reasonable redundancy range, it indicates that there may be overfitting noise caused by geometric reconstruction failure in the region. For example, a reasonable range for the density value is greater than or equal to 50 and less than or equal to 500, or other suitable values may be used based on the application scenario requirements; this paper does not impose any restrictions on this.
[0048] The gradient value of appearance change between adjacent frames is determined by statistical aggregation based on the color change amplitude of each pixel in the preset sub-region between temporally adjacent frames. It is used to represent the temporal stability and appearance coherence of the three-dimensional video prior sequence in the preset sub-region.
[0049] As an example, for the consecutive video frames in the aforementioned 3D video prior sequence determined by a preset camera motion trajectory, firstly, the set of pixels belonging to the same preset sub-region in the current evaluation frame and its temporally adjacent previous frame (or several previous frames) is determined. The difference measure of each corresponding pixel position in this pixel set in the color space is calculated. The color space includes, but is not limited to, RGB space and LAB color space. The difference measure includes, but is not limited to, absolute difference and mean square error. Subsequently, the inter-frame difference values of all pixels within the preset sub-region are statistically aggregated. This statistical aggregation includes, but is not limited to, calculating the arithmetic mean, the maximum value, or calculating the percentage of pixels with differences exceeding a preset threshold. The resulting aggregation result is the inter-frame appearance change gradient value corresponding to the preset sub-region.
[0050] It should be noted that the gradient value is positively correlated with the instantaneous motion amplitude of the preset camera trajectory. If the gradient value is within a reasonable range that matches the current camera speed, the preset sub-region is considered to have good temporal dynamic consistency. Conversely, if the gradient value is higher than the upper threshold, it indicates that the preset sub-region has experienced unexpected and drastic pixel jumps between adjacent frames, i.e., high-frequency flickering or jitter artifacts exist. If the gradient value is lower than the lower threshold, it indicates that the preset sub-region failed to produce the expected parallax changes or light and shadow flow when the camera moves, i.e., rendering freeze or texture sticking artifacts exist. For example, a reasonable gradient value range is greater than or equal to 0.002 and less than or equal to 0.080, or other suitable values may be used based on application scenario requirements; this article does not impose any limitations on this.
[0051] The cross-frame geometric position offset is a quantitative indicator used to characterize the temporal stability of the geometric structure of the three-dimensional Gaussian space representation at the corresponding local surface of the preset sub-region, based on the statistical aggregation of the three-dimensional spatial displacement of all Gaussian points projected into the preset sub-region between temporally adjacent rendering frames.
[0052] As an example, for any target Gaussian point in the 3D Gaussian space representation, the pairing relationship between it and the corresponding local surface region in the current rendering frame and the adjacent previous rendering frame is first determined by optical flow estimation or the nearest neighbor matching algorithm. The 3D center coordinates of the target Gaussian point are obtained during the rendering of each frame, and the Euclidean distance of the center coordinates between adjacent frames is calculated as the individual cross-frame geometric position offset of the target Gaussian point.
[0053] Based on this, for any preset sub-region on a 2D image frame, a set of all Gaussian points projected onto that sub-region through rasterization rendering is determined. A statistical aggregation operation is performed on the individual offsets corresponding to each Gaussian point in this set. This statistical aggregation includes, but is not limited to, calculating the arithmetic mean, weighted average, or maximum value. The resulting aggregation is the cross-frame geometric position offset corresponding to that preset sub-region.
[0054] It should be noted that if the offset matches the expected geometric displacement calculated based on the current camera's inter-frame pose transformation, the preset sub-region is considered to have good 3D geometric dynamic consistency. Conversely, if the offset is higher than the upper limit of the expected offset interval, it indicates that the 3D structure corresponding to the preset sub-region has undergone non-rigid abnormal deformation or positional jitter between adjacent frames, i.e., structural tearing artifacts exist. If the offset is lower than the lower limit of the expected offset interval, it indicates that the 3D structure corresponding to the preset sub-region has failed to generate the necessary geometric parallax in response to changes in camera viewpoint, i.e., structural solidification or depth loss artifacts exist. For example, the expected offset interval is greater than or equal to 0.001 and less than or equal to 0.100, or it can be other suitable values depending on the application scenario requirements; this paper does not impose any restrictions on this.
[0055] The degree of reconstruction uncertainty is a quantitative indicator used to characterize the credibility of the reconstructed content in the preset sub-region, which is determined by statistical aggregation of the individual uncertainty labels of all Gaussian points projected onto the preset sub-region.
[0056] As an example, for each Gaussian point in a three-dimensional Gaussian space representation, the electronic device pre-computes or calculates its individual uncertainty label in real time, the acquisition of which includes performing at least one of the following two decisions: Viewpoint Coverage Determination: Count the number of valid viewing angles in all input images used to construct the 3D Gaussian spatial representation where the 3D center coordinates of the Gaussian point are within the camera's view frustum and the line connecting the Gaussian point to the camera's optical center is not obscured by other Gaussian points. If the number of valid observations is lower than a preset minimum threshold, the Gaussian point is marked as a viewpoint blind spot. Projection consistency determination: The Gaussian point is projected onto the two-dimensional pixel plane corresponding to each effective observation viewpoint, and the image feature difference metric at the projected pixel position under any two effective observation viewpoints is calculated. If the difference metric exceeds a preset consistency threshold, it indicates that although the Gaussian point is within the visible range, its represented geometry or appearance conflicts under different viewpoints, and the Gaussian point is marked as a projection inconsistency point.
[0057] Based on the above determination, each Gaussian point is assigned a binary or continuous uncertainty label. For example, Gaussian points marked as blind spots or projection inconsistencies are marked with an uncertainty label of 1 (high uncertainty); the remaining Gaussian points are marked with an uncertainty label of 0 (low uncertainty).
[0058] Based on this, for any preset sub-region on any frame of the three-dimensional video prior sequence, the set of Gaussian points in the three-dimensional Gaussian space representation corresponding to the two-dimensional pixels in the preset sub-region is statistically analyzed, or the set of all Gaussian points projected onto the sub-region and effectively contributing to the pixel color is determined through the rasterization rendering process. A statistical aggregation operation is performed on the individual uncertainty markers corresponding to each Gaussian point in this set. This statistical aggregation includes, but is not limited to: calculating the proportion of Gaussian points marked as high uncertainty to the total number of Gaussian points in the set; calculating the weighted average of the uncertainty markers using the luminance contribution weight of each Gaussian point to the sub-region as the weight; and calculating the percentage of pixel area covered by Gaussian points marked as high uncertainty. The result is the degree of reconstruction uncertainty corresponding to the preset sub-region.
[0059] It should be noted that the greater the degree of uncertainty in the reconstruction, the more the image content presented by the preset sub-region relies on the extrapolation or guessing of the three-dimensional Gaussian space representation in the absence of sufficient observation data. The lower its visual credibility, the more likely it is to produce artifacts such as structural distortion, fictitious content, or loss of details.
[0060] S103, the electronic device converts multiple quality indicators corresponding to each preset sub-region in the above quality assessment results into a single comprehensive confidence level and generates a spatial confidence target map.
[0061] Among them, confidence level is positively correlated with image quality.
[0062] Specifically, the electronic device converts multiple quality indicators corresponding to each preset sub-region into a single comprehensive confidence level to obtain the aforementioned spatial confidence target map. This spatial confidence target map includes spatial confidence information corresponding to each two-dimensional image in the three-dimensional video prior sequence.
[0063] As an example, such as Figure 3 As shown, step S103 specifically includes: S1031, normalize each quality indicator corresponding to the preset sub-region in the quality assessment results; For example, five values with completely different physical meanings are uniformly mapped to the interval [0, 1], where 1.0 represents the best quality and 0.0 represents the worst quality.
[0064] For example, for multi-view projection error values, the corresponding confidence level can be calculated as follows: Set a preset optimal center value and a preset maximum tolerance upper limit value; determine the absolute value of the difference between the measured value and the optimal center value, and the difference between the maximum tolerance upper limit and the optimal center value; divide the absolute value by this difference to obtain the normalized deviation distance; subtract this normalized deviation distance from a constant 1, compare the result with 0, and take the larger value to obtain the confidence level of the single indicator. This calculation method can be understood as distance attenuation based on the optimal interval center. The same distance attenuation mapping function can also be used for local Gaussian element density, appearance change gradient, and geometric position offset indicators.
[0065] For the reconstruction uncertainty index, its confidence level can be calculated as follows: set a preset uncertainty tolerance threshold, calculate the quotient obtained by dividing the measured value by the tolerance threshold, subtract the quotient by the constant 1, and compare the result with 0 and take the larger value to obtain the confidence level of the single index. This calculation method indicates that the higher the proportion of uncertainty, the lower the confidence level; when the proportion of uncertainty reaches or exceeds the tolerance threshold, the confidence level is 0.
[0066] S1032, by using weighted averaging or a learning network, multiple indicators belonging to the same preset sub-region are fused and mapped into a single comprehensive confidence score to obtain a spatial confidence reference map. As an example, after obtaining the five individual confidence scores mentioned above, a single comprehensive confidence score corresponding to the preset sub-region is generated by multiplying them element by element. Alternatively, a weighted average fusion or a minimum value constraint fusion method can also be used; this paper does not limit the specific approach.
[0067] S1033, Based on spatial smoothing constraints and temporal consistency constraints, the confidence of adjacent sub-regions in the same frame of the above spatial confidence reference map and the confidence of sub-regions that have a corresponding relationship with adjacent frames are smoothed and adjusted respectively to obtain the spatial confidence target map.
[0068] Specifically, based on spatial smoothing constraints and temporal consistency constraints, the confidence levels of adjacent sub-regions and corresponding sub-regions within the same frame of the spatial confidence reference map are smoothed and adjusted. This includes: applying spatial smoothing filtering to adjacent sub-regions within the same frame of the spatial confidence reference map to maintain continuous confidence changes; applying temporal smoothing filtering to corresponding sub-regions within the spatial confidence reference map to ensure smooth evolution of confidence at the same physical spatial location over time; and generating the aforementioned spatial confidence target map based on the results of these spatial and temporal smoothing filters. The spatial smoothing and temporal smoothing are performed sequentially: spatial smoothing is performed first, followed by temporal smoothing adjustment based on the spatial smoothing adjustment.
[0069] As an example, spatial smoothing can be implemented using a simple weighted average, or other differentiated processing methods for flat areas and edges. Temporal smoothing can be implemented by determining the confidence of a preset sub-region after smoothing adjustment in the current frame based on the confidence of the preset sub-region in the current frame, the confidence of the previous frame (if smoothing was performed in the previous frame), and a smoothing factor.
[0070] Understandably, imposing spatial smoothing constraints on the spatial confidence target map, ensuring continuous confidence changes between adjacent spatial locations, can prevent local abrupt changes from affecting generation stability. Furthermore, constraining the confidence changes of the same spatial region between adjacent frames in the video's temporal dimension ensures the consistency of spatial generation control over the temporal dimension.
[0071] S104, the electronic device generates a spatially adaptive target diffusion time step field based on the above-mentioned spatial confidence target map.
[0072] In this embodiment, the target diffusion time step field includes a diffusion time step corresponding to each preset sub-region. The diffusion time step is used to determine the noise perturbation amplitude applied to the corresponding preset sub-region of the three-dimensional video prior sequence. The diffusion time step is negatively correlated with the confidence level.
[0073] In the embodiments of this application, the target diffusion time step field includes the spatial diffusion time step information corresponding to each two-dimensional image in the three-dimensional video prior sequence.
[0074] It should be noted that the term 'diffusion time step field' is a new definition in this application. In traditional diffusion models, the diffusion time step is a single scalar parameter used to uniformly control the noise injection and denoising process of the entire sample. It assumes that all regions within the sample have the same reconstruction reliability, which is difficult to adapt to the actual situation of highly uneven spatial quality distribution in 3D videos.
[0075] For example, such as Figure 4 As shown, step S104 above specifically includes: S1041, based on a preset mapping function, the confidence level corresponding to each preset sub-region in the spatial confidence target map is mapped to a diffusion time step, generating a reference diffusion time step field.
[0076] In this embodiment, the aforementioned preset mapping function is a mathematical transformation relationship that maps the value of spatial confidence (e.g., the value range of spatial confidence is greater than or equal to 0 and less than or equal to 1) to the value of diffusion time step (the value range of diffusion time step is greater than or equal to 0 and less than or equal to the preset maximum time step). The preset mapping function satisfies a monotonically decreasing law: the higher the confidence, the smaller the diffusion time step; the lower the confidence, the larger the diffusion time step.
[0077] Specifically, it can be a linear mapping, a power function mapping, an exponential decay mapping, etc. This paper does not limit it to symmetry. For example, when the spatial confidence value ranges from greater than or equal to 0 to less than or equal to 1, the product of the difference between the diffusion time step t of the preset sub-region and the spatial confidence c and the preset maximum time step Tmax is t = (1-c) * Tmax.
[0078] S1042, based on the continuity constraints in the spatial and temporal dimensions, the confidence of adjacent sub-regions in the same frame image in the reference diffusion time step field and the diffusion time step of the sub-regions that have a corresponding relationship with adjacent frames are smoothly adjusted to obtain the target diffusion time step field.
[0079] For example, in the spatial smoothing layer, for a reference diffusion time step field in the same frame image, for any spatial location, the diffusion time step values of its neighboring sub-regions are weighted and averaged, and the calculated weighted average value is used as the updated diffusion time step value at that location.
[0080] At the temporal smoothing level, for sub-regions with corresponding relationships between adjacent frames, the pixel-level correspondence between frames is first established through optical flow estimation or 3D projection matching, ensuring accurate association of the projected positions of the same physical point at different times. Then, causal moving average filtering or temporal Gaussian filtering is applied to the diffusion time step sequence of the same physical point along the time axis. Specifically, the smoothed value of the current frame is obtained by weighting the original value of the current frame with the smoothed value of the previous frame according to a preset weight, or by weighted averaging of the original values of several frames before and after. This suppresses abrupt changes in the diffusion time step caused by random jitter or occlusion changes between frames, ensuring the smooth evolution of diffusion control commands for the same physical point in consecutive frames. Finally, the diffusion time step field after spatial and temporal smoothing is used as the target diffusion time step field for subsequent spatial non-uniform noise injection steps.
[0081] S105, the electronic device performs spatially non-uniform noise injection on the above-mentioned three-dimensional video prior sequence based on the above-mentioned target diffusion time step field, and obtains the noise-perturbed three-dimensional video prior sequence.
[0082] Specifically, the electronic device determines the noise perturbation amplitude to be injected into each preset sub-region based on the diffusion time step and a preset noise scheduling table. The preset noise scheduling table indicates the mapping relationship between the diffusion time step and the noise perturbation amplitude. Based on time consistency constraints, the noise perturbation amplitude of the same preset sub-region is smoothly adjusted between adjacent frames. Based on the smoothed noise perturbation amplitude, noise perturbation is performed on the three-dimensional video prior sequence to obtain the noise-perturbed three-dimensional video prior sequence. This noise-perturbed three-dimensional video prior sequence is regarded as an intermediate diffusion state after spatially non-uniform noise injection processing, and this state serves as the direct input to the single-step diffusion generation network.
[0083] Understandably, in the diffusion model, there is a one-to-one correspondence between time step t and noise intensity, and the time step field can be used to indicate the noise injection intensity. The time step field serves as an index value, while the actual noise intensity is obtained through the noise scheduling table of the diffusion model.
[0084] In other words, based on the aforementioned target diffusion time step field, the electronic device first obtains the noise injection intensity corresponding to the preset sub-region through a predefined noise scheduling table, then performs time-consistent smoothing adjustment, and performs noise injection based on the smoothed noise disturbance amplitude, thereby achieving the application of differentiated noise disturbance to different regions.
[0085] For example, the aforementioned preset noise scheduling table is used to indicate the following mapping relationship: for a preset sub-region of a preset high-confidence region where the diffusion time step is less than or equal to a first threshold, the noise disturbance amplitude to be injected is a first amplitude noise; for a preset sub-region of a preset medium-confidence region where the diffusion time step is greater than the first threshold and less than a second threshold, the noise disturbance amplitude to be injected is a second amplitude noise, the second amplitude noise being greater than the first amplitude noise; for a preset sub-region of a preset low-confidence region where the diffusion time step is greater than or equal to the second threshold, the noise disturbance amplitude to be injected is a third amplitude noise, the third amplitude noise being greater than the second amplitude noise.
[0086] That is, when the diffusion time step corresponding to a certain spatial location is less than or equal to a first threshold, the region is determined to be a preset high-confidence region; when the diffusion time step is greater than the first threshold and less than a second threshold, the region is determined to be a preset medium-confidence region; when the diffusion time step is greater than or equal to the second threshold, the region is determined to be a preset low-confidence region. The noise perturbation amplitude to be injected for the preset high-confidence region is the first amplitude noise, the noise perturbation amplitude to be injected for the preset medium-confidence region is the second amplitude noise, and the noise perturbation amplitude to be injected for the preset low-confidence region is the third amplitude noise.
[0087] For example, the time-consistent smoothing adjustment can be implemented by determining the noise amplitude of the preset sub-region after smoothing adjustment in the current frame based on the noise amplitude of the preset sub-region in the current frame, the noise amplitude of the previous frame (if the previous frame was smoothed, then it is the noise amplitude after smoothing in the previous frame), and the smoothing factor.
[0088] S106, the electronic device generates a local redraw mask based on the aforementioned target diffusion time step field.
[0089] In this embodiment of the application, the local redraw mask is a continuous value mask (that is, not a binary mask), which is used to indicate the degree of participation of each spatial region (preset sub-region) in the denoising process (that is, understood as a single-step diffusion process).
[0090] For example, a preset sub-region with a diffusion time step less than or equal to a first threshold is determined to be a structurally reliable region and labeled with a first mask value (e.g., 0.2); a preset sub-region with a diffusion time step greater than the first threshold and less than a second threshold is determined to be a lightly modified region and labeled with a second mask value (e.g., 0.5), where the second mask value is greater than the first mask value, and a higher mask value indicates a higher degree of participation; a preset sub-region with a diffusion time step greater than or equal to the second threshold is determined to be a redraw priority region and labeled with a third mask value (e.g., 0.8), where the third mask value is greater than the second mask value, and the second threshold is greater than the first threshold.
[0091] The above-mentioned region division does not form a hard boundary, but exists in the form of continuous numerical values through mask values, which are used for continuous control of the subsequent noise injection intensity.
[0092] In this embodiment, the local redraw mask corresponds one-to-one with the 3D video in both spatial and temporal dimensions, and its value is used to indicate the degree of participation of each spatial location in the diffusion process. Unlike traditional local inpainting methods based on binary masks, the local redraw mask in this embodiment is a continuous value mask, used to adjust noise intensity rather than directly cropping the generated region, thereby avoiding boundary breaks or semantic discontinuities.
[0093] In the embodiments of this application, the mask value in the local redraw mask can also be understood as noise weight, indicating the weight (or proportion) of noise retention during the denoising process, and indirectly indicating the weight of the original image in the region to be retained. The effective noise intensity that finally participates in the diffusion process after denoising is determined by the noise amplitude corresponding to the time step and the mask value.
[0094] S107, the electronic device inputs the above-mentioned noise-perturbed 3D video prior sequence, the above-mentioned local redraw mask, and the above-mentioned target diffusion time step field into the single-step diffusion generation network for single-step denoising to generate the target 3D video.
[0095] In the embodiments of this application, the single-step diffusion generation network can realize the generation of a high-quality 3D video result from a noisy perturbation state in one forward propagation.
[0096] Specifically, the input to this single-step diffusion generation network includes the noise-perturbed 3D video prior sequence and the target diffusion time step field generated in step S04 above. A spatially aware coding mechanism can be introduced within the network to embed diffusion control signals at different spatial locations, enabling the network to distinguish generation strategies for different regions at the feature level. This mechanism allows the network to adaptively adjust denoising and generation behaviors based on corresponding diffusion intensity and confidence information when facing different spatial regions within the same video frame. This allows the network to simultaneously perceive 3D video spatial structure information and diffusion control information, thereby completing the joint modeling of complex spatial structures in a single forward computation. Through this design, the system significantly reduces inference latency while ensuring generation quality, enabling 3D video generation to meet the requirements of real-time or near-real-time applications.
[0097] Specifically, for structurally reliable regions, the diffusion time step is small and the noise injection intensity is low. The denoising network mainly compensates for and smooths the details. This denoising network (also called the denoising generation network) is the same as the single-step diffusion generation network. The name of the single-step diffusion generation network emphasizes the network structure: completed in one forward propagation. The name of the denoising generation network emphasizes the network function: removing noise and generating content.
[0098] For regions with slight modifications, the diffusion time step is larger, resulting in a significant increase in noise injection intensity. The denoising network then regenerates the geometry and appearance structure under weaker prior constraints. For the priority redrawing region, the diffusion time step is larger, the noise injection intensity is significantly improved, and the denoising network regenerates the geometry and appearance structure under weaker prior constraints. For spatial boundary regions, the diffusion time step varies continuously, allowing the generated results to maintain a natural transition between different regions and avoiding structural breaks.
[0099] The spatial intelligent 3D video generation method with spatial adaptive noise injection provided in this application introduces a spatial confidence perception mechanism and a non-uniform noise injection strategy, enabling differentiated control of the diffusion generation process in the spatial dimension. This allows for the maintenance of geometric and appearance stability in high-confidence areas while fully redrawing low-confidence areas, thereby improving the overall quality, consistency, and generation efficiency of 3D video generation.
[0100] In some other possible implementations, such as Figure 5 As shown, the electronic device can also perform distillation optimization and model update steps during the training phase: determining a reference diffusion model and a single-step generation model, wherein the reference diffusion model adopts a multi-step diffusion generation strategy during training to progressively denoise the noise-perturbed 3D video prior sequence to generate a first 3D video, and the single-step generation model adopts the single-step diffusion generation network during training to perform single-step denoising on the same noise-perturbed 3D video prior sequence to generate a second 3D video; training the reference diffusion model and the single-step generation model, minimizing the difference between the output of the single-step generation model and the output of the reference diffusion model (minimizing the difference between the first 3D video and the second 3D video), updating the parameters of the single-step generation model, so that the single-step generation model can reproduce the spatial generation behavior of the reference diffusion model in single-step generation, specifically aligning the 3D spatial features of the reference diffusion model and the single-step generation model in the intermediate layer; and, temporal consistency constraints and spatial structure preservation constraints can also be introduced (refer to the relevant technical implementation descriptions of temporal consistency and spatial structure constraints above).
[0101] Accordingly, step S107 may specifically include: inputting the noise-perturbed 3D video prior sequence, the local redraw mask, and the target diffusion time step field into the single-step generation model to generate the target 3D video.
[0102] Alternatively, the parameters related to the target 3D video generation may also include generation model parameters, which indicate the 3D video to be generated using the reference diffusion model or the single-step generation model. These parameters are derived from user-specified or default settings. Specifically, step S107 includes: inputting the noise-perturbed 3D video prior sequence into the model indicated by the generation model parameters for single-step denoising to generate the target 3D video.
[0103] By adopting this distillation optimization method, the instability caused by relying solely on pixel-level loss can be avoided, enabling the single-step generation model to have stronger generation capabilities in complex occluded regions, dynamic regions, and detailed structures.
[0104] Furthermore, while existing methods have incorporated teacher-student distillation strategies to improve generation efficiency, most distillation processes focus on compression of the temporal or feature dimensions, lacking explicit constraints on 3D spatial structure and cross-frame consistency, making it difficult to fully inherit the ability of multi-step diffusion models to repair complex regions. The distillation optimization method provided in this application can also introduce a spatial confidence target map derived from a 3D Gaussian spatial representation and a target diffusion time step field as explicit spatial structure priors in the distillation process. This allows the distilled single-step diffusion generation network to accurately identify and learn the repair behavior of the multi-step model (reference diffusion model) in geometrically ambiguous regions and viewpoint blind spots. Unlike indiscriminate global feature distillation, this application encourages the single-step generation network to apply higher attention weights to low-confidence sub-regions (such as floating object clusters and temporal flickering regions) during training. This allows the lightweight architecture, which only requires single-step inference, to inherit the reference diffusion model's ability to finely repair complex 3D artifacts with high fidelity. Furthermore, it effectively suppresses geometric tearing and cross-frame appearance jitter in the generated target 3D video, significantly improving the spatial geometric fidelity and temporal visual coherence of the distilled model in video generation tasks.
[0105] The advantages of the spatially adaptive noise injection-based intelligent 3D video generation method provided in this application compared to existing solutions are described below: Compared with existing 3D video generation methods that mainly rely on diffusion models or multi-step inversion generation strategies with globally consistent noise injection, this application has achieved significant technical improvements in spatial differentiation generation capability, local structure repair quality, temporal consistency maintenance, and generation efficiency, and can better meet the actual needs of spatial intelligent 3D video generation under sparse view conditions.
[0106] 1. Spatial differentiation generation capability is significantly enhanced. This application introduces a spatial confidence-driven diffusion control mechanism, which enables the diffusion noise injection process to no longer be limited to using a uniform time step for the entire 3D video, but to allocate different diffusion intensities according to the reconstruction reliability of different spatial regions.
[0107] In high-confidence regions, the system introduces only low-intensity noise for detail correction; in low-confidence regions, the system triggers full redrawing with higher-intensity noise, thereby achieving differentiated processing of different spatial regions in the same generation process.
[0108] This mechanism effectively avoids the problems of "overall noise reduction leading to degradation of high-quality areas" or "slight noise reduction leading to unrepairable erroneous areas" in existing technologies, making the generated 3D video results more reasonable in terms of spatial dimensions.
[0109] 2. The restoration quality of obstructed and under-observed areas has been significantly improved. Under sparse view input conditions, occlusion areas, blind spots, and dynamic areas are often the main sources of quality degradation in 3D video generation.
[0110] This application explicitly identifies the aforementioned low-reliability regions using a spatial confidence target map and applies a higher noise injection intensity to them during the diffusion generation stage, enabling the denoising generation network to regenerate the geometric and appearance structures under weaker prior constraints.
[0111] Experiments show that this method can effectively reduce geometric distortion, texture mismatch, and temporal flicker in occluded areas, significantly improving the generation integrity and realism of complex areas.
[0112] 3. The temporal consistency and spatial continuity of 3D video are effectively guaranteed. This application introduces spatial continuity and temporal consistency constraints when constructing a spatially adaptive diffusion time step field, so that the diffusion control parameters change smoothly between adjacent spatial locations and adjacent time frames.
[0113] This design enables the generated video to maintain a stable appearance evolution over time, avoiding abrupt changes or jumps introduced by local redrawing, thus ensuring overall temporal consistency while guaranteeing local repair effects. This effect is particularly important for long-term 3D video generation and scenarios with continuously changing camera trajectories.
[0114] 4. Achieving a balance between single-step generation efficiency and generation quality. Compared with existing methods that rely on multi-step diffusion inversion, this application enables the model to complete high-quality reconstruction of complex 3D videos in a single generation process through spatial adaptive noise injection and 3D perceptual distillation mechanism.
[0115] This scheme significantly reduces the number of inference steps and computational latency while maintaining generation results similar to or even better than multi-step diffusion models, solving the problem of "difficulty in achieving both generation quality and inference efficiency" in existing technologies.
[0116] Therefore, this application is more suitable for spatial intelligence application scenarios that are sensitive to real-time performance or computing resources.
[0117] 5. Robustness to uneven input prior quality is significantly enhanced. Existing 3D video generation methods typically assume that the overall quality of the input prior is relatively balanced. When the prior quality is not evenly distributed in space, the generation result is easily affected by local errors.
[0118] This application decouples spatial confidence modeling and diffusion control, enabling the system to automatically adapt to prior inputs with different quality distributions. Even when there are both high-quality backgrounds and severely distorted foreground regions in the prior video, it can still maintain stable generation performance.
[0119] This feature significantly improves the applicability and robustness of the system in real-world, complex scenarios.
[0120] 6. Simultaneous improvement in the ability to maintain three-dimensional structure and the ability to express details. By performing structure-preserving processing on high-confidence regions during the generation process, this application can effectively retain the parts of the original 3D prior that already have correct geometric relationships, avoiding the structural drift problem caused by repeated generation.
[0121] Meanwhile, by introducing full redrawing in low-confidence areas, the ability to generate details is not constrained by the overall structure, thereby achieving a synergistic improvement in "structural stability" and "detail reconstruction capability" in the same generated result.
[0122] 7. Applicable to various spatial intelligent 3D generation application scenarios Based on the above technical effects, the method described in this application can be widely applied to the following scenarios, but is not limited to them: sparse views Figure 3 3D video synthesis, robot environmental perception and simulation, augmented reality and virtual reality content generation, digital twin modeling, 3D film and television production, and spatial intelligent systems that need to generate high-quality 3D temporal content under limited input conditions.
[0123] Its spatial adaptive generation characteristics enable it to adapt well to different scene scales and input quality conditions.
[0124] The following combination Figure 6 This paper describes the overall architecture of the spatially adaptive noise injection-based intelligent 3D video generation system provided in this application, including sparse view... Figure 3The module includes a priori generation module 601, a 3D video spatial quality assessment module 602, a spatial confidence map construction module 603, a spatial adaptive diffusion time step field generation module 604, a local redraw mask and spatial noise injection module 605, and a single-step diffusion 3D video generation module 606.
[0125] Among them, sparse view Figure 3 The prior generation module 601 is used to construct a three-dimensional video prior sequence with basic geometric consistency and temporal continuity under the condition that the input view is extremely sparse (e.g., only two images). The output of this module is: a set of coarse three-dimensional video frame sequences containing the temporal dimension and their corresponding three-dimensional Gaussian space representations, which can be referred to in the relevant description of S101 above.
[0126] The 3D video spatial quality assessment module 602 is used to quantitatively assess the reconstruction quality of coarse 3D video (3D video prior sequence) in the spatial dimension, providing a basis for subsequent diffusion control. The module outputs a set of quality index measurement values corresponding to the spatial location of the 3D video, which can be referred to in the relevant description of S102 above.
[0127] The spatial confidence map construction module 603 is used to transform the spatial quality assessment results into a spatial confidence target map that can be used to generate control. The spatial confidence target map can be understood as a continuous function field aligned with the three-dimensional video prior sequence in the spatial and temporal dimensions. For details, please refer to the relevant description of S103 above.
[0128] The spatial adaptive diffusion time step field generation module 604 is used to generate a diffusion time step control field related to spatial location based on the spatial confidence target map, thereby realizing the adaptive adjustment of the diffusion noise injection intensity in the spatial dimension. For a detailed description, please refer to the relevant description of S104 above.
[0129] The local redraw mask and spatial noise injection module 605 is used to perform position-related noise injection operations on the three-dimensional video prior according to the diffusion time step corresponding to the spatial position, and to further construct a local redraw mask after obtaining the spatially adaptive target diffusion time step field, in order to help control the degree of participation of different preset sub-regions in the diffusion process. For a detailed description, please refer to the relevant descriptions of S105 and S106 above.
[0130] The single-step diffusion 3D video generation module 606 is used to realize the single-step generation from coarse 3D video prior to high-quality 3D video result under spatial adaptive noise control conditions (target diffusion time step field and local redraw mask). For a detailed description, please refer to the relevant description in S107 above.
[0131] In some possible implementations, the system further includes a 3D-aware jump-flow distillation and model optimization module 607, used to transfer the generation capability of the multi-step diffusion model to the single-step generation model during the training phase, thereby improving the stability and spatial consistency of the single-step generation results. For a detailed description, please refer to the section above regarding... Figure 5 The relevant methods and steps are explained.
[0132] It should be noted that any specific implementation of the above method embodiments in this application is applicable to the corresponding module in the spatial intelligent 3D video generation system. For a detailed description, please refer to the relevant descriptions of the above method embodiments, which will not be elaborated further in this article.
[0133] This application also provides a spatially adaptive noise injection-based intelligent 3D video generation apparatus, including a unit for executing any of the spatially adaptive noise injection-based intelligent 3D video generation methods in the above method embodiments.
[0134] In the embodiments of this application, any of the implementation methods mentioned in the method embodiments are also applicable to the spatial intelligent 3D video generation device with spatial adaptive noise injection provided in this application. For specific execution steps, please refer to the description of the foregoing method embodiments, which will not be detailed here.
[0135] This application also provides a spatially adaptive noise injection-based intelligent 3D video generation apparatus, including a processor. The processor is used to execute any of the spatially adaptive noise injection-based intelligent 3D video generation methods described in the above method embodiments. Specific execution steps can be found in the descriptions of the foregoing method embodiments, and will not be detailed here.
[0136] This application also provides a computer storage medium that can store multiple instructions. These instructions are adapted to be loaded and executed by a processor using the spatial adaptive noise injection spatial intelligent 3D video generation method provided in this application. For details of the execution process, please refer to the specific description of the method embodiments shown above, which will not be elaborated here.
[0137] This application also provides a computer program product containing instructions that, when run on an electronic device, cause the electronic device to execute the method steps of the method embodiments shown above.
[0138] This application also provides a chip module, including a transceiver component and a chip, wherein the chip is used to execute the method steps of the above-described method embodiments.
[0139] It is understood that the spatially adaptive noise injection-based intelligent 3D video generation system, spatially adaptive noise injection-based intelligent 3D video generation device, computer storage medium, computer program, computer program product, and chip provided above are all used to execute the method shown in any implementation of the corresponding aspect of the embodiments of this application. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding method, and will not be detailed here.
[0140] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes the processes of the embodiments of the above methods.
[0141] The term "at least one" in this application refers to one or more items. "More than one item" means two or more items. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, it should be understood that although the terms "first," "second," etc., may be used to describe objects in this application, these objects should not be limited to these terms. These terms are only used to distinguish the objects from each other.
[0142] The terms “including” and “having”, and any variations thereof, mentioned above are intended to cover non-exclusive inclusion.
[0143] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A spatially adaptive noise-injected intelligent 3D video generation method, characterized in that, The method includes: Based on at least two input images, a three-dimensional Gaussian space representation is constructed, and based on the three-dimensional Gaussian space representation and a preset camera motion trajectory, a three-dimensional video prior sequence is rendered. The reconstruction quality of the three-dimensional video prior sequence in the spatial dimension is analyzed, and the values of at least two quality indicators corresponding to the preset sub-region are obtained. The quality evaluation results corresponding to the three-dimensional video prior sequence in time and space are obtained. The preset sub-region refers to any sub-region obtained by regularly dividing any two-dimensional image in the three-dimensional video prior sequence. The at least two quality indicators corresponding to each preset sub-region in the quality assessment results are converted into a single comprehensive confidence score to generate a spatial confidence target map, wherein the confidence score is positively correlated with image quality; Based on the spatial confidence target map, a spatially adaptive target diffusion time step field is generated. The target diffusion time step field includes a diffusion time step corresponding to each preset sub-region. The diffusion time step is used to determine the noise perturbation amplitude applied to the corresponding preset sub-region of the three-dimensional video prior sequence. The diffusion time step is negatively correlated with the confidence level. Based on the target diffusion time step field, spatially non-uniform noise injection is performed on the three-dimensional video prior sequence to obtain a noise-perturbed three-dimensional video prior sequence; and a local redraw mask is generated based on the target diffusion time step field, the local redraw mask being used to indicate the degree of participation of each preset sub-region in the denoising process. The noise-perturbed 3D video prior sequence, the local redraw mask, and the target diffusion time step field are input into a single-step diffusion generation network for single-step denoising to generate the target 3D video.
2. The method as described in claim 1, characterized in that, The construction of a three-dimensional Gaussian space representation based on at least two input images includes: Based on the at least two input images, a set of Gaussian elements are initialized in three-dimensional space to obtain an initial three-dimensional Gaussian field; By acquiring the projected images of the initial three-dimensional Gaussian field under different input viewpoints and aligning them with the corresponding input images under the input viewpoints, the spatial position and appearance attributes of the Gaussian primitives are optimized to obtain the three-dimensional Gaussian space representation.
3. The method as described in claim 1 or 2, characterized in that, The at least two quality indicators include at least two of the following: The multi-view projection error value is determined by statistical aggregation of the individual error values of all Gaussian points projected onto the preset sub-region. The error value of a single Gaussian point is used to quantitatively characterize the geometric and appearance consistency of the Gaussian point under different observation views. The local Gaussian element density value is determined statistically based on the number of all Gaussian points projected onto the preset sub-region. It is used to quantitatively characterize the geometric richness and redundancy of the three-dimensional Gaussian space representation in the corresponding local space of the preset sub-region. The gradient value of appearance change between adjacent frames is determined by statistical aggregation based on the color change amplitude of each pixel in the preset sub-region between temporally adjacent frames. It is used to quantitatively characterize the temporal stability and appearance coherence of the three-dimensional video prior sequence in the preset sub-region. The cross-frame geometric position offset is a quantitative index that is determined by statistical aggregation of the three-dimensional spatial displacement of all Gaussian points projected into the preset sub-region between temporally adjacent rendering frames. It is used to quantify the temporal stability of the geometric structure represented by the three-dimensional Gaussian space at the corresponding local surface of the preset sub-region. The degree of reconstruction uncertainty is a quantitative indicator used to characterize the credibility of the reconstructed content in the preset sub-region, which is determined by statistical aggregation of the individual uncertainty labels of all Gaussian points projected onto the preset sub-region.
4. The method as described in claim 1 or 2, characterized in that, The step of converting at least two quality indicators corresponding to each preset sub-region in the quality assessment results into a single comprehensive confidence level and generating a spatial confidence target map includes: Normalize each quality index corresponding to each preset sub-region in the quality assessment results; By using weighted averaging or a learning network, multiple normalized indicators belonging to the same preset sub-region are fused and mapped into a single comprehensive confidence level to obtain a spatial confidence reference map. Based on spatial smoothing constraints and temporal consistency constraints, the confidence of adjacent sub-regions in the same frame of the spatial confidence reference map and the confidence of sub-regions that have a corresponding relationship with adjacent frames are smoothed and adjusted respectively to obtain the spatial confidence target map.
5. The method as described in claim 1 or 2, characterized in that, The step of generating a spatially adaptive target diffusion time step field based on the spatial confidence target map includes: Based on a preset mapping function, the confidence level corresponding to each preset sub-region in the spatial confidence target map is mapped to a diffusion time step to generate a reference diffusion time step field. Based on the continuity constraints in the spatial and temporal dimensions, the confidence of adjacent sub-regions in the same frame image in the reference diffusion time step field and the diffusion time step of the sub-regions that have a corresponding relationship with adjacent frames are smoothly adjusted to obtain the target diffusion time step field.
6. The method as described in claim 1 or 2, characterized in that, The step of injecting spatially non-uniform noise into the three-dimensional video prior sequence based on the target diffusion time step field to obtain a noise-perturbed three-dimensional video prior sequence includes: Based on the diffusion time step and the preset noise scheduling table corresponding to each preset sub-region, the noise disturbance amplitude to be injected into the preset sub-region is determined. The preset noise scheduling table is used to indicate the mapping relationship between the diffusion time step and the noise disturbance amplitude. Based on time consistency constraints, the noise perturbation amplitude of the same preset sub-region between adjacent frames is smoothly adjusted. Based on the smoothed noise perturbation amplitude, noise injection is performed on the three-dimensional video prior sequence to obtain the noise-perturbed three-dimensional video prior sequence.
7. The method as described in claim 1 or 2, characterized in that, The generation of a local redraw mask based on the target diffusion time step field includes: For a preset sub-region where the diffusion time step is less than or equal to the first threshold, it is determined to be a structurally reliable region and labeled with the first mask value; For a preset sub-region where the diffusion time step is greater than the first threshold and less than the second threshold, it is determined to be a slightly modified region and marked with a second mask value. The second mask value is greater than the first mask value. The higher the mask value, the higher the degree of participation in the diffusion process. The second threshold is greater than the first threshold. For a preset sub-region where the diffusion time step is greater than or equal to the second threshold, it is determined as a redraw priority region and marked with a third mask value, wherein the third mask value is greater than the second mask value.
8. The method as described in claim 1 or 2, characterized in that, The method further includes: A reference diffusion model and a single-step generation model are determined. The reference diffusion model adopts a multi-step diffusion generation strategy during training to progressively denoise the noise-perturbed 3D video prior sequence. The single-step generation model adopts the single-step diffusion generation network during training to perform single-step denoising on the same noise-perturbed 3D video prior sequence. The reference diffusion model and the single-step generation model are trained, and during the training process, the difference between the output of the single-step generation model and the output of the reference diffusion model is minimized. The step of inputting the noise-perturbed 3D video prior sequence, the local redraw mask, and the target diffusion time step field into a single-step diffusion generation network for single-step denoising to generate the target 3D video includes: The noise-perturbed 3D video prior sequence, the local redraw mask, and the target diffusion time step field are input into the single-step generation model to generate the target 3D video.
9. A spatially adaptive noise-injection-based intelligent 3D video generation system, characterized in that, include: The sparse view 3D prior generation module is used to construct a 3D Gaussian space representation based on at least two input images, and to render a 3D video prior sequence based on the 3D Gaussian space representation and a preset camera motion trajectory. The three-dimensional video spatial quality assessment module is used to analyze the reconstruction quality of the three-dimensional video prior sequence in the spatial dimension, obtain the values of at least two quality indicators corresponding to the preset sub-region, and obtain the quality assessment result corresponding to the three-dimensional video prior sequence in time and space. The preset sub-region refers to any sub-region obtained after regularly dividing any two-dimensional image in the three-dimensional video prior sequence. The spatial confidence map construction module is used to convert the at least two quality indicators corresponding to each preset sub-region in the quality assessment result into a single comprehensive confidence score, and generate a spatial confidence target map, wherein the confidence score is positively correlated with the image quality; The spatially adaptive diffusion time step field generation module is used to generate a spatially adaptive target diffusion time step field based on the spatial confidence target map. The target diffusion time step field includes a diffusion time step corresponding to each preset sub-region. The diffusion time step is used to determine the noise perturbation amplitude applied to the corresponding preset sub-region of the three-dimensional video prior sequence. The diffusion time step is negatively correlated with the confidence level. The local redraw mask and spatial noise injection module is used to perform spatially non-uniform noise injection on the three-dimensional video prior sequence based on the target diffusion time step field to obtain the noise-perturbed three-dimensional video prior sequence. The local redraw mask and spatial noise injection module is also used to generate a local redraw mask based on the target diffusion time step field, which is used to indicate the degree of participation of each preset sub-region in the denoising process. The single-step diffusion 3D video generation module is used to input the noise-perturbed 3D video prior sequence, the local redraw mask, and the target diffusion time step field into the single-step diffusion generation network for single-step denoising to generate the target 3D video.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed, performs the method according to any one of claims 1 to 8.