Text-driven three-dimensional editing method and device based on generative video prior

Through a text-driven three-dimensional editing method based on a generative video prior, the pre-trained video generation model and attention mechanism are used to solve the problem of perspective consistency and coherence in the existing 3D editing methods, and a fast and accurate 3D editing effect is achieved.

CN120147509APending Publication Date: 2025-06-13Artificial Intelligence and Robotics Innovation Center of Hong Kong Institute of Innovation, Chinese Academy of Sciences +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111460.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing 3D editing methods are difficult to maintain the consistency and coherence of edited 3D content from all perspectives, hindering the development and application promotion of 3D editing technology.

Method used

The text-driven three-dimensional editing method based on the generative video prior is adopted. By obtaining the pre-trained generated video model and the three-dimensional model to be edited, the training viewing image of multiple continuous perspectives is captured, the original video is generated, the potential noise is mixed with random Gaussian noise, inversion is performed, the spatial and temporal attention map is extracted, and the three-dimensional model is updated according to the text editing instructions.

Benefits of technology

It realizes fast and precise editing of video content, ensuring that the edited video is visually natural and smooth, and can maintain the consistency and coherence of 3D content from different perspectives, reducing editing complexity and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147509A_ABST
    Figure CN120147509A_ABST
Patent Text Reader

Abstract

The invention provides a text-driven three-dimensional editing method and a text-driven three-dimensional editing device based on generative video prior. The text-driven three-dimensional editing method comprises the following steps: generating an original video based on training view angle images of a three-dimensional model to be edited captured from a plurality of continuous view angles; and extracting potential noise from the original video based on the pre-training generated video model, and mixing the potential noise with the random Gaussian noise to obtain mixed noise. Performing inversion on the original video based on the pre-training generated video model and the mixed noise, and simultaneously extracting a space attention map and a time attention map in the inversion process; according to the obtained text editing instruction, the space and time attention maps are used for covering or guiding the editing video to generate the corresponding attention map in the denoising process, and the edited video is obtained. And updating the three-dimensional model based on the video to obtain an edited three-dimensional model. According to the method and the device, the consistency and coherence of the 3D contents under different visual angles can be ensured in the editing process, and the rapid and accurate editing of the video contents is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D editing technology, and in particular, to a text-driven three-dimensional editing method and device based on generative video prior. Background Art

[0002] Traditional 3D editing methods are the exclusive domain of professionals who need to spend a lot of time and effort on fine operations in professional software, which is both complex and time-consuming for most ordinary users. Text-driven 3D editing, with its user-friendly editing interface and clear editing instructions, provides an unprecedented editing experience for ordinary users. However, the current mainstream 3D editing methods still face severe challenges. Most of them rely on existing image editing models and adopt a per-view iterative editing method, which, although realizing the possibility of 3D editing to some extent, makes it difficult for the edited 3D content to maintain consistency and coherence in all views. This problem seriously hinders the development and application promotion of 3D editing technology. Summary of the Invention

[0003] The present invention provides a text-driven three-dimensional editing method and device based on generative video prior to solve the defect that the edited 3D content in the prior art is difficult to maintain consistency and coherence in all views, and realizes fast and accurate editing of video content. The technical solutions proposed by the present invention are as follows: In a first aspect, the present invention provides a text-driven three-dimensional editing method based on generative video prior, including: Obtaining a pre-trained generative video model and a three-dimensional model to be edited; the pre-trained generative video model is an inversion-based video editing model; Capturing training perspective images of the three-dimensional model to be edited from multiple consecutive perspectives, and generating an original video based on the training perspective images of the multiple consecutive perspectives; Extracting latent noise from the original video based on the pre-trained generative video model, and mixing the latent noise with random Gaussian noise to obtain mixed noise; Inverting the original video based on the pre-trained generative video model and the mixed noise, and simultaneously extracting a spatial attention map and a temporal attention map during the inversion process; According to the obtained text editing instruction, using the spatial attention map and the temporal attention map to cover or guide the corresponding attention map in the denoising process of generating the edited video, to obtain the edited video; Updating the three-dimensional model to be edited based on the edited video to obtain an edited three-dimensional model.

[0004] Optionally, capturing training perspective images of the 3D model to be edited from multiple consecutive perspectives and generating an original video based on the training perspective images of multiple consecutive perspectives includes: Capturing training perspective images of the 3D model to be edited from multiple consecutive perspectives based on a pre-established original camera trajectory; Obtaining the camera coordinates corresponding to each training perspective image, sorting all the training perspective images according to the camera coordinates, and obtaining a sorted video frame sequence; Calculating the physical distance between adjacent camera coordinates in the sorted video frame sequence, and determining whether there is discontinuity between adjacent camera coordinates based on the physical distance; For each pair of adjacent camera coordinates with discontinuity, performing spherical interpolation on the rotation matrix of the adjacent camera coordinates and performing linear interpolation on the translation matrix of the adjacent camera coordinates to obtain corresponding intermediate camera trajectories; Integrating each of the intermediate camera trajectories into the original camera trajectory to obtain a continuous camera trajectory; Inputting the continuous camera trajectory into a rendering engine to generate an image sequence; Encoding the image sequence generated by rendering into a video file using a video encoding library to obtain the original video.

[0005] Optionally, mixing the latent noise with random Gaussian noise to obtain mixed noise includes: Selecting an original anchor frame from the original video, and editing the original anchor frame using a text instruction-driven pixel conversion method to obtain an edited anchor frame; Calculating the structural similarity between the original anchor frame and the edited anchor frame, and determining a mixing weight according to the structural similarity; Mixing the latent noise with random Gaussian noise according to the mixing weight to obtain mixed noise.

[0006] Optionally, selecting the original anchor frame from the original video includes: Obtaining a text prompt, and calculating the graphic-text matching score between each frame image in the original video and the text prompt; Selecting the image with the highest graphic-text matching score as the original anchor frame.

[0007] Optionally, the pre-trained video generation model includes an encoder and an inversion module; inverting the original video based on the pre-trained video generation model and the mixed noise, and simultaneously extracting a spatial attention map and a temporal attention map during the inversion process includes: Adding the mixed noise to the original video to generate a noisy video; Inputting the noisy video into the encoder to convert it into a preliminary representation in the latent space; Input the preliminary representation into the inversion module for optimization to obtain the latent representation of the noisy video. During the inversion process, a spatial attention mechanism and a temporal attention mechanism are used to extract a spatial attention map and a temporal attention map respectively.

[0008] Optionally, the pre-trained video generation model further includes a decoder and a denoising network. The step of using the spatial attention map and the temporal attention map to cover or guide the corresponding attention maps in the denoising process of editing video generation according to the obtained text editing instructions to obtain the edited video includes: Edit the latent representation of the noisy video in the latent space according to the text editing instructions to obtain the edited latent representation; Input the edited latent representation into the decoder to generate an initial edited video; Input the initial edited video into the denoising network. The denoising network uses the spatial attention map and the temporal attention map as additional input information to cover or guide the corresponding attention maps in the denoising process of editing video generation, and outputs the edited video.

[0009] In a second aspect, the present invention further provides a text-driven 3D editing device based on a generative video prior, including the following modules: A data acquisition module, configured to acquire a pre-trained video generation model and a 3D model to be edited. The pre-trained video generation model is an inversion-based video editing model; A video generation module, configured to capture training perspective images of the 3D model to be edited from multiple consecutive perspectives, and generate an original video based on the training perspective images of the multiple consecutive perspectives; A noise extraction module, configured to extract latent noise from the original video based on the pre-trained video generation model, and mix the latent noise with random Gaussian noise to obtain mixed noise; A video inversion module, configured to invert the original video based on the pre-trained video generation model and the mixed noise, and simultaneously extract a spatial attention map and a temporal attention map during the inversion process; A video editing module, configured to use the spatial attention map and the temporal attention map to cover or guide the corresponding attention maps in the denoising process of editing video generation according to the obtained text editing instructions to obtain the edited video; A model update module, configured to update the 3D model to be edited based on the edited video to obtain the edited 3D model.

[0010] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the method for text-driven 3D editing based on generative video prior as described in the first aspect above is implemented.

[0011] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for text-driven 3D editing based on generative video prior as described in the first aspect above is implemented.

[0012] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for text-driven 3D editing based on generative video prior as described in the first aspect above is implemented.

[0013] Based on the above technical solutions, the beneficial effects of the present invention compared with the prior art are as follows: The method and device for text-driven 3D editing based on generative video prior provided by the present invention introduce a generative video prior, that is, use a pre-trained video generation model to capture and understand the spatio-temporal structure and dynamic changes of 3D scenes. This helps to ensure that the 3D content under different perspectives can maintain consistency and coherence during the editing process. The spatial attention map and temporal attention map extracted during the inversion process help to maintain the continuity and consistency of the video content. This makes the edited video more natural and smooth visually. It is possible to update the 3D model to be edited based on the edited video content, thereby achieving precise editing of the 3D model. Compared with traditional video editing methods, this method does not require cumbersome manual operations or complex algorithm designs. It uses technologies such as pre-trained video generation models and attention mechanisms to achieve fast and precise editing of video content.

[0014] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0015] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of a text-driven 3D editing method based on a generative video prior provided by the present invention.

[0018] Figure 2 It is a schematic ablation illustration provided by the present invention.

[0019] Figure 3 It is a schematic structural diagram of a pre-trained generative video model provided by the present invention.

[0020] Figure 4 It is a schematic structural diagram of a text-driven 3D editing device based on a generative video prior provided by the present invention.

[0021] Figure 5 It is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.

[0023] The following will describe Figures 1 - 4 the text-driven 3D editing method and device based on a generative video prior of the present invention.

[0024] The text-driven 3D editing method based on generative video prior utilizes a pre-trained generative video model, which has powerful video generation capabilities. This capability can be regarded as a kind of prior knowledge, that is, the model has mastered how to generate coherent and realistic video content during the learning process. Through this prior knowledge, the method can maintain the inter-frame consistency of the video during the editing process, ensuring that the edited video is still visually coherent and natural. The method edits according to specific text editing instructions, which means that users can guide the video editing process by inputting text instructions, thus achieving personalized editing effects. The text-driven method makes the editing process more intuitive and flexible. Users can adjust the editing instructions at any time according to their needs to obtain satisfactory editing results and spread the editing effects throughout the video. In this process, the text editing instructions play a key guiding role, ensuring the consistency and coherence of the editing effects in the video. The method finally uses the edited video frames to update the original 3D model, thereby obtaining the edited 3D model. This means that the method not only focuses on the editing of video content but also on the update and change of the 3D model. In this way, the method can achieve two-way editing from video to 3D model, further expanding the flexibility and application scenarios of editing. During the editing process, the method captures key information and important regions in the video by extracting spatial attention maps and temporal attention maps. These attention maps are used to cover the corresponding attention maps in the denoising process of the edited video generation, thus ensuring that the edited video is consistent with the original video both spatially and temporally. This precision is crucial for 3D editing because it can maintain the integrity and accuracy of the structure and details of the 3D model during the editing process.

[0025] Refer to Figure 1 As shown, the text-driven 3D editing method based on generative video prior includes the following: Step S110, obtain a pre-trained generative video model and a 3D model to be edited.

[0026] The above-mentioned pre-trained video generation model is an inversion-based video editing model that can generate videos or video frames from the latent space and capture the statistical laws and structural information of video data. The model has been trained with a large amount of data and can generate high-quality videos. The pre-trained video generation model adopts a deep learning architecture, which can be a pre-trained Singular Value Decomposition (SVD) model. The SVD model is a video generation model trained based on the latent diffusion principle, which can generate coherent videos from given static images. During the training process, the SVD model learns the spatio-temporal structure and statistical laws of videos from a large-scale video dataset, thereby acquiring the ability to generate high-quality videos. The pre-trained video generation model can also be a model based on the Variational Autoencoder (VAE). After sufficient training, it can also be regarded as a pre-trained video generation model. These models have broad application prospects in the fields of video generation, video repair, video super-resolution, etc.

[0027] The video generation model is trained using a large amount of high-quality video data, which covers a wide range of scenarios, lighting conditions, and action variations. During the training process, the model learns to capture the statistical laws and structural information of video data, such as the shape of objects, motion trajectories, and texture features. During the training process, the model also learns to invert the original latent noise or features from the generated videos, which is a key step for subsequent video editing.

[0028] The above-mentioned 3D model to be edited is a 3D model that the user hopes to edit and can be in any 3D file format, such as OBJ, FBX, or STL. These files contain 3D geometric information (such as vertices, edges, and faces) and possible texture and material information. Before editing, the 3D model can be preprocessed, such as removing redundant data, optimizing the geometric structure, or adjusting the texture resolution.

[0029] Step S120: Capture training perspective images of the 3D model to be edited from multiple consecutive perspectives, and generate an original video based on the training perspective images of multiple consecutive perspectives.

[0030] Use a camera or other image capture device to capture images of the 3D model to be edited from multiple consecutive perspectives. These perspectives should be evenly distributed around the 3D model to ensure that complete information of the model is captured. Camera parameters (such as focal length, exposure time, and white balance) should be adjusted as needed. At each perspective, capture a series of consecutive images to form video frames. Ensure that the time interval between images is consistent to maintain the smoothness of the video.

[0031] Arrange the captured training perspective images in camera coordinates and chronological order to form an image sequence. Use video encoding software (such as FFmpeg) to encode the image sequence into a raw video. Ensure that the video encoding format is compatible with the pre-trained video generation model, and adjust parameters such as video resolution, frame rate, and bitrate to meet the requirements of subsequent processing. This video will serve as a benchmark for subsequent editing.

[0032] Step S130: Extract latent noise from the raw video based on the pre-trained video generation model, and mix the latent noise with random Gaussian noise to obtain mixed noise.

[0033] Input the raw video into the pre-trained video generation model. The model analyzes the input video and extracts the latent noise in the video. This noise reflects the statistical characteristics and structural information of the video data, such as the edges, textures, and dynamic changes of objects. The extracted latent noise is represented in the form of vectors or matrices and stored in the computer memory for subsequent processing. Use a random number generator to generate random Gaussian noise with the same dimension as the latent noise. Mix the extracted latent noise with the random Gaussian noise. The mixing ratio can be adjusted according to specific requirements to balance the diversity and stability of the video. Store the mixed noise in the computer memory for use in subsequent video inversion and editing processes.

[0034] Step S140: Invert the raw video based on the pre-trained video generation model and the mixed noise, and simultaneously extract the spatial attention map and the temporal attention map during the inversion process.

[0035] Use the pre-trained video generation model and the mixed noise to invert the raw video, that is, try to recover the content of the raw video from the mixed noise. During the inversion process, simultaneously extract the spatial attention map and the temporal attention map. Specifically, input the mixed noise into the generator of the pre-trained video generation model. The generator recovers the content of the raw video from the mixed noise and tries to retain the potential spatial and temporal information. Use the discriminator to evaluate the generated video to ensure that it is consistent with the raw video in terms of statistical laws and structural information. Introduce an attention mechanism in the generator to extract the spatial attention map and the temporal attention map. The spatial attention map reflects the importance or salience of different regions in the video frame. By calculating the weights of each pixel or region, the areas that need to be edited can be located. The temporal attention map reflects the temporal dependence and continuity between video frames. By calculating the weights between adjacent frames, the smoothness of the edited video can be maintained. Store the extracted spatial and temporal attention maps in the computer memory for use in subsequent video editing processes.

[0036] Step S150: According to the obtained text editing instructions, use the spatial attention map and the temporal attention map to cover or guide the corresponding attention map in the denoising process of the edited video generation, and obtain the edited video; According to the user's editing requirements, select appropriate editing operations, such as color adjustment, texture replacement, or action modification. Utilize the extracted spatial and temporal attention maps to cover or guide the corresponding regions during the editing process, thereby achieving precise editing of the video content. For example, when adjusting colors, only focus on the regions with higher weights in the spatial attention map; when modifying actions, maintain the continuity between frames with higher weights in the temporal attention map. During the editing process, preview the editing results in real time and adjust the editing parameters as needed. Combine the edited video frames into a complete video. Perform post-processing on the synthesized video, such as denoising, sharpening, or color correction, to improve the video quality. Output the edited video in common video file formats (such as MP4, AVI, or MKV) for playback on players or social media. This video is similar to the original video in content but has been modified in certain aspects (such as color, texture, action, etc.) according to the user's editing requirements.

[0037] Step S160: Update the to-be-edited 3D model based on the edited video to obtain the edited 3D model.

[0038] Use video processing software or programming tools (such as FFmpeg, OpenCV, etc.) to split the edited video into independent image frames. These image frames now contain the editing effects required by the user, such as color adjustment, texture replacement, or action modification. Obtain the camera position and parameters used when capturing the training perspective images in step S120. This information is crucial for mapping the 2D image frames back to the 3D space. Utilize the camera position information to map the pixels or feature points in each image frame back to the corresponding positions of the original 3D model. According to the mapping results, perform 3D updating to update the geometry, texture, or material of the original 3D model. This can be done by manually adjusting the model in 3D modeling software or using automated tools or scripts to apply these changes. During the updating process, ensure the consistency of the image frames between different perspectives when mapping to the 3D space to avoid misalignment or inconsistent visual effects. Render the updated 3D model using a 3D rendering engine and compare it with the edited video frames to verify the accuracy of the mapping and updating. Output the verified 3D model in common 3D file formats (such as OBJ, FBX, or STL) for use in 3D modeling, animation production, or virtual reality applications.

[0039] The text-driven 3D editing method based on generative video prior provided by the present invention captures and understands the spatio-temporal structure and dynamic changes of 3D scenes by introducing a generative video prior, that is, using a pre-trained video generation model. This helps to ensure that the 3D content under different perspectives can maintain consistency and coherence during the editing process. The spatial attention map and temporal attention map extracted during the inversion process contribute to maintaining the continuity and consistency of the video content. This makes the edited video visually more natural and smooth. By introducing a mixture of generative video prior and random Gaussian noise, this method can generate edited videos with diversity and flexibility. Users can make various modifications to the video according to their needs, such as color adjustment, texture replacement, action modification, etc. This method can update the 3D model to be edited based on the edited video content, thereby achieving precise editing of the 3D model. This provides users with more creative space and possibilities. Compared with traditional video editing methods, this method does not require cumbersome manual operations or complex algorithm designs. It utilizes technologies such as pre-trained video generation models and attention mechanisms to achieve fast and precise editing of video content, thereby reducing the complexity and cost of video editing.

[0040] Moreover, the text-driven 3D editing method based on generative video prior adopts an efficient editing process. Users only need to provide editing instructions, and the system can automatically edit the original 3D according to these instructions. This avoids the cumbersome process of manual operation and software iteration required by traditional methods. This method uses a pre-trained video generation model to accelerate the editing process. These models have been trained on large-scale datasets and possess powerful generation and editing capabilities. Therefore, during the editing process, the prior knowledge of these models can be fully utilized to quickly generate and modify 3D content. Since this method adopts an efficient editing process and pre-trained models, the number of iterations and computational volume can be significantly reduced. This helps to shorten the editing time and improve the editing efficiency. This enables ordinary users to easily perform 3D editing tasks without the need for professional knowledge and skills.

[0041] In an optional embodiment, the step of capturing training perspective images of the 3D model to be edited from multiple consecutive perspectives in step S120 and generating an original video based on the training perspective images of multiple consecutive perspectives includes: S1201. Capture training perspective images of the 3D model to be edited from multiple consecutive perspectives based on a pre-established original camera trajectory, and these images will be used for subsequent video generation.

[0042] In this step, according to a pre-planned original camera trajectory, training perspective images of the 3D model to be edited are captured from multiple consecutive perspectives. These images are the basic materials for subsequent video generation. To achieve this, a system capable of precisely controlling the camera position and orientation is required to ensure that high-quality images can be captured from the required perspectives. These images cover all key features and details of the 3D model for use in subsequent video generation and editing processes.

[0043] S1202. Obtain the camera coordinates corresponding to each training perspective image, sort all the training perspective images according to the camera coordinates, and obtain a sorted video frame sequence. Calculate the physical distance between adjacent camera coordinates in the sorted video frame sequence, and determine whether there is discontinuity between adjacent camera coordinates based on the physical distance.

[0044] Obtain the camera coordinates corresponding to each training perspective image. This coordinate information is crucial for understanding the camera's movement trajectory. Sort all the training perspective images according to the camera coordinates to obtain a sorted video frame sequence consistent with the camera's movement trajectory. After sorting, further calculate the physical distance between adjacent camera coordinates in the sorted video frame sequence. Through this physical distance, it can be detected whether there is discontinuity between adjacent frames, that is, whether the camera movement suddenly changes direction or speed, which may cause jumps or missing parts in the generated video. By comparing the physical distances between adjacent camera coordinates, these potential problem areas can be identified. If it is found that the distance between some adjacent cameras far exceeds the set threshold, it is considered that there is discontinuity between these two frames.

[0045] S1203. For each pair of adjacent camera coordinates with discontinuity, perform spherical interpolation on the rotation matrix of the adjacent camera coordinates and perform linear interpolation on the translation matrix of the adjacent camera coordinates to obtain the corresponding intermediate camera trajectory.

[0046] For each pair of adjacent camera coordinates with detected discontinuity, perform camera interpolation (CameraInterpolation), use spherical interpolation to process its rotation matrix, and use linear interpolation to process its translation matrix at the same time. These two interpolation methods are used to smooth the rotation and translation movements of the camera respectively. Through interpolation, additional intermediate camera trajectory points can be generated in the discontinuous area, and these points will help to smoothly transition the camera movement, thus eliminating the sense of jump. After obtaining all the necessary intermediate camera trajectory points, integrate them into the original camera trajectory to obtain a continuous camera trajectory. This continuous camera trajectory now contains all the necessary motion information and can be used to generate a high-quality image sequence. Refer to Figure 3As shown, black represents the original pose and red represents the interpolated poses.

[0047] Specifically, convert the rotation matrix to a quaternion, perform spherical linear interpolation on the quaternions between adjacent cameras with discontinuities to generate intermediate poses. Convert the interpolated quaternions back to the rotation matrix. Perform linear interpolation on the position vectors between adjacent cameras with discontinuities to generate intermediate positions. Integrate the generated intermediate poses and intermediate positions into the original camera trajectory to obtain a continuous and smooth camera trajectory. The continuity of the trajectory can be verified by calculating the distances and angles between adjacent interpolated cameras.

[0048] S1204. Integrate each of the intermediate camera trajectories into the original camera trajectory to obtain a continuous camera trajectory; input the continuous camera trajectory into a rendering engine to generate an image sequence. These image sequences will form the final video.

[0049] Input this continuous camera trajectory into a rendering engine (such as OpenGL, DirectX, etc.). Each camera pose and position corresponds to a rendering view. Utilize the powerful functions of the rendering engine to render each camera pose to generate corresponding images, obtaining an image sequence. These image sequences will be arranged and rendered according to the movement trajectory of the camera, thereby ensuring visual consistency and coherence of the final video.

[0050] S1205. Use a video encoding library to encode the image sequence generated by rendering into a video file to obtain the original video, which ensures the quality and compatibility of the video.

[0051] Use a video encoding library (such as FFmpeg) to encode the image sequence generated by rendering into a video file. This process involves converting a series of static images into a dynamic video file while ensuring the quality and compatibility of the video. The video encoding library provides the necessary algorithms and tools to compress and optimize video data for playback on different devices and platforms. During the encoding process, appropriate video format, resolution, frame rate, and other parameters can be selected to meet specific playback requirements. Once the encoding is completed, a high-quality original video file is obtained, which contains the training perspective images of the 3D model to be edited captured from multiple consecutive perspectives and has been smoothed and optimized to ensure the viewing experience.

[0052] By capturing images from multiple consecutive perspectives, the present invention can obtain comprehensive information of the 3D model to be edited from different perspectives, providing a rich data basis for subsequent video generation. The multi-angle image capture helps to present the three-dimensional sense and dynamic effects of the 3D model in the final video, enhancing the audience's immersion and visual experience. By sorting the training perspective images, the coherence of the video frame sequence can be ensured, thus generating a smooth video. Detecting the discontinuity between adjacent camera coordinates helps to timely discover and handle potential perspective jumps or image misalignments, avoiding abrupt switches in the final video. Through spherical interpolation and linear interpolation, a smooth intermediate camera trajectory can be generated between adjacent camera coordinates with discontinuities, ensuring a smooth transition of the video perspective. The interpolation method can retain the key features of the original camera trajectory while filling in the missing or jumping parts, making the generated camera trajectory more continuous and realistic. Integrating the intermediate camera trajectory into the original camera trajectory can ensure the continuity of the entire camera trajectory, providing a basis for generating high-quality videos. Based on the continuous camera trajectory, the rendering engine can generate a high-quality image sequence, which will visually present a smooth transition and realistic effects. Using a video encoding library to encode the image sequence into a video file can ensure that the generated video file has wide compatibility and can be played on different devices and platforms. The encoding process can optimize the size and quality of the video file, so that the finally generated video has a small file size while maintaining high image quality and viewing experience.

[0053] In an optional embodiment, mixing the potential noise with random Gaussian noise to obtain mixed noise in step S130 described above includes: S1301. Select an original anchor frame from the original video, and use a text instruction-driven pixel conversion method to edit the original anchor frame to obtain an edited anchor frame.

[0054] Select one or more key frames from the original video as anchor frames. These anchor frames will serve as the starting point for editing, and their selection can be based on video content, editing requirements, and the convenience of subsequent processing.

[0055] Use a text instruction-driven pixel conversion (InsctructPix2Pix) method to edit the anchor frame. This method allows users to describe the desired modifications through natural language, convert the text instructions provided by the user into specific modification operations on the pixels of the anchor frame, and then automatically implement these modifications using image editing techniques (such as diffusion models, generative adversarial networks, etc.) to obtain the edited anchor frame .

[0056] Use EDM Inversion to obtain the latent noise corresponding to the original video. Then, select the anchor frame from the original video and use InsctructPix2Pix to obtain the edited anchor frame that meets the editing instructions. Refer to Figure 2 As shown, Reference is the original video (from the first frame 1 st frame to the last frame Last frame), and the input text editing instruction is Make him smlie, obtaining (a). The generation of video editing uses the latent noise as the starting noise and the edited anchor frame as the condition. As Figure 2 shown in (b), the obtained image has a poor editing effect for large color changes. This is because the latent noise (i.e., the inverted noise) contains a large amount of color information of the original video. To solve this problem, the present invention proposes redundancy decay. Add linearly to the random Gaussian noise, which effectively alleviates the color shift. Figure 2 In (c), W / RR represents the method with redundancy reduction, W / AMO represents the method with attention map overriding, and W / o RR represents the method without redundancy reduction. represents the hyperparameter coefficient in the redundancy reduction method. represents the hyperparameter coefficient in the attention map overriding. Set to 0.6 means that the attention map overriding is only used within the denoising steps of 0.6 - 1, and the attention map overriding is not used within the denoising steps of 0 - 0.6. Figure 2 In (d)-(f), it is mainly for ablation visualization of the effects of RR and AMO. (d) shows that when = 0 in RR, the method of the present invention degenerates into an img2video generation process starting from a random noise, so the video motion of the generated video does not match the source video motion. (e) shows that when AMO acts on all denoising steps, it will be over-smoothed, resulting in artifacts. (f) shows adding RR and AMO and adjusting =SSIM(source image, edited image), The best effect is achieved when it is 0.6. The source image is the original anchor frame, and the edited image is the edited anchor frame.

[0057] S1302. Calculate the structural similarity between the original anchor frame and the edited anchor frame, and determine the mixing weight according to the structural similarity.

[0058] Use image quality assessment algorithms such as Structural Similarity (SSIM) to calculate the similarity between the original anchor frame and the edited anchor frame. The SSIM algorithm comprehensively considers the brightness, contrast, and structural information of the image, and can more accurately reflect the perceptual differences between images. By calculating the SSIM value, the impact degree of the editing operation on the anchor frame can be quantified, providing a basis for the subsequent weight allocation of the mixed noise. According to the calculated SSIM value, determine the mixing weight of the potential noise and the random Gaussian noise. The higher the SSIM value, the more similar the edited anchor frame is to the original anchor frame. Therefore, a smaller mixing weight can be assigned to the potential noise; otherwise, a larger mixing weight is assigned. Through reasonable weight allocation, the desired editing effect can be achieved while maintaining the image quality.

[0059] S1303. Mix the potential noise and the random Gaussian noise according to the mixing weight to obtain the mixed noise.

[0060] According to the determined mixing weight , , refer to Figure 3 as shown, mix the potential noise (Inverted noise ) and the random Gaussian noise (Random roise) Z to obtain the mixed noise (Redundancy-reduced noise) . The potential noise may come from the inherent errors of image processing, interference during transmission, etc., while the random Gaussian noise is used to simulate the random fluctuations in nature. By mixing these two kinds of noises, a mixed noise that contains both the characteristics of the potential noise and certain randomness can be obtained. This mixed noise can be used in subsequent image processing tasks to enhance the robustness of the image or achieve specific visual effects.

[0061] The present invention drives a pixel conversion method through text instructions, enabling users to more intuitively control the image editing process and achieve diverse editing effects. This method reduces the dependence on professional image processing knowledge, allowing non-professionals to easily perform image editing. By calculating the structural similarity between the original anchor frame and the edited anchor frame and determining the blending weight based on the similarity, the desired editing effect can be achieved while maintaining the image quality. This method avoids problems such as image distortion or quality degradation caused by over-editing. Mixing potential noise with random Gaussian noise can enhance the image's adaptability to noise. This mixed noise can be used to simulate image noise situations in different environments, providing richer data support for subsequent image processing tasks. This process is not only applicable to the field of video editing but can also be extended to multiple fields such as image processing, computer vision, and augmented reality. By combining different image processing techniques and algorithms, more diverse application scenarios and effects can be achieved.

[0062] In an optional embodiment, selecting the original anchor frame from the original video in step S1301 includes: S13011. Obtain a text prompt and calculate the graphic-text matching score between each frame image in the original video and the text prompt. Select the image with the highest graphic-text matching score as the original anchor frame.

[0063] First, receive the text prompt provided by the user. These text prompts usually contain descriptions of video content, keywords, or information about specific scenarios, and are used to guide subsequent video frame selection. The accuracy and specificity of the text prompt are crucial for the success of subsequent steps. Traverse each frame image in the original video and calculate the graphic-text matching score between these images and the text prompt. The calculation of the graphic-text matching score can be based on various factors, such as the matching degree between the objects and scenes in the image and the keywords in the text prompt, and whether the emotional color of the image is consistent with the emotional tendency of the text prompt. After calculating the graphic-text matching scores of all frames, select the image with the highest matching score as the original anchor frame. This frame image is considered to best represent the content or scenario described by the text prompt.

[0064] Specifically, first, the text prompt provided by the user is received. Next, the graphic-text matching scores between each frame image in the original video and the text prompt are calculated. The Contrastive Language–Image Pre-training (CLIP) model is adopted in the present invention to calculate the scores. The CLIP model is a method for multi-modal vision and text learning. It can learn a shared latent space from a large number of image-text pairs, such that similar images and texts have similar representations in this space. Specifically, each frame image in the video and the text prompt are respectively input into the image encoder and text encoder of the CLIP model to obtain their representations in the latent space. Then, the cosine similarity or Euclidean distance between these representations is calculated as the graphic-text matching score (CLIP score). The higher the score, the higher the matching degree between the image and the text prompt. After calculating the graphic-text matching scores for all frames, the frame image with the highest matching score is found. This frame image is considered to be the video frame that best represents the content described by the text prompt. Finally, this frame image is selected as the original anchor frame. The original anchor frame will play a key role in subsequent video editing or processing, such as serving as the starting point for editing, a reference frame, or a key frame, etc.

[0065] By selecting the image with the highest graphic-text matching score as the original anchor frame, the present invention can ensure that the editing operation targets the most relevant or representative content in the video. This helps to improve the accuracy and pertinence of video editing. This process allows the user to guide the selection of video frames through text prompts, thereby enhancing the interactivity between the user and the video editing system. The user can more intuitively express their editing intentions and obtain more satisfactory editing results. The application of the graphic-text matching algorithm enables the system to quickly and accurately find the video frame that best matches the text prompt. This avoids the cumbersome process of manually browsing and selecting video frames, thus improving the efficiency of video editing.

[0066] In an optional embodiment, at the initial stage of video editing, latent noise is extracted from the original video through EDM inversion. This latent noise can be regarded as the representation of the video content in the latent space. By utilizing this latent representation, the overall structure and style of the video content can be maintained during the editing process, while allowing for fine editing of details.

[0067] Refer to Figure 3As described above, the pre-trained video generation model in step S140 includes an encoder (VAE Encoder) and an inversion module (EDM Inversion). The encoder is responsible for converting the video from the original space to the latent space, while the inversion module is responsible for restoring the video from the latent space. These two modules together constitute a bridge between the latent space and the original space of the video. Based on this model, multi-frame consistent editing can be achieved.

[0068] Taking the pre-trained video generation model SVD (Stable Video Diffusion) as an example, the design and implementation of the inversion module need to consider the characteristics and outputs of the SVD model. The SVD model is a video generation model trained based on the principle of latent diffusion. It can generate coherent video sequences from given static images. During the inversion process, the goal is to restore the original information or perform some form of optimization from the videos or latent representations generated by the SVD model.

[0069] The above inversion module can be an inversion module based on latent space optimization. The SVD model performs video generation in the latent space, so the inversion module can first attempt to optimize in the latent space. This structure of the inversion module includes the following parts: a latent representation extraction part, a latent space search part, and a latent representation decoding part. The latent representation extraction part is used to extract the latent representation from the videos generated by the SVD model, which can be achieved through an encoder or a similar structure. The latent space search part is used to search for the representation in the latent space that is most similar to the original video. This can be achieved through iterative optimization algorithms such as gradient descent. In each iteration, the parameters in the latent representation are adjusted according to the quality of the currently restored video. The latent representation decoding part is used to decode the optimized latent representation back into video form, which can be achieved through the decoder of the SVD model or a similar structure.

[0070] The SVD model utilizes a spatial attention mechanism and a temporal attention mechanism when generating videos. The above inversion module also includes: an attention map extraction part, an attention map optimization part, and a video restoration part based on the attention map. The attention map extraction part is used to extract the spatial attention map and the temporal attention map from the SVD model. These maps can capture the importance of different regions within the video frames and the temporal correlation between video frames. The attention map optimization part is used to optimize the extracted attention maps to make them more accurately reflect the information in the original video. This can be achieved through iterative optimization algorithms or deep learning models. The video restoration part based on the attention map is used to guide the video restoration process using the optimized attention maps. For example, the pixel values of different regions can be weighted according to the spatial attention map, or the transition between video frames can be smoothed according to the temporal attention map.

[0071] The above inversion module includes a generator and a discriminator. The generator is responsible for generating a restored video from a latent representation or an optimized attention map. The structure of the generator can be similar to the decoder of the SVD model. The discriminator is responsible for distinguishing between real videos and restored videos. The structure of the discriminator can be designed according to specific tasks and needs to be able to capture key features in the video. Through adversarial training between the generator and the discriminator, the quality of the restored video is gradually improved. This usually involves an iterative optimization process, where the generator tries to deceive the discriminator, while the discriminator tries to more accurately identify real videos and restored videos.

[0072] In addition to the above inversion module with a single structure, it is also possible to consider combining multiple structures to fully utilize their respective advantages. For example, an inversion module based on latent space optimization can be combined with an inversion module based on the attention mechanism. This hybrid-structured inversion module has stronger robustness and higher restoration quality.

[0073] Inverting the original video based on the pre-trained video generation model and the hybrid noise, and simultaneously extracting spatial attention maps and temporal attention maps during the inversion process, includes: S1401. Adding the hybrid noise to the original video to generate a noisy video.

[0074] Select multiple types of noise for mixing. These noises can include Gaussian noise (a type of random noise that simulates random interference in signal transmission), latent noise (noise sampled from the latent space, used to simulate potential changes in video content), and other possible noise types (such as salt-and-pepper noise, Poisson noise, etc., selected according to specific application scenarios). Mix these noises according to a certain ratio or distribution to generate hybrid noise with complex characteristics. Add the hybrid noise frame by frame to each frame image of the original video. The way of adding noise can be direct superposition, multiplication, or other mathematical operations, depending on the noise type and the desired effect. The video after adding noise will have a different appearance from the original video, but its basic structure and content information should be retained. The intensity and type of noise should be adjusted according to specific applications and expected effects.

[0075] S1402. Inputting the noisy video into the encoder (VAE Encoder) to convert it into a preliminary representation in the latent space.

[0076] Input the video added with hybrid noise into the encoder of the pre-trained video generation model. The encoder will process the video and convert it into a preliminary representation in the latent space. This preliminary representation is a low-dimensional and efficient representation form of the video in the latent space.

[0077] Specifically, before inputting the noisy video into the encoder, some preprocessing operations can be performed, such as adjusting the video size, frame rate, color space, etc., to ensure the compatibility of the video with the encoder. The preprocessed noisy video is input into the encoder of the pre-trained video generation model. The encoder will perform multi-level feature extraction and transformation on the video, converting it from the original space to the latent space. In the latent space, the video is represented in the form of a low-dimensional and efficient vector or tensor, which is the preliminary representation. The structure and parameters of the encoder should be selected according to the pre-trained video generation model.

[0078] S1403. Input the preliminary representation into the inversion module for optimization to obtain the latent representation of the noisy video; wherein, during the inversion process, the spatial attention mechanism and the temporal attention mechanism are used to extract the spatial attention map and the temporal attention map respectively.

[0079] Input the encoded preliminary representation into the inversion module for optimization. The task of the inversion module is to restore a video as similar as possible to the original video from the preliminary representation in the latent space. During the optimization process, the spatial attention mechanism and the temporal attention mechanism are used to extract the spatial attention map and the temporal attention map respectively. The spatial attention mechanism focuses on the importance of different regions within a video frame and highlights the key regions by calculating the spatial attention map. This helps to maintain the spatial structure of the video during the editing process. The temporal attention mechanism focuses on the temporal correlation between video frames and captures the dynamic information in the video by calculating the temporal attention map. This helps to maintain the temporal continuity of the video during the editing process.

[0080] Specifically, input the preliminary representation into the inversion module and initialize the inversion process. During the inversion process, the spatial attention mechanism and the temporal attention mechanism are used to extract the spatial attention map and the temporal attention map respectively. The spatial attention mechanism focuses on the importance of different regions within a video frame and highlights the key regions by calculating the spatial attention map. This helps to maintain the spatial structure of the video during the editing process. The temporal attention mechanism focuses on the temporal correlation between video frames and captures the dynamic information in the video by calculating the temporal attention map. This helps to maintain the temporal continuity of the video during the editing process. According to these attention maps, the inversion module will adjust the parameters in the preliminary representation to gradually optimize the restored video. The optimization process may involve multiple iterations, and each iteration will be adjusted according to the quality of the currently restored video. When the optimization process reaches a predetermined stop condition (such as reaching the maximum number of iterations, the quality of the restored video meets the requirements, etc.), the inversion module will output the final latent representation. This latent representation contains both the basic structure and content information of the original video and has been optimized through the interference of the mixed noise and the encoding-inversion process. The structure and parameters of the inversion module can be selected according to the pre-trained video generation model.

[0081] The present invention can achieve local adjustment of video content without affecting the overall structure and style by extracting potential noise as a representation of the video content and using spatial attention mechanism and temporal attention mechanism for fine editing. Editing using the representation in the latent space can significantly reduce the computational complexity and improve the efficiency of video editing. At the same time, since the representation in the latent space has the characteristics of low-dimensional and high-efficiency, it can be more easily stored and transmitted. This method allows users to perform diverse editing on video content, such as adding special effects, changing colors, adjusting frame rates, etc. At the same time, due to the use of spatial attention mechanism and temporal attention mechanism, complex scenes and dynamic information in the video can be processed more flexibly. Through EDM inversion and the representation in the latent space, this method can achieve stable editing of video content. Even in the case of substantial modification of the video, the overall consistency and coherence of the video can be maintained.

[0082] In an optional embodiment, redundancy decay will to some extent destroy the dynamic information of the original video, as shown in (c) of Figure 2 To solve this problem, the present invention proposes attention map overriding, which extracts the spatial attention map and temporal attention map obtained in the inversion process for overriding the corresponding attention maps in the denoising process of the edited video generation. This can effectively ensure the alignment of the dynamics with the original video. The pre-trained video generation model further includes a decoder and a denoising network; the step of obtaining the edited video according to the acquired text editing instruction in step S150 by using the spatial attention map and the temporal attention map to override or guide the corresponding attention maps in the denoising process of the edited video generation includes: S1501. Edit the latent representation of the noisy video in the latent space according to the text editing instruction to obtain an edited latent representation.

[0083] In this step, the latent space representation of the pre-trained video generation model (such as SVD) is used. This latent space contains low-dimensional and abstract features of the video data, which can capture the main content and structure of the video. According to the text editing instruction provided by the user (such as "Turn him into a clown" etc.), the latent representation of the noisy video is edited accordingly in the latent space. This editing can be a direct numerical modification of the latent vector or a certain form of interpolation or transformation in the latent space.

[0084] S1502. Input the edited latent representation into the decoder (Video Decoder) to generate an initial edited video.

[0085] After the latent space editing, the edited latent representation is obtained. Next, this edited latent representation is input into the decoder of the pre-trained video generation model. The role of the decoder is to convert the latent representation back into high-dimensional video data, that is, to generate an initial edited video. This initial edited video may still contain some noise or artifacts because it directly comes from the editing result of the latent space and has not undergone subsequent denoising processing.

[0086] S1503. Input the initial edited video into a denoising network. The denoising network uses the spatial attention map and the temporal attention map as additional input information to overwrite or guide the corresponding attention maps in the denoising process of the edited video, and outputs an edited video.

[0087] In this step, the initial edited video is input into a denoising network. This denoising network is a deep learning model specifically designed for video denoising, such as the attention-based image denoising network ADNet, etc. The denoising network uses the spatial attention map and the temporal attention map as additional input information. These attention maps can capture the importance of different regions within video frames and the temporal correlation between video frames. By overwriting or guiding the corresponding attention maps in the denoising process of the edited video, the denoising network can more effectively remove the noise and artifacts in the video while keeping the main content and structure of the video unaffected. Finally, the denoising network outputs the edited video, which contains both the user-specified editing content and has high quality.

[0088] Specifically, the user inputs an initial edited video into the denoising network. This video may contain the editing content specified by the user, but may also be interfered by noise and artifacts. Before the video is input into the network, some preprocessing operations, such as format conversion, size adjustment, etc., are required to ensure that the video meets the input requirements of the network. The denoising network first extracts a spatial attention map and a temporal attention map from the input video. The spatial attention map can reflect the importance of different regions within a video frame, while the temporal attention map can capture the temporal correlation between video frames. Using the extracted spatial attention map and temporal attention map, the denoising network starts to perform denoising processing on the input video. During the processing, the network will targetedly remove the noise and artifacts in the video according to the information in the attention map. During the denoising process, the attention map also plays a role in covering or guiding the editing. Specifically, the network will determine which regions are the main content of the video and need to be retained, and which regions are noise or artifacts and need to be removed according to the information in the spatial attention map. At the same time, the temporal attention map helps the network understand the temporal correlation between video frames, so as to more accurately remove the temporal noise and artifacts. While denoising, the denoising network also tries to keep the main content and structure of the video unaffected. This is achieved by precisely controlling the denoising process and using the key information provided by the attention map. After the above processing flow, the denoising network will output a denoised video. This video contains both the editing content specified by the user and has high quality, and the noise and artifacts are effectively removed. The user can further edit and process (Videorefinement) the output high-quality video to meet specific requirements.

[0089] By performing editing in the latent space in the present invention, we can avoid performing complex operations directly on high-dimensional video data, thereby greatly improving the editing efficiency. Using the denoising network and the attention mechanism, we can effectively remove the noise and artifacts introduced during the editing process while keeping the main content and structure of the video unaffected. Since the editing is performed in the latent space, various complex editing tasks, such as color change, object addition or deletion, text caption addition, etc., can be flexibly achieved. Combining the text editing instructions and the latent space editing technology, cross-modal video editing tasks can be realized, that is, modifying the video content according to the user's text description.

[0090] The text-driven 3D editing device based on the generative video prior provided by the present invention will be described below. The text-driven 3D editing device based on the generative video prior described below can be mutually corresponding and referred to the text-driven 3D editing method described above.

[0091] The text-driven 3D editing device based on the generative video prior provided by the present invention, refer to Figure 4As shown, it includes: A data acquisition module 210, configured to acquire a pre-trained video generation model and a three-dimensional model to be edited; the pre-trained video generation model is an inversion-based video editing model; A video generation module 220, configured to capture training perspective images of the three-dimensional model to be edited from multiple consecutive perspectives, and generate an original video based on the training perspective images of the multiple consecutive perspectives; A noise extraction module 230, configured to extract latent noise from the original video based on the pre-trained video generation model, and mix the latent noise with random Gaussian noise to obtain mixed noise; A video inversion module 240, configured to invert the original video based on the pre-trained video generation model and the mixed noise, and simultaneously extract a spatial attention map and a temporal attention map during the inversion process; A video editing module 250, configured to, according to the acquired text editing instruction, use the spatial attention map and the temporal attention map to cover or guide the editing of the corresponding attention map during the denoising process of the generated video, so as to obtain an edited video; A model update module 260, configured to update the three-dimensional model to be edited based on the edited video to obtain an edited three-dimensional model.

[0092] Figure 5 An example of a schematic physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete mutual communication through the communication bus 340. The processor 310 may call logical instructions in the memory 330 to execute a text-driven three-dimensional editing method based on a generative video prior.

[0093] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0094] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the text-driven 3D editing method based on the generative video prior provided by the above-mentioned various methods.

[0095] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the text-driven 3D editing method based on the generative video prior provided by the above-mentioned various methods.

[0096] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A text-driven 3D editing method based on generative video priors, characterized in that: include: Obtaining a pre-trained generated video model and a three-dimensional model to be edited; the pre-trained generated video model is a video editing model based on inversion; Capturing training view images of the to-be-edited three-dimensional model from a plurality of continuous view angles, and generating an original video based on the training view images from the plurality of continuous view angles; Extracting potential noise from the original video based on the pre-trained generated video model, and mixing the potential noise with random Gaussian noise to obtain mixed noise; Inverting the original video based on the pre-trained generated video model and the mixed noise, and extracting a spatial attention map and a temporal attention map simultaneously during the inversion process; According to the acquired text editing instruction, the spatial attention map and the temporal attention map are used to cover or guide the corresponding attention map in the denoising process of the edited video to obtain an edited video; The to-be-edited three-dimensional model is updated based on the edited video to obtain an edited three-dimensional model.

2. The text-driven 3D editing method based on generative video prior according to claim 1, characterized in that: The step of capturing training view images of the to-be-edited three-dimensional model from a plurality of continuous view angles and generating an original video based on the training view images of the plurality of continuous view angles comprises: Capturing training perspective images of the to-be-edited three-dimensional model from a plurality of continuous perspectives based on a pre-established original camera trajectory; Obtain the camera coordinates corresponding to each training view image, sort all the training view images according to the camera coordinates, and obtain a sorted video frame sequence; Calculating a physical distance between adjacent camera coordinates in the sorted video frame sequence, and determining whether there is discontinuity between adjacent camera coordinates based on the physical distance; For each adjacent camera coordinate having discontinuity, the rotation matrix of the adjacent camera coordinate is spherically interpolated, and the translation matrix of the adjacent camera coordinate is linearly interpolated to obtain a corresponding intermediate camera trajectory; Integrate each of the intermediate camera trajectories into the original camera trajectory to obtain a continuous camera trajectory; Inputting the continuous camera trajectory into a rendering engine to generate an image sequence; The rendered image sequence is encoded into a video file using a video encoding library to obtain the original video.

3. The text-driven 3D editing method based on generative video prior according to claim 1, characterized in that: The step of mixing the potential noise with random Gaussian noise to obtain mixed noise includes: Selecting an original anchor frame from the original video, and editing the original anchor frame using a text instruction driven pixel conversion method to obtain an edited anchor frame; Calculating the structural similarity between the original anchor frame and the edited anchor frame, and determining a mixing weight according to the structural similarity; The latent noise is mixed with random Gaussian noise according to the mixing weight to obtain the mixed noise.

4. The text-driven 3D editing method based on generative video prior according to claim 3, characterized in that: The selecting an original anchor frame from the original video comprises: Get the text prompt and calculate the image-text matching score between each frame of the original video and the text prompt; The image with the highest image-text matching score is selected as the original anchor frame.

5. The text-driven 3D editing method based on generative video prior according to claim 1, characterized in that: The pre-trained video generation model includes an encoder and an inversion module; the inversion of the original video based on the pre-trained video generation model and the mixed noise, and the extraction of the spatial attention map and the temporal attention map during the inversion process, including: Adding the mixed noise to the original video to generate a noisy video; Inputting the noisy video into the encoder and converting it into a preliminary representation of the latent space; The preliminary representation is input into the inversion module for optimization to obtain a potential representation of the noisy video; wherein, during the inversion process, a spatial attention mechanism and a temporal attention mechanism are used to extract a spatial attention map and a temporal attention map respectively.

6. The text-driven 3D editing method based on generative video prior according to claim 5, characterized in that: The pre-trained video generation model also includes a decoder and a denoising network; the method of using the spatial attention map and the temporal attention map to cover or guide the corresponding attention map in the denoising process of the edited video generation according to the acquired text editing instruction to obtain the edited video includes: Editing the latent representation of the noisy video in a latent space according to the text editing instruction to obtain an edited latent representation; Inputting the edited latent representation into the decoder to generate an initial edited video; The initial edited video is input into a denoising network. The denoising network uses the spatial attention map and the temporal attention map as additional input information to cover or guide the edited video to generate a corresponding attention map in the denoising process, and outputs the edited video.

7. A text-driven 3D editing device based on generative video priors, characterized in that: include: A data acquisition module, used to acquire a pre-trained generated video model and a three-dimensional model to be edited; the pre-trained generated video model is a video editing model based on inversion; A video generation module, used to capture training view images of the to-be-edited three-dimensional model from multiple continuous view angles, and generate an original video based on the training view images of the multiple continuous view angles; A noise extraction module, configured to extract potential noise from the original video based on the pre-trained generated video model, and mix the potential noise with random Gaussian noise to obtain mixed noise; A video inversion module, used to invert the original video based on the pre-trained generated video model and the mixed noise, and simultaneously extract a spatial attention map and a temporal attention map during the inversion process; A video editing module, configured to use the spatial attention map and the temporal attention map to cover or guide the corresponding attention map in the denoising process of the edited video according to the acquired text editing instruction, so as to obtain an edited video; The model updating module is used to update the to-be-edited three-dimensional model based on the edited video to obtain the edited three-dimensional model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the text-driven three-dimensional editing method based on generative video prior as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text-driven three-dimensional editing method based on generative video prior as claimed in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the text-driven three-dimensional editing method based on generative video prior as claimed in any one of claims 1 to 6 is implemented.