Interactive three-dimensional scene generation and motion editing method and device, electronic equipment and storage medium
Through the multi-view optical flow estimation and the decoupling representation strategy of the motion diffusion stage, the accuracy and consistency of complex motions in image editing in the prior art are solved, efficient three-dimensional scene generation and motion editing are achieved, and the consumption of computing resources is reduced.
Patent Information
- Application Number
- CN202510032589.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to achieve accurate understanding of spatial content and fine motion editing in image editing, especially in the case of complex motion such as scaling or rotation, and most methods require a large number of computing resources for data set retraining or model fine-tuning.
The multi-view optical flow estimation stage and multi-view motion diffusion stage are used to segment the three-dimensional scenes through MaskClustering, the point cloud dynamics model PKM is used to estimate the optical flow, and the optical flow guidance strategy FGS, hidden space fusion LSF and background grid point constraint BGC are designed to decouple the motion process to guide the diffusion model for editing.
It improves the robustness and versatility of the zero-sample reasoning paradigm, enhances the accuracy and multi-view consistency of motion editing, and reduces dependence on computing resources.
Smart Images

Figure CN120355839A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to an interactive three-dimensional scene generation and motion editing method, apparatus, electronic device, and storage medium. Background Art
[0002] Recently, text-to-image methods based on diffusion models have made rapid progress. Thanks to a large amount of training data, these models have strong image prior properties. Related techniques include: the DiffusionCLIP method, which makes the editing of image content more reasonable by using the powerful image-text matching ability of CLIP. The UniTune method further improves the editing quality by focusing on a single image representation during the fine-tuning stage and sampling image details from the original image. The Dreambooth method fine-tunes a pre-trained text-to-image model by introducing a specific class of prior-preserving loss, enabling it to learn a unique text subject identifier. This unique identifier can be used to synthesize new images of the subject in different scenarios. The SDEdit method utilizes the generative prior of stochastic differential equations and combines it with the diffusion model to iteratively obtain the edited image. The Imagic method fine-tunes the text embedding without changing the backbone diffusion architecture, and then fine-tunes the denoising model with the newly obtained text embedding, capable of obtaining complex (such as non-rigid) text-guided semantic editing results. However, these methods only rely on text and are difficult to achieve accurate understanding of spatial content and fine motion editing.
[0003] In addition, there are also generative image editing methods based on physical priors, which propose to use physical priors (dragging, moving points, optical flow, etc.) for fine spatial motion editing. For example, DragGAN achieves dragging editing through the movement of control points and the point tracking mechanism. However, the generality of GAN-based methods is greatly limited. Related techniques propose the DiffEditor method, which introduces the collaboration of image prompts and text prompts to better describe the editing object. At the same time, the combination of stochastic differential equation and ordinary differential equation sampling improves the consistency and flexibility. However, the above-mentioned methods based on dragging points are good at handling simple translational motions and are difficult to handle complex motions such as scaling or rotation. The MotionGuidance method utilizes optical flow as a prior and guides the diffusion model for various complex motion edits. However, its structure lacks consideration of the texture of moving objects, resulting in a change in the texture of its editing results. The MagicFixup method achieves texture detail fidelity by designing a detail extractor and synthesizer. However, it lacks effective multi-view consistency constraints, resulting in limited multi-view motion editing performance. In addition, most methods require collecting datasets to retrain or fine-tune the diffusion model, which consumes a huge amount of computing resources. Summary of the Invention
[0004] The main objective of the embodiments of the present invention is to propose an interactive three-dimensional scene generation and motion editing method, device, electronic device, and storage medium, which can improve the robustness and generality of the zero-shot inference paradigm and effectively improve the motion accuracy and multi-view consistency during the motion editing process.
[0005] To achieve the above objective, on the one hand, the embodiments of the present invention propose an interactive three-dimensional scene generation and motion editing method, including the following steps:
[0006] Multi-view optical flow estimation stage, specifically: given a static scene and multi-view images under the static scene, use MaskClustering to segment the three-dimensional scene and export the query codes of each object; interactively generate the optical flow of the selected single-view image through a graphical operation interface; estimate the multi-view optical flow through the point cloud dynamics model PKM;
[0007] Multi-view motion diffusion stage, specifically: based on the diffusion inference paradigm without training based on optical flow, guide the diffusion model to complete motion editing by decoupling the representation of the motion process and designing corresponding strategies;
[0008] According to the processing of the multi-view optical flow estimation stage and the multi-view motion diffusion stage, complete the interactive three-dimensional scene generation and motion editing.
[0009] In some embodiments, after obtaining the optical flow of the selected single-view image, use the optical flow to estimate the point cloud P after motion m , and then project the point cloud after motion and the original point cloud P o to obtain the multi-view optical flow. The expression of this process is:
[0010]
[0011] Among them, foc is the camera focal length; [R|T] is the camera rotation and translation matrix; d is the depth value.
[0012] In some embodiments, the point cloud dynamics model PKM is used to calculate the sparse point cloud P after motion according to the single-view optical flow sm , and the expression of this process is:
[0013]
[0014] Among them, K is the camera intrinsic matrix; c x,y is the 2D coordinate position; pp is the optical center of the camera; f sx represents the x component of the optical flow; f sy represents the y component of the optical flow.
[0015] In some embodiments, the method further includes: for different motion modes, constructing corresponding point motion mechanics models, specifically:
[0016] For the translational motion mode, represent the translational motion in three-dimensional space as an offset P between the original point cloud and the post-motion point cloud off , and then calculate the motion point cloud;
[0017] For the scaling motion mode, there is a scaling factor s between P m and P o : for the shrinking motion, use the magnitude of the single-view optical flow to statistically calculate the scaling factor; for the enlarging motion, calculate the area ratio of the collision region to the original region; f : for the shrinking motion, use the magnitude of the single-view optical flow to statistically calculate the scaling factor; for the enlarging motion, calculate the area ratio of the collision region to the original region;
[0018] For the rotational motion mode, based on the same rotation angle θ of all three-dimensional points of the rotational motion, use the spatial rotation matrix Rot and the centroid p of the three-dimensional point cloud c to represent the post-motion point cloud;
[0019] For the stretching motion mode, represent the stretching motion in space as a stretching plane, and the original point cloud is stretched by the stretching plane to different degrees.
[0020] In some embodiments, in the translational motion mode, the calculation formula for the motion point cloud is:
[0021] P m =P o +p off ,
[0022]
[0023] where N(P so ) represents the number of the original sparse point cloud P so ;
[0024] In the scaling motion mode, the expression for using the magnitude of the single-view optical flow to statistically calculate the scaling factor is:
[0025]
[0026] where N represents the number of non-zero optical flow values;
[0027] In the scaling motion mode, the expression for calculating the area ratio of the collision region to the original region is:
[0028]
[0029] o r =L(c x ,c y, f s ) ∪ r(f s ),
[0030] where L represents the linear sampling function, r(f s ) and o r represent the areas of the original and collision regions;
[0031] In the described rotational motion mode, the spatial rotation matrix Rot and the centroid p of the three-dimensional point cloud c are calculated as follows:
[0032] P m = Rot × (P o - p c ) + p c ,
[0033]
[0034] In the described stretching motion mode, the expression for stretching the original point cloud to different degrees is:
[0035] P m = P o + t f × Max(P sm - P so ),
[0036]
[0037] where dis represents the distance from the point cloud Po to the plane "AX + BY + D"; t f represents the ratio of a certain point distance to the maximum point distance.
[0038] In some embodiments, the multi-view motion diffusion stage includes the following steps:
[0039] According to the multi-view optical flow, represent the motion as a combination of a static background, a moving object, and an occlusion area, and decouple these three parts and adopt corresponding strategies to achieve motion editing; specifically:
[0040] Optical flow guidance strategy FGS process: For the static background, use DDIM inversion to replace the output of each step of the diffusion model;
[0041] Latent space fusion LSF process: Use optical flow to perform fusion in the latent space and supervise the texture details of the moving object;
[0042] Background grid constraint BGC process: Perform background grid constraint on the collision area during the diffusion process to maintain the consistency of multi-view images.
[0043] In some embodiments, the process of the optical flow guidance strategy (FGS) is specifically as follows: Use an optical flow estimator (RAFT) to predict the optical flow f between the generated image and the input image p , and calculate the optical flow loss L by comparing it with the input optical flow f i ; Use warp transformation to calculate the loss L between the generated image and the input image flow ; Then take the derivative of these losses to change the predicted noise of the diffusion model: color ;
[0044]
[0045] The process of the latent space fusion (LSF) is specifically as follows: Perform inverse warp transformation on the input image and the optical flow, then use a variational autoencoder (VAE) to encode the warped image and add noise to it; Finally, fuse the noisy image with its mask in the latent space to obtain the noise output at the current moment;
[0046] The process of the background grid constraint (BGC) is specifically as follows: First, perform grid transformation on the output of the diffusion model, then denoise the tensor, and finally use the inverse grid transformation to obtain the final tensor as the final output.
[0047] Another aspect of the embodiments of the present invention also provides an interactive three-dimensional scene generation and motion editing device, including:
[0048] A first module for performing the multi-view optical flow estimation stage, specifically: Given a static scene and multi-view images of the static scene, use MaskClustering to segment the three-dimensional scene and export the query codes of each object; Interactively generate the optical flow of the selected single-view image through a graphical user interface; Estimate the multi-view optical flow through a point cloud kinetics model (PKM);
[0049] A second module for performing the multi-view motion diffusion stage, specifically: Based on a diffusion inference paradigm without training based on optical flow, guide the diffusion model to complete motion editing by decoupling the representation of the motion process and designing corresponding strategies;
[0050] A third module for completing interactive three-dimensional scene generation and motion editing according to the processing of the multi-view optical flow estimation stage and the multi-view motion diffusion stage.
[0051] To achieve the above object, another aspect of the embodiments of the present invention proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method described above.
[0052] To achieve the above object, on the other hand, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method described above.
[0053] An embodiment of the present invention also discloses a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the method described above.
[0054] The embodiments of the present invention at least include the following beneficial effects: The present invention provides an interactive three-dimensional scene generation and motion editing method, device, electronic device, and storage medium. The solution includes a multi-view optical flow estimation stage, specifically: given a static scene and multi-view images under the static scene, using MaskClustering to segment the three-dimensional scene and export the query codes of each object; generating the optical flow of the selected single-view image through a graphical operation interface interaction; estimating the multi-view optical flow through a point cloud dynamics model PKM; a multi-view motion diffusion stage, specifically: based on a diffusion inference paradigm without training based on optical flow, by decoupling the representation of the motion process and designing corresponding strategies, guiding the diffusion model to complete motion editing; according to the processing of the multi-view optical flow estimation stage and the multi-view motion diffusion stage, completing the interactive three-dimensional scene generation and motion editing. The embodiments of the present invention can improve the robustness and generality of the zero-shot inference paradigm, and can effectively improve the motion accuracy and multi-view consistency in the motion editing process. Description of the Drawings
[0055] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention;
[0056] Figure 2 is a flowchart of the overall steps provided by an embodiment of the present invention;
[0057] Figure 3 is a structural diagram of MFES provided by an embodiment of the present invention;
[0058] Figure 4 is a structural diagram of MMDS provided by an embodiment of the present invention;
[0059] Figure 5 is a schematic diagram of the motion decoupling representation provided by an embodiment of the present invention;
[0060] Figure 6 is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0061] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention detailed in the appended claims.
[0062] It can be understood that the terms "first", "second", etc. used in the present invention can be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".
[0063] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present invention, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0065] The interactive three-dimensional scene generation and motion editing method, device, electronic device, and storage medium provided by the embodiments of the present invention relate to the field of computer technology. The interactive three-dimensional scene generation and motion editing method provided by the embodiments of the present invention can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side may be configured as an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server may also be a node server in a blockchain network; the software may be an application that implements the interactive three-dimensional scene generation and motion editing method, etc., but is not limited to the above forms.
[0066] The present invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0067] As Figure 1 shown, it is a schematic diagram of an implementation environment provided by the embodiments of the present invention. Referring to Figure 1 , this implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be network-connected through wireless or wired means to complete data transmission and exchange.
[0068] The server 101 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0069] In addition, the server 101 can also be a node server in a blockchain network. Among them, the blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms.
[0070] The terminal 102 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc. Among them, the terminal 102 can also be an in-vehicle terminal of various device types exemplified above, but is not limited thereto. The terminal 102 and the server 101 can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present invention do not limit this here.
[0071] Exemplarily based on Figure 1 the implementation environment shown, the embodiments of the present invention provide an interactive three-dimensional scene generation and motion editing method. Taking the application of this interactive three-dimensional scene generation and motion editing method in the server 101 as an example for illustration, it can be understood that this method can also be applied to the terminal 102.
[0072] Referring to Figure 2 , Figure 2 is a flowchart of the interactive three-dimensional scene generation and motion editing method applied to the server provided by the embodiments of the present invention. The execution subject of this method can be any of the aforementioned computer devices (including servers or terminals). Referring to Figure 2 , this method may include the following steps:
[0073] The multi-view optical flow estimation stage, specifically: given a static scene and multi-view images under the static scene, use MaskClustering to segment the three-dimensional scene and export the query codes of each object; generate the optical flow of the selected single-view image through the graphical operation interface interaction; estimate the multi-view optical flow through the point cloud dynamics model PKM;
[0074] The multi-view motion diffusion stage, specifically: based on the diffusion inference paradigm without training of the optical flow, by decoupling the representation of the motion process and designing corresponding strategies, guide the diffusion model to complete motion editing;
[0075] Based on the processing of the multi-view optical flow estimation stage and the multi-view motion diffusion stage, interactive 3D scene generation and motion editing are completed.
[0076] In some embodiments, after obtaining the optical flow of the selected single-view image, the point cloud P after motion is estimated using the optical flow. m , and then the point cloud after motion and the original point cloud P o are projected to obtain the multi-view optical flow. The expression for this process is:
[0077]
[0078] where foc is the camera focal length; [R|T] is the camera rotation and translation matrix; and d is the depth value.
[0079] In some embodiments, the point cloud kinetic model PKM is used to calculate the sparse point cloud P after motion based on the single-view optical flow. sm , and the expression for this process is:
[0080]
[0081] where K is the camera intrinsic matrix; c x,y is the 2D coordinate position; pp is the optical center of the camera; f sx represents the x component of the optical flow; and f sy represents the y component of the optical flow.
[0082] In some embodiments, the method further includes: for different motion modes, constructing corresponding point motion mechanics models, specifically:
[0083] For the translational motion mode, the translational motion is represented in 3D space as an offset p between the original point cloud and the point cloud after motion off , and then the moving point cloud is calculated;
[0084] For the scaling motion mode, there is a scaling factor s between P m and P o : for the shrinking motion, the amplitude of the single-view optical flow is used to statistically calculate this scaling factor; for the enlarging motion, the ratio of the area of the collision region to the area of the original region is calculated; f For the rotational motion mode, based on the same rotation angle θ of all 3D points in the rotational motion, the point cloud after motion is represented using the spatial rotation matrix Rot and the centroid p of the 3D point cloud
[0085] c ; For the stretching motion mode, the stretching motion is represented in space as a stretching plane, and the original point cloud is stretched by the stretching plane to different degrees.
[0086]
[0087] In some embodiments, in the translational motion mode, the calculation formula for the moving point cloud is:
[0088] P m = P o + p off ,
[0089]
[0090] where N(P so ) represents the number of the original sparse point cloud P so .
[0091] In the scaling motion mode, the expression for statistically calculating the scaling factor using the magnitude of the single-view optical flow is:
[0092]
[0093] where N represents the number of non-zero optical flow values;
[0094] In the scaling motion mode, the expression for calculating the area ratio of the collision region to the original region is:
[0095]
[0096] o r = L(c x , c y , f s ) ∪ r(f s ),
[0097] where L represents the linear sampling function, r(f s ) and o r represent the areas of the original and collision regions;
[0098] In the rotational motion mode, the calculation formulas for the spatial rotation matrix Rot and the centroid p c of the three-dimensional point cloud are:
[0099] P m = Rot × (P o - p c ) + p c ,
[0100]
[0101] In the stretching motion mode, the expression for stretching the original point cloud to different degrees is:
[0102] P m = P o + t f × Max(Psm -P so ),
[0103]
[0104] where dis represents the distance from the point cloud Po to the plane "AX + BY + D"; t f represents the ratio of the distance of a certain point to the maximum point distance.
[0105] In some embodiments, the multi-view motion diffusion stage includes the following steps:
[0106] According to the multi-view optical flow, the motion is represented as a combination of a static background, a moving object, and an occluded area. After decoupling these three parts, corresponding strategies are adopted to achieve motion editing; specifically:
[0107] Optical flow guidance strategy FGS process: For the static background, use DDIM inversion to replace the output of each step of the diffusion model;
[0108] Latent space fusion LSF process: Use optical flow to perform fusion in the latent space and supervise the texture details of the moving object;
[0109] Background grid constraint BGC process: Perform background grid constraint on the collision area during the diffusion process to maintain the consistency of multi-view images.
[0110] In some embodiments, the optical flow guidance strategy FGS process is specifically as follows: Use an optical flow estimator RAFT to predict the optical flow f between the generated image and the input image p and calculate the optical flow loss L i by comparing it with the input optical flow f flow ; Use warp deformation to calculate the loss L color between the generated image and the input image; Then take the derivative of these losses to change the predicted noise of the diffusion model:
[0111]
[0112] The latent space fusion LSF process is specifically as follows: Perform inverse warp deformation on the input image and the optical flow, then use VAE to encode the deformed image and add noise to it; Finally, fuse the noisy image with its mask in the latent space to obtain the noise output at the current moment;
[0113] The background grid constraint BGC process is specifically as follows: First, perform grid transformation on the output of the diffusion model, then denoise the tensor, and finally use the inverse grid transformation to obtain the final tensor as the final output.
[0114] Next, taking a specific application scenario as an example, the specific implementation process of the embodiments of the present invention will be described in detail:
[0115] In view of the problems existing in the prior art, an embodiment of the present invention proposes a zero-shot generative interactive motion editing method MotionDiff without training. According to the monocular optical flow obtained from user interaction, the multi-view optical flow is estimated, and the diffusion model is guided to complete the multi-view motion editing task. To improve the robustness and generality of this zero-shot inference paradigm, the embodiment of the present invention decouples and represents the motion process as a combination of a static background, a moving object, and a collision area, making it immune to restrictions such as scenes, object categories, and text features, and allowing the diffusion model to focus on the understanding and processing of spatial information. In this embodiment, it is found in experiments that each decoupling module is necessary, and by combining these decouplings organically, the diffusion model can be guided to obtain motion editing results with texture fidelity and multi-view consistency.
[0116] The implementation process of the embodiment of the present invention in a specific scenario is as Figure 3 and Figure 4 shown, and specifically includes the following steps:
[0117] 1. Point cloud dynamics model PKM:
[0118] After obtaining the monocular optical flow, the embodiment of the present invention aims to estimate the point cloud P m after motion with it, and then project the point cloud after motion and the original point cloud P o to obtain the multi-view optical flow f m , as follows:
[0119]
[0120] Among them, foc is the camera focal length, [R|T] is the camera rotation and translation matrix, and d is the depth value.
[0121] For this reason, the embodiment of the present invention designs PKM. First, calculate the sparse point cloud P sm after motion according to the monocular optical flow:
[0122]
[0123] Among them, K is the camera intrinsic matrix, c x,y is the 2D coordinate position, and pp is the optical center of the camera.
[0124] For different motion modes, we design corresponding kinematic models:
[0125] Translation. The embodiment of the present invention represents the translation motion in three-dimensional space as an offset p off between the original point cloud and the point cloud after motion, and then the moving point cloud can be calculated:
[0126] P m = Po +p off ,
[0127] where N(P so ) represents the number of the original sparse point cloud P si .
[0128] Scaling. The embodiments of the present invention believe that there is a scaling factor s m between P o and P f . For reduction, the embodiments of the present invention use the magnitude of the single-view optical flow to statistically calculate this scaling factor:
[0129]
[0130] where N represents the number of non-zero optical flow values.
[0131] For enlargement, the embodiments of the present invention calculate the area ratio of the collision area to the original area:
[0132] o r =L(c x , c y , f s ) ∪ r(f s ),
[0133] where L represents the linear sampling function, r(f s ) and o r represent the areas of the original and collision areas.
[0134] Rotation. All three-dimensional points of the rotational motion have the same rotation angle θ. Therefore, the point cloud after motion can be represented by the spatial rotation matrix Rot and the centroid p c of the three-dimensional point cloud as follows:
[0135] P m =Rot×(P o -p c ) + p c ,
[0136] Stretching. The stretching motion is characterized by a stretching plane in space, and the original point cloud is stretched by this plane to different degrees, which can be represented as follows:
[0137] P m =P o +t f ×Max(P sm -P so ),
[0138] Through the established point cloud dynamics model above, we can obtain the point cloud after motion and multi-view optical flow.
[0139] 2. Motion decoupling representation and corresponding strategies
[0140] As Figure 5 shown, in the embodiment of the present invention, according to the above optical flow, the motion is represented as a combination of a static background, a moving object, and an occlusion area. The embodiment of the present invention decouples these three parts and designs corresponding strategies to achieve motion editing:
[0141] Optical flow guidance strategy FGS. For the static background, in the embodiment of the present invention, DDIM inversion is used to replace the output of each step of the diffusion model. Then, as Figure 4 (a) of p shown, in the embodiment of the present invention, an optical flow estimator RAFT is used to predict the optical flow f i between the generated image and the input image, and calculate the optical flow loss L flow by comparing it with the input optical flow f color . At the same time, using warp deformation, calculate the loss L
[0142]
[0143] between the generated image and the input image. Then, in the embodiment of the present invention, the derivatives of these losses can be calculated to change the predicted noise of the diffusion model: Figure 4 Latent space fusion LSF. To better supervise the texture details of the moving object, the embodiment of the present invention proposes to use optical flow for fusion in the latent space. As
[0144] (b) of
[0145] shown, first, inverse warp deformation is performed on the input image and the optical flow, then the deformed image is encoded by VAE and noise is added to it. Finally, the noisy image and its mask are fused in the latent space to obtain the noise output at the current moment.
[0146] The present invention constructs a zero-shot generative interactive motion editing method MotionDiff without training. For static scene input, the user can interactively obtain motion optical flow, and then obtain motion editing results, including translation, scaling, stretching, rotation, etc. Among them, the point cloud dynamics model PKM can estimate the three-dimensional point cloud after motion and multi-view optical flow based on single-view optical flow. The motion decoupling representation and the corresponding diffusion model guidance strategy proposed by the present invention obtain multi-view motion editing results in a convenient inference paradigm. After comparison with other generative motion editing methods, the method of the present invention is 5.5 and 1.31 higher than the current state-of-the-art methods in terms of motion accuracy and multi-view consistency, reaching 6.8 and 34.92 (the top1 two indicators are about 12.3 and 33.61). And the method of the present invention has obvious superiority in visual effects. This is sufficient to prove the effectiveness of the method of the present invention.
[0147] In addition, the method provided by the embodiments of the present invention can be applied to the following scenarios: 1. Point cloud processing and multi-view editing in user interactive scenarios; 2. A resource-friendly zero-shot diffusion model inference paradigm; 3. Making a high-quality input-motion paired data set. The embodiments of the present invention can improve the robustness and generality of the zero-shot inference paradigm, and can effectively improve the motion accuracy and multi-view consistency in the motion editing process.
[0148] Another aspect of the embodiments of the present invention also provides an interactive three-dimensional scene generation and motion editing device, including:
[0149] A first module for performing a multi-view optical flow estimation stage, specifically: given a static scene and multi-view images under the static scene, using MaskClustering to segment the three-dimensional scene and export query codes of each object; interactively generating the optical flow of the selected single-view image through a graphical operation interface; estimating multi-view optical flow through the point cloud dynamics model PKM;
[0150] A second module for performing a multi-view motion diffusion stage, specifically: based on a diffusion inference paradigm without training of optical flow, guiding the diffusion model to complete motion editing by decoupling the representation of the motion process and designing corresponding strategies;
[0151] A third module for completing interactive three-dimensional scene generation and motion editing according to the processing of the multi-view optical flow estimation stage and the multi-view motion diffusion stage.
[0152] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0153] An embodiment of the present invention further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned interactive three-dimensional scene generation and motion editing method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0154] It can be understood that the content in the above method embodiments is applicable to this device embodiment. The functions specifically implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0155] Please refer to Figure 6 , Figure 6 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0156] A processor 601, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0157] A memory 602, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 602 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 602, and the processor 601 is used to call and execute the interactive three-dimensional scene generation and motion editing method of the embodiments of the present invention;
[0158] An input / output interface 603, which is used to implement information input and output;
[0159] A communication interface 604, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);
[0160] A bus 605, which transmits information between the various components of the device (such as the processor 601, the memory 602, the input / output interface 603, and the communication interface 604);
[0161] Among them, the processor 601, the memory 602, the input / output interface 603, and the communication interface 604 are communicatively connected to each other inside the device through the bus 605.
[0162] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned interactive three-dimensional scene generation and motion editing method is implemented.
[0163] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0164] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0165] It should be noted that in various specific embodiments of the present invention, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present invention need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present invention will be obtained.
[0166] The embodiments described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0167] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than those shown, or combine some steps, or different steps.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solutions of this embodiment.
[0169] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0170] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0171] It should be understood that in the present invention, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or similar expressions below refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0172] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0173] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0174] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0175] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. And the aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0176] The preferred embodiments of the embodiments of the present invention have been described above with reference to the drawings, but this does not limit the scope of the rights of the embodiments of the present invention. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.
Claims
1. An interactive three-dimensional scene generation and motion editing method, characterized in that It includes the following steps: The multi-view optical flow estimation stage, specifically: Given a static scene and multi-view images under the static scene, use MaskClustering to segment the 3D scene and export the query codes of each object; Generate the optical flow of the selected single-view image through interaction with the graphical operation interface; Estimate the multi-view optical flow through the point cloud kinetics model PKM; The multi-view motion diffusion stage, specifically: Based on the diffusion inference paradigm without training based on optical flow, by decoupling the representation of the motion process and designing corresponding strategies, guide the diffusion model to complete motion editing; According to the processing of the multi-view optical flow estimation stage and the multi-view motion diffusion stage, complete the interactive 3D scene generation and motion editing.
2. The interactive three-dimensional scene generation and motion editing method according to claim 1, wherein After obtaining the optical flow of the selected single-view image, use the optical flow to estimate the point cloud P after motion m , and then project the point cloud after motion and the original point cloud P o to obtain the multi-view optical flow. The expression for this process is: Wherein, foc is the camera focal length; [R|T] is the camera rotation and translation matrix; d is the depth value.
3. An interactive three-dimensional scene generation and motion editing method according to claim 2, characterized in that, The point cloud dynamics model PKM is used to calculate the sparse point cloud P after motion according to the single-view optical flow sm , and the expression of this process is: Among them, K is the camera internal parameter matrix; c x,y is the 2D coordinate position; pp is the optical center of the camera; f sx represents the x component of the optical flow; f sy represents the y component of the optical flow.
4. An interactive three-dimensional scene generation and motion editing method according to claim 3, characterized in that The method further includes: For different motion modes, construct corresponding point motion mechanics models, specifically: For the translational motion pattern, the translational motion is represented in three-dimensional space as an offset p between the original point cloud and the point cloud after motion, off and then the moving point cloud is calculated. For the scaled motion pattern, at P m and P o there is a scaling factor s f : For the shrinking motion, the amplitude of the monocular optical flow is used to statistically calculate the scaling factor; for the enlarging motion, the ratio of the area of the collision region to the area of the original region is calculated; For the rotational motion pattern, based on the same rotation angle θ of all three-dimensional points in the rotational motion, the spatial rotation matrix Rot and the centroid p of the three-dimensional point cloud are used c to represent the point cloud after motion; For the stretching motion mode, represent the stretching motion as a stretching plane in space, and the original point cloud is stretched by the stretching plane to different degrees.
5. An interactive 3D scene generation and motion editing method according to claim 4, characterized in that In the translational motion mode, the calculation formula for the moving point cloud is: P m = P o + p off , Among them, N(P so ) represents the number of the original sparse point cloud P so ; In the scaling motion mode, the expression for statistically calculating the scaling factor using the magnitude of the single-view optical flow is: Wherein, N represents the number of non-zero optical flow values; In the scaling motion mode, the expression for calculating the area ratio of the collision region to the original region is: o r = L(c x , c y , f s ) ∪ r(f s ), where L represents the linear sampling function, r(f s ) and o r represent the areas of the original and collision regions; In the described rotational motion mode, the spatial rotation matrix Rot and the centroid p of the three-dimensional point cloud c are calculated as follows: P m = Rot×(P o - p c ) + p c , In the stretching motion mode, the expression for stretching the original point cloud to different degrees is: P m = P o + t f × Max(P sm - P so ), where dis represents the distance from the point cloud Po to the plane "AX + BY + D"; t f represents the ratio of the distance of a certain point to the maximum point distance.
6. An interactive three-dimensional scene generation and motion editing method according to claim 1, characterized in that, The multi-view motion diffusion stage includes the following steps: According to the multi-view optical flow, represent the motion as a combination of a static background, moving objects, and occlusion regions, decouple these three parts, and adopt corresponding strategies to achieve motion editing; Specifically: The optical flow guidance strategy FGS process: For the static background, use DDIM inversion to replace the output of each step of the diffusion model; The latent space fusion LSF process: Use the optical flow to perform fusion in the latent space to supervise the texture details of the moving objects; The background grid constraint BGC process: Perform background grid constraints on the collision regions during the diffusion process to maintain the consistency of the multi-view images.
7. An interactive 3D scene generation and motion editing method according to claim 6, characterized in that The specific process of the optical flow guidance strategy FGS is as follows: Use an optical flow estimator RAFT to predict the optical flow f between the generated image and the input image p , and calculate the optical flow loss L with it and the input optical flow f i ; Use warp transformation to calculate the loss L between the generated image and the input image flow ; Then take the derivative of these losses to change the predicted noise of the diffusion model: color The latent space fusion LSF process, specifically: Perform inverse warp deformation on the input image and the optical flow, then use VAE to encode the deformed image and add noise to it; Finally, fuse the noisy image with its mask in the latent space to obtain the noise output at the current moment; The background grid constraint BGC process, specifically: First perform grid transformation on the output of the diffusion model, then denoise the tensor, and finally use the inverse grid transformation to obtain the final tensor as the final output.
8. An interactive three-dimensional scene generation and motion editing device, characterized in that, It includes: The first module is used to execute the multi-view optical flow estimation stage, specifically: given a static scene and multi-view images under the static scene, use MaskClustering to segment the 3D scene and export the query codes of each object; generate the optical flow of the selected single-view image through interaction with the graphical operation interface; Estimate the multi-view optical flow through the point cloud dynamics model PKM; The second module is used to execute the multi-view motion diffusion stage, specifically: based on the diffusion inference paradigm without training of optical flow, through decoupling the representation of the motion process and designing corresponding strategies, guide the diffusion model to complete motion editing; The third module is used to complete the interactive 3D scene generation and motion editing according to the processing of the multi-view optical flow estimation stage and the multi-view motion diffusion stage.
9. An electronic device, characterized in that, It includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program, and the program is executed by the processor to implement the method according to any one of claims 1 to 7.