3D Model Editing Method, Device, Storage Medium, and Program Product
By obtaining multiple source images from different perspectives from the source three-dimensional model and using the multi-view image editing model that has been trained to converge, combined with the target text prompt information, the problem of the inability to handle the state changes of the target object in the existing technology is solved, non-rigid editing and multi-view consistency are achieved, and editing speed and model accuracy are improved.
Patent Information
- Application Number
- CN202510192935.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art cannot effectively handle the state changes of the target object, and can only realize partial feature changes or replacement of objects of the same type in the source three-dimensional model in the same state.
Obtain multiple source images from different perspectives from the source three-dimensional model, use the multi-view image editing model that has been trained to convergence, combine the target text prompt information, output the target multi-view image to achieve non-rigid editing, and reconstruct the target three-dimensional model through the target multi-view image.
The non-rigid editing of the target object from the source state to the target state is achieved, which improves editing speed and multi-view consistency, and the obtained three-dimensional model is more accurate.
Smart Images

Figure CN119693593B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a three-dimensional model editing method, device, storage medium, and program product. Background Art
[0002] Currently, creating three-dimensional models plays a key role in many applications and industries. Three-dimensional models can be applied in film / game production to enhance visualization and interactivity, etc. Therefore, creating high-quality three-dimensional models is particularly important.
[0003] In related technologies, a pre-trained image editing model is used to edit the text description of the source image corresponding to the source three-dimensional model to obtain an edited image, and then the edited image is lifted to a three-dimensional model by reconstruction to obtain a target three-dimensional model, thereby realizing three-dimensional editing. However, the above-mentioned related technologies can handle partial feature changes of an object in the same state or replacement of the same type of object in the source three-dimensional model. However, the state change of the target object cannot be processed. Summary of the Invention
[0004] The present application provides a three-dimensional model editing method, device, storage medium, and program product to solve the problem of partial feature changes of an object in the same state or replacement of the same type of object in related technologies, and achieve the effect of state change of the target object.
[0005] In a first aspect, the present application provides a three-dimensional model editing method, including:
[0006] Obtain multiple source images from different perspectives in the source three-dimensional model, where the source images are images corresponding to the three-dimensional model from different perspectives; each source image includes a target object; each target object is in a source state;
[0007] Determine the target text prompt information corresponding to each target multi-perspective image to be generated; the target text prompt information includes the target state information of the target object;
[0008] Input the multiple source images and the target text prompt information into a multi-perspective image editing model that has been trained to convergence, and output the target multi-perspective images of the multiple source images from the corresponding perspectives to achieve non-rigid editing; the target multi-perspective image is an image in which the target object executes a target action from the source state to form a target state from the source state in the perspective of the corresponding source image;
[0009] Reconstruct the target three-dimensional model based on each target multi-perspective image.
[0010] In a second aspect, the present application further provides an editing device, including: a memory and a processor;
[0011] The memory stores a computer program;
[0012] When the processor executes the computer program, the steps of any of the above three-dimensional model editing methods are implemented.
[0013] In a third aspect, the present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above three-dimensional model editing methods are implemented.
[0014] In a fourth aspect, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above three-dimensional model editing methods are implemented.
[0015] The present application provides a three-dimensional model editing method, device, storage medium, and program product. First, multiple source images from different perspectives are obtained from the source three-dimensional model, and the target object in the source images is in the source state. Then, the target text prompt information is determined. Since the target text prompt information includes the target state information of the target object, the target state information refers to the state information of the target object in the target multi-perspective images obtained after editing based on the source images. Further, the multiple source images and the target text prompt information are input into the multi-perspective image editing model that has been trained to convergence, so that multiple target multi-perspective images corresponding to the multiple source images can be output at one time. In the present application, in the perspective of the corresponding source image, the target object in the target multi-perspective image performs the target action and then forms an image of the target state from the source state, thereby completing non-rigid editing. It can be seen that in the present application, based on the target text prompt information, the editing of the state of the target object is realized, and the target object can be changed from one state to another state, that is, from the source state to the target state. In addition, in the present application, the multi-perspective image editing model that has been trained to convergence can input multiple source images for simultaneous processing and output multiple target multi-perspective images corresponding to the multiple source images at the same time, so the editing speed is improved. In the present application, the source images from multiple perspectives are used to implement editing, so multiple perspectives are combined, making the obtained three-dimensional model improve the multi-perspective consistency and be more accurate. Description of the Drawings
[0016] In order to more clearly illustrate the embodiments of the present application, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a schematic diagram of an application scenario of a three-dimensional model editing method provided by the present application;
[0018] Figure 2Schematic diagram of a 3D model editing method provided for Embodiment 1;
[0019] Figure 3 Schematic diagram of a 3D model editing method provided for Embodiment 2;
[0020] Figure 4 Schematic diagram of a 3D model editing method provided for Embodiment 3;
[0021] Figure 5 Schematic diagram of a 3D model editing method provided for Embodiment 4;
[0022] Figure 6 Schematic diagram of a 3D model editing method provided for Embodiment 6;
[0023] Figure 7 Schematic diagram of a 3D model editing method provided for Embodiment 8;
[0024] Figure 8 Schematic diagram of a 3D model editing method provided for Embodiment 10;
[0025] Figure 9 Schematic diagram of a 3D model editing method provided for Embodiment 11;
[0026] Figure 10 Overall process schematic diagram of a 3D model editing method provided for Embodiment 13;
[0027] Figure 11 Schematic diagram of the structure of a 3D model editing device provided for Embodiment 14;
[0028] Figure 12 Schematic diagram of the structure of an editing device provided for Embodiment 14. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0030] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0031] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail with reference to the accompanying drawings and specific embodiments.
[0032] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the 3D model editing method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0033] Low-Rank Adaptation (LoRA for short): By decomposing the weight matrix into the product of low-rank matrices, the number of parameters is reduced, thereby achieving the purpose of reducing hardware resources and accelerating the fine-tuning and training processes.
[0034] In the related art, a pre-trained image editing model is used to edit the text description of the source image corresponding to the source 3D model to obtain an edited image, and then the edited image is lifted to the 3D mode by reconstruction to obtain a target 3D model, thereby realizing 3D editing. However, the above-mentioned related art can handle partial feature changes or replacement of the same type of objects in the source 3D model in the same state. However, it cannot handle the state changes of the target object.
[0035] To address the deficiencies of related technologies, the inventors of this solution conducted creative research and designed a new solution. This solution provides a 3D model editing method. To solve the problem in related technologies where only partial feature changes of a target object in the same state or replacement of the same type of object can be performed on the source image, in this application, multiple source images from different perspectives are obtained from the source 3D model, and the source images are in the source state. Then, a trained-to-converge multi-view image editing model is used to output target multi-view images corresponding to the multiple source images. In the target multi-view images, it is reflected that the target object forms the target state from the source state after performing the target action. Thus, it can be seen that the target object changes from the source state to the target state, thereby realizing the change of the state and achieving non-rigid editing of the source image. In addition, the trained-to-converge multi-view image editing model in this application is convergent, so the target multi-view images output based on this are more accurate. In addition, the trained-to-converge multi-view image editing model inputs multiple source images simultaneously and outputs multiple target multi-view images of the source object after editing simultaneously, thereby improving the editing speed. In this application, non-rigid editing is performed on the source images from different perspectives, improving multi-view consistency and making the target 3D model reconstructed based on multiple target multi-view images more accurate.
[0036] Figure 1 Schematic diagram of an application scenario of a 3D model editing method provided by this application. As Figure 1 shown, it includes an editing device 101.
[0037] Among them, the editing device 101 can be any electronic device, such as a computer or a server, etc., and no limitation is made here.
[0038] In this scenario, the editing device 101 obtains multiple source images from different perspectives from the source 3D model. Further, the editing device 101 determines the target text prompt information.
[0039] Further, the editing device 101 inputs the source images and the target text prompt information into the trained-to-converge multi-view image editing model, and then outputs the target multi-view images of the multiple source images in the corresponding perspectives.
[0040] Further, the editing device 101 reconstructs the target 3D model based on each target multi-view image.
[0041] Embodiments of this application provide a 3D model editing method. Combining the execution process of the 3D model editing method, the method is described in detail.
[0042] Embodiment 1 The execution subject of Embodiment 1 to Embodiment 13 of this application is a 3D model editing device, and this 3D model editing device is located in the editing device.
[0043] Figure 2Schematic diagram of a 3D model editing method provided for Embodiment 1. As Figure 2 shown, it specifically includes:
[0044] S201, obtaining multiple source images from different perspectives of the source 3D model, where the source images are the images corresponding to the 3D model from different perspectives; each source image includes a target object; each target object is in a source state.
[0045] Among them, the source 3D model refers to the 3D model to be edited for the target object therein.
[0046] Among them, the target object refers to the object included in the source 3D model or source image, and can be any thing. For example, the target object is a cat or a dog. It should be noted that the source 3D model is a three-dimensional spatial structure.
[0047] Among them, the source state refers to the state in which the target object is before editing.
[0048] Among them, the source image refers to the image corresponding to the source 3D model from different perspectives. It should be noted that the source image is two-dimensional, that is, planar.
[0049] S202, determining the target text prompt information corresponding to each target multi-perspective image to be generated; the target text prompt information includes the target state information of the target object.
[0050] Among them, the target text prompt information refers to the information for text prompt processing of the source image.
[0051] Among them, the target state refers to the state in which the target object is after editing.
[0052] Exemplarily, if the target text prompt information includes a sitting yellow kitten, then the target text prompt information is information such as "sitting", "yellow", and "small", and the target state information is "sitting".
[0053] It should be noted that the target text prompt information also includes perspective information. For example, the target text prompt information can also include the perspective corresponding to the source image. Exemplarily, the target text prompt information can be a sitting yellow kitten in the perspective corresponding to the source image.
[0054] S203, inputting the multiple source images and the target text prompt information into a multi-perspective image editing model that has been trained to convergence, and outputting the target multi-perspective images of the multiple source images in the corresponding perspectives to achieve non-rigid editing; the target multi-perspective images are images in which the target object performs a target action and forms a target state from the source state in the perspective corresponding to the corresponding source image.
[0055] It should be noted that in this embodiment, multiple source images are input into a multi-view image editing model that has been trained to convergence, and the target multi-view images of the multiple source images in the corresponding views are output simultaneously.
[0056] Among them, the multi-view image editing model that has been trained to convergence refers to a model that performs multi-view editing on multiple source images based on target text prompt information, so that the target object in the multiple source images changes from the source state to the target state under multi-view control.
[0057] In one way, the multi-view image editing model that has been trained to convergence can receive multiple source images simultaneously, but processes the multiple source images independently and outputs the target multi-view images corresponding to each source image simultaneously.
[0058] In another way, the multi-view image editing model that has been trained to convergence can combine the view information between multiple source images, fully combine the multi-view consistency, and achieve non-rigid editing with multi-view consistency, so that the target object changes from the source state to the target state. It can be seen that each target multi-view image is obtained by combining the views between each source image.
[0059] It should be noted that assuming there are four source images, the above four source images are input into the multi-view image editing model that has been trained to convergence, and the multi-view image editing model that has been trained to convergence is used to output the target multi-view images of the four source images in the corresponding views simultaneously, so as to obtain four target multi-view images.
[0060] Among them, the view corresponding to the target multi-view image is the view corresponding to the source image, and the target multi-view image combines the corresponding views of each source image, so as to achieve non-rigid editing with multi-view consistency.
[0061] Among them, the target object in the target multi-view image changes from the source state to the target state.
[0062] Among them, non-rigid editing refers to the editing in which the state of the target object changes after performing the target action.
[0063] Exemplarily, assume that one source image is a front source image and a left-side source image, the target object is a cat, the source state is "sitting", the target state is "standing", and the target object also has other original features. For example, the color is yellow. Regarding the original features, they will not be elaborated here. In this example, the target text prompt information includes the target state information (that is, it means changing the target object from the source state to the target state, so the state of the target object changes), and also includes other features. In one way, the other features can be the same as the original features in the source image, indicating that the original features remain unchanged at this time.
[0064] Further, input the above-mentioned source image of the front side, the source image of the left side, and the target text prompt information into the multi-view image editing model that has been trained to convergence. The multi-view image editing model that has been trained to convergence combines the perspective information of the two source images and outputs the target multi-view image of the front side corresponding to the source image of the front side, and the target multi-view image of the left side corresponding to the source image of the left side. Among them, the target multi-view image of the front side refers to an image in which the target object remains unchanged in other original features and changes from the source state to the target state under the front view. Similarly, the target multi-view image of the left side refers to an image in which the target object remains unchanged in other original features and changes from the source state to the target state under the left side view.
[0065] It should be noted that in this application, in addition to the state change of the target object, the color and size of the target object in the source image can also be changed.
[0066] It should be noted that in this application, the target state can be a local state change of the target object from the source state, or a global state change of the target object from the source state. Thus, it can be seen that global or local non-rigid editing of the target object can be achieved in this application; the editing device in this application performs global or local non-rigid editing on the source image, that is, it can enable non-rigid editing of the global or local state change of the target object in the source image. It should be noted that the local state change means that some parts of the target object change in state. For example, the target object in the source image is a cat sitting with both hands vertical, and the target object in the target multi-view image of the source image is a cat sitting with its right hand raised and its left hand vertical. The global state change means that all parts of the target object change in state. For example, the target object is a cat sitting with both hands vertical, and the target object in the target multi-view image is a standing cat.
[0067] S204. Reconstruct the target three-dimensional model based on each target multi-view image.
[0068] Specifically, an initial three-dimensional model is obtained by using three-dimensional Gaussian representation according to each target multi-view image, and then the reconstruction loss function is applied to optimize the initial three-dimensional model to obtain the non-rigidly edited target three-dimensional model.
[0069] This embodiment provides a 3D model editing method. In this embodiment, first, multiple source images from different perspectives are obtained from the source 3D model, and the target object in the source images is in the source state. Then, the target text prompt information is determined. Since the target text prompt information includes the target state information of the target object, this target state information refers to the state information of the target object in the target multi-perspective images obtained by editing based on the source images. Further, the multiple source images and the target text prompt information are input into the multi-perspective image editing model that has been trained to convergence, so that multiple target multi-perspective images corresponding to the multiple source images can be output at one time. In this embodiment, in the perspective of the corresponding source image of the target multi-perspective image, the target object performs the target action and forms an image of the target state from the source state to achieve non-rigid editing. It can be seen that in this embodiment, based on the target text prompt information, the editing of the state of the target object is realized, and the target object can be changed from one state to another state, that is, from the source state to the target state. In addition, in this embodiment, the multi-perspective image editing model that has been trained to convergence can input multiple source images for simultaneous processing and output multiple target multi-perspective images corresponding to the multiple source images at the same time, so the editing speed is improved. In this embodiment, the source images from multiple perspectives are used to achieve editing, so multiple perspectives are combined, making the obtained 3D model more accurate.
[0070] Embodiment 2
[0071] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to input multiple source images and the target text prompt information into the multi-perspective image editing model that has been trained to convergence, and output the target multi-perspective images of the multiple source images in the corresponding perspectives to achieve non-rigid editing.
[0072] Figure 3 Schematic diagram of a 3D model editing method provided for Embodiment 2. As Figure 3 shown, it specifically includes:
[0073] S301, obtain the camera pose information of multiple source images.
[0074] Among them, the camera pose information refers to the position and orientation of the camera that captures the source images in space. It should be noted that the camera pose information can reflect the perspective of the source images.
[0075] S302, perform feature extraction on multiple source images at a preset time step to obtain the source image feature set at the preset time step.
[0076] Among them, the preset time step refers to the time step preset for the multi-perspective image editing model that has been trained to convergence.
[0077] Among them, the source image feature set refers to the feature set obtained by performing preliminary feature extraction on the source image. Among them, the source image feature set includes the original features of multiple source images, that is, the source image features. It can be seen that the source image feature set includes multiple source image features.
[0078] It should be noted that in this step, the source image feature set can be represented by the following formula (1).
[0079] (1)
[0080] Among them, , N represents the number of images in the source image feature set, is the size of the source image, C represents the source image feature dimension, t is the preset time step, is the source image feature set, is the source image feature of the i-th source image.
[0081] S303. Obtain a three-dimensional cost frustum at a preset time step based on each camera pose information and the source image feature set.
[0082] Among them, the three-dimensional cost frustum is used to represent the range of the target object visible in a three-dimensional space and is a perspective model.
[0083] Among them, the three-dimensional cost frustum can represent the frustum after combining the source images from different perspectives with multi-view consistency.
[0084] Specifically, the source image features corresponding to each source image are transformed or distorted according to the camera pose information, and finally fused to obtain the three-dimensional cost frustum corresponding to the preset time step, that is, the 3D cost volume.
[0085] S304. Use a multi-view image editing model trained to convergence to output target multi-view images corresponding to the perspectives of multiple source images based on the three-dimensional cost frustum and the target text prompt information to achieve non-rigid editing.
[0086] This embodiment provides a three-dimensional model editing method. In this embodiment, first, camera pose information is obtained, and then feature extraction is performed on multiple source images to obtain a source image feature set corresponding to a preset time step. Further, a three-dimensional cost frustum is obtained. Since the three-dimensional cost frustum can represent the multi-view consistency of multiple source images, the target multi-view images output based on the three-dimensional cost frustum are more accurate.
[0087] Embodiment Three
[0088] This embodiment is a further refinement of any of the above embodiments and is an optional way to obtain a three-dimensional cost frustum at a preset time step based on each camera pose information and the source image feature set.
[0089] Figure 4 Schematic diagram of a 3D model editing method provided for Embodiment 3. As Figure 4 shown, it includes:
[0090] S401, determining each camera pose transformation matrix based on each camera pose information and preset conversion perspective information; the camera pose transformation matrix is a transformation matrix for converting from the perspective corresponding to the camera pose information to the preset conversion perspective information on the source depth; the camera pose information includes the perspective information corresponding to the source image.
[0091] Among them, the preset conversion perspective information refers to the information for converting from the corresponding perspective to the reference perspective set in advance. Exemplarily, assuming there are four source images, namely the source image of the front, the source image of the left side, the source image of the right side, and the source image of the back, among which, the preset conversion perspective information can be the first perspective in the perspective image set. Among them, the perspective image set includes multiple source images. Among them, the first perspective can be the front.
[0092] It should be noted that assuming the camera pose information of the source image of the left side reflects its perspective as the left side, then, determining the camera pose transformation matrix corresponding to the source image of the left side is to determine the transformation matrix for converting from the left side perspective to the reference perspective (the first perspective) on the source depth.
[0093] It should be noted that the camera pose information includes the perspective information corresponding to the source image, so that the perspective corresponding to the source image can be determined from the perspective information in the camera pose information.
[0094] S402, inputting the camera pose transformation matrix and the source image feature set into the preset conversion perspective algorithm to output the conversion feature set of each source image from the corresponding perspective to the preset conversion perspective at the preset time step.
[0095] Among them, the preset conversion perspective algorithm refers to the algorithm for converting the perspective of the source image set in advance.
[0096] Among them, the conversion feature set refers to the set of each source image feature converted from the corresponding perspective to the preset conversion perspective.
[0097] Exemplarily, the preset conversion perspective algorithm is as shown in (2):
[0098] (2)
[0099] Among them, is the warping transformation graph at the source depth at the preset time step t, represents a pixel position in the reference perspective, is at the source depth The camera pose transformation matrix for converting from the corresponding perspective to the reference view at this point.
[0100] Among them, represents the conversion feature, and each conversion feature is combined into a conversion feature set. Among them, D represents the source depth range.
[0101] S403. Calculate the variance between each conversion feature set to obtain the three-dimensional cost frustum corresponding to the preset time step
[0102] Exemplarily, the variance calculation formula between each conversion feature set in this step is as shown in (3):
[0103] (3)
[0104] Among them, Var represents variance calculation. Among them, .
[0105] This embodiment provides a three-dimensional model editing method. In this embodiment, first, the camera pose transformation matrix is determined, and then it is input into the preset conversion perspective algorithm to calculate the conversion feature set. It can be seen that different perspectives are combined in the conversion feature set, thus fusing multi-perspective features and further improving multi-perspective consistency; in addition, the three-dimensional cost frustum includes the source depth range, the source image size, and the source image feature dimension, making the three-dimensional cost frustum fully complete and more accurate.
[0106] Embodiment 4
[0107] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to implement non-rigid editing by using a multi-perspective image editing model trained to convergence to output target multi-perspective images corresponding to multiple source images based on the three-dimensional cost frustum and the target text prompt information.
[0108] Figure 5 Schematic diagram of a three-dimensional model editing method provided for Embodiment 4. As Figure 5 shown, it includes:
[0109] S501. Calculate the depth fusion feature at the preset time step based on the three-dimensional cost frustum reshaping; the depth fusion feature is the fusion feature on the source depth.
[0110] Exemplarily, calculate the depth fusion feature based on the preset reshaping algorithm in formula (4).
[0111] (4)
[0112] Among them, represents the depth fusion feature, .
[0113] It should be noted that the depth fusion features fuse features with the same depth together.
[0114] S502, Calculate the query vector, key vector, and value vector using a linear transformation layer based on the depth fusion features.
[0115] Exemplarily, the calculation method is as shown in (5):
[0116] (5)
[0117] where are the query vector linear transformation layer, key vector linear transformation layer, and value vector linear transformation layer respectively. Among them, 。
[0118] S503, Input the query vector, key vector, and value vector into a preset 3D self-attention mechanism algorithm to output multi-view fusion features at a preset time step; the multi-view fusion features are features after fusing multiple views through the 3D self-attention mechanism.
[0119] Among them, the preset 3D self-attention mechanism algorithm is a preset 3D self-attention mechanism algorithm, as shown in (6):
[0120] (6)
[0121] where is the 3D self-attention mechanism, is the multi-view fusion feature, where T represents transpose.
[0122] It should be noted that in this embodiment, multi-view self-attention calculation for depth perception is performed on the 3D cost view frustum, that is, multi-view self-attention calculation is performed at positions with the same source depth. Specifically, its principle is: different pixels of the same image will perform self-attention mechanism with other images at the same depth, and a new multi-view image is obtained after fusing features of other views.
[0123] It should be noted that in this embodiment, first, depth fusion features are obtained, then the query vector, key vector, and value vector are calculated respectively using a linear transformation layer, and further, self-attention mechanism is used to fuse multi-view features to obtain multi-view fusion features, thereby improving multi-view consistency.
[0124] S504, Use a multi-view image editing model that has been trained to convergence to output target multi-view images corresponding to the perspectives of multiple source images based on the multi-view fusion features and target text prompt information to achieve non-rigid editing.
[0125] Furthermore, use the multi-view fusion features and target text prompt information to output target multi-view images corresponding to the perspectives of multiple source images.
[0126] This embodiment provides a three-dimensional model editing method. In this embodiment, multi-view self-attention calculation based on depth perception is performed on the three-dimensional cost view frustum, and then a multi-view fusion feature is obtained. In this embodiment, the multi-view fusion feature is fused from the query vector, key vector, and value vector, so as to achieve multi-view fusion and improve multi-view consistency.
[0127] Embodiment 5
[0128] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to output target multi-view images corresponding to the perspectives of multiple source images based on the multi-view fusion feature and the target text prompt information by using a multi-view image editing model that has been trained to convergence, so as to achieve non-rigid editing, including:
[0129] Step 1: According to the camera pose information corresponding to each source image, project the multi-view fusion feature onto the perspective corresponding to each source image to obtain a multi-view image feature set with depth perception at a preset time step; the multi-view image feature set is a feature set obtained by calculating the three-dimensional self-attention mechanism of multiple source images through depth perception; the multi-view image feature set includes multi-view image features corresponding to multiple source images.
[0130] Among them, the multi-view image feature set is calculated by the three-dimensional self-attention mechanism with depth perception. In this embodiment, for the perspective corresponding to the source image, the editing device projects the multi-view fusion feature onto each perspective, so as to obtain the multi-view image features of each perspective at a preset time step, and multiple multi-view image features form a multi-view image feature set.
[0131] Exemplarily, assume that there are a total of four source images from different perspectives, and the perspectives are the front, left side, right side, and back. The editing device projects the multi-view fusion feature onto the front to obtain the multi-view image feature of the front. Similarly, project the multi-view fusion feature onto the left side, right side, and back to obtain the multi-view image features of the corresponding perspectives respectively. It can be understood that in this example, four multi-view image features are obtained, and the above four multi-view image features form a multi-view image feature set.
[0132] Step 2: Use the multi-view image editing model that has been trained to convergence to output target multi-view images corresponding to the perspectives of multiple source images based on the multi-view image feature set and the target text prompt information, so as to achieve non-rigid editing.
[0133] This embodiment provides a three-dimensional model editing method. In this embodiment, multi-view fusion features are projected onto the perspectives corresponding to each source image, so that a multi-view image feature set for depth perception at a preset time can be obtained. Since the multi-view fusion features improve multi-view consistency, the multi-view image feature set also fuses multi-view information, making the target multi-view image output by the trained-to-converge multi-view image editing model more accurate.
[0134] Embodiment Six
[0135] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to implement non-rigid editing by using a trained-to-converge multi-view image editing model to output target multi-view images corresponding to the perspectives of multiple source images based on the multi-view image feature set and target text prompt information.
[0136] Figure 6 Schematic diagram of a three-dimensional model editing method provided for Embodiment Six. As Figure 6 shown, it specifically includes:
[0137] S601, obtain the corresponding perspective information from the camera pose information of each source image.
[0138] S602, use a preset multi-layer perceptron to encode each perspective information to obtain each perspective encoding.
[0139] Among them, the preset multi-layer perceptron can encode the perspective information.
[0140] S603, use a trained-to-converge multi-view image editing model to output target multi-view images corresponding to the perspectives of multiple source images based on each perspective encoding, the multi-view image feature set, and target text prompt information to achieve non-rigid editing.
[0141] Specifically, determine the denoising condition based on the perspective encoding and target text prompt information, and control the multi-view image feature set to be edited and denoised based on the target text prompt information and denoising condition, so as to obtain the target multi-view image.
[0142] This embodiment provides a three-dimensional model editing method. In this embodiment, perspective information is first obtained based on the camera pose information, and then encoded to obtain perspective encoding. Further, based on the perspective encoding, the multi-view image feature set, and the target text prompt information, the target multi-view image is output. Since the multi-view image feature set fuses multi-view features, the target multi-view image is more accurate and the multi-view consistency is improved.
[0143] Embodiment Seven
[0144] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to output target multi-view images corresponding to multiple source images from each perspective based on a multi-view image editing model that has been trained to convergence, using each perspective encoding, a multi-view image feature set, and target text prompt information, so as to achieve non-rigid editing. Specifically, it includes:
[0145] Step 1: Encode and optimize the target text prompt information to obtain optimized target text embedding data.
[0146] Specifically, a preset text editor can be used to encode the target text prompt information to obtain target text embedding data, and then the target text embedding data is optimized. During the optimization process, semantic compression is performed. Specifically: decompose the semantic common vector and the semantic difference vector based on the source text embedding data corresponding to the source image, and calculate the optimized target text embedding data based on the semantic common vector, the semantic difference vector, and the target text embedding data, so as to obtain optimized target text embedding data that combines semantic commonality and semantic difference.
[0147] Step 2: Based on the multi-view image feature set, each perspective encoding, and the optimized target text embedding data, obtain target multi-view images corresponding to multiple source images from each perspective, so as to achieve non-rigid editing.
[0148] It should be noted that in this embodiment, the optimized target text embedding data combines semantic commonality and semantic difference, which can accurately control the degree of non-rigid editing of the source image, so that the obtained target multi-view images are more accurate.
[0149] In one way, the multi-view image editing model that has been trained to convergence obtains a certain multi-view image feature in the multi-view image feature set. This multi-view image feature corresponds to the source image Is1. Then, based on the perspective encoding corresponding to the source image Is1 and the optimized target text embedding data, the target multi-view image corresponding to the source image Is1 is output; at the same time, for the remaining source images, the same method is also used to output the corresponding target multi-view images.
[0150] This embodiment provides a three-dimensional model editing method. In this embodiment, the optimized target text embedding data can more accurately control non-rigid editing, making the finally obtained target multi-view images more accurate.
[0151] Embodiment 8
[0152] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to obtain target multi-view images corresponding to multiple source images from each perspective based on a multi-view image feature set, each perspective encoding, and the optimized target text embedding data, so as to achieve non-rigid editing.
[0153] Figure 7 Schematic diagram of a 3D model editing method provided for the eighth embodiment. As Figure 7 shown, it includes:
[0154] S701, obtaining a preset time step.
[0155] S702, inputting the encoded views, preset time step, optimized target text embedding data, and multiple source images into a preset convolutional module to output the denoising conditions of the multiple source images.
[0156] Among them, the preset convolutional module can be a ResBlock basic module.
[0157] It should be noted that in this embodiment, the encoded views, preset time step, optimized target text embedding data, and multiple source images are used as conditional information and input into the preset convolutional module, so as to output the denoising conditions of the multiple source images.
[0158] S703, controlling the multi-view image feature set to perform target actions according to the optimized target text embedding data to obtain the edited multi-view images under the views corresponding to the view encodings.
[0159] It should be noted that the multi-view image feature set includes multi-view image features corresponding to multiple source images. For each multi-view image feature, perform target actions according to the corresponding optimized target text embedding data, so as to obtain the multi-view images under the views corresponding to the view encodings.
[0160] Among them, the multi-view image refers to the un-denoised and multi-view combined image after editing.
[0161] Among them, the denoising condition refers to the condition that needs to be denoised during the editing process.
[0162] It should be noted that in this embodiment, when obtaining multi-view images through the cross-attention mechanism, the trained-to-converge multi-view image editing model can dynamically adjust the image features according to the relationship between each part of the source image and the conditional information, so that the trained-to-converge multi-view image editing model controls the generation of relevant parts of the multi-view image according to the view encoding and the optimized target text embedding data, so as to realize that the target object performs target actions in the source state to form a target state, and generate multi-view images under the views corresponding to the view encodings.
[0163] S704, denoising each edited multi-view image based on the denoising conditions to obtain the target multi-view images under the corresponding views, so as to realize non-rigid editing.
[0164] Among them, the denoising conditions can include view encoding and a preset time step, and thus the time step and view encoding can be controlled during the editing process to control the noise level and complete denoising, thereby realizing the view control design. Driven by the optimized target text embedding data, the multi-view image editing model trained to convergence can effectively combine temporal dynamics, view information, and text prompts to generate high-quality and view-controllable images.
[0165] This embodiment provides a three-dimensional model editing method. In this embodiment, denoising conditions are obtained, and then, based on the denoising conditions, denoising is performed on the generated multi-view images, so that the obtained target multi-view images have higher quality; and in this embodiment, the multi-view image editing model trained to convergence can effectively combine temporal dynamics, view information, and the optimized target text embedding data.
[0166] Embodiment Nine
[0167] This embodiment is a further refinement of any of the above embodiments. In this embodiment, the multi-view image editing model trained to convergence is a model with a three-dimensional self-attention module added to the diffusion model trained to convergence.
[0168] Low-rank adaptation layers are respectively added to the calculation layers of the query vector, key vector, and value vector in the three-dimensional self-attention module.
[0169] Among them, the diffusion model trained to convergence can be a pre-trained diffusion model.
[0170] It should be noted that in this embodiment, the three-dimensional self-attention module can process the source image from multiple views, so that the obtained target multi-view image combines multiple views and improves multi-view consistency.
[0171] This embodiment provides a three-dimensional model editing method. In this embodiment, since the multi-view image editing model trained to convergence is a model with a three-dimensional self-attention module added to the diffusion model trained to convergence, due to the addition of the three-dimensional self-attention module, the multi-view image editing model trained to convergence can combine multiple views, calculate the source images with different views input through the three-dimensional self-attention mechanism, thereby fusing different view information, outputting effective target multi-view images, and outputting multi-view consistency.
[0172] Embodiment Ten
[0173] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way before non-rigid editing by inputting multiple source images and target text prompt information into the multi-view image editing model trained to convergence and outputting the target multi-view images of the multiple source images at the corresponding views.
[0174] Figure 8 Schematic diagram of a 3D model editing method provided for Example X. As Figure 8 shown, it includes:
[0175] S801, obtaining a training data set; the training data set includes at least one source training image from different perspectives and its corresponding edited image.
[0176] Among them, the training data set can be obtained from an open-source database, and at least one image from different perspectives is determined as the source training image.
[0177] In one way, at least one source training image can be respectively input into a diffusion model that has been trained to convergence to obtain the corresponding edited image.
[0178] S802, keeping the diffusion model that has been trained to convergence unchanged, training the training parameters of the low-rank adaptation layer until the convergence condition is reached.
[0179] Among them, the convergence condition refers to the condition for the multi-perspective image editing model to reach convergence. The convergence condition can include the number of training times or whether the training result is preset to be consistent with the edited image in step S801, or the multi-perspective diffusion loss value output satisfies the preset loss value.
[0180] Exemplarily, when training the multi-perspective image editing model in this step, a multi-perspective diffusion loss function is adopted, as shown in (7):
[0181] (7)
[0182] Among them, represents the noise ground truth of N perspectives, is the input image with noise for N perspectives, that is, the image formed by adding noise to the edited image, represents the parameter weights of the multi-perspective image editing model, represents the optimized training text embedding data, represents the preset time step of the multi-perspective image editing model, represents the perspective encoding of N perspectives, represents the image to be edited, (1:N) represents from the first perspective to the Nth perspective, represents the multi-perspective diffusion loss value, refers to the multi-perspective image editing model.
[0183] Specifically, input the training data set into the multi-perspective image editing model, and use the above (7) to obtain the multi-perspective diffusion loss value corresponding to the training data set, and train the training parameters of the low-rank adaptation layer in the multi-perspective image editing model based on the multi-perspective diffusion loss value until the convergence condition is reached.
[0184] S803, determine the multi-view image editing model that meets the convergence condition as the multi-view image editing model trained to convergence.
[0185] This embodiment provides a 3D model editing method. During training in this embodiment, since the diffusion model trained to convergence has been trained, only the training parameters of the low-rank adaptive layer are changed for training, and the diffusion model trained to convergence remains unchanged, thereby accelerating network convergence and the training process.
[0186] Embodiment Eleven
[0187] This embodiment is a further refinement of any of the above embodiments and is an optional way to obtain the training dataset.
[0188] Figure 9 It is a schematic diagram of a 3D model editing method provided for Embodiment Eleven. As Figure 9 shown, it specifically includes:
[0189] S901, obtain a preset 3D object dataset; the preset 3D object dataset includes at least one preset 3D object.
[0190] Among them, the preset 3D object dataset can be an open-source dataset or a user's private dataset.
[0191] Among them, the preset 3D object dataset also includes at least one text description of the preset 3D object.
[0192] S902, for each preset 3D object, obtain at least one to-be-edited image corresponding to the preset elevation angle and the preset training view angle.
[0193] Among them, the preset elevation angle refers to the elevation angle when facing the center of the preset 3D object set in advance. For example, the preset elevation angle can be between 0 - 30°.
[0194] Among them, the preset training view angle refers to the view angle facing the preset 3D object set in advance. It should be noted that the preset training view angle can be the front, left side, right side, and back, or other view angles, which are not limited here.
[0195] Exemplarily, assume that M preset 3D objects are obtained, and for each preset 3D object, at least one to-be-edited image corresponding to a preset elevation angle and a preset training view is obtained. Specifically, assume that the preset elevation angle is 10°, and the preset training views are the front, left side, right side, and back. Then, for one preset 3D object, one to-be-edited image with a preset elevation angle of 10° and a preset training view of the front is obtained, one to-be-edited image with a preset elevation angle of 10° and a preset training view of the left side is obtained, one to-be-edited image with a preset elevation angle of 10° and a preset training view of the right side is obtained, and one to-be-edited image with a preset elevation angle of 10° and a preset training view of the back is obtained. Thus, four to-be-edited images can be obtained.
[0196] Exemplarily, for M preset 3D objects, 4M to-be-edited images can be obtained.
[0197] S903. Obtain a training dataset based on at least one to-be-edited image.
[0198] In one way, the 4M to-be-edited images obtained in S902 above are input into any image editing model that has been trained to convergence or a single-view image editing model that has been trained to convergence to output the corresponding edited images. It can be understood that the number of edited images is 4M.
[0199] Among them, the single-view image editing model that has been trained to convergence refers to a model that performs single-view editing on a source image based on target text prompt information, so that the target object in the source image changes from the source state to the target state under single-view control.
[0200] Among them, the image editing model that has been trained to convergence can be any diffusion model for editing images.
[0201] Furthermore, a training dataset is determined from the above 4M to-be-edited images and the corresponding 4M edited images.
[0202] This embodiment provides a three-dimensional model editing method. In this embodiment, at least one to-be-edited image is specifically determined from at least one preset 3D object according to a preset elevation angle and a preset training view, and then a training dataset is obtained based on the to-be-edited image.
[0203] Embodiment Twelve
[0204] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to obtain a training dataset based on at least one to-be-edited image, including:
[0205] Step 1: Input at least one image to be edited into a single-view image editing model that has been trained to convergence, respectively, to output corresponding subsets of edited images; the subsets of edited images include at least one edited image corresponding to each image to be edited.
[0206] Among them, the single-view image editing model trained to convergence is a model with a low-rank adaptation layer added to the key vector and value vector of the cross-attention module of the single-view image generation model trained to convergence. Among them, the single-view image generation model trained to convergence is a model obtained by continuing to train on the basis of a pre-trained basic diffusion model by optimizing the input and training parameters included.
[0207] Exemplarily, assume 4M images to be edited, which are respectively input into the single-view image editing model trained to convergence, and the corresponding edited images are output respectively. Finally, 4M edited images are obtained, and the 4M edited images are determined as the subsets of edited images.
[0208] Step 2: Continue to input at least one edited image into the single-view image editing model trained to convergence respectively to output a new subset of edited images until the obtained set of edited images meets the data preparation conditions. Determine at least one source training image and its corresponding edited image from at least one image to be edited and the set of edited images under different perspectives to obtain a training dataset; the set of edited images is composed of at least one subset of edited images; the source training images are determined from the images to be edited.
[0209] Further, to ensure accuracy and multi-view consistency, the at least one edited image obtained is respectively input into the single-view image editing model trained to convergence, so as to respectively output new edited images and obtain new subsets of edited images, and so on until the data preparation conditions are met.
[0210] Among them, the data preparation conditions refer to the conditions for preparing the training dataset. Exemplarily, it can be that the number of edited images included in the set of edited images reaches a preset number, thus meeting the data preparation conditions. Exemplarily, the data preparation conditions can be the condition that the number of edited images is greater than or equal to 64M. If the set of edited images includes 64M edited images, it is determined that the set of edited images meets the data preparation conditions.
[0211] It should be noted that the set of edited images includes at least one edited image (i.e., the subset of edited images) output by the single-view image editing model trained to convergence each time. Exemplarily, assume 16 cycles, then 16 subsets of edited images are obtained, and the set of edited images is composed of 16 subsets of edited images.
[0212] Exemplarily, there are a total of 4M images to be edited. At least one image to be edited is selected from the 4M images to be edited as the source training image, and the source training image includes images from different perspectives. The edited image corresponding to the source training image is selected from the edited image set, and the training object in the edited image can more accurately represent the change from the initial training state to the final training state after the training action.
[0213] It should be noted that each time the single - perspective image editing model trained to convergence inputs an image to be edited, it is one image.
[0214] This embodiment provides a three - dimensional model editing method. In this embodiment, at least one image to be edited is input into the single - perspective image editing model trained to convergence to obtain at least one corresponding edited image. Then, the at least one edited image obtained each time is repeatedly input into the single - perspective image editing model trained to convergence to obtain a new subset of edited images. This is repeated until the edited image set meets the data preparation conditions. Then, at least one source training image is determined from the at least one image to be edited, and at least one edited image is determined from the edited image set. Finally, the obtained training data set improves the multi - perspective consistency and is more accurate.
[0215] Embodiment Thirteen
[0216] Figure 10 It is a schematic diagram of the overall process of a three - dimensional model editing method provided for Embodiment Thirteen. As Figure 10 shown, it includes a source three - dimensional model 1001, multiple source images 1002, a multi - perspective image editing model 1003 trained to convergence, a target multi - perspective image 1004, and a target three - dimensional model 1005.
[0217] Among them, in the multi - perspective image editing model 1003 trained to convergence, a three - dimensional cost frustum is first obtained, and then a multi - perspective image feature set is obtained based on the three - dimensional cost frustum through a depth - aware multi - perspective self - attention mechanism. After passing through a cross - attention mechanism, view encoding, and optimized target text - embedded data, the target multi - perspective image 1004 corresponding to the multiple source images 1002 is obtained.
[0218] Figure 10 In, taking two source images 1002 as an example for illustration, one of the source images 1002 is a front - view source image 1002, and the other source image 1002 is a left - view source image 1002. As Figure 10 shown, finally, the target multi - perspective images 1004 corresponding to the two source images are output simultaneously.
[0219] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0220] Embodiment Fourteen
[0221] The following is an embodiment of the device of the present application. Figure 11 A schematic structural diagram of a three-dimensional model editing device provided for Embodiment Fourteen. The three-dimensional model editing device 1100 includes the following modules:
[0222] An acquisition module 1101, configured to acquire multiple source images from a source three-dimensional model from different perspectives. The source images are images corresponding to the three-dimensional model from different perspectives; each source image includes a target object; each target object is in a source state;
[0223] A determination module 1102, configured to determine the target text prompt information corresponding to each target multi-perspective image to be generated; the target text prompt information includes the target state information of the target object;
[0224] An output module 1103, configured to input the multiple source images and the target text prompt information into a multi-perspective image editing model that has been trained to convergence, and output the target multi-perspective images of the multiple source images from the corresponding perspectives to implement non-rigid editing; the target multi-perspective images are images in which the target object performs a target action from the source state to form a target state from the perspective of the corresponding source image;
[0225] A reconstruction module 1104, configured to reconstruct a target three-dimensional model based on each target multi-perspective image.
[0226] Optionally, when the output module 1103 inputs the multiple source images and the target text prompt information into a multi-perspective image editing model that has been trained to convergence and outputs the target multi-perspective images of the multiple source images from the corresponding perspectives to implement non-rigid editing, it is specifically configured to:
[0227] Obtain the camera pose information of the multiple source images;
[0228] Extract features from the multiple source images at a preset time step to obtain a source image feature set at the preset time step;
[0229] Obtain a three-dimensional cost frustum at the preset time step based on each camera pose information and the source image feature set;
[0230] Use the multi-perspective image editing model trained to convergence to output the target multi-perspective images of the multiple source images from the corresponding perspectives based on the three-dimensional cost frustum and the target text prompt information to implement non-rigid editing.
[0231] Optionally, when the output module 1103 obtains the three-dimensional cost frustum at a preset time step based on the camera pose information and the source image feature set, it is specifically used for:
[0232] Determine the camera pose transformation matrix based on the camera pose information and the preset conversion perspective information; the camera pose transformation matrix is the transformation matrix from the perspective corresponding to the camera pose information to the preset conversion perspective information on the source depth; the camera pose information includes the perspective information corresponding to the source image;
[0233] Input the camera pose transformation matrix and the source image feature set into the preset conversion perspective algorithm to output the conversion feature set of each source image from the corresponding perspective to the preset conversion perspective at the preset time step;
[0234] Calculate the variance between the conversion feature sets to obtain the three-dimensional cost frustum corresponding to the preset time step.
[0235] Optionally, when the output module 1103 uses the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of multiple source images based on the three-dimensional cost frustum and the target text prompt information to achieve non-rigid editing, it is specifically used for:
[0236] Calculate the depth fusion feature at the preset time step based on the reshaping of the three-dimensional cost frustum; the depth fusion feature is the fusion feature on the source depth;
[0237] Calculate the query vector, key vector, and value vector based on the depth fusion feature using a linear transformation layer;
[0238] Input the query vector, key vector, and value vector into the preset three-dimensional self-attention mechanism algorithm to output the multi-view fusion feature at the preset time step; the multi-view fusion feature is the feature after fusing multiple views through the three-dimensional self-attention mechanism;
[0239] Use the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of multiple source images based on the multi-view fusion feature and the target text prompt information to achieve non-rigid editing.
[0240] Optionally, when the output module 1103 uses the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of multiple source images based on the multi-view fusion feature and the target text prompt information to achieve non-rigid editing, it is specifically used for:
[0241] According to the camera pose information corresponding to each source image, project the multi-view fusion features onto the viewpoints corresponding to each source image to obtain a multi-view image feature set for depth perception at a preset time step; the multi-view image feature set is a feature set obtained by calculating the three-dimensional self-attention mechanism of multiple source images through depth perception; the multi-view image feature set includes multi-view image features corresponding to multiple source images.
[0242] Use the multi-view image editing model trained to convergence to output the target multi-view images under the viewpoints corresponding to multiple source images based on the multi-view image feature set and the target text prompt information, so as to achieve non-rigid editing.
[0243] Optionally, when the output module 1103 uses the multi-view image editing model trained to convergence to output the target multi-view images under the viewpoints corresponding to multiple source images based on the multi-view image feature set and the target text prompt information to achieve non-rigid editing, it is specifically used for:
[0244] Obtain the corresponding viewpoint information from the camera pose information of each source image.
[0245] Encode each viewpoint information using a preset multi-layer perceptron to obtain each viewpoint encoding.
[0246] Use the multi-view image editing model trained to convergence to output the target multi-view images under the viewpoints corresponding to multiple source images based on each viewpoint encoding, the multi-view image feature set, and the target text prompt information, so as to achieve non-rigid editing.
[0247] Optionally, when the output module 1103 uses the multi-view image editing model trained to convergence to output the target multi-view images under the viewpoints corresponding to multiple source images based on each viewpoint encoding, the multi-view image feature set, and the target text prompt information to achieve non-rigid editing, it is specifically used for:
[0248] Encode and optimize the target text prompt information to obtain optimized target text embedding data.
[0249] Based on the multi-view image feature set, each viewpoint encoding, and the optimized target text embedding data, obtain the target multi-view images under the viewpoints corresponding to multiple source images, so as to achieve non-rigid editing.
[0250] Optionally, when the output module 1103 obtains the target multi-view images under the viewpoints corresponding to multiple source images based on the multi-view image feature set, each viewpoint encoding, and the optimized target text embedding data to achieve non-rigid editing, it is specifically used for:
[0251] Obtain the preset time step.
[0252] Encode each perspective, preset time steps, optimized target text embedding data, and multiple source images into a preset convolutional module to output the denoising conditions of the multiple source images;
[0253] Control the multi-perspective image feature set to perform target actions according to the optimized target text embedding data to obtain the edited multi-perspective images under the perspectives corresponding to the perspective encodings;
[0254] Denoise each edited multi-perspective image based on the denoising conditions to obtain the target multi-perspective images under the corresponding perspectives.
[0255] Optionally, the multi-perspective image editing model trained to convergence is a model with a three-dimensional self-attention module added to the diffusion model trained to convergence;
[0256] Low-rank adaptation layers are respectively added to the calculation layers of the query vector, key vector, and value vector in the three-dimensional self-attention module.
[0257] Optionally, before inputting multiple source images and target text prompt information into the multi-perspective image editing model trained to convergence to output the target multi-perspective images of the multiple source images under the corresponding perspectives to achieve non-rigid editing, this embodiment provides a three-dimensional model editing device, further including: a training module;
[0258] An acquisition module 1101, further configured to acquire a training data set; the training data set includes at least one source training image under different perspectives and its corresponding edited image;
[0259] A training module, configured to keep the diffusion model trained to convergence unchanged, and train the training parameters of the low-rank adaptation layer until the convergence condition is reached;
[0260] A determination module 1102, further configured to determine the multi-perspective image editing model that reaches the convergence condition as the multi-perspective image editing model trained to convergence.
[0261] Optionally, when the acquisition module 1101 acquires the training data set, it is specifically configured to:
[0262] Acquire a preset 3D object data set; the preset 3D object data set includes at least one preset 3D object;
[0263] For each preset 3D object, acquire at least one image to be edited corresponding to a preset elevation angle and a preset training perspective;
[0264] Acquire a training data set based on at least one image to be edited.
[0265] Optionally, when the acquisition module 1101 acquires the training data set based on at least one image to be edited, it is specifically configured to:
[0266] Input at least one image to be edited into a single-view image editing model that has been trained to convergence, respectively, to output corresponding subsets of edited images; the subsets of edited images include at least one edited image corresponding to the image to be edited.
[0267] Continue to input at least one edited image into the single-view image editing model that has been trained to convergence, respectively, to output a new subset of edited images until the set of edited images obtained meets the data preparation conditions. Determine at least one source training image and its corresponding edited image from at least one image to be edited and the set of edited images under different perspectives to obtain a training data set; the set of edited images is composed of at least one subset of edited images; the source training image is determined from the images to be edited.
[0268] For the description of the features in the corresponding embodiments of the three-dimensional model editing device, reference may be made to the relevant descriptions in the corresponding embodiments of the three-dimensional model editing method, which will not be elaborated here one by one.
[0269] Embodiments of the present application also provide an editing device. Figure 12 It is a schematic structural diagram of the editing device provided in Embodiment XIV. As shown, the editing device 1200 includes a processor 1201 and a memory 1202. Among them, the processor 1201, the memory 1202, and the communication component 1203 are connected through a bus 1204. The memory 1202 stores a computer program, and the processor 1201 is configured to run the computer program to execute the steps in any of the above embodiments of the three-dimensional model editing method.
[0270] Embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is configured to execute the steps in any of the above embodiments of the three-dimensional model editing method when running.
[0271] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0272] Embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the three-dimensional model editing method.
[0273] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the three-dimensional model editing method.
[0274] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0275] The above has introduced in detail a three-dimensional model editing method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A three-dimensional model editing method, characterized in that Including: Obtain multiple source images from different perspectives of the source 3D model, where the source images are the images corresponding to the 3D model from different perspectives, and each source image includes a target object in the source state; Determine the target text prompt information corresponding to each target multi-perspective image to be generated; The target text prompt information includes the target state information of the target object; Obtain the camera pose information of the multiple source images; Extract features from the multiple source images at a preset time step to obtain a source image feature set at the preset time step; Based on each of the camera pose information and the source image feature set, obtain a 3D cost frustum at the preset time step; Use a multi-perspective image editing model trained to convergence to output the target multi-perspective images corresponding to the perspectives of the multiple source images based on the 3D cost frustum and the target text prompt information, so as to achieve non-rigid editing; the target multi-perspective images are images in which the target object performs a target action and forms a target state from the source state in the perspective of the corresponding source image; Reconstruct a target 3D model based on each target multi-perspective image.
2. The method according to claim 1, characterized in that The obtaining of the 3D cost frustum at the preset time step based on each of the camera pose information and the source image feature set includes: Determine each camera pose transformation matrix based on each of the camera pose information and preset transformation perspective information; the camera pose transformation matrix is a transformation matrix for converting from the perspective corresponding to the camera pose information to the preset transformation perspective information at the source depth; the camera pose information includes the perspective information corresponding to the source image; Input the camera pose transformation matrix and the source image feature set into a preset transformation perspective algorithm to output a transformation feature set of each source image from the corresponding perspective to the preset transformation perspective at the preset time step; Calculate the variance between each transformation feature set to obtain the 3D cost frustum corresponding to the preset time step.
3. The method according to claim 1, characterized in that, The using of a multi-perspective image editing model trained to convergence to output the target multi-perspective images corresponding to the perspectives of the multiple source images based on the 3D cost frustum and the target text prompt information to achieve non-rigid editing includes: Calculate the depth fusion feature at the preset time step based on the 3D cost frustum reshaping; the depth fusion feature is a fusion feature at the source depth; Calculate a query vector, a key vector, and a value vector based on the depth fusion feature using a linear transformation layer; Input the query vector, the key vector, and the value vector into a preset 3D self-attention mechanism algorithm to output a multi-perspective fusion feature at the preset time step; the multi-perspective fusion feature is a feature after fusing multiple perspectives through the 3D self-attention mechanism; Use a multi-perspective image editing model trained to convergence to output the target multi-perspective images corresponding to the perspectives of the multiple source images based on the multi-perspective fusion feature and the target text prompt information to achieve non-rigid editing.
4. The method according to claim 3, characterized in that, Using the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of the multiple source images based on the multi-view fusion features and the target text prompt information to achieve non-rigid editing, including: According to the camera pose information corresponding to each source image, project the multi-view fusion features to the perspective corresponding to each source image to obtain a multi-view image feature set with depth perception at the preset time step; the multi-view image feature set is the feature set obtained by calculating the three-dimensional self-attention mechanism of the multiple source images through depth perception; the multi-view image feature set includes multi-view image features corresponding to the multiple source images; Using the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of the multiple source images based on the multi-view image feature set and the target text prompt information to achieve non-rigid editing.
5. The method according to claim 4, wherein Using the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of the multiple source images based on the multi-view image feature set and the target text prompt information to achieve non-rigid editing, including: Obtain the corresponding perspective information from the camera pose information of each source image; Encode each perspective information using a preset multi-layer perceptron to obtain each perspective encoding; Using the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of the multiple source images based on each perspective encoding, the multi-view image feature set, and the target text prompt information to achieve non-rigid editing.
6. The method according to claim 5, wherein Using the multi-view image editing model trained to convergence to output the target multi-view images corresponding to the perspectives of the multiple source images based on each perspective encoding, the multi-view image feature set, and the target text prompt information to achieve non-rigid editing, including: Encode and optimize the target text prompt information to obtain optimized target text embedding data; Based on the multi-view image feature set, each perspective encoding, and the optimized target text embedding data, obtain the target multi-view images corresponding to the perspectives of the multiple source images to achieve non-rigid editing.
7. The method according to claim 6, wherein Based on the multi-view image feature set, each perspective encoding, and the optimized target text embedding data to obtain the target multi-view images corresponding to the perspectives of the multiple source images to achieve non-rigid editing, including: Obtain the preset time step; Input each perspective encoding, the preset time step, the optimized target text embedding data, and the multiple source images into a preset convolutional module to output the denoising conditions of the multiple source images; Control the multi-view image feature set to perform a target action according to the optimized target text embedding data to obtain the edited multi-view images under the perspective corresponding to the perspective encoding; Denoise each of the edited multi-view images based on the denoising conditions to obtain the target multi-view images under the corresponding perspectives to achieve non-rigid editing.
8. The method according to claim 1, wherein The multi-view image editing model trained to convergence is a model with a three-dimensional self-attention module added to the diffusion model trained to convergence; Low-rank adaptation layers are respectively added to the calculation layers of the query vector, key vector, and value vector in the three-dimensional self-attention module.
9. The method according to claim 8, wherein Before inputting the multiple source images and the target text prompt information into the multi-view image editing model that has been trained to convergence and outputting the target multi-view images of the multiple source images from corresponding perspectives to achieve non-rigid editing, it further includes: Obtaining a training data set; the training data set includes at least one source training image from different perspectives and its corresponding edited image; Keeping the diffusion model that has been trained to convergence unchanged, training the training parameters of the low-rank adaptation layer until the convergence condition is reached; Determining the multi-view image editing model that reaches the convergence condition as the multi-view image editing model that has been trained to convergence.
10. The method according to claim 9, characterized in that, The obtaining of the training data set includes: Obtaining a preset 3D object data set; the preset 3D object data set includes at least one preset 3D object; For each of the preset 3D objects, obtaining at least one to-be-edited image corresponding to a preset elevation angle and a preset training perspective; Obtaining a training data set based on the at least one to-be-edited image.
11. The method according to claim 10, characterized in that, The obtaining of the training data set based on the at least one to-be-edited image includes: Respectively inputting the at least one to-be-edited image into the single-view image editing model that has been trained to convergence to respectively output corresponding subsets of edited images; the subsets of edited images include the edited images corresponding to the at least one to-be-edited image; Continuing to respectively input at least one edited image into the single-view image editing model that has been trained to convergence to output new subsets of edited images until the obtained set of edited images meets the data preparation condition, and determining at least one source training image from different perspectives and its corresponding edited image from the at least one to-be-edited image and the set of edited images to obtain a training data set; the set of edited images is composed of at least one subset of edited images; the source training image is determined from the to-be-edited images.
12. An editing device, characterized in that, It includes: A memory and a processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1-11.
13. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1-11.
14. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Three-dimensional model data processing method and system, product, equipment and medium
CN118864741A