3D Model Editing Method, Device, Storage Medium, and Program Product
By acquiring source images and text embedding data from different perspectives, the trained single-view image editing model is used to generate target single-view images, which solves the problem that the state changes of three-dimensional model in the prior art are not possible, and high-quality non-rigid editing and multi-view consistency are achieved.
Patent Information
- Application Number
- CN202510192923.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art cannot effectively realize non-rigid editing of objects in a three-dimensional model when changing from one state to another. Especially in augmented/virtual reality, movie/game production and artistic creation, existing editors can only perform the same type of replacement or minor feature changes.
By obtaining multiple source images from different perspectives from the source three-dimensional model, determining the initial and target text prompt information, using the single-view image editing model trained to convergence, the target single-view image is generated based on the source and target text embedding data, non-rigid editing is achieved, and the target three-dimensional model is finally reconstructed.
The non-rigid editing of the target object in the state is achieved, and the consistency of multi-view angles is improved, making the reconstructed target three-dimensional model more accurate.
Smart Images

Figure CN119672275B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a three-dimensional model editing method, device, storage medium, and program product. Background Art
[0002] Currently, three-dimensional dynamic changes are widely used in augmented / virtual reality, movie / game production, and artistic creation. Therefore, how to achieve high-quality three-dimensional dynamic changes is particularly important.
[0003] In related technologies, an editor can perform same-type replacement on objects in a source three-dimensional model or make changes to other minor features in the same state. However, when changing an object in the source three-dimensional model from one state to another, related technologies cannot handle it. Summary of the Invention
[0004] This application provides a three-dimensional model editing method, device, storage medium, and program product to achieve the effect of changing the state of a target object.
[0005] In a first aspect, this application provides a three-dimensional model editing method, including:
[0006] Obtain multiple source images from different perspectives of a source three-dimensional model, where the source images are images corresponding to the three-dimensional model from different perspectives; each source image includes a target object; the target object is in a source state;
[0007] Determine the initial text prompt information corresponding to each source image and the target text prompt information corresponding to each target single-perspective image to be generated. The initial text prompt information includes the source state information of the target object, and the target text prompt information includes the target state information of the target object;
[0008] Determine the source text embedding data corresponding to each source image according to each initial text prompt information; and determine the target text embedding data corresponding to each target single-perspective image to be generated according to each target text prompt information. Each target single-perspective image is an image in which the target object performs a target action from the source state to form a target state in the perspective of the corresponding source image;
[0009] For each source image, use a single-perspective image editing model that has been trained to convergence and determine the target single-perspective image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image to achieve non-rigid editing;
[0010] Reconstruct a target three-dimensional model based on each target single-perspective image.
[0011] In a second aspect, an embodiment of this application provides an editing device, including: a memory and a processor;
[0012] The memory stores computer-executable instructions;
[0013] The processor executes the computer-executable instructions stored in the memory, such that the processor performs the above first aspect and / or various possible implementation manners of the first aspect.
[0014] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above first aspect and / or various possible implementation manners of the first aspect.
[0016] The three-dimensional model editing method, device, storage medium, and program product provided by the embodiments of the present application. In the present application, first, source images corresponding to different perspectives are obtained from a source three-dimensional model, and then initial text prompt information and target text prompt information corresponding to each source image are determined. Among them, the initial text prompt information includes source state information of a target object, and the target text prompt information includes target state information of the target object. Then, source text embedding data corresponding to the source images are determined according to the initial text prompt information. Further, target text embedding data is determined according to the target text prompt information for generating a target single-perspective image, where the target single-perspective image is an image in which the target object forms a target state from a source state after performing a target action in the perspective of the corresponding source image. Further, according to each source image, a trained-to-converge single-perspective image editing model is used to respectively determine the corresponding target single-perspective image, so that the edited target single-perspective images corresponding to each source image can be obtained, and then the target three-dimensional model is edited and reconstructed based on each target single-perspective image. It can be seen that in the present application, with the help of the target text prompt information, a target single-perspective image in which the source image forms a target state after performing a target action in the corresponding source state can be obtained; in addition, the present application can achieve non-rigid editing in which the target object changes in state under the action of the target text prompt information, and further achieve text-driven image non-rigid editing; in the present application, for the source three-dimensional model, source images are respectively obtained from multiple perspectives, and edited target single-perspective images are respectively obtained from multiple source images. Since the present application considers from multiple perspectives, the multi-perspective consistency is improved, so that the target three-dimensional model obtained based on multiple target single-perspective images is more accurate. Description of the Drawings
[0017] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 Schematic diagram of the three-dimensional model editing scene provided by the present application;
[0019] Figure 2 Schematic diagram of a three-dimensional model editing method provided for Embodiment 1;
[0020] Figure 3 Schematic diagram of a three-dimensional model editing method provided for Embodiment 2;
[0021] Figure 4 Schematic diagram of a three-dimensional model editing method provided for Embodiment 6;
[0022] Figure 5 Schematic diagram of a three-dimensional model editing method provided for Embodiment 7;
[0023] Figure 6 Schematic diagram of a three-dimensional model editing method provided for Embodiment 10;
[0024] Figure 7 Schematic diagram of a three-dimensional model editing method provided for Embodiment 11;
[0025] Figure 8 Schematic diagram of a three-dimensional model editing method provided for Embodiment 12;
[0026] Figure 9 Overall process schematic diagram of a three-dimensional model editing method provided for Embodiment 13;
[0027] Figure 10 Schematic diagram of the structure of a three-dimensional model editing device provided for Embodiment 14;
[0028] Figure 11 Schematic diagram of the structure of the editing device provided for Embodiment 14. Detailed implementation manners
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0030] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0031] In order to enable those skilled in the art of this technical field to better understand the solution of this application, the following further details this application in conjunction with the accompanying drawings and specific embodiments.
[0032] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the three-dimensional model editing method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0033] Currently, three-dimensional dynamic changes are widely used in augmented / virtual reality, film / game production, and artistic creation. Therefore, how to achieve high-quality three-dimensional dynamic changes is particularly important.
[0034] In the related art, an editor can perform same-type replacement or minor feature changes on objects in a source three-dimensional model, for example: color change. However, when faced with state changes, it cannot be processed.
[0035] To address the deficiencies of the related art, the inventors of this solution conducted creative research and designed a new solution. This solution provides a 3D model editing method. To solve the problem in the related art that only partial feature changes of the target object can be achieved in the same state or the target object can be replaced with the same type of object, in this application, non-rigid editing is realized based on text driving. Specifically, in this application, multiple source images from different perspectives are obtained from the source 3D model, and then the initial text prompt information and the target text prompt information are determined. Among them, the target text prompt information includes the target state information of the target object, that is, it means that the target object forms a target state after performing a target action from the source state. Therefore, this application can achieve changes in the state of the target object, and then realizes non-rigid editing of the source images, increasing the application scenarios. To solve the accuracy problem, in this application, multiple source images from different perspectives can be obtained. For each source image, a single-view image editing model that has been trained to convergence is used to determine the corresponding target single-view image of the source image based on the source text embedding data and the target text embedding data corresponding to the source image, and then the target 3D model is reconstructed according to multiple target single-view images. At this time, the target object in the target 3D model is a 3D model formed after performing the target action. Since this application considers the source images from different perspectives, the target 3D model can be made more accurate. It should be noted that in this application, the single-view image editing model that has been trained to convergence edits each source image respectively to obtain the corresponding target single-view image after editing.
[0036] Figure 1 FIG. is a schematic diagram of the 3D model editing scenario provided by this application, as Figure 1 shown, the specific application scenario of this application includes an editing device 101.
[0037] Among them, the editing device 101 can be a computer or a server, and no limitation is made here.
[0038] In this scenario, the editing device 101 obtains source images from different perspectives from the source 3D model, and obtains the target text prompt information, and inputs the source images and the target text prompt information into the single-view image editing model that has been trained to convergence respectively, and outputs the target single-view images corresponding to the perspectives. Repeating this process, target single-view images from multiple perspectives are obtained.
[0039] Furthermore, the editing device 101 reconstructs the target 3D model based on multiple target single-view images.
[0040] As can be seen from the above scenarios, in the related art, it is possible to achieve changes in other features of the target object in the same state, such as color changes, but it is impossible to achieve changes in the state of the target object. In this application, with the help of the target text prompt information, the target text prompt information includes the target state information formed after the target object executes the target action, so that the source image can be non-rigidly edited according to the target text prompt information in the source state; in this application, a single-view image editing model that has been trained to convergence is used, and the target single-view images of each source image after editing can be output respectively in the corresponding view, realizing non-rigid editing; since the single-view image editing model is trained to convergence, the target single-view image is more accurate, that is, it accurately represents the target state formed after the target object executes the target action from the source state; in addition, in this application, since different views are considered, the target three-dimensional model reconstructed based on multiple target single-view images is more accurate.
[0041] An embodiment of the present application provides a three-dimensional model editing method, and the method will be described in detail in combination with the execution process of the three-dimensional model editing method.
[0042] Embodiment 1
[0043] The execution subject of Embodiment 1 to Embodiment 13 of the present application is a three-dimensional model editing device (referred to as the editing device for short), and the editing device is located in the editing device.
[0044] Figure 2 A schematic diagram of a three-dimensional model editing method provided for Embodiment 1. As Figure 2 shown, it specifically includes:
[0045] S201, obtain multiple source images from different perspectives in the source three-dimensional model, where the source images are the images corresponding to the three-dimensional model from different perspectives; each source image includes a target object; the target object is in the source state.
[0046] Among them, the source three-dimensional model refers to the three-dimensional model to be edited for the target object therein.
[0047] Among them, the source image refers to the image corresponding to the source three-dimensional model from different perspectives. It should be noted that the source image is two-dimensional.
[0048] Among them, the perspective can refer to the side, the front or the back, etc., which is not limited here. It should be noted that if the perspective is the side, a source image of the side is the image corresponding to the source three-dimensional model in the front.
[0049] Among them, the target object refers to the object included in the source three-dimensional model or the source image. For example, if a cat is included in the source three-dimensional model, the cat is the target object.
[0050] Among them, the source state refers to the pose of the target object in the source 3D model or the source image. For example, if a standing cat is included in the source 3D model, the source state is "standing". It should be noted that the source states of the target object in the source 3D model and the source image are the same.
[0051] In one way, in this application, rendering technology is used to render source images from the source 3D model from different perspectives. For example, the front source image, the left-side source image, the right-side source image, and the back source image.
[0052] It should be noted that in each source image, only the perspective of the target object is different, but the source states of the target object are the same.
[0053] S202, determine the initial text prompt information corresponding to each source image and the target text prompt information corresponding to each target single-perspective image to be generated. The initial text prompt information includes the source state information of the target object, and the target text prompt information includes the target state information of the target object.
[0054] Among them, the initial text prompt information includes the source state information of the target object. Among them, the source state information includes the information of the target object in terms of state. The initial text prompt information also includes the size, color, etc. of the target object. For example, if the initial text prompt information includes a standing yellow cat, the initial text prompt information is information such as "standing" and "yellow", and the source state information is "standing"; or, if the initial text prompt information includes a standing yellow kitten, the initial text prompt information is information such as "standing", "yellow", and "small", and the source state information is "standing".
[0055] Among them, the target text prompt information includes the target state information of the target object after performing the target action. Among them, the target text prompt information refers to the information such as the state, color, and size of the target object, and the target state information includes the pose information of the target object.
[0056] Among them, if the target text prompt information includes a sitting yellow kitten, the target text prompt information is information such as "sitting", "yellow", and "small", and the target state information is "sitting".
[0057] It should be noted that there is common information in the target text prompt information and the initial text prompt information, that is, the remaining information other than the state of the target object can be common information. For example, the initial text prompt information is "a standing yellow kitten", and the target text prompt information is "a sitting yellow kitten". Among them, only the state of the target object changes, and the remaining information does not change. For example, "yellow" and "small" do not change, which is the common information.
[0058] In this application, the editing device can enable the target object to change from one state to another while the other features remain unchanged. That is, the color and size of the target object can remain unchanged, only the state of the target object changes. It can be understood that the target text prompt information is based on the initial text prompt information, and the generated target single-view image can keep the other features of the target object unchanged and only change the state of the target object.
[0059] In one way, the size, color, or type of the target object can also be changed. It should be noted that color, size, and the type of the target object are also features of the image.
[0060] Among them, the target single-view image is the image of the target object after performing the target action and changing from the source state to the target state.
[0061] S203. Determine the source text embedding data corresponding to each source image according to each initial text prompt information; and determine the target text embedding data corresponding to each to-be-generated target single-view image according to each target text prompt information. Each target single-view image is the image of the target object after performing the target action and changing from the source state to the target state under the perspective of the corresponding source image.
[0062] Among them, the perspective of the target single-view image corresponds to that of the source image. For example, a front source image generates a front target single-view image.
[0063] Among them, the source text embedding data and the target text embedding data are obtained after encoding.
[0064] It should be noted that the source text embedding data includes the source state embedding data of the target object, which can represent the source state of the source image; the target text embedding data includes the target state embedding data of the target object, which can represent the target state of the target single-view image.
[0065] S204. For each source image, use the single-view image editing model that has been trained to convergence and determine the target single-view image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image, so as to achieve non-rigid editing.
[0066] Among them, the single-view image editing model that has been trained to convergence refers to a model that performs single-view editing on a source image based on the target text prompt information, so that the target object in the source image changes from the source state to the target state under single-view control.
[0067] It should be noted that the single-view image editing model that has been trained to convergence processes each source image respectively, and thus sequentially outputs the target single-view images corresponding to each source object.
[0068] It should be noted that in a single editing task, a single-view image editing model that has been trained to convergence receives a source image, edits the above-mentioned single source image, and obtains a target single-view image corresponding to the perspective to complete the current editing task. For the source image in the next different perspective, the corresponding target single-view image is obtained by using the same method as above.
[0069] In this embodiment, using the source text embedding data to output the target single-view image can ensure non-rigid editing under the original features of the source image, that is, ensure that the target single-view image realizes non-rigid editing from the source state to the target state based on the original features of the source image. Among them, the original features may include the color, size, and other information features of the target object in the source image.
[0070] S205. Reconstruct the target three-dimensional model based on each target single-view image.
[0071] Among them, the target three-dimensional model includes the target object. At this time, the target object in the target three-dimensional model is in the target state, that is, the target state formed after the target action is executed.
[0072] Among them, the target three-dimensional model refers to the three-dimensional model formed after the target object in the source three-dimensional model executes the target action.
[0073] Method 1: Construct a three-dimensional model based on the target single-view images corresponding to different perspectives, which is the target three-dimensional model.
[0074] Method 2: Specifically, initialize the edited initial three-dimensional model with the source three-dimensional model. For the edited initial three-dimensional model, render it from a preset perspective to obtain a rendered image , obtain the target single-view image from the target single-view images in different perspectives obtained in step S204 for the preset perspective, denoted as , and and are input into the reconstruction loss function to calculate the iterative reconstruction loss value. Based on the iterative reconstruction loss value, optimize the edited initial three-dimensional model to obtain the optimized three-dimensional model after editing, and complete the first iteration. Among them, the reconstruction loss function is , among which, refers to the image after non-rigid editing, refers to the rendered image, among which, represents the iterative reconstruction loss value. Among them, the source three-dimensional model is the same as the edited initial three-dimensional model.
[0075] Furthermore, for the optimized three-dimensional model after editing, continue the iteration. Specifically: Render from any perspective of the optimized three-dimensional model after editing to obtain the rendered image in the second iteration , input the rendered image into a single-view image editing model that has been trained to convergence, and output the edited image to obtain , and continue to input as well as into the reconstruction loss function to calculate the iterative reconstruction loss value, and continue to optimize the above-optimized 3D model based on this iterative reconstruction loss value. Iterate continuously according to the above steps until the iterative reconstruction loss value or the number of iterations corresponding to a 3D model from any perspective meets the preset iteration condition, then determine the 3D model at this time as the target 3D model. Among them, the preset iteration condition includes a preset reconstruction loss value and / or a preset number of iterations.
[0076] Exemplarily, if the iterative reconstruction loss value calculated from any perspective is less than the preset reconstruction loss value, it is determined that the preset iteration condition is met; or, if the number of iterations of the 3D model has reached the preset number of iterations, it is determined that the preset iteration condition is met.
[0077] In the second method, when iteratively optimizing the 3D model, iterative optimization can be performed from different perspectives, thereby combining multiple perspectives and making the target 3D model more accurate and reasonable.
[0078] It should be noted that in this application, the target state can be a local state change of the target object from the source state, or a global state change of the target object from the source state. Thus, it can be seen that global or local non-rigid editing of the target object can be achieved in this application; in this application, the editing device performs global or local non-rigid editing on the source image, that is, it can enable non-rigid editing of the target object in the source image to achieve global or local state changes.
[0079] It should be noted that the editing device in this application can also achieve changes in the color and size of the target object in the source image. Exemplarily, in one method, in addition to the target object state being inconsistent with the source state, the target text prompt information can also include features to be changed. For example, the initial text prompt information includes "a sitting yellow kitten", and the target text prompt information includes "a standing red kitten". Then, from the source image to the target single-view image, not only does the state of the target object change, but the color also changes, where the color is the feature to be changed. Thus, in this method, a single-view image editing model that has been trained to convergence is used to determine the target single-view image corresponding to the source image, and both the state and color of the target object in the target single-view image change based on the source image.
[0080] This embodiment provides a 3D model editing method. In this embodiment, source images corresponding to different perspectives are first obtained from the source 3D model. The source images include a target object in a source state. Then, the initial text prompt information and the target text prompt information corresponding to each source image are determined. Among them, the initial text prompt information includes the source state information of the target object, and the target text prompt information includes the target state information of the target object. Then, the source text embedding data corresponding to the source image is determined according to each initial text prompt information. Further, the target text embedding data is determined according to the target text prompt information for generating a target single-view image. The target single-view image is an image in which the target object performs a target action in the perspective of the corresponding source image and forms a target state from the source state. Further, according to each source image, the trained-to-converge single-view image editing model is used to determine the corresponding target single-view image based on the source text embedding data and the target text embedding data corresponding to the source image, so that the edited target single-view images corresponding to each source image can be obtained, and then the target 3D model is edited and reconstructed based on each target single-view image. It can be seen that in this application, with the help of the target text prompt information, the source image can form a target single-view image in the target state after performing the target action in the corresponding source state. In this application, the target state can be a local state change of the target object from the source state, or a global state change of the target object from the source state. It can be seen that non-rigid editing of the target object can be realized in this application; in addition, this application can realize non-rigid editing under the action of the target text prompt information, and then realize text-driven non-rigid image editing; in this application, for the source 3D model, source images are obtained from multiple perspectives respectively, and edited target single-view images are obtained from multiple source images respectively. Since this application considers from multiple perspectives, the target 3D model obtained based on multiple target single-view images is more accurate.
[0081] Embodiment 2
[0082] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to determine the initial text prompt information corresponding to each source image.
[0083] Figure 3 Schematic diagram of a 3D model editing method provided for Embodiment 2. As Figure 3 shown, it specifically includes:
[0084] S301, input each source image into a preset text prediction algorithm respectively, and use the preset text prediction algorithm to determine the image description information of the source image respectively.
[0085] Among them, the preset text prediction algorithm can be a multi-modal algorithm.
[0086] Specifically, each source image is input into a preset text prediction algorithm to output image description information for the source image.
[0087] The image description information includes descriptions of the features of the source image. For example, it includes description information about the color, size, and pose of the target object.
[0088] S302: Determine the image description information of each source image as each initial text prompt information.
[0089] This embodiment provides a three-dimensional model editing method. In this embodiment, a preset text prediction algorithm can be used to output the image description information corresponding to each source image, and thus the image description information can be accurately obtained.
[0090] Embodiment Three
[0091] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to determine the target text prompt information corresponding to each target single-view image, including:
[0092] Generate target text prompt information according to each initial text prompt information and the target state.
[0093] Among them, the initial text prompt information includes the color, size of the target object, and the perspective information of the corresponding source image. Thus, the editing device can generate the target text prompt information based on the initial text prompt information and in combination with the target state.
[0094] Thus, the target text prompt information can include the color, size, corresponding perspective, and target state information of the target object.
[0095] It should be noted that the initial text prompt information and the target text prompt information may also include color, size, corresponding perspective information, target object type, target state, and other feature information, which is not limited here.
[0096] It should be noted that the target text prompt information can include the color, size, target state information, and perspective information of the target object. Among them, the color and size of the target object can be obtained from the initial text prompt information. The perspective information is configuration information, and the perspective information can be "perform the target action from the perspective corresponding to the source image", that is, control to perform the target action from the perspective corresponding to the source image to form the target state from the source state. If the color and size of the target object in the target text prompt information remain unchanged, the target text prompt information represents controlling to ensure that the color and size of the target object remain unchanged from the perspective corresponding to the source image and performing the target action to form the target state.
[0097] This embodiment provides a 3D model editing method. In this embodiment, according to the initial text prompt information and the target state, target text prompt information is generated, so that the target text prompt information includes some information in the initial text prompt information. Furthermore, the source image can perform a target action under the original features to form the target state.
[0098] Embodiment 4
[0099] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to determine the source text embedding data corresponding to each source image according to each initial text prompt information, including:
[0100] Step 1: Input each initial text prompt information into a preset text editor, and use the preset text editor to encode each initial text prompt information respectively to obtain each initial text embedding data.
[0101] Among them, the preset text editor is a pre-set text editor.
[0102] Among them, the initial text embedding data can be expressed as, where m is the number of tokens in the text encoding and d is the dimension of each token.
[0103] It should be noted that text embedding data is to convert text data into high-dimensional and dense vector data, and these vector data can measure the semantic and syntactic similarities between different texts.
[0104] Step 2: Optimize the initial text embedding data to obtain each source text embedding data when the optimization conditions are met.
[0105] Among them, the optimization condition refers to the condition for pre-setting and optimizing the initial text embedding data.
[0106] This embodiment provides a 3D model editing method. In this embodiment, using the preset text editor can accurately obtain the initial text embedding data.
[0107] Embodiment 5
[0108] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to determine the target text embedding data corresponding to each target single-view image to be generated according to each target text prompt information, including:
[0109] Input each target text prompt information into a preset text editor, and use the preset text editor to encode each target text prompt information respectively to obtain each target text embedding data.
[0110] In this embodiment, the target text embedding data can be expressed as.
[0111] This embodiment provides a 3D model editing method. In this embodiment, a preset text editor can be used to accurately obtain target text embedding data.
[0112] Embodiment Six
[0113] This embodiment is a further refinement of any of the above embodiments. For each source image, an optional way to achieve non-rigid editing is to use a single-view image editing model that has been trained to convergence and determine the target single-view image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image.
[0114] Figure 4 Schematic diagram of a 3D model editing method provided for Embodiment Six. As Figure 4 shown, it includes:
[0115] S401, determine the optimized target text embedding data based on the source text embedding data and the target text embedding data corresponding to the source image.
[0116] In one way, for each source image, combine the source text embedding data and the target text embedding data corresponding to it to obtain the optimized target text embedding data.
[0117] It should be noted that the target text embedding data characterizes the features corresponding to the source image and also combines the information included in the target text prompt information.
[0118] S402, input the optimized target text embedding data corresponding to each source image into the single-view image editing model that has been trained to convergence, and use the single-view image editing model that has been trained to convergence to determine the target single-view image corresponding to each view respectively to achieve non-rigid editing.
[0119] Specifically, for each source image, assuming there are four source images from different views, namely the front source image, the left-side source image, the right-side source image, and the back source image, then based on the above steps, the optimized target text embedding data corresponding to the four source images can be obtained respectively, namely the first optimized target text embedding data, the second optimized target text embedding data, the third optimized target text embedding data, and the fourth optimized target text embedding data.
[0120] Further, the first optimized target text embedding data is input into the single-view image editing model that has been trained to convergence to output the target single-view image corresponding to the front source image. Exemplarily, it is assumed that the target text prompt information includes the same target object color, size, view information, and target state in the source image, and this target state is inconsistent with the source state in the initial text prompt information. Furthermore, the edited target single-view image is an image in which, under the front view, the color and size of the target object are the same as those in the source image, and the state of the target object changes from the source state to the target state.
[0121] This embodiment provides a three-dimensional model editing method. In this embodiment, the target single-view images output by the single-view image editing model that has been trained to convergence for each source image can be obtained respectively. Since the single-view image editing model has been trained to convergence, the target single-view images are more accurate. In addition, in this embodiment, the optimized target text embedding data is used, and the semantic commonality and semantic difference are included in the optimized target text embedding data. Therefore, this embodiment considers semantic compression, so that in the non-rigid editing process, not only the guidance of the semantic difference part to the non-rigid editing is considered, but also the semantic commonality is used to maintain the original features of the source image, so as to achieve controllable non-rigid editing on the basis of the source image and obtain more accurate target single-view images.
[0122] Embodiment Seven
[0123] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to determine the optimized target text embedding data based on the source text embedding data and the target text embedding data corresponding to the source image.
[0124] Figure 5 It is a schematic diagram of a three-dimensional model editing method provided for Embodiment Seven. As Figure 5 shown, it includes:
[0125] S501, decompose the target text embedding data based on the source text embedding data corresponding to the source image to obtain a projection component and a vertical component. The projection component represents the component of the semantic commonality between the source text embedding data and the target text embedding data, and the vertical component represents the semantic difference component between the source text embedding data and the target text embedding data.
[0126] S502, determine the optimized target text embedding data based on the target text embedding data, the projection component, and the vertical component.
[0127] In this embodiment, the projection component and the vertical component are calculated. Since the projection component can represent semantic commonalities, the semantic commonalities of the source text embedding data and the target text embedding data can be combined. The vertical component can represent semantic differences, and thus the semantic differences between the source text embedding data and the target text embedding data can be combined, so that the determined optimized target text embedding data takes into account semantic commonalities and semantic differences.
[0128] This embodiment provides a three-dimensional model editing method. In this embodiment, the source text embedding data and the target text embedding data are decomposed to calculate the projection component and the vertical component, and then the optimized target text embedding data is calculated. Since the projection component takes into account semantic commonalities and the vertical component takes into account semantic differences, the optimized target text embedding data takes into account semantic commonalities and differences, and thus the optimized target text embedding data is more accurate.
[0129] Embodiment VIII
[0130] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to decompose the target text embedding data based on the source text embedding data corresponding to the source image to obtain the projection component and the vertical component, including:
[0131] Input the source text embedding data and the target text embedding data into a preset text embedding fusion algorithm to output the projection component and the vertical component.
[0132] Exemplarily, the preset text embedding fusion algorithm is as shown in (1):
[0133] (1)
[0134] Wherein, represents the inner product of vectors, represents the projection component, represents the vertical component, represents the source text embedding data, represents the target text embedding data.
[0135] Specifically, input the source text embedding data and the target text embedding data into the preset text embedding fusion algorithm, and the projection component and the vertical component are obtained through the inner product of vectors respectively.
[0136] This embodiment provides a three-dimensional model editing method. In this embodiment, a preset text embedding fusion algorithm is used to calculate the projection component and the vertical component respectively through the inner product of vectors.
[0137] Embodiment IX
[0138] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to determine the optimized target text embedding data based on the target text embedding data, the projection component, and the vertical component, and includes:
[0139] Input the target text embedding data, the projection component, and the vertical component into a preset text embedding optimization algorithm to output the optimized target text embedding data.
[0140] Among them, the preset text embedding optimization algorithm is as shown in (2):
[0141] (2)
[0142] Among them, represents the optimized target text embedding data; is the semantic difference parameter, which is a preset weight coefficient used to control the importance of semantic differences; W is the semantic commonality parameter used to control the importance of semantic commonalities. The larger W is, the smaller the influence of semantic commonalities is.
[0143] In this embodiment, the optimized target text embedding data can be obtained by interpolation.
[0144] This embodiment provides a 3D model editing method. In this embodiment, the preset text embedding optimization algorithm is used to calculate the optimized target text embedding data based on the target text embedding data, the projection component, and the vertical component, so that the optimized target text embedding data is more accurate.
[0145] Embodiment Ten
[0146] This embodiment is a further refinement of any of the above embodiments. In this embodiment, the single-view image editing model trained to convergence is a model with a low-rank adaptation layer added to the key vector and value vector of the cross-attention module of the single-view image generation model trained to convergence.
[0147] Among them, the single-view image generation model trained to convergence is used to perform preparatory recognition on the source image, and can read out the original features in the source image. For example, the target object type, the target object color, the target object size, and the source state of the target object, and generate an edited image to-be-optimized image based on the original features of the source image. The edited image to-be-optimized image is similar to the features in the source image. Since it is a single-view image generation model trained to convergence, it is obtained after training, so that the edited image is similar to the features in the source image.
[0148] This embodiment is an optional method before non-rigid editing for each source image, which uses a single-view image editing model that has been trained to convergence and determines the target single-view image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image.
[0149] Figure 6 It is a schematic diagram of a 3D model editing method provided for Embodiment Ten. As Figure 6 shown, it includes:
[0150] S601, obtain a single-view image generation model that has been trained to convergence.
[0151] Among them, the single-view image generation model that has been trained to convergence can be obtained by further controlling the input and its own training parameters on the basis of a pre-trained and converged basic diffusion model.
[0152] S602, keep the single-view image generation model that has been trained to convergence unchanged, and train the training parameters of the low-rank adaptation layer until the first convergence condition is reached.
[0153] It should be noted that in this embodiment, a low-rank adaptation layer is added to the key vector and value vector on the basis of the single-view generation model that has been trained to convergence. Therefore, only the training parameters in the low-rank adaptation layer need to be adjusted, and the parameters in the single-view image generation model that has been trained to convergence are not trained, so as to keep the single-view image generation model that has been trained to convergence unchanged. Input the training images in the training samples into the low-rank adaptation layer, and adjust the training parameters in the low-rank adaptation layer until the edited images output from the single-view image editing model are consistent with the actual images in the training samples, then it is determined that the first convergence condition is reached. Among them, the training samples include training images and their corresponding actual images. The actual image refers to the image of a certain state formed after the training object performs a training action based on the training features of the training image. The training object in the training image is in the previous state, and the training object in the actual image is in the subsequent state.
[0154] Among them, the first convergence condition refers to the convergence condition that the training parameters of the adjusted low-rank adaptation layer make the edited image reach a certain set consistency with the actual image in the training sample.
[0155] S603, determine the single-view image editing model that reaches the first convergence condition as the single-view image editing model that has been trained to convergence.
[0156] It should be noted that when the first convergence condition is reached, at this time, the training parameters of the low-rank adaptation layer have been adjusted and adjusted to the target parameters. Determine the single-view image editing model corresponding to the target parameters at this time as the single-view image editing model that has been trained to convergence.
[0157] This embodiment provides a 3D model editing method. In this embodiment, during the training process, the single-view image generation model that has been trained to convergence remains unchanged, and only the training parameters of the low-rank adaptive layer are adjusted, thereby reducing the number of optimization parameters and accelerating the network training speed.
[0158] Example XI
[0159] This embodiment is a further refinement of any of the above embodiments. In this embodiment, the single-view image generation model that has been trained to convergence is a diffusion model, which is an optional way of the training process of the single-view image generation model that has been trained to convergence.
[0160] Figure 7 It is a schematic diagram of a 3D model editing method provided for Example XI. As Figure 7 shown, it includes:
[0161] S701, obtaining an initial single-view image generation model.
[0162] Among them, the initial single-view image generation model can be a basic diffusion model that has been trained to convergence.
[0163] S702, keeping the source text embedding data corresponding to the source image unchanged, inputting the source text embedding data into the initial single-view image generation model, and training the training parameters of the initial single-view image generation model until the second convergence condition is reached.
[0164] Exemplarily, the training process of the single-view image generation model can be as shown in (3):
[0165] (3)
[0166] Among them, is the input image with noise, that is, the image formed by adding noise to the source image; is the noise ground truth, represents the single-view image generation model, represents the training parameter weight of the single-view image generation model, is the source text embedding data, t is the time step of the single-view image generation model, refers to the diffusion loss value during the training of the single-view image generation model.
[0167] In this step, keeping the input source text embedding data unchanged, inputting it into (3), and training the training parameters of the initial single-view image generation model until the training parameters adjusted make the image to be optimized output by the single-view image generation model satisfy the second convergence condition with the source image.
[0168] Among them, the second convergence condition means that the adjusted training parameters make the consistency between the to-be-optimized image output by the single-view image generation model and the source image greater than or equal to the first preset consistency degree. Exemplarily, assuming that the first preset consistency degree is 95%, then the consistency between the to-be-optimized image and the source image is greater than or equal to 95%, that is, the second convergence condition is reached.
[0169] It should be noted that, to ensure accuracy, the first preset consistency degree can be set larger. For example, it can be set to 98% or 100%, and there is no limit here.
[0170] S703, determine the single-view image generation model that reaches the second convergence condition as the single-view image generation model trained to convergence.
[0171] Exemplarily, if, after the training parameters of the single-view image generation model are adjusted, the second convergence condition is reached, then the single-view image generation model corresponding to the adjusted training parameters at this time is determined as the single-view image generation model trained to convergence.
[0172] This embodiment provides a 3D model editing method. In this embodiment, first, an initial single-view image generation model is obtained. Then, while keeping the source text embedding data unchanged, the source text embedding data is input into the initial single-view image generation model, and the training parameters therein are trained until the second convergence condition is reached, thereby determining the single-view image generation model trained to convergence. This application uses the source image to train the single-view image generation model, which can improve generalization. In addition, in this embodiment, the single-view image generation model is trained with the source image to maintain the original features of the source image.
[0173] Embodiment Twelve
[0174] This embodiment is a further refinement of any of the above embodiments. This embodiment is an optional way to optimize the initial text embedding data to obtain each source text embedding data that meets the optimization conditions.
[0175] Figure 8 It is a schematic diagram of a 3D model editing method provided for Embodiment Twelve. As Figure 8 shown, it includes:
[0176] S801, input the initial text embedding data into the initial single-view image generation model to output the to-be-optimized image.
[0177] Exemplarily, the process of optimizing the initial text embedding data is as shown in (4):
[0178] (4)
[0179] Among them, is the input image with noise, that is, the image formed by adding noise to the source image; is the ground truth of the noise, represents the single-view image generation model, represents the training parameter weights of the single-view image generation model, represents the text embedding data, and t is the time step of the single-view image generation model, represents the diffusion loss value under the initial single-view image generation model.
[0180] It should be noted that in step S801, is the initial single-view image generation model, including the initial text embedding data and the new text embedding data after adjusting the initial text embedding data.
[0181] In this step, the initial text embedding data is input into (4), and the initial single-view image generation model is used to output the image to be optimized.
[0182] S802. In response to the image to be optimized meeting the optimization condition, determine the initial text embedding data as the source text embedding data.
[0183] Among them, the optimization condition can be the condition that the consistency between the image to be optimized and the source image is greater than or equal to the second preset consistency. Among them, the second preset consistency is less than the first preset consistency. Exemplarily, assuming that the first preset consistency is 98%, the second preset consistency can be set to 95%.
[0184] Exemplarily, assuming that the consistency between the image to be optimized and the source image is 96%, it is determined that the image to be optimized meets the optimization condition, and then the initial text embedding data is determined as the source text embedding data.
[0185] It should be noted that the exemplary examples in this application are only used to illustrate the solution and do not represent the actual situation.
[0186] S803. In response to the image to be optimized not meeting the optimization condition, continue to adjust the initial text embedding data and keep the initial single-view image generation model unchanged until the output image to be optimized meets the optimization condition, and determine the text embedding data corresponding to when the optimization condition is met as the source text embedding data.
[0187] Exemplarily, if the consistency between the image to be optimized and the source image is 80%, and it is determined that it is less than the second preset consistency (95%), it is determined that the optimization condition is not met.
[0188] Further, continue to adjust the initial text embedding data , so as to obtain the new text embedding data , and maintain the initial single-view image generation model, that is, keep the training parameters therein unchanged, and then input the new text embedding data into (4) to output a new image to be optimized. Continue to determine whether the image to be optimized meets the optimization conditions. If not, continue to adjust the initial text embedding data until the adjusted text embedding data makes the image to be optimized output by the initial single-view image generation model meet the optimization conditions. Determine the text embedding data corresponding to the optimization conditions at this time as the source text embedding data.
[0189] This embodiment provides a three-dimensional model editing method. In this embodiment, the initial single-view image generation model remains unchanged, and the initial text embedding data is adjusted so that the output image to be optimized meets the optimization conditions. Then, the text embedding data corresponding to the optimization conditions is determined as the source text embedding data. Thus, a text embedding optimization strategy is proposed in this embodiment. By optimizing the initial text embedding data, accurate source text embedding data can be obtained, and the source text embedding data can more accurately describe the source image.
[0190] Embodiment Thirteen
[0191] Figure 9 is a schematic diagram of the overall process of a three-dimensional model editing method provided for Embodiment Thirteen. As Figure 9 shown, it includes: a source three-dimensional model 901, a source image 902, a single-view image editing model 903 that has been trained to convergence, a target single-view image 904, and a target three-dimensional model 905.
[0192] In Figure 9 , the target object in the source three-dimensional model 901 is a cat, and the target objects have their respective characteristics. A front source image 902 is rendered from the source three-dimensional model 901, and the source image 902 is input into the single-view image editing model 903 that has been trained to convergence, and the target single-view image 904 in the source image view corresponding to the front source image 902 is output based on the target text prompt information.
[0193] It should be noted that the source images 902 from other viewpoints can continue to be obtained from the source three-dimensional model 901, and the single-view image editing model 903 that has been trained to convergence is used to output the target single-view images 904 corresponding to the corresponding viewpoints.
[0194] Furthermore, the target three-dimensional model is reconstructed based on multiple target single-view images 904.
[0195] It should be noted that the single-view image generation model trained to convergence in this application is trained with source images. There can be multiple source images. The single-view image generation model can be trained each time a source image is obtained until it converges to obtain the single-view image generation model trained to convergence. Alternatively, a single source image can be used to obtain the single-view image generation model trained to convergence, and for the remaining source images, the target single-view images can be output by using the single-view image generation model trained to convergence obtained from the above-mentioned single source image.
[0196] It should be noted that text embedding data fusion is achieved in this application, that is, the source text embedding data and the target text embedding data are fused to obtain the optimized target embedding data. In addition, in this application, the source image is used to train the single-view image generation model until it converges to obtain the single-view image generation model trained to convergence. Furthermore, a low-rank adaptation layer is added based on the single-view image generation model trained to convergence, and only the training parameters of the low-rank adaptation layer are trained to obtain the single-view image editing model trained to convergence. This single-view image editing model trained to convergence can edit the source image into the target single-view image, thereby completing non-rigid editing. In addition, the target state in the target single-view image can be a local state change or a global state change, so that global or local non-rigid editing can be realized.
[0197] It should be noted that in this application, based on the optimization of text embedding data, feature extraction is performed on the source image to obtain the source text embedding data, where the source text embedding data includes the original characteristics of the source image.
[0198] It should be noted that in this application, the optimized target text embedding data is obtained based on the source text embedding data and the target text embedding data, so that the target single-view image can complete the change from the source state to the target state while keeping the original features of the source image unchanged, and complete non-rigid editing.
[0199] It should be noted that since the single-view image editing model refers to adding a low-rank adaptation layer based on the single-view image generation model trained to convergence, when training the single-view image editing model in this application, only the low-rank adaptation layer needs to be trained, which can thus accelerate the training speed.
[0200] It should be noted that in this application, a semantic compression-based text embedding fusion strategy is proposed. The source text embedding data and the target text embedding data are decomposed to obtain the components of semantic commonality and the components of semantic difference, so that during the non-rigid editing process, the guidance of semantic difference to non-rigid editing is considered, and at the same time, the original features of the source image are maintained by using semantic commonality, thereby realizing non-rigid editing.
[0201] In summary, the present application proposes text-driven non-rigid editing of the source three-dimensional model, thereby achieving global or local three-dimensional non-rigid editing, and thus providing accuracy, consistency, and controllability of the editing result.
[0202] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0203] Embodiment Fourteen
[0204] The following is an embodiment of the apparatus of the present application. Figure 10 It is a schematic structural diagram of a three-dimensional model editing apparatus provided for Embodiment Fourteen. As Figure 10 shown, the three-dimensional model editing apparatus 1000 includes the following modules:
[0205] An acquisition module 1001, configured to acquire multiple source images from different perspectives in the source three-dimensional model, where the source images are images corresponding to the three-dimensional model from different perspectives; each source image includes a target object; the target object is in a source state;
[0206] A determination module 1002, configured to determine the initial text prompt information corresponding to each source image and the target text prompt information corresponding to each target single-perspective image to be generated. The initial text prompt information includes the source state information of the target object, and the target text prompt information includes the target state information of the target object;
[0207] The determination module 1002 is further configured to determine the source text embedding data corresponding to each source image according to each initial text prompt information; the determination module 1002 is further configured to determine the target text embedding data corresponding to each target single-perspective image to be generated according to each target text prompt information. Each target single-perspective image is an image in which the target object performs a target action in the perspective of the corresponding source image and forms a target state from the source state;
[0208] The determination module 1002 is further configured to, for each source image, adopt a single-perspective image editing model that has been trained to convergence and determine the target single-perspective image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image, so as to achieve non-rigid editing;
[0209] A reconstruction module 1003, configured to reconstruct the target three-dimensional model based on each target single-perspective image.
[0210] Optionally, when the determination module 1002 determines the initial text prompt information corresponding to each source image, it is specifically configured to:
[0211] Input each source image into a preset text prediction algorithm, and use the preset text prediction algorithm to determine the image description information of the source image respectively;
[0212] Determine the image description information of each source image as each initial text prompt information.
[0213] Determination module 1002, when determining the target text prompt information corresponding to each target single-view image to be generated, specifically used for:
[0214] Generate target text prompt information according to each initial text prompt information and the target state.
[0215] Determination module 1002, when determining the source text embedding data corresponding to each source image according to each initial text prompt information, specifically used for:
[0216] Input each initial text prompt information into a preset text editor, and use the preset text editor to encode each initial text prompt information respectively to obtain each initial text embedding data;
[0217] Optimize the initial text embedding data to obtain each source text embedding data when meeting the optimization conditions.
[0218] Determination module 1002, when determining the target text embedding data corresponding to each target single-view image to be generated according to each target text prompt information, specifically used for:
[0219] Input each target text prompt information into a preset text editor, and use the preset text editor to encode each target text prompt information respectively to obtain each target text embedding data.
[0220] Determination module 1002, for each source image, use a single-view image editing model that has been trained to convergence and based on the source text embedding data and target text embedding data corresponding to the source image to determine the target single-view image corresponding to the source image, so as to achieve non-rigid editing, specifically used for:
[0221] Determine the optimized target text embedding data based on the source text embedding data and target text embedding data corresponding to the source image;
[0222] Input the optimized target text embedding data corresponding to each source image into a single-view image editing model that has been trained to convergence respectively, and use the single-view image editing model that has been trained to convergence to determine the target single-view image corresponding to each view respectively, so as to achieve non-rigid editing.
[0223] Determination module 1002, when determining the optimized target text embedding data based on the source text embedding data and target text embedding data corresponding to the source image, specifically used for:
[0224] Decompose the target text embedding data based on the source text embedding data corresponding to the source image to obtain a projection component and a vertical component. The projection component represents the component of the semantic commonality between the source text embedding data and the target text embedding data, and the vertical component represents the component of the semantic difference between the source text embedding data and the target text embedding data;
[0225] Determine the optimized target text embedding data based on the target text embedding data, the projection component, and the vertical component.
[0226] Determination module 1002, when decomposing the target text embedding data based on the source text embedding data corresponding to the source image to obtain a projection component and a vertical component, is specifically used for:
[0227] Input the source text embedding data and the target text embedding data into a preset text embedding fusion algorithm to output the projection component and the vertical component.
[0228] Determination module 1002, when determining the optimized target text embedding data based on the target text embedding data, the projection component, and the vertical component, is specifically used for:
[0229] Input the target text embedding data, the projection component, and the vertical component into a preset text embedding optimization algorithm to output the optimized target text embedding data.
[0230] The single-view image editing model trained to convergence is a model that adds a low-rank adaptation layer to the key vector and value vector of the cross-attention module of the single-view image generation model trained to convergence;
[0231] Determination module 1002, before, for each source image, using the single-view image editing model trained to convergence and based on the source text embedding data and the target text embedding data corresponding to the source image to determine the target single-view image corresponding to the source image to achieve non-rigid editing, is specifically further used for:
[0232] Obtain the single-view image generation model trained to convergence;
[0233] Keep the single-view image generation model trained to convergence unchanged, and train the training parameters of the low-rank adaptation layer until the first convergence condition is reached;
[0234] Determine the single-view image editing model that reaches the first convergence condition as the single-view image editing model trained to convergence.
[0235] The single-view image generation model trained to convergence is a diffusion model. This embodiment provides a three-dimensional model editing device, further including: a training module;
[0236] The training module, during the training process of the single-view image generation model that has been trained to convergence, is specifically configured to:
[0237] Obtain an initial single-view image generation model;
[0238] Keep the source text embedding data corresponding to the source image unchanged, input the source text embedding data into the initial single-view image generation model, and train the training parameters of the initial single-view image generation model until the second convergence condition is reached;
[0239] Determine the single-view image generation model that reaches the second convergence condition as the single-view image generation model that has been trained to convergence.
[0240] The determination module 1002, when optimizing the initial text embedding data to obtain each source text embedding data that meets the optimization conditions, is specifically configured to:
[0241] Input the initial text embedding data into the initial single-view image generation model to output an image to be optimized;
[0242] In response to the image to be optimized meeting the optimization conditions, determine the initial text embedding data as the source text embedding data;
[0243] In response to the image to be optimized not meeting the optimization conditions, continue to adjust the initial text embedding data and keep the initial single-view image generation model unchanged until the output image to be optimized meets the optimization conditions, and determine the text embedding data corresponding to when the optimization conditions are met as the source text embedding data.
[0244] For the description of the features in the corresponding embodiment of the three-dimensional model editing device, reference can be made to the relevant description in the corresponding embodiment of the three-dimensional model editing method, which will not be elaborated here one by one.
[0245] The embodiment of the present application also provides an editing device, Figure 11 It is a schematic structural diagram of the editing device provided in Embodiment Fourteen. As Figure 11 shown, the editing device 1100 includes a processor 1101 and a memory 1102. Among them, the processor 1101, the memory 1102, and the communication component 1103 are connected through a bus 1104. The memory 1102 stores a computer program, and the processor 1101 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the three-dimensional model editing method.
[0246] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is configured to execute the steps in any of the above-mentioned embodiments of the three-dimensional model editing method when running.
[0247] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.
[0248] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the three-dimensional model editing method are implemented.
[0249] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the three-dimensional model editing method are implemented.
[0250] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0251] The above provides a detailed introduction to a three-dimensional model editing method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A three-dimensional model editing method, characterized in that, Including: Obtain multiple source images from different perspectives of a source 3D model, where the source images are the images corresponding to the 3D model from different perspectives; the target objects in each source image are in the same source state, and this source state is the same as the state of the target object in the source 3D model; Input each of the source images into a preset text prediction algorithm respectively, and use the preset text prediction algorithm to determine the image description information of the source images respectively; Determine the image description information of each of the source images as the initial text prompt information corresponding to each of the source images, and the initial text prompt information includes the source state information of the target object; Generate the target text prompt information corresponding to each target single - perspective image to be generated according to each of the initial text prompt information and the target state, and the target text prompt information includes the target state information of the target object; determine the source text embedding data corresponding to each of the source images according to each of the initial text prompt information, and determine the target text embedding data corresponding to each target single - perspective image to be generated according to each of the target text prompt information. Each of the target single - perspective images is an image in which the target object forms a target state from the source state after performing a target action at the perspective of the corresponding source image; For each source image, use a single - perspective image editing model that has been trained to convergence and determine the target single - perspective image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image, so as to achieve non - rigid editing; Reconstruct a target 3D model based on each target single - perspective image.
2. The method according to claim 1, characterized in that, The determining the source text embedding data corresponding to each of the source images according to each of the initial text prompt information includes: Input each of the initial text prompt information into a preset text editor, and use the preset text editor to encode each of the initial text prompt information respectively to obtain each initial text embedding data; Optimize the initial text embedding data to obtain each source text embedding data when the optimization conditions are met.
3. The method according to claim 1, characterized in that, The determining the target text embedding data corresponding to each target single - perspective image to be generated according to each of the target text prompt information includes: Input each of the target text prompt information into a preset text editor, and use the preset text editor to encode each of the target text prompt information respectively to obtain each target text embedding data.
4. The method according to claim 1, wherein The for each source image, using a single - perspective image editing model that has been trained to convergence and determining the target single - perspective image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image to achieve non - rigid editing includes: Determine the optimized target text embedding data based on the source text embedding data and the target text embedding data corresponding to the source image; Input the optimized target text embedding data corresponding to each source image into the single - perspective image editing model that has been trained to convergence respectively, and use the single - perspective image editing model that has been trained to convergence to determine the target single - perspective image corresponding to each perspective respectively to achieve non - rigid editing.
5. The method according to claim 4, wherein The determining the optimized target text embedding data based on the source text embedding data and the target text embedding data corresponding to the source image includes: Decompose the target text embedding data based on the source text embedding data corresponding to the source image to obtain a projection component and a vertical component, where the projection component represents the component of the semantic commonality between the source text embedding data and the target text embedding data, and the vertical component represents the component of the semantic difference between the source text embedding data and the target text embedding data; Determine the optimized target text embedding data based on the target text embedding data, the projection component, and the vertical component.
6. The method according to claim 5, characterized in that The decomposing the target text embedding data based on the source text embedding data corresponding to the source image to obtain a projection component and a vertical component includes: Input the source text embedding data and the target text embedding data into a preset text embedding fusion algorithm to output the projection component and the vertical component.
7. The method according to claim 5, wherein The determining the optimized target text embedding data based on the target text embedding data, the projection component, and the vertical component includes: Input the target text embedding data, the projection component, and the vertical component into a preset text embedding optimization algorithm to output the optimized target text embedding data.
8. The method according to claim 1, wherein The single-view image editing model trained to convergence is a model with a low-rank adaptive layer added to the key vector and value vector of the cross-attention module of the single-view image generation model trained to convergence; Before, for each source image, using the single-view image editing model trained to convergence and determining the target single-view image corresponding to the source image based on the source text embedding data and the target text embedding data corresponding to the source image, further includes: Obtain a single-view image generation model trained to convergence; Keep the single-view image generation model trained to convergence unchanged, and train the training parameters of the low-rank adaptive layer until the first convergence condition is reached; Determine the single-view image editing model that reaches the first convergence condition as the single-view image editing model trained to convergence.
9. The method according to claim 8, wherein The single-view image generation model trained to convergence is a diffusion model, and the training process of the single-view image generation model trained to convergence includes: Obtain an initial single-view image generation model; Keep the source text embedding data corresponding to the source image unchanged, input the source text embedding data into the initial single-view image generation model, and train the training parameters of the initial single-view image generation model until the second convergence condition is reached; Determine the single-view image generation model that reaches the second convergence condition as the single-view image generation model trained to convergence.
10. The method according to claim 9, characterized in that, The optimizing the initial text embedding data to obtain each source text embedding data that meets the optimization condition includes: Input the initial text embedding data into the initial single-view image generation model to output an image to be optimized; In response to the image to be optimized meeting the optimization condition, determine the initial text embedding data as the source text embedding data; In response to the image to be optimized not meeting the optimization conditions, continue to adjust the initial text embedding data and keep the initial single-view image generation model unchanged until the output image to be optimized meets the optimization conditions, and determine the text embedding data corresponding to when the optimization conditions are met as the source text embedding data.
11. An editing device, characterized in that, Comprising: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, such that the processor executes the method according to any one of claims 1-10.
12. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1-10.
13. A computer program product, characterized in that, Comprising a computer program, which when executed by a processor implements the method according to any one of claims 1-10.
Citation Information
Patent Citations
Image editing method and device, equipment, storage medium and program product
CN117611709A
Three-dimensional model data processing method and system, product, equipment and medium
CN118887348A