Three-dimensional model generation method and device, equipment, storage medium and program product
By generating viewpoint-aligned reference 3D models through text editing prompts and multi-view rendering, the problem of low efficiency and poor consistency in existing 3D model editing is solved, achieving highly efficient and automated 3D model editing and reducing manpower and time costs.
Patent Information
- Application Number
- CN202511562173.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-02-24
AI Technical Summary
The existing 3D model editing is inefficient, has poor consistency, and incurs high manpower and time costs.
A reference 3D model with viewpoint alignment is generated using text editing prompts. The target 3D model is then generated through multi-view rendering and stitching of contour image sets, combined with text editing and contour prompt image set processing.
It improves model editing efficiency, reduces manpower and time costs, enhances model consistency, and achieves automated model editing.
Smart Images

Figure CN121564191A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, storage medium, and program product for generating three-dimensional models. Background Technology
[0002] With the continuous advancement of computer graphics, artificial intelligence, and computing hardware technologies, 3D modeling has become an indispensable core technology in cutting-edge fields such as virtual reality, augmented reality, film and television special effects, game development, industrial design, and digital twins. High-fidelity, interactive 3D content not only enhances the realism and immersion of the user experience but also promotes the intelligent and automated process of digital content creation. Driven by the rise of the metaverse concept and the AIGC (AI-generated content) wave, the demand for the generation and editing of high-quality 3D assets has exploded, prompting academia and industry to continuously explore more efficient and intelligent methods for 3D content creation.
[0003] Current mainstream 3D model editing techniques mainly rely on professional software (such as Maya, Blender, ZBrush, etc.) to modify the model's geometry, texture mapping, and pose by manually adjusting vertices, control points, bones, or material parameters.
[0004] Currently, 3D model editing is inefficient, has poor model consistency, and incurs high labor and time costs. Summary of the Invention
[0005] This disclosure provides a method, apparatus, device, storage medium, and program product for generating three-dimensional models, in order to at least solve the problems of low efficiency, poor model consistency, and high labor and time costs in existing three-dimensional model editing.
[0006] The technical solution disclosed herein is as follows: This disclosure provides a method for generating a three-dimensional model, including: Using text editing prompts, a reference 3D model is generated, wherein the reference 3D model is a model that is viewpoint aligned with the original 3D model; Render a first two-dimensional image of the original three-dimensional model and a second two-dimensional image of the reference three-dimensional model from multiple perspectives; Based on the text editing prompts, a first set of contour images is extracted from the first two-dimensional image and a second set of contour images is extracted from the second two-dimensional image; The first contour image set and the second contour image set are stitched together to obtain a contour prompt image set; The original 3D model is processed based on the text editing prompts and the set of outline prompt images to obtain the target 3D model.
[0007] Optionally, extracting a first set of contour images from the first two-dimensional image and extracting a second set of contour images from the second two-dimensional image according to the text editing prompt includes: Contour extraction is performed on each of the first two-dimensional images to obtain a first contour set; and Contour extraction is performed on each of the second two-dimensional images to obtain a second contour set; Based on the text editing prompts, a first region mask is input from the first contour set; contours outside the first region mask are removed based on the first region mask to obtain a first contour image set of the unedited region; and Based on the text editing prompt, input the second region mask of the editing prompt from the first contour set; remove the contours outside the second region mask based on the second region mask to obtain the second contour image set of the editing region.
[0008] Optionally, the step of concatenating the first contour image set and the second contour image set to obtain a contour hint image set includes: The first contour image set and the second contour image set are stitched together to obtain a stitched contour set; The spliced contour set is processed to obtain a contour hint image set.
[0009] Optionally, processing the stitched contour set to obtain a contour hint image set includes: The spliced contour set is fused to obtain a fused spliced contour set; The fused spliced contour set is smoothed to obtain a smoothed spliced contour set. The smoothed and stitched contour set is processed using an edge enhancement algorithm or a contour optimization algorithm to obtain the contour prompt image set.
[0010] Optionally, the edge enhancement algorithm is Canny edge detection, and the contour optimization algorithm is Bezier curve fitting. The step of processing the smoothed, stitched contour set using either the edge enhancement algorithm or the contour optimization algorithm to obtain the contour hint image set includes: The smoothed spliced contour set is processed using the Canny edge detection or the Bezier curve fitting to obtain the contour prompt image set.
[0011] Optionally, the step of processing the original 3D model based on the text editing prompts and the contour prompt image set to obtain the target 3D model includes: The first loss function is determined based on the time step, the time step weight function, the input prompts, and the noisy image. The second loss function is determined based on the 3D mask of the unedited region, the parameters of the original 3D model, and the image features; Determine the total loss function based on the first loss function and the second loss function; Based on the total loss function, a cue adapter based on decoupled cross-attention is used to input the set of text editing cue and the set of contour cue images into each cross-attention module of the diffusion model for model iteration, so as to obtain the target 3D model.
[0012] This disclosure also provides a three-dimensional model generation apparatus, including: The generation module is used to generate a reference 3D model using text editing prompts, wherein the reference 3D model is a model that is viewpoint aligned with the original 3D model; The rendering module is used to render a first two-dimensional image of the original three-dimensional model and a second two-dimensional image of the reference three-dimensional model from multiple perspectives; The extraction module is used to extract a first set of contour images from the first two-dimensional image and a second set of contour images from the second two-dimensional image according to the text editing prompts. The stitching module is used to stitch the first contour image set and the second contour image set together to obtain a contour prompt image set; The processing module is used to process the original 3D model according to the text editing prompts and the outline prompt image set to obtain the target 3D model.
[0013] This disclosure also provides an electronic device, including: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the steps in the above method.
[0014] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0015] This disclosure also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method described above.
[0016] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: In some embodiments of this disclosure, a reference 3D model is generated using text editing prompts. This reference 3D model is a model with a viewpoint aligned with the original 3D model. A 3D model under the target semantics is generated based on user natural language instructions, reducing manual operations and improving model editing efficiency. A first 2D image of the original 3D model and a second 2D image of the reference 3D model are rendered from multiple viewpoints. A first contour image set is extracted from the first 2D image, and a second contour image set is extracted from the second 2D image, based on the text editing prompts. The first and second contour image sets are then stitched together to obtain a contour prompt image set. The original 3D model is processed based on the text editing prompts and the contour prompt image set to obtain the target 3D model. This disclosure effectively constrains the differences between the edited model and the original 3D model, improving model consistency, automatically performing model editing, increasing model editing efficiency, and reducing labor and time costs.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0019] Figure 1 A flowchart illustrating a three-dimensional model generation method provided as an exemplary embodiment of this disclosure; Figure 2 A schematic diagram of the structure of a three-dimensional model generation apparatus provided for an exemplary embodiment of this disclosure; Figure 3 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this disclosure. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0021] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure.
[0022] It should be noted that the user information involved in this disclosure includes, but is not limited to, user device information and user personal information; the collection, storage, use, processing, transmission, provision and disclosure of user information in this disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0023] To address the aforementioned technical problems, in some embodiments of this disclosure, a reference 3D model is generated using text editing prompts. This reference 3D model is a model with a viewpoint aligned with the original 3D model. Based on user natural language instructions, a 3D model under the target semantics is generated, reducing manual operations and improving model editing efficiency. A first 2D image of the original 3D model and a second 2D image of the reference 3D model are rendered from multiple viewpoints. Based on the text editing prompts, a first contour image set is extracted from the first 2D image, and a second contour image set is extracted from the second 2D image. The first and second contour image sets are then stitched together to obtain a contour prompt image set. The original 3D model is then processed based on the text editing prompts and the contour prompt image set to obtain the target 3D model. This disclosure effectively constrains the differences between the edited model and the original 3D model, improving model consistency, automatically performing model editing, increasing model editing efficiency, and reducing labor and time costs.
[0024] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0025] Figure 1 This is a flowchart illustrating a three-dimensional model generation method provided as an exemplary embodiment of this disclosure. Figure 1 As shown, the method includes: S101: Using text editing prompts, generate a reference 3D model, which is a model that is viewpoint aligned with the original 3D model; S102: Render the first two-dimensional image of the original 3D model and the second two-dimensional image of the reference 3D model from multiple perspectives; S103: Based on the text editing prompts, extract a first set of contour images from the first two-dimensional image and extract a second set of contour images from the second two-dimensional image; S104: The first contour image set and the second contour image set are stitched together to obtain the contour prompt image set; S105: Process the original 3D model based on the text editing prompts and the outline prompt image set to obtain the target 3D model.
[0026] In this embodiment, the entity executing the above method can be a terminal device or a server.
[0027] The terminal device includes, but is not limited to, mobile stations (MS), mobile terminals, mobile phones, handsets, and portable equipment. This terminal device can communicate with one or more core networks via a radio access network (RAN). For example, the terminal device can be a mobile phone (or "cellular" phone), a computer with wireless communication capabilities, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an AR terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical care, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The operating systems installed on the terminal device include, but are not limited to, iOS, Android, Windows, Linux, and Mac OS. In different networks, terminals may be called by different names, such as: user equipment, mobile station, user unit, station, cellular phone, personal digital assistant, wireless modem, wireless communication device, handheld device, laptop, cordless phone, wireless local loop station, television, etc. For ease of description, this embodiment will simply refer to it as terminal device.
[0028] In this embodiment, the implementation form of the server is not limited. For example, the server can be a conventional server, a cloud server, a cloud host, a virtual center, or other server devices. The server mainly consists of a processor, hard disk, memory, system bus, and other common computer architecture types.
[0029] In this embodiment, a reference 3D model is generated using text editing prompts. This reference 3D model is a model with a viewpoint aligned with the original 3D model. Based on user natural language instructions, a 3D model with the target semantics is generated, reducing manual operations and improving model editing efficiency. The first 2D image of the original 3D model and the second 2D image of the reference 3D model are rendered from multiple viewpoints. Based on the text editing prompts, a first contour image set is extracted from the first 2D image, and a second contour image set is extracted from the second 2D image. The first and second contour image sets are then stitched together to obtain a contour prompt image set. The original 3D model is then processed based on the text editing prompts and the contour prompt image set to obtain the target 3D model. This disclosure effectively constrains the differences between the edited model and the original 3D model, improving model consistency, automatically performing model editing, increasing model editing efficiency, and reducing labor and time costs.
[0030] It should be noted that the 3D model uses a Nerf (Neural RadianceFields) model with a differentiable renderer or a 3D GS (3D Gaussian Splatting) model to facilitate subsequent fractional distillation editing and iterative optimization.
[0031] In some embodiments of this disclosure, a reference 3D model is generated using text editing prompts. One possible approach is to first receive a text editing prompt from the user, such as: generate a person in a seated posture. Based on the text editing prompt, a pre-trained text-to-3D model generation system (such as DreamFusion, Text2Mesh, or a similar diffusion model-driven generation framework) is invoked to generate multiple candidate 3D models that conform to the semantic description. These candidate models embody the semantic features described in the text in terms of posture, shape, or structure. Subsequently, the system presents these generated results to the user for interactive selection. The user selects the model that best matches their editing intention as the reference 3D model. This reference 3D model semantically represents the target state that the original 3D model needs to be transformed into, providing geometric priors for subsequent multi-view contour extraction and structure-guided editing. By introducing a text-to-3D generation model to generate a reference 3D model and selecting the reference 3D model that best matches the editing intention, embodiments of this disclosure achieve semantic-level control and structure-level guidance of the editing direction.
[0032] For example, consider a 3D character model D1, which is standing and needs to be edited to a sitting position. First, the user inputs a text editing prompt: "Generate a sitting person." Using existing text-to-3D generation models (such as DreamFusion and Text2Mesh), multiple sitting character models with different styles or postures are automatically generated, such as cross-legged sitting, sitting in a chair, and leaning forward. After reviewing the generated results, the user selects one of the naturally sitting character models as a reference 3D model D2.
[0033] In some embodiments of this disclosure, the original 3D model and the generated reference 3D model are aligned in terms of viewpoint. For example, the front view of the original 3D model is aligned with that of the reference 3D model. During the editing of the reference 3D model, the original 3D model and the generated reference 3D model may have different poses, positions, or scales. To ensure that subsequent editing operations can be performed accurately, the two models must first be aligned in terms of viewpoint. The purpose of viewpoint alignment is to make the two models present a consistent pose and position under the same camera viewpoint, facilitating subsequent comparison and editing.
[0034] It should be noted that multiple perspectives include, but are not limited to: front, left side, right side, back, top, and bottom.
[0035] In some embodiments of this disclosure, a first two-dimensional image of the original 3D model and a second two-dimensional image of the reference 3D model are rendered from multiple perspectives. First, a set of camera viewpoints evenly distributed in 3D space (e.g., front, left, right, back, top, and bottom, a total of six viewpoints, or multiple viewpoint points sampled on a sphere) are set to ensure comprehensive coverage of all surfaces of the model. Then, a computer graphics rendering engine (such as OpenGL, Blender, PyTorch3D, etc.) is used to render the original 3D model and the reference 3D model from these same viewpoints. For the original 3D model D1, a two-dimensional image is generated at each viewpoint, forming the first two-dimensional image set; similarly, for the reference 3D model D2, a corresponding two-dimensional image is generated at the same viewpoint, forming the second two-dimensional image set. These images may contain RGB color information, or may be rendered as intermediate representations such as depth maps, normal maps, or silhouette maps as needed.
[0036] In some embodiments of this disclosure, a first contour image set is extracted from a first two-dimensional image and a second contour image set is extracted from a second two-dimensional image based on text editing prompts. One possible implementation involves extracting contours from each first two-dimensional image to obtain a first contour set; and extracting contours from each second two-dimensional image to obtain a second contour set; inputting a first region mask from the first contour set based on text editing prompts; removing contours outside the first region mask based on the first region mask to obtain a first contour image set of the unedited region; and inputting a second region mask from the first contour set based on text editing prompts; removing contours outside the second region mask based on the second region mask to obtain a second contour image set of the edited region. Specifically, for each image in the first two-dimensional image set, an edge detection algorithm (such as Canny edge detection, Sobel operator, etc.) or more complex computer vision techniques (such as deep learning-based instance segmentation methods) is applied to extract the boundaries of objects in the image, resulting in a series of contour images constituting the first contour set. Similarly, the same operation is performed on the second two-dimensional image set to generate the second contour set. Based on the provided text editing prompts, determine the first region to be edited and create a corresponding first region mask based on the first contour set. This mask identifies the areas the user wants to retain or pay special attention to; it can be viewed as a binary image, where pixel values within the region of interest are set to 1, and pixel values within the background or other areas are set to 0. Similarly, define a second region mask based on the text editing prompts, but this time for the editable region. This mask is also a binary image, marking the parts that need to be modified or replaced. Using the first region mask, remove all contours outside the mask from the first contour set, leaving the contours of the unedited region, forming the first contour image set of the unedited region. Next, using the second region mask, filter out the contours within the editable region from the first contour set, obtaining the second contour image set of the editable region. The emphasis here is on operating on contours within a specified editable region so that subsequent editing can be carried out in a targeted manner.
[0037] For example, rendering a set of two-dimensional images from multiple perspectives onto the original 3D model D1. = { , , · ·· , Contour extraction is performed on each 2D image to obtain the first contour set. = { , , · · · , Based on the text editing prompt c, from the first contour set Enter a first area mask for an editing prompt in the = field. Based on the first region mask Remove the outlines outside the mask to obtain a first set of outline images of the unedited regions. = { , , · · · , }
[0038] Render a set of two-dimensional images from the generated reference 3D model D2. = { , , · · · , Contour extraction is performed on each 2D image to obtain a second contour set. = { , , · · · , Based on the text editing hint c, from the second contour set Input a second region mask M2, and remove the contours outside the mask based on the second region mask M2 to obtain the second contour image set of the edited region. = { , , · · · , The mask can be generated using text-based automatic segmentation algorithms (such as CLIPSeg, GroupViT, etc.), other interactive tools, or manually drawn by the user. The mask defines the range of the editable area and is usually represented as a binary image, where the area inside the mask is the part to be retained, and the area outside the mask is the part to be removed.
[0039] In some embodiments of this disclosure, a first contour image set and a second contour image set are stitched together to obtain a contour hint image set. One possible approach is to stitch the first contour image set and the second contour image set together to obtain a stitched contour set; then, the stitched contour set is processed to obtain the contour hint image set. In the embodiments of this disclosure, the first contour image set and the second contour image set represent different parts of the model. To generate a complete edited contour, these two contour sets need to be stitched together. The purpose of stitching is to combine the unedited and edited regions, providing a foundation for forming a coherent contour later.
[0040] For example, a collection of contour images and By stitching the pieces together, a set of stitched outlines is obtained. = { , , ·· · , During the 3D model editing process, the first contour image set of the unedited area is generated. = { , , · · · , } and the second contour image set of the editing area = { , , · · · , The two sets of outlines (}) represent different parts of the model. To generate a complete edited outline, these two outline sets need to be stitched together. The purpose of stitching is to combine the unedited and edited areas, providing a foundation for forming a coherent outline later.
[0041] In the above embodiments, the stitched contour set is processed to obtain a contour hint image set. One possible approach is to fuse the stitched contour set to obtain a fused stitched contour set; then, smooth the fused stitched contour set to obtain a smoothed stitched contour set; finally, process the smoothed stitched contour set using an edge enhancement algorithm or a contour optimization algorithm to obtain the contour hint image set. The edge enhancement algorithm is Canny edge detection, and the contour optimization algorithm is Bezier curve fitting. The smoothed stitched contour set is then processed using Canny edge detection or Bezier curve fitting to obtain the contour hint image set.
[0042] For example, based on alignment, image fusion algorithms (such as Poisson fusion, alpha blending, etc.) are used to stitch together the contour sets. A blending process is then performed. The purpose of blending is to ensure a natural transition at the seams, avoiding obvious seams or discontinuities. After blending, the contours at the seams are smoothed using Gaussian filtering or morphological operations to remove jagged edges, ensuring a smooth contour line after blending. Following smoothing, edge enhancement techniques (such as Canny edge detection) or contour optimization algorithms (such as Bezier curve fitting) can be used for post-processing of the contour lines to ensure a more natural and smooth geometry, ultimately obtaining a set of contour hints. = { , , · · · , These methods ensure a seamless transition of the outline at the joints, avoiding noticeable seams or discontinuities.
[0043] In some embodiments of this disclosure, the original 3D model is processed based on a set of text editing prompts and contour prompt images to obtain a target 3D model. One possible approach is to determine a first loss function based on a time step, a time step weight function, the input prompts, and a noisy image; determine a second loss function based on a 3D mask of the unedited region, the parameters of the original 3D model, and image features; determine a total loss function based on the first and second loss functions; and, based on the total loss function, use a prompt adapter based on decoupled cross-attention to input the set of text editing prompts and contour prompt images into each cross-attention module of the diffusion model for model iteration to obtain the target 3D model. Specifically, the text editing prompts and the contour prompt image set are used... = { , , · · · , The 3D model is iteratively optimized using SDS fractional distillation technology. SDS fractional distillation sampling is a technique that uses a pre-trained diffusion model as prior knowledge to facilitate 3D generation. By optimizing the parameters of the 3D model, SDS makes the rendered image closer to the high-density regions in the real image distribution, thus aligning with the cues from the diffusion model. A cue adapter based on decoupled cross-attention is introduced, inputting both text cues and contour information into the diffusion model to achieve multimodal interaction between text and image. This improves the matching accuracy between text cues and 3D model editing, ensuring that the editing results are highly consistent with the text description.
[0044] For example, in SDS fractional distillation sampling, let θ be a 3D model with learnable parameters, g be a differentiable rendering function that can render an image x=g(θ;c) from the 3D model, and c represent the camera viewpoint. SDS introduces a first loss function. To optimize parameter θ:
[0045] Its gradient is defined as:
[0046] in, This represents the gradient with respect to the parameter θ. For the expectation value operator, w(t) is a weight function for time step t. This is achieved by adding Gaussian noise to the rendered image x, corresponding to the t-th step of the forward diffusion process. The resulting noisy image, This indicates a prompt for input.
[0047] In this embodiment, the text editing prompt c and the outline prompt image set are included. The prompts, which serve as input together, are combined using a decoupled cross-attention-based prompt adapter to integrate the text editing prompts (c) and the contour prompt image set. The inputs are fed into each cross-attention block of the diffusion model.
[0048] To ensure that the edited model maintains consistency with the original 3D model, I incorporated a consistency preservation loss (i.e., the second loss function):
[0049] in, and These are preset hyperparameters. A 3D mask representing the unedited region. These are the parameters of the original 3D model D1. This indicates the image features extracted using the VGG network.
[0050] This disclosure introduces a consistency preservation loss, which effectively constrains the differences between the edited model and the original 3D model during the 3D model editing process. This ensures that the edited result maintains consistency with the original model in geometry, texture, and visual effects, while also meeting the requirements of text prompts. Editing is performed using a 3D model with a differentiable renderer, supporting real-time interaction and rapid feedback, thus enhancing the user experience of the editing process.
[0051] Therefore, the total loss function is expressed as: ; After iteratively optimizing the 3D model using the total loss function, the edited target 3D model is obtained. The target 3D model not only maintains the same features as the original 3D model in the unedited area as much as possible, but also achieves a smooth connection between the edited and unedited areas, and the edited area has good consistency from various viewpoints.
[0052] This disclosure achieves efficient and flexible 3D model editing by combining text prompts, image fusion, and 3D model optimization techniques. First, by generating a 3D model from text and aligning the viewpoint, a reference model highly matching the user's intent can be quickly generated. Second, image fusion algorithms and smoothing techniques ensure seamless stitching between edited and unedited areas, avoiding obvious seams or discontinuities. Finally, through a prompt adapter based on decoupled cross-attention and SDS (fractional distillation sampling) technology, iterative optimization of the 3D model is achieved, ensuring that the edited model not only meets the requirements of the text prompts but also maintains a high degree of consistency with the original model.
[0053] The 3D model generation method disclosed herein performs excellently in terms of editing accuracy, visual effects, and operational flexibility, and can be widely applied in fields such as 3D modeling, animation production, and virtual reality.
[0054] Figure 2 This is a schematic diagram of the structure of a three-dimensional model generation apparatus 20 provided for an exemplary embodiment of this disclosure. (See diagram below.) Figure 2 As shown, the 3D model generation device 20 includes: a generation module 21, a rendering module 22, an extraction module 23, a stitching module 24, and a processing module 25.
[0055] Among them, the generation module 21 is used to generate a reference 3D model using text editing prompts, wherein the reference 3D model is a model that is aligned with the original 3D model in terms of perspective; Rendering module 22 is used to render a first two-dimensional image of the original three-dimensional model and a second two-dimensional image of the reference three-dimensional model from multiple perspectives; Extraction module 23 is used to extract a first set of contour images from a first two-dimensional image and a second set of contour images from a second two-dimensional image based on text editing prompts; The stitching module 24 is used to stitch together the first contour image set and the second contour image set to obtain a contour prompt image set. The processing module 25 is used to process the original 3D model based on the text editing prompts and the outline prompt image set to obtain the target 3D model.
[0056] Optionally, when the extraction module 23 extracts the first contour image set from the first two-dimensional image and the second contour image set from the second two-dimensional image according to the text editing prompts, it is used to: Contour extraction is performed on each of the first two-dimensional images to obtain a first contour set; and Contour extraction is performed on each second two-dimensional image to obtain a second contour set; Based on the text editing prompts, input the first region mask from the first contour set; remove the contours outside the first region mask based on the first region mask to obtain the first contour image set of the unedited region; and Based on the text editing prompts, input the second region mask from the first contour set; remove the contours outside the second region mask based on the second region mask to obtain the second contour image set of the editing area.
[0057] Optionally, when stitching the first contour image set and the second contour image set to obtain the contour prompt image set, the stitching module 24 is used to: The first contour image set and the second contour image set are stitched together to obtain the stitched contour set; The spliced contour set is processed to obtain a set of contour hint images.
[0058] Optionally, when processing the stitched contour set to obtain the contour hint image set, the stitching module 24 is used to: The spliced contour set is fused to obtain the fused spliced contour set; The merged spliced contour set is smoothed to obtain a smoothed spliced contour set. By using edge enhancement algorithms or contour optimization algorithms to process the smoothed and stitched contour set, a set of contour hint images is obtained.
[0059] Optionally, the edge enhancement algorithm is Canny edge detection, and the contour optimization algorithm is Bezier curve fitting. When the stitching module 24 processes the smoothed stitched contour set using the edge enhancement algorithm or the contour optimization algorithm to obtain the contour prompt image set, it is used for: The smoothed, stitched contour set is processed using Canny edge detection or Bezier curve fitting to obtain a set of contour hint images.
[0060] Optionally, when processing the original 3D model based on the set of text editing prompts and contour prompt images to obtain the target 3D model, the processing module 25 is used to: The first loss function is determined based on the time step, the time step weight function, the input prompts, and the noisy image. The second loss function is determined based on the 3D mask of the unedited region, the parameters of the original 3D model, and image features; Determine the total loss function based on the first loss function and the second loss function; Based on the total loss function, a cue adapter based on decoupled cross-attention is used to input the set of text editing cue and contour cue images into each cross-attention module of the diffusion model for model iteration, so as to obtain the target 3D model.
[0061] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0062] Figure 3 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of the present disclosure. For example... Figure 3 As shown, the electronic device includes a memory 31 and a processor 32. Additionally, the electronic device also includes a power supply component 33 and a communication component 34.
[0063] Memory 31 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device.
[0064] The memory 31 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0065] Communication component 34 is used for data transmission with other devices.
[0066] The processor 32 is executable computer instructions stored in the memory 31 for: generating a reference 3D model using text editing prompts, wherein the reference 3D model is a model that is viewpoint aligned with the original 3D model; rendering a first 2D image of the original 3D model and a second 2D image of the reference 3D model from multiple viewpoints; extracting a first set of contour images from the first 2D image and a second set of contour images from the second 2D image according to the text editing prompts; stitching the first set of contour images and the second set of contour images together to obtain a set of contour prompt images; and processing the original 3D model according to the text editing prompts and the set of contour prompt images to obtain a target 3D model.
[0067] Accordingly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program. When the computer-readable storage medium stores a computer program, and the computer program is executed by one or more processors, it causes one or more processors to perform... Figure 1 Each step in the method embodiment.
[0068] Accordingly, embodiments of this disclosure also provide a computer program product, which includes a computer program / instructions that are executed by a processor. Figure 1 Each step in the method embodiment.
[0069] The above Figure 3The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0070] The above Figure 3 The power supply component provides power to the various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.
[0071] The aforementioned electronic devices also include a display screen and audio components.
[0072] The display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also the duration and pressure associated with the touch or swipe operation.
[0073] An audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals may be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0074] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0075] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0076] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0077] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0078] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0079] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0080] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0081] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0082] The above are merely specific embodiments of this disclosure, enabling those skilled in the art to understand or implement this disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating a three-dimensional model, characterized in that, include: Using text editing prompts, a reference 3D model is generated, wherein the reference 3D model is a model that is viewpoint aligned with the original 3D model; Render a first two-dimensional image of the original three-dimensional model and a second two-dimensional image of the reference three-dimensional model from multiple perspectives; Based on the text editing prompts, a first set of contour images is extracted from the first two-dimensional image and a second set of contour images is extracted from the second two-dimensional image; The first contour image set and the second contour image set are stitched together to obtain a contour prompt image set; The original 3D model is processed based on the text editing prompts and the set of outline prompt images to obtain the target 3D model.
2. The method according to claim 1, characterized in that, The step of extracting a first set of contour images from the first two-dimensional image and extracting a second set of contour images from the second two-dimensional image according to the text editing prompt includes: Contour extraction is performed on each of the first two-dimensional images to obtain a first contour set; and Contour extraction is performed on each of the second two-dimensional images to obtain a second contour set; Based on the text editing prompts, a first region mask is input from the first contour set; contours outside the first region mask are removed based on the first region mask to obtain a first contour image set of the unedited region; and Based on the text editing prompt, input the second region mask of the editing prompt from the first contour set; remove the contours outside the second region mask based on the second region mask to obtain the second contour image set of the editing region.
3. The method according to claim 1, characterized in that, The step of concatenating the first contour image set and the second contour image set to obtain a contour hint image set includes: The first contour image set and the second contour image set are stitched together to obtain a stitched contour set; The spliced contour set is processed to obtain a contour hint image set.
4. The method according to claim 3, characterized in that, The process of processing the stitched contour set to obtain a contour hint image set includes: The spliced contour set is fused to obtain a fused spliced contour set; The fused spliced contour set is smoothed to obtain a smoothed spliced contour set. The smoothed and stitched contour set is processed using an edge enhancement algorithm or a contour optimization algorithm to obtain the contour prompt image set.
5. The method according to claim 4, characterized in that, The edge enhancement algorithm is Canny edge detection, and the contour optimization algorithm is Bezier curve fitting. The smoothed, stitched contour set is processed using either the edge enhancement algorithm or the contour optimization algorithm to obtain the contour hint image set, including: The smoothed spliced contour set is processed using the Canny edge detection or the Bezier curve fitting to obtain the contour prompt image set.
6. The method according to claim 1, characterized in that, The step of processing the original 3D model based on the text editing prompts and the contour prompt image set to obtain the target 3D model includes: The first loss function is determined based on the time step, the time step weight function, the input prompts, and the noisy image. The second loss function is determined based on the 3D mask of the unedited region, the parameters of the original 3D model, and the image features; Determine the total loss function based on the first loss function and the second loss function; Based on the total loss function, a cue adapter based on decoupled cross-attention is used to input the set of text editing cue and the set of contour cue images into each cross-attention module of the diffusion model for model iteration, so as to obtain the target 3D model.
7. A three-dimensional model generation device, characterized in that, include: The generation module is used to generate a reference 3D model using text editing prompts, wherein the reference 3D model is a model that is viewpoint aligned with the original 3D model; The rendering module is used to render a first two-dimensional image of the original three-dimensional model and a second two-dimensional image of the reference three-dimensional model from multiple perspectives; The extraction module is used to extract a first set of contour images from the first two-dimensional image and a second set of contour images from the second two-dimensional image according to the text editing prompts. The stitching module is used to stitch the first contour image set and the second contour image set together to obtain a contour prompt image set; The processing module is used to process the original 3D model according to the text editing prompts and the outline prompt image set to obtain the target 3D model.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the steps of the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-6.