Text-guided 3D knitted product model generation and editing method and device

By introducing 3D Gaussian representation and text guidance methods, combined with 3D diffusion model and 2D semantic segmentation technology, efficient generation and local precise editing of three-dimensional knitted products are achieved, solving the problems of low generation efficiency and insufficient local semantic control in the existing technology.

CN120198627BActive Publication Date: 2025-08-12HUAQIAO UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510685431.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-12
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The prior art has problems such as low generation efficiency, insufficient local semantic control and difficulty in editing in the generation of three-dimensional knitted product models, especially in the insufficient adaptability of complex textures and geometric features of knitted products.

Method used

Using the 3D Gaussian representation method, combined with text-guided 3D diffusion model and 2D semantic segmentation technology, multi-view projection rendering and semantic segmentation are performed by generating initial point clouds, optimizing density and color, 3D Gaussians with semantic labels are obtained, and iterative updates are used to achieve efficient local editing.

Benefits of technology

The generation and editing speed of three-dimensional knitted products is improved, precise control of local areas and high-precision editing are achieved, and the problems of slow rendering speed and lack of semantic information in the prior art are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198627B_ABST
    Figure CN120198627B_ABST
Patent Text Reader

Abstract

The present invention discloses a text-guided 3D knitted product model generation and editing method and device, relating to the field of computer vision. The method comprises: S1, using a 3D diffusion model to generate an initial point cloud from knitted product prompt text; S2, optimizing and initializing the initial point cloud into a 3D Gaussian; S3, projecting and rendering the initial 3D Gaussian to obtain a multi-view image sequence, semantically segmenting the image sequence based on the editing prompt text to obtain a mask sequence; S4, obtaining a semantically labeled 3D Gaussian based on the mask sequence; S5, projecting the semantically labeled 3D Gaussian onto a 2D plane to obtain a rendered image; inputting the rendered image and the editing prompt text into a 2D diffusion model to output a loss gradient; and using the loss gradient to guide the iteration of the 3D Gaussian. The 3D Gaussian obtained after the iteration is the final 3D knitted product model. By introducing a 3D Gaussian and using the editing prompt text to add semantic labels to the 3D Gaussian, the present invention balances generation efficiency, local semantic control, and high-precision editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a text-guided 3D knitted product model generation and editing method and device. Background Art

[0002] The generation and editing of 3D knitted product models is a crucial foundation for digital design and intelligent manufacturing in the textile industry. Traditional methods rely on computer-aided design software, constructing 3D models through manual drawing and parameterized adjustments. However, these methods require extensive manual intervention, are inefficient, and struggle to rapidly respond to the demands of generating complex knitted textures and structures.

[0003] In recent years, with the development of deep learning technology, data-driven 3D model generation methods have gradually become a research hotspot. Among existing methods, some use implicit neural representations or point cloud generation networks, combined with text or image guidance to generate 3D models. However, implicit representations suffer from slow rendering speeds and high video memory consumption, making them difficult to meet the requirements of real-time interactive editing. Furthermore, existing methods focus primarily on global shape generation and lack the ability to precisely control local semantic regions of the model. This results in inefficient editing of specific knitted product structures, requiring manual annotation or complex post-processing processes.

[0004] When it comes to semantic editing of 3D models, existing technologies primarily achieve local adjustments through a rough mapping of 2D image segmentation results onto the 3D model. For example, projection methods based on multi-view 2D masks can reverse-map user-annotated 2D semantic information into 3D space. However, due to the discrete nature of 3D representations (such as point clouds or voxels), mapping accuracy and editing effectiveness are limited. Furthermore, existing editing methods based on diffusion models typically employ global optimization strategies, making it difficult to achieve refined adjustments to local details while maintaining the stability of unedited areas. This is particularly problematic for the complex texture structures and geometric features found in knitted products.

[0005] Therefore, the existing technology urgently needs a three-dimensional knitted product modeling method that can take into account generation efficiency, local semantic control and high-precision editing, so as to solve technical bottlenecks such as slow implicit representation rendering speed, lack of semantic information and difficulty in accurate positioning of local editing. Summary of the Invention

[0006] To address the above problems, the present invention proposes a text-guided 3D knitted product model generation and editing method and device. By introducing 3D Gaussian to generate 3D knitted product models, and using the editing prompt text after 2D semantic segmentation to add semantic labels to the 3D Gaussian, three-dimensional knitted product modeling is achieved that takes into account generation efficiency, local semantic control and high-precision editing.

[0007] On the one hand, the text-guided 3D knitted product model generation and editing method has the following specific steps:

[0008] S1, input the knitted product prompt text into the pre-trained 3D diffusion model to generate the initial point cloud;

[0009] S2, density growth and color optimization are performed on the initial point cloud to obtain an optimized point cloud; the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian;

[0010] S3, performs multi-view projection rendering on the initial 3D Gaussian to obtain a multi-view image sequence; performs semantic segmentation on the multi-view image sequence based on the input editing prompt text to obtain a mask sequence;

[0011] S4, obtains 3D Gaussian with semantic labels based on the mask sequence;

[0012] S5, projects the semantically labeled 3D Gaussian from 3D to 2D to obtain a rendered image; inputs the rendered image and editing prompt text into a pre-trained 2D diffusion model, and outputs the loss gradient; uses the loss gradient to guide the projection rendering process of the semantically labeled 3D Gaussian to iteratively update the semantically labeled 3D Gaussian, and uses the 3D Gaussian obtained after the iteration as the final 3D knitted product model.

[0013] Preferably, the knitted product prompt text is input into a pre-trained 3D diffusion model to generate an initial point cloud, specifically as follows:

[0014] Use CLIP text encoder to generate knitted product hint text embedding from knitted product hint text;

[0015] Embed the knitted product prompt text into the pre-trained 3D diffusion model to obtain a 3D asset represented by a triangular mesh;

[0016] Extract the triangular mesh vertices of the 3D asset as basic points in the initial point cloud, and construct the initial point cloud based on the basic points.

[0017] Preferably, the density growth and color optimization are performed on the initial point cloud to obtain an optimized point cloud; and the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian, specifically as follows:

[0018] Calculating the size of a bounding box of the entire point cloud data based on the three-dimensional coordinate information of the initial point cloud; the bounding box is formed by taking the largest three-dimensional coordinate and the smallest three-dimensional coordinate in the point cloud as diagonal endpoints to form a cubic area as the bounding box;

[0019] Based on the initial point cloud, the point cloud is uniformly grown in the bounding box to obtain the growing point cloud;

[0020] Build a KD tree using the positions of the points in the initial point cloud;

[0021] Calculate the normalized distance from each point in the growing point cloud to the nearest point in the KD tree, and select the point whose normalized distance is less than the preset normalized distance threshold as the point cloud growth point;

[0022] Retrieve the closest point between the initial point cloud and the point cloud growth point, and assign a color to the point cloud growth point based on the closest point in the initial point cloud, expressed as:

[0023] ;

[0024] in, Indicates the color attribute of the point cloud growth point, represents the interference value of random sampling; Indicates the color attribute of the closest point between the initial point cloud and the point cloud growth point;

[0025] Integrate the initial point cloud and the point cloud growth points to form an optimized point cloud;

[0026] According to the position attributes and color attributes of the optimized point cloud, the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian.

[0027] Preferably, the initial 3D Gaussian is subjected to multi-view projection rendering to obtain a multi-view image sequence; and the multi-view image sequence is subjected to semantic segmentation based on the input editing prompt text to obtain a mask sequence, as follows:

[0028] Based on a predefined camera parameter set, the pinhole camera model is used to project the initial 3D Gaussian onto a two-dimensional plane at several camera viewpoints to obtain the 2D Gaussian at the corresponding viewpoints.

[0029] Differentiable rendering technology is used to render 2D Gaussian images at different viewing angles to obtain multi-view image sequences;

[0030] Based on the editing prompt text, a pre-trained semantic segmentation model is used to perform pixel-level semantic segmentation on the multi-view image sequence, extracting the areas related to the editing target and generating an initial mask sequence.

[0031] The initial mask is morphologically expanded to cover the target edge area, and the internal holes are filled through connected component analysis to obtain a mask sequence.

[0032] Preferably, the method of obtaining a 3D Gaussian with a semantic label based on a mask sequence is as follows:

[0033] Inversely project the mask sequence back to the 3D Gaussian space;

[0034] Traverse the mask sequence for each 3D Gaussian point of the initial 3D Gaussian and calculate the association weight of each 3D Gaussian point with the semantic label, expressed as:

[0035] ;

[0036] in, represents the association weight between the i-th initial 3D Gaussian point and the j-th semantic label; Indicates the transparency of the i-th initial 3D Gaussian point at pixel p; represents the transmittance of the i-th initial 3D Gaussian point at pixel p; Represents the semantic mask value of the j-th semantic label of pixel p; Indicates summation.

[0037] The average weight of each 3D Gaussian point and the semantic label of each category is calculated based on the association weight of each 3D Gaussian point and the semantic label. If the average weight of a Gaussian point and a certain category of semantic label exceeds the preset weight threshold, the semantic label of the category is assigned to the Gaussian point. The set of all 3D Gaussian points with semantic labels is the 3D Gaussian with semantic labels.

[0038] Preferably, the loss gradient is a comprehensive gradient of the mask loss gradient and the noise prediction error gradient; the comprehensive gradient is expressed as:

[0039] ;

[0040] in, and represents a hyperparameter used to balance the strength of global and local constraints; Represents the comprehensive gradient, that is, the comprehensive loss gradient; Represents the noise prediction error gradient, that is, the noise prediction error loss gradient; Represents the gradient of mask loss, i.e., mask loss gradient; Represents the gradient.

[0041] Preferably, the noise prediction error gradient is expressed as:

[0042] ;

[0043] in, Represents the rendered image after adding noise; represents the noise predicted by the 2D diffusion model, Indicates is the main input, conditioned on y and t; y represents the editing prompt text; ε represents real random noise; represents the weight function at time step t, which is used to balance the contribution of different noise levels; represents the partial derivative of x with respect to θ, x represents the rendered image, and θ represents the optimizable parameters of the 3D Gaussian; represents the expected value of t and ε.

[0044] Preferably, the mask loss is expressed as:

[0045] ;

[0046] in, represents mask loss; M represents semantic mask; ⊙ represents element-wise multiplication; represents the noise predicted by the 2D diffusion model, Indicates is the main input, conditioned on y and t; Calculates the sum of squares of all elements.

[0047] On the other hand, the text-guided 3D knitted product model generation and editing device includes the following:

[0048] Point cloud generation module, used to input knitted product prompt text into the pre-trained 3D diffusion model to generate the initial point cloud;

[0049] The initial 3D Gaussian generation module is used to perform density growth and color optimization on the initial point cloud to obtain an optimized point cloud; the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian;

[0050] The mask sequence generation module is used to perform multi-view projection rendering on the initial 3D Gaussian to obtain a multi-view image sequence; and perform semantic segmentation on the multi-view image sequence based on the editing prompt text to obtain a mask sequence;

[0051] 3D Gaussian generation module with semantic labels, used to obtain 3D Gaussian with semantic labels based on mask sequences;

[0052] The 3D knitted product model acquisition module is used to project the semantically labeled 3D Gaussian from 3D to 2D to obtain a rendered image; the rendered image and editing prompt text are input into a pre-trained 2D diffusion model and the loss gradient is output; the loss gradient is used to guide the projection rendering process of the semantically labeled 3D Gaussian to iteratively update the semantically labeled 3D Gaussian. The 3D Gaussian obtained after the iteration is the final 3D knitted product model.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] (1) The present invention introduces the explicit representation of 3D Gaussian, which greatly improves the generation and editing speed of 3D knitted products and solves the shortcoming of the slow speed of implicit neural expression methods;

[0055] (2) The present invention performs image semantic segmentation based on editing prompt text, and then obtains 3D Gaussians with semantic labels. It introduces Gaussian semantic tracking to achieve accurate local positioning and editing of 3D knitted products, which greatly improves local editable optimization compared with traditional global optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The present invention will be described in further detail below with reference to the accompanying drawings;

[0057] Figure 1 Flowchart of a text-guided 3D knitted product model generation and editing method according to an embodiment of the present invention;

[0058] Figure 2 Schematic diagram of the flow of a text-guided 3D knitted product model generation and editing method according to an embodiment of the present invention;

[0059] Figure 3 This is a structural block diagram of a text-guided 3D knitted product model generation and editing device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0060] The present invention is further described below through specific embodiments.

[0061] like Figure 1 and Figure 2 As shown in FIG, the text-guided 3D knitted product model generation and editing method includes the following specific steps:

[0062] S1, input the knitted product prompt text into the pre-trained 3D diffusion model to generate the initial point cloud.

[0063] The knitted product prompt text is the text used to guide the generation of the 3D knitted product model. The 3D diffusion model of this embodiment adopts Shap-E, which is used to generate 3D assets based on text. The knitted product prompt text is encoded into a high-dimensional semantic vector through the CLIP text encoder, and then the vector is input as a conditional signal into the noise prediction network of the 3D diffusion model Shap-E to guide the inverse denoising process in the implicit parameter space. Starting from the initial Gaussian noise, the noise distribution is gradually corrected through multi-step iterations, so that the implicit signed distance field parameters gradually converge to a geometric form that matches the text semantics during the denoising process, and finally generate signed distance field parameters that conform to the text semantics, and then output a 3D asset ℳ in the form of a triangular mesh. Extract the mesh vertices of the 3D asset ℳ and use it as the initial point cloud The basic points in the point cloud can be represented as ,in Represents the position attributes of the point cloud in three-dimensional space, Represents the color attribute of a point in a point cloud.

[0064] S2, density growth and color optimization are performed on the initial point cloud to obtain an optimized point cloud; the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian.

[0065] Based on the initial point cloud The 3D coordinate information of the point cloud is used to calculate the bounding box size of the entire point cloud data, that is, a cubic area is formed with the largest 3D coordinate and the smallest 3D coordinate in the point cloud as the diagonal endpoints as the bounding box. Then the point cloud is uniformly grown in the bounding box to obtain the point cloud. , or expressed as ,in, and Represents point clouds The position and color attributes of the initial point cloud. The position of the point in Constructing a KD tree To improve the speed of the subsequent retrieval steps. Calculate point cloud Each point in Normalized distance to the nearest point , the point cloud Calculated in The points are filtered out as the required point cloud growth points, and these filtered points are represented as point clouds , or expressed as ,in, and Represent point clouds The position and color attributes of the point cloud. Points in the original point cloud are retrieved according to the nearest point ( Center and point cloud The color information of the point closest to the point in the image is used to assign a color to it. The specific assignment process is expressed as follows:

[0066] ;

[0067] in, Representing point clouds The color attribute, is a randomly sampled interference value between 0 and 0.2, Representing point clouds Color property of the midpoint.

[0068] Point Cloud With point cloud Integrate the two point sets into one set to form a new point cloud , or expressed as ,in, and Represent point clouds The position and color attributes of the point cloud. Location attributes and color attributes Initialize 3D Gaussian ( 、 and Represent the position, color and opacity properties of the 3D Gaussian, represents the covariance matrix of the 3D Gaussian) and color .

[0069] The so-called 3D Gaussian refers to a set of 3D Gaussian ellipsoids. This set forms a 3D model. Each Gaussian sphere has attributes such as position and size, and the gradient of the loss is used to guide the direction of updating these attributes.

[0070] S3, performs multi-view projection rendering on the initial 3D Gaussian to obtain a multi-view image sequence; performs semantic segmentation on the multi-view image sequence based on the input editing prompt text to obtain a mask sequence.

[0071] Based on a predefined set of camera parameters, the initial 3D Gaussian model is projected onto a two-dimensional plane under several camera perspectives through a pinhole camera model to obtain a 2D Gaussian under the corresponding perspective. The 2D Gaussian under each perspective is then rendered using differentiable rendering technology to obtain a multi-perspective RGB image sequence (or multi-perspective image sequence). Subsequently, based on the input editing prompt text, the pre-trained semantic segmentation model SAM (Segment Anything Model) is used to perform pixel-level semantic segmentation on the multi-perspective images, extract the areas related to the editing target to generate an initial mask sequence, and perform a morphological dilation operation on the mask to cover the target edge area. At the same time, the internal holes are filled through connected domain analysis to form a refined mask sequence. In this process, the camera pose parameters corresponding to each 2D mask are recorded, and a position mapping relationship between the mask and the 3D Gaussian space is established to ensure the geometric consistency of subsequent inverse projection.

[0072] S4, obtains 3D Gaussian with semantic labels based on the mask sequence.

[0073] For the i-th initial 3D Gaussian point, traverse the multi-view mask sequence and calculate the association weight of each 3D Gaussian point with the semantic label, which is expressed as:

[0074] ;

[0075] in, represents the association weight between the i-th initial 3D Gaussian point and the j-th semantic label; Indicates the transparency of the i-th initial 3D Gaussian point at pixel p; represents the transmittance of the i-th initial 3D Gaussian point at pixel p; Represents the semantic mask value of the j-th semantic label of pixel p; Indicates summation.

[0076] Obtain the association weight of the i-th Gaussian point in each category of semantic labels, sum these association weights, and the proportion of the j-th category semantic label in the sum of the association weights is the average weight of the semantics of this category. If the average weight of a certain category of semantics exceeds the preset weight threshold, the corresponding semantic label is assigned to the Gaussian point to form a 3D Gaussian set with semantic attributes, that is, a 3D Gaussian with semantic labels.

[0077] S5, projects the semantically labeled 3D Gaussian from 3D to 2D to obtain a rendered image; inputs the rendered image and editing prompt text into a pre-trained 2D diffusion model, and outputs the loss gradient; uses the loss gradient to guide the projection rendering process of the semantically labeled 3D Gaussian to iteratively update the semantically labeled 3D Gaussian, and uses the 3D Gaussian obtained after the iteration as the final 3D knitted product model.

[0078] The 3D Gaussian sputtering technique is used to project the semantically labeled 3D Gaussian set onto the 2D image plane to generate a rendered image. ,in, Represents a set of 3D Gaussian parameters, including position μ, color c, covariance Σ, and opacity α.

[0079] The input editing hint text y is converted into a semantic embedding vector through the CLIP text encoder. Input the pre-trained 2D diffusion model ϕ with the text embedding y and calculate the noisy prediction error gradient:

[0080] ;

[0081] in is the rendered image after adding noise, is the noise predicted by the diffusion model, is the main input, conditioned on y and t, and ε is true random noise; is the weight function at time step t, which is used to balance the contribution of different noise levels. represents the partial derivative of x with respect to θ, where x is the rendered image and θ represents the optimizable parameters of the 3D Gaussian.

[0082] The 2D diffusion model in this embodiment uses the stabilityai / stable-diffusion-2-1-base model, which generates images based on text prompts and guides the update of the 3D Gaussian model by calculating the noise prediction error gradient, thereby achieving precise editing of the 3D model.

[0083] The mask loss is used to limit the editing to the specified semantic area to avoid global distortion. The mask loss is expressed as:

[0084] ;

[0085] Where M is the semantic mask, ⊙ represents element-wise multiplication, represents the calculation of the square sum of all elements of the masked noise difference matrix. ε is the true random noise.

[0086] The noise prediction error gradient and the mask loss gradient are weighted and fused to obtain the comprehensive gradient, which is expressed as:

[0087] ;

[0088] in, and is a hyperparameter that balances the strength of global and local constraints. Represents the gradient. Represents comprehensive loss. represents the noise prediction error loss. Represents mask loss.

[0089] According to the comprehensive gradient Back propagation updates the 3D Gaussian parameters, and after multiple rounds of iterations until the fitting is achieved, the final 3D knitted product model is obtained.

[0090] like Figure 3 As shown, the present invention also discloses a text-guided 3D knitted product model generation and editing device, comprising:

[0091] The point cloud generation module 301 is used to input the knitted product prompt text into the pre-trained 3D diffusion model to generate an initial point cloud;

[0092] The initial 3D Gaussian generation module 302 is used to perform density growth and color optimization on the initial point cloud to obtain an optimized point cloud; and initialize the optimized point cloud to a 3D Gaussian to obtain an initial 3D Gaussian.

[0093] The mask sequence generation module 303 is used to perform multi-view projection rendering on the initial 3D Gaussian to obtain a multi-view image sequence; perform semantic segmentation on the multi-view image sequence based on the input editing prompt text to obtain a mask sequence;

[0094] A 3D Gaussian with semantic labels generation module 304 is used to obtain a 3D Gaussian with semantic labels based on a mask sequence;

[0095] The 3D knitted product model acquisition module 305 is used to project the 3D Gaussian with semantic labels from 3D to 2D to obtain a rendered image; input the rendered image and editing prompt text into a pre-trained 2D diffusion model and output a loss gradient; use the loss gradient to guide the projection rendering process of the 3D Gaussian with semantic labels to iteratively update the 3D Gaussian with semantic labels, and use the 3D Gaussian obtained after the iteration as the final 3D knitted product model.

[0096] The specific implementation of the text-guided 3D knitted product model generation and editing device is the same as the text-guided 3D knitted product model generation and editing method, and will not be repeated in this embodiment.

[0097] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited to this. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.

Claims

1. A text-guided 3D knitted product model generation and editing method, characterized in that: The steps include: S1, input the knitted product prompt text into the pre-trained 3D diffusion model to generate the initial point cloud; S2, density growth and color optimization are performed on the initial point cloud to obtain an optimized point cloud; the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian; S3, performs multi-view projection rendering on the initial 3D Gaussian to obtain a multi-view image sequence; performs semantic segmentation on the multi-view image sequence based on the input editing prompt text to obtain a mask sequence; S4, obtains 3D Gaussian with semantic labels based on the mask sequence; S5: Project the semantically labeled 3D Gaussian from 3D to 2D to obtain a rendered image; input the rendered image and the editing prompt text into a pre-trained 2D diffusion model and output a loss gradient; use the loss gradient to guide the projection rendering process of the semantically labeled 3D Gaussian to iteratively update the semantically labeled 3D Gaussian, and use the 3D Gaussian obtained after the iteration as the final 3D knitted product model; The density growth and color optimization are performed on the initial point cloud to obtain an optimized point cloud; the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian, as follows: Calculating the size of a bounding box of the entire point cloud data based on the three-dimensional coordinate information of the initial point cloud; the bounding box is formed by taking the largest three-dimensional coordinate and the smallest three-dimensional coordinate in the point cloud as diagonal endpoints to form a cubic area as the bounding box; Based on the initial point cloud, the point cloud is uniformly grown in the bounding box to obtain the growing point cloud; Build a KD tree using the positions of the points in the initial point cloud; Calculate the normalized distance from each point in the growing point cloud to the nearest point in the KD tree, and select the point whose normalized distance is less than the preset normalized distance threshold as the point cloud growth point; Retrieve the closest point between the initial point cloud and the point cloud growth point, and assign a color to the point cloud growth point based on the closest point in the initial point cloud, expressed as: c′ r =c m +a; Among them, c r ′ represents the color attribute of the point cloud growth point, a represents the interference value of random sampling; c m Indicates the color attribute of the closest point between the initial point cloud and the point cloud growth point; Integrate the initial point cloud and the point cloud growth points to form an optimized point cloud; According to the position attributes and color attributes of the optimized point cloud, the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian.

2. The text-guided 3D knitted product model generation and editing method according to claim 1, characterized in that: The knitted product prompt text is input into the pre-trained 3D diffusion model to generate the initial point cloud, as follows: Use CLIP text encoder to generate knitted product hint text embedding from knitted product hint text; Embed the knitted product prompt text into the pre-trained 3D diffusion model to obtain a 3D asset represented by a triangular mesh; Extract the triangular mesh vertices of the 3D asset as basic points in the initial point cloud, and construct the initial point cloud based on the basic points.

3. The text-guided 3D knitted product model generation and editing method according to claim 1, characterized in that: The initial 3D Gaussian is subjected to multi-view projection rendering to obtain a multi-view image sequence; the multi-view image sequence is subjected to semantic segmentation based on the input editing prompt text to obtain a mask sequence, as follows: Based on a predefined camera parameter set, the pinhole camera model is used to project the initial 3D Gaussian onto a two-dimensional plane at several camera viewpoints to obtain the 2D Gaussian at the corresponding viewpoints. Differentiable rendering technology is used to render 2D Gaussian images at different viewing angles to obtain multi-view image sequences; Based on the editing prompt text, a pre-trained semantic segmentation model is used to perform pixel-level semantic segmentation on the multi-view image sequence, extracting the areas related to the editing target and generating an initial mask sequence. The initial mask is morphologically expanded to cover the target edge area, and the internal holes are filled through connected component analysis to obtain a mask sequence.

4. The text-guided 3D knitted product model generation and editing method according to claim 1, characterized in that: The method of obtaining a 3D Gaussian with semantic labels based on a mask sequence is as follows: Inversely project the mask sequence back to the 3D Gaussian space; Traverse the mask sequence for each 3D Gaussian point of the initial 3D Gaussian and calculate the association weight of each 3D Gaussian point with the semantic label, expressed as: w ij =∑ p α i (p)T i (p)M j (p); Among them, w ij represents the association weight between the i-th initial 3D Gaussian point and the j-th semantic label; α i (p) represents the transparency of the i-th initial 3D Gaussian point at pixel p; T i (p) represents the transmittance of the i-th initial 3D Gaussian point at pixel p; M j (p) represents the semantic mask value of the j-th semantic label of pixel p; ∑ p Indicates summation; The average weight of each 3D Gaussian point and the semantic label of each category is calculated based on the association weight of each 3D Gaussian point and the semantic label. If the average weight of a Gaussian point and a certain category of semantic labels exceeds the preset weight threshold, the semantic label of the category is assigned to the Gaussian point. The set of all 3D Gaussian points with semantic labels is the 3D Gaussian with semantic labels.

5. The text-guided 3D knitted product model generation and editing method according to claim 1, characterized in that: The loss gradient is the comprehensive gradient of the mask loss gradient and the noise prediction error gradient; the comprehensive gradient is expressed as: Among them, ζ1 and ζ2 represent hyperparameters used to balance the strength of global and local constraints; Represents the comprehensive gradient, that is, the comprehensive loss L total gradient; Represents the noise prediction error gradient, that is, the noise prediction error loss L SDS gradient; Represents the mask loss gradient, that is, the mask loss L mask gradient; Represents the gradient.

6. The text-guided 3D knitted product model generation and editing method according to claim 5, characterized in that: The noise prediction error gradient is expressed as: Among them, z t Represents the rendered image after adding noise; represents the noise predicted by the 2D diffusion model, (z t ; y,t) represents z t is the main input, conditioned on y and t; y represents the edit prompt text; ε represents the true random noise; w(t) represents the weight function at time step t, which is used to balance the contribution of different noise levels; represents the partial derivative of x with respect to θ, x represents the rendered image, and θ represents the optimizable parameters of the 3D Gaussian; represents the expected value of t and ε.

7. The text-guided 3D knitted product model generation and editing method according to claim 5, characterized in that: The mask loss is expressed as: Among them, L mask represents mask loss; M represents semantic mask; ⊙ represents element-wise multiplication; represents the noise predicted by the 2D diffusion model, (z t ; y,t) represents z t is the main input, conditioned on y and t; Calculates the sum of squares of all elements.

8. A text-guided 3D knitted product model generation and editing device using the text-guided 3D knitted product model generation and editing method according to any one of claims 1 to 7, comprising: Point cloud generation module, used to input knitted product prompt text into the pre-trained 3D diffusion model to generate the initial point cloud; The initial 3D Gaussian generation module is used to perform density growth and color optimization on the initial point cloud to obtain an optimized point cloud; the optimized point cloud is initialized to a 3D Gaussian to obtain an initial 3D Gaussian; The mask sequence generation module is used to perform multi-view projection rendering on the initial 3D Gaussian to obtain a multi-view image sequence; based on the input editing prompt text, the multi-view image sequence is semantically segmented to obtain a mask sequence; 3D Gaussian generation module with semantic labels, used to obtain 3D Gaussian with semantic labels based on mask sequences; The 3D knitted product model acquisition module is used to project the semantically labeled 3D Gaussian from 3D to 2D to obtain a rendered image; the rendered image and editing prompt text are input into a pre-trained 2D diffusion model and the loss gradient is output; the loss gradient is used to guide the projection rendering process of the semantically labeled 3D Gaussian to iteratively update the semantically labeled 3D Gaussian, and the 3D Gaussian obtained after the iteration is used as the final 3D knitted product model.

Citation Information

Patent Citations

  • Garment secondary design system and method

    CN117576246A

  • Semantic picture editing method based on diffusion model

    CN117671084A