Three-dimensional vector graph generation method and device, electronic equipment and storage medium
By iteratively optimizing the 3D Gaussian splash representation to generate 3D vector graphics, the problems of multi-view consistency and insufficient occlusion perception in the existing technology are solved, and efficient and accurate multi-view graphics generation is achieved.
Patent Information
- Application Number
- CN202510574618.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-09
AI Technical Summary
Existing vector graphics cannot support multi-view consistency and occlusion awareness, which limits their use in applications such as multi-view sketch design and cartoon animation.
By obtaining the descriptive text of the target object, the initial three-dimensional Gaussian splash representation is determined, and the target three-dimensional vector graphics are generated through iterative optimization processing, including determining latent variables, optimizing the Gaussian splash representation, generating supervision images and intermediate consistency trajectory samples, and finally generating three-dimensional vector graphics that meet multi-view consistency and occlusion perception.
It enables the generation of three-dimensional vector graphics that meet multi-view consistency and occlusion perception without manual processing, improving the accuracy and efficiency of generation.
Smart Images

Figure CN120612423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for generating three-dimensional vector graphics. Background Art
[0002] Currently, vector graphics are widely used in fields such as graphic design, artistic creation, and conceptual illustration. However, existing vector graphics are primarily designed from a specific perspective and cannot support arbitrary viewing angles. Furthermore, they lack multi-view consistency and occlusion awareness, limiting their use in applications such as multi-view sketching and cartoon animation.
[0003] Therefore, how to provide a method that can meet the requirements of multi-view consistency, occlusion perception and detail optimization is a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The present invention provides a method, device, electronic device and storage medium for generating three-dimensional vector graphics, which are used to solve the problem that vector graphics are insufficient in generating multi-view consistency and occlusion perception.
[0005] The present invention provides a method for generating three-dimensional vector graphics, comprising: Get the description text of the target object; Determining an initial three-dimensional Gaussian splash representation based on the description text; the initial three-dimensional Gaussian splash representation is a three-dimensional visual representation of the target object; An iterative optimization process is performed on the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object.
[0006] According to a method for generating a three-dimensional vector graphic provided by the present invention, the iterative optimization processing of the initial three-dimensional Gaussian splatter representation to generate a target three-dimensional vector graphic corresponding to the target object includes: Step A: In each iteration, based on the initial three-dimensional Gaussian splatter representation, determining a latent variable of the target object corresponding to different camera poses; the latent variable represents the noise potential representation of the target object under the camera pose; Step B: Based on each of the latent variables, optimizing the initial three-dimensional Gaussian splash representation to obtain an optimized three-dimensional Gaussian splash representation; the optimized three-dimensional Gaussian splash representation is used as a new initial three-dimensional Gaussian splash representation for the next iteration; Step C: Processing the optimized three-dimensional Gaussian splatter representation to obtain supervision images of the target object at different perspectives and intermediate consistent trajectory samples corresponding to the target object; the intermediate consistent trajectory samples represent the trajectory image of the target object; Step D: Based on the supervisory image and the intermediate consistency trajectory sample, optimizing the initial three-dimensional vector graphics corresponding to the supervisory image to obtain an optimized three-dimensional vector graphics; the optimized three-dimensional vector graphics is used as a new three-dimensional vector graphics for the next iteration; Step E: Repeat steps A to D. When the number of iterations meets the first target number of iterations, the finally optimized three-dimensional vector graphic is used as the target three-dimensional vector graphic corresponding to the target object.
[0007] According to a method for generating three-dimensional vector graphics provided by the present invention, determining latent variables of the target object corresponding to different camera postures based on the initial three-dimensional Gaussian splash representation includes: Determining, based on the initial three-dimensional Gaussian splatter representation, a first rendered image of the target object corresponding to different camera poses; Each of the first rendered images is sampled to obtain the latent variables of the target object at viewing angles corresponding to different camera postures.
[0008] According to a method for generating three-dimensional vector graphics provided by the present invention, the method of optimizing the initial three-dimensional Gaussian splash representation based on each of the latent variables to obtain an optimized three-dimensional Gaussian splash representation includes: determining a first loss value based on each of the latent variables; Based on the first loss value, the initial three-dimensional Gaussian splash representation is optimized to obtain an optimized three-dimensional Gaussian splash representation.
[0009] According to a method for generating a three-dimensional vector graphic provided by the present invention, the method optimizes an initial three-dimensional vector graphic corresponding to the supervisory image based on the supervisory image and the intermediate consistent trajectory sample to obtain an optimized three-dimensional vector graphic, including: determining a guiding image of the target object based on the intermediate consistent trajectory samples, the latent variable, and a preset update direction; Rendering a projection of the initial three-dimensional vector graphics corresponding to the supervision image to obtain a second rendered image at a corresponding viewing angle; Calculating a second loss value based on the guide image and the second rendered image; Based on the second loss value, the initial three-dimensional vector graphics are optimized to obtain the optimized three-dimensional vector graphics.
[0010] According to a method for generating a three-dimensional vector graphic provided by the present invention, the method further includes: Based on the projection results of the target three-dimensional vector graphics, a multi-layer perceptron is used to determine the average importance of each curve of the target object under different camera postures; Based on the average importance, the multilayer perceptron is optimized to obtain a target multilayer perceptron.
[0011] According to a method for generating three-dimensional vector graphics provided by the present invention, the method of optimizing the multilayer perceptron based on the average importance to obtain a target multilayer perceptron includes: If the average importance is not less than a preset threshold, adding the curve to the importance set; Calculating a third loss value based on the three-dimensional vector graphics corresponding to the curves in the importance set and the optimized three-dimensional Gaussian splash representation; Optimizing the multilayer perceptron based on the third loss value to obtain an optimized multilayer perceptron; Repeat the above steps to obtain the optimized multilayer perceptron. When the number of iterations meets the second target number of iterations, the finally optimized multilayer perceptron is used as the target multilayer perceptron.
[0012] The present invention also provides a device for generating three-dimensional vector graphics, comprising: The acquisition module is used to obtain the description text of the target object; A first determining module is configured to determine an initial three-dimensional Gaussian splash representation based on the description text; the initial three-dimensional Gaussian splash representation is a three-dimensional visual representation of the target object; A generation module is used to perform iterative optimization processing on the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described methods for generating three-dimensional vector graphics when executing the computer program.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for generating three-dimensional vector graphics.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the above-mentioned methods for generating three-dimensional vector graphics.
[0016] The present invention provides a method, device, electronic device, and storage medium for generating three-dimensional vector graphics. The method obtains a descriptive text of a target object; based on the descriptive text, determines an initial three-dimensional Gaussian splatter representation; the initial three-dimensional Gaussian splatter representation is a three-dimensional visual representation of the target object; and the initial three-dimensional Gaussian splatter representation is iteratively optimized to generate a target three-dimensional vector graphic corresponding to the target object. By iteratively optimizing the initial three-dimensional Gaussian splatter representation, the generated target three-dimensional vector graphic can meet the requirements of multi-view consistency and occlusion perception without manual processing, thereby improving the accuracy and efficiency of three-dimensional vector graphic generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is one of the flow charts of the method for generating three-dimensional vector graphics provided by the present invention.
[0019] Figure 2 This is the second flow chart of the method for generating three-dimensional vector graphics provided by the present invention.
[0020] Figure 3 This is one of the schematic diagrams of the target three-dimensional vector graphics provided by the present invention.
[0021] Figure 4 This is the second schematic diagram of the target three-dimensional vector graphics provided by the present invention.
[0022] Figure 5 It is a structural schematic diagram of the device for generating three-dimensional vector graphics provided by the present invention.
[0023] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0025] The following combination Figure 1-Figure 4The method for generating three-dimensional vector graphics of the present invention is described.
[0026] Figure 1 This is one of the flow charts of the method for generating three-dimensional vector graphics provided by the present invention, such as Figure 1 As shown, the method includes the following steps 101 to 103.
[0027] Step 101: Get the description text of the target object.
[0028] Specifically, a description text of a target object can be obtained through user input. The target object is the object for which a 3D vector graphic is to be generated. For example, target objects include, but are not limited to, a cat, an airplane, and a helmet. The description text is a detailed description of the desired 3D vector graphic. The description text can be input in the form of a string, which can be recognized by the computer and used in the subsequent generation process. For example, the text description is "a cat."
[0029] Step 102: Determine an initial three-dimensional Gaussian splash representation based on the description text; the initial three-dimensional Gaussian splash representation is a three-dimensional visual presentation of the target object.
[0030] Specifically, the initial three-dimensional Gaussian splash representation is a three-dimensional visual representation of the target object, that is, the initial three-dimensional Gaussian splash representation is a three-dimensional geometry and appearance representation of the parameterized scene.
[0031] Based on the description text, the three-dimensional Gaussian splash representation is initialized, so that the initial three-dimensional Gaussian splash representation can be determined.
[0032] Step 103: performing iterative optimization processing on the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object.
[0033] Specifically, since the three-dimensional visual presentation of the target object represented by the initial three-dimensional Gaussian splash is relatively rough and not detailed enough, an iterative optimization method is used to iteratively optimize the initial three-dimensional Gaussian splash representation to finally generate a target three-dimensional vector graphic corresponding to the target object.
[0034] The present invention provides a method for generating 3D vector graphics. The method comprises obtaining a descriptive text of a target object; determining an initial 3D Gaussian splatter representation based on the descriptive text; using the initial 3D Gaussian splatter representation as a 3D visual representation of the target object; and iteratively optimizing the initial 3D Gaussian splatter representation to generate a target 3D vector graphic corresponding to the target object. By iteratively optimizing the initial 3D Gaussian splatter representation, the generated target 3D vector graphic can meet the requirements of multi-view consistency and occlusion perception without manual processing, thereby improving the accuracy and efficiency of 3D vector graphics generation.
[0035] Optionally, a specific implementation of step 103 includes the following steps: Step A: In each iteration, based on the initial three-dimensional Gaussian splatter representation, determine the latent variables of the target object corresponding to different camera poses; the latent variables represent the noise potential representation of the target object under the camera pose.
[0036] Specifically, this application constructs a three-dimensional Gaussian splash distillation model to process the initial three-dimensional Gaussian splash representation corresponding to the description text, that is, using a differentiable renderer to render the initial three-dimensional Gaussian splash representation into multi-perspective rendered images of the target object under different camera postures; then based on the fractional distillation sampling method, the inverse process of the denoising implicit model (DDIM) is used to sample latent variables with different noise levels from the rendered image, and then the target three-dimensional Gaussian splash representation is obtained.
[0037] Step B: Based on each of the latent variables, the initial three-dimensional Gaussian splash representation is optimized to obtain an optimized three-dimensional Gaussian splash representation; the optimized three-dimensional Gaussian splash representation is used as a new initial three-dimensional Gaussian splash representation for the next iteration.
[0038] Specifically, based on each latent variable, the initial three-dimensional Gaussian splash representation is optimized to obtain an optimized three-dimensional Gaussian splash representation. The optimized three-dimensional Gaussian splash representation is used as a new initial three-dimensional Gaussian splash representation for the next iteration.
[0039] Step C: Processing the optimized three-dimensional Gaussian splatter representation to obtain supervision images of the target object at different perspectives and intermediate consistency trajectory samples corresponding to the target object; the intermediate consistency trajectory samples represent the trajectory image of the target object.
[0040] Specifically, this application uses a three-dimensional vector graphics optimization model to generate target three-dimensional vector graphics. The optimized three-dimensional Gaussian splatter representation is sampled to obtain supervisory images of the target object at different perspectives. The supervisory images are used to guide the generation of the three-dimensional vector graphics. During the sampling process of the optimized three-dimensional Gaussian splatter representation, intermediate consistency trajectory samples corresponding to the target object are generated, where the intermediate consistency trajectory samples represent the trajectory image of the target object.
[0041] Step D: Based on the supervisory image and the intermediate consistency trajectory sample, optimizing the initial three-dimensional vector graphics corresponding to the supervisory image to obtain an optimized three-dimensional vector graphics; the optimized three-dimensional vector graphics is used as a new three-dimensional vector graphics for the next iteration.
[0042] Specifically, based on the supervision image and the intermediate consistent trajectory samples, the initial 3D vector graphics corresponding to the supervision image can be further optimized to obtain an optimized 3D vector graphics. The optimized 3D vector graphics is used as the new 3D vector graphics for the next iteration.
[0043] Step E: Repeat steps A to D. When the number of iterations meets the first target number of iterations, the finally optimized three-dimensional vector graphic is used as the target three-dimensional vector graphic corresponding to the target object.
[0044] Specifically, the first target number of iterations is a preset maximum number of iterative optimizations. If the number of iterations meets the first target number of iterations, i.e., the current number of iterations is greater than or equal to the first target number of iterations, the optimized 3D vector graphic at that time is determined as the final optimized 3D vector graphic, and the final optimized 3D vector graphic is used as the target 3D vector graphic corresponding to the target object.
[0045] Optionally, when the number of iterations meets the first target number of iterations, the optimized three-dimensional Gaussian splattering representation at this time is determined as the final optimized three-dimensional Gaussian splattering representation.
[0046] Optionally, determining the latent variables of the target object corresponding to different camera poses based on the initial three-dimensional Gaussian splatter representation includes: Based on the initial three-dimensional Gaussian splatter representation, first rendered images of the target object at different camera poses and corresponding viewing angles are determined; and each of the first rendered images is sampled to obtain the latent variables of the target object at different camera poses and corresponding viewing angles.
[0047] Specifically, based on the initial three-dimensional Gaussian splash representation, a differentiable renderer is used to render the initial three-dimensional Gaussian splash representation into the first rendered image of the target object at different camera poses corresponding to the perspective, that is, a multi-perspective two-dimensional image; based on the fractional distillation sampling method, the inverse process of the distillation pre-trained diffusion model (DDIM) is applied, and combined with the description text, the first rendered image is sampled to obtain the latent variables corresponding to different diffusion time steps of the target object at different camera poses corresponding to the perspective, that is, the noise latent representation.
[0048] Optionally, optimizing the initial three-dimensional Gaussian splash representation based on each of the latent variables to obtain an optimized three-dimensional Gaussian splash representation includes: Based on each of the latent variables, a first loss value is determined; based on the first loss value, the initial three-dimensional Gaussian splash representation is optimized to obtain an optimized three-dimensional Gaussian splash representation.
[0049] Specifically, based on each latent variable, the update direction of the fractional distillation sampling is calculated; then, based on the calculated update direction, time-dependent weight, and the gradient of the differentiable renderer g with respect to the three-dimensional Gaussian splash representation parameters, the first loss value is calculated; wherein the first loss value is the Interpretative Structural Modeling Method (ISM) loss value; based on the first loss value, the gradient of the first loss value with respect to the three-dimensional Gaussian splash representation parameters is determined. , as shown in formula (1); then the initial three-dimensional Gaussian splash representation is optimized according to the gradient to obtain the optimized three-dimensional Gaussian splash representation, so that the rendering result of the three-dimensional Gaussian splash representation is consistent with the input description text in terms of semantics and geometric structure.
[0050] (1) in, Represents the description text, represents the denoiser of the pre-trained diffusion model, Indicates that over time The weight of the change; and All represent latent variables; represents a three-dimensional Gaussian splatter representation; represents the camera pose, Indicates that over time The expectation of randomly sampled camera poses.
[0051] It should be noted that the ISM loss searches for deterministic trajectories in the latent space of the diffusion model, ensuring the consistency of semantics and geometric structures during the optimization process.
[0052] Optionally, optimizing the initial three-dimensional vector graphics corresponding to the supervisory image based on the supervisory image and the intermediate consistent trajectory samples to obtain the optimized three-dimensional vector graphics includes: Based on the intermediate consistency trajectory samples, the latent variables and the preset update direction, a guidance image of the target object is determined; the projection of the initial three-dimensional vector graphics corresponding to the supervision image is rendered to obtain a second rendered image at a corresponding perspective; based on the guidance image and the second rendered image, a second loss value is calculated; based on the second loss value, the initial three-dimensional vector graphics is optimized to obtain the optimized three-dimensional vector graphics.
[0053] Specifically, based on the intermediate consistency trajectory samples, latent variables and the preset update direction, the guidance image of the target object is determined using formula (2) ; Using a time annealing scheduler, by gradually reducing the sampling range of t, a differentiable vector rasterizer is used to render the projection of the initial 3D vector graphics corresponding to the supervision image obtained based on the optimized 3D Gaussian splash representation, and a second rendered image at the corresponding perspective is obtained; then, based on the guidance image and the second rendered image, the second loss value is calculated using formula (3) ; Among them, the second loss value is obtained based on the learned perceptual image patch similarity (LPIPS) loss value and the contrastive language-image pre-training (CLIP) loss value. The LPIPS loss is used for geometric structure constraints, and the CLIP loss is used for high-level semantic perception similarity. Based on the second loss value, the initial three-dimensional vector graphics are optimized to obtain the optimized three-dimensional vector graphics.
[0054] (2) in, , represents the diffusion coefficient, represents the maximum time step of the pre-trained diffusion model, Based on updated directions from ISM The non-classifier guided construction, and All represent interpolation for bootstrapping without a classifier.
[0055] It should be noted that a larger t usually produces a smoother ; As t decreases, it will Add more details that are consistent with the description text.
[0056] (3) in, represents the cosine distance, represents the expectation of a randomly sampled camera pose, represents a rendered image, Represents three-dimensional vector graphics, Represents the projection of a 3D vector graphic.
[0057] Figure 2 This is the second flow chart of the method for generating three-dimensional vector graphics provided by the present invention, such as Figure 2 As shown, the method includes steps 201 to 213.
[0058] Step 201: Get the description text of the target object.
[0059] Step 202 : determining an initial three-dimensional Gaussian splash representation based on the described text.
[0060] Step 203 : determining first rendered images of the target object at different camera poses and corresponding viewing angles based on the initial three-dimensional Gaussian splash representation.
[0061] Step 204 : sampling each first rendered image to obtain latent variables of the target object at different camera postures corresponding to viewing angles.
[0062] Step 205: Determine a first loss value based on each latent variable.
[0063] Step 206 : Based on the first loss value, optimize the initial three-dimensional Gaussian splash representation to obtain an optimized three-dimensional Gaussian splash representation; the optimized three-dimensional Gaussian splash representation is used as a new initial three-dimensional Gaussian splash representation for the next iteration.
[0064] Step 207 : Process the optimized three-dimensional Gaussian splash representation to obtain supervision images of the target object at different perspectives and intermediate consistency trajectory samples corresponding to the target object.
[0065] Step 208 : Determine a guiding image of the target object based on the intermediate consistent trajectory samples, the latent variables, and the preset update direction.
[0066] Step 209 : Render the projection of the initial three-dimensional vector graphics corresponding to the supervision image to obtain a second rendered image at a corresponding viewing angle.
[0067] Step 210 : Calculate a second loss value based on the guide image and the second rendered image.
[0068] Step 211 : Optimize the initial three-dimensional vector graphics based on the second loss value to obtain an optimized three-dimensional vector graphics.
[0069] Step 212: Determine whether the current number of iterations meets the first target number of iterations. If the current number of iterations does not meet the first target number of iterations, go to step 203; if the current number of iterations meets the first target number of iterations, go to step 213.
[0070] In step 213 , the optimized three-dimensional Gaussian splash representation is used as the target three-dimensional Gaussian splash representation, and the finally optimized three-dimensional vector graphic is used as the target three-dimensional vector graphic corresponding to the target object.
[0071] Figure 3 is one of the schematic diagrams of the target three-dimensional vector graphics provided by the present invention, such as Figure 3 As shown, Figure 3The following are the target 3D vector graphics generated in different camera poses when the description texts are “a flying dragon”, “an alpaca”, “a crab” and “an airplane” respectively.
[0072] Optionally, the method further includes: Based on the projection result of the target three-dimensional vector graphics, a multilayer perceptron is used to determine the average importance of each curve of the target object under different camera postures; based on the average importance, the multilayer perceptron is optimized to obtain a target multilayer perceptron.
[0073] Specifically, in each iteration, based on the projection result of the target 3D vector graphics, a multi-layer perceptron is used to determine the importance of each 3D point on each curve of the target object under different camera postures, and then the importance of all 3D points on the curve is averaged to obtain the average importance of each curve; wherein, the multi-layer perceptron MLP is represented by the importance function f, that is, ,in, Representing a 3D point Based on the average importance, the multi-layer perceptron can be further optimized to obtain the target multi-layer perceptron.
[0074] Optionally, optimizing the multilayer perceptron based on the average importance to obtain a target multilayer perceptron includes: When the average importance is not less than a preset threshold, the curve is added to the importance set; based on the three-dimensional vector graphics corresponding to the curve in the importance set and the optimized three-dimensional Gaussian splash representation, a third loss value is calculated; based on the third loss value, the multilayer perceptron is optimized to obtain an optimized multilayer perceptron; the steps of obtaining the optimized multilayer perceptron are repeated, and when the number of iterations meets the second target number of iterations, the final optimized multilayer perceptron is used as the target multilayer perceptron.
[0075] Specifically, the second target number of iterations is the preset maximum number of iterations of the multilayer perceptron. When the average importance is not less than the preset threshold, the curve is added to the importance set; based on the three-dimensional vector graphics corresponding to the curve in the importance set and the optimized three-dimensional Gaussian splash representation, a third loss value can be calculated; wherein the third loss value is obtained based on the LPIPS loss value and the CLIP loss value. Based on the third loss value, the multilayer perceptron is optimized to obtain an optimized multilayer perceptron; when the number of iterations does not meet the second target number of iterations, the above steps of obtaining the optimized multilayer perceptron are repeated until the number of iterations meets the second target number of iterations. When the number of iterations meets the second target number of iterations, the final optimized multilayer perceptron is used as the target multilayer perceptron.
[0076] Optionally, when the average importance is less than a preset threshold, the curve is added to the non-important set, and the curves in the non-important set are voted on for radial depth visibility to recalibrate the visibility, that is, for the three-dimensional points sampled on the non-important curves, the three-dimensional points are projected to the front image plane seen by the camera and the rear image plane seen by the opposite camera of the camera to obtain the projected points and depths; the depth map obtained by rendering the underlying three-dimensional Gaussian splash representation under the two camera perspectives is queried; by comparing the distance between the depth of the sampling point in the camera space and the queried front surface depth, as well as the distance between the depth of the sampling point in the opposite camera space and the queried rear surface depth, it is determined whether the point is closer to the front surface; if the number of visible sampling points on a curve is greater than the threshold, the curve is determined to be visible. Finally, the final visible curve is obtained by combining the results of importance filtering and depth voting. That is, during the inference process, if the predicted importance of a curve is greater than the threshold or satisfies the depth visibility voting test, the curve is rendered with a fixed higher opacity; otherwise, the curve is rendered with a fixed lower opacity. This achieves a faithful expression of the inherent perspective-dependent occlusion relationship in three-dimensional space and completes the generation of perspective-occlusion-aware vector graphics.
[0077] Figure 4 This is the second schematic diagram of the target three-dimensional vector graphics provided by the present invention, such as Figure 4 As shown, Figure 4 The target 3D vector graphics shown are target 3D vector graphics of the target object in different camera poses obtained by filtering out non-important curves and voting based on depth visibility when the average importance is less than a preset threshold.
[0078] Optionally, the method further comprises the following steps: Step S1, create a unique user identification code for the user; the user identification code is used to identify the user; the creation of the user identification code can be generated by combining the user name and a number. For example, the user abbreviation KF combined with the number 01 is used to generate the user identification code KF01.
[0079] Step S2, obtaining the work number of the designer who processes the 3D vector graphics generation task; wherein, after the management user receives the 3D vector graphics generation task from the user, the management user assigns the 3D vector graphics generation task to the designer for processing.
[0080] Step S3, obtaining the work number of the management user; the work numbers of the management user and the designer are unique numbers within the design company.
[0081] Step S4: Create a unique task code for the user's 3D vector graphics generation task. The task code can be generated by combining the task number and task attributes. The task attributes can mark the attribute characteristics of the task.
[0082] Step S5: generating a user task code based on the user identification code and the task code; specifically, the user identification code and the task code may be concatenated to obtain the user task code.
[0083] Step S6, generating a design code based on the work number of the management user and the work number of the designer; specifically, the work number of the management user and the work number of the designer can be concatenated, or a specific character can be added between the work number of the management user and the work number of the designer, and then concatenated to obtain the unique design code mentioned above.
[0084] Step S7: Generate a marking code for the 3D vector graphics generation task based on the task code and the design code, where the marking code is used to mark the 3D vector graphics generation task.
[0085] In this embodiment, after the above-mentioned user's design requirements are adapted and added to the 3D vector graphics generation model, it is necessary to create a 3D vector graphics generation task and assign a specific designer to handle it. In order to facilitate the subsequent tracking of the task, the above-mentioned 3D vector graphics generation task needs to be marked with a marking code. At the same time, when marking, in order to associate the above-mentioned 3D vector graphics generation task with the user, designer, and management user, when generating the marking code for marking the above-mentioned 3D vector graphics generation task, the user identification code, the designer's work number, the management user's work number, and the task code of the 3D vector graphics generation task can be combined. The above-mentioned information is combined in the above-mentioned marking code for generation, so that the personnel information and task information related to the 3D vector graphics generation task can be easily identified from the above-mentioned marking code; the generation of the above-mentioned marking code is convenient for both tracking the task and subsequent auditing.
[0086] In one embodiment, generating a marking code for a 3D vector graphics generation task based on the task code and the design code includes: S71, determining whether the task code and the design code have the same characters.
[0087] S72, if yes, obtain the number of identical characters and record it as the target number; if not, obtain a preset number as the target number; for example, the preset number is 3.
[0088] S73, obtain a standard Base64 encoding table, and uniformly shift the codes in the Base64 encoding table backward in a circular motion by a preset number of bits to obtain a new Base64 encoding table; wherein the value of the preset number of bits is equal to the value of the target number; it can be understood that the codes arranged at the end of the standard Base64 encoding table are moved to the head of the encoding table.
[0089] S74, concatenating the task code and the design code to obtain a concatenated code.
[0090] S75, using a new Base64 encoding table to encode the concatenated code to obtain a marker code for the 3D vector graphics generation task.
[0091] In this embodiment, it is necessary to use a Base64 encoding table to generate the task's tag code. If a standard Base64 encoding table is used, the tag code can be easily tampered with. Therefore, the standard Base64 encoding table can be rearranged. When rearranging the Base64 encoding table, in order to associate its arrangement with the task code and the design code, the same number of characters in the task code and the design code can be obtained; and based on the numerical value of the same number of characters, the standard Base64 encoding table can be moved and rearranged. It is understandable that even if the Base64 encoding table is moved by only one position, the encoding table will undergo significant changes, and the encoding result will be completely different. Therefore, the new Base64 encoding table obtained by rearrangement is unique, and the concatenated code obtained by concatenating the task code and the design code is also unique and cannot be easily tampered with, thereby ensuring the security of the data.
[0092] In another embodiment, after generating a markup code for the three-dimensional vector graphics generation task based on the task code and the design code, the method further includes: Step S8: converting the format of the design requirement text prompt file into a specific format.
[0093] Step S9: adding a marker code of the three-dimensional vector graphics generation task at a designated position of the converted design requirement text prompt file.
[0094] Step S10: Perform a hash calculation on the design requirement text prompt file to obtain a corresponding standard hash value. The standard hash value and the design requirement text prompt file are stored in a database. An index relationship between the standard hash value and the tag code is established in the database. When the standard hash value in the database is called, the number and time of the call are recorded. The above text prompt file is the text prompt file after the tag code is added.
[0095] In this embodiment, after establishing the index relationship between the hash value and the marking code in the database, the method further includes: Step S11, when task verification is required during the processing of the 3D vector graphics generation task, the design requirement text prompt file stored in the database is obtained, and the mark code corresponding to the task that needs to be verified is obtained, which is recorded as the verification mark code.
[0096] Step S12: performing hash calculation on the design requirement text prompt file stored in the database to obtain a corresponding hash value.
[0097] Step S13: According to the verification mark code, based on the index relationship between the standard hash value and the mark code established in the database, obtain the standard hash value corresponding to the verification mark code.
[0098] Step S14, verify whether the hash value is the same as the standard hash value; if not, verify that the task has been tampered with; if they are the same, obtain the number and time of calls to the standard hash value, and verify whether they are completely consistent with the number and time of task verification; if not completely consistent, determine that the task has been tampered with; if completely consistent, determine that the task has not been tampered with.
[0099] This embodiment also provides a method for determining whether a design task has been tampered with. Specifically, the design requirement text prompt file corresponding to the design task can be converted into a specific format (storage format, content editing format, image attributes, etc.), and a marker code is added to a specified location. A hash calculation is then performed to obtain a corresponding standard hash value, which is used as the basis for subsequent determination of whether the design task has been tampered with.
[0100] When verifying whether the task has been tampered with subsequently, it is only necessary to obtain the design requirement text prompt file stored in the database and perform a hash calculation to obtain the corresponding hash value; then obtain the standard hash value corresponding to the design requirement text prompt file mark code from the database; verify whether the hash value is the same as the standard hash value; if not, verify that the task has been tampered with; if they are the same, obtain the number and time of calls to the standard hash value, and verify whether they are completely consistent with the number and time of task verification; if not completely consistent, determine that the task has been tampered with; if completely consistent, determine that the task has not been tampered with.
[0101] The following describes a device for generating three-dimensional vector graphics provided by the present invention. The device for generating three-dimensional vector graphics described below and the method for generating three-dimensional vector graphics described above can refer to each other.
[0102] Figure 5 This is a schematic diagram of the structure of the device for generating three-dimensional vector graphics provided by the present invention. Figure 5 As shown, the three-dimensional vector graphics generating device 500 includes an acquisition module 501, a first determination module 502 and a generation module 503; wherein, Acquisition module 501, used to obtain the description text of the target object; A first determining module 502 is configured to determine an initial three-dimensional Gaussian splash representation based on the description text; the initial three-dimensional Gaussian splash representation is a three-dimensional visual representation of the target object; The generating module 503 is configured to perform iterative optimization processing on the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object.
[0103] The present invention provides a device for generating 3D vector graphics. The device obtains a descriptive text of a target object; based on the descriptive text, determines an initial 3D Gaussian splatter representation; the initial 3D Gaussian splatter representation is a 3D visual representation of the target object; and iteratively optimizes the initial 3D Gaussian splatter representation to generate a target 3D vector graphic corresponding to the target object. By iteratively optimizing the initial 3D Gaussian splatter representation, the generated target 3D vector graphic can meet the requirements of multi-view consistency and occlusion perception without manual processing, thereby improving the accuracy and efficiency of 3D vector graphics generation.
[0104] Optionally, the generating module 503 is specifically configured to: Step A: In each iteration, based on the initial three-dimensional Gaussian splatter representation, determining a latent variable of the target object corresponding to different camera poses; the latent variable represents the noise potential representation of the target object under the camera pose; Step B: Based on each of the latent variables, optimizing the initial three-dimensional Gaussian splash representation to obtain an optimized three-dimensional Gaussian splash representation; the optimized three-dimensional Gaussian splash representation is used as a new initial three-dimensional Gaussian splash representation for the next iteration; Step C: Processing the optimized three-dimensional Gaussian splatter representation to obtain supervision images of the target object at different perspectives and intermediate consistent trajectory samples corresponding to the target object; the intermediate consistent trajectory samples represent the trajectory image of the target object; Step D: Based on the supervisory image and the intermediate consistency trajectory sample, optimizing the initial three-dimensional vector graphics corresponding to the supervisory image to obtain an optimized three-dimensional vector graphics; the optimized three-dimensional vector graphics is used as a new three-dimensional vector graphics for the next iteration; Step E: Repeat steps A to D. When the number of iterations meets the first target number of iterations, the finally optimized three-dimensional vector graphic is used as the target three-dimensional vector graphic corresponding to the target object.
[0105] Optionally, the generating module 503 is further configured to: Determining, based on the initial three-dimensional Gaussian splatter representation, a first rendered image of the target object corresponding to different camera poses; Each of the first rendered images is sampled to obtain the latent variables of the target object at viewing angles corresponding to different camera postures.
[0106] Optionally, the generating module 503 is further configured to: determining a first loss value based on each of the latent variables; Based on the first loss value, the initial three-dimensional Gaussian splash representation is optimized to obtain an optimized three-dimensional Gaussian splash representation.
[0107] Optionally, the generating module 503 is further configured to: determining a guiding image of the target object based on the intermediate consistent trajectory samples, the latent variable, and a preset update direction; Rendering a projection of the initial three-dimensional vector graphics corresponding to the supervision image to obtain a second rendered image at a corresponding viewing angle; Calculating a second loss value based on the guide image and the second rendered image; Based on the second loss value, the initial three-dimensional vector graphics are optimized to obtain the optimized three-dimensional vector graphics.
[0108] Optionally, the three-dimensional vector graphics generating device 500 further includes: A second determination module is configured to determine, based on the projection result of the target three-dimensional vector graphics, the average importance of each curve of the target object under different camera postures using a multi-layer perceptron; An optimization module is used to optimize the multilayer perceptron based on the average importance to obtain a target multilayer perceptron.
[0109] The optimization module is specifically used to: When the average importance is not less than a preset threshold, adding the curve to the importance set; Calculating a third loss value based on the three-dimensional vector graphics corresponding to the curves in the importance set and the optimized three-dimensional Gaussian splash representation; Optimizing the multilayer perceptron based on the third loss value to obtain an optimized multilayer perceptron; Repeat the above steps to obtain the optimized multilayer perceptron. When the number of iterations meets the second target number of iterations, the finally optimized multilayer perceptron is used as the target multilayer perceptron.
[0110] Figure 6 This is a schematic diagram of the physical structure of an electronic device provided by the present invention, such as Figure 6As shown, the electronic device 600 may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute a method for generating a three-dimensional vector graphic, the method comprising: obtaining a description text of a target object; determining an initial three-dimensional Gaussian splash representation based on the description text; the initial three-dimensional Gaussian splash representation being a three-dimensional visual representation of the target object; and performing iterative optimization processing on the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object.
[0111] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0112] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the three-dimensional vector graphics generation method provided by the above-mentioned methods, which includes: obtaining a description text of the target object; based on the description text, determining an initial three-dimensional Gaussian splash representation; the initial three-dimensional Gaussian splash representation is a three-dimensional visual presentation of the target object; and iteratively optimizing the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphics corresponding to the target object.
[0113] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the three-dimensional vector graphics generation method provided by the above-mentioned methods, the method comprising: obtaining a description text of the target object; based on the description text, determining an initial three-dimensional Gaussian splash representation; the initial three-dimensional Gaussian splash representation is a three-dimensional visual presentation of the target object; and iteratively optimizing the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphics corresponding to the target object.
[0114] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0115] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating three-dimensional vector graphics, characterized in that: include: Get the description text of the target object; Based on the description text, determining an initial three-dimensional Gaussian splatter representation; The initial three-dimensional Gaussian splash representation is a three-dimensional visual representation of the target object; An iterative optimization process is performed on the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object.
2. The method for generating three-dimensional vector graphics according to claim 1, wherein: The iterative optimization processing of the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object includes: Step A: In each iteration, based on the initial three-dimensional Gaussian splatter representation, determining a latent variable of the target object corresponding to different camera poses; the latent variable represents the noise potential representation of the target object under the camera pose; Step B: Based on each of the latent variables, optimizing the initial three-dimensional Gaussian splash representation to obtain an optimized three-dimensional Gaussian splash representation; the optimized three-dimensional Gaussian splash representation is used as a new initial three-dimensional Gaussian splash representation for the next iteration; Step C: Processing the optimized three-dimensional Gaussian splatter representation to obtain supervision images of the target object at different perspectives and intermediate consistent trajectory samples corresponding to the target object; the intermediate consistent trajectory samples represent the trajectory image of the target object; Step D: Based on the supervisory image and the intermediate consistency trajectory sample, optimizing the initial three-dimensional vector graphics corresponding to the supervisory image to obtain an optimized three-dimensional vector graphics; the optimized three-dimensional vector graphics is used as a new three-dimensional vector graphics for the next iteration; Step E: Repeat steps A to D. When the number of iterations meets the first target number of iterations, the finally optimized three-dimensional vector graphic is used as the target three-dimensional vector graphic corresponding to the target object.
3. The method for generating three-dimensional vector graphics according to claim 2, wherein: The determining of the latent variables of the target object corresponding to different camera poses based on the initial three-dimensional Gaussian splatter representation includes: Determining, based on the initial three-dimensional Gaussian splatter representation, a first rendered image of the target object corresponding to different camera poses; Each of the first rendered images is sampled to obtain the latent variables of the target object at viewing angles corresponding to different camera postures.
4. The method for generating three-dimensional vector graphics according to claim 2, wherein: The step of optimizing the initial three-dimensional Gaussian splash representation based on each of the latent variables to obtain an optimized three-dimensional Gaussian splash representation includes: determining a first loss value based on each of the latent variables; Based on the first loss value, the initial three-dimensional Gaussian splash representation is optimized to obtain an optimized three-dimensional Gaussian splash representation.
5. The method for generating three-dimensional vector graphics according to claim 2, wherein: The step of optimizing the initial three-dimensional vector graphics corresponding to the supervisory image based on the supervisory image and the intermediate consistent trajectory samples to obtain the optimized three-dimensional vector graphics includes: determining a guiding image of the target object based on the intermediate consistent trajectory samples, the latent variable, and a preset update direction; Rendering a projection of the initial three-dimensional vector graphics corresponding to the supervision image to obtain a second rendered image at a corresponding viewing angle; Calculating a second loss value based on the guide image and the second rendered image; Based on the second loss value, the initial three-dimensional vector graphics are optimized to obtain the optimized three-dimensional vector graphics.
6. The method for generating three-dimensional vector graphics according to any one of claims 1 to 5, characterized in that: The method further comprises: Based on the projection results of the target three-dimensional vector graphics, a multi-layer perceptron is used to determine the average importance of each curve of the target object under different camera postures; Based on the average importance, the multilayer perceptron is optimized to obtain a target multilayer perceptron.
7. The method for generating three-dimensional vector graphics according to claim 6, wherein: The step of optimizing the multilayer perceptron based on the average importance to obtain a target multilayer perceptron includes: If the average importance is not less than a preset threshold, adding the curve to the importance set; Calculating a third loss value based on the three-dimensional vector graphics corresponding to the curves in the importance set and the optimized three-dimensional Gaussian splash representation; Optimizing the multilayer perceptron based on the third loss value to obtain an optimized multilayer perceptron; Repeat the above steps to obtain the optimized multilayer perceptron. When the number of iterations meets the second target number of iterations, the finally optimized multilayer perceptron is used as the target multilayer perceptron.
8. A device for generating three-dimensional vector graphics, characterized in that: include: The acquisition module is used to obtain the description text of the target object; A first determining module, configured to determine an initial three-dimensional Gaussian splash representation based on the description text; The initial three-dimensional Gaussian splash representation is a three-dimensional visual representation of the target object; A generation module is used to perform iterative optimization processing on the initial three-dimensional Gaussian splash representation to generate a target three-dimensional vector graphic corresponding to the target object.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating three-dimensional vector graphics according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating three-dimensional vector graphics according to any one of claims 1 to 7 is implemented.