A three-dimensional model data processing method, system, product, equipment and medium

By generating edited images to be selected from multiple perspectives during the three-dimensional model editing process and selecting target edited images based on the similarity score, the problem of inconsistency in multiple perspectives is solved, and editing efficiency and quality are improved.

CN118887348BActive Publication Date: 2025-05-02SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411366005.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-05-02
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

During the editing process of three-dimensional models, multiple iterations can easily lead to inconsistent multi-view angles, affecting editing efficiency and quality.

Method used

By determining the source 3D model of the current iteration, we generate edited images to be selected from multiple perspectives, and select target edited images based on the similarity scores, and process the three-dimensional model to ensure the consistency of multiple perspectives.

Benefits of technology

It improves the consistency and iteration efficiency of multi-view angles during the iteration process, reduces the number of iterations, and improves the quality of three-dimensional editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887348B_ABST
    Figure CN118887348B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method, system, product, device and medium for a three-dimensional model, and relates to the field of three-dimensional modeling. In order to solve the problem of inconsistency of multiple perspectives of three-dimensional editing caused by multiple iterative editing, the method includes generating multiple perspectives of selected editing images based on the current single-perspective editing image; for each perspective, obtaining the first similarity score between the selected editing image corresponding to the perspective in the current iteration and the guiding text, obtaining the second similarity score between the selected editing image corresponding to the perspective in the previous iteration and the guiding text, determining the selected editing image corresponding to the larger of the first similarity score and the second similarity score as the target editing image of the perspective; processing the source three-dimensional model according to all the target editing images, and obtaining the edited three-dimensional model corresponding to the current iteration. The present invention can improve the consistency of multiple perspectives and the iteration efficiency in the iterative process, and is helpful to efficiently reconstruct a three-dimensional model that conforms to the text description.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional modeling, and in particular to a data processing method, system, product, equipment and medium for a three-dimensional model. Background Art

[0002] Three-dimensional models play a key role in applications and industries such as augmented / virtual reality, film / game production, and artistic creation. Currently, the most direct and efficient way to create a three-dimensional model is to change the geometry or appearance of an existing three-dimensional model through a three-dimensional editing method based on text description, thereby generating a variety of three-dimensional models. In order to complete the editing of the target editing area while retaining the non-editing area, it is necessary to edit the image of the three-dimensional model in accordance with the text description, and then promote the image editing result to a three-dimensional model to achieve three-dimensional editing. When promoting the image editing result to a three-dimensional model, multiple iterative editing is required. Since only one image is edited in each iterative process, multiple iterations will cause the problem of multi-perspective inconsistency in three-dimensional editing, affecting the editing efficiency and quality of the three-dimensional model.

[0003] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. Summary of the invention

[0004] The purpose of the present invention is to provide a data processing method, system, product, device and medium for a three-dimensional model, which can improve the multi-perspective consistency and iteration efficiency in the iteration process, and help to efficiently reconstruct a three-dimensional model that conforms to the text description.

[0005] In order to solve the above technical problems, the present invention provides a data processing method of a three-dimensional model, comprising:

[0006] Determine a source three-dimensional model corresponding to a current iteration, and obtain a current single-view edited image of the source three-dimensional model edited according to the guidance text;

[0007] Generate multiple perspectives of to-be-selected editing images based on the current single-perspective editing image;

[0008] For each of the perspectives, obtaining a first similarity score between the selected editing image corresponding to the perspective in the current iteration and the guiding text, obtaining a second similarity score between the selected editing image corresponding to the perspective in the previous iteration and the guiding text, and determining the selected editing image corresponding to the larger one of the first similarity score and the second similarity score as the target editing image of the perspective;

[0009] The source three-dimensional model is processed according to all the target edited images to obtain the edited three-dimensional model corresponding to the current iteration.

[0010] The process of obtaining the first similarity score between the selected editing image corresponding to the perspective in the current iteration and the guiding text includes:

[0011] Obtaining a first image title text of the to-be-selected editing image corresponding to the viewing angle in the current iteration;

[0012] A first similarity score between the first image title text and the guidance text is obtained.

[0013] The process of obtaining the first similarity score between the first image title text and the guidance text includes:

[0014] A baseline similarity score between the first image title text and the guidance text is obtained; and the baseline similarity score is adjusted based on feature information in the first image title text and feature information in the guidance text to obtain the first similarity score.

[0015] The process of adjusting the reference similarity score based on the feature information in the first image title text and the feature information in the guidance text to obtain the first similarity score includes:

[0016] If the feature information in the first image title text includes the feature information in the guide text, increasing the baseline similarity score by a first preset score to obtain the first similarity score;

[0017] If the feature information in the first image title text does not include the feature information in the guide text, the reference similarity score is reduced by a second preset score to obtain the first similarity score.

[0018] The process of obtaining the first similarity score between the first image title text and the guidance text includes:

[0019] Obtaining a baseline similarity score between the first image title text and the guidance text;

[0020] Determining whether there is a difference text between the source image text and the guide text in the first image title text;

[0021] If so, the base similarity score is increased by a third preset score to obtain the first similarity score.

[0022] The process of obtaining the second similarity score between the selected editing image corresponding to the perspective in the previous iteration and the guiding text includes:

[0023] Rendering the edited three-dimensional model in the previous iteration based on the viewing angle to obtain a selected editing image corresponding to the viewing angle in the previous iteration;

[0024] Obtaining a second image title text of the image to be selected for editing corresponding to the viewing angle in the previous iteration;

[0025] A second similarity score between the second image title text and the guidance text is calculated.

[0026] Wherein, after the source three-dimensional model is processed according to all the target edited images to obtain the edited three-dimensional model corresponding to the current iteration, the data processing method of the three-dimensional model further includes:

[0027] Determine whether the current iteration satisfies the iteration end condition;

[0028] If not, the edited three-dimensional model is used as the source three-dimensional model of the next iteration for the next iteration;

[0029] If so, the edited three-dimensional model is used as a target three-dimensional model that conforms to the description of the guidance text.

[0030] The process of processing the source three-dimensional model according to all the target edited images to obtain the edited three-dimensional model corresponding to the current iteration includes:

[0031] Constructing a hybrid loss function, wherein the hybrid loss function includes an editing loss and a reconstruction loss;

[0032] The source three-dimensional model is processed based on the mixed loss function and all the target edited images to obtain an edited three-dimensional model corresponding to the current iteration.

[0033] Among them, the mixed loss function is ;

[0034] Wherein, L is the value of the hybrid loss function, α is the first hyperparameter corresponding to the editing loss, β is the second hyperparameter corresponding to the reconstruction loss, and L edit is the editing loss, L recon To rebuild the losses.

[0035] Wherein, the data processing method of the three-dimensional model further includes:

[0036] Obtaining a similarity score between each of the target editing images and the guiding text;

[0037] The first hyperparameter and the second hyperparameter are adjusted according to the sum of all the similarity scores.

[0038] The process of adjusting the first hyperparameter and the second hyperparameter according to the sum of all the similarity scores includes:

[0039] If the sum of the similarity scores is greater than a preset semantic value, adjusting the first super parameter to be less than the second super parameter; the preset semantic value is determined according to the total number of the perspectives;

[0040] If the sum of the similarity scores is less than or equal to the preset semantic value, the first hyperparameter is adjusted to be equal to the second hyperparameter.

[0041] Wherein, the data processing method of the three-dimensional model further includes:

[0042] The editing loss is calculated based on the first relational expression, which is: ,in, , ;

[0043] The reconstruction loss is calculated based on the second relation, which is: ;

[0044] Wherein, i is the index of the viewing angle, E img is the image encoder, E txt For text encoder, represents the cosine similarity, I t For the edited image, I S is the source image, T t is the edited text, T S is the source text, is the candidate editing image of the i-th perspective, is the image of the i-th perspective obtained by rendering the source 3D model, The difference between the images before and after editing. is the text difference before and after editing, and M is the total number of perspectives.

[0045] The process of obtaining the current single-view edited image of the source three-dimensional model edited according to the guiding text includes:

[0046] Determine the selected perspective for the current iteration;

[0047] Acquire a current single-view image of the source three-dimensional model at the selected view angle;

[0048] The current single-view image is edited according to the guide text to obtain a current single-view edited image.

[0049] The process of obtaining the current single-view image of the source three-dimensional model at the selected view angle includes:

[0050] Based on the selected viewing angle, determining a query point corresponding to each pixel point in the three-dimensional space of the source three-dimensional model;

[0051] For each query point, obtaining the color of the pixel point corresponding to the query point according to a three-dimensional Gaussian function and a distance between the query point and a center point of the source three-dimensional model;

[0052] A current single-view image of the source three-dimensional model at the selected viewing angle is obtained based on the colors of all the pixel points.

[0053] The process of generating multiple perspectives of selected editing images based on the current single-perspective editing image includes:

[0054] The single-view editing image is input into a dense-view generator to obtain a dense-view candidate editing image.

[0055] Wherein, the data processing method of the three-dimensional model further includes:

[0056] Establishing a dense view generation network; the dense view generation network includes a video generation diffusion network and a view adaptation module;

[0057] Acquire a first training data set obtained by rotating a camera around a target object based on a fixed trajectory and a second training data set obtained by rotating a camera around the target object based on a random trajectory; the target object is an object in the source three-dimensional model;

[0058] The dense perspective generation network is trained using the first training data set. When the number of training times reaches a preset number, the video generation diffusion network is fixed, and the perspective adaptation module is trained using the second training data set until the training requirements are met, thereby obtaining the dense perspective generator.

[0059] Wherein, the data processing method of the three-dimensional model further includes:

[0060] Sampling the fixed trajectory to obtain a plurality of sampling points;

[0061] For each of the sampling points, adding noise to the azimuth angle corresponding to the sampling point to obtain a new azimuth angle, and / or adding noise to the elevation angle corresponding to the sampling point to obtain a new elevation angle;

[0062] At least one random trajectory is obtained according to all the new azimuth angles and / or the new elevation angles.

[0063] In order to solve the above technical problems, the present invention also provides a three-dimensional model data processing system, comprising:

[0064] A first determination module is used to determine a source three-dimensional model corresponding to a current iteration, and obtain a current single-view edited image of the source three-dimensional model edited according to the guidance text;

[0065] A first generating module, configured to generate multiple perspectives of to-be-selected editing images based on the current single-perspective editing image;

[0066] A first acquisition module is used to acquire, for each of the perspectives, a first similarity score between the selected editing image corresponding to the perspective in the current iteration and the guiding text, acquire a second similarity score between the selected editing image corresponding to the perspective in the previous iteration and the guiding text, and determine the selected editing image corresponding to the larger one of the first similarity score and the second similarity score as the target editing image of the perspective;

[0067] The editing and reconstruction module is used to process the source three-dimensional model according to all the target editing images to obtain the edited three-dimensional model corresponding to the current iteration.

[0068] In order to solve the above technical problems, the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the three-dimensional model data processing method described in any one of the above items.

[0069] In order to solve the above technical problems, the present invention further provides an electronic device, comprising:

[0070] Memory for storing computer programs;

[0071] A processor is used to implement the steps of any of the above three-dimensional model data processing methods when executing the computer program.

[0072] In order to solve the above technical problems, the present invention also provides a non-volatile storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data processing method of the three-dimensional model described in any one of the above items are implemented.

[0073] The present invention provides a data processing method for a three-dimensional model. First, a current single-view editing image obtained by rendering a source three-dimensional model at any viewing angle is used to generate selected editing images of multiple viewing angles. Then, the three-dimensional model is edited based on the selected editing images of multiple viewing angles, which can ensure multi-view consistency and thus reduce the number of iterations. A target editing image that best meets the description of a guiding text is selected from the selected editing images in the current iteration process and the edited images corresponding to the previous iteration to reconstruct the source three-dimensional model, thereby achieving a smooth transition between the two iteration processes and improving the three-dimensional editing quality, thereby improving the multi-view consistency and iteration efficiency in the iteration process, and facilitating efficient reconstruction of a three-dimensional model that meets the text description.

[0074] The present invention also provides a three-dimensional model data processing system, a computer program product, an electronic device and a non-volatile storage medium, which have the same beneficial effects as the above-mentioned three-dimensional model data processing method. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0076] Figure 1 A flowchart of the steps of a three-dimensional model editing method provided by the present invention;

[0077] Figure 2 A schematic diagram of a source three-dimensional model of a toy dog ​​provided by the present invention;

[0078] Figure 3 A schematic diagram of a single-view image obtained by rendering a source three-dimensional model provided by the present invention at a certain viewing angle;

[0079] Figure 4 A schematic diagram of a single-view editing image provided by the present invention;

[0080] Figure 5 A schematic diagram of generating a generalized dense perspective image provided by the present invention;

[0081] Figure 6 A flowchart of editing a three-dimensional model based on dense perspective generation provided by the present invention;

[0082] Figure 7 A schematic diagram of a three-dimensional model editing interactive interface provided by the present invention;

[0083] Figure 8 A schematic diagram of the structure of a three-dimensional model editing system provided by the present invention;

[0084] Fig. 9 A schematic diagram of the structure of an electronic device provided by the present invention;

[0085] Fig.10 A schematic diagram of the structure of a non-volatile storage medium provided by the present invention. DETAILED DESCRIPTION

[0086] The core of the present invention is to provide a data processing method, system, product, equipment and medium for a three-dimensional model, which can improve the multi-perspective consistency and iteration efficiency in the iteration process, and help to efficiently reconstruct a three-dimensional model that conforms to the text description.

[0087] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0088] First, please refer to Figure 1 The present invention provides a three-dimensional model data processing method, comprising:

[0089] S101: Determine a source 3D model corresponding to a current iteration, and obtain a current single-view edited image of the source 3D model edited according to a guidance text;

[0090] In this embodiment, it is considered that in the three-dimensional editing based on text description, multiple iterations can be performed to gradually refine the three-dimensional model and improve the precision and realism of the three-dimensional model. For the convenience of explanation, this embodiment takes any one of the iterations as an example to illustrate the data processing method of the three-dimensional model, and the data processing methods of the three-dimensional models of other iterations are similar. For the current iteration, if the current iteration is the first iteration, the source three-dimensional model corresponding to the current iteration is the initial three-dimensional model, that is, the three-dimensional model that has not been iteratively processed. If the current iteration is not the first iteration, then the source three-dimensional model corresponding to the current iteration is the edited three-dimensional model corresponding to the previous iteration. Each iteration is based on the result of the previous iteration, which ensures the continuity of the editing operation and reduces the repeated workload. Through step-by-step iteration, the three-dimensional model can be adjusted and optimized more accurately. Exemplarily, assuming that the current iteration is the second iteration, the source three-dimensional model corresponding to the second iteration is the edited three-dimensional model corresponding to the first iteration.

[0091] In an exemplary embodiment, the process of obtaining a current single-perspective edited image of a source three-dimensional model edited according to a guidance text includes: determining a selected perspective of a current iteration, obtaining a current single-perspective image of the source three-dimensional model at the selected perspective, and editing the current single-perspective image according to the guidance text to obtain the current single-perspective edited image.

[0092] It can be understood that first, a perspective is selected as the designated perspective of the current iteration. The designated perspective can be randomly selected, and the two-dimensional image (single-perspective image) of the source three-dimensional model at the designated perspective is edited in accordance with the description of the guidance text to generate the current single-perspective editing image corresponding to the current iteration, so as to facilitate the precise adjustment and optimization of the detail features at the perspective according to the guidance text in the current iteration, thereby improving the overall refinement of the model editing.

[0093] Specifically, the process of obtaining the current single-view image of the source three-dimensional model at a selected viewing angle includes: based on the selected viewing angle, determining the query point corresponding to each pixel point in the three-dimensional space of the source three-dimensional model; for each query point, obtaining the color of the pixel point corresponding to the query point according to the three-dimensional Gaussian function and the distance between the query point and the center point of the source three-dimensional model; and obtaining the current single-view image of the source three-dimensional model at the selected viewing angle based on the colors of all pixel points.

[0094] In this embodiment, the edited three-dimensional model g is represented by a three-dimensional Gaussian. s ={g 1 , …, g n}, where the i-th three-dimensional Gaussian point g i ={μ, λ, c, a}, i∈{1,…,N}, Indicates the location of the center point, Represents the covariance matrix, and the color of each pixel is given by Indicates that transparency , represents the three-dimensional real space, represents the seven-dimensional real space, then the three-dimensional Gaussian representation is:

[0095] ;

[0096] Where x represents the distance between the query point and the center point.

[0097] Then, the color C of a certain pixel of the single-view image Is of the source 3D model obtained based on rendering is:

[0098] , ;c i Represents the color vector of the i-th three-dimensional Gaussian point, a i represents the transparency of the i-th three-dimensional Gaussian point, G(x) represents the Gaussian function, Indicates an intermediate parameter.

[0099] The current single-view image Is of the source 3D model can be generated based on the colors of all pixels. Figure 2 and Figure 3 As shown, Figure 2 A source three-dimensional model including a toy dog ​​in this embodiment is shown. Figure 3 A single-view image of a toy dog ​​rendered from the source three-dimensional model at a certain viewing angle is shown.

[0100] Then, the current single-view image is edited to obtain a single-view edited image that meets the description of the instruction text. According to the input edited text, the single-view image of the source 3D model is edited using the instruction text and the image editing method InstructPix2Pix to obtain a single-view edited image that meets the description of the instruction text. Assuming that the instruction text is a toy dog ​​with a ball under its feet, the single-view edited image can refer to Figure 4 shown.

[0101] S102: generating multiple perspectives of selected editing images based on the current single-perspective editing image;

[0102] In this embodiment, after obtaining a single-perspective editing image, multiple perspectives of selected editing images are generated based on the single-perspective editing image, and the editing of the three-dimensional model is completed based on the multiple perspectives of selected editing images in one iteration, which facilitates improving the consistency of model details under different perspectives, thereby improving iteration efficiency.

[0103] As an optional embodiment, the multiple perspectives may be dense perspectives. Dense perspectives generally refer to a series of perspectives distributed at a high density within a certain range. These perspectives are very close to each other, forming a continuous perspective coverage, so as to be able to capture and display the appearance and details of the three-dimensional model at different angles in detail. Dense perspectives provide a smooth transition from one perspective to another. After obtaining the candidate edit images of dense perspectives, a dense perspective image set corresponding to the current iteration is generated based on all the candidate edit images.

[0104] S103: for each perspective, obtaining a first similarity score between the candidate editing image corresponding to the perspective in the current iteration and the guiding text, obtaining a second similarity score between the candidate editing image corresponding to the perspective in the previous iteration and the guiding text, and determining the candidate editing image corresponding to the larger one of the first similarity score and the second similarity score as the target editing image of the perspective;

[0105] In this embodiment, considering that in the process of obtaining edited images of multiple perspectives through multiple iterations, there will be editing inconsistencies in multiple iterations due to the lack of semantics, perspectives, etc. constraints between multiple iterations, and multiple candidate edited images in the dense perspective image set corresponding to the current iteration will have semantic drift or multi-perspective editing inconsistencies, and may also cause semantic gaps with the dense perspective image set generated in the previous iteration, further leading to training shocks and slowing down optimization efficiency. Based on this, this embodiment first selects the candidate edited image I corresponding to the zth perspective in the dense perspective image set in the current iteration. z,b , determine the candidate editing image I z,b The first similarity score with the guidance text is used to characterize the candidate edited image I z,bThe matching degree between the selected and guided text is then used to render the source 3D model corresponding to the current iteration based on the perspective to obtain the candidate edited image I in the previous iteration. z,b-1 , determine the candidate editing image I z,b-1 The second similarity score with the guidance text is used to characterize the selected editing image I z,b-1 Determine the larger value of the first similarity score and the second similarity score, and if the first similarity score is greater than the second similarity score, select the to-be-selected editing image I z,b As the target edited image of the z-th perspective, if the first similarity score is less than the second similarity score, the candidate edited image I is selected. z,b-1 As the target edited image of the z-th perspective, if the first similarity score is equal to the second similarity score, the candidate edited image I is selected. z,b As the target editing image of the z-th perspective. It can be understood that in two adjacent iterations, the candidate editing image with a higher similarity score is selected as the target editing image for each angle. On the one hand, it can ensure that the selected image best matches the guidance text description, and on the other hand, it can ensure that the model changes between two adjacent iterations are smooth and seamless, thereby improving the accuracy and reliability of model editing, thereby improving iteration efficiency.

[0106] S104: Processing the source three-dimensional model according to all target edited images to obtain an edited three-dimensional model corresponding to the current iteration.

[0107] In this embodiment, the source three-dimensional model is optimized using all target edited images obtained in S103 to obtain the edited three-dimensional model corresponding to the current iteration. It can be understood that if the iteration is not completed, the edited three-dimensional model corresponding to the current iteration is the source three-dimensional model corresponding to the next iteration.

[0108] It can be seen that in this embodiment, the current single-perspective editing image obtained by rendering the source three-dimensional model at any perspective is first used to generate candidate editing images of multiple perspectives, and then the three-dimensional model is edited based on the candidate editing images of multiple perspectives, which can ensure multi-perspective consistency and thus reduce the number of iterations. The target editing image that best meets the description in the guidance text is selected from the candidate editing images in this iteration and the edited images corresponding to the previous iteration to reconstruct the source three-dimensional model, thereby achieving a smooth transition between the two iterative processes and improving the three-dimensional editing quality, thereby improving the multi-perspective consistency and iteration efficiency in the iterative process, and facilitating the efficient reconstruction of a three-dimensional model that meets the text description.

[0109] Based on the above embodiments:

[0110] In an exemplary embodiment, the process of obtaining the first similarity score between the selected editing image and the guiding text corresponding to the perspective in the current iteration includes:

[0111] Obtain the first image title text of the image to be selected for editing corresponding to the viewing angle in the current iteration;

[0112] A first similarity score between the first image title text and the guidance text is obtained.

[0113] In this embodiment, when determining the edited image I of the zth viewing angle in the current iteration, z,b When the matching degree between the selected editing image and the guiding text is determined, the editing image I is first obtained. z,b Specifically, the first image title text is obtained based on a preset image title acquisition method to obtain the selected editing image I z,b The first image title text is obtained, wherein the preset image title acquisition method includes but is not limited to the BLIP2 (Behavioral Language-Image Pre-trained 2, a large-scale pre-trained model for computer vision and natural language processing tasks) model. Then the first similarity score between the first image title text and the guidance text is calculated. The first similarity score can reflect the degree of semantic association between the first image title text and the guidance text. The higher the similarity, the closer the first image title text and the guidance text are in semantics, that is, the more similar or related the contents they describe, that is, the more the selected editing image corresponding to the first image title text conforms to the description of the guidance text.

[0114] In an exemplary embodiment, the process of obtaining a first similarity score between the first image title text and the guidance text includes:

[0115] A reference similarity score between the first image title text and the guidance text is obtained; and the reference similarity score is adjusted based on feature information in the first image title text and feature information in the guidance text to obtain a first similarity score.

[0116] Specifically, when obtaining the first similarity score between the first image title text and the guidance text, the baseline similarity score between the first image title text and the guidance text is first calculated based on a preset similarity algorithm. The preset similarity algorithm includes but is not limited to a cosine similarity algorithm, a Jaccard similarity algorithm, etc., which can be selected according to actual engineering needs and is not limited in this embodiment.

[0117] After obtaining the baseline similarity score, the baseline similarity score is adjusted according to the feature information in the first image title text and the feature information in the guide text to obtain the final first similarity score. It can be understood that the baseline similarity score is based on a direct comparison of the text content, and the adjustment of the feature information can take into account the important parts or key information in the text, thereby improving the accuracy of the similarity calculation. In addition, through the adjustment of the feature information, the context information of the text can be more comprehensively considered, which is suitable for more complex and specific application scenarios.

[0118] As an optional embodiment, a text encoder may be used to encode the prompt text and the first image title text respectively, and the similarity between the two may be calculated using a cosine function, and the similarity score may be transformed into a value range of 0-10 by multiplying by 10, as a benchmark similarity score. By multiplying the similarity score by 10 and scaling it to a range of 0-10, a standardized processing of the similarity score may be achieved, and the standardized processing facilitates the comparison of similarities between different texts, so that they are all within the same value range, which is convenient for subsequent comparison and analysis, and can improve the efficiency of calculation and the quality of decision-making.

[0119] In an exemplary embodiment, the process of adjusting the reference similarity score based on the feature information in the first image title text and the feature information in the guidance text to obtain the first similarity score includes:

[0120] If the feature information in the first image title text includes the feature information in the guidance text, increasing the baseline similarity score by a first preset score to obtain a first similarity score;

[0121] If the feature information in the first image title text does not include the feature information in the guidance text, the reference similarity score is reduced by a second preset score to obtain a first similarity score.

[0122] In this embodiment, if the first image title text includes the characteristic information in the guidance text, the first preset score is correspondingly increased to the reference similarity score to obtain the first similarity score. If the first image title text does not include the characteristic information in the guidance text, the first similarity score is reduced by the second preset score. The first preset score and the second preset score can be the same or different, and can be set according to actual project needs.

[0123] For example, if the first image title text contains the correct feature information in the guidance text, the baseline similarity score should be considered to be increased by 1 point, and if the first image title text does not contain the correct feature information in the guidance text, the baseline similarity score should be considered to be reduced by 0.5 points. As another optional embodiment, if more feature information is included, multiple first preset scores can be superimposed on the baseline similarity score.

[0124] In an exemplary embodiment, the process of obtaining a first similarity score between the first image title text and the guidance text includes:

[0125] Obtaining a baseline similarity score between the first image title text and the guidance text;

[0126] Determining whether there is a difference text between the source image text and the guidance text in the first image title text;

[0127] If so, the base similarity score is increased by a third preset score to obtain a first similarity score.

[0128] In this embodiment, the extent to which the first image title text covers the guidance text and the extent to which the difference information between the guidance text and the source text (i.e., the image title text of the single-view image) is covered is also considered. If the first image title text contains information that does not appear in the guidance text, it should not be considered as a poor prediction. Evaluate the first image title text. If the first image title text covers the difference text between the source text and the guidance text prompt, it should be considered to increase the baseline similarity score by 1.

[0129] The following is an example of the guidance text, source text and first image title text. The guidance text may be a stuffed dog wearing sunglasses, the source text is a stuffed dog, and the first image title text is a yellow stuffed dog wearing red-rimmed sunglasses sitting on a table. By scoring them, the first similarity score can be 8.

[0130] In an exemplary embodiment, the process of obtaining the second similarity score between the candidate editing image and the guiding text corresponding to the perspective in the previous iteration includes:

[0131] Rendering the edited 3D model in the previous iteration based on the view angle to obtain a candidate edited image corresponding to the view angle in the previous iteration;

[0132] Obtain the second image title text of the image to be selected for editing corresponding to the viewing angle in the previous iteration;

[0133] A second similarity score between the second image title text and the guidance text is calculated.

[0134] In this embodiment, the edited 3D model in the previous iteration is rendered based on the z-th perspective to obtain the selected edited image I corresponding to the z-th perspective in the previous iteration. z,b-1 , in determining the candidate edited image I of the z-th perspective in the previous iteration z,b-1 When the matching degree between the selected editing image and the guiding text is determined, the editing image I is first obtained. z,b-1Specifically, based on a preset image title acquisition method, the image to be edited I is acquired. z,b-1 The second image title text is obtained, and then the second similarity score between the second image title text and the guidance text is calculated. The second similarity score can reflect the degree of semantic association between the second image title text and the guidance text. The higher the similarity, the closer the second image title text and the guidance text are semantically, that is, the more similar or related the contents they describe, that is, the more the selected editing image corresponding to the second image title text conforms to the guidance text description.

[0135] The process of calculating the second similarity score between the second image title text and the guidance text is the same as the process of calculating the first similarity score between the first image title text and the guidance text, and will not be repeated here. After the first similarity score between the selected editing image and the guidance text corresponding to each perspective in the dense perspective image set of the current iteration is calculated, the result is obtained. ,in , is the first similarity score between the candidate editing image and the guidance text corresponding to the first perspective, is the first similarity score between the candidate editing image and the guidance text corresponding to the Mth perspective, and M is the total number of perspectives.

[0136] After the second similarity scores between the candidate editing image and the guidance text corresponding to each perspective in the dense perspective image set of the previous iteration are calculated, we get ,in , is the second similarity score between the candidate editing image and the guidance text corresponding to the first perspective, is the second similarity score between the candidate editing image and the guidance text corresponding to the Mth perspective. , compare the similarity scores and , the image with higher semantic similarity score in the corresponding perspective i is taken as the target editing image corresponding to perspective i, and the target editing images of M perspectives constitute the editing image set.

[0137] As another optional embodiment, the above-mentioned process of obtaining the first similarity score between the selected edited image and the guiding text at each viewing angle in the dense viewing angle image set of the current iteration, and the process of obtaining the second similarity score between the selected edited image and the guiding text at each viewing angle corresponding to the previous iteration, can be implemented by a large prediction model. The large language model can imitate the capabilities of human experts in data annotation and evaluation. This embodiment proposes a semantic similarity calculation method based on a large language model. In order to evaluate the semantic similarity between the generated selected edited image and the guiding text, the large language model is called using the guiding text to calculate the similarity between the image title text of the selected edited image and the guiding text, and the similarity between the image title text of the selected edited image and the source text corresponding to the source three-dimensional model. It can be understood that the input of the large language model mainly includes the first image title text, the second image title text, the guidance text and the source text, and the output of the large prediction model mainly includes the first similarity score and the second similarity score. During the iteration process, the use of the large language model can quickly provide similarity scores and speed up the iteration speed of three-dimensional model editing. Moreover, for a large number of dense perspective image sets, the large language model can process similarity calculations on a large scale and improve processing capabilities.

[0138] In an exemplary embodiment, after the source three-dimensional model is processed according to all target edited images to obtain the edited three-dimensional model corresponding to the current iteration, the three-dimensional model data processing method further includes:

[0139] Determine whether the current iteration meets the iteration end condition;

[0140] If not, the edited 3D model is used as the source 3D model of the next iteration for the next iteration;

[0141] If so, the edited three-dimensional model is used as the target three-dimensional model that meets the description of the guidance text.

[0142] In this embodiment, considering that multiple iterations are required to obtain the target three-dimensional model that meets the description of the guidance text, it is necessary to determine whether the current iteration meets the iteration end condition in this embodiment. If the iteration end condition is not met, the edited three-dimensional model is used as the source three-dimensional model for the next iteration. If the iteration end condition is met, the edited three-dimensional model of the current iteration is directly output as the target three-dimensional model. In this embodiment, the determination of the iteration end condition avoids unnecessary excessive iterations, saves computing resources and time, and improves editing efficiency. On the other hand, it can ensure that the output three-dimensional model meets the description of the guidance text.

[0143] In an exemplary embodiment, the process of processing the source 3D model according to all target edited images to obtain the edited 3D model corresponding to the current iteration includes:

[0144] Construct a hybrid loss function, which includes editing loss and reconstruction loss;

[0145] The source 3D model is processed based on the hybrid loss function and all target edited images to obtain the edited 3D model corresponding to the current iteration.

[0146] In order to ensure that the edited 3D model conforms to the description of the guidance text and ensure the editing quality, this embodiment constructs a hybrid loss function based on the editing loss and the reconstruction loss, wherein the editing loss can be used to characterize the difference between the edited 3D model and the editing target, and the reconstruction loss can be used to characterize the difference between the edited 3D model and the source 3D model. In this embodiment, the hybrid loss function value can be obtained by rendering all target edited images and the source 3D model to obtain images of corresponding perspectives and the hybrid loss function, and the parameters in the source 3D model are adjusted based on the hybrid loss function value to obtain the edited 3D model.

[0147] Specifically, the edited 3D model is still represented by a 3D Gaussian, and its structure is referred to above, and this embodiment will not be repeated here. It can be understood that if the current iteration is the first iteration, before the source 3D model is processed based on the mixed loss function and all target edited images, the edited 3D model is also initialized.

[0148] Among them, the mixed loss function is ;

[0149] Where L is the value of the hybrid loss function, α is the first hyperparameter corresponding to the editing loss, β is the second hyperparameter corresponding to the reconstruction loss, and L edit is the editing loss, L recon To rebuild the losses.

[0150] It can be understood that by adjusting the first hyperparameter and the second hyperparameter, the importance of the editing goal and the reconstruction goal can be balanced, and it is convenient to adjust the parameters according to different editing tasks and model characteristics to optimize specific goals, thereby achieving high-quality editing effects while maintaining the original characteristics of the model.

[0151] In an exemplary embodiment, the data processing method of the three-dimensional model further includes:

[0152] The editing loss is calculated based on the first relation, which is ,in, , ;

[0153] The reconstruction loss is calculated based on the second relation, which is ;

[0154] Among them, i is the number of the viewing angle, E img is the image encoder, E txt For text encoder, represents the cosine similarity, I t For the edited image, I S is the source image, T t is the edited text, T S is the source text, is the candidate editing image of the i-th perspective, is the image of the i-th perspective obtained by rendering the source 3D model, The difference between the images before and after editing. is the text difference before and after editing, and M is the total number of perspectives.

[0155] The editing loss is explained. This embodiment is based on the first relation: To calculate the editing loss, , It can be understood that the editing loss is calculated by comparing the difference between the image before and after editing and the corresponding text. Specifically, the source image I S and a text T describing the source image S , get the edited image I t and text describing the edited image t , using the image encoder E img The source image Is and the edited image T are S Encode and get the feature vector E img (I s ) and E img (I t ), use the text editor E txt For source text T S and the edited text T t Encode and get the feature vector E txt (T s ) and E txt (T t ). Then calculate the difference ΔI between the feature vectors of the edited image and the source image, and the difference ΔT between the feature vectors of the edited text and the source text. For images, the image difference , for text, text differences After calculating the image difference and text difference, calculate the cosine similarity between the image difference and the text difference , and then the editing loss can be calculated using the first relation.

[0156] The reconstruction loss is described. This embodiment is based on the second relation: To calculate the reconstruction loss, it can be understood that the reconstruction loss is calculated by the image difference under different perspectives. Specifically, for each perspective i, the candidate edited image of the i-th perspective is and the image of the i-th perspective obtained by rendering the source 3D model The squares of the differences between them are summed and then divided by the total number of viewpoints M to obtain the average error as the reconstruction loss.

[0157] In an exemplary embodiment, the data processing method of the three-dimensional model further includes:

[0158] Obtaining the similarity score between each target edited image and the guidance text;

[0159] The first hyperparameter and the second hyperparameter are adjusted according to the sum of all similarity scores.

[0160] In this embodiment, considering that the target editing image of the i-th perspective may be the selected editing image corresponding to the current iteration or the selected editing image corresponding to the previous iteration, therefore, when the target editing image is the selected editing image corresponding to the current iteration, the similarity score here is the first similarity score, and when the target editing image is the selected editing image corresponding to the previous iteration, the similarity score here is the second similarity score. In this embodiment, adjusting the first hyperparameter and the second hyperparameter according to the semantic similarity between the selected editing image and the guiding text can better meet the editing target, ensuring that the similarity between the editing image and the editing target is gradually improved while maintaining the reconstruction quality.

[0161] In an exemplary embodiment, the process of adjusting the first hyperparameter and the second hyperparameter according to the sum of all similarity scores includes:

[0162] If the sum of the similarity scores is greater than a preset semantic value, adjusting the first hyperparameter to be less than the second hyperparameter; the preset semantic value is determined according to the total number of perspectives;

[0163] If the sum of the similarity scores is less than or equal to the preset semantic value, the first hyperparameter is adjusted to be equal to the second hyperparameter.

[0164] It can be understood that if the sum of the similarity scores of all perspectives is greater than the preset semantic value, it indicates that the semantic similarity between the image set created by the target edited image in the current iteration and the guidance text is high, that is, the target edited image is closer to the content described in the guidance text. In this case, the reconstruction loss L reconIt should dominate, because the edited image already reflects the editing target well, and the reconstruction loss helps to maintain the consistency of the edited image with the source 3D model, ensuring the quality and stability of the editing result. On the contrary, if the sum of the similarity scores of all perspectives is less than or equal to the preset semantic value, it indicates that the semantic similarity between the image set created by the target edited image in the current iteration and the guidance text is low, and there is a large difference between the edited image and the editing target. In this case, the editing loss L should be increased. edit The weights of are used to better guide the editing process and ensure that the editing result is closer to the editing target. Of course, the reconstruction loss still needs to be maintained to keep the edited image consistent with the source 3D model.

[0165] In this embodiment, by dynamically adjusting the values ​​of α and β, real-time control of the editing process can be achieved, ensuring that the similarity between the edited image and the edited target is gradually improved while maintaining the reconstruction quality. This method makes the editing process more flexible and intelligent, and can automatically adjust the optimization strategy according to the current editing effect and semantic similarity score, thereby improving the quality and accuracy of the editing results.

[0166] In an exemplary embodiment, the process of generating multiple perspectives of selected editing images based on the current single-perspective editing image includes:

[0167] The single-view editing image is input into the dense view generator to obtain the candidate editing image of dense view.

[0168] In this embodiment, a dense perspective generator is pre-established, and a single-perspective editing image is input into the dense perspective generator, so that multiple perspectives of candidate editing images can be quickly generated, and there is no need to perform separate perspective rendering for each single-perspective editing image, thereby improving editing efficiency. Generating candidate editing images under different perspectives can better evaluate the performance of a single-perspective editing image under different perspectives, ensure perspective consistency, reduce the number of iterations, and speed up the editing process.

[0169] In an exemplary embodiment, the data processing method of the three-dimensional model further includes:

[0170] Establish a dense view generation network; the dense view generation network includes a video generation diffusion network and a view adaptation module;

[0171] Acquire a first training data set obtained by rotating the camera around the target object for one cycle based on a fixed trajectory and a second training data set obtained by rotating the camera around the target object for one cycle based on a random trajectory; the target object is an object in the source three-dimensional model;

[0172] The dense view generation network is trained using the first training data set. When the training times reach a preset number, the video is fixed to generate the diffusion network, and the view adaptation module is trained using the second training data set until the training requirements are met to obtain a dense view generator.

[0173] In this embodiment, the process of establishing a dense view generator is described. First, a dense view generation network with adaptive view is constructed to optimize the traditional video diffusion model so that it can adapt to different view inputs. The network structure follows the three-dimensional UNet (U-shaped Network) structure of the video diffusion model. Each layer is a spatial convolution layer, a temporal convolution layer, a self-attention layer, and a cross-attention layer. Add a view adaptation module, that is, add a LoRA (Low-Rank Adaptation) layer to each cross-attention layer to construct a dense view generation model, such as Figure 5 shown.

[0174] On a fixed trajectory, the camera is controlled to rotate around the object of the source 3D model at a regularly spaced azimuth angle at the same elevation angle as the single-view edited image. The fixed trajectory is represented by S = (s 1 ,…,s g ), images at each azimuth angle can be obtained, and the first training data set is constructed based on the images at each azimuth angle on the fixed trajectory. Generally, 18 azimuth angles can be selected. Considering that the fixed trajectory cannot obtain information about the top or bottom of the object, this embodiment also sets a random trajectory. The spacing of the azimuth angles of the random trajectory is irregular, and the elevation angle of each observation is also randomly selected. The random trajectory is expressed as D=(d 1 , …, d g ), and then acquire images at different azimuths and elevations based on each random trajectory to obtain a second training data set. The constructed generalized dense view generation network (including the video generation diffusion network and the view adaptation module) is trained using the first training data set based on the fixed trajectory to reach a certain number of training times. Then, the fixed video generation diffusion network part remains unchanged, and the view adaptation module is trained with the second training data set generated by the random trajectory, thereby ensuring that the editing adaptation module focuses on learning the generation of free viewpoints while reducing the amount of training. After training, given an input image of any viewpoint, the generalized dense view generator can generate a corresponding set of dense view images.

[0175] In an exemplary embodiment, the data processing method of the three-dimensional model further includes:

[0176] Sampling a fixed trajectory to obtain multiple sampling points;

[0177] For each sampling point, adding noise to the azimuth angle corresponding to the sampling point to obtain a new azimuth angle, and / or adding noise to the elevation angle corresponding to the sampling point to obtain a new elevation angle;

[0178] At least one random trajectory is obtained according to all new azimuth angles and / or new elevation angles.

[0179] In this embodiment, when generating a random trajectory, a fixed trajectory can be sampled and a small random noise can be added to the azimuth. Specifically, a small random number can be added to the original azimuth. For example, if the original azimuth is 120°, a random number between [-10°, 10°] can be added, and the new azimuth is a value between 110° and 130°. A random weighted combination of sine waves of different frequencies can also be added to the elevation, that is, there may be multiple sine waves superimposed in each direction, each sine wave has a different frequency and amplitude. Suppose there is a sine wave with a frequency of 0.5Hz and a sine wave with a frequency of 1Hz, and their amplitudes at each sampling point are random. By adjusting the parameters of the random noise and the sine wave, it can be ensured that the trajectory of the camera ends when it loops to the same azimuth and elevation as the conditional image, that is, the scene captured by the camera in the last image is similar to the scene in the conditional image. Due to the addition of random noise and sine waves, the motion trajectory of the camera is smooth in time, avoiding obvious jumps or unnatural changes in the generated trajectory. Through different combinations of random noise and sine waves, diverse trajectories can be generated, increasing the diversity of generated images.

[0180] Reference Figure 6 , the process of iteratively editing single-view images, generating dense view image sets and reconstructing to obtain the edited 3D model is described. During training, a view is randomly selected, and the edited 3D model reconstructed in the previous iteration is rendered to obtain a rendered single-view image. The above processing steps are repeated many times (including using a view-adaptive dense view generation method to generate a 360° dense view edited image set given any input view image, using a semantically consistent view selection method to select semantically consistent dense view edited images to reconstruct the edited 3D model, and using a 3D Gaussian reconstruction based on a mixed loss to reconstruct the edited 3D model according to a given dense view edited image set), until a certain number of cycles are met, and finally a 3D model that conforms to the edited text description is obtained. During inference, a single-view image is rendered given a source 3D model, and then input into the trained 3D scene editing model to obtain the edited 3D model.

[0181] As an optional embodiment, an editing platform that is easy for users to operate is also provided, which is used to respond to the operation instructions and guidance text input by the user, and generate a target editing image that conforms to the description of the guidance text based on the source 3D model. Specifically, a Web-based 3D model editing platform is constructed to provide an operable UI interface to assist users in completing editing. The interface diagram is shown in FIG. Figure 7 As shown. Click "Please enter the source 3D model" and a 3D model input box will pop up, which supports the input of a 3D Gaussian model; click "Please enter the edit text prompt" and a text input box will pop up to enter text; after that, the background will automatically perform single-view image editing and training of the dense view generation model with adaptive view. After the training is completed, a prompt will pop up saying "Please select a semantically consistent view or use the algorithm's default view selection method". Click "Please select a semantically consistent view" and the user can automatically select a semantically consistent view; click "Use the algorithm's default view selection method" and the background will automatically use the semantically consistent view selection model to select the appropriate view. After that, the 3D Gaussian reconstruction based on mixed loss is automatically completed and the edited 3D model is output.

[0182] In summary, the present invention proposes a three-dimensional scene model editing scheme based on efficient iteration. In one iterative training process, single-view editing and dense view generation are integrated to improve the iteration speed and editing quality. The iterative editing method is adopted to improve the editing quality. At the same time, the edited image of the dense view can ensure the consistency of multiple views in the three-dimensional editing process, reduce the number of iterations, and accelerate the optimization speed of the three-dimensional model editing. The present invention also proposes a dense view generation scheme with adaptive view to realize the generation of dense view for any input view image. The existing video generation diffusion model is used as the backbone network, and a view adaptive module is added to construct a dense view generation model. According to the fixed trajectory and the random trajectory, a data set is constructed, and the dense view generation model is trained so that it can generate a dense view image set of 360° around the object for any input view. The present invention also proposes a semantically consistent view selection method based on a large language model. The large language model is used to select multiple view images that best match the edit text description from the generated dense view images and are semantically consistent with the reconstruction model obtained in the previous iteration process, and form a dense view edit image set for iterative reconstruction, thereby improving the multi-view consistency and iteration efficiency in the iterative process. The present invention also proposes a three-dimensional Gaussian reconstruction based on mixed loss, which improves the three-dimensional editing quality based on a dense perspective editing image set according to a constructed multi-loss function and adaptive loss function hyperparameter learning.

[0183] Second, please refer to Figure 8 The present invention also provides a three-dimensional model data processing system, comprising:

[0184] A first determination module 11 is used to determine a source 3D model corresponding to a current iteration, and obtain a current single-view edited image of the source 3D model edited according to the guidance text;

[0185] A first generating module 12, configured to generate multiple perspectives of to-be-selected editing images based on the current single-perspective editing image;

[0186] The first acquisition module 13 is used to acquire, for each perspective, a first similarity score between the selected editing image corresponding to the perspective in the current iteration and the guidance text, acquire a second similarity score between the selected editing image corresponding to the perspective in the previous iteration and the guidance text, and determine the selected editing image corresponding to the larger one of the first similarity score and the second similarity score as the target editing image of the perspective;

[0187] The editing and reconstruction module 14 is used to process the source three-dimensional model according to all target editing images to obtain the edited three-dimensional model corresponding to the current iteration.

[0188] It can be seen that in this embodiment, the current single-perspective editing image obtained by rendering the source three-dimensional model at any perspective is first used to generate candidate editing images of multiple perspectives, and then the three-dimensional model is edited based on the candidate editing images of multiple perspectives, which can ensure multi-perspective consistency and thus reduce the number of iterations. The target editing image that best meets the description in the guidance text is selected from the candidate editing images in this iteration and the edited images corresponding to the previous iteration to reconstruct the source three-dimensional model, thereby achieving a smooth transition between the two iterative processes and improving the three-dimensional editing quality, thereby improving the multi-perspective consistency and iteration efficiency in the iterative process, and facilitating the efficient reconstruction of a three-dimensional model that meets the text description.

[0189] In an exemplary embodiment, the process of obtaining the first similarity score between the selected editing image and the guiding text corresponding to the perspective in the current iteration includes:

[0190] Obtain the first image title text of the image to be selected for editing corresponding to the viewing angle in the current iteration;

[0191] A first similarity score between the first image title text and the guidance text is obtained.

[0192] In an exemplary embodiment, the process of obtaining a first similarity score between the first image title text and the guidance text includes:

[0193] A reference similarity score between the first image title text and the guidance text is obtained; and the reference similarity score is adjusted based on feature information in the first image title text and feature information in the guidance text to obtain a first similarity score.

[0194] In an exemplary embodiment, the process of adjusting the reference similarity score based on the feature information in the first image title text and the feature information in the guidance text to obtain the first similarity score includes:

[0195] If the feature information in the first image title text includes the feature information in the guidance text, increasing the baseline similarity score by a first preset score to obtain a first similarity score;

[0196] If the feature information in the first image title text does not include the feature information in the guidance text, the reference similarity score is reduced by a second preset score to obtain a first similarity score.

[0197] In an exemplary embodiment, the process of obtaining a first similarity score between the first image title text and the guidance text includes:

[0198] Obtaining a baseline similarity score between the first image title text and the guidance text;

[0199] Determining whether there is a difference text between the source image text and the guidance text in the first image title text;

[0200] If so, the base similarity score is increased by a third preset score to obtain a first similarity score.

[0201] In an exemplary embodiment, the process of obtaining the second similarity score between the candidate editing image and the guiding text corresponding to the perspective in the previous iteration includes:

[0202] Rendering the edited 3D model in the previous iteration based on the view angle to obtain a candidate edited image corresponding to the view angle in the previous iteration;

[0203] Obtain the second image title text of the image to be selected for editing corresponding to the viewing angle in the previous iteration;

[0204] A second similarity score between the second image title text and the guidance text is calculated.

[0205] In an exemplary embodiment, the data processing system for the three-dimensional model further includes:

[0206] The judgment module is used to process the source three-dimensional model according to all target editing images to obtain the edited three-dimensional model corresponding to the current iteration, and then judge whether the current iteration meets the iteration end condition. If not, the edited three-dimensional model is used as the source three-dimensional model for the next iteration for the next iteration. If so, the edited three-dimensional model is used as the target three-dimensional model that meets the description of the guidance text.

[0207] In an exemplary embodiment, the process of processing the source 3D model according to all target edited images to obtain the edited 3D model corresponding to the current iteration includes:

[0208] Construct a hybrid loss function, which includes editing loss and reconstruction loss;

[0209] The source 3D model is processed based on the hybrid loss function and all target edited images to obtain the edited 3D model corresponding to the current iteration.

[0210] In an exemplary embodiment, the hybrid loss function is ;

[0211] Where L is the value of the hybrid loss function, α is the first hyperparameter corresponding to the editing loss, β is the second hyperparameter corresponding to the reconstruction loss, and L edit is the editing loss, L recon To rebuild the losses.

[0212] In an exemplary embodiment, the data processing system for the three-dimensional model further includes:

[0213] The second acquisition module is used to obtain the similarity score between each target edited image and the guidance text;

[0214] An adjustment module is used to adjust the first hyperparameter and the second hyperparameter according to the sum of all similarity scores.

[0215] In an exemplary embodiment, the process of adjusting the first hyperparameter and the second hyperparameter according to the sum of all similarity scores includes:

[0216] If the sum of the similarity scores is greater than a preset semantic value, adjusting the first hyperparameter to be less than the second hyperparameter; the preset semantic value is determined according to the total number of perspectives;

[0217] If the sum of the similarity scores is less than or equal to the preset semantic value, the first hyperparameter is adjusted to be equal to the second hyperparameter.

[0218] In an exemplary embodiment, the data processing system for the three-dimensional model further includes:

[0219] The first calculation module is used to calculate the editing loss based on the first relationship. The first relationship is: ,in, , ;

[0220] The second calculation module is used to calculate the reconstruction loss based on the second relationship. The second relationship is: ;

[0221] Among them, i is the number of the viewing angle, E img is the image encoder, E txt For text encoder, represents the cosine similarity, I tFor the edited image, I S is the source image, T t is the edited text, T S is the source text, is the candidate editing image of the i-th perspective, is the image of the i-th perspective obtained by rendering the source 3D model, The difference between the images before and after editing. is the text difference before and after editing, and M is the total number of perspectives.

[0222] In an exemplary embodiment, the process of obtaining a current single-view image of a source 3D model at a selected viewpoint includes:

[0223] Based on the selected viewing angle, determining a query point corresponding to each pixel point in the three-dimensional space of the source three-dimensional model;

[0224] For each query point, the color of the pixel corresponding to the query point is obtained according to the three-dimensional Gaussian function and the distance between the query point and the center point of the source three-dimensional model;

[0225] The current single-view image of the source 3D model at the selected viewpoint is obtained based on the colors of all pixels.

[0226] In an exemplary embodiment, the process of generating multiple perspectives of selected editing images based on the current single-perspective editing image includes:

[0227] The single-view editing image is input into the dense view generator to obtain the candidate editing image of dense view.

[0228] In an exemplary embodiment, the data processing system for the three-dimensional model further includes:

[0229] A pre-creation module for establishing a dense view generation network; the dense view generation network includes a video generation diffusion network and a view adaptation module;

[0230] The third acquisition module is used to acquire a first training data set obtained by the camera rotating around the target object for one cycle based on a fixed trajectory and a second training data set obtained by the camera rotating around the target object for one cycle based on a random trajectory; the target object is an object in the source three-dimensional model;

[0231] The training module is used to train the dense view generation network using the first training data set. When the training times reach the preset times, the video is fixed to generate the diffusion network, and the view adaptation module is trained using the second training data set until the training requirements are met to obtain a dense view generator.

[0232] In an exemplary embodiment, the data processing system for the three-dimensional model further includes:

[0233] A sampling module, used to sample a fixed trajectory to obtain multiple sampling points;

[0234] A fourth acquisition module is used for adding noise to the azimuth angle corresponding to each sampling point to obtain a new azimuth angle, and / or adding noise to the elevation angle corresponding to the sampling point to obtain a new elevation angle;

[0235] The fifth acquisition module is used to obtain at least one random trajectory according to all new azimuth angles and / or new elevation angles.

[0236] In a third aspect, the present invention further provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the three-dimensional model data processing method described in any of the above embodiments.

[0237] For an introduction to a computer program product provided by the present invention, please refer to the above embodiments, and the present invention will not be described in detail here.

[0238] A computer program product provided by the present invention has the same beneficial effects as the data processing method of the three-dimensional model.

[0239] Third, please refer to Fig. 9 The present invention also provides an electronic device, comprising:

[0240] A memory 21, used for storing computer programs;

[0241] The processor 22 is used to implement the steps of the three-dimensional model data processing method described in any of the above embodiments when executing the computer program.

[0242] The electronic device also includes:

[0243] The input interface 23 is connected to the processor 22 via the communication bus 26, and is used to obtain the computer programs, parameters and instructions imported from the outside, and save them in the memory 21 under the control of the processor 22. The input interface can be connected to an input device to receive parameters or instructions manually input by the user. The input device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the terminal housing.

[0244] The display unit 24 is connected to the processor 22 via the communication bus 26 and is used to display the data sent by the processor 22. The display unit can be a liquid crystal display or an electronic ink display.

[0245] The network port 25 is connected to the processor 22 via the communication bus 26, and is used to communicate with various external terminal devices. The communication technology used in the communication connection can be a wired communication technology or a wireless communication technology, such as mobile high-definition link technology, universal serial bus, high-definition multimedia interface, wireless fidelity technology, Bluetooth communication technology, low-power Bluetooth communication technology, communication technology based on IEEE802.11s, etc.

[0246] For an introduction to an electronic device provided by the present invention, please refer to the above embodiments, and the present invention will not be described in detail here.

[0247] The electronic device provided by the present invention has the same beneficial effects as the data processing method of the three-dimensional model.

[0248] Fifth, please refer to Fig.10 The present invention further provides a non-volatile storage medium 30 on which a computer program 31 is stored. When the computer program 31 is executed by a processor, the steps of the three-dimensional model data processing method described in any of the above embodiments are implemented.

[0249] Among them, the non-volatile storage medium 30 may include: ROM (Read-only memory), PROM (Programmable read-only memory), EAROM (Electrically alterable read only memory), EPROM (Erasable programmable read only memory), EEPROM (Electrically erasable programmable read only memory), Flash memory and other media that can store program codes.

[0250] For an introduction to a non-volatile storage medium provided by the present invention, please refer to the above embodiments, and the present invention will not be described in detail here.

[0251] The non-volatile storage medium provided by the present invention has the same beneficial effects as the data processing method of the three-dimensional model.

[0252] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0253] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method for a three-dimensional model, characterized in that: include: Determine a source three-dimensional model corresponding to a current iteration, and obtain a current single-view edited image of the source three-dimensional model edited according to the guidance text; Generate multiple perspectives of to-be-selected editing images based on the current single-perspective editing image; For each of the viewing angles, obtaining a first similarity score between the first image title text of the to-be-selected editing image corresponding to the viewing angle in the current iteration and the guiding text, obtaining a second similarity score between the second image title text of the to-be-selected editing image corresponding to the viewing angle in the previous iteration and the guiding text, and determining the to-be-selected editing image corresponding to the larger one of the first similarity score and the second similarity score as the target editing image of the viewing angle; Processing the source three-dimensional model according to all the target edited images to obtain an edited three-dimensional model corresponding to the current iteration; The process of processing the source three-dimensional model according to all the target edited images to obtain the edited three-dimensional model corresponding to the current iteration includes: Constructing a hybrid loss function, wherein the hybrid loss function includes an editing loss and a reconstruction loss; The source three-dimensional model is processed based on the mixed loss function and all the target edited images to obtain an edited three-dimensional model corresponding to the current iteration.

2. The data processing method of the three-dimensional model according to claim 1, characterized in that: The process of obtaining a first similarity score between the first image title text of the selected editing image corresponding to the perspective in the current iteration and the guiding text includes: A baseline similarity score between the first image title text and the guidance text is obtained; and the baseline similarity score is adjusted based on feature information in the first image title text and feature information in the guidance text to obtain the first similarity score.

3. The data processing method of the three-dimensional model according to claim 2, characterized in that: The process of adjusting the reference similarity score based on the feature information in the first image title text and the feature information in the guidance text to obtain the first similarity score includes: If the feature information in the first image title text includes the feature information in the guide text, increasing the baseline similarity score by a first preset score to obtain the first similarity score; If the feature information in the first image title text does not include the feature information in the guide text, the reference similarity score is reduced by a second preset score to obtain the first similarity score.

4. The data processing method of the three-dimensional model according to claim 1, characterized in that: The process of obtaining a first similarity score between the first image title text of the selected editing image corresponding to the perspective in the current iteration and the guiding text includes: Obtaining a baseline similarity score between the first image title text and the guidance text; Determining whether there is a difference text between the source image text and the guide text in the first image title text; If so, the base similarity score is increased by a third preset score to obtain the first similarity score.

5. The data processing method of the three-dimensional model according to claim 1, characterized in that: The process of obtaining a second similarity score between the second image title text of the selected editing image corresponding to the viewing angle in the previous iteration and the guiding text includes: Rendering the edited three-dimensional model in the previous iteration based on the viewing angle to obtain a selected editing image corresponding to the viewing angle in the previous iteration; Obtaining a second image title text of the image to be selected for editing corresponding to the viewing angle in the previous iteration; A second similarity score between the second image title text and the guidance text is calculated.

6. The data processing method of the three-dimensional model according to claim 1, characterized in that: After the source three-dimensional model is processed according to all the target edited images to obtain the edited three-dimensional model corresponding to the current iteration, the three-dimensional model data processing method further includes: Determine whether the current iteration satisfies the iteration end condition; If not, the edited three-dimensional model is used as the source three-dimensional model of the next iteration for the next iteration; If so, the edited three-dimensional model is used as a target three-dimensional model that conforms to the description of the guidance text.

7. The data processing method of the three-dimensional model according to claim 1, characterized in that: The hybrid loss function is ; Wherein, L is the value of the hybrid loss function, α is the first hyperparameter corresponding to the editing loss, β is the second hyperparameter corresponding to the reconstruction loss, and L edit is the editing loss, L recon To rebuild the losses.

8. The data processing method of the three-dimensional model according to claim 7, characterized in that: The data processing method of the three-dimensional model also includes: Obtaining a similarity score between each of the target editing images and the guiding text; The first hyperparameter and the second hyperparameter are adjusted according to the sum of all the similarity scores.

9. The data processing method of the three-dimensional model according to claim 8, characterized in that: The process of adjusting the first hyperparameter and the second hyperparameter according to the sum of all the similarity scores comprises: If the sum of the similarity scores is greater than a preset semantic value, adjusting the first super parameter to be less than the second super parameter; the preset semantic value is determined according to the total number of the perspectives; If the sum of the similarity scores is less than or equal to the preset semantic value, the first hyperparameter is adjusted to be equal to the second hyperparameter.

10. The data processing method of the three-dimensional model according to claim 7, characterized in that: The data processing method of the three-dimensional model also includes: The editing loss is calculated based on the first relational expression, which is: ,in, , ; The reconstruction loss is calculated based on the second relation, which is: ; Wherein, i is the index of the viewing angle, E img is the image encoder, E txt For text encoder, represents the cosine similarity, I t For the edited image, I S is the source image, T t is the edited text, T S is the source text, is the candidate editing image of the i-th perspective, is the image of the i-th perspective obtained by rendering the source 3D model, The difference between the images before and after editing. is the text difference before and after editing, and M is the total number of perspectives.

11. The data processing method of the three-dimensional model according to claim 1, characterized in that: The process of obtaining the current single-view edited image of the source 3D model edited according to the guidance text includes: Determine the selected perspective for the current iteration; Acquire a current single-view image of the source three-dimensional model at the selected view angle; The current single-view image is edited according to the guide text to obtain a current single-view edited image.

12. The data processing method of the three-dimensional model according to claim 11, characterized in that: The process of acquiring the current single-view image of the source three-dimensional model at the selected view angle includes: Based on the selected viewing angle, determining a query point corresponding to each pixel point in the three-dimensional space of the source three-dimensional model; For each query point, obtaining the color of the pixel point corresponding to the query point according to a three-dimensional Gaussian function and a distance between the query point and a center point of the source three-dimensional model; A current single-view image of the source three-dimensional model at the selected viewing angle is obtained based on the colors of all the pixel points.

13. The data processing method of a three-dimensional model according to any one of claims 1 to 12, characterized in that: The process of generating multiple perspectives of selected editing images based on the current single-perspective editing image includes: The current single-view editing image is input into a dense view generator to obtain a dense view candidate editing image.

14. The data processing method of the three-dimensional model according to claim 13, characterized in that: The data processing method of the three-dimensional model also includes: Establishing a dense view generation network; the dense view generation network includes a video generation diffusion network and a view adaptation module; Acquire a first training data set obtained by rotating a camera around a target object based on a fixed trajectory and a second training data set obtained by rotating a camera around the target object based on a random trajectory; the target object is an object in the source three-dimensional model; The dense perspective generation network is trained using the first training data set. When the number of training times reaches a preset number, the video generation diffusion network is fixed, and the perspective adaptation module is trained using the second training data set until the training requirements are met, thereby obtaining the dense perspective generator.

15. The data processing method of the three-dimensional model according to claim 14, characterized in that: The data processing method of the three-dimensional model also includes: Sampling the fixed trajectory to obtain a plurality of sampling points; For each of the sampling points, adding noise to the azimuth angle corresponding to the sampling point to obtain a new azimuth angle, and / or adding noise to the elevation angle corresponding to the sampling point to obtain a new elevation angle; At least one random trajectory is obtained according to all the new azimuth angles and / or the new elevation angles.

16. A three-dimensional model data processing system, characterized in that: include: A first determination module is used to determine a source three-dimensional model corresponding to a current iteration, and obtain a current single-view edited image of the source three-dimensional model edited according to the guidance text; A first generating module, configured to generate multiple perspectives of to-be-selected editing images based on the current single-perspective editing image; A first acquisition module is used to acquire, for each of the perspectives, a first similarity score between a first image title text of a candidate editing image corresponding to the perspective in the current iteration and the guide text, acquire a second similarity score between a second image title text of a candidate editing image corresponding to the perspective in the previous iteration and the guide text, and determine the candidate editing image corresponding to the larger one of the first similarity score and the second similarity score as the target editing image of the perspective; An editing and reconstruction module, used for processing the source three-dimensional model according to all the target editing images to obtain an edited three-dimensional model corresponding to the current iteration; The process of processing the source three-dimensional model according to all the target edited images to obtain the edited three-dimensional model corresponding to the current iteration includes: Constructing a hybrid loss function, wherein the hybrid loss function includes an editing loss and a reconstruction loss; The source three-dimensional model is processed based on the mixed loss function and all the target edited images to obtain an edited three-dimensional model corresponding to the current iteration.

17. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the three-dimensional model data processing method described in any one of claims 1-15 are implemented.

18. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the three-dimensional model data processing method according to any one of claims 1 to 15 when executing the computer program.

19. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, which, when executed by a processor, implements the steps of the three-dimensional model data processing method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Data mining system, method and device based on graphics and text information combination

    CN116595064A

  • Target behavior monitoring method and system

    CN117953544A