Model image editing method and device, equipment, medium and product
By acquiring the model representation and rendering data of the 3D model, determining the pose of the distribution points and performing binding rendering, the problem of inaccurate 3D content caused by mesh geometry dependency in the existing technology is solved, and better rendering effect and image quality are achieved.
Patent Information
- Application Number
- CN202410613313.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-18
AI Technical Summary
Existing 3D content editing solutions based on 3DGS rely heavily on the accuracy of mesh geometry and cannot effectively repair or delete inaccurate mesh parts, resulting in inaccurate 3D content after editing.
By acquiring the model representation data and rendering data of the 3D model, the first pose of the distribution points in the local coordinate space is determined, and the second pose of the structural units in the global coordinate space is determined based on the editing operation. The machine learning model is then used for binding and rendering to realize the editing of the 3D model.
It improves the rendering effect and image quality of the edited 3D model image, ensures real-time control and adjustment of the distribution points of structural units after editing, and reduces the impact on the original model.
Smart Images

Figure CN120976382A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to a model image editing method and device, equipment, medium and product. BACKGROUND
[0002] Operating and editing three-dimensional content is an important task in the field of computer graphics and three-dimensional modeling, involving the creation, modification and optimization of three-dimensional models, scenes and animations. Three-dimensional Gaussian splatting (3D Gaussian Splatting, 3DGS) technology is a technique for three-dimensional data visualization and processing, mainly applied in the fields of computer graphics, volume rendering and volume data processing. This method realizes the smoothing processing and reconstruction of volume data by applying a Gaussian function (i.e. normal distribution function) to the data points.
[0003] In related technologies, in a three-dimensional content editing scheme based on 3DGS, the Sparse Gaussian Representations (SuGaR) technology is implemented, i.e. SuGaR proposes a method for extracting a mesh from 3DGS, adds additional regularization, binds 3DGS on the extracted mesh and uses Gaussian splatting rendering for animation production.
[0004] However, for the three-dimensional content editing scheme based on 3DGS, SuGaR relies heavily on the accuracy of mesh geometry, inheriting the defects of mesh rendering. That is, for the inaccurate part of the extracted mesh, SuGaR cannot repair the missing part or delete the redundant part, resulting in inaccurate edited three-dimensional content generated. SUMMARY
[0005] The embodiments of the present application provide a model image editing method, device, equipment, medium and product. The technical solution is as follows:
[0006] On the one hand, a model image editing method is provided, which comprises:
[0007] When receiving an editing operation for a first model image, obtaining model representation data and model rendering data corresponding to a three-dimensional model in the first model image, the model representation data comprising a structural unit constructing the three-dimensional model, and the model rendering data comprising rendering distribution features corresponding to a plurality of distribution points on the three-dimensional model;
[0008] Determine the first pose corresponding to the rendering distribution features of the plurality of distribution points in the first space, the first space being a local coordinate space constructed based on the structural unit;
[0009] determine a second pose of the edited structure unit in a second space based on the editing operation and the model representation data, the second space being a global coordinate space constructed based on the first model image;
[0010] render a second model image edited for the three-dimensional model in the first model image based on the first pose and the second pose.
[0011] In another aspect, an editing device of a model image is provided, the device comprising:
[0012] an obtaining module configured to, when an editing operation for a first model image is received, obtain model representation data corresponding to a three-dimensional model in the first model image and model rendering data, the model representation data comprising structure units constructing the three-dimensional model, and the model rendering data comprising rendering distribution features corresponding to a plurality of distribution points on the three-dimensional model;
[0013] a first determining module configured to determine first poses of the rendering distribution features corresponding to the plurality of distribution points in a first space, the first space being a local coordinate space constructed based on the structure units;
[0014] a second determining module configured to determine a second pose of an edited structure unit in a second space based on the editing operation and the model representation data, the second space being a global coordinate space constructed based on the first model image;
[0015] a rendering module configured to render a second model image edited for the three-dimensional model in the first model image based on the first pose and the second pose.
[0016] In some optional embodiments, the first determining module further comprises:
[0017] a constructing unit configured to construct the first space corresponding to the structure units;
[0018] a binding unit configured to bind the plurality of distribution points to the first space of the structure units respectively, and determine the first poses of the distribution points in the first space.
[0019] In some optional embodiments, the method is implemented through a pre-trained machine learning model;
[0020] the binding unit is further configured to extract pose features of the distribution points in the first space according to the model rendering data through the machine learning model, to obtain the first poses of the distribution points.
[0021] In some optional embodiments, the obtaining module is configured to obtain sample model images of a sample three-dimensional model at multiple viewing angles, training editing instructions, and sample edited model images corresponding to the training editing instructions.
[0022] The device further includes a training module configured to input the sample model images and the training editing instructions into the machine learning model to obtain predicted edited model images, and perform iterative learning on a relationship between a first pose of a sample distribution point in the first space, a second pose of the sample distribution point in the second space, and a third pose of a sample structure unit in the second space based on a difference between the predicted edited model images and the sample edited model images, to obtain the pre-trained machine learning model.
[0023] In some optional embodiments, the model representation data includes a plurality of structure units, and the model rendering data includes N distribution points, where N is a positive integer.
[0024] The binding unit is further configured to take a spatial origin of the first space corresponding to each of the plurality of structure units as a clustering center, to cluster the N distribution points to obtain a clustering cluster corresponding to each of the plurality of structure units; for an i-th clustering cluster, to take K distribution points in the i-th clustering cluster as distribution points corresponding to an i-th structure unit, where i and K are positive integers, and K
[0025] In some optional embodiments, the structure unit includes a polygonal mesh.
[0026] The constructing unit is further configured to, for an i-th polygon in the polygonal mesh, take a specified point on a plane on which the i-th polygon is located as a coordinate origin of a first spatial coordinate system corresponding to the first space, where i is a positive integer; take a specified edge in the i-th polygon as a first axis direction of the first spatial coordinate system; take a normal direction of the i-th polygon as a second axis direction of the first spatial coordinate system; and determine a third axis direction of the first spatial coordinate system according to the first axis direction and the second axis direction, where the third axis direction is perpendicular to the first axis direction, and the third axis direction is perpendicular to the second axis direction.
[0027] In some optional embodiments, the model representation data includes a plurality of structure units.
[0028] The second determining module is further configured to adjust the structural relationships between the plurality of structural units in the model representation data based on the editing operation to obtain adjusted model representation data; and to determine the second pose of the plurality of adjusted structural units in the second space based on the adjusted model representation data.
[0029] In some alternative embodiments, the plurality of structural units are implemented as a polygonal mesh composed of a plurality of polygons, each polygon including at least three vertices;
[0030] The second determining module is further configured to, for the i-th polygon, generate the second pose of the i-th polygon in the second space based on the coordinate values of the vertices corresponding to the i-th polygon in the second spatial coordinate system corresponding to the second space, where i is a positive integer.
[0031] In some alternative embodiments, the apparatus further includes:
[0032] The third determining module is used to determine the third pose of the edited distribution points in the second space based on the first pose and the second pose;
[0033] The rendering module is also used to render and generate the second model image based on the third pose corresponding to the plurality of distribution points.
[0034] In some optional embodiments, the first pose includes a first rotation matrix, a first scaling matrix, and a first position matrix of the distribution points in the first space, and the second pose includes a second rotation matrix and a second position matrix of the edited structural unit in the second space;
[0035] The third determining module is further configured to determine a third rotation matrix of the edited distribution points in the second space based on the first rotation matrix and the second rotation matrix; determine a second scaling matrix of the edited distribution points in the second space based on the first scaling matrix; determine a third position matrix of the edited distribution points in the second space based on the first position matrix and the second position matrix; and form the third pose of the edited distribution points in the second space by the third rotation matrix, the second scaling matrix and the third position matrix.
[0036] In some optional embodiments, the distribution points correspond to a color density distribution;
[0037] The rendering module is further configured to determine the position distribution of the edited distribution points in the second space based on the third pose corresponding to the plurality of distribution points respectively; and to render the edited 3D model based on the color density distribution corresponding to the distribution points and the position distribution to obtain the second model image.
[0038] In some optional embodiments, the acquisition module includes:
[0039] The acquisition unit is used to acquire first model images of the three-dimensional model from multiple perspectives;
[0040] The generation unit is used to generate the model rendering data based on the color distribution density of pixels in the first model image under the multiple viewpoints; extract the model structure corresponding to the three-dimensional model based on the model rendering data, determine multiple structural units, and form the model representation data by the multiple structural units.
[0041] In some optional embodiments, the generation unit is further configured to extract multiple polygons corresponding to the three-dimensional model based on the model rendering data, and use the polygon mesh composed of the multiple polygons as the model representation data, wherein the polygon mesh is used to indicate the geometric structure of the three-dimensional model.
[0042] In some optional embodiments, the generation unit is further configured to extract point cloud data corresponding to the 3D model based on the model rendering data, and use the point cloud data as the model representation data, wherein the point cloud data is used to indicate the geometric coordinates of the points that make up the 3D model;
[0043] In some optional embodiments, the generation unit is further configured to extract multiple voxels corresponding to the three-dimensional model based on the model rendering data, and use the voxel model composed of the multiple voxels as the model representation data, wherein the voxel model is used to indicate the geometric distribution of the volume pixels that make up the three-dimensional model.
[0044] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the model image editing method as described in any of the embodiments of this application above.
[0045] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the model image editing method as described in any of the embodiments of this application above.
[0046] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the model image editing method described in any of the above embodiments.
[0047] The technical solution provided in this application includes at least the following beneficial effects:
[0048] When editing the 3D model in the first model image, the distribution points corresponding to the rendering distribution features of the 3D model are associated and bound to the structural units that make up the 3D model. This allows the first pose of the rendering distribution features in the local coordinate space corresponding to the structural unit to be obtained. Then, the second pose of the structural unit in the global coordinate space after editing is determined based on the impact of the editing operation on the model structure of the 3D model. Since the first pose can preserve the relative positions between the distribution points bound to adjacent structural units, the corresponding bound distribution points can be manipulated and adjusted in real time after the structural unit changes according to the editing operation, thus obtaining the adjusted rendering distribution features. This results in a better rendering effect for the rendered second model image. Furthermore, since the first pose preserves the relative positions between the distribution points on adjacent structural units, the structural units only serve as a conversion medium for the editing operation. Therefore, the structural units obtained from the first model image have little impact on the image quality of the edited second model image, thus ensuring the image quality of the generated edited second model image. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a structural block diagram of a computer system provided in an exemplary embodiment of this application;
[0051] Figure 2 This is a flowchart of a model image editing method provided in an exemplary embodiment of this application;
[0052] Figure 3 This is a flowchart of a model image editing method provided in an exemplary embodiment of this application;
[0053] Figure 4This is a schematic diagram of the spatial coordinate system corresponding to a triangle provided in an exemplary embodiment of this application;
[0054] Figure 5 This is a schematic flowchart of a model image editing method provided in an exemplary embodiment of this application;
[0055] Figure 6 This is a flowchart of a model image editing method provided in an exemplary embodiment of this application;
[0056] Figure 7 This is a schematic diagram comparing the editing effects of a model image editing method provided in an exemplary embodiment of this application;
[0057] Figure 8 This is a schematic diagram comparing the editing effects of a model image editing method provided in an exemplary embodiment of this application;
[0058] Figure 9 This is a schematic diagram illustrating the effect of editing a model image provided in an exemplary embodiment of this application;
[0059] Figure 10 This is a structural block diagram of a model image editing device provided in an exemplary embodiment of this application;
[0060] Figure 11 This is a structural block diagram of a model image editing device provided in an exemplary embodiment of this application;
[0061] Figure 12 This is a schematic diagram of the structure of a server provided in an exemplary embodiment of this application. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0063] First, a brief introduction to the terms used in the embodiments of this application will be given.
[0064] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0065] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0066] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.
[0067] Computer Vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision. Pre-trained models in the field of vision, such as Swin-Transformer, Vision Transformer (ViT), Vision Mixture of Experts (V-MOE), and Masked Autoencoder (MAE), can be quickly and widely applied to specific downstream tasks after fine-tuning. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality and map building, as well as common biometric recognition technologies.
[0068] 3D Gaussian Splatting (3DGS) is a technique for 3D data visualization and processing, enabling volumetric rendering. It is primarily used in computer graphics, volumetric drawing, and volumetric data processing. This method smooths and reconstructs volumetric data by applying a Gaussian function (i.e., a normal distribution function) to data points. In other words, it achieves this by "spraying" volumetric data into a series of Gaussian-distributed points (splats) in 3D space. Each splat represents a local region of the volumetric data and has specific density and color attributes.
[0069] A triangular mesh is a common data structure in computer graphics and 3D modeling used to represent the surfaces of 3D objects. As the name suggests, a triangular mesh is composed of many triangles connected by vertices and edges to form a continuous mesh structure. Each vertex in the mesh has coordinates (x, y, z) in 3D space and is the foundation of the mesh. An edge is a straight line segment connecting two vertices; each edge in the mesh represents a connection between two vertices. A triangular face in the mesh is formed by sequentially connecting three vertices, and these triangular faces define the surface of the mesh.
[0070] Manipulation: In computer graphics, manipulation typically involves transforming, editing, and modifying 2D images or 3D models. These manipulations help achieve desired visual effects, create complex scenes or animations, and perform real-time rendering. Manipulating 3D content is a crucial task in computer graphics and 3D modeling, involving the creation, modification, and optimization of 3D models, scenes, and animations. These tasks have wide applications in game development, film and television production, architectural design, virtual reality, and augmented reality.
[0071] In 3D modeling tasks, various methods can be used to create and edit 3D models, such as geometry-based modeling, parametric modeling, surface modeling, and voxel modeling. Mesh operations involve adding, deleting, connecting, and splitting the mesh structure of a 3D model to change its shape and topology. Furthermore, basic transformations such as translation, rotation, and scaling, as well as more complex deformation operations like twisting, bending, and expanding, can be performed on 3D models.
[0072] Figure 1 A structural block diagram of a computer system provided in an exemplary embodiment of this application is shown. The computer system 100 includes a terminal 120 and a server 140.
[0073] The device type of terminal 120 includes at least one of the following: desktop computer, smartphone, tablet computer, e-book reader, Moving Picture Experts Group Audio Layer III (MP3) player, Moving Picture Experts Group Audio Layer IV (MP4) player, and laptop computer. The following embodiments use a desktop computer as an example.
[0074] Terminal 120 is connected to server 140 via a wireless network or a wired network.
[0075] Those skilled in the art will understand that the number of the aforementioned devices can be more or less. For example, there may be only one device, or there may be dozens or hundreds of devices, or even more. This application does not limit the number or type of devices.
[0076] Server 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 140 is used to provide background services for applications supporting a three-dimensional virtual environment. Optionally, server 140 undertakes the primary computing task, and terminal 120 undertakes the secondary computing task; or, server 140 undertakes the secondary computing task, and terminal 120 undertakes the primary computing task; or, server 140 and terminal 120 collaborate on computing using a distributed computing architecture.
[0077] It is worth noting that the aforementioned server 140 can be implemented as a physical server or as a cloud server. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied in the cloud computing business model. It can form resource pools, allowing for flexible and convenient on-demand use.
[0078] In some embodiments, the server 140 described above can also be implemented as a node in a blockchain system.
[0079] The following is an illustrative example of implementing the model image editing method provided in this application embodiment in server 140. Illustratively, server 140 provides model image editing functionality to terminal 120. Terminal 120 sends a model image editing request to server 140, requesting server 140 to perform editing operations on a first model image, wherein the model editing request includes the first model image. After receiving the model image editing request, server 140 obtains model representation data and model rendering data corresponding to the 3D model in the first model image. The model representation data includes structural units that construct the 3D model, and the model rendering data includes rendering distribution features corresponding to multiple distribution points on the 3D model. Server 140 inputs the aforementioned model representation data and model rendering data into a pre-trained machine learning model for model image editing. This machine learning model determines the first pose in the first space corresponding to the rendering distribution features of multiple distribution points. Based on the editing operation and model representation data, it determines the second pose of the edited structural unit in the second space. Based on the first and second poses, it renders the second model image after editing the 3D model in the first model image. The first space is a local coordinate space constructed based on the structural unit, and the second space is a global coordinate space constructed based on the first model image. After the machine learning model outputs the second model image, server 140 sends the second model image to terminal 120, which displays the second model image.
[0080] It is worth noting that the model image editing method provided in this application embodiment can also be implemented in the terminal 120. Schematic, the terminal 120 is equipped with a hardware module / software module for implementing the model image editing function. When the user needs to edit the first model image, the terminal 120 reads the first model image indicating the editing operation from the storage area, processes the first model image to obtain the model representation data and model rendering data corresponding to the first model image, and then inputs the model representation data and model rendering data to the hardware module / software module corresponding to the model image editing function, outputting a second model image, which the terminal 120 then displays.
[0081] Based on the above-described terminology and application scenarios, the method for editing model images provided in this application will be explained, taking the execution of this method by a server as an example. Figure 2 As shown, the method includes the following steps 210 to 240.
[0082] Step 210: When an editing operation is received for the first model image, obtain the model representation data and model rendering data corresponding to the 3D model in the first model image.
[0083] In this embodiment of the application, the first model image is an image containing a three-dimensional model, that is, the three-dimensional model is displayed in the first model image. Optionally, the first model image can be implemented as an image of the three-dimensional model observed from multiple perspectives, or the first model image can be implemented as an image of the three-dimensional model observed from a specified perspective.
[0084] Optionally, the above-mentioned three-dimensional model can be implemented as a model of a real object. For example, the above-mentioned first model image can be implemented as a photograph, which includes the three-dimensional model corresponding to the real object; or, the above-mentioned three-dimensional model can be implemented as a three-dimensional model designed by modeling software, such as virtual objects or virtual buildings in games.
[0085] Optionally, the above editing operations include translation, rotation, scaling, stretching, deformation, topology change, and viewpoint change operations, etc., which are not limited here. Specifically, translation indicates moving the position of the 3D model within the image; rotation indicates rotating the 3D model within the image; scaling indicates adjusting the size of the 3D model within the image; stretching indicates scaling the 3D model non-uniformly in a specified direction within the image; deformation indicates changing the shape of the 3D model within the image; topology change indicates changing the topology of the 3D model within the image; and viewpoint change indicates changing the viewpoint of the 3D model displayed in the image.
[0086] Schematic, the model representation data described above is data used to describe the geometric representation of a 3D model; it is data used to represent a 3D object in a computer using mathematical and geometric methods. This model representation data includes structural units that construct the 3D model; that is, the model representation data constructs the 3D model through the geometric structure indicated by these structural units.
[0087] In some embodiments, the model representation data includes multiple structural units.
[0088] Optionally, the model representation data mentioned above can be implemented as polygon / mesh representation data, voxel representation data, surface representation data, constructive solid geometry (CSG) data, boundary representation (B-Rep) data, point cloud data, subdivision surface representation data, etc.
[0089] The aforementioned polygon / mesh representation data approximates the surface of a 3D model as multiple polygons, that is, data composed of polygons as structural units. Optionally, the aforementioned polygons can be implemented as triangles, quadrilaterals, pentagons, etc.
[0090] The voxel representation data mentioned above uses voxels as structural units to represent 3D models. A voxel is a pixel (a 3D version of a pixel) used to represent volume data in 3D space. A voxel is the basic unit in volumetric imaging technology; it is a cubic unit in 3D space with three dimensions: width, height, and depth, and is typically stored as a 3D array.
[0091] The surface representation data mentioned above uses parametric or implicit surfaces to represent smooth 3D models. Parametric surfaces include Bézier surfaces and B-spline surfaces, while implicit surfaces are defined using mathematical equations such as the equation of a sphere. Therefore, structural elements are implemented as parametric or implicit surfaces.
[0092] The above-mentioned constructing entity geometric data uses Boolean operations (such as union, intersection, difference) to combine basic geometric shapes (such as cubes, cylinders, spheres), and then uses the combination of the above basic geometric shapes to obtain the required three-dimensional model data. Thus, the structural unit is implemented as a basic geometric shape represented by Boolean operations.
[0093] The boundary representation data mentioned above is used to represent the data of a 3D model by defining the boundaries of an object; therefore, the structural unit is implemented as the boundary of the object.
[0094] The point cloud data mentioned above consists of a large number of discretely distributed points. Each point has its own coordinates in three-dimensional space. Usually, the points also include color and intensity information. The set of points is used to indicate the three-dimensional model. The points mentioned above are realized as structural units in the data.
[0095] The above-mentioned subdivision surface representation data creates a smooth surface by subdividing the edges and vertices in a polygonal mesh. That is, by subdividing the original coarse polygonal mesh into smaller patches, a smoother surface is generated. These patches are the structural units in the data.
[0096] Schematic, model rendering data is used to indicate the rendering status of a 3D model. The model rendering data includes rendering distribution features corresponding to multiple distribution points on the 3D model. In some embodiments, the distribution points can be implemented as Gaussian points, and the rendering distribution features indicated by the Gaussian points are implemented as 3D Gaussian sputtering (3DGS), wherein 3DGS uses explicit 3D Gaussian points as its primary rendering primitives.
[0097] Step 220: Determine the first pose of the rendering distribution features corresponding to the multiple distribution points in the first space.
[0098] The first space is a local coordinate space constructed based on structural units; that is, the first space is the space corresponding to the first spatial coordinate system constructed based on structural units.
[0099] In some embodiments, a first space corresponding to a structural unit is constructed, and multiple distribution points are respectively bound to the first space of the structural unit to determine the first pose of each distribution point in the first space. That is, using the structural unit as the construction reference for the spatial coordinate system, a first spatial coordinate system corresponding to the structural unit is generated, and multiple distribution points are respectively bound to the first spatial coordinate system to determine the first pose of each distribution point in the first space.
[0100] In this embodiment, a local coordinate space is defined for each structural unit, and the rendering distribution features corresponding to the distribution points are bound to the local coordinate space. When the structural unit is edited, the properties of the local coordinate space remain stable, thus enabling the rendering distribution features of the bound distribution points to be manipulated and adjusted in real time.
[0101] In some embodiments, the model representation data includes multiple structural units. Optionally, for each structural unit, a binding process for multiple distribution points is implemented, that is, the first pose of each distribution point in the first spatial coordinate system corresponding to the multiple structural units is generated. Illustratively, the model representation data includes M structural units, and the model rendering data includes N distribution points. For the i-th structural unit among the M structural units, the first pose of N distribution points in the first spatial coordinate system corresponding to the i-th structural unit is generated, where M, N, and i are positive integers, and i ≤ M.
[0102] Optionally, for each structural unit, a specified number of distribution points are determined from the multiple distribution points mentioned above, and the first pose of the specified number of distribution points in the first spatial coordinate system corresponding to the structural unit is generated. Illustratively, the model representation data corresponds to M structural units, and the model rendering data corresponds to N distribution points. For the i-th structural unit among the M structural units, K distribution points corresponding to the i-th structural unit are obtained, and the first pose of each of the K distribution points in the first spatial coordinate system corresponding to the i-th structural unit is generated, where K is a positive integer and K < N.
[0103] Optionally, the determination of the above K distribution points can be implemented in at least one of the following ways:
[0104] The first method is to determine the distance between the distribution point and the origin of the first spatial coordinate system.
[0105] In a schematic manner, for the i-th structural unit, the origin of the first spatial coordinate system corresponding to the i-th structural unit is determined, the distribution distance between multiple distribution points and the origin is determined respectively, and the distribution points whose distance is less than a preset distance threshold are determined as the distribution points corresponding to the i-th structural unit.
[0106] The second method involves clustering multiple distribution points to determine the location.
[0107] In a schematic manner, the origin of the first space corresponding to each of the multiple structural units is used as the cluster center, and N distribution points are clustered to obtain clusters corresponding to the multiple structural units. For the i-th cluster, K distribution points in the i-th cluster are used as the distribution points corresponding to the i-th structural unit, where i and K are positive integers and K < N. The K distribution points are bound to the first space of the i-th structural unit. The first pose of the K distribution points in the first space of the i-th structural unit is determined respectively.
[0108] Optionally, the aforementioned first pose includes a first rotation matrix, a first scaling matrix, and a first position matrix for the distribution point in the first space. The first rotation matrix indicates the orientation of the distribution point relative to the first spatial coordinate system, representing the pose of the distribution point in the first spatial coordinate system. The first scaling matrix indicates the shape and size of the rendering distribution feature corresponding to the distribution point. When the distribution point is a Gaussian point, the first scaling matrix determines the width of the Gaussian function, that is, the distribution range of the rendering distribution feature in space. A smaller scaling ratio produces a more concentrated rendering distribution, while a larger scaling ratio results in a more dispersed rendering distribution. The first position matrix indicates the coordinate values of the distribution point in the first spatial coordinate system.
[0109] Optionally, binding multiple distribution points to the first space of the structural unit can be implemented in at least one of the following ways:
[0110] The first method is coordinate system transformation.
[0111] Schematic example: The fourth pose of multiple distribution points in a specified coordinate system is obtained, along with the coordinate system transformation relationship between the first spatial coordinate system and the specified coordinate system. Based on the aforementioned coordinate system transformation relationship and the fourth pose, the first pose of the distribution points in the first spatial coordinate system is determined. Here, the specified coordinate system is a coordinate system constructed based on the 3D model, such as the world coordinate system.
[0112] In one example, the specified coordinate system can be based on the centroid of the 3D model as the origin, the horizontal direction of the first model image as the x-axis, the vertical direction of the first model image as the y-axis, and the direction perpendicular to the first model image as the z-axis. It is worth noting that the origin of the specified coordinate system can also be the upper left or lower right corner of the first model image, or the upper left or lower right corner of the 3D model, etc., without specific limitations.
[0113] The second method is model prediction.
[0114] In a schematic manner, a pre-trained machine learning model is obtained, which is used to implement the editing function of the model image. The machine learning model extracts the pose features of the distribution points in the first space based on the model rendering data to obtain the first pose of the distribution points.
[0115] Optionally, the aforementioned machine learning model can be implemented as at least one neural network model among Convolutional Neural Networks (CNN), Feedforward Neural Network (FNN), Residual Network (ResNet), Transformer, Swin-Transformer, ViT, V-MOE, MAE, etc.
[0116] Optionally, the construction process of the first space described above can be implemented as at least one of the following:
[0117] The first type is when the structural unit includes a polygonal mesh, that is, the model representation data is a polygonal mesh composed of multiple polygons, and a first spatial coordinate system is constructed for each polygon.
[0118] In a schematic way, for the i-th polygon in the polygonal mesh, a specified point on the plane where the i-th polygon is located is taken as the origin of the first spatial coordinate system corresponding to the first space, where i is a positive integer. A specified edge in the i-th polygon is taken as the first axis direction of the first spatial coordinate system corresponding to the first space. The normal direction of the i-th polygon is taken as the second axis direction of the first spatial coordinate system. The third axis direction of the first spatial coordinate system is determined based on the first axis direction and the second axis direction. The third axis direction is perpendicular to the first axis direction and the second axis direction.
[0119] Optionally, the specified point on the plane where the i-th polygon is located can be a specified vertex of the i-th polygon, the center of the polygon, the centroid of the polygon, the centroid of the polygon, etc., and is not limited here.
[0120] In one example, taking the polygon as a triangle, the first space is the local space corresponding to the triangle. The triangle includes the first side, the second side, and the third side. The coordinate system definition process of the first space coordinate system corresponding to the triangle is as follows: the centroid of the triangle is taken as the origin of the first space coordinate system, the first axis direction of the first space coordinate system is defined as the direction of the first side, the second axis direction of the first space coordinate system is defined as the normal direction of the triangle, and the third axis direction of the first space coordinate system is defined as the cross product of the first axis direction and the second axis direction.
[0121] Using polygonal meshes to indicate the structure of a 3D model can transform editing of the 3D model into editing of the polygonal mesh, thus simplifying the process of editing the structure of the 3D model. Furthermore, the simple polygonal structure allows for the construction of the local coordinate space corresponding to the polygons with fewer data processing steps, reducing the consumption of computing resources.
[0122] The second approach involves implementing structural units using voxels. The model representation data includes volume data composed of multiple voxels, with each voxel corresponding to a first spatial coordinate system.
[0123] Schematic illustration: For the i-th voxel, the centroid of the i-th voxel is taken as the origin of the first spatial coordinate system. The cutting direction of the i-th voxel in the first voxel is taken as the first axis direction of the first spatial coordinate system, where i is a positive integer. The cutting direction of the i-th voxel in the second voxel is taken as the second axis direction of the first spatial coordinate system. The cutting direction of the i-th voxel in the third voxel is taken as the third axis direction of the first spatial coordinate system. The first, second, and third axis directions are mutually perpendicular. Taking a cube as an example, the voxel cutting direction can be implemented as a direction perpendicular to the front and back, left and right, or top and bottom faces of the cube.
[0124] The third type, where the structural unit is implemented through points, involves model representation data consisting of point cloud data composed of multiple points, with each point corresponding to a first spatial coordinate system.
[0125] Indicatively, for the i-th point, the local point cloud corresponding to the i-th point is obtained. The local point cloud is a point cloud composed of points within the neighborhood of the i-th point. The normal direction corresponding to the local point cloud is determined, and the normal direction is defined as the first axis direction of the first spatial coordinate system. The covariance matrix corresponding to the local point cloud is determined, and the eigenvector with the largest eigenvalue in the covariance matrix is defined as the second axis direction of the first spatial coordinate system. The third axis direction of the first spatial coordinate system is defined as the cross product of the first axis direction and the second axis direction.
[0126] Step 230: Based on the editing operation and model representation data, determine the second pose of the edited structural unit in the second space.
[0127] The aforementioned second space is a global coordinate space constructed based on the first model image; that is, the second space is the space corresponding to the second space coordinate system constructed based on the three-dimensional model of the first model image. In some embodiments, the aforementioned second space coordinate system can be implemented as a world coordinate system, a Gaussian coordinate system, etc.
[0128] Optionally, the second spatial coordinate system can be implemented with a specified point of the 3D model as the origin, a first specified direction of the 3D model as the first axis direction of the second spatial coordinate system, a second specified direction as the second axis direction of the second spatial coordinate system, and the cross product of the first and second axis directions as the third axis direction. The specified point can be at least one of the model's centroid, center of mass, or center. In one example, taking a table as the 3D model, the first specified direction can be perpendicular to the tabletop, and the second specified direction can be parallel to the tabletop.
[0129] Optionally, the aforementioned second spatial coordinate system can be implemented with a specified angle of the first model image as the origin, the horizontal direction of the first model image as the first axis direction, the vertical direction of the first model image as the second axis direction, and the cross product of the first and second axis directions as the third axis direction. In some embodiments, when multiple first model images are included, the second spatial coordinate system can be constructed based on the first model image corresponding to a specified viewpoint among the multiple first model images.
[0130] In some embodiments, determining the second pose of the edited structural unit in the second space based on the editing operation and the model representation data can be achieved by: adjusting the structural relationship between multiple structural units in the model representation data based on the editing operation to obtain the adjusted model representation data, and determining the second pose of multiple adjusted structural units in the second space based on the adjusted model representation data.
[0131] Schematic example, taking the aforementioned structural unit as a polygon and the model representation data as a polygon mesh, the polygon includes at least three vertices. The edited polygon mesh is determined by the positions of each vertex in the polygon mesh based on the editing operation. The edited polygon mesh includes multiple edited polygons. The second pose is determined based on the distribution of the edited polygons in the second spatial coordinate system. That is, for the i-th polygon, the second pose of the i-th polygon in the second spatial coordinate system is generated based on the coordinate values of the corresponding vertices of the i-th polygon in the second spatial coordinate system, where i is a positive integer.
[0132] That is, structural units are used as proxies for distribution point editing. The structural relationships between structural units are adjusted according to the editing operation. The second pose is generated according to the distribution of the adjusted structural units in the second space. Since the combination of structural units can indicate the geometry of the 3D model, and the editing operation of the 3D model is actually the editing of the geometry of the 3D model, using structural units as proxies for distribution point editing can quickly realize the editing of the 3D model.
[0133] Optionally, the second pose includes a second rotation matrix and a second position matrix for the edited structural unit in the second space. The second rotation matrix indicates the orientation of the edited structural unit relative to the second spatial coordinate system, representing the orientation of the edited structural unit in the second spatial coordinate system. The second position matrix indicates the coordinate values of the edited structural unit in the second spatial coordinate system.
[0134] In some embodiments, when the structural unit is implemented as a polygon and the model representation data is implemented as a polygonal mesh composed of multiple polygons, the direction of the above-edited structural unit relative to the second spatial coordinate system is implemented as the angle between the direction from the origin of the second spatial coordinate system to the center of the polygon and the specified rotation surface corresponding to the origin, or as the angle between the horizontal direction of the polygon and the direction from the origin to the center.
[0135] Step 240: Based on the first pose and the second pose, render the second model image after editing the 3D model in the first model image.
[0136] In this embodiment, the first pose indicates the relative distribution relationship between the distribution points and the structural units, and the second pose indicates the distribution of the edited structural units in the global coordinate space. Through the first pose and the second pose, the rendering of the second model image obtained after editing can be achieved.
[0137] In some embodiments, based on the first pose and the second pose, the third pose of the edited distribution points in the second space is determined, and a second model image is generated by rendering based on the third poses corresponding to multiple distribution points. That is, the distribution of the edited distribution points in the global coordinate space can be determined based on the first pose and the second pose, and rendering is performed based on the distribution of the distribution points in the global coordinate space to obtain the second model image.
[0138] Optionally, the aforementioned third pose includes a third rotation matrix, a second scaling matrix, and a third position matrix. The third rotation matrix indicates the orientation of the edited distribution points relative to the second spatial coordinate system, representing the pose of the edited distribution points in the second spatial coordinate system. The second scaling matrix indicates the shape and size of the rendered distribution features corresponding to the edited distribution points. When the distribution points are Gaussian points, the second scaling matrix determines the width of the edited Gaussian function, i.e., the distribution range of the rendered distribution features in space. A smaller scaling ratio produces a more concentrated rendered distribution, while a larger scaling ratio results in a more dispersed rendered distribution. The third position matrix indicates the coordinate values of the edited distribution points in the second spatial coordinate system.
[0139] Since the properties of the first spatial coordinate system corresponding to the local coordinate space of the structural unit do not change before and after editing, after binding the rendering distribution features of the distribution points to the local coordinate space, the pose of the rendering distribution features of the edited distribution points in the global coordinate space can be determined by the pose of the distribution points in the local coordinate space and the pose of the edited structural unit in the global coordinate space. Thus, rendering can be achieved based on the pose of the rendering distribution features of the edited distribution points in the global coordinate space and the rendering distribution features corresponding to the distribution points. This transforms the editing of the 3D model into the editing of the rendering distribution features in the global coordinate space. At the same time, since the rendering result is based on the original rendering distribution features, and the structural unit only acts as a proxy for the editing operation, the accuracy of the structural unit will not affect the final model rendering effect, ensuring the model rendering effect of the second model image after editing.
[0140] In summary, when implementing editing operations on the 3D model in the first model image, the distribution points corresponding to the rendering distribution features of the 3D model are associated and bound with the structural units that make up the 3D model. This allows the acquisition of the first pose of the rendering distribution features in the local coordinate space corresponding to the structural unit. Then, the second pose of the structural unit in the global coordinate space after editing is determined based on the impact of the editing operation on the model structure of the 3D model. Since the first pose can preserve the relative positions between the distribution points bound to adjacent structural units, the corresponding bound distribution points can be manipulated and adjusted in real time after the structural unit changes according to the editing operation, thus obtaining the adjusted rendering distribution features. This results in a better rendering effect for the rendered second model image. Furthermore, since the first pose preserves the relative positions between the distribution points on adjacent structural units, the structural units only serve as a conversion medium for the editing operation. Therefore, the structural units obtained from the first model image have a relatively small impact on the image quality of the edited second model image, thus ensuring the image quality of the generated edited second model image.
[0141] In some embodiments, the process of the model image editing method provided in this application can be abstracted into three parts: 1. Structural unit extraction, 2. Distribution point binding, and 3. Distribution point editing. Please refer to... Figure 2 The diagram illustrates a flowchart of a method for editing a model image provided in an exemplary embodiment of this application, which includes the following steps 311 to 332.
[0142] Step 311: When an editing operation is received for the first model image, the first model image of the 3D model is obtained from multiple perspectives.
[0143] Optionally, the first model image under multiple perspectives can be implemented as including the image content of the 3D model under multiple perspectives in the first model image, that is, the first model image includes multiple perspective regions, and each perspective region corresponds to the image content of the 3D model under that perspective; Optionally, multiple first model images of the 3D model are obtained, and the multiple first model images correspond to the image content of the 3D model under different perspectives.
[0144] Optionally, the above editing operations include translation, rotation, scaling, stretching, deformation, topological structure change, and viewing angle change, etc., which are not limited here.
[0145] Step 312: Generate model rendering data based on the color distribution density of pixels in the first model image from multiple perspectives.
[0146] Schematic, model rendering data is used to indicate the rendering status of a 3D model. The model rendering data includes rendering distribution features corresponding to multiple distribution points on the 3D model. Optionally, the aforementioned rendering distribution features can be implemented as color distribution density; that is, each distribution point in the model rendering data corresponds to a distribution range, and within this distribution range, the model rendering data records the color distribution density corresponding to the distribution point.
[0147] In some embodiments, the aforementioned distribution points can be implemented as Gaussian points, and the rendering distribution features indicated by the Gaussian points are implemented as 3D Gaussian sputtering (3DGS), wherein 3DGS uses explicit 3D Gaussian points as its primary rendering primitives. A mathematically defined 3D Gaussian point is shown in Equation 1.
[0148] Formula 1:
[0149] Where μ is the 3D mean position coordinate of the 3D Gaussian point, i.e. the center of the Gaussian, x is any point in space, x-μ represents the relative position between the two, and G(x) is the density corresponding to x.
[0150] In addition, each Gaussian has an opacity o and a view-dependent color c represented by a set of spherical harmonics (SH).
[0151] Each 3D Gaussian point is described by a 3D mean position coordinate μ and a covariance matrix Σ. To ensure that the covariance matrix Σ retains its meaningful interpretation, it is parameterized as a unit quaternion q and a 3D scaling vector s, as defined in Equation 2.
[0152] Formula 2: ∑=RSS T R T
[0153] Where R is the rotation matrix of the 3D Gaussian point, and R and q can be uniquely determined to be mutually convertible, and S is the scaling matrix of the 3D Gaussian point, which is a matrix composed of 3D scaling vectors s.
[0154] To render the image from a specific viewpoint, a 3D Gaussian is projected onto the image plane, resulting in a 2D Gaussian. The 2D covariance matrix is approximated by Equation 3.
[0155] Formula 3: ∑′=JW∑W T J T
[0156] Where W represents the Jacobian matrix of the affine approximation of the view transformation, J represents the Jacobian matrix of the affine approximation of the perspective projection transformation, and the 2D mean is calculated through the projection matrix. Then, pixel colors are synthesized by alpha mixing of N ordered 2D Gaussians, the mixing process being represented by Equation 4.
[0157] Formula 4:
[0158] Here, α is obtained by multiplying the opacity o by the 2D covariance probability calculated from ∑′ and the pixel coordinates in the image space.
[0159] Step 313: Extract the model structure corresponding to the 3D model based on the model rendering data, determine multiple structural units, and form model representation data from multiple structural units.
[0160] In this embodiment of the application, after extracting the model rendering data corresponding to the three-dimensional model from the first model image, the model structure can be extracted from the summary of the model rendering data to obtain the model representation data.
[0161] Alternatively, the process of extracting data from the model representation can be implemented as at least one of the following:
[0162] The first method involves extracting multiple polygons corresponding to the 3D model based on the model rendering data, and using the polygon mesh composed of these polygons as model representation data. The polygon mesh is used to indicate the geometric structure of the 3D model.
[0163] That is, the structural units in the model representation data are implemented as polygons, and a polygon mesh composed of multiple polygons can indicate the geometry of the three-dimensional model.
[0164] In some embodiments, the extracted model representation data includes a list of vertices corresponding to multiple polygons, which contains the coordinates of all vertices of the multiple polygons that make up the 3D model in a second spatial coordinate system.
[0165] In some embodiments, the extracted model representation data also includes a list of polygons. For each face of a polygon in the 3D model, a polygon object is created, which contains the indices of all vertices that make up the polygon, thus obtaining the list of polygons.
[0166] In some embodiments, the extracted model representation data further includes a list of normals. Optionally, the list of normals includes smooth shading normals and / or planar shading normals, wherein the smooth shading normals are the normals corresponding to the vertices, and the planar shading normals are the normals calculated based on the entire polygon.
[0167] Alternatively, the polygon described above can be implemented as a triangle, quadrilateral, or other shape, without limitation.
[0168] Alternatively, the method for extracting polygon meshes from model rendering data can be implemented as at least one of the following:
[0169] 1. Screening for Poisson reconstruction.
[0170] 3DGS can be considered a type of point cloud, which allows the Poisson reconstruction algorithm to be used to extract the mesh. The 3D Gaussian used for 3DGS does not have a normal vector for reconstruction. In this embodiment, an additional Gaussian attribute, namely the normal vector n, is assigned to the 3D Gaussian, as shown in Formula 5.
[0171] Formula 5:
[0172] Where, d i n represents the depth of each Gaussian point. i The normal vector, α, is obtained by multiplying the opacity o by the 2D covariance probability calculated from ∑′ and pixel coordinates in image space, T. i As shown in Formula 4.
[0173] To optimize the normal vector, render the normal vector. With pseudo normal vector graph Alignment, this pseudo-normal map is based on the local plane assumption from the rendering depth The mesh is calculated. After training 3DGS using normal vector attributes, a filtered Poisson reconstruction algorithm is used to extract the mesh.
[0174] In some embodiments, due to the rendering normal vector of the 3D Gaussian point It is not accurate enough, especially for invisible points. The position of the polygon corresponding to the polygon in the extracted polygon mesh may not be aligned with the input Gaussian point. Polygons that are more than a certain threshold away from their nearest Gaussian point can be removed.
[0175] 2. Marching Cubes for Gaussian Sputtering.
[0176] In the DreamGaussian method, the alpha values of neighboring Gaussian points are aggregated into a composite density value for the sampling points of the traveling cube. Gaussian points classified as neighbors are located within predefined local voxels. DreamGaussian is a generative framework for efficient 3D content creation, capable of rapidly generating high-quality 3D models from single-view images or text descriptions. This method primarily utilizes 3D Gaussian sputtering technology, achieving 3D model generation by optimizing the score-based distillation sampling (SDS) loss function. DreamGaussian significantly improves the efficiency of 3D content generation while maintaining generation quality, enabling the generation of realistic 3D models with explicit meshes and texture maps from single images.
[0177] However, DreamGaussian suffers from neglecting small and weak structures. To address this issue, this embodiment employs a sphere query nearest neighbor search method, searching only for the nearest Gaussian point of each grid sampling point. If a sampling point has a neighbor within a predefined distance, it is assigned a specified density value (e.g., 1); otherwise, the density value is 0. Then, a moving cube algorithm is used to extract the geometry. This supplementary sphere query nearest neighbor search method enables DreamGaussian to extract small and weak structures.
[0178] 3. Neural hidden aspects.
[0179] In this embodiment, the method proposed by Neural Surface Reconstruction (NeuS) is used to extract high-quality surfaces from the implicit representation. NeuS extracts the surface of a 3D model as a zero-level set of the signed distance function (SDF), which has been a robust and widely used method in the field of neural surface reconstruction.
[0180] However, surfaces derived from NeuS may contain too many polygonal surfaces, resulting in a large amount of data in the resulting polygonal mesh. A large number of polygons can negatively impact training and inference speed. Therefore, in this embodiment, a mesh cleaning process is used to eliminate noise fluctuations, while mesh clipping is used to reduce the number of polygons (e.g., reducing the number of polygons to approximately 300K). This avoids the problems that may arise from editing a massive number of polygons, while maintaining high-fidelity rendering effects after reducing the number of polygons.
[0181] The second method involves extracting point cloud data corresponding to the 3D model based on the model rendering data, and using the point cloud data as model representation data. The point cloud data is used to indicate the geometric coordinates of the points that make up the 3D model.
[0182] That is, the structural units in the model representation data are realized as points in the point cloud data. These points have geometric coordinates in the second spatial coordinate system corresponding to the three-dimensional model. The relative positional relationship between the points can indicate the model structure of the three-dimensional model.
[0183] In some embodiments, extracting point cloud data corresponding to a 3D model based on model rendering data can be achieved by: determining the mean and standard deviation of the 3D Gaussian points, determining the Gaussian distribution of the 3D Gaussian points, generating a point set using a random number generator based on the Gaussian distribution, and creating a data structure to store the coordinates (x, y, z) of each point in the point set in the second spatial coordinate system, thereby obtaining the model representation data.
[0184] Since 3DGS uses data represented by 3D Gaussian points, which is also represented in the form of point clouds, the generation efficiency of model representation data can be improved by using point cloud data that indicates the model structure as editing proxies for 3D Gaussian points that indicate the distribution of rendering features.
[0185] The third method involves extracting multiple voxels corresponding to the 3D model based on the model rendering data, and using the voxel model composed of multiple voxels as model representation data. The voxel model is used to indicate the geometric distribution of the volume pixels that make up the 3D model.
[0186] That is, the structural units in the model representation data are implemented as voxels, and a voxel model composed of multiple voxels is used to indicate the geometric structure of the three-dimensional model.
[0187] In some embodiments, extracting multiple voxels corresponding to a 3D model based on model rendering data can be implemented as follows: determining the voxel size; defining an initial 3D voxel mesh based on the voxel size, wherein the voxel mesh includes multiple voxels, and the size of the voxel determines the resolution of the voxel mesh, with smaller voxels providing higher detail; for each 3D Gaussian point, determining which voxel in the voxel mesh it maps to based on its position in the second spatial coordinate system, which typically involves rounding or truncating the coordinates of the cutoff point to the nearest integer to match the voxel boundary; after determining the voxel to which the 3D Gaussian point maps, the value of the voxel can be updated; voxels in the initial voxel mesh that have 3D Gaussian point mappings are retained, while voxels that do not have 3D Gaussian point mappings are filtered out, finally obtaining the voxel model as model representation data.
[0188] Voxel models can more accurately reflect the geometry of a model because they provide denser geometric information. Therefore, using them as proxies for editing 3D Gaussian points ensures that the editing results conform to the geometry of the original model, guaranteeing the stability of the editing results.
[0189] Step 321: Determine the first pose of the rendering distribution features corresponding to the multiple distribution points in the first space.
[0190] The first space is a local coordinate space constructed based on structural units; that is, the first space is the space corresponding to the first spatial coordinate system constructed based on structural units.
[0191] Schematic illustration: A first space corresponding to a structural unit is constructed, and multiple distribution points are respectively bound to the first space of the structural unit, determining the first pose of the distribution points in the first space. In this embodiment, taking the distribution points as 3D Gaussian points and the structural unit as a polygon as an example, it is necessary to attach the 3D Gaussian points to the polygon mesh.
[0192] To illustrate, in order to maintain high-fidelity rendering results after editing a 3D model, the key is to maintain local rigidity and preserve the relative positions between Gaussian distributions, whether for the mean or rotation of the 3D Gaussian points. In this embodiment, a local coordinate system is defined for each polygon in the polygonal mesh, i.e., a first spatial coordinate system.
[0193] In one example, taking the polygon as a triangle, the first axis direction of the first space corresponding to the triangle is defined as the direction of the first side, the second axis direction of the first space is defined as the normal direction of the triangle, and the third axis direction of the first space is defined as the cross product of the first and second axes. Figure 4 As shown, it illustrates a schematic diagram of the spatial coordinate system corresponding to a triangle provided in an exemplary embodiment of this application. For a target triangle 410 in the triangular mesh, the x-axis 411 is defined according to the direction of the first side of the target triangle 410, the y-axis 412 is defined according to the normal direction of the target triangle 410, and the z-axis 413 is determined according to the cross product between the x-axis 411 and the y-axis 412.
[0194] Optionally, the first pose mentioned above includes a first rotation matrix, a first scaling matrix, and a first position matrix of the distribution points in the first space.
[0195] Step 322: Based on the editing operation and the model representation data, determine the second pose of the edited structural unit in the second space.
[0196] The aforementioned second space is a global coordinate space constructed based on the first model image; that is, the second space is the space corresponding to the second space coordinate system constructed based on the three-dimensional model of the first model image. In some embodiments, the aforementioned second space coordinate system can be implemented as a world coordinate system, a Gaussian coordinate system, etc.
[0197] In some embodiments, determining the second pose of the edited structural unit in the second space based on the editing operation and the model representation data can be achieved by: adjusting the structural relationship between multiple structural units in the model representation data based on the editing operation to obtain the adjusted model representation data, and determining the second pose of multiple adjusted structural units in the second space based on the adjusted model representation data.
[0198] After constructing the first spatial coordinate system, the pose of the triangle in the first spatial coordinate system relative to the second spatial coordinate system can be used as the pose of the triangle in the second space. Optionally, the aforementioned second pose includes the second rotation matrix and the second position matrix of the edited structural unit in the second space.
[0199] Specifically, the second rotation matrix R of the triangle in the second space t This can be represented by Formula Six.
[0200] Formula Six:
[0201] Where v1 represents the first vertex of the triangle, v2 represents the second vertex of the triangle, and n t The normal vector of the triangle. These represent rotations corresponding to different axes. Wherein, n... t It is obtained by calculation using Formula 7.
[0202] Public Notice 7:
[0203] Here, v3 represents the third vertex of the triangle.
[0204] Optionally, the second position matrix of the triangle is used to indicate at least one of the following: the coordinates of the center of the triangle in the second space, the coordinates of the centroid of the triangle in the second space, the coordinates of the centroid of the triangle in the second space, the coordinates of the vertices of the triangle in the second space, and the coordinates of the origin of the first spatial coordinate system corresponding to the triangle.
[0205] Step 331: Based on the first pose and the second pose, determine the third pose of the edited distribution points in the second space.
[0206] In this embodiment, the first pose includes a first rotation matrix, a first scaling matrix, and a first position matrix of the distribution points in the first space, and the second pose includes a second rotation matrix and a second position matrix of the edited structural units in the second space.
[0207] In a schematic manner, a third rotation matrix of the edited distribution points in the second space is determined based on a first rotation matrix and a second rotation matrix; a second scaling matrix of the edited distribution points in the second space is determined based on a first scaling matrix; and a third position matrix of the edited distribution points in the second space is determined based on a first position matrix and a second position matrix. The third pose of the edited distribution points in the second space is composed of the third rotation matrix, the second scaling matrix, and the third position matrix.
[0208] In one example, taking the distribution points as 3D Gaussian points and the structural units as triangles, the third rotation matrix R of the edited 3D Gaussian points in the second space is shown in Formula 8.
[0209] Formula 8: R = R t R l
[0210] Among them, R l Let R be the first rotation matrix. t This is the second rotation matrix.
[0211] The second scaling matrix s of the edited 3D Gaussian points in the second space is shown in Equation 9.
[0212] Formula 9: s = s l
[0213] Among them, s lThis is the first scaling matrix.
[0214] The third position matrix of the edited 3D Gaussian points in the second space is shown in Formula 10.
[0215] Formula 10: μ = R t μ l +μ t
[0216] Among them, R t Let μ be the second rotation matrix. l Let μ be the first position matrix. t This is the second position matrix.
[0217] Step 332: Render and generate the second model image based on the third pose corresponding to multiple distribution points.
[0218] Optionally, the distribution points correspond to color density distributions. Schematic, the position distribution of the edited distribution points in the second space is determined based on the third pose corresponding to multiple distribution points. The edited 3D model is rendered based on the color density distribution and position distribution corresponding to the distribution points to obtain the second model image.
[0219] Optionally, the distribution points correspond to light intensity distributions. Schematic, the position distribution of the edited distribution points in the second space is determined based on the third pose corresponding to multiple distribution points. The edited 3D model is then rendered based on the light intensity distribution and position distribution corresponding to the distribution points to obtain the second model image.
[0220] Optionally, the distribution points correspond to the model texture distribution. Schematic, the position distribution of the edited distribution points in the second space is determined based on the third pose corresponding to multiple distribution points. The edited 3D model is rendered based on the model texture distribution and position distribution corresponding to the distribution points to obtain the second model image.
[0221] Please refer to Figure 5This illustration shows a schematic flowchart of a model image editing method provided in an exemplary embodiment of this application. The model image editing method provided in this embodiment can be abstracted into three parts: mesh extraction 510, 3D Gaussian point-mesh binding 520, and 3D Gaussian point editing 530. Mesh extraction 510 involves extracting 3D Gaussian points from a first model image 511 under multiple viewpoints, obtaining 3D Gaussian points and their corresponding Gaussian distributions 512. Triangular meshes are then extracted based on the 3D Gaussian points and their corresponding Gaussian distributions 512, resulting in triangular meshes 513. 3D Gaussian point-mesh binding 520 involves binding the triangular meshes 513 and the 3D Gaussian points and their corresponding Gaussian distributions 512, that is, binding the 3D Gaussian points and their corresponding Gaussian distributions 512 to the triangular space corresponding to the triangles in the triangular meshes 513. The 3D Gaussian point editing 530 is implemented by editing the pose corresponding to the triangular mesh 513 to obtain the edited 3D model structure 531, and by editing the 3D Gaussian points bound to the triangular mesh 513 and their corresponding Gaussian distributions 512 according to the edited triangular mesh, and by rendering the edited 3D Gaussian points and their corresponding Gaussian distributions to obtain the corresponding second model image 532.
[0222] In summary, when implementing editing operations on the 3D model in the first model image, the distribution points corresponding to the rendering distribution features of the 3D model are associated and bound with the structural units that make up the 3D model. This allows the acquisition of the first pose of the rendering distribution features in the local coordinate space corresponding to the structural unit. Then, the second pose of the structural unit in the global coordinate space after editing is determined based on the impact of the editing operation on the model structure of the 3D model. Since the first pose can preserve the relative positions between the distribution points bound to adjacent structural units, the corresponding bound distribution points can be manipulated and adjusted in real time after the structural unit changes according to the editing operation, thus obtaining the adjusted rendering distribution features. This results in a better rendering effect for the rendered second model image. Furthermore, since the first pose preserves the relative positions between the distribution points on adjacent structural units, the structural units only serve as a conversion medium for the editing operation. Therefore, the structural units obtained from the first model image have a relatively small impact on the image quality of the edited second model image, thus ensuring the image quality of the generated edited second model image.
[0223] In some embodiments, the model image editing method provided in this application is implemented through a pre-trained machine learning model. The machine learning model is used to predict the model editing effect of the editing operation based on model representation data and model rendering data, thereby obtaining a second model image. The pre-trained machine learning model includes a rendering feature extraction subnetwork, a mesh extraction subnetwork, and an editing subnetwork. The rendering feature extraction subnetwork extracts 3D Gaussian points and their corresponding Gaussian distributions from the first model image to obtain model rendering data. The mesh extraction subnetwork extracts polygon meshes from the model rendering data to obtain model representation data. The editing subnetwork binds the 3D Gaussian points and their corresponding Gaussian distributions to the local space corresponding to the polygons, and adjusts the topological structure between polygons in the polygon mesh based on the editing operation, thereby adjusting the 3D Gaussian points and their corresponding Gaussian distributions bound to the polygon mesh, and outputting the edited second model image. Please refer to... Figure 6 The diagram illustrates a flowchart of a model image editing method provided in an exemplary embodiment of this application, which includes the following steps 601 to 605.
[0224] Step 601: When an editing operation is received for the first model image, the first model image of the 3D model is obtained from multiple perspectives.
[0225] Optionally, the above editing operations include translation, rotation, scaling, stretching, deformation, topological structure change, and viewing angle change, etc., which are not limited here.
[0226] Optionally, the first model image under multiple perspectives can be implemented as including the image content of the 3D model under multiple perspectives in the first model image, that is, the first model image includes multiple perspective regions, and each perspective region corresponds to the image content of the 3D model under that perspective; Optionally, multiple first model images of the 3D model are obtained, and the multiple first model images correspond to the image content of the 3D model under different perspectives.
[0227] Step 602: Input the editing instruction text corresponding to the editing operation and the first model image into the pre-trained machine learning model.
[0228] This is illustrative of generating corresponding editing instruction text based on an editing operation. In some embodiments, the editing instruction text corresponding to the editing operation is retrieved from a preset instruction set (prompt) based on the editing operation, wherein the instruction text in the preset instruction set is content preset during the training of the machine learning model.
[0229] Step 603: Extract the rendering distribution features in the first model image through the rendering feature extraction sub-network in the machine learning model to obtain the model rendering data.
[0230] Indicatively, the rendering feature extraction subnetwork is used to extract rendering distribution features in the first model image. In this embodiment, the rendering feature extraction subnetwork extracts 3D Gaussian points and their corresponding Gaussian distributions from the first model image using 3DGS technology.
[0231] Optionally, the above-mentioned rendering feature extraction subnetwork can be implemented as Neural Radiance Fields (NeRF), 3D Deformation Models, etc., without limitation.
[0232] Step 604: Extract the model structure from the model rendering data using the grid extraction sub-network in the machine learning model to obtain the model representation data.
[0233] Indicatively, the mesh extraction subnetwork extracts the model structure from the model rendering data to obtain the polygonal mesh corresponding to the 3D model. Optionally, the mesh extraction subnetwork extracts the polygonal mesh by selecting at least one of Poisson reconstruction, a traveling cube for Gaussian sputtering, and a neural hidden surface.
[0234] Optionally, the above-mentioned mesh extraction subnetwork can be implemented as DreamGaussian, SuGaR, Neural Implicit Surface Reconstruction with 3D GaussianSplatting Guidance (NeuSG) network, etc., and is not limited here.
[0235] Step 605: The editing sub-network in the machine learning model, based on the editing instruction text, model rendering data, and model representation data, binds the distribution points in the model rendering data to the structural units in the model representation data to edit the rendering distribution features in the first model image, thereby generating the second model image.
[0236] In this embodiment, the editing sub-network in the machine learning model extracts the pose features of the distribution points in the first space based on the model rendering data to obtain the first pose of the distribution points. Based on the extracted first pose of the distribution points and the second pose of the edited structural units in the second space, the editing sub-network determines the third pose corresponding to the edited distribution points, and generates a second model image based on the rendered distribution features and the third pose.
[0237] In some embodiments, the training process of the above-mentioned machine learning model includes: acquiring sample model images of the sample 3D model from multiple perspectives, training editing instructions, and sample editing model images corresponding to the training editing instructions; inputting the sample model images and training editing instructions into the machine learning model to obtain a predicted editing model image; and iteratively learning the correlation between the first pose of the sample distribution points in the first space, the second pose of the sample distribution points in the second space, and the third pose of the sample structural units in the second space, based on the difference between the predicted editing model image and the sample editing model image, to obtain a pre-trained machine learning model.
[0238] In one example, taking the structural unit as a triangle and the distribution points as 3D Gaussian points, the optimization objective for training the edit subnetwork during the machine learning model training process is the first pose of the 3D Gaussian point in the first space corresponding to the triangle, i.e., the first position matrix μ. l and the first rotation matrix R l The aforementioned optimization objectives are indicated by Formulas 11, 12, and 13.
[0239] Formula 11: R = R t R l
[0240] Among them, R t R is the second rotation matrix of the edited triangle in the second spatial coordinate system. l R is the first rotation matrix of the 3D Gaussian point in the first spatial coordinate system corresponding to the triangle, and R is the third rotation matrix of the edited 3D Gaussian point in the second spatial coordinate system.
[0241] Formula 12: s = βes l
[0242] Among them, s l Let be the first scaling matrix of the 3D Gaussian points before editing, s be the second scaling matrix of the 3D Gaussian points after editing, β be a hyperparameter, and e be an adaptive vector, where e = [e1, e2, e3]. The design of e is to ensure that the global scaling s is proportional to the shape of the triangle. The first axis is along the first side, so e1 is designed to be the length l1 of the first side of the triangle. The second axis is along the normal direction, so e2 is set to (0.5 * (e1 + e3)). The third axis is perpendicular to the first side, and e3 is set to the average length of the second and third sides (0.5 * (l1 + l3)).
[0243] Formula 13: μ=eR t μ l +μ t
[0244] Where, μ l Let μ be the first position matrix of the 3D Gaussian point in the first spatial coordinate system corresponding to the triangle. t is the second position matrix of the edited triangle in the second spatial coordinate system, and μ is the third position matrix of the edited 3D Gaussian point in the second spatial coordinate system.
[0245] During the training of the edited subnetwork, R will be used. l s l and μ l Adjust these parameters as needed for optimization.
[0246] Optionally, the training processes of the rendering feature extraction subnetwork, grid extraction subnetwork, and editing subnetwork included in the machine learning model can be independent of each other, or the machine learning model can be trained as a whole.
[0247] In one example, taking the training of a machine learning model as a whole as an example, the input of the machine learning model during the training process includes sample model images and training editing instructions. During the training process, the machine learning model generates a predicted editing model image based on the sample model images and training editing instructions. The difference between the sample editing model image corresponding to the sample model image and the predicted editing model image is used as a reference to iteratively adjust the model parameters of the machine learning model, thereby obtaining a pre-trained machine learning model.
[0248] Schematic, the difference between the sample edited model image and the predicted edited model image can be determined by a first loss function. Optionally, the first loss function can be implemented as at least one of the following: cross-entropy loss, mean squared error loss, log loss, absolute value loss (L1 loss), structural similarity index Measure (SSIM), etc., without limitation.
[0249] In some embodiments, in order to enhance the effect of the model image output by the machine learning model, the image quality of the predicted edit model image is added to the loss evaluation objective during the training process. That is, the difference between the sample edit model image corresponding to the sample model image and the predicted edit model image, as well as the image quality of the predicted edit model image, are used as references to iteratively adjust the model parameters of the machine learning model, thereby obtaining a pre-trained machine learning model.
[0250] Schematic illustration: The image quality of the predicted edit model image can be determined by a second loss function. Optionally, the second loss function can be implemented as at least one of the following: cross-entropy loss function, mean squared error loss function, log loss function, absolute value loss function, structural similarity index loss function, etc., without limitation.
[0251] In one example, the training loss of the machine learning model is implemented as L1 loss, SSIM loss, and additional masked cross-entropy loss in 3DGS.
[0252] In summary, when implementing editing operations on the 3D model in the first model image, the distribution points corresponding to the rendering distribution features of the 3D model are associated and bound with the structural units that make up the 3D model. This allows the acquisition of the first pose of the rendering distribution features in the local coordinate space corresponding to the structural unit. Then, the second pose of the structural unit in the global coordinate space after editing is determined based on the impact of the editing operation on the model structure of the 3D model. Since the first pose can preserve the relative positions between the distribution points bound to adjacent structural units, the corresponding bound distribution points can be manipulated and adjusted in real time after the structural unit changes according to the editing operation, thus obtaining the adjusted rendering distribution features. This results in a better rendering effect for the rendered second model image. Furthermore, since the first pose preserves the relative positions between the distribution points on adjacent structural units, the structural units only serve as a conversion medium for the editing operation. Therefore, the structural units obtained from the first model image have a relatively small impact on the image quality of the edited second model image, thus ensuring the image quality of the generated edited second model image.
[0253] In this embodiment, a pre-trained machine learning model is used to edit the 3D model in the model image. Since the machine learning model can learn the pose features of the rendering distribution features of different 3D models in the local coordinate space corresponding to the structural unit during the training process, the machine learning model can quickly extract the pose of the rendering distribution features of the 3D model in the model image in the local coordinate space of its corresponding structural unit during the application process, thereby improving the editing efficiency of the model image.
[0254] The editing effect achieved by the model image editing method provided in this application embodiment is compared with the editing effect achieved by the model image editing method implemented by SuGaR in related technologies. Figure 7As shown, this diagram illustrates the editing effect of a model image editing method provided in this embodiment. Rotation and stretching operations are performed on the original chair model image A701, and SuGaR outputs model image A702. The method of this embodiment outputs model image B703. Because SuGaR cannot adapt to compensate for geometric errors, the output model image A702 suffers from geometric expansion at the chair's boundaries, resulting in poor image quality. In contrast, the rendering result of model image B703 has more accurate model boundaries. Similarly, rotation and narrowing operations are performed on the original Lego model image B704, and SuGaR outputs model image C705. The method of this embodiment outputs model image D706. Model image C705, like model image A702, suffers from geometric expansion at its boundaries. Compared to model image C705, the Lego in model image D706 has more accurate model boundaries.
[0255] In another example, such as Figure 8 As shown, this diagram illustrates another editing effect of the model image editing method provided in this embodiment. A bending operation is performed on the original model image C801 of the potted plant, and SuGaR outputs model image E802. The method of this embodiment outputs model image F803. Because SuGaR cannot adapt to compensate for geometric errors, the output model image E802 has geometric missing parts at the boundaries of areas such as the flowerpot, resulting in poor image quality. In contrast, the rendering result corresponding to model image F803 has more accurate model boundaries. Similarly, a rotation and upward bending operation are performed on the original model image D804 of the microphone, and SuGaR outputs model image G805. The method of this embodiment outputs model image H806. Model image G805, like model image E802, suffers from boundary geometric expansion. Compared to model image G805, the microphone in model image H806 has more accurate model boundaries.
[0256] It is worth noting that the model image editing method provided in this application embodiment can also achieve good results in local operations and physical simulation. For example... Figure 9 As shown, it illustrates an example of the application of the method of this application to software simulation. Figure 9The method for editing model images provided in this application attempts to mix the red and yellow sauces of the hot dog in the original model image A911. The resulting edited model image A912 is shown in the figure. The corresponding editing effect is satisfactory and the rendering result is reasonable. In another example, in the original model image B921, one of the cymbals in the drum kit is rearranged and the other cymbal is elastically deformed to obtain the edited model image B922. After rearrangement and elastic deformation, the model image B922 still retains a photorealistic rendering effect.
[0257] Optionally, the model image editing method provided in this application embodiment can be applied to computer products in the following ways:
[0258] The first type is 3D modeling and design: The model image editing method provided in this application embodiment can be used to create and modify 3D models in the model image based on the provided model image, enabling designers to easily adjust the shape, texture and color. It has broad application prospects in product prototype design, architectural visualization and game asset creation, and can bring great value and convenience to these fields.
[0259] The second type is virtual reality and augmented reality: In virtual reality (VR) and augmented reality (AR) applications, the model image editing method provided in this application embodiment can be used to modify and adjust the three-dimensional model in the scene image in real time after acquiring the scene image. This allows users to interact more naturally and intuitively in the virtual environment provided by the device, improving the immersive experience.
[0260] Thirdly, in film and animation production: the model image editing method provided in this application can be used for scene design, character modeling, and animation in film and animation production. By using the model image editing method provided in this application, artists can easily create and modify complex 3D scenes based on images, bringing a more realistic visual effect to the audience.
[0261] Fourthly, product display and marketing: For online product display and marketing, the model image editing method provided in this application embodiment can be used to create realistic 3D renderings, thereby obtaining 3D product images, enabling potential customers to better understand the product's appearance and function. Furthermore, the model image editing method provided in this application embodiment can provide customers with product customization functions, allowing them to easily customize and personalize the provided product model images, thereby increasing customer engagement.
[0262] Fifth, Computer-Aided Design (CAD): In the field of computer-aided design, the model image editing method provided in this application embodiment can be used to create and modify three-dimensional models of imported model images.
[0263] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0264] Please refer to Figure 10 The diagram illustrates a structural block diagram of a model image editing device provided in an exemplary embodiment of this application. The device includes the following modules:
[0265] The acquisition module 1010 is used to acquire model representation data and model rendering data corresponding to the three-dimensional model in the first model image when an editing operation is received for the first model image. The model representation data includes structural units that construct the three-dimensional model, and the model rendering data includes rendering distribution features corresponding to multiple distribution points on the three-dimensional model.
[0266] The first determining module 1020 is used to determine the first pose of the rendering distribution features corresponding to the plurality of distribution points in a first space, wherein the first space is a local coordinate space constructed based on the structural unit.
[0267] The second determining module 1030 is used to determine the second pose of the edited structural unit in the second space based on the editing operation and the model representation data, wherein the second space is a global coordinate space constructed based on the first model image;
[0268] The rendering module 1040 is used to render a second model image, which is an edited version of the three-dimensional model in the first model image, based on the first pose and the second pose.
[0269] In some alternative embodiments, such as Figure 11 As shown, the first determining module 1020 further includes:
[0270] Construction unit 1021 is used to construct the first space corresponding to the structural unit;
[0271] Binding unit 1022 is used to bind the plurality of distribution points to the first space of the structural unit respectively, and to determine the first pose of the distribution points in the first space.
[0272] In some alternative embodiments, the method is implemented using a pre-trained machine learning model;
[0273] The binding unit 1022 is further configured to extract the pose features of the distribution points in the first space based on the model rendering data through the machine learning model, and obtain the first pose of the distribution points.
[0274] In some optional embodiments, the acquisition module 1010 is used to acquire sample model images of the sample 3D model from multiple perspectives, training editing instructions, and sample editing model images corresponding to the training editing instructions;
[0275] The device further includes: a training module 1050, used to input the sample model image and the training editing instructions into the machine learning model to obtain a predicted editing model image; based on the difference between the predicted editing model image and the sample editing model image, iteratively learning the correlation between the first pose of the sample distribution points corresponding to the sample model image in the first space, the second pose of the sample distribution points in the second space, and the third pose of the sample structural units in the second space to obtain the pre-trained machine learning model.
[0276] In some optional embodiments, the model representation data includes multiple structural units, and the model rendering data includes N distribution points, where N is a positive integer;
[0277] The binding unit 1022 is further configured to use the spatial origin of the first space corresponding to the plurality of structural units as the cluster center to cluster the N distribution points to obtain clusters corresponding to the plurality of structural units; for the i-th cluster, take K distribution points in the i-th cluster as distribution points corresponding to the i-th structural unit, where i and K are positive integers and K < N; bind the K distribution points to the first space of the i-th structural unit; and determine the first pose of the K distribution points in the first space of the i-th structural unit.
[0278] In some alternative embodiments, the structural unit includes a polygonal mesh;
[0279] The construction unit 1021 is further configured to, for the i-th polygon in the polygonal mesh, take a specified point on the plane where the i-th polygon is located as the origin of the first spatial coordinate system corresponding to the first space, where i is a positive integer; take a specified edge in the i-th polygon as the first axis direction of the first spatial coordinate system; take the normal direction of the i-th polygon as the second axis direction of the first spatial coordinate system; and determine the third axis direction of the first spatial coordinate system based on the first axis direction and the second axis direction, wherein the third axis direction is perpendicular to the first axis direction and the third axis direction is perpendicular to the second axis direction.
[0280] In some optional embodiments, the model representation data includes multiple structural units;
[0281] The second determining module 1030 is further configured to adjust the structural relationships between the plurality of structural units in the model representation data based on the editing operation to obtain adjusted model representation data; and to determine the second pose of the plurality of adjusted structural units in the second space based on the adjusted model representation data.
[0282] In some alternative embodiments, the plurality of structural units are implemented as a polygonal mesh composed of a plurality of polygons, each polygon including at least three vertices;
[0283] The second determining module 1030 is further configured to, for the i-th polygon, generate the second pose of the i-th polygon in the second space based on the coordinate values of the vertices corresponding to the i-th polygon in the second spatial coordinate system corresponding to the second space, where i is a positive integer.
[0284] In some alternative embodiments, the apparatus further includes:
[0285] The third determining module 1060 is used to determine the third pose of the edited distribution points in the second space based on the first pose and the second pose;
[0286] The rendering module 1040 is further configured to render and generate the second model image based on the third pose corresponding to the plurality of distribution points.
[0287] In some optional embodiments, the first pose includes a first rotation matrix, a first scaling matrix, and a first position matrix of the distribution points in the first space, and the second pose includes a second rotation matrix and a second position matrix of the edited structural unit in the second space;
[0288] The third determining module 1060 is further configured to determine a third rotation matrix of the edited distribution points in the second space based on the first rotation matrix and the second rotation matrix; determine a second scaling matrix of the edited distribution points in the second space based on the first scaling matrix; determine a third position matrix of the edited distribution points in the second space based on the first position matrix and the second position matrix; and form the third pose of the edited distribution points in the second space by the third rotation matrix, the second scaling matrix and the third position matrix.
[0289] In some optional embodiments, the distribution points correspond to a color density distribution;
[0290] The rendering module 1040 is further configured to determine the position distribution of the edited distribution points in the second space based on the third pose corresponding to the plurality of distribution points respectively; and to render the edited three-dimensional model based on the color density distribution corresponding to the distribution points and the position distribution to obtain the second model image.
[0291] In some optional embodiments, the acquisition module 1010 includes:
[0292] The acquisition unit 1011 is used to acquire a first model image of the three-dimensional model from multiple perspectives;
[0293] The generation unit 1012 is used to generate the model rendering data based on the color distribution density of pixels in the first model image under the multiple viewpoints; extract the model structure corresponding to the three-dimensional model based on the model rendering data, determine multiple structural units, and form the model representation data by the multiple structural units.
[0294] In some optional embodiments, the generation unit 1012 is further configured to extract multiple polygons corresponding to the three-dimensional model based on the model rendering data, and use the polygon mesh composed of the multiple polygons as the model representation data, wherein the polygon mesh is used to indicate the geometric structure of the three-dimensional model.
[0295] In some optional embodiments, the generation unit 1012 is further configured to extract point cloud data corresponding to the three-dimensional model based on the model rendering data, and use the point cloud data as the model representation data, wherein the point cloud data is used to indicate the geometric coordinates of the points that make up the three-dimensional model;
[0296] In some optional embodiments, the generation unit 1012 is further configured to extract multiple voxels corresponding to the three-dimensional model based on the model rendering data, and use the voxel model composed of the multiple voxels as the model representation data, wherein the voxel model is used to indicate the geometric distribution of the volume pixels that make up the three-dimensional model.
[0297] It should be noted that the model image editing device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the model image editing device and the model image editing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0298] Figure 12 This illustration shows a schematic diagram of the structure of a server provided in an exemplary embodiment of this application. Specifically, it includes the following structure.
[0299] Server 1200 includes a Central Processing Unit (CPU) 1201, a system memory 1204 including Random Access Memory (RAM) 1202 and Read Only Memory (ROM) 1203, and a system bus 1205 connecting the system memory 1204 and the CPU 1201. Server 1200 also includes a mass storage device 1206 for storing an operating system 1213, application programs 1214, and other program modules 1215.
[0300] Mass storage device 1206 is connected to central processing unit 1201 via a mass storage controller (not shown) connected to system bus 1205. Mass storage device 1206 and its associated computer-readable media provide non-volatile storage for server 1200. That is, mass storage device 1206 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drives.
[0301] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1204 and mass storage device 1206 described above can be collectively referred to as memory.
[0302] According to various embodiments of this application, server 1200 can also be connected to a remote computer on a network, such as the Internet. That is, server 1200 can be connected to network 1212 via network interface unit 1211 connected to system bus 1205, or network interface unit 1211 can be used to connect to other types of networks or remote computer systems (not shown).
[0303] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.
[0304] Embodiments of this application also provide a computer device including a processor and a memory. The memory stores at least one instruction, at least one program, a code set, or an instruction set. The processor loads and executes the at least one instruction, at least one program, a code set, or an instruction set to implement the model image editing method provided in the above-described method embodiments. Optionally, the computer device may be a terminal or a server.
[0305] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the model image editing method provided in the above-described method embodiments.
[0306] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the model image editing methods described in the above embodiments.
[0307] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0308] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0309] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for editing model images, characterized in that, The method includes: When an editing operation is received for the first model image, model representation data and model rendering data corresponding to the three-dimensional model in the first model image are obtained. The model representation data includes structural units that construct the three-dimensional model, and the model rendering data includes rendering distribution features corresponding to multiple distribution points on the three-dimensional model. The first pose of the rendering distribution features corresponding to the multiple distribution points in the first space is determined, where the first space is a local coordinate space constructed based on the structural unit. Based on the editing operation and the model representation data, the second pose of the edited structural unit in the second space is determined, where the second space is a global coordinate space constructed based on the first model image. Based on the first pose and the second pose, render a second model image after editing the 3D model in the first model image.
2. The method according to claim 1, characterized in that, Determining the first pose of the rendering distribution features corresponding to the plurality of distribution points in the first space includes: Construct the first space corresponding to the structural unit; The plurality of distribution points are respectively bound to the first space of the structural unit, and the first pose of the distribution points in the first space is determined.
3. The method according to claim 2, characterized in that, The method is implemented using a pre-trained machine learning model; The step of binding the plurality of distribution points to the first space of the structural unit and determining the first pose of the distribution points in the first space includes: The machine learning model extracts the pose features of the distribution points in the first space based on the model-rendered data to obtain the first pose of the distribution points.
4. The method according to claim 3, characterized in that, The training process of the machine learning model includes: Acquire sample model images of the 3D model from multiple perspectives, training editing instructions, and sample edit model images corresponding to the training editing instructions; The sample model image and the training editing instructions are input into the machine learning model to obtain the predicted editing model image; Based on the difference between the predicted editing model image and the sample editing model image, the correlation between the first pose of the sample distribution points in the first space, the second pose of the sample distribution points in the second space, and the third pose of the sample structural units in the second space corresponding to the sample model image is iteratively learned to obtain the pre-trained machine learning model.
5. The method according to claim 2, characterized in that, The model representation data includes multiple structural units, and the model rendering data includes N distribution points, where N is a positive integer. The step of binding the plurality of distribution points to the first space of the structural unit and determining the first pose of the distribution points in the first space includes: Using the spatial origin of the first space corresponding to each of the multiple structural units as the cluster center, the N distribution points are clustered to obtain the clusters corresponding to each of the multiple structural units. For the i-th cluster, the K distribution points in the i-th cluster are taken as the distribution points corresponding to the i-th structural unit, where i and K are positive integers and K < N; Bind the K distribution points to the first space of the i-th structural unit; The first poses of the K distribution points in the first space of the i-th structural unit are determined respectively.
6. The method according to any one of claims 2 to 5, characterized in that, The structural unit includes a polygonal grid; The construction of the first space corresponding to the structural unit includes: For the i-th polygon in the polygonal mesh, a designated point on the plane where the i-th polygon is located is taken as the origin of the first spatial coordinate system corresponding to the first space, where i is a positive integer; The specified edge in the i-th polygon is taken as the first axis direction of the first spatial coordinate system; The normal direction of the i-th polygon is taken as the second axis direction of the first spatial coordinate system; The third axis direction of the first spatial coordinate system is determined based on the first axis direction and the second axis direction, wherein the third axis direction is perpendicular to the first axis direction and the second axis direction.
7. The method according to any one of claims 1 to 5, characterized in that, The model represents data including multiple structural units; The step of determining the second pose of the edited structural unit in the second space based on the editing operation and the model representation data includes: Based on the editing operation, the structural relationships between the multiple structural units in the model representation data are adjusted to obtain the adjusted model representation data; Based on the adjusted model representation data, the second pose of multiple adjusted structural units in the second space is determined.
8. The method according to claim 7, characterized in that, The plurality of structural units are implemented as a polygonal mesh composed of a plurality of polygons, wherein each polygon includes at least three vertices; The step of determining the second pose of multiple adjusted structural units in the second space based on the adjusted model representation data includes: For the i-th polygon, the second pose of the i-th polygon in the second space is generated based on the coordinate values of the vertices corresponding to the i-th polygon in the second space coordinate system, where i is a positive integer.
9. The method according to any one of claims 1 to 5, characterized in that, The step of rendering a second model image, edited based on the 3D model in the first model image, according to the first pose and the second pose, includes: Based on the first pose and the second pose, determine the third pose of the edited distribution points in the second space; The second model image is generated by rendering based on the third pose corresponding to the multiple distribution points.
10. The method according to claim 9, characterized in that, The first pose includes a first rotation matrix, a first scaling matrix, and a first position matrix of the distribution points in the first space; the second pose includes a second rotation matrix and a second position matrix of the edited structural unit in the second space. The step of determining the third pose of the edited distribution points in the second space based on the first pose and the second pose includes: Based on the first rotation matrix and the second rotation matrix, determine the third rotation matrix of the edited distribution points in the second space; Based on the first scaling matrix, determine the second scaling matrix of the edited distribution points in the second space; Based on the first position matrix and the second position matrix, determine the third position matrix of the edited distribution points in the second space; The third pose of the edited distribution points in the second space is composed of the third rotation matrix, the second scaling matrix, and the third position matrix.
11. The method according to claim 9, characterized in that, The distribution points correspond to color density distributions; The step of rendering and generating the second model image based on the third pose corresponding to the multiple distribution points includes: The positional distribution of the edited distribution points in the second space is determined based on the third pose corresponding to the plurality of distribution points; The edited 3D model is rendered based on the color density distribution corresponding to the distribution points and the position distribution to obtain the second model image.
12. The method according to any one of claims 1 to 5, characterized in that, The step of obtaining the model representation data and model rendering data corresponding to the 3D model in the first model image includes: Obtain first model images of the 3D model from multiple perspectives; The model rendering data is generated based on the color distribution density of pixels in the first model image from the multiple perspectives. Based on the model rendering data, the model structure corresponding to the 3D model is extracted, multiple structural units are determined, and the model representation data is composed of the multiple structural units.
13. The method according to claim 12, characterized in that, The step of extracting the model structure corresponding to the 3D model based on the model rendering data, determining multiple structural units, and composing the model representation data from the multiple structural units includes: Based on the model rendering data, multiple polygons corresponding to the 3D model are extracted, and the polygon mesh composed of the multiple polygons is used as the model representation data. The polygon mesh is used to indicate the geometric structure of the 3D model.
14. The method according to claim 12, characterized in that, The step of extracting the model structure corresponding to the 3D model based on the model rendering data, determining multiple structural units, and composing the model representation data from the multiple structural units includes: Based on the model rendering data, the point cloud data corresponding to the 3D model is extracted, and the point cloud data is used as the model representation data. The point cloud data is used to indicate the geometric coordinates of the points that make up the 3D model.
15. The method according to claim 12, characterized in that, The step of extracting the model structure corresponding to the 3D model based on the model rendering data, determining multiple structural units, and composing the model representation data from the multiple structural units includes: Based on the model rendering data, multiple voxels corresponding to the 3D model are extracted, and the voxel model composed of the multiple voxels is used as the model representation data. The voxel model is used to indicate the geometric distribution of the volume pixels that make up the 3D model.
16. An editing device for a model image, characterized in that, The device includes: The acquisition module is used to acquire model representation data and model rendering data corresponding to the three-dimensional model in the first model image when an editing operation is received for the first model image. The model representation data includes structural units that construct the three-dimensional model, and the model rendering data includes rendering distribution features corresponding to multiple distribution points on the three-dimensional model. The first determining module is used to determine the first pose of the rendering distribution features corresponding to the plurality of distribution points in the first space, wherein the first space is a local coordinate space constructed based on the structural unit. The second determining module is used to determine the second pose of the edited structural unit in the second space based on the editing operation and the model representation data, wherein the second space is a global coordinate space constructed based on the first model image; The rendering module is used to render a second model image, which is an edited version of the 3D model in the first model image, based on the first pose and the second pose.
17. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the model image editing method as described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the method for editing a model image as described in any one of claims 1 to 15.
19. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the method for editing a model image as described in any one of claims 1 to 15.