Three-dimensional model generation method and device, equipment, storage medium and program product

By generating viewpoint images from multiple perspectives and updating color parameters, the problems of image detail loss and poor consistency in 3D model generation are solved, thereby improving the color performance and overall quality of the 3D model.

CN121767535APending Publication Date: 2026-03-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, 3D model generation methods suffer from problems such as loss of image details and poor consistency among multiple views, resulting in low-quality generated 3D models.

Method used

By acquiring model materials, multiple perspective images are generated, and the color parameters of the 3D model are updated based on the shading parameters of each perspective image to ensure consistent color representation under different perspectives.

Benefits of technology

It improves the color detail and quality of 3D models, avoids color distortion, and achieves higher quality 3D model generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767535A_ABST
    Figure CN121767535A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional model generation method and device, equipment, a storage medium and a program product. The method comprises the steps of obtaining model materials; generating view angle images of a plurality of view angles based on the model material; a first three-dimensional model is generated based on the view angle images of the multiple view angles, and the first three-dimensional model comprises a first color parameter; obtaining a coloring parameter of each visual angle image; and updating the first color parameter of the first three-dimensional model based on the coloring parameter of each view angle image to obtain a second three-dimensional model. According to the invention, the quality of the generated three-dimensional model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, storage medium, and program product for generating three-dimensional models. Background Technology

[0002] In the 3D model generation methods of related technologies, the multi-view cross-domain attention mechanism is used to promote information exchange across views and modalities, thereby generating a 3D model from a single-view image. However, compared with the input single-view image, there are still problems such as loss of image details, resulting in a lower quality of the generated 3D model. Summary of the Invention

[0003] This application provides a method, apparatus, device, storage medium, and program product for generating three-dimensional models, which can improve the quality of the generated three-dimensional models.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a method for generating a three-dimensional model, the method comprising:

[0006] Obtain model materials;

[0007] Multiple perspective images are generated based on the model materials;

[0008] A first three-dimensional model is generated based on the view images from the multiple viewpoints, wherein the first three-dimensional model includes a first color parameter;

[0009] Obtain the shading parameters for each of the aforementioned viewpoint images;

[0010] Based on the coloring parameters of each viewpoint image, the first color parameters of the first three-dimensional model are updated to obtain the second three-dimensional model.

[0011] This application provides a three-dimensional model generation apparatus, the apparatus comprising:

[0012] The data acquisition module is used to acquire model materials;

[0013] A multi-view generation module is used to generate multiple perspective images based on the model material;

[0014] A 3D model generation module is used to generate a first 3D model based on the view images from the multiple viewpoints, wherein the first 3D model includes a first color parameter;

[0015] The 3D model generation module is also used to obtain the shading parameters of each viewpoint image;

[0016] The 3D model generation module is further configured to update the first color parameter of the first 3D model based on the coloring parameters of each viewpoint image, thereby obtaining a second 3D model.

[0017] This application provides an electronic device, the electronic device comprising:

[0018] Memory is used to store executable instructions or computer programs.

[0019] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the three-dimensional model generation method provided in the embodiments of this application.

[0020] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the three-dimensional model generation method provided in this application.

[0021] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the three-dimensional model generation method provided in this application.

[0022] The embodiments of this application have the following beneficial effects:

[0023] After generating the first 3D model using model materials and multiple perspective images based on the model materials, the first color parameter of the first 3D model is updated based on the color parameters of each perspective image. This ensures that the color representation of the 3D model is consistent across different perspectives, thereby avoiding color distortion and optimizing the color parameters of the 3D model. This further enhances the color details of the second 3D model, thus improving the quality of the generated 3D model. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the architecture of the 3D model generation system provided in the embodiments of this application;

[0025] Figure 2 This is a schematic diagram of the server structure provided in an embodiment of this application;

[0026] Figure 3A This is a first flowchart illustrating the three-dimensional model generation method provided in this application embodiment;

[0027] Figure 3B This is a schematic diagram of the second process of the three-dimensional model generation method provided in the embodiments of this application;

[0028] Figure 3C This is a schematic diagram of the third process of the three-dimensional model generation method provided in the embodiments of this application;

[0029] Figure 3D This is a schematic diagram of the fourth process of the three-dimensional model generation method provided in the embodiments of this application;

[0030] Figure 3E This is a schematic diagram of the fifth process of the three-dimensional model generation method provided in the embodiments of this application;

[0031] Figure 3F This is a schematic diagram of the sixth process of the three-dimensional model generation method provided in the embodiments of this application;

[0032] Figure 3G This is a schematic diagram of the seventh process of the three-dimensional model generation method provided in the embodiments of this application;

[0033] Figure 3H This is the eighth flowchart of the three-dimensional model generation method provided in the embodiments of this application;

[0034] Figure 3I This is the ninth flowchart of the three-dimensional model generation method provided in the embodiments of this application;

[0035] Figure 3J This is a schematic diagram of the tenth process of the three-dimensional model generation method provided in the embodiments of this application;

[0036] Figure 3K This is a schematic diagram of the eleventh step of the three-dimensional model generation method provided in the embodiments of this application;

[0037] Figure 3L This is a schematic diagram of the twelfth step of the three-dimensional model generation method provided in the embodiments of this application;

[0038] Figure 3M This is a schematic diagram of the thirteenth step of the three-dimensional model generation method provided in the embodiments of this application;

[0039] Figure 3N This is a schematic diagram of the fourteenth step of the three-dimensional model generation method provided in the embodiments of this application;

[0040] Figure 3O This is the fifteenth flowchart of the three-dimensional model generation method provided in the embodiments of this application;

[0041] Figure 3P This is a sixteenth flowchart illustrating the three-dimensional model generation method provided in this application embodiment;

[0042] Figure 4 This is a schematic diagram illustrating an application scenario of the three-dimensional model generation method provided in the embodiments of this application;

[0043] Figure 5 This is a schematic diagram illustrating the principle framework of the three-dimensional model generation method provided in the embodiments of this application;

[0044] Figure 6 This is a flowchart illustrating the method for generating a 3D model of a game scene according to an embodiment of this application;

[0045] Figure 7A This is a first schematic diagram of the three-dimensional model generation interface provided in the embodiments of this application;

[0046] Figure 7B This is a second schematic diagram of the three-dimensional model generation interface provided in the embodiments of this application;

[0047] Figure 8 This is a schematic diagram of the architecture of the 3D model generation platform provided in the embodiments of this application;

[0048] Figure 9 This is a comparative schematic diagram of the rendering effects of a 3D model under different lighting parameters provided in the embodiments of this application;

[0049] Figure 10 This is a schematic diagram of a fixed viewing angle provided in an embodiment of this application;

[0050] Figure 11 This is a schematic diagram illustrating the construction principle of the fine-tuning weight parameters provided in the embodiments of this application;

[0051] Figure 12 This is a schematic diagram comparing the image material generation effects of the original image generation model and the trained image generation model provided in the embodiments of this application;

[0052] Figure 13 This is a schematic diagram of multiple viewpoint sampling provided in the embodiments of this application;

[0053] Figure 14 This is a schematic diagram of the data structure of the second dataset provided in the embodiments of this application;

[0054] Figure 15A This is a first schematic diagram illustrating the effect of generating viewpoint images from multiple perspectives according to embodiments of this application;

[0055] Figure 15B This is a second schematic diagram illustrating the effect of generating viewpoint images from multiple perspectives provided in the embodiments of this application;

[0056] Figure 15C This is a third schematic diagram illustrating the effect of generating viewpoint images from multiple perspectives according to embodiments of this application;

[0057] Figure 15D This is a fourth schematic diagram illustrating the effect of generating viewpoint images from multiple perspectives provided in the embodiments of this application;

[0058] Figure 16AThis is a first schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application;

[0059] Figure 16B This is a second schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application;

[0060] Figure 16C This is a third schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application;

[0061] Figure 16D This is a fourth schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application;

[0062] Figure 17A This is a first comparison diagram of the first three-dimensional model and the second three-dimensional model provided in the embodiments of this application;

[0063] Figure 17B This is a second comparison diagram of the first three-dimensional model and the second three-dimensional model provided in the embodiments of this application;

[0064] Figure 17C This is a third comparative diagram of the first three-dimensional model and the second three-dimensional model provided in the embodiments of this application;

[0065] Figure 17D This is a fourth comparison diagram of the first three-dimensional model and the second three-dimensional model provided in the embodiments of this application;

[0066] Figure 18 This is a schematic diagram illustrating the effect of texture optimization provided in the embodiments of this application;

[0067] Figure 19 This is a schematic diagram illustrating the effect of topology optimization provided in the embodiments of this application;

[0068] Figure 20 This is a schematic diagram illustrating the effect of the surface reduction processing provided in the embodiments of this application;

[0069] Figure 21A This is a schematic diagram illustrating the principle of post-processing provided in the embodiments of this application;

[0070] Figure 21B This is a schematic diagram of the post-processing flow provided in the embodiments of this application.

[0071] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0072] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0073] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0074] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0075] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0076] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0077] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0078] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0079] 1) Model material is a two-dimensional model image used to provide an image as a reference to intuitively indicate the appearance of a three-dimensional model, thereby guiding the style of the generated three-dimensional model. Model material can be, for example, a hand-drawn sketch, a photograph, a design drawing, or a screenshot of another three-dimensional model. Model material can also be generated from text data. Here, text data uses natural language to describe the characteristics of the three-dimensional model to be generated, which can include shape, color, texture, size, proportion, style, etc., thereby performing image material generation processing (text-to-image processing) based on text data to obtain the corresponding model material.

[0080] 2) A 3D model is a mathematical or computer graphics representation of an object in three-dimensional space. 3D models are widely used in computer graphics, animation, virtual reality, game development, industrial design, architecture, medical imaging, and other fields. A 3D model can be composed of vertices, edges, and faces; these basic elements define the shape and structure of the model.

[0081] 3) Camera viewpoint refers to the position and orientation of the camera in three-dimensional space, which determines the content of the image of the three-dimensional model captured by the camera. Camera viewpoint can include parameters such as pitch and azimuth, which together define the camera's orientation and viewing angle.

[0082] The following is an explanation of two main parameters in the camera's field of view:

[0083] Pitch angle: This is the angle at which the camera tilts up or down, determining whether the camera is pointing towards the sky or the ground. The pitch angle is calculated from 0 degrees, where 0 degrees means the camera is horizontal and facing forward. Positive values ​​indicate the camera is tilted upward, and negative values ​​indicate the camera is tilted downward. For example, 0 degrees: camera horizontal and facing forward; 90 degrees: camera pointing directly at the sky; -90 degrees: camera pointing directly at the ground.

[0084] Azimuth: This is the angle by which the camera rotates left and right, determining the direction the camera faces. The azimuth angle is usually between 0 and 360 degrees. 0 degrees and 360 degrees represent the same direction. For example, 0 degrees / 360 degrees: the camera faces north; 90 degrees: the camera faces east; 180 degrees: the camera faces south; 270 degrees: the camera faces west.

[0085] These angles collectively define the camera's field of view, allowing for precise control over the camera's orientation in three-dimensional space, thereby capturing images of the desired three-dimensional model.

[0086] 4) Perspective images from multiple viewpoints refer to a collection of images obtained by observing the same 3D model from different camera perspectives. In the fields of 3D modeling and computer vision, such image collections are often used to reconstruct, analyze, or render 3D models.

[0087] 5) The first color parameter refers to the parameter used to describe and specify the color characteristics of the surface of the first 3D model. These parameters can affect the appearance and texture of the model under different lighting conditions.

[0088] 6) Shading parameters refer to the parameters used to recolor the 3D model. Shading parameters include viewpoint coordinates and shading values ​​(i.e., second color parameters). The viewpoint coordinates refer to the position of the observer or camera, which defines the position and orientation of the observer or camera relative to the 3D model and is used to determine the perspective from which the observer or camera observes the 3D model. The shading values ​​refer to the values ​​of the second color parameters of the pixels in the viewpoint image mapped to the first 3D model through the viewpoint coordinates.

[0089] 7) The first vertex refers to the vertex located at the tip of the pyramid (i.e., not on the base). The first vertex is the highest point of the pyramid, and its position determines the height of the pyramid. For example, if the base of the pyramid is located on the XY plane, with the center of the base of the pyramid as the origin, the coordinates of the first vertex can be represented as (0, 0, h), where h is the height of the pyramid.

[0090] 8) The second vertex is one of the basic units that make up a 3D model. A vertex is a point in 3D space, usually defined by its coordinates. These coordinates can be Cartesian coordinates (x, y, z) or coordinates in other coordinate systems. Each vertex can include corresponding color parameters to construct a 3D model with multiple colors.

[0091] In the related 3D model generation methods, the multi-view cross-domain attention mechanism is used to promote information exchange across views and modalities, thereby generating 3D models from single-view images. However, there are still problems such as poor consistency among multiple views and loss of image details compared to the input single-view image, resulting in lower quality of the generated 3D models.

[0092] In addition, in the related technologies, the method of generating 3D models based on text data is to convert text data into single-view images through text image processing, and then generate 3D models based on single-view images. However, due to the instability of the generation style of the images obtained by text image processing, the appearance of the 3D models generated based on text data has an uncontrollable style and cannot be directly applied to business scenarios.

[0093] This application provides a method, apparatus, device, computer-readable storage medium, and computer program product for generating three-dimensional models, which can improve the quality of the generated three-dimensional models.

[0094] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. The electronic devices provided in the embodiments of this application can be implemented as various types of terminal devices such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and vehicle terminals, or they can be implemented as servers.

[0095] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the 3D model generation system provided in the embodiments of this application. Figure 1 The system involves server 100, terminal device 200, and network 300. Terminal device 200 is connected to server 100 through network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.

[0096] In some embodiments, the embodiments of this application can be implemented collaboratively by a server and a terminal device. For example, the terminal device 200 sends model materials to the server 100, and the server 100 obtains a second three-dimensional model based on the model materials using the three-dimensional model generation method provided in the embodiments of this application, and sends the second three-dimensional model to the terminal device 200.

[0097] In other embodiments, the embodiments of this application can be implemented by a terminal device alone. For example, the terminal device 200 acquires model materials and generates a second three-dimensional model based on the model materials using the three-dimensional model generation method provided in the embodiments of this application.

[0098] In other embodiments, the embodiments of this application can be implemented by the server alone. For example, server 100 obtains model materials and generates a second three-dimensional model based on the model materials using the three-dimensional model generation method provided in the embodiments of this application.

[0099] In some embodiments, server 100 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal devices and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0100] Taking the server used for 3D model generation as an example, see [link to relevant documentation]. Figure 2 , Figure 2 This is a schematic diagram of the server structure provided in an embodiment of this application. Figure 2The server 100 shown includes at least one processor 110, memory 130, and at least one network interface 120. The various components of server 100 are coupled together via a bus system 140. It is understood that the bus system 140 is used to implement communication between these components. In addition to a data bus, the bus system 140 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 140.

[0101] The processor 110 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0102] The memory 130 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 130 may optionally include one or more storage devices physically located away from the processor 110.

[0103] The memory 130 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 130 described in this application embodiment is intended to include any suitable type of memory.

[0104] In some embodiments, memory 130 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0105] Operating system 131 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks.

[0106] The network communication module 132 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 120, exemplary network interfaces 120 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0107] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A 3D model generation device 133 stored in memory 130 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a data acquisition module 1331, a multi-view generation module 1332, and a 3D model generation module 1333. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0108] In some embodiments, the terminal device or server can implement the 3D model generation method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as 3D modeling APPs or game development APPs (used to create 3D game characters, props, and scenes, etc.); or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0109] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the three-dimensional model generation method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0110] The following will describe the 3D model generation method provided in this application embodiment, with the server as the execution subject, using exemplary applications and implementations of the server provided in the embodiments of this application. See also Figure 3A , Figure 3AThis is a first flowchart illustrating the three-dimensional model generation method provided in this application embodiment, which will be combined with... Figure 3A The steps shown are explained.

[0111] In step 101, model materials are obtained.

[0112] In some embodiments, the model material is a two-dimensional model image, see [link to relevant documentation]. Figure 3B , Figure 3A Step 101 shown can be implemented through steps 1011 to 1012, as explained in detail below.

[0113] In step 1011, text data is acquired, wherein the text data is used to describe the three-dimensional model.

[0114] In some embodiments, text data submitted by the terminal device to the 3D model generation platform is obtained. The text data may include the shape, color, texture, size, proportion, style, etc. of the 3D model to be generated. For example, the text data may be expressed as "a blue cylinder with a red ring on top, a height of 10 cm and a diameter of 5 cm". The embodiments of this application do not limit the specific content contained in the text data.

[0115] In step 1012, image material generation processing is performed based on text data to obtain model material.

[0116] In some embodiments, image materials are generated based on text data using a pre-trained image generation model to obtain model materials.

[0117] In some embodiments, see Figure 3C The image generation model is trained through the following steps 201 to 207, which are explained in detail below.

[0118] In step 201, a first dataset for the image material generation task is constructed, wherein the first dataset includes multiple image samples and a text description for each image sample, and the multiple image samples have the same style.

[0119] In some embodiments, see Figure 3D , Figure 3C Step 201 shown can be achieved through steps 2011 to 2014, as explained in detail below.

[0120] In step 2011, multiple 3D model samples are obtained, wherein the multiple 3D model samples have the same style but different shapes.

[0121] In some embodiments, multiple 3D model samples have the same style but different shapes. For example, a classical style 3D model sample can be a "classical style sofa" or "classical style chair". For example, you can access professional 3D model libraries or online resource platforms, such as TurboSquid or CGTrader, and use the search function to find models with the same style but different shapes. For example, you can use keyword filtering, such as specifying a cartoon style, to find 3D model samples with similar styles but different shapes. The embodiments of this application do not limit the specific way of obtaining 3D model samples.

[0122] In step 2012, each three-dimensional model sample is rasterized according to a pre-set fixed viewpoint to obtain an image sample corresponding to each three-dimensional model sample.

[0123] In some embodiments, rendering parameters such as lighting parameters (e.g., setting the light source to ambient light) and fixed viewing angle (e.g., setting the camera viewing angle to 30° azimuth and 20° top angle) can be set in the rendering engine or software to simulate the lighting and shadow effects of the real world, thereby performing rasterization processing to obtain image samples corresponding to the 3D model samples. Here, rasterization processing is part of the rendering process and is used to convert the 3D model samples into 2D image samples. That is, rasterization processing is a process of transforming the geometric data of the 3D model samples and converting them into pixels to be presented on the display device.

[0124] See Figure 9 , Figure 9 This is a comparative schematic diagram of the rendering effect of a 3D model under different lighting parameters provided in the embodiments of this application. Here, since the 3D model itself does not have lighting effects, setting the light source to ambient light can avoid the appearance of specular highlights in the image sample, thereby ensuring the unlit nature of the subsequently generated 3D model.

[0125] For example, rasterization can be implemented as follows: First, set the camera position, viewpoint (i.e., fixed viewpoint), and field of view (FOV). The camera position defines the observer's position, the viewpoint defines the observer's orientation, and the FOV defines the area the camera can observe. Next, transform the 3D model sample to the world coordinate system, converting the camera's position and orientation to the camera coordinate system. Apply a view transformation matrix to transform the 3D model sample from the world coordinate system to the camera coordinate system. Then, apply a projection transformation matrix to further transform the 3D model sample from the camera coordinate system to clip space, for example, through perspective projection or orthographic projection. Next, in clip space, identify and remove 3D model portions that exceed the view's range, as well as occluded portions. Finally, transform the coordinates in clip space to screen space (Normalized Device Space). Coordinates (NDC) is a step that maps coordinates to a range of -1 to 1, corresponding to the bottom left to top right corner of the screen. Next, primitives (usually triangles) in screen space are converted into pixels by a rasterizer. This process includes determining which pixels belong to primitives and how to interpolate vertex attributes (such as color, texture coordinates, normals, etc.) on these pixels. Then, each pixel is shaded, such as lighting calculations, texture mapping, and shadow processing. Finally, the rendering result is output to an image buffer, and the contents of the image buffer are saved as a file as an image sample, such as saving the image sample as an image format like PNG, JPEG, and HDR.

[0126] By fixing the camera's perspective (setting the azimuth angle to 30° and the top angle to 20°), the features of multiple faces of the 3D model sample can be observed, capturing more details. This improves the accuracy of image material generation and processing, thus better guiding the generation of 3D models. For an example, see [link to example]. Figure 10 , Figure 10 This is a schematic diagram of a fixed viewing angle provided in an embodiment of this application. The same viewing angle is used for rasterization processing of each three-dimensional model sample.

[0127] In step 2013, text generation processing is performed on each image sample to obtain the text description corresponding to each image sample.

[0128] In some embodiments, text descriptions of image samples can be generated using a large language model (or graph-to-text model), with the image samples used as input to the large language model to generate the text descriptions.

[0129] Taking the image-to-text model Pix2Text as an example, generating text descriptions for image samples can be achieved in the following ways: First, the image samples are preprocessed, such as by grayscale conversion and normalization. Next, feature extraction is performed on the preprocessed image samples. For example, by using a Convolutional Neural Network (CNN) or other image processing techniques, features such as texture, shape, and color in the image samples are identified, and this information is encoded into high-dimensional vectors to obtain image features. Next, the processed image features are converted into fixed-length vectors, which contain rich information about the image samples and are used for subsequent text generation. Finally, text mapping is performed based on the vectors to obtain the text description. For example, the vectors are mapped using a Long Short-Term Memory (LSTM) network to generate a text sequence. After obtaining the text sequence, it can be further organized and corrected to improve the accuracy of the text description.

[0130] For example, for an image sample containing a game character, the output text description could be: "In this image, you can see a male character wearing a blue shirt. The male character's head uses a small polygonal mesh to reduce rendering complexity, while the body uses a high-precision polygonal mesh to accurately depict the character's muscle structure and physique. The character's clothing is composed of polygonal meshes, with a detailed design that reflects the texture and pattern of the clothing."

[0131] In other embodiments, a text description for each 3D model sample can also be generated through manual annotation.

[0132] In step 2014, each image sample and its corresponding text description are combined into a first dataset.

[0133] In some embodiments, the first dataset may include image data of image samples (such as file paths, binary data, pixel matrices, etc.), text descriptions of image samples (in string form), and other metadata (such as image source information, tags, etc.). This application embodiment does not limit the data structure and data storage method of the first dataset.

[0134] See also Figure 3C In step 202, a pre-trained original image generation model is obtained, and the following processing is performed using the original image generation model.

[0135] In step 203, feature extraction processing is performed on the text description to obtain text features.

[0136] In some embodiments, the text description is segmented to obtain multiple input units; the multiple input units are embedded and encoded to obtain embedded features; and the embedded features are then subjected to attention encoding to obtain text features.

[0137] For example, suppose the input text description is "thank you very much". After word segmentation, multiple input units (tokens) are obtained, such as ["thank", "you", "very", "much"]. Next, embedding encoding is performed on multiple tokens. Assuming that each token is mapped to a 2048-dimensional vector, the text description "thank you verymuch" is converted into 4*2048 embedding features. Then, attention encoding is performed on the 4*2048 embedding features to obtain text features. Here, the input unit (token) refers to the basic processing unit of the text description. A token can be a text, word, subword, character, or other meaningful element, depending on the granularity and requirements of text processing. This application embodiment does not impose any restrictions.

[0138] For example, multiple input units can be embedded and encoded using an embedding layer to obtain embedded features. For example, the embedding layer can use embedding methods such as word embeddings, character embeddings, and subword embeddings to perform embedding and encoding processes to obtain embedded features.

[0139] For example, attention weights of embedded features are calculated using an attention mechanism (such as self-attention). These attention weights represent the importance of the embedded features to the current task. Multiplying the attention weights of the embedded features by the embedded features yields a weighted embedded feature, which is the text feature.

[0140] In other embodiments, text descriptions can also be processed by network models such as recurrent neural networks (RNNs) to obtain text features. This application does not limit the specific implementation of the feature extraction process for text descriptions.

[0141] For example, feature extraction from text descriptions using RNNs can be achieved as follows: Before inputting the text description into the RNN, preprocessing can be performed, including tokenization, part-of-speech tagging, and stop word removal. Next, each word in the text description is converted into a numerical vector, which is then combined to obtain a numerical sequence, for example, through a bag-of-words model or pre-trained word embeddings (such as Word2Vec). Finally, the numerical sequence is input into the RNN, which learns how to capture the sequence dependencies in the text description and generates a fixed-size vector representation, which is the text feature.

[0142] In step 204, image generation processing is performed based on text features to obtain the predicted image.

[0143] In some embodiments, the extracted text features are mapped to the image space to obtain a predicted image. For example, a Generative Adversarial Network (GAN) or other image generation model is used to generate a predicted image based on the text features. The GAN consists of a generator and a discriminator. The generator is responsible for generating the predicted image, and the discriminator is responsible for determining whether the generated predicted image conforms to the distribution of the real data. The embodiments of this application do not limit the image generation model used for image generation processing.

[0144] In step 205, the first loss value of the predicted image and the image sample is obtained by using a pre-set loss function.

[0145] In some embodiments, a first loss value between the predicted image and the image sample is calculated using a pre-defined loss function (e.g., cross-entropy loss).

[0146] In step 206, fine-tuning weight parameters are constructed by adding fine-tuning weight parameters to the network parameters of the original image generation model.

[0147] In some embodiments, see Figure 3E , Figure 3C The construction of fine-tuned weight parameters in step 206 shown can be achieved through the following steps 2061 to 2062, which are explained in detail below.

[0148] In step 2061, a first low-rank matrix initialized based on a Gaussian distribution is generated, and a second low-rank matrix initialized based on all zeros is generated.

[0149] In some embodiments, the dimension of the first low-rank matrix can be determined based on the size of the weight matrix of the fully connected layer of the original image generation model. The dimension of the second low-rank matrix is ​​the same as that of the transpose of the first low-rank matrix. The first low-rank matrix is ​​initialized with random numbers generated using a Gaussian distribution (normal distribution), which is defined by the mean and standard deviation (STD) parameters. All elements of the second low-rank matrix are initialized to zero.

[0150] For example, if the weight matrix of the fully connected layer of the original image generation model is m×n, then the first low-rank matrix can be an m×k matrix, where k is the rank in the low-rank approximation, and the dimension of the second low-rank matrix is ​​the same as the transpose of the first low-rank matrix.

[0151] In step 2062, fine-tuning weight parameters are constructed based on the first low-rank matrix and the second low-rank matrix.

[0152] In some embodiments, fine-tuning weight parameters are constructed based on the product of the first low-rank matrix and the second low-rank matrix, see [link to relevant documentation]. Figure 11 , Figure 11 This is a schematic diagram illustrating the construction principle of the fine-tuning weight parameters provided in the embodiments of this application, as shown below. Figure 11 As shown, the product of the first low-rank matrix (A) and the second low-rank matrix (B) constitutes the fine-tuning weight parameter (ΔW), wherein the dimension of the second low-rank matrix is ​​the same as that of the transpose of the first low-rank matrix.

[0153] See also Figure 3C In step 207, the fine-tuned weight parameters are updated using the first loss value to obtain the trained image generation model.

[0154] In some embodiments, the parameters in the network parameters other than the first low-rank matrix and the second low-rank matrix are kept unchanged, and the first low-rank matrix and the second low-rank matrix are updated using the first loss value.

[0155] For example, the original network parameters (weight parameters) of the original image generation model can be represented as W0∈R d×k The image generation model is fine-tuned using a weight parameter (ΔW) so that the weight parameters after fine-tuning (i.e., after training) are W′=W0+ΔW. Then, ΔW is decomposed into the product of two lower-rank matrices ΔW=B×A, where B∈R. d ×r (corresponding to the second low-rank matrix) and A∈R r×k(corresponding to the first low-rank matrix), and r << min(k,d), during the reverse update process, W0 is frozen (not updated, corresponding to keeping the parameters in the network parameters unchanged except for the first low-rank matrix and the second low-rank matrix), only the parameters of A and B are updated. Thus, the forward propagation of the trained image generation model can be expressed as h = W′x = W0x + ΔWx = W0x + BAx, where x represents the input of the image generation model.

[0156] See Figure 12 , Figure 12 This is a comparative diagram of the model material generation effects of the original image generation model and the trained image generation model provided in this application embodiment. In this diagram, keywords are used to characterize the type of model objects in the model material. Keywords can be obtained by extracting keywords from text descriptions. Here, the style of multiple image samples in the first dataset used for training is cartoon style. It can be seen that the style of the model material obtained by the trained image generation model is more consistent than that of the original image generation model. By using the trained image generation model to generate image materials, the style of the model materials can be controlled, thereby better guiding the generation of 3D models.

[0157] For example, the original image generation model can be a potential diffusion model for improved high-resolution image synthesis (Stable Diffusion XL, SDXL), an image generation model based on GAN networks, etc. This application does not limit the specific original image generation model.

[0158] In other embodiments, image data submitted by the terminal device to the 3D model generation platform is directly used as model material.

[0159] See also Figure 3A In step 102, multiple perspective images are generated based on the model material.

[0160] In some embodiments, the model material is a two-dimensional model image, see [link to relevant documentation]. Figure 3F , Figure 3A Step 102 shown can be implemented through steps 1021 to 1022, which will be explained in detail below.

[0161] In step 1021, the model material is processed by feature extraction using a pre-trained multi-view generation model to obtain model image features.

[0162] In some embodiments, see Figure 3G The multi-view generation model is trained through the following steps 301 to 305, which are explained in detail below.

[0163] In step 301, a second dataset for the multi-view generation task is constructed, wherein the second dataset includes multiple single-image samples and view image samples of multiple perspectives corresponding to each single-image sample.

[0164] In some embodiments, see Figure 3H , Figure 3G Step 301 shown can be implemented through steps 3011 to 3014, as explained in detail below.

[0165] In step 3011, multiple 3D model samples are obtained, and the following processing is performed on each 3D model sample.

[0166] In some embodiments, a professional 3D model library or online resource platform, such as TurboSquid or CGTrader, can be accessed to obtain multiple 3D model samples. This application embodiment does not limit the method of obtaining 3D model samples. Here, multiple 3D model samples can be of the same category, for example, all of them are 3D models of the "furniture" category.

[0167] In step 3012, the three-dimensional model sample is sampled according to the pre-set viewpoint parameters and lighting parameters of multiple viewpoints to obtain viewpoint image samples of the three-dimensional model sample from multiple viewpoints.

[0168] In some implementations, lighting parameters (e.g., setting the lighting parameters (light source) to ambient light) and multiple viewpoints (e.g., setting multiple camera viewpoints to corresponding azimuth angles of 30°, 90°, 150°, 210°, 270°, 330°, and pitch angles of 20°, -10°, 20°, -10°, 20°, -10°, -10°, etc.) can be set in the rendering engine or software to simulate the lighting and shadow effects of the real world. This allows for rasterization processing based on each camera viewpoint (or taking screenshots in the rendering engine or software) to obtain viewpoint image samples of the 3D model sample from multiple perspectives. Here, rasterization is part of the rendering process, used to convert the 3D model sample into a 2D viewpoint image sample. In other words, rasterization is a process of transforming the geometric data of the 3D model sample and converting it into pixels to be displayed on the display device. Specific details of rasterization can be found in step 2012 above, and will not be repeated here.

[0169] For example, see six-view sampling. Figure 13 , Figure 13 This is a schematic diagram of multiple viewpoint sampling provided in the embodiments of this application. Figure 13 The sampled viewpoints of multiple viewpoint image samples from different camera positions are shown.

[0170] In step 3013, the three-dimensional model sample is sampled according to the pre-set single-viewpoint perspective parameters and lighting parameters to obtain a single-image sample of the three-dimensional model sample.

[0171] In some embodiments, referring to the sampling method in step 3012, the three-dimensional model sample is sampled according to the preset single-view perspective parameters and lighting parameters (only the perspective parameters are changed, that is, the single-view perspective parameters are different from the perspective parameters of multiple perspectives) to obtain a single image sample of the three-dimensional model sample.

[0172] In step 3014, the viewpoint image samples and single image samples from multiple perspectives of each 3D model sample are combined into a second dataset.

[0173] In some embodiments, each 3D model sample corresponds to multiple viewpoint image samples and a single image sample.

[0174] For example, see Figure 14 , Figure 14 This is a schematic diagram of the data structure of the second dataset provided in the embodiments of this application. For the 3D model sample named "800015_1_R30", there is a single image sample "000.png" and multiple view (six views) view image samples "001.png-006.png".

[0175] In some embodiments, different initial rotation angles can be set for the same 3D model sample to sample multiple sets of images (including single image samples and view image samples from multiple perspectives) based on a 3D model sample, thereby improving the compatibility of the multi-view generation model with different perspectives.

[0176] In some embodiments, the viewpoint parameters of multiple viewpoints include camera distance parameters, see [link to relevant documentation]. Figure 3I After obtaining multiple perspective image samples of the 3D model sample, the following steps 3015 to 3016 can be performed, which are explained in detail below.

[0177] In step 3015, edge detection is performed on each viewpoint image sample to obtain the edge detection result.

[0178] In some embodiments, edge detection can be implemented as follows: First, image samples for each viewpoint are loaded using an image processing library (such as OpenCV, Pillow, etc.). Next, image preprocessing is performed, such as adjusting the viewpoint image sample size, performing grayscale conversion, and denoising. Then, edge detection is performed using a pre-set edge detection algorithm, such as Canny edge detection (which includes four steps: noise removal, gradient calculation, non-maximum suppression, and hysteresis thresholding), Sobel operator (which detects edges based on the magnitude of the gradient), and Laplacian operator (which detects edges based on the zero-crossing of the second derivative). After obtaining the edge detection results of the image samples, the edge detection results can be refined, such as removing small isolated points or filling in discontinuities in the edges. Finally, the edges are connected to form a closed contour, resulting in the final edge detection result.

[0179] In step 3016, in response to the edge detection result indicating that the edge of the model object in the view image sample exceeds the boundary of the view image sample, the camera distance parameter is increased according to the preset parameter, and the three-dimensional model sample is sampled according to the increased camera distance parameter to obtain view image samples of multiple views of the new three-dimensional model sample.

[0180] In some embodiments, when the edge of the model object in the edge detection result exceeds the boundary of the viewpoint image sample, it means that the current camera distance parameter is too small to completely capture the model object (i.e., the model object is not fully displayed within the camera's field of view). In this case, the camera distance is increased according to a preset parameter. This parameter can be a fixed value or calculated according to a certain rule (such as a certain proportion of the current camera distance). The new camera distance parameter is used to update the camera settings. For example, the camera position and / or focal length are adjusted, and the 3D model sample is sampled according to the updated camera settings to generate new viewpoint image samples. For example, the screenshot is re-executed according to the new camera settings until the model object is inside all image samples, ensuring that the new image samples contain the complete edge of the model object.

[0181] See also Figure 3G In step 302, feature extraction processing is performed on the single image sample to obtain image feature samples.

[0182] In some embodiments, a pre-trained convolutional neural network (e.g., ResNet, Inception, etc.) can be used to extract features from a single image sample to obtain image features. In a convolutional neural network, there are usually multiple fully connected layers or multiple global average pooling layers, which generate a final fixed-size feature vector at the end of the convolutional neural network. This is the image feature, which contains high-level semantic information of the single image sample.

[0183] In step 303, multi-view generation processing is performed based on image feature samples to obtain view prediction images from multiple perspectives.

[0184] In some embodiments, conditional information features are obtained by embedding viewpoint information (e.g., a pre-set camera pose, including the camera's rotation matrix R and translation matrix T, where the camera pose determines the camera's viewpoint) based on image feature samples. Based on the conditional information features and the noisy image, viewpoint prediction images of multiple viewpoints are generated through a stepwise denoising process. For example, attention encoding processing (e.g., Cross Attention) is performed on the conditional information features and the noisy image to obtain attention encoding features. A denoising network (e.g., a U-Net) performs noise prediction based on the attention encoding features. The noisy image is then denoised according to the predicted noise to obtain the viewpoint prediction image corresponding to the viewpoint information.

[0185] For example, R and T can be concatenated with image features as conditional information features. In CrossAttention, the conditional information features can be used as the query, and the noisy image features as the key and value. Attention-encoded features are calculated, and the denoising network performs noise prediction based on these features. For instance, the attention-encoded features can be fused (concatenated) with the noisy image features, and the fused features can be mapped using a mapping function (such as Softmax) to obtain the noise prediction result. The purpose of noise prediction is to determine how to adjust each pixel of the noisy image during the denoising process. Based on the noise prediction result of the denoising network, the noisy image is progressively denoised. The denoising process includes multiple time steps, each of which reduces noise and increases image clarity. At each time step, the denoising network updates the noisy image to make it closer to the final target image (viewpoint prediction image).

[0186] In step 304, a second loss value is obtained for the view prediction images from multiple viewpoints and the view image samples from multiple viewpoints through a pre-set loss function.

[0187] In some embodiments, a second loss value is obtained for the view prediction image of multiple viewpoints and the view image sample of multiple viewpoints through a pre-set loss function (e.g., mean squared error loss function). The embodiments of this application do not limit the specific loss function used.

[0188] In step 305, the network parameters of the original multi-view generation model are updated using the second loss value to obtain the trained multi-view generation model.

[0189] In some embodiments, the gradient information of the original multi-view generation model is obtained through the second loss value to obtain the trained multi-view generation model.

[0190] For example, the gradient information of the second loss value with respect to each parameter of the original multi-view generation model is obtained through the backpropagation algorithm. The parameters of the original multi-view generation model are updated using the obtained gradient information according to the gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.). The above process is repeated until a certain number of iterations are reached or the original multi-view generation model converges, thereby obtaining the trained multi-view generation model.

[0191] By constructing a second dataset for the multi-view generation task, the original multi-view generation model is incrementally trained based on the second dataset. This allows the multi-view generation model to learn the model features of multiple perspectives in the second dataset, thereby improving the generation effect of multi-view images and making the multi-view images more consistent in perspective, so as to better guide the generation of 3D models.

[0192] See Figure 15A , Figure 15B , Figure 15C and Figure 15D , Figure 15A This is a first schematic diagram illustrating the effect of generating viewpoint images from multiple perspectives according to embodiments of this application. Figures 15A to 15D It can be seen that the view consistency of the multi-view generation model trained on the second dataset has been improved.

[0193] See also Figure 3F In step 1022, a multi-view generation process is performed based on the model image features using a pre-trained multi-view generation model to obtain view images from multiple perspectives.

[0194] In some embodiments, multi-view generation processing is performed based on model image features through forward inference of the multi-view generation model (see the description of steps 302 to 303) to obtain view images from multiple perspectives.

[0195] For example, the multi-view generation model can be a generation model that generates multiple views from a single graph based on a diffusion model (such as Stable Diffusion), including but not limited to zero123, zero123++, etc. The embodiments of this application do not limit the specific multi-view generation model.

[0196] See also Figure 3A In step 103, a first three-dimensional model is generated based on view images from multiple perspectives, wherein the first three-dimensional model includes a first color parameter.

[0197] In some embodiments, see Figure 3J , Figure 3A Step 103 shown can be implemented through steps 21031 to 1034, which will be explained in detail below.

[0198] In step 1031, feature extraction processing is performed on the viewpoint image of each viewpoint to obtain the image features of each viewpoint.

[0199] In some embodiments, feature extraction processing can be performed on the viewpoint image of each viewpoint using a pre-trained convolutional neural network (e.g., ResNet, Inception, etc.) to obtain the image features of each viewpoint. In a convolutional neural network, there are usually multiple fully connected layers or multiple global average pooling layers, thereby generating a final fixed-size feature vector at the end of the convolutional neural network. This is the image feature, which contains high-level semantic information of the viewpoint image.

[0200] In step 1032, feature encoding processing is performed on the image features of each viewpoint to obtain the encoded features of each viewpoint.

[0201] In some embodiments, feature encoding processing of image features for each viewpoint can be achieved by: dividing the image features into multiple image feature blocks, arranging the multiple image feature blocks into an image feature block sequence, and encoding the image feature block sequence to obtain encoded features.

[0202] For example, assuming the feature map size of the image features is 224×224, and it is divided into 16×16 patches, the result will be (224 / 16). 2 = 196 patches, and then arrange the 196 patches in the order from the top left to the bottom right in the view image to obtain the image feature block sequence.

[0203] In some embodiments, encoding the image feature block sequence to obtain encoded features can be achieved by: performing embedding encoding on the image feature block sequence to obtain embedded features; normalizing the embedded features to obtain normalized features; performing attention encoding on the normalized features to obtain attention encoded features; fusing the attention encoded features and the embedded features (e.g., concatenating) to obtain fused features; performing feedforward mapping on the normalized features to obtain mapped features; and fusing the mapped features and the fused features to obtain encoded features.

[0204] For example, an embedding layer can be used to embed the image feature block sequence to obtain embedded features. For instance, in the embedding layer, a convolution operation is performed on the image feature block sequence, and position embedding is then performed on the convolutional image feature block sequence to obtain embedded features.

[0205] For example, normalization can be achieved by: obtaining the mean and variance of the embedded features in each feature dimension; then, standardizing the feature data in each feature dimension, for example, by subtracting the mean from the feature data and then dividing by the square root of the variance; finally, further scaling and shifting the standardized embedded features using learnable scaling and shifting factors to obtain normalized features.

[0206] For example, normalized features can be encoded using a multi-head attention mechanism to obtain attention-encoded features. For instance, for the input normalized features, a linear transformation is first performed to generate three matrices: a query vector (Q), a key vector (K), and a value vector (V). This linear transformation is implemented using learnable weight matrices W_Q, W_K, and W_V. Next, for each attention head, the dot product of Q and K is calculated to obtain an attention score. This score is then normalized, for example, by applying a normalization function (such as the softmax function), transforming the attention score into a probability distribution. The attention probability distribution is then used to weight V, generating a new feature representation. Finally, the outputs of all attention heads are concatenated and subjected to a linear transformation to obtain a rich representation that includes the correlations between different positions within the normalized features—that is, the attention-encoded features.

[0207] For example, normalized features can be processed by a feed-forward neural network (FFN) layer to obtain mapped features. Taking a multilayer perceptron (MLP) structure as an example, firstly, the normalized features are linearly transformed by a first linear layer. This first linear layer can be a fully connected layer with a weight matrix W1 and a bias vector b1. Next, the normalized features after linear transformation are nonlinearly transformed by an activation function (such as ReLU, Sigmoid, etc.) to obtain nonlinearly transformed features. Finally, the nonlinearly transformed features are linearly transformed by a second linear layer to obtain mapped features. This second linear layer can be a fully connected layer with a weight matrix W2 and a bias vector b2.

[0208] For example, the mapped features and fused features are concatenated to obtain the encoded features.

[0209] In step 1033, the encoded features of each viewpoint are decoded to obtain three-dimensional features.

[0210] In some embodiments, the feature maps of the encoded features of each viewpoint are stitched together into a three-dimensional structure in a specific way, such as stacking the feature maps of each viewpoint along an axis (e.g., the depth axis). During the stacking process, the feature maps can be transformed, such as rotated or flipped, to ensure that they are aligned in space. Through the above stitching operation, a three-dimensional feature (or triplane feature) is generated.

[0211] For example, Triplane consists of three orthogonal planes, each representing a feature of the 3D model as seen from a different perspective: the horizontal plane represents a top view feature; the vertical plane represents a side view feature; and the diagonal plane represents an oblique view feature.

[0212] In step 1034, a three-dimensional reconstruction process is performed based on the three-dimensional features to obtain the first three-dimensional model.

[0213] In some embodiments, the 3D reconstruction process can be implemented as follows: Sampling points are generated in the feature map on each plane of the 3D feature (Triplane). These sampling points can be keypoints, edge points, or points selected according to a strategy (such as uniform sampling, importance sampling, etc.). Next, the sampling points on different planes are fused, for example, by projecting the sampling points into a common space, and then using a neural network or other methods to merge these features. Next, a cube is constructed using the fused sampling point features. This cube can be a fixed-size voxel grid, a dynamically sized point cloud, or other structure. A single cube is expanded into a set of multiple cubes (FlexiCubes), which can cover the entire Triplane space. Each cube represents a local 3D feature region. Next, the cubes in the FlexiCubes are fused and optimized to improve the continuity and accuracy of the 3D features, such as through spatial filtering, feature smoothing, and denoising. Finally, the optimized FlexiCubes are used for 3D reconstruction, for example, by converting the cubes into voxel grids, point clouds, or other 3D representations, and then visualized or further processed to obtain a first 3D model.

[0214] In other embodiments, a first 3D model can be generated based on viewpoint images from multiple perspectives using a pre-trained 3D model generation model, such as an Instant Mesh generator. The embodiments of this application do not limit the specific 3D model generation model used.

[0215] See also Figure 3A In step 104, the coloring parameters of each viewpoint image are obtained.

[0216] In some embodiments, the viewpoint image includes a model object, see [link to relevant documentation]. Figure 3K , Figure 3A Step 104 shown can be achieved by performing steps 1041 to 1043 for each viewpoint image, as explained in detail below.

[0217] In step 1041, the bounding box of the model object in the view image is obtained.

[0218] In some embodiments, the bounding box of the model object in the view image is obtained, that is, the smallest box surrounding the model object is obtained.

[0219] For example, the viewpoint image is preprocessed, such as by grayscale conversion, filtering, and denoising, to improve the accuracy of subsequent processing. Next, object detection is performed, for example, using object detection algorithms such as Histogram of Oriented Gradients (HOG) and Single Shot MultiBox Detector (SSD) to detect model objects in the image, obtain the position and size of the model objects, obtain the coordinates of all pixels of the model objects through object detection algorithms, find the minimum x and y values ​​(top left corner) and the maximum x and y values ​​(bottom right corner) among all the pixel coordinates of the model objects, and thus define a minimum bounding box that can enclose all the pixels of the model objects.

[0220] In step 1042, the bounding box is used as the bottom surface, the first vertex is determined based on the preset distance parameters, and a quadrangular pyramid is constructed based on the bottom surface and the first vertex.

[0221] In some embodiments, the bounding box is used as the bottom surface, and the first vertex (the vertex at the tip of the pyramid (i.e., not on the bottom surface)) is determined based on a preset distance parameter, and the pyramid is constructed based on the bottom surface and the first vertex.

[0222] For example, see Figure 16A , Figure 16A This is a first schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application, as shown below. Figure 16A As shown, the bounding box is able to enclose all model objects. Figure 16A The smallest rectangular bounding box of the pixels of the sphere shown in the figure is used as the base, and the first vertex is determined based on the preset distance parameters to construct a quadrangular pyramid.

[0223] Here, the preset distance parameter can be determined by the camera pose in the view information set in the multi-view generation model when generating view images from multiple perspectives. That is, the position of the model object is transformed from the world coordinate system to the camera coordinate system, thereby obtaining the distance parameter between the camera (first vertex) and the model object.

[0224] In step 1043, the shading parameters of the viewpoint image are determined by the positional relationship between the face of the pyramid and the first three-dimensional model.

[0225] In some embodiments, the shading parameters include viewpoint coordinates, with each viewpoint image corresponding to a camera viewpoint; see [link to documentation]. Figure 3L , Figure 3K Step 1043 shown can be implemented through steps 10431 to 10432, which will be explained in detail below.

[0226] In step 10431, in the object coordinate system where the first three-dimensional model is located, the face of the quadrangular pyramid is moved along the direction of the camera viewpoint corresponding to the viewpoint image until the face of the quadrangular pyramid is tangent to the boundary of the first three-dimensional model.

[0227] In some embodiments, each camera viewpoint corresponds to a quadrangular pyramid. The face of the quadrangular pyramid is moved along the direction of the camera viewpoint corresponding to the viewpoint image until the normal vector of the face of the quadrangular pyramid is parallel to the normal vector of the boundary of the first three-dimensional model. At this time, the face of the quadrangular pyramid is tangent to the boundary of the first three-dimensional model.

[0228] For example, to obtain the boundary of the first 3D model (e.g., represented by a rectangular bounding box; see the explanation of the bounding box in step 1041 above), move the face of the pyramid and compare the normal vector of the pyramid face with the normal vector of the model boundary. If the angle between them is close to 0 degrees or 180 degrees (i.e., they are almost parallel or antiparallel), then the face of the pyramid can be considered tangent to the boundary of the first 3D model. See [link to documentation]. Figure 16B , Figure 16B This is a second schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application, as shown below. Figure 16B As shown, moving the quadrangular pyramid until its four faces are tangent to the boundary of the first three-dimensional model can also be understood as moving (translating) the quadrangular pyramid so that the quadrangular pyramid just "locks" the first three-dimensional model.

[0229] In step 10432, the intersection point of the faces of the square pyramid tangent to the boundary of the first three-dimensional model is determined, and the coordinates of the intersection point in the object coordinate system are used as the viewpoint coordinates.

[0230] In some embodiments, the intersection points of the faces of the square pyramid tangent to the boundary of the first three-dimensional model are determined, for example, by obtaining the least squares solution. Figure 16BThe intersection points of the four faces shown are used as the viewpoint coordinates in the object coordinate system. For an example, see [example image / description]. Figure 16C , Figure 16C This is a third schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application. Figure 16C The viewpoint coordinates are shown in relation to the position of the first 3D model.

[0231] See also Figure 3A In step 105, the first color parameters of the first three-dimensional model are updated based on the coloring parameters of each viewpoint image to obtain the second three-dimensional model.

[0232] In some embodiments, the first three-dimensional model includes a plurality of second vertices, and the first color parameter of the first three-dimensional model can be set to zero (i.e., the first color parameter of each second vertex of the first three-dimensional model is removed) before updating the first color parameter of the first three-dimensional model.

[0233] In other embodiments, the first three-dimensional model may not include the first color parameter; for example, the first three-dimensional model does not include the first color parameter when it is a point cloud model.

[0234] In some embodiments, the shading parameters further include a second color parameter for each pixel of the model object in the view image. The first 3D model includes a plurality of second vertices, each second vertex corresponding to a first color parameter. See [link to documentation]. Figure 3M , Figure 3A Step 105 shown can be achieved by performing steps 1051 to 1052 on each pixel of the model object in the view image, as explained in detail below.

[0235] In step 1051, a ray is generated in the object coordinate system with the viewpoint coordinate as the endpoint and passing through the pixel point. The second vertex on the first 3D model that intersects with the ray is determined as the shading point.

[0236] In some embodiments, a ray is formed with the viewpoint coordinates as the endpoints and the pixels of the model object in the bounding box, extending to the first three-dimensional model, and the second vertex on the first three-dimensional model that intersects with the ray is used as a shading point.

[0237] In step 1052, the first color parameter of the shading point in the first three-dimensional model is replaced with the second color parameter of the pixel to obtain the second three-dimensional model.

[0238] In some embodiments, the second color parameter of the pixel is mapped to the shading point to obtain the second three-dimensional model.

[0239] For example, see Figure 16D , Figure 16DThis is a fourth schematic diagram illustrating the principle of obtaining the color parameters of each viewpoint image provided in the embodiments of this application, as shown below. Figure 16D As shown, the pixel values ​​of the model object within the bounding box of a viewpoint image are mapped to the first 3D model. Here, for ease of understanding, only the shading process from one viewpoint is shown.

[0240] In some embodiments, see Figure 3N , Figure 3A Step 105 shown can be achieved by performing steps 1053 to 1054 on each second vertex in the first three-dimensional model, as explained in detail below.

[0241] In step 1053, in response to the second vertex being in at least two camera views, the second color parameters of at least two pixels corresponding to the second vertex in the two camera views are weighted and averaged to obtain the third color parameter.

[0242] In some embodiments, when the second vertex of the first three-dimensional model corresponds to the second color parameter of multiple viewpoint image mappings, a weight value is determined based on the angle between the ray formed by the viewpoint coordinates and the center of the bounding box and the normal of the second vertex (the normal is perpendicular to the tangent plane of the surface of the first three-dimensional model), and a weighted average is performed to obtain the third color parameter.

[0243] For example, suppose the second vertex corresponds to the second color parameters (denoted as A and B) of the pixels from the two viewpoint images, and the angles between the ray formed by the viewpoint coordinates of the two viewpoint images and the center of the bounding box and the normal of the second vertex are 60° and 30° respectively. Then the third color parameter (C) can be expressed as C = cos60°A + cos30°B, where cos represents the cosine function.

[0244] In step 1054, the first color parameter of the second vertex in the first three-dimensional model is replaced with the third color parameter to obtain the second three-dimensional model.

[0245] In some embodiments, see Figure 3O , Figure 3A Step 105 shown can be achieved by performing steps 1055 to 1057 on each second vertex in the first three-dimensional model, as explained in detail below.

[0246] In step 1055, in response to the second vertex not being in any camera viewpoint, at least two neighboring vertices adjacent to the second vertex within a preset range are acquired.

[0247] In some embodiments, when the second vertex does not correspond to the second color parameter of any viewpoint image mapping, at least two neighboring vertices adjacent to the second vertex within a preset range are obtained.

[0248] For example, the first three-dimensional model includes the coordinate information of each second vertex. A preset range can be determined by a preset radius value. At least two neighboring vertices within the preset range can be obtained based on the coordinate information of each second vertex, or at least two neighboring vertices that are closest to each other can be selected. The embodiments of this application do not limit the specific method of obtaining neighboring vertices.

[0249] In step 1056, the average value of the second color parameter corresponding to at least two neighboring vertices is obtained.

[0250] Here, the second color parameter corresponding to the neighboring vertex can be obtained through steps 1051 to 1052, or through steps 1053 to 1054. The specific value of the second color parameter corresponding to the neighboring vertex varies depending on the number of neighboring vertices in the camera's view.

[0251] In step 1057, the first color parameter of the second vertex in the first three-dimensional model is replaced with the average value to obtain the second three-dimensional model.

[0252] By updating the first color parameters of the first 3D model through steps 104 and 1051 to 1057, the color parameters of the 3D model are optimized, which further enhances the color details of the second 3D model, thereby improving the quality of the generated 3D model.

[0253] For example, see Figure 17A , Figure 17B , Figure 17C and Figure 17D , Figure 17A This is a first comparison diagram of the first three-dimensional model and the second three-dimensional model provided in the embodiments of this application. Figures 17A to 17D It can be seen that after updating the first color parameter of the first 3D model, the color detail of the second 3D model is enhanced, and the texture detail of the 3D model can be better represented by color (due to...). Figures 17A to 17D This is a grayscale image, where different shades of gray correspond to different colors.

[0254] In some embodiments, see Figure 3P After obtaining the second three-dimensional model, the following steps 106 to 108 can be performed, which are explained in detail below.

[0255] In step 106, the texture of the second 3D model is optimized to obtain the model texture.

[0256] In some embodiments, texture detection processing is performed on the second 3D model to obtain texture detection results; texture optimization is performed based on the texture detection results to obtain model textures.

[0257] For example, in response to a texture detection result indicating that the second 3D model does not include textures (lacks material information), a model texture can be generated for the second 3D model using a 3D model rendering tool (such as Blender). For instance, a 2D texture image can be selected or created using a 3D model rendering tool; this will serve as the surface pattern of the model. The texture image can be edited, such as resizing, color correction, and applying filters, to achieve specific visual effects. Next, the surface of the second 3D model is unfolded into a 2D plane (also known as UV mapping). UV mapping defines how points on the model's surface correspond to points on the texture image. The UV coordinates are adjusted in the UV editor to ensure that the texture distribution on the model is uniform and reasonable. The texture image is then applied to the UV-unwrapped model. Texture refinement techniques, such as seamless textures, texture stitching, and detail mapping, are used to enhance the realism and detail of the texture. Based on the model's lighting model, the lighting and shadow effects of the texture are adjusted to simulate real-world lighting conditions. Material properties, such as diffuse, glossiness, and transparency, are set for the model. Test the texture effect in the rendering view of the 3D model rendering tool, adjust the parameters as needed, and generate a texture suitable for the second 3D model to make it more realistic and visually appealing during rendering.

[0258] For example, in response to the texture detection result indicating that the second 3D model includes textures, but the textures are not conducive to secondary editing, such as when the textures of the second 3D model cannot be directly edited due to the 3D model rendering tool not supporting a specific texture format, or when the texture file is damaged due to encoding problems during saving or transmission, the textures are recalculated and generated. The geometric features, materials, and lighting information of the second 3D model are saved as texture files. As explained above, model textures can be generated for the second 3D model using the 3D model rendering tool.

[0259] See Figure 18 , Figure 18 This is a schematic diagram illustrating the effect of texture optimization provided in the embodiments of this application. Figure 18 As can be seen, the model textures obtained after texture optimization are easier to edit, and these effects can be quickly applied in subsequent rendering or real-time display without recalculation, thus achieving the beneficial effect of improving the rendering efficiency of 3D models.

[0260] In step 107, topology optimization is performed on the second three-dimensional model to obtain the third three-dimensional model.

[0261] In some embodiments, the second three-dimensional model can be topologically optimized through retopology processing to obtain a third three-dimensional model.

[0262] For example, retopology involves rebuilding the surface geometry of a complex or high-resolution 3D model to create a version with fewer polygons while maintaining the visual appearance of the original model. Retopology tools such as QuadRemesh and QuadriFlow can be used to optimize the topology of a second 3D model to obtain a third 3D model. See [link to documentation]. Figure 19 , Figure 19 This is a schematic diagram illustrating the effect of topology optimization provided in the embodiments of this application. Figure 19 As can be seen, the third-dimensional model has a simple model structure, which can reduce the computing power required for model rendering and is more conducive to production applications.

[0263] For example, when topology optimization (retopology) cannot be directly performed due to the excessive number of faces in the second 3D model, face reduction can be achieved through optimization methods such as voxel reconstruction. After reducing the face count of the second 3D model, topology optimization can then be performed. See [example example]. Figure 20 , Figure 20 This is a schematic diagram illustrating the effect of the face reduction processing provided in the embodiments of this application. Figure 20 As can be seen, after the face reduction process, the number of faces in the second 3D model is significantly reduced.

[0264] For example, when topology optimization cannot be performed directly due to structural anomalies in the second 3D model (e.g., topology failure caused by non-manifold geometry, which is a 3D shape that cannot be unfolded into a 2D surface with all normals pointing in the same direction, such as multiple faces connecting to the same vertex), scripts can be used with the help of a graphics editor to edit isolated or redundant vertices and faces in the second 3D model (e.g., using scripts to detect isolated or redundant vertices and faces to delete or merge isolated vertices or faces), while flipping faces to unify the direction of normals (e.g., selecting all faces, using scripts to check the direction of normals, flipping faces with inconsistent normals to ensure that normals uniformly point outwards), thereby transferring the anomaly-corrected second 3D model into topology optimization processing.

[0265] Topology optimization enables the reconstruction of 3D models that lack usable mesh structures, have irregular sizes (such as point cloud models, point-shaded models, etc.), or have excessively high polygon counts. This allows for the standardization and simplification of the model structure, making it more suitable for application in business environments.

[0266] In step 108, the model texture and the third 3D model are baked to obtain the fourth 3D model.

[0267] In some embodiments, a 3D model rendering tool (such as Blender) can be used to bake a fourth 3D model based on the model texture and a third 3D model.

[0268] For example, in Blender, you set the baking type and parameters, such as the resolution, sampling rate, and margins of the baked texture. You also select baking attributes (e.g., color, lighting, shadows, AO, etc.) and set the output path and filename for the baked texture. The baking process begins; Blender renders the information and applies it to the target model (the fourth-dimensional model). It then checks the baked textures and the fourth-dimensional model to ensure all details are correctly transferred. Finally, it exports the baked target model to the desired format, such as FBX, OBJ, or GLTF.

[0269] The three-dimensional model generation method provided in this application embodiment can be applied to various scenarios that require three-dimensional model generation, some of which include: (1) game development: for example, game designers use three-dimensional model generation to create game characters, environments and props; (2) film and television production: for example, in film and animation production, three-dimensional model generation is used to create special effects and animation scenes, thereby bringing visual shock to the audience; (3) industrial design: for example, engineers use three-dimensional model generation to design product prototypes, thereby improving the appearance of the product.

[0270] See Figure 4 , Figure 4 This is a schematic diagram illustrating an application scenario of the 3D model generation method provided in the embodiments of this application, such as... Figure 4 As shown, in a game scene, a game environment and game items (such as...) can be created using the 3D model generation method provided in this application embodiment. Figure 4 The following explanation uses the generation of 3D models in a game scene as an example, including the "furniture" shown in the image.

[0271] See Figure 5 , Figure 5 This is a schematic diagram illustrating the principle framework of the 3D model generation method provided in this application embodiment. The method for obtaining model materials can be to directly use the input image as the model material (corresponding to...). Figure 5 Option 1 ("Input Image") refers to the input text data. The image generation module processes the input text data to generate corresponding images, which are then used as model materials. Figure 5 Option 2: Input text data) is used to generate multiple perspective images based on model materials. The 3D model generation module is used to obtain a second 3D model based on the multi-view images. The post-processing module is used to optimize the texture and topology of the second 3D model to obtain a fourth 3D model.

[0272] For example, see Figure 8 , Figure 8This is a schematic diagram of the architecture of the 3D model generation platform provided in this application embodiment, including a first server, a second server and multiple third servers.

[0273] The first server (or front-end page server) is the front-end part of the user interaction. For example, Nginx can be used as the server and Vue3 as the front-end framework to build and render the 3D model generation interface. The first server is used to receive user requests and forward them to the second server for processing.

[0274] The second server (or routing server) is the system middleware of the 3D model generation platform, comprising three modules: model material storage, 3D model storage, and remesh. The second server handles requests from the first server. Specifically, image storage stores user-uploaded images, object storage stores generated 3D model data, and remesh performs post-processing such as retopology on the 3D models. In detail, the second server receives requests from the first server (e.g., the first server forwards requests via an IP proxy), sends parameters including keywords and image links to the third server, and receives the returned 3D model download link (e.g., a Uniform Resource Locator, URL).

[0275] The third server (or model server cluster) is used to handle specific 3D model generation tasks. It receives upper-layer parameter information, processes it, and returns a 3D model download URL. Each third server can deploy a model corresponding to a different 3D model generation algorithm. Specifically, users send requests through the first server, which are then passed to the second server via an IP proxy. The second server passes the parameters and other information in the request to the specific backend third servers (third server 1, third server 2, third server n). Each backend third server returns a 3D model download URL, which is then passed back to the second server for further processing or storage.

[0276] The three-tier architecture effectively separates the user interface, business logic, and data processing, ensuring the system's scalability and maintainability.

[0277] See Figure 6 , Figure 6 This is a flowchart illustrating the method for generating 3D models of game scenes according to embodiments of this application. The following is a summary of the process. Figure 6 Please provide an explanation.

[0278] In step 401, game model materials are obtained.

[0279] In some embodiments, the first server responds to a trigger operation on the 3D model generation method selection control of the 3D model generation interface to obtain game model materials (corresponding to the model materials mentioned above).

[0280] For example, see Figure 7A , Figure 7A This is a first schematic diagram of the three-dimensional model generation interface provided in the embodiments of this application. In response to the trigger operation of the three-dimensional model generation method selection control-002 (corresponding to text to 3D), text data is obtained. The text data is used to describe the three-dimensional model. Based on the text data, image material generation processing is performed to obtain game model materials (see the description of step 1012 above).

[0281] For example, in response to the trigger operation of the 3D model generation method selection control-001 (corresponding to image to 3D), the image data is directly obtained as game model material.

[0282] For example, see Figure 7B , Figure 7B This is a second schematic diagram of the 3D model generation interface provided in this application embodiment. The 3D model generation interface includes five modules: image-based 3D model generation, text-based 3D model generation, personal history record, queue management, and retopology. When triggered... Figure 7A When the 3D model generation method selection control-001 (corresponding to image to 3D) is activated, the image-based 3D model generation module is called, allowing the user to upload or select image data. The user can specify the 3D model generation algorithm (e.g., InstantMesh) through the "Generation Algorithm Selection." Figure 7A When selecting the 3D model generation method control -002 (corresponding to text-to-3D), the 3D model generation module based on text data is invoked for the user to input text data. The personal history module is used to display information about the generated 3D model, the queue management module is used to display the model generation progress, and the retopology module is used to further optimize the generated 3D model to reduce the number of faces in the 3D model or adjust the lighting intensity of the rendered 3D model.

[0283] For example, the first server receives a user request and forwards it to the second server for processing. The request carries game model assets. The second server stores the game model assets in the request and sends parameter information including keywords (such as text-based 3D, image-based 3D, etc.) and image links (game model asset links) to the third server. The third server generates the corresponding 3D model based on the parameter information and returns the 3D model download URL back to the second server for further processing or storage.

[0284] In step 402, perspective images from multiple viewpoints are generated based on the game model assets.

[0285] In some embodiments, the game model material is a two-dimensional game model image. The third server performs feature extraction processing on the game model material through a pre-trained multi-view generation model to obtain game model image features (corresponding to the model image features mentioned above). Through the pre-trained multi-view generation model, multi-view generation processing is performed based on the game model image features to obtain view images from multiple perspectives. Here, the implementation method of multi-view generation can be found in the description of steps 1021 to 1022 above, and will not be repeated here.

[0286] For example, the training of the multi-view generation model can be achieved as follows: Construct a second dataset for the multi-view generation task, wherein the second dataset includes multiple single-image samples and view image samples from multiple perspectives corresponding to each single-image sample (see step 301 above); perform feature extraction processing on the single-image samples to obtain image feature samples (see step 302 above); perform multi-view generation processing based on the image feature samples to obtain view prediction images from multiple perspectives (see step 303 above); obtain the second loss value of the view prediction images from multiple perspectives and the view image samples from multiple perspectives through a pre-set loss function (see step 304 above); update the network parameters of the original multi-view generation model through the second loss value to obtain the trained multi-view generation model (see step 305 above).

[0287] In step 403, a first three-dimensional model is generated based on the view images from multiple viewpoints, wherein the first three-dimensional model includes a first color parameter.

[0288] In some embodiments, the third server performs feature extraction processing on the viewpoint image of each viewpoint to obtain image features of each viewpoint (see the description of step 1031 above); performs feature encoding processing on the image features of each viewpoint to obtain encoded features of each viewpoint (see the description of step 1032 above); performs decoding processing based on the encoded features of each viewpoint to obtain three-dimensional features (see the description of step 1033 above); and performs three-dimensional reconstruction processing based on the three-dimensional features to obtain a first three-dimensional model (see the description of step 1034 above).

[0289] In other embodiments, the third server can generate a first three-dimensional model based on view images from multiple perspectives using a pre-trained three-dimensional model generation model, such as an instant mesh generator. This application does not limit the specific three-dimensional model generation model used. The first three-dimensional model can be a point cloud model, a point shading model, a mesh model, etc. This application does not limit the three-dimensional representation of the first three-dimensional model.

[0290] In step 404, the shading parameters of each viewpoint image are obtained.

[0291] In some embodiments, the third server obtains the bounding box of the model object in the view image (see the description of step 1041 above); using the bounding box as the bottom surface, the first vertex is determined based on a preset distance parameter, and a quadrangular pyramid is constructed based on the bottom surface and the first vertex (see the description of step 1042 above); the shading parameters of the view image are determined by the positional relationship between the face of the quadrangular pyramid and the first three-dimensional model (see the description of step 1043 above).

[0292] In step 405, the first color parameters of the first 3D model are updated based on the coloring parameters of each viewpoint image to obtain the second 3D model.

[0293] In some embodiments, the shading parameters further include a second color parameter for each pixel of the model object in the view image. The first 3D model includes a plurality of second vertices, each second vertex corresponding to a first color parameter. The third server generates a ray in the object coordinate system with the viewpoint coordinate as the endpoint and passing through the pixel, and determines the second vertex on the first 3D model that intersects the ray as a shading point (see the description of step 1051 above); the first color parameter of the shading point in the first 3D model is replaced with the second color parameter of the pixel to obtain the second 3D model (see the description of step 1052 above).

[0294] In step 406, the second three-dimensional model is post-processed to obtain the fourth three-dimensional model.

[0295] In some embodiments, the second server optimizes the texture of the second 3D model to obtain a model texture (see the description of step 106 above); performs topology optimization on the second 3D model to obtain a third 3D model (see the description of step 107 above); and bakes the model texture and the third 3D model to obtain a fourth 3D model (see the description of step 108 above).

[0296] For example, see Figure 21A , Figure 21A This is a schematic diagram of the post-processing principle provided in the embodiments of this application. Topology optimization includes reducing the number of faces in the second three-dimensional model (e.g., using a face reduction method based on voxel reconstruction), and then performing retopology processing on the second three-dimensional model after face reduction to obtain a third three-dimensional model. Texture optimization may include UV unwrapping the texture image to obtain a model texture, and then baking the model texture onto the third three-dimensional model to obtain a fourth three-dimensional model that can be used in the business environment.

[0297] For example, see Figure 21B , Figure 21B This is a schematic diagram of the post-processing flow provided in the embodiments of this application. The following is in conjunction with... Figure 21B The steps are explained below.

[0298] In step 501, the second three-dimensional model is imported.

[0299] In some embodiments, a second 3D model is imported into a 3D model editing tool (such as Blender).

[0300] In step 502, it is determined whether the second three-dimensional model has material. If the second three-dimensional model has material, the process proceeds to steps 506 to 508. If the second three-dimensional model does not have material, the process proceeds to steps 503 to 505.

[0301] In some embodiments, the material information of the second 3D model can be obtained through Blender's material editor. It is also possible to determine whether the second 3D model has materials by its category. For example, when the second 3D model is a point cloud model or a point shading model, the second 3D model does not have materials.

[0302] Here, material is a property parameter used to define how the surface of a 3D model interacts with light. The material determines the appearance of the 3D model during rendering, including characteristics such as color, gloss, and transparency.

[0303] In step 503, add materials.

[0304] In some embodiments, new materials can be added to a second 3D model using Blender's material editor.

[0305] In step 504, the second three-dimensional model is copied.

[0306] In some embodiments, a second 3D model with added materials is copied as a replica model.

[0307] In step 505, color attributes are added to the original second 3D model.

[0308] In some embodiments, Blender adds a color attribute to each vertex of the original model (the original second 3D model without materials), so that the vertex color will be used during rendering. This may override the material color unless the material uses a specific node to blend these colors. Specifically, if the vertex color is incompatible with the material color (e.g., the vertex color completely overrides the material color), the rendering will only display the vertex color. In complex material and lighting settings, the vertex color can create complex interactions with other material properties (such as texture, transparency, reflectivity, etc.).

[0309] In step 506, the second three-dimensional model is copied.

[0310] In some embodiments, a second three-dimensional model is copied as a replica model.

[0311] In step 507, the textures of the original second 3D model are removed.

[0312] In some embodiments, material information on the original model can be removed using Blender's material editor.

[0313] In step 508, duplicate nodes are removed from the copied model.

[0314] In some embodiments, redundant nodes in the replicated model are filtered out, for example, by viewing the node links of the material through the "Node View" of Blender's material editor. For each unnecessary node, select "Delete" or "Remove Node". For example, if a node is connected to another node, but its output is not connected anywhere, then it may be a redundant node. Before deleting a node, it is necessary to understand the function of the node to avoid accidentally deleting important nodes.

[0315] In step 509, the number of faces that need to be optimized is calculated.

[0316] In some embodiments, you can use Blender's Properties panel to find the model's Geometry property, which displays the model's face count. Record the current face count and determine whether to reduce or increase the face count based on the model's size and detail requirements.

[0317] In step 510, topology optimization is performed.

[0318] In some embodiments, the replicated model is retopologically redesigned to optimize the model's geometry.

[0319] In step 511, UV mapping is performed.

[0320] In some embodiments, UV mapping is performed using a UV unwrapping algorithm (such as automatic or manual unwrapping). UV mapping refers to creating a two-dimensional coordinate map that converts vertex coordinates on the copied model into coordinates on a two-dimensional plane. This process is crucial because it allows texture images (such as texture maps, material maps, etc.) to be correctly applied to the 3D model.

[0321] In step 512, the baking texture is specified.

[0322] In some embodiments, a model texture map is generated for baking, specifying the texture type and attributes to be baked, such as color, bump, or lighting.

[0323] In step 513, the model is superimposed.

[0324] In some embodiments, a copy model that has undergone topology optimization and UV mapping is overlaid onto the original second 3D model.

[0325] In step 514, the extrusion parameters are calculated.

[0326] In some embodiments, the extrusion parameters of the replica model are calculated to ensure that the replica model does not overlap during extrusion, thereby avoiding abnormalities in the baked texture.

[0327] In step 515, the baking parameters are set.

[0328] In some embodiments, parameters required for baking, such as resolution and lighting conditions, are set to ensure that the baking effect meets expectations. The baking process reduces the computational burden of real-time rendering by pre-compiling lighting and texture information, meaning that the setting of baking parameters directly affects the quality and effect of the final baked texture. By adjusting the baking parameters, the level of detail in the baked texture and the quality of the rendering effect can be controlled. For example, increasing the resolution can increase the detail of the texture, but it will also increase the baking time and the required data storage. By setting specific baking parameters, specific visual effects can be achieved, such as soft shadows, strong highlights, or specific styles of lighting effects.

[0329] For example, resolution determines the level of detail in the baked texture; higher resolution provides more detailed textures but increases baking time and file size. Lighting conditions include light source type, intensity, and direction, which affect lighting and shadow effects in the baked texture. Sampling rate determines the accuracy of calculations during baking; a high sampling rate reduces noise but also increases baking time. Baking mode selects the type of texture to bake, such as diffuse mapping, normal mapping, or ambient occlusion mapping. In short, setting baking parameters is about finding a balance between model quality, visual appeal, and performance to ensure the baked result meets the expected artistic effect and practical application requirements.

[0330] In step 516, baking takes place.

[0331] In some embodiments, the generated model texture is baked onto the topology-optimized copy model.

[0332] In step 517, the fourth three-dimensional model is exported.

[0333] In some embodiments, the retopologically and baked fourth 3D model is exported in the desired format (such as FBX, OBJ, or GLTF) for subsequent use and application.

[0334] Through steps 401 to 406 and steps 501 to 517, a first 3D model based on multiple viewpoint images is generated. By obtaining the shading parameters of each viewpoint image, the first color parameters of the first 3D model are updated, thus optimizing the color parameters of the 3D model and further enhancing the color details of the second 3D model, thereby improving the quality of the generated 3D model. Post-processing the second 3D model allows for easier secondary editing of the textured maps, enabling quick application of these effects in subsequent rendering or real-time display without recalculation, thus improving the rendering efficiency of the 3D model. Topology optimization enables the reconstruction of the model structure for 3D models lacking usable mesh structures or with irregular sizes (such as point cloud models, point-shaded models, etc.) or with excessively high polygon counts, achieving standardization and simplification of the model structure for better application in business environments.

[0335] The following description continues to illustrate the exemplary structure of the three-dimensional model generation device 133 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the 3D model generation device 133 in the memory 130 may include:

[0336] Data acquisition module 1331 is used to acquire model materials.

[0337] The multi-view generation module 1332 is used to generate multiple perspective images based on the model material.

[0338] The 3D model generation module 1333 is used to generate a first 3D model based on the view images of the multiple viewpoints, wherein the first 3D model includes a first color parameter.

[0339] In some embodiments, the 3D model generation module 1333 is further configured to obtain the shading parameters for each of the viewpoint images.

[0340] In some embodiments, the three-dimensional model generation module 1333 is further configured to update the first color parameter of the first three-dimensional model based on the coloring parameters of each viewpoint image to obtain a second three-dimensional model.

[0341] In some embodiments, the 3D model generation module 1333 is further configured to perform the following processing on each of the view images: obtain the bounding box of the model object in the view image; use the bounding box as the base, determine the first vertex based on a preset distance parameter, construct a quadrangular pyramid based on the base and the first vertex; and determine the shading parameters of the view image through the positional relationship between the face of the quadrangular pyramid and the first 3D model.

[0342] In some embodiments, the shading parameters include viewpoint coordinates, each viewpoint image corresponds to a camera viewpoint, and the 3D model generation module 1333 is further configured to move the face of the quadrangular pyramid along the direction of the camera viewpoint corresponding to the viewpoint image in the object coordinate system where the first 3D model is located, until the face of the quadrangular pyramid is tangent to the boundary of the first 3D model; determine the intersection point of the face of the quadrangular pyramid that is tangent to the boundary of the first 3D model, and use the coordinates of the intersection point in the object coordinate system as the viewpoint coordinates.

[0343] In some embodiments, the shading parameters further include a second color parameter for each pixel of the model object in the viewpoint image. The first three-dimensional model includes a plurality of second vertices. The three-dimensional model generation module 1333 is further configured to perform the following processing on each pixel of the model object in the viewpoint image: in the object coordinate system, generate a ray with the viewpoint coordinate as the endpoint and passing through the pixel; determine the second vertex on the first three-dimensional model that intersects with the ray as a shading point; replace the first color parameter of the shading point in the first three-dimensional model with the second color parameter of the pixel to obtain a second three-dimensional model.

[0344] In some embodiments, the 3D model generation module 1333 is further configured to perform the following processing on each second vertex in the first 3D model: in response to the second vertex being in at least two camera views, perform a weighted average of the second color parameters of at least two pixels corresponding to the second vertex in the two camera views to obtain a third color parameter; replace the first color parameter of the second vertex in the first 3D model with the third color parameter to obtain a second 3D model.

[0345] In some embodiments, the 3D model generation module 1333 is further configured to perform the following processing on each second vertex in the first 3D model: in response to the second vertex not being in any of the camera viewpoints, obtain at least two neighboring vertices adjacent to the second vertex within a preset range; obtain the average value of the second color parameter corresponding to the at least two neighboring vertices; replace the first color parameter of the second vertex in the first 3D model with the average value to obtain a second 3D model.

[0346] In some embodiments, the data acquisition module 1331 is further configured to acquire text data, wherein the text data is used to describe the three-dimensional model; and to perform image material generation processing based on the text data to obtain the model material.

[0347] In some embodiments, the image material generation process is implemented through a pre-trained image generation model. The data acquisition module 1331 is further configured to construct a first dataset for the image material generation task, wherein the first dataset includes multiple image samples and text descriptions for each image sample, and the multiple image samples have the same style; acquire a pre-trained original image generation model, and perform the following processing through the original image generation model: perform feature extraction processing on the text description to obtain text features; perform image generation processing based on the text features to obtain a predicted image; obtain a first loss value for the predicted image and the image sample through a pre-set loss function; construct fine-tuning weight parameters and add the fine-tuning weight parameters to the network parameters of the original image generation model; update the fine-tuning weight parameters through the first loss value to obtain a trained image generation model.

[0348] In some embodiments, the data acquisition module 1331 is further configured to generate a first low-rank matrix initialized based on a Gaussian distribution, and generate a second low-rank matrix initialized based on all zeros; and construct fine-tuning weight parameters based on the first low-rank matrix and the second low-rank matrix.

[0349] In some embodiments, the data acquisition module 1331 is further configured to keep the parameters in the network parameters other than the first low-rank matrix and the second low-rank matrix unchanged, and update the first low-rank matrix and the second low-rank matrix using the first loss value.

[0350] In some embodiments, the data acquisition module 1331 is further configured to acquire multiple three-dimensional model samples, wherein the multiple three-dimensional model samples have the same style but different shapes; perform rasterization processing on each three-dimensional model sample according to a preset fixed viewpoint to obtain an image sample corresponding to each three-dimensional model sample; perform text generation processing on each image sample to obtain a text description corresponding to each image sample; and combine each image sample and the corresponding text description into the first dataset.

[0351] In some embodiments, the multi-view generation module 1332 is further configured to perform feature extraction processing on the model material using a pre-trained multi-view generation model to obtain model image features; and to perform multi-view generation processing based on the model image features using the pre-trained multi-view generation model to obtain view images of the multiple perspectives.

[0352] In some embodiments, the multi-view generation module 1332 is further configured to construct a second dataset for the multi-view generation task, wherein the second dataset includes multiple single-image samples and multiple view image samples corresponding to each single-image sample; and to perform the following processing through a pre-trained original multi-view generation model: performing feature extraction processing on the single-image samples to obtain image feature samples; performing multi-view generation processing based on the image feature samples to obtain view prediction images for multiple viewpoints; obtaining a second loss value for the view prediction images for multiple viewpoints and the view image samples for multiple viewpoints through a pre-set loss function; and updating the network parameters of the original multi-view generation model through the second loss value to obtain a trained multi-view generation model.

[0353] In some embodiments, the multi-view generation module 1332 is further configured to acquire multiple three-dimensional model samples and perform the following processing on each three-dimensional model sample: sampling the three-dimensional model sample according to pre-set view parameters and lighting parameters of multiple viewpoints to obtain view image samples of the three-dimensional model sample from multiple viewpoints; sampling the three-dimensional model sample according to the pre-set view parameters and lighting parameters of a single viewpoint to obtain a single image sample of the three-dimensional model sample; and combining the view image samples of multiple viewpoints and the single image samples of each three-dimensional model sample into the second dataset.

[0354] In some embodiments, the viewpoint parameters of the multiple viewpoints include camera distance parameters. The multi-view generation module 1332 is further configured to perform edge detection on each viewpoint image sample to obtain an edge detection result; in response to the edge detection result indicating that the edge of the model object in the viewpoint image sample exceeds the boundary of the viewpoint image sample, the camera distance parameter is increased by a preset parameter, and the three-dimensional model sample is sampled according to the increased camera distance parameter to obtain a new viewpoint image sample of the three-dimensional model sample from multiple viewpoints.

[0355] In some embodiments, the three-dimensional model generation module 1333 is further configured to perform feature extraction processing on the viewpoint image of each viewpoint to obtain image features of each viewpoint; perform feature encoding processing on the image features of each viewpoint to obtain encoded features of each viewpoint; perform decoding processing based on the encoded features of each viewpoint to obtain three-dimensional features; and perform three-dimensional reconstruction processing based on the three-dimensional features to obtain the first three-dimensional model.

[0356] In some embodiments, the 3D model generation module 1333 is further configured to perform texture optimization on the second 3D model to obtain a model texture; perform topology optimization on the second 3D model to obtain a third 3D model; and bake based on the model texture and the third 3D model to obtain a fourth 3D model.

[0357] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the three-dimensional model generation method described above in this application.

[0358] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will cause the processor to perform the 3D model generation provided in this application embodiment. For example, ... Figure 3A The method for generating a 3D model is shown.

[0359] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0360] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0361] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0362] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0363] In summary, through the embodiments of this application, after generating a first three-dimensional model using model materials and viewpoint images generated from multiple perspectives based on the model materials, the first color parameters of the first three-dimensional model are updated based on the color parameters of each viewpoint image, so that the color performance of the three-dimensional model under different perspectives is consistent, thereby avoiding color distortion and optimizing the color parameters of the three-dimensional model. This further enhances the color details of the second three-dimensional model, thereby achieving the beneficial effect of improving the quality of the generated three-dimensional model.

[0364] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for generating a three-dimensional model, characterized in that, The method includes: Obtain model materials; Multiple perspective images are generated based on the model materials; A first three-dimensional model is generated based on the view images from the multiple viewpoints, wherein the first three-dimensional model includes a first color parameter; Obtain the shading parameters for each of the aforementioned viewpoint images; Based on the coloring parameters of each viewpoint image, the first color parameters of the first three-dimensional model are updated to obtain the second three-dimensional model.

2. The method according to claim 1, characterized in that, The step of obtaining the shading parameters for each viewpoint image includes: Perform the following processing on each of the aforementioned viewpoint images: Obtain the bounding box of the model object in the view image; Using the bounding box as the base, the first vertex is determined based on a preset distance parameter, and a quadrangular pyramid is constructed based on the base and the first vertex; The shading parameters of the viewpoint image are determined by the positional relationship between the face of the quadrangular pyramid and the first three-dimensional model.

3. The method according to claim 2, characterized in that, The shading parameters include viewpoint coordinates, and each viewpoint image corresponds to a camera viewpoint. Determining the shading parameters of the viewpoint image based on the positional relationship between the faces of the pyramid and the first 3D model includes: In the object coordinate system where the first three-dimensional model is located, the face of the quadrangular pyramid is moved along the direction of the camera viewpoint corresponding to the viewpoint image until the face of the quadrangular pyramid is tangent to the boundary of the first three-dimensional model; Determine the intersection point of the faces of the quadrangular pyramid that are tangent to the boundary of the first three-dimensional model, and use the coordinates of the intersection point in the object coordinate system as the viewpoint coordinates.

4. The method according to claim 3, characterized in that, The shading parameters also include a second color parameter for each pixel of the model object in the view image. The first 3D model includes multiple second vertices. Updating the first color parameter of the first 3D model based on the shading parameters of each view image to obtain a second 3D model includes: For each pixel of the model object in the viewpoint image, the following processing is performed: In the object coordinate system, a ray is generated with the viewpoint coordinates as the endpoint and passing through the pixel point. The second vertex on the first three-dimensional model that intersects with the ray is determined as the shading point. The first color parameter of the shading point in the first three-dimensional model is replaced with the second color parameter of the pixel to obtain the second three-dimensional model.

5. The method according to claim 4, characterized in that, The step of updating the first color parameters of the first 3D model based on the color parameters of each viewpoint image to obtain the second 3D model includes: Perform the following processing on each of the second vertices in the first 3D model: In response to the second vertex being in at least two camera views, the second color parameters of at least two pixels corresponding to the second vertex in the two camera views are weighted and averaged to obtain the third color parameter; The first color parameter of the second vertex in the first three-dimensional model is replaced with the third color parameter to obtain the second three-dimensional model.

6. The method according to claim 4 or 5, characterized in that, The step of updating the first color parameters of the first 3D model based on the color parameters of each viewpoint image to obtain the second 3D model includes: Perform the following processing on each of the second vertices in the first 3D model: In response to the second vertex not being in any of the camera views, acquire at least two neighboring vertices adjacent to the second vertex within a preset range; Obtain the average value of the second color parameter corresponding to the at least two neighboring vertices; The first color parameter of the second vertex in the first three-dimensional model is replaced with the average value to obtain the second three-dimensional model.

7. The method according to any one of claims 1 to 5, characterized in that, The acquisition of model materials includes: Acquire text data, wherein the text data is used to describe the three-dimensional model; The image material is generated based on the text data to obtain the model material.

8. The method according to claim 7, characterized in that, The image material generation process is implemented through a pre-trained image generation model, which is trained in the following way: A first dataset for an image material generation task is constructed, wherein the first dataset includes multiple image samples and a text description of each image sample, and the multiple image samples have the same style; Obtain a pre-trained original image generation model, and perform the following processing using the original image generation model: The text description is subjected to feature extraction processing to obtain text features; Based on the text features, image generation processing is performed to obtain a predicted image; The first loss value of the predicted image and the image sample is obtained by using a pre-set loss function; Construct fine-tuning weight parameters by adding the fine-tuning weight parameters to the network parameters of the original image generation model; The fine-tuned weight parameters are updated using the first loss value to obtain the trained image generation model.

9. The method according to claim 8, characterized in that, The construction of fine-tuned weight parameters includes: Generate a first low-rank matrix initialized based on a Gaussian distribution, and generate a second low-rank matrix initialized based on all zeros; Fine-tuning weight parameters are constructed based on the first low-rank matrix and the second low-rank matrix; The step of updating the fine-tuned weight parameters using the first loss value includes: Keeping all network parameters except the first low-rank matrix and the second low-rank matrix unchanged, the first low-rank matrix and the second low-rank matrix are updated using the first loss value.

10. The method according to claim 8, characterized in that, The first dataset for constructing the image material generation task includes: Obtain multiple 3D model samples, wherein the multiple 3D model samples have the same style but different shapes; Each of the three-dimensional model samples is rasterized according to a pre-set fixed viewpoint to obtain an image sample corresponding to each of the three-dimensional model samples; Perform text generation processing on each of the image samples to obtain a text description corresponding to each of the image samples; Each image sample and its corresponding text description are combined to form the first dataset.

11. The method according to any one of claims 1 to 5, characterized in that, The model material is a two-dimensional model image, and the process of obtaining multiple viewpoint images based on the model material includes: The model material is processed by feature extraction using a pre-trained multi-view generation model to obtain model image features; The multi-view generation model, pre-trained, generates multiple view images based on the model's image features.

12. The method according to claim 11, characterized in that, The multi-view generation model was trained in the following way: A second dataset is constructed for the multi-view generation task, wherein the second dataset includes multiple single-image samples and multiple view image samples corresponding to each single-image sample. The following processing is performed using the pre-trained original multi-view generation model: The single image sample is subjected to feature extraction processing to obtain image feature samples; Based on the image feature samples, multi-view generation processing is performed to obtain view prediction images from multiple perspectives. The second loss value of the predicted view image and the view image sample of the multiple viewpoints is obtained by using a pre-set loss function; The network parameters of the original multi-view generation model are updated using the second loss value to obtain the trained multi-view generation model.

13. The method according to claim 12, characterized in that, The second dataset used to construct the multi-view generation task includes: Obtain multiple 3D model samples, and perform the following processing on each of the 3D model samples: The three-dimensional model sample is sampled according to the pre-set view parameters and lighting parameters of multiple viewpoints to obtain view image samples of the three-dimensional model sample from multiple viewpoints. The three-dimensional model sample is sampled according to the pre-set single-viewpoint parameters and lighting parameters to obtain a single-image sample of the three-dimensional model sample; The viewpoint image samples and the single image samples from multiple perspectives of each of the three-dimensional model samples are combined to form the second dataset.

14. The method according to claim 13, characterized in that, The viewpoint parameters of the multiple viewpoints include camera distance parameters. After obtaining the viewpoint image samples of the 3D model sample from multiple viewpoints, the method further includes: Edge detection is performed on each of the aforementioned viewpoint image samples to obtain edge detection results; In response to the edge detection result indicating that the edge of the model object in the view image sample exceeds the boundary of the view image sample, the camera distance parameter is increased by a preset parameter, and the three-dimensional model sample is sampled according to the increased camera distance parameter to obtain view image samples of multiple views of the new three-dimensional model sample.

15. The method according to any one of claims 1 to 5, characterized in that, The generation of the first 3D model based on the view images from the multiple viewpoints includes: Feature extraction processing is performed on the viewpoint image of each viewpoint to obtain the image features of each viewpoint; The image features of each viewpoint are subjected to feature encoding processing to obtain the encoded features of each viewpoint; The coded features of each viewpoint are decoded to obtain three-dimensional features; Based on the aforementioned three-dimensional features, a three-dimensional reconstruction process is performed to obtain the first three-dimensional model.

16. The method according to claim 1, characterized in that, After obtaining the second three-dimensional model, the method further includes: The texture of the second 3D model is optimized to obtain the model texture; The second three-dimensional model is topologically optimized to obtain the third three-dimensional model; The fourth 3D model is obtained by baking based on the model texture and the third 3D model.

17. A three-dimensional model generation device, characterized in that, The device includes: The data acquisition module is used to acquire model materials; A multi-view generation module is used to generate multiple perspective images based on the model material; A 3D model generation module is used to generate a first 3D model based on the view images from the multiple viewpoints, wherein the first 3D model includes a first color parameter; The 3D model generation module is also used to obtain the shading parameters of each viewpoint image; The 3D model generation module is further configured to update the first color parameter of the first 3D model based on the coloring parameters of each viewpoint image, thereby obtaining a second 3D model.

18. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, configured to execute computer-executable instructions or computer programs stored in the memory, implements the three-dimensional model generation method according to any one of claims 1 to 16.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the three-dimensional model generation method according to any one of claims 1 to 16 is implemented.

20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the three-dimensional model generation method according to any one of claims 1 to 16 is implemented.