Three-dimensional model generation method and apparatus, device, storage medium, and program product

By generating viewpoint images from multiple perspectives and updating the color parameters of the 3D model, the problems of lost details and unstable shape style in the existing technology of 3D model images are solved, thereby improving the generation quality and color consistency of the 3D model.

WO2026066811A1PCT designated stage Publication Date: 2026-04-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing technologies, 3D models generated through multi-view cross-domain attention mechanisms suffer from problems such as loss of image details and low quality. In particular, when generating 3D models based on text data, the shape style is unstable and cannot be directly applied to business scenarios.

Method used

By acquiring model materials, multiple perspective images are generated. The color parameters of the 3D model are updated based on the color parameters of the perspective images to ensure the consistency of color representation under any viewing angle and improve model quality.

Benefits of technology

It enhances the color details of the 3D model, improves the quality of the generated model, ensures high consistency of color representation from any viewing angle, and avoids local color deviation caused by a single-view image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025115445_02042026_PF_FP_ABST
    Figure CN2025115445_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a three-dimensional model generation method and apparatus, a device, a storage medium, and a program product. The method comprises: acquiring a model asset; generating view images from multiple views on the basis of the model asset; generating a first three-dimensional model on the basis of the view images from the multiple views, wherein the first three-dimensional model comprises a first color parameter; acquiring a coloring parameter of each view image; and on the basis of the coloring parameter of each view image, updating the first color parameter of the first three-dimensional model, to obtain a second three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional model generation method, device, equipment, storage medium and program product

[0001] Cross-reference to related applications

[0002] The present application is based on the Chinese patent application No. 2024114010924, filed on September 30, 2024, and claims priority to the Chinese patent application, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] The present application relates to the technical field of computers, and in particular to a three-dimensional model generation method, device, equipment, storage medium and program product. BACKGROUND

[0004] In the three-dimensional model generation method of the related art, a multi-view cross-domain attention mechanism is used to promote information exchange across views and modalities, and then a three-dimensional model is generated from a single-view image. However, compared with the input single-view image, there are still problems such as loss of image details, which makes the quality of the generated three-dimensional model low. SUMMARY

[0005] The embodiments of the present application provide a three-dimensional model generation method, device, equipment, storage medium and program product, which can improve the quality of the generated three-dimensional model.

[0006] The technical solutions of the embodiments of the present application are implemented as follows:

[0007] The embodiments of the present application provide a three-dimensional model generation method applied to an electronic device, and the method comprises the following steps:

[0008] Obtaining model materials;

[0009] Generating view images of multiple views based on the model materials;

[0010] Generating a first three-dimensional model based on the view images of the multiple views, wherein the first three-dimensional model comprises a first color parameter;

[0011] Obtaining a coloring parameter of each view image;

[0012] Updating the first color parameter of the first three-dimensional model based on the coloring parameter of each view image to obtain a second three-dimensional model.

[0013] The embodiments of the present application provide a three-dimensional model generation device, and the device comprises:

[0014] A data acquisition module configured to acquire model materials;

[0015] a multi-view generation module configured to generate perspective images of multiple perspectives based on the model material;

[0016] a three-dimensional model generation module configured to generate a first three-dimensional model based on the perspective images of the multiple perspectives, wherein the first three-dimensional model comprises a first color parameter;

[0017] the three-dimensional model generation module is further configured to obtain a shading parameter of each of the perspective images;

[0018] the three-dimensional model generation module is further configured to update the first color parameter of the first three-dimensional model based on the shading parameter of each of the perspective images to obtain a second three-dimensional model.

[0019] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:

[0020] a memory configured to store computer executable instructions or computer programs;

[0021] a processor configured to execute the computer executable instructions or computer programs stored in the memory to implement the three-dimensional model generation method provided in the embodiments of the present application.

[0022] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores computer programs or computer executable instructions, and is configured to be executed by a processor to implement the three-dimensional model generation method provided in the embodiments of the present application.

[0023] A computer program product is provided in an embodiment of the present application, and the computer program product comprises computer programs or computer executable instructions, and the computer programs or computer executable instructions are executed by a processor to implement the three-dimensional model generation method provided in the embodiments of the present application.

[0024] The embodiments of the present application have the following beneficial effects:

[0025] After the model material and the perspective images of multiple perspectives generated based on the model material are used to generate a first three-dimensional model, the first color parameter of the first three-dimensional model is updated based on the shading parameter of each perspective image, the first color parameter of the first three-dimensional model in different perspectives is updated and calibrated, the first color parameter is more consistent with the real color attribute of the model material after the update, the color performance of the second three-dimensional model at any observation angle is kept highly consistent, local color deviation caused by the limitation of a single perspective image is avoided, the color details of the second three-dimensional model are further enhanced, and the quality of the generated three-dimensional model is improved. BRIEF DESCRIPTION OF DRAWINGS

[0026] FIG. 1 is a structural schematic diagram of a three-dimensional model generation system architecture provided in an embodiment of the present application;

[0027] FIG. 2 is a structural schematic diagram of a server according to an embodiment of the present application;

[0028] FIG. 3A is a first flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0029] FIG. 3B is a second flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0030] FIG. 3C is a third flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0031] FIG. 3D is a fourth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0032] FIG. 3E is a fifth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0033] FIG. 3F is a sixth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0034] FIG. 3G is a seventh flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0035] FIG. 3H is an eighth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0036] FIG. 3I is a ninth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0037] FIG. 3J is a tenth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0038] FIG. 3K is an eleventh flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0039] FIG. 3L is a twelfth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0040] FIG. 3M is a thirteenth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0041] FIG. 3N is a fourteenth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0042] FIG. 3O is a fifteenth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0043] FIG. 3P is a sixteenth flowchart of a three-dimensional model generation method according to an embodiment of the present application;

[0044] FIG. 4 is an application scenario diagram of a three-dimensional model generation method according to an embodiment of the present application;

[0045] FIG. 5 is a schematic diagram of a principle framework of a three-dimensional model generation method according to an embodiment of the present application;

[0046] FIG. 6 is a schematic diagram of a flow of a three-dimensional model generation method for a game scene according to an embodiment of the present application;

[0047] FIG. 7A is a first schematic diagram of a three-dimensional model generation interface according to an embodiment of the present application;

[0048] FIG. 7B is a second schematic diagram of a three-dimensional model generation interface according to an embodiment of the present application;

[0049] FIG. 8 is a schematic diagram of an architecture of a three-dimensional model generation platform according to an embodiment of the present application;

[0050] FIG. 9 is a comparative schematic diagram of rendering effects of three-dimensional models under different lighting parameters according to an embodiment of the present application;

[0051] FIG. 10 is a schematic diagram of a fixed view angle according to an embodiment of the present application;

[0052] FIG. 11 is a schematic diagram of a construction principle of fine-tuning weight parameters according to an embodiment of the present application;

[0053] FIG. 12 is a comparative schematic diagram of image material generation effects of an original image generation model and a trained image generation model according to an embodiment of the present application;

[0054] FIG. 13 is a schematic diagram of multiple view angle sampling according to an embodiment of the present application;

[0055] FIG. 14 is a schematic diagram of a data structure of a second data set according to an embodiment of the present application;

[0056] FIG. 15A is a first schematic diagram of view angle image generation effects of multiple view angles according to an embodiment of the present application;

[0057] FIG. 15B is a second schematic diagram of view angle image generation effects of multiple view angles according to an embodiment of the present application;

[0058] FIG. 15C is a third schematic diagram of view angle image generation effects of multiple view angles according to an embodiment of the present application;

[0059] FIG. 15D is a fourth schematic diagram of view angle image generation effects of multiple view angles according to an embodiment of the present application;

[0060] FIG. 16A is a first schematic diagram of a principle of obtaining shading parameters of each view angle image according to an embodiment of the present application;

[0061] FIG. 16B is a second schematic diagram of a principle of obtaining shading parameters of each view angle image according to an embodiment of the present application;

[0062] FIG. 16C is a third schematic diagram of a principle of obtaining the shading parameter of each view image according to an embodiment of the present application;

[0063] FIG. 16D is a fourth schematic diagram of a principle of obtaining the shading parameter of each view image according to an embodiment of the present application;

[0064] FIG. 17A is a first comparison schematic diagram of a first three-dimensional model and a second three-dimensional model according to an embodiment of the present application;

[0065] FIG. 17B is a second comparison schematic diagram of the first three-dimensional model and the second three-dimensional model according to an embodiment of the present application;

[0066] FIG. 17C is a third comparison schematic diagram of the first three-dimensional model and the second three-dimensional model according to an embodiment of the present application;

[0067] FIG. 17D is a fourth comparison schematic diagram of the first three-dimensional model and the second three-dimensional model according to an embodiment of the present application;

[0068] FIG. 18 is a schematic diagram of an effect of texture optimization according to an embodiment of the present application;

[0069] FIG. 19 is a schematic diagram of an effect of topology optimization according to an embodiment of the present application;

[0070] FIG. 20 is a schematic diagram of an effect of surface reduction according to an embodiment of the present application;

[0071] FIG. 21A is a schematic diagram of a principle of post-processing according to an embodiment of the present application;

[0072] FIG. 21B is a flow schematic diagram of post-processing according to an embodiment of the present application;

[0073] FIG. 22 is a schematic diagram of a network structure of a multi-view generation model according to an embodiment of the present application.

[0074] It should be noted that the above-mentioned “first” and “second” are only used to distinguish different schemes, and do not represent the advantages or disadvantages of the schemes or the priority in the implementation process. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be described in further detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by a person of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.

[0076] In the following description, “some embodiments” are related to a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0077] In the following description, the terms "first\second\third" are merely distinguished similar objects, and do not represent a specific order of the objects. It can be understood that the "first\second\third" can be interchanged with a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0078] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.

[0079] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by one skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0080] The relevant data collection process in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and within the scope of authorization of laws and regulations and the personal information subject, carry out subsequent data use and processing.

[0081] Before further detailing the embodiments of the present application, the terms and terms involved in the embodiments of the present application are explained, and the terms and terms involved in the embodiments of the present application are applicable to the following explanations.

[0082] 1) Model material, a two-dimensional model image, is used to provide an image as a reference to visually indicate the appearance of a three-dimensional model, thereby guiding the style of the generated three-dimensional model. The model material may, for example, be a hand-drawn sketch, a photograph, a design drawing, or a screenshot of another three-dimensional model. The model material can also be generated by text data. Here, the text data describes the features of the three-dimensional model to be generated using natural language, which can include shape, color, texture, size, proportion, style, etc. so that image material generation processing (text-to-image processing) is performed based on the text data to obtain the corresponding model material.

[0083] 2) Three-Dimensional Model (3D Model), refers to a mathematical representation or computer graphics representation of an object in three-dimensional space. Three-dimensional models are widely used in computer graphics, animation, virtual reality, game development, industrial design, architecture, medical imaging, and other fields. A three-dimensional model can be composed of vertices, edges, and faces, which define the shape and structure of the model.

[0084] 3) Camera View, refers to the position and orientation of a camera in three-dimensional space, which determines the image content of the three-dimensional model captured by the camera. The camera view can include parameters such as pitch and azimuth, which together define the orientation and view angle of the camera.

[0085] The following is an explanation of the two main parameters in the camera view:

[0086] Pitch: is the angle between the camera and the horizontal plane (or reference horizontal plane), used to represent the degree of camera tilt upwards or downwards. The size of the pitch angle determines the pitch state of the camera relative to the horizontal plane. The pitch angle is calculated from 0 degrees, 0 degrees represents the camera pointing horizontally forward, positive values represent the camera tilting upwards, and negative values represent the camera tilting downwards, for example, 0 degrees: the camera points horizontally forward; 90 degrees: the camera points straight up; -90 degrees: the camera points straight down.

[0087] Azimuth: is the angle of the camera's left and right rotation, which determines the direction the camera faces. The degree of azimuth is usually between 0 degrees and 360 degrees, 0 degrees and 360 degrees represent the same direction, for example, 0 degrees / 360 degrees: the camera faces north; 90 degrees: the camera faces east; 180 degrees: the camera faces south; 270 degrees: the camera faces west.

[0088] These angles together define the camera's view, allowing precise control of the camera's orientation in three-dimensional space, thus capturing the desired image of the three-dimensional model.

[0089] 4) Multiple View Images, refers to a collection of images obtained by observing the same three-dimensional model from different camera views. In the field of three-dimensional modeling and computer vision, this image collection is often used to reconstruct, analyze, or render three-dimensional models.

[0090] 5) First Color Parameter, refers to parameters used to describe and specify the color characteristics of the surface of the first three-dimensional model. These parameters can affect the appearance and texture of the model under different lighting conditions.

[0091] 6) Color parameter, refers to the parameter for re-coloring the three-dimensional model, the color parameter includes the viewpoint coordinate and the color value (i.e., the second color parameter), wherein the viewpoint coordinate refers to the position of the observer or the camera, which defines the position and direction of the observer or the camera relative to the three-dimensional model, and is used to determine the viewing angle of the observer or the camera observing the three-dimensional model, and the color value refers to the value of the second color parameter of the pixel point in the view image mapped to the first three-dimensional model through the viewpoint coordinate.

[0092] 7) First vertex, refers to the vertex at the tip of the four-sided pyramid (i.e., not on the bottom surface), which is the highest point of the four-sided pyramid, and its position determines the height of the four-sided pyramid. For example, if the bottom surface of the four-sided pyramid is located on the XY plane and the center of the bottom surface of the four-sided pyramid is the origin, then the coordinates of the first vertex can be represented as (0, 0, h), where h is the height of the four-sided pyramid.

[0093] 8) Second vertex, is one of the basic units that make up the three-dimensional model. A vertex is a point in three-dimensional space, usually defined by its coordinates, which can be Cartesian coordinates (x, y, z) or coordinates in other coordinate systems. Each vertex can include corresponding color parameters to construct a three-dimensional model with multiple colors.

[0094] 9) Diffusion Model, is a generative model that simulates the process of gradually covering data samples with noise (forward diffusion), and then learns how to remove these noises in reverse to recover the original data (reverse diffusion).

[0095] 10) Forward Process, refers to gradually adding Gaussian noise to the original data (such as single image samples), and after multiple steps, the data gradually becomes pure noise (noise image). This process is similar to "destroying" the data, and ultimately obtaining a random distribution unrelated to the original data.

[0096] 11) Reverse Process, refers to the process of gradually removing noise to reconstruct high-quality samples from random noise. At each step of this process, the diffusion model attempts to recover some of the original data (such as view image samples) from the current noisy data. After multiple iterations, a new sample similar to the original data (such as a view prediction image similar to the view image sample) is recovered.

[0097] In the related art three-dimensional model generation method, through the multi-view cross-domain attention mechanism, the information exchange across views and modalities is promoted, and then the three-dimensional model is generated from the single view image, but there are still problems such as poor consistency of multi-view and loss of image details compared to the input single view image, which makes the quality of the generated three-dimensional model low.

[0098] In addition, in the method for generating a three-dimensional model based on text data in the related art, text data is converted into a single-view image through text-to-image processing, and then a three-dimensional model is generated based on the single-view image. However, due to the instability of the generation style of the image obtained through text-to-image processing, the appearance of the three-dimensional model generated based on the text data has an uncontrollable style, and thus cannot be directly applied to a business scenario.

[0099] The embodiment of the present application provides a three-dimensional model generation method, device, equipment, computer readable storage medium and computer program product, which can improve the quality of the generated three-dimensional model.

[0100] The following describes an exemplary application of the electronic device provided by the embodiment of the present application. The electronic device provided by the embodiment of the present application can be implemented as a notebook computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart television, a vehicle-mounted terminal and various types of terminal devices, and can also be implemented as a server.

[0101] Referring to FIG. 1, FIG. 1 is a structural schematic diagram of a three-dimensional model generation system architecture provided by the embodiment of the present application, and FIG. 1 involves a server 100, a terminal device 200 and a network 300. The terminal device 200 is connected to the server 100 through the network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two.

[0102] In some embodiments, the embodiment of the present application can be implemented by the server and the terminal device in cooperation. For example, the terminal device 200 sends model materials to the server 100, and the server 100 obtains a second three-dimensional model based on the model materials through the three-dimensional model generation method provided by the embodiment of the present application, and sends the second three-dimensional model to the terminal device 200.

[0103] In some other embodiments, the embodiment of the present application can be implemented by the terminal device alone. For example, the terminal device 200 obtains model materials, and generates a second three-dimensional model based on the model materials through the three-dimensional model generation method provided by the embodiment of the present application.

[0104] In some other embodiments, the embodiment of the present application can be implemented by the server alone. For example, the server 100 obtains model materials, and generates a second three-dimensional model based on the model materials through the three-dimensional model generation method provided by the embodiment of the present application.

[0105] In some embodiments, the server 100 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal device and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the embodiments of the present application.

[0106] Taking the server for generating a three-dimensional model as an example, referring to FIG. 2, which is a structural schematic diagram of the server provided in the embodiments of the present application, the server 100 shown in FIG. 2 includes at least one processor 110, a memory 130, and at least one network interface 120. The various components in the server 100 are coupled together by a bus system 140. It can be understood that the bus system 140 is used to realize the connection and communication between the components. In addition to the data bus, the bus system 140 also includes a power bus, a control bus, and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 140 in FIG. 2.

[0107] The processor 110 can be an integrated circuit chip with signal processing capability, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor.

[0108] The memory 130 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard drives, optical drives, and the like. The memory 130 optionally includes one or more storage devices remotely located from the processor 110 in a physical location.

[0109] The memory 130 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 130 described in the embodiments of the present application is intended to include any suitable type of memory.

[0110] In some embodiments, the memory 130 can store data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, which are exemplarily illustrated below.

[0111] The operating system 131 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks;

[0112] The network communication module 132 is configured to communicate with other electronic devices via one or more (wired or wireless) network interfaces 120, such as Bluetooth, Wireless Fidelity (WiFi), Universal Serial Bus (USB), and the like.

[0113] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in a software manner. FIG. 2 shows a three-dimensional model generation apparatus 133 stored in the memory 130, which can be a software in the form of a program and a plug-in, and includes the following software modules: a data acquisition module 1331, a multi-view generation module 1332, and a three-dimensional model generation module 1333. These modules are logical, and thus can be combined or further split according to the implemented functions. The functions of the various modules will be described below.

[0114] In some embodiments, the terminal device or the server can implement the three-dimensional model generation method provided by the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. The computer program can be a native program in the operating system or a software module; can be a native application (APP), i.e., a program that needs to be installed in the operating system to run, such as a three-dimensional modeling APP or a game development APP (for creating three-dimensional game characters, props, and scenes, etc.); or can be a small program that can be embedded into any APP, i.e., a program that only needs to be downloaded into a browser environment to run. In summary, the above computer-executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules, or plug-ins.

[0115] In some embodiments, the device provided by the embodiments of the present application can be implemented in a hardware manner. For example, the device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the three-dimensional model generation method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can be implemented by using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic elements.

[0116] The three-dimensional model generation method provided by the embodiments of the present application will be described below in combination with an exemplary application and implementation of a server provided by the embodiments of the present application, taking the server as an execution subject. Referring to FIG. 3A, FIG. 3A is a first flowchart of the three-dimensional model generation method provided by the embodiments of the present application. The steps shown in FIG. 3A will be described below.

[0117] In step 101, model materials are obtained.

[0118] In some embodiments, the model materials are two-dimensional model images. Referring to FIG. 3B, the step 101 shown in FIG. 3A can be implemented by the following steps 1011 to 1012, which will be described in detail below.

[0119] In step 1011, text data is obtained, wherein the text data is used to describe a three-dimensional model.

[0120] In some embodiments, the text data submitted by a terminal device to a three-dimensional model generation platform is obtained. The text data can include the shape, color, texture, size, proportion, style, etc. of a three-dimensional model expected to be generated. For example, the text data can be represented as “a blue cylinder with a red ring at the top, 10 cm in height, and 5 cm in diameter”. The embodiments of the present application do not limit the specific content contained in the text data.

[0121] In step 1012, image material generation processing is performed based on the text data to obtain model materials.

[0122] In some embodiments, the image material generation processing is performed based on the text data by using a pre-trained image generation model to obtain the model materials.

[0123] Through steps 1011 to 1012, a two-dimensional model image conforming to the semantics of the text data can be generated as a model material based on the text data, realizing direct conversion from a literal description to a visual material and providing image data basis that meets the needs for generating perspective images of multiple perspectives.

[0124] In some embodiments, referring to FIG. 3C, the image generation model is trained through steps 201 to 207, which are specifically explained as follows.

[0125] In step 201, a first data set for an image material generation task is constructed, where the first data set includes a plurality of image samples and a text description of each image sample, and the plurality of image samples are of the same style.

[0126] In some embodiments, referring to FIG. 3D, step 201 shown in FIG. 3C can be implemented through steps 2011 to 2014, which are specifically explained as follows.

[0127] In step 2011, a plurality of three-dimensional model samples are obtained, where the plurality of three-dimensional model samples are of the same style and different shapes.

[0128] In some embodiments, the plurality of three-dimensional model samples are of the same style and different shapes, for example, three-dimensional model samples of a classical style can be a "classical style sofa", a "classical style chair", and the like. For example, a professional three-dimensional model library or an online resource platform such as TurboSquid, CGTrader, and the like can be accessed to find models with the same style but different shapes by using a search function, for example, by using a keyword filter, for example, by specifying a cartoon style, to find three-dimensional model samples of the same style but different shapes. The embodiments of the present application do not limit the specific acquisition method of the three-dimensional model samples.

[0129] In step 2012, each three-dimensional model sample is rasterized according to a pre-set fixed perspective to obtain an image sample corresponding to each three-dimensional model sample.

[0130] In some embodiments, the rendering parameters such as the lighting parameter (for example, setting the light source (lighting parameter) as an ambient light), the fixed perspective (for example, setting the camera perspective as an azimuth angle of 30° and an elevation angle of 20°), and the like can be set in a rendering engine or software for rendering three-dimensional models to simulate the lighting and shadow effects in the real world, so as to perform rasterization to obtain the image sample corresponding to the three-dimensional model sample. Here, the rasterization is part of the rendering process, which is used to convert the three-dimensional model sample into a two-dimensional image sample, that is, the rasterization is a process of converting the geometric data of the three-dimensional model sample into pixels on a display device.

[0131] Referring to FIG. 9, FIG. 9 is a comparison diagram of rendering effects of a three-dimensional model under different lighting parameters, provided by an embodiment of the present application. Here, since the three-dimensional model itself does not have a lighting effect, setting the light source as an ambient light can avoid the appearance of highlight effects in the image sample, thereby ensuring the non-lighting property of the subsequently generated three-dimensional model. As shown in FIG. 9, using a point light source (left figure) will cause obvious highlights on the surface of the three-dimensional model, presenting a light and dark contrast effect formed by a light with a specific direction, similar to the light and shadow of a single light source (such as a light bulb) in reality; using an ambient light (right figure), the lighting is uniform, without obvious highlights, and the light and dark transition on the surface of the three-dimensional model is natural, which can present a non-specific direction and softer overall lighting effect, and can ensure the non-lighting property related requirements of the three-dimensional model, similar to the lighting of global diffuse light (such as an overcast environment) in reality.

[0132] For example, the rasterization process can be implemented in the following way: first, set the position, view angle (i.e., fixed view angle), and field of view (FOV) of the camera, wherein the camera position defines the position of the observer, the view angle defines the orientation of the observer, and the field of view defines the range that can be observed by the camera; next, convert the three-dimensional model sample to the world coordinate system, convert the position and direction of the camera to the camera coordinate system, apply a view transformation matrix to convert the three-dimensional model sample in the world coordinate system to the camera coordinate system, and apply a projection transformation matrix to further convert the three-dimensional model sample in the camera coordinate system to the clip space, for example, by perspective projection or orthogonal projection; next, in the clip space, identify and remove those parts of the three-dimensional model that are outside the view range and those parts that are occluded; convert the coordinates in the clip space to the screen space (Normalized Device Coordinates, NDC), which maps the coordinates to the range of -1 to 1, corresponding to the lower left corner to the upper right corner of the screen; next, convert the primitives (usually triangles) in the screen space to pixels through a rasterizer, which includes determining which pixels belong to the primitive and how to interpolate the vertex attributes (such as color, texture coordinates, normal, etc.) on these pixels; next, perform shading processing on each pixel, such as lighting calculation, texture mapping, shadow processing, etc.; finally, output the rendering result to an image buffer, save the content in the image buffer as a file as an image sample, for example, save the image sample as a PNG, JPEG, and HDR image format.

[0133] Here, when the rendering engine renders or displays the three-dimensional model, the rendering engine constructs a local virtual coordinate system centered on the three-dimensional model (which can be understood as a "virtual observation space" around the three-dimensional model), and the camera position is defined as the three-dimensional coordinate point (such as the X, Y, Z axis coordinates with the center of the three-dimensional model as the origin) where the "observer" (i.e., the virtual camera) is located in this local coordinate system. The camera position is used to control the relative spatial relationship between the camera and the three-dimensional model, ensuring that the two-dimensional image obtained by rasterization processing accurately presents the form of the three-dimensional model under the corresponding viewing angle.

[0134] By fixing the viewing angle (setting the camera viewing angle to an azimuth angle of 30° and a vertical angle of 20°), the features of multiple faces of the three-dimensional model sample can be observed, and more details can be captured, thereby achieving the beneficial effect of improving the accuracy of image material generation processing, thereby better guiding the generation of the three-dimensional model. For example, refer to FIG. 10, which is a schematic diagram of a fixed viewing angle provided by an embodiment of the present application. The same viewing angle is used for rasterization processing for each three-dimensional model sample.

[0135] In step 2013, text generation processing is performed on each image sample to obtain a text description corresponding to each image sample.

[0136] In some embodiments, the text description of the image sample can be generated by a large language model (or a graph-to-text model). The image sample is input into the large language model to generate the text description.

[0137] Taking the graph-to-text model Pix2Text as an example, the text description of the image sample can be generated in the following manner: first, the image sample is preprocessed, such as grayscale, normalization, etc.; next, feature extraction processing is performed on the preprocessed image sample, such as using a convolutional neural network (CNN) or other image processing techniques to identify features such as texture, shape, color, etc. in the image sample, and encoding these information into a high-dimensional vector to obtain image features; next, the processed image features are converted into a fixed-length vector, which contains rich information of the image sample and is used for subsequent text generation; finally, feature mapping processing is performed based on the fixed-length vector to obtain the text description, such as mapping the vector through a long short-term memory (LSTM) network to generate the text description. After obtaining the text description, it can also be organized and corrected to improve the accuracy of the text description.

[0138] For example, for an image sample containing a game character, the output text description can be: "In this image, a male character wearing a blue shirt can be seen, the character's head adopts a small polygon mesh to reduce rendering complexity, the body part adopts a high-precision polygon mesh to accurately depict the character's muscle structure and posture, the character's clothes are composed of polygon meshes, and the design is meticulous, which can reflect the texture and pattern of the clothes."

[0139] In some embodiments, the text description of each three-dimensional model sample can also be generated by manual annotation.

[0140] In step 2014, each image sample and the corresponding text description are combined into a first data set.

[0141] In some embodiments, the first data set can include image data (such as file path, binary data, pixel matrix, etc.) of the image sample, text description (string form) of the image sample, and other metadata (such as image source information, label, etc.), and the embodiments of the present application do not limit the data structure and data storage method of the first data set.

[0142] Through steps 2011 to 2014, based on three-dimensional model samples of the same style and different shapes, image samples with consistent features are obtained through fixed-view rasterization processing, and combined with accurate text descriptions, the strong association and style uniformity of images and texts in the first data set are ensured, and the details captured by the fixed view are improved. The data quality provides high-quality data support for the training of subsequent image generation models, thereby enhancing the ability of the image generation model to generate model materials that meet the text description, style uniformity and accurate details.

[0143] Referring to FIG. 3C, in step 202, a pre-trained original image generation model is obtained, and the following processing is performed through the original image generation model.

[0144] In step 203, the text description is subjected to feature extraction processing to obtain text features.

[0145] In some embodiments, the text description is subjected to Tokenization processing to obtain a plurality of input units (tokens); the plurality of input units are subjected to embedding encoding processing to obtain embedding features; and the embedding features are subjected to attention encoding processing to obtain text features.

[0146] For example, assuming that the input text description is "thank you very much" and the tokenization processing results in multiple input units, such as [“thank”, “you”, “very”, “much”], the multiple input units are then subjected to embedding encoding processing. Assuming that each input unit is mapped to a 2048-dimensional vector, the text description "thank you very much" is converted into 4*2048 embedding features. The 4*2048 embedding features are then subjected to attention encoding processing to obtain text features. Here, an input unit refers to a basic processing unit of a text description. An input unit can be a character, a word, a subword, a character, or another meaningful element, depending on the granularity and requirements of text processing. The embodiments of the present application do not impose any limitation.

[0147] For example, the multiple input units can be subjected to embedding encoding processing by an embedding layer to obtain embedding features. For example, the embedding layer can use word embedding, character embedding, subword embedding, or other embedding methods to perform embedding encoding processing to obtain embedding features.

[0148] For example, the attention weights of the embedding features are calculated by an attention mechanism (such as a self-attention mechanism), which represent the importance of the embedding features for the current task. The attention weights of the embedding features are multiplied by the embedding features to obtain a weighted embedding feature, i.e., a text feature.

[0149] In some other embodiments, the text description can also be subjected to feature extraction processing by a recurrent neural network (RNN) or other network model to obtain text features. The embodiments of the present application do not impose any limitation on the specific implementation of the feature extraction processing of the text description.

[0150] For example, the text description can be subjected to feature extraction processing by an RNN in the following manner. Before the text description is input into the RNN, preprocessing can be performed, including tokenization, part-of-speech tagging, removal of stop words, and other steps. Next, each word in the text description is converted into a numerical vector to form a numerical sequence, for example, by using a bag-of-words model, a pre-trained word embedding (such as Word2Vec), or other methods. Finally, the numerical sequence is input into the RNN, which learns how to capture the sequence dependency in the text description and produces a fixed-size vector representation, which is the text feature.

[0151] For example, word segmentation can be implemented in the following manner: according to the language characteristics of the text description (for example, Chinese uses jieba and other tools for word or word-level splitting, English is naturally divided by spaces and punctuation), the continuous text description is divided into independent input units, for example, "red sofa" is divided into "red", "of", and "sofa"; for text containing special symbols, numbers or professional terms, the complete semantics is retained (for example, "3D model" is not split into "3", "D", and "model"), to ensure that the split token accurately reflects the meaning of the original text.

[0152] For example, part-of-speech tagging can be implemented in the following manner: each input unit obtained by word segmentation is assigned a corresponding part-of-speech tag (such as noun, verb, adjective, etc.) by a part-of-speech tagging tool (such as NLTK, spaCy, etc.), for example, "red" (adjective) and "sofa" (noun) are tagged with part of speech.

[0153] For example, removing stop words can be implemented in the following manner: according to a pre-set stop word list (containing words with no actual semantics or very small contribution to text features, such as "of" and "in" in Chinese; "the" and "is" in English), the stop words appearing after word segmentation are filtered and removed, for example, "of" is removed from "red sofa is very beautiful", and only core input units such as "red", "sofa", and "very beautiful" are retained, thereby compressing the length of the text sequence and improving the learning efficiency of RNN on key semantic information.

[0154] In step 204, image generation processing is performed based on the text features to obtain a predicted image.

[0155] In some embodiments, the extracted text features are mapped to an image space to obtain a predicted image, for example, a generative adversarial network (GAN) or other image generation model is used to generate a predicted image based on the text features, where the GAN consists of a generator and a discriminator, the generator is responsible for generating a predicted image, and the discriminator is responsible for judging whether the generated predicted image conforms to the distribution of real data. The embodiments of the present application do not limit the image generation model used for image generation processing.

[0156] In step 205, a first loss value of the predicted image and the image sample is obtained by a pre-set loss function.

[0157] In some embodiments, the first loss value between the predicted image and the image sample is calculated by a pre-set loss function (such as cross-entropy loss).

[0158] For example, the first loss value between the predicted image and the image sample can be calculated by formula (1):

[0159] Wherein, L1 represents the first loss value, N is the total number of pixels of the image sample or the predicted image; yi is the pixel value of the i th pixel in the image sample; is the pixel value of the i th pixel in the predicted image; log is the logarithmic function.

[0160] In step 206, the fine-tuning weight parameter is constructed, and the fine-tuning weight parameter is added to the network parameter of the original image generation model.

[0161] In some embodiments, referring to FIG. 3E, the construction of the fine-tuning weight parameter in step 206 shown in FIG. 3C can be realized by the following steps 2061 to 2062, which are described in detail below.

[0162] In step 2061, a first low-rank matrix initialized based on Gaussian distribution is generated, and a second low-rank matrix initialized based on all-zero is generated.

[0163] In some embodiments, the dimension of the first low-rank matrix can be determined according to the size of the weight matrix of the fully connected layer of the original image generation model, and the dimension of the second low-rank matrix is the same as the transpose of the first low-rank matrix, wherein the first low-rank matrix is initialized by using random numbers generated by Gaussian distribution (normal distribution), and the Gaussian distribution is defined by mean (Mean) and standard deviation (Standard Deviation, STD) parameters, and all elements of the second low-rank matrix are initialized to zero.

[0164] For example, if the weight matrix of the fully connected layer of the original image generation model is m x n, then the first low-rank matrix can be an m x k matrix, where k is the rank in the low-rank approximation, and the dimension of the second low-rank matrix is the same as the transpose of the first low-rank matrix.

[0165] In step 2062, the fine-tuning weight parameter is constructed based on the first low-rank matrix and the second low-rank matrix.

[0166] In some embodiments, the fine-tuning weight parameter is constructed based on the product of the first low-rank matrix and the second low-rank matrix, referring to FIG. 11, which is a schematic diagram of the construction principle of the fine-tuning weight parameter provided by the embodiments of the present application, as shown in FIG. 11, the product of the first low-rank matrix (A) and the second low-rank matrix (B) constitutes the fine-tuning weight parameter (ΔW), wherein the dimension of the second low-rank matrix is the same as the transpose of the first low-rank matrix.

[0167] Through steps 2061 to 2062, the initialization of the first low-rank matrix based on the Gaussian distribution can give the parameter initial exploration ability, and the all-zero initialization of the second low-rank matrix can ensure the minimum interference of the fine-tuning initial stage to the original image generation model; and the fine-tuning weight parameter constructed through the low-rank matrix product can not only greatly reduce the parameter dimension (reduce the training calculation amount), but also accurately capture the key adjustment direction of the full connection layer weight of the original image generation model. In this way, the fine-tuning process is more targeted, which can efficiently adapt to specific style or task requirements while retaining the core image generation capability of the original image generation model, thereby laying a foundation for the accurate optimization of the subsequent image generation model.

[0168] With reference back to FIG. 3C, in step 207, the fine-tuning weight parameter is updated through the first loss value to obtain the trained image generation model.

[0169] In some embodiments, the parameters in the network parameters except the first low-rank matrix and the second low-rank matrix remain unchanged, and the first low-rank matrix and the second low-rank matrix are updated through the first loss value.

[0170] In an example, the original network parameters (weight parameters) of the original image generation model are denoted as W0∈R d×k The original image generation model is fine-tuned through a fine-tuning weight parameter (ΔW), so that the weight parameter of the fine-tuned (i.e., trained) image generation model is W′=W0+ΔW. Then, ΔW is decomposed into the product of two low-rank matrices through rank decomposition, i.e., ΔW=B×A, where B∈R d×r (corresponding to the second low-rank matrix) and A∈R r×k (corresponding to the first low-rank matrix), and r<<min(k,d). In the backward update process, W0 is frozen (not updated, corresponding to keeping the parameters in the network parameters except the first low-rank matrix and the second low-rank matrix unchanged), and only the parameters of A and B are updated. In this way, the forward propagation of the trained image generation model can be represented as h=W′x=W0x+ΔWx=W0x+BAx, where x represents the input of the image generation model.

[0171] Referring to FIG. 12, FIG. 12 is a comparison diagram of the model material generation effects of the original image generation model and the trained image generation model according to an embodiment of the present application. In the diagram, the keywords are used to represent the types of the model objects in the model materials, which can be obtained by keyword extraction processing on the text description. Here, the styles of the plurality of image samples in the first data set used for training are cartoon styles. It can be seen that the style of the model materials generated by the trained image generation model is more uniform than that of the original image generation model. The image material generation processing through the trained image generation model can achieve the effect of controlling the style of the model materials, thereby better guiding the generation of the three-dimensional model.

[0172] Referring to FIG. 12, the model materials generated by the trained image generation model relative to the original image generation model have differences in style uniformity and content consistency, which are described in detail as follows:

[0173] 1) Style uniformity

[0174] The trained image generation model: taking "cartoon style" as the core training target, the generated images of all categories (such as bed, TV, car, etc.) present a low polygon, soft color, and simple line cartoon visual language. For example, the retro style of TV and the round contour of car all follow the "cartoon rendering logic" and have a high degree of style consistency.

[0175] The original image generation model: the style has no clear direction, and the visual language of images of different categories is chaotic (such as "realistic style" of TV), lacking a unified style framework.

[0176] 2) Content consistency

[0177] The trained image generation model: the generated images under the same keyword (such as "bed", "TV", etc.) have a high degree of focus on functional form. Taking "bed" as an example, the three images are all designed around the "basic structure of cartoonized bed (bed frame, mattress, bed head)", and only the details (bed head style, bed leg shape) are varied; the TV is uniformly "retro cartoon TV outline", and the functional form is clear and consistent.

[0178] The original image generation model: the images under the same keyword have great differences in functional form. For example, under the keyword "TV", there are images of "wall-mounted modern TV" and "retro TV", which are completely different in form, and the content consistency is poor.

[0179] For example, the original image generation model can be a potential diffusion model for improving high-resolution image synthesis (Stable Diffusion XL, SD XL), an image generation model based on GAN network, etc., and the present application does not limit the specific original image generation model.

[0180] Through steps 201 to 207, an image generation model with unified style and high generation accuracy can be trained, that is, first, based on three-dimensional model samples with consistent style and different shapes, image samples are generated through fixed-view rasterization processing, and a high-quality first data set is constructed by combining accurate text descriptions, thereby providing a learning basis for image generation model training with unified style and rich details; then, by using a pre-trained original image generation model, through processes such as feature extraction, image generation, and loss calculation, and in combination with fine-tuning weight parameters constructed by a low-rank matrix, targeted optimization is performed, only fine-tuning parameters are updated, and core parameters of the original image generation model are frozen, thereby greatly reducing training costs and improving optimization efficiency while retaining the basic generation capability of the image generation model. Finally, the trained image generation model can accurately generate model materials with consistent style and rich details that meet the text description, thereby providing reliable support for subsequent high-quality generation of three-dimensional models.

[0181] In some embodiments, the image data submitted by the terminal device to the three-dimensional model generation platform is directly used as model materials. The image data can include surface texture details, color distribution characteristics, and key structure outlines of the three-dimensional model to be generated. For example, the image data uploaded by the user can cover the pattern texture, color gradient effect, and light reflection difference of different materials such as metal and plastic on the surface of the three-dimensional model to be generated, as well as the structure information such as protrusions and depressions of the three-dimensional model.

[0182] Referring back to FIG. 3A, in step 102, view images of multiple views are generated based on the model materials.

[0183] In some embodiments, the model materials are two-dimensional model images. Referring to FIG. 3F, step 102 shown in FIG. 3A can be implemented through the following steps 1021 to 1022, which are described in detail as follows.

[0184] In step 1021, the model materials are subjected to feature extraction processing by a pre-trained multi-view generation model to obtain model image features.

[0185] In some embodiments, referring to FIG. 3G, the multi-view generation model is trained through the following steps 301 to 305, which are described in detail as follows.

[0186] In step 301, a second data set for a multi-view generation task is constructed, wherein the second data set includes a plurality of single-image samples and a plurality of view images corresponding to each single-image sample.

[0187] In some embodiments, referring to FIG. 3H, step 301 shown in FIG. 3G can be implemented through the following steps 3011 to 3014, which are described in detail as follows.

[0188] In step 3011, a plurality of three-dimensional model samples are obtained, and the following processing is performed on each three-dimensional model sample.

[0189] In some embodiments, a professional three-dimensional model library or online resource platform such as TurboSquid, CGTrader, etc. can be accessed to obtain a plurality of three-dimensional model samples. The embodiments of the present application do not limit the manner of obtaining the three-dimensional model samples. Here, the plurality of three-dimensional model samples can be of the same category, for example, all being three-dimensional models of the "furniture" category.

[0190] In step 3012, the three-dimensional model samples are sampled according to the view angle parameters and the lighting parameters of the plurality of view angles that are set in advance, to obtain view angle image samples of the plurality of view angles of the three-dimensional model samples.

[0191] In some embodiments, the lighting parameters (for example, the lighting parameters (light source) are set to ambient light), the plurality of view angles (for example, the plurality of camera view angles are respectively set to the corresponding azimuth angles of 30°, 90°, 150°, 210°, 270°, 330°, and the corresponding pitch angles of 20°, -10°, 20°, -10°, 20°, -10°), and other rendering parameters can be set in the rendering engine or software to simulate the lighting and shadow effects in the real world, so that rasterization processing is respectively performed based on each camera view angle (or photographing and screenshot are performed in the rendering engine or software), to obtain the view angle image samples of the plurality of view angles of the three-dimensional model samples. Here, the rasterization processing is part of the rendering process, which is used to convert the three-dimensional model samples into two-dimensional view angle image samples. The specific details of the rasterization processing can be referred to the description in step 2012 above, and will not be described here.

[0192] Taking sampling of six azimuth angles as an example, referring to FIG. 13, which is a schematic diagram of sampling of a plurality of view angles provided by an embodiment of the present application, FIG. 13 shows the sampling view angles of the view angle image samples of the plurality of view angles (such as two camera view angles corresponding to the azimuth angle 1, two camera view angles corresponding to the azimuth angle 2, etc.) of different camera positions.

[0193] In step 3013, the three-dimensional model samples are sampled according to the view angle parameters and the lighting parameters of the single view angle that are set in advance, to obtain single image samples of the three-dimensional model samples.

[0194] In some embodiments, referring to the sampling manner in step 3012, the three-dimensional model samples are sampled according to the view angle parameters and the lighting parameters of the single view angle that are set in advance (only the view angle parameter is changed, that is, the view angle parameters of the single view angle are different from the view angle parameters of the plurality of view angles), to obtain the single image samples of the three-dimensional model samples.

[0195] In step 3014, the perspective image samples and the single image sample of each three-dimensional model sample corresponding to multiple perspectives are combined into a second data set.

[0196] In some embodiments, each three-dimensional model sample corresponds to perspective image samples of multiple perspectives and a single image sample.

[0197] For example, referring to FIG. 14, which is a schematic diagram of a data structure of the second data set provided by the embodiments of the present application, for a three-dimensional model sample named “800015_1_R30”, the corresponding single image sample is “000.png”, and the perspective image samples corresponding to multiple perspectives are “001.png-006.png”.

[0198] In some embodiments, different initial rotation angles can be set for the same three-dimensional model sample to sample multiple groups of pictures (including single image samples and perspective image samples of multiple perspectives) based on one three-dimensional model sample, so as to improve the compatibility of the multi-view generation model for different perspectives.

[0199] In some embodiments, the perspective parameters of multiple perspectives include camera distance parameters. Referring to FIG. 31, after obtaining the perspective image samples of multiple perspectives of the three-dimensional model sample, the following steps 3015 to 3016 can also be performed, which are described in detail below.

[0200] In step 3015, edge detection is performed on each perspective image sample to obtain an edge detection result.

[0201] In some embodiments, edge detection can be achieved in the following way: first, load each perspective image sample through an image processing library (such as OpenCV, Pillow, etc.); next, perform image preprocessing, such as size standardization (e.g., uniform adjustment to 256x256 pixels), grayscale conversion, noise reduction, etc.; next, perform edge detection through a pre-set edge detection algorithm, such as Canny edge detection (including four steps of noise removal, gradient calculation, non-maximum suppression, and hysteresis thresholding), Sobel operator (detecting edges based on the size of the gradient), Laplacian operator (detecting edges based on the zero crossing of the second derivative), etc.; after obtaining the edge detection result of the image sample, the edge detection result can also be refined, such as removing small isolated points (e.g., when the number of connected pixels contained in an isolated point is less than a pre-set threshold (e.g., 3 pixels), it is determined to be a “small isolated point” and removed, and here, an isolated point refers to a pixel point or a collection of pixel points in the edge detection result that has no any connected relationship with the surrounding edge contour and exists independently) or filling in the gaps in the edge, etc.; finally, connect the edges to form a closed contour to obtain the final edge detection result.

[0202] For example, edge detection is performed on the view angle image sample by Canny edge detection to obtain an edge detection result, which can be achieved in the following manner: grayscale processing is performed on the pre-processed view angle image sample to obtain a grayscale image sample; gradient calculation is performed on each pixel point in the grayscale image sample to obtain a gradient value of each pixel point (for example, the Sobel operator is used to calculate the gradient value of the pixel point); in response to the gradient value of the pixel point being greater than a first gradient threshold, the pixel point is regarded as a first edge point; in response to the gradient value of the pixel point being greater than or equal to a second gradient threshold and less than the first gradient threshold, and the pixel point being adjacent to the first edge point, the pixel point is regarded as a second edge point, wherein the first gradient threshold is greater than the second gradient threshold; and the edge detection result is determined by the first edge point and the second edge point.

[0203] In step 3016, in response to the edge detection result representing that the edge of the model object in the view angle image sample exceeds the boundary of the view angle image sample, the camera distance parameter is increased by a preset parameter, and the three-dimensional model sample is sampled according to the increased camera distance parameter to obtain view angle image samples of multiple view angles of a new three-dimensional model sample.

[0204] In some embodiments, when the edge of the model object in the edge detection result exceeds the boundary of the view angle image sample, it means that the current camera distance parameter is too small to completely capture the model object (i.e., the model object is not fully displayed within the viewfinder range of the camera), at which time the camera distance is increased according to a preset parameter, which can be a fixed value or calculated according to a certain rule (such as a certain proportion of the current camera distance), and the camera settings are updated using the new camera distance parameter. For example, the position and / or focal length of the camera are adjusted, the three-dimensional model sample is sampled according to the updated camera settings, and new view angle image samples are generated, for example, the photographing and screenshot are re-executed according to the new camera settings, until the model object in all image samples is inside the image sample, ensuring that the new image samples contain the complete edge of the model object.

[0205] Continuing to refer to FIG. 3G, in step 302, feature extraction processing is performed on the single-image sample to obtain an image feature sample.

[0206] In some embodiments, the single-image sample can be subjected to feature extraction processing by a pre-trained convolutional neural network (such as ResNet, Inception, etc.) to obtain an image feature, in which the convolutional neural network usually has multiple fully connected layers or multiple global average pooling layers, thereby generating a final fixed-size feature vector at the end of the convolutional neural network, which is the image feature. The image feature contains high-level semantic information of the single-image sample.

[0207] In step 303, multi-view generation processing is performed based on the image feature sample to obtain view prediction images of multiple view angles.

[0208] In some embodiments, step 303 can be implemented by performing the following processing for each view: fusing the image feature sample and the view information (for example, a pre-set camera pose including a rotation matrix R and a translation matrix T of the camera, the camera pose determines the camera view) to obtain conditional information features, performing denoising processing on the noise map based on the conditional information features to obtain a view prediction image corresponding to the view, for example, performing attention encoding processing (for example, Cross Attention, etc.) based on the conditional information features and the noise map to obtain attention encoding features, and a denoising network (for example, a U-Net, etc.) performs noise prediction based on the attention encoding features, and performs denoising processing on the noise map according to the predicted noise to obtain a view prediction image corresponding to the view information.

[0209] Here, the noise map can be obtained by randomly generating noise (the size of the noise map is the same as the size of the image feature sample), and can also be obtained by forward diffusion based on the image feature sample of the single image sample. The embodiments of the present application are not limited.

[0210] For example, for each view, R and T (view information) can be spliced with the image feature sample to obtain conditional information features, the conditional information features are taken as a query, and the noise map is taken as a key and a value for attention encoding to obtain attention encoding features, and a denoising network performs noise prediction according to the attention encoding features, for example, fuses (splices) the attention encoding features and the noise map, maps the fused features by a mapping function (for example, Softmax, etc.) to obtain predicted noise, removes the predicted noise from the noise map to obtain a view prediction image. The denoising processing can include t time steps, and noise prediction and reduction of predicted noise from the noise map are performed at each time step to increase the clarity of the image, that is, the denoising network updates the noise map at each time step.

[0211] For example, referring to FIG. 22, FIG. 22 is a network structure schematic diagram of a multi-view generation model provided by an embodiment of the present application. After the single image sample is image-encoded by an encoder, an image feature sample is obtained, a Gaussian noise matrix (that is, a random noise map) is superimposed on the image feature sample to obtain a diffusion starting point Z0, and forward diffusion is performed based on the diffusion starting point Z0 (that is, a noise map Z T of pure noise is obtained by superimposing T time steps) to obtain a noise map Z T Z TAs an input of the denoising network, the image feature sample and the view information are injected into each stage of the denoising network as conditional information. For example, after the image feature sample and the view information are fused, they are respectively fused with the features of each stage of the denoising network (such as a plurality of stages corresponding to upsampling and a plurality of stages corresponding to downsampling in FIG. 22) to guide the denoising network to perform noise prediction. The denoising process (i.e., the reverse diffusion process of the diffusion model) of the denoising network is implemented through multiple iterations. The output (the feature representation) of each iteration is used as the input of the next iteration until a preset condition (such as the number of iterations) is met. The feature output by the last iteration (i.e., the feature representation Z 0′ in FIG. 22) is decoded by the decoder to obtain the view prediction image.

[0212] Here, the conditional information is additional input data used to guide the denoising process and control the image content. The core goal of the conditional information is to incorporate specific constraints or guidance information in the denoising process to obtain the expected view prediction image.

[0213] In step 304, a second loss value of the view prediction images of the plurality of views and the view image samples of the plurality of views is obtained through a pre-set loss function.

[0214] In some embodiments, the second loss value of the view prediction images of the plurality of views and the view image samples of the plurality of views is obtained through a pre-set loss function (such as a mean square error loss function). The embodiments of the present application do not limit the specific loss function used.

[0215] For example, the second loss value of the view prediction images of the plurality of views and the view image samples of the plurality of views can be represented by formula (2):

[0216] where L2 represents the second loss value; M is the total number of views; N is the total number of pixel points of a single view image sample or a view prediction image; represents the pixel value of the i-th pixel point in the m-th view prediction image of the view; I m,i represents the pixel value of the i-th pixel point in the m-th view image sample of the view.

[0217] In step 305, the network parameters of the original multi-view generation model are updated through the second loss value to obtain a trained multi-view generation model.

[0218] In some embodiments, the gradient of the parameters of the original multi-view generation model is obtained through the second loss value to obtain a trained multi-view generation model.

[0219] For example, the second loss value is obtained by a back propagation algorithm to obtain gradient information of each parameter of the original multi-view generation model, the obtained gradient information is used to update the parameters of the original multi-view generation model according to a gradient descent optimization algorithm (for example, batch gradient descent, stochastic gradient descent, etc.), the above process is repeated until a certain number of iterations is reached or the original multi-view generation model converges, thereby obtaining the trained multi-view generation model.

[0220] By constructing the second data set for the multi-view generation task, the original multi-view generation model is incrementally trained based on the second data set, so that the multi-view generation model learns the model features of multiple perspectives in the second data set, thereby improving the generation effect of the perspective images of multiple perspectives, making the perspective images of multiple perspectives more consistent in perspective, and better guiding the generation of the three-dimensional model.

[0221] Referring to FIGS. 15A, 15B, 15C and 15D, FIG. 15A is a first schematic diagram of the effect of generating perspective images of multiple perspectives provided by an embodiment of the present application. As can be seen from FIGS. 15A to 15D, in the images generated by the original multi-view generation model, the modeling details (such as armrest shape, backrest angle, body structure) of the same object (such as a chair, a blender, etc.) are quite different under different perspectives, as if the object is “deformed” (for example, in the blender image generated by the original multi-view generation model, there are two blender handles displayed in one image, as shown in FIG. 15D), the perspective consistency is poor, and each view lacks unified presentation of the same object (for example, in the multi-view double chair image generated by the original multi-view generation model, there is a phenomenon that the connecting structure between the double chairs is missing in the images corresponding to multiple perspectives, as shown in FIG. 15C); while the trained multi-view generation model, the modeling details of the same object under different perspectives are highly unified, such as chair armrest, chair backrest, blender body, etc. The structure is stable in each perspective, consistent with the perspective logic of the same object in actual observation, the perspective consistency is greatly improved, and the reasonable visual effect of the same object under multiple perspectives can be clearly presented, that is, the perspective consistency of the multi-view generation model trained based on the second data set is improved.

[0222] Specifically, referring to FIGS. 15A-15D, the images generated by the original multi-view generation model have large differences in modeling details (such as armrest shape, backrest angle, and body structure) of the same object (such as a chair, a blender, etc.) under different viewing angles, as if the object is “transforming”, (for example, the image of the double chair generated by the original multi-view generation model in FIG. 15C lacks the connection between the double chairs), the viewing angle consistency is poor, and each view lacks a unified presentation of the same object (for example, two blender handles appear under the same viewing angle in the image of the blender generated by the original multi-view generation model in FIG. 15D); while the trained multi-view generation model has highly unified modeling details of the same object under different viewing angles, such as stable structures of chair armrests, chair backrests, and blender bodies under different viewing angles, which conforms to the viewing angle logic of the same object in actual observation, greatly improves the viewing angle consistency, and can clearly show the reasonable visual effect of the same object under multiple viewing angles.

[0223] Referring back to FIG. 3F, in step 1022, a multi-view generation processing is performed based on the model image features by the pre-trained multi-view generation model to obtain viewing angle images of multiple viewing angles.

[0224] In some embodiments, the multi-view generation processing based on the model image features is performed by forward inference of the multi-view generation model (see the description of steps 302-303).

[0225] For example, the multi-view generation model can be a generation model for generating multiple views from a single image based on a diffusion model (such as Stable Diffusion, etc.), including but not limited to zero123, zero123++, etc., and the embodiments of the present application do not limit the specific multi-view generation model.

[0226] Through steps 1021-1022, multiple viewing angle images with consistent viewing angles and complete details can be efficiently generated based on two-dimensional model materials. Specifically, the model image features are obtained by using the multi-view generation model trained by the second data set to extract features from the model materials; and then the multi-view generation processing is performed based on the model image features, and the multi-view generation model learns the correlation rules of different viewing angles to generate multiple viewing angle images that are consistent in style with the original model materials and have coherent viewing angles. This not only reduces the cost of obtaining multiple viewing angle images without relying on three-dimensional models for direct sampling, but also ensures that the generated viewing angle images are highly matched in feature details, avoiding the fragmentation of model features caused by viewing angle deviation, and providing reliable and consistent visual basis for subsequent construction of high-quality three-dimensional models based on multiple viewing angle images.

[0227] Referring back to FIG. 3A, in step 103, a first three-dimensional model is generated based on the viewing angle images of multiple viewing angles, wherein the first three-dimensional model includes a first color parameter.

[0228] In some embodiments, referring to FIG. 3J, the step 103 shown in FIG. 3A can be implemented by the following steps 1031 to 1034, which are explained in detail as follows.

[0229] In step 1031, feature extraction processing is performed on the view image of each view to obtain the image feature of each view.

[0230] In some embodiments, the feature extraction processing on the view image of each view can be performed by a pre-trained convolutional neural network (such as ResNet, Inception, etc.), in which there are usually multiple fully connected layers or multiple global average pooling layers, so that a final fixed-size feature vector is generated at the end of the convolutional neural network, which is the image feature. The image feature contains high-level semantic information of the view image.

[0231] In step 1032, feature encoding processing is performed on the image feature of each view to obtain the encoded feature of each view.

[0232] In some embodiments, the feature encoding processing on the image feature of each view can be implemented by the following way: dividing the image feature into multiple image feature blocks, arranging the multiple image feature blocks into an image feature block sequence; and performing encoding processing on the image feature block sequence to obtain the encoded feature.

[0233] For example, assuming that the feature map size of the image feature is 224x224, the image feature is divided into blocks (Patches) with a size of 16x16, and after division, (224 / 16) 2 = 196 Patches, and the 196 Patches are arranged in the order from the top left to the bottom right in the view image to obtain the image feature block sequence.

[0234] In some embodiments, the encoding processing on the image feature block sequence to obtain the encoded feature can be implemented by the following way: performing embedding encoding processing on the image feature block sequence to obtain embedding features; performing normalization processing on the embedding features to obtain normalized features; performing attention encoding processing on the normalized features to obtain attention encoding features; fusing (e.g., splicing) the attention encoding features and the embedding features to obtain fused features; performing feedforward mapping processing on the normalized features to obtain mapping features; and fusing the mapping features and the fused features to obtain the encoded feature.

[0235] In an example, the image feature block sequence can be embedded and encoded by an embedding layer to obtain embedded features. For example, the image feature block sequence can be convoluted by a convolution operation in the embedding layer, and the convoluted image feature block sequence can be position embedded to obtain the embedded features.

[0236] In an example, the normalization processing can be implemented by obtaining the mean and variance of the embedded features in each feature dimension, then normalizing the feature data in each feature dimension, for example, subtracting the mean and dividing by the square root of the variance, and finally further scaling and translating the normalized embedded features by a learnable scaling factor and a learnable offset factor to obtain normalized features.

[0237] In an example, the normalized features can be attention encoded by a multi-head attention mechanism to obtain attention encoded features. For example, for the input normalized features, first perform linear transformation to generate three matrices of query vector (Query, Q), key vector (Key, K) and value vector (Value, V). Linear transformation is implemented by learnable weight matrices W Q , W K and W V , then for each attention head, calculate the dot product of Q and K to obtain attention scores, then normalize the attention scores, for example, apply a normalization function (such as a softmax function) to make the attention scores into a probability distribution, use the attention probability distribution to weight V to generate a new feature representation, finally concatenate the outputs of all attention heads and perform linear transformation to obtain a rich representation containing the correlation of different positions in the normalized features, i.e. the attention encoded features.

[0238] The normalized features can be mapped by a feed-forward neural network layer (FFN) to obtain mapped features. Taking a multilayer perceptron (MLP) structure as an example, the normalized features are first linearly transformed by a first linear layer, which can be a fully connected layer with a weight matrix W1 and a bias vector b1. Next, the linearly transformed normalized features are non-linearly transformed by an activation function (e.g., ReLU, Sigmoid, etc.) to obtain non-linearly transformed features. Finally, the non-linearly transformed features are linearly transformed by a second linear layer to obtain the mapped features, which can be a fully connected layer with a weight matrix W2 and a bias vector b2.

[0239] The mapped features and the fused features are spliced to obtain encoded features.

[0240] In step 1033, the encoded features of each view are decoded to obtain three-dimensional features.

[0241] In some embodiments, the feature maps of the encoded features of each view are spliced into a three-dimensional structure in a specific manner, for example, the feature maps of each view are stacked along an axis (e.g., a depth axis). During the stacking process, the feature maps can be transformed, such as rotated or flipped, to ensure that they are spatially aligned. Through the above splicing operation, three-dimensional features, also known as triplane features, are generated.

[0242] The Triplane includes three orthogonal planes, each representing the features of a three-dimensional model observed from different views. The horizontal plane represents the features of a top view, the vertical plane represents the features of a side view, and the diagonal plane represents the features of an oblique view.

[0243] In step 1034, three-dimensional reconstruction is performed based on the three-dimensional features to obtain a first three-dimensional model.

[0244] In some embodiments, the three-dimensional reconstruction processing can be implemented by generating sampling points in feature maps on each plane of a three-dimensional feature (Triplane), which can be key points, edge points on the feature map, or points selected according to a certain strategy (such as uniform sampling, importance sampling, etc.), then fusing the sampling points on different planes, for example, projecting the sampling points into a common space, and then using a neural network or other method to combine the features, then using the fused sampling point features to construct a cube (Cube), which can be a fixed-size voxel grid, or a dynamically-sized point cloud or other structure, expanding the single cube into a set of FlexiCubes composed of multiple cubes, which can cover the entire Triplane space, and each cube represents the three-dimensional features of a local area, then fusing and optimizing the cubes in the FlexiCubes to improve the continuity and accuracy of the three-dimensional features, such as spatial filtering, feature smoothing, and denoising steps, and finally using the optimized FlexiCubes for three-dimensional reconstruction, such as converting the cubes to voxel grids, point clouds, or other three-dimensional representations, then visualizing or further processing to obtain the first three-dimensional model.

[0245] In other embodiments, a pre-trained three-dimensional model generation model can be used to generate the first three-dimensional model based on the perspective images of multiple perspectives, such as an InstantMesh or the like, and the present application does not limit the specific three-dimensional model generation model used.

[0246] Through steps 1031 to 1034, first, feature extraction is performed on each perspective image to capture its high-level semantic information; then, through feature encoding processing, the image features are converted into encoded features containing spatial correlations, strengthening the internal relationship of different perspective features; then, based on the encoded features, three-dimensional features (such as Triplane features) are decoded to realize the mapping of two-dimensional features to three-dimensional space; finally, through three-dimensional reconstruction processing, the three-dimensional features are converted into a structurally complete first three-dimensional model. The complementary information of multiple perspective images is fully utilized, and the limitations of a single perspective are avoided through feature fusion and spatial alignment, and the original image details such as color can be preserved during the reconstruction process, so that the generated first three-dimensional model is guaranteed in terms of structural accuracy, spatial coherence, and color richness, laying a high-quality foundation model for subsequent optimization of the first color parameter.

[0247] Referring back to FIG. 3A, in step 104, the coloring parameters of each perspective image are obtained.

[0248] In some embodiments, the view image includes a model object, referring to FIG. 3K, the step 104 shown in FIG. 3A can be implemented by performing steps 1041 to 1043 for each view image, which are described in detail below.

[0249] In step 1041, a bounding box of the model object in the view image is obtained.

[0250] In some embodiments, the bounding box of the model object in the view image is obtained, that is, the smallest square box surrounding the model object is obtained.

[0251] In some embodiments, the view image is pre-processed, such as at least one of grayscale and denoising (which can be implemented by filtering, etc.), to improve the accuracy of subsequent processing. Next, object detection is performed, such as using an object detection algorithm such as Histogram of Oriented Gradients (HOG), Single Shot MultiBox Detector (SSD), etc. to detect the model object from the image, to obtain the position and size of the model object. Through the object detection algorithm, the coordinates of all pixel points of the model object are obtained, and the minimum x and y values (top left corner) and maximum x and y values (bottom right corner) in all pixel point coordinates of the model object are found, thereby defining a minimum rectangular box that can enclose all pixel points of the model object.

[0252] In step 1042, the bounding box is taken as the bottom surface, a first vertex is determined based on a preset distance parameter, and a quadrangular pyramid is constructed based on the bottom surface and the first vertex.

[0253] In some embodiments, the bounding box is taken as the bottom surface, a first vertex (a vertex in the quadrangular pyramid located at the tip (i.e., not located on the bottom surface)) is determined based on a preset distance parameter, and a quadrangular pyramid is constructed based on the bottom surface and the first vertex.

[0254] For example, referring to FIG. 16A, which is a first schematic diagram of the principle of obtaining the shading parameter of each view image according to an embodiment of the present application, as shown in FIG. 16A, the bounding box is the smallest rectangular box that can enclose all pixel points of the model object (a sphere shown in FIG. 16A). The bounding box is taken as the bottom surface, a first vertex is determined based on a preset distance parameter, and a quadrangular pyramid is constructed.

[0255] Here, the preset distance parameter can be determined by the camera pose in the view information set in the multi-view generation model when the view images of multiple views are generated, that is, the position of the model object is converted from the world coordinate system to the camera coordinate system, thereby obtaining the distance parameter between the camera (the first vertex) and the model object.

[0256] In step 1043, the shading parameters of the perspective image are determined by the positional relationship between the face of the quadrangular pyramid and the first three-dimensional model.

[0257] In some embodiments, the shading parameters include a viewpoint coordinate, and each perspective image corresponds to a camera viewpoint, as shown in FIG. 3L, the step 1043 shown in FIG. 3K can be implemented by the following steps 10431 to 10432, which are described in detail below.

[0258] In step 10431, the face of the quadrangular pyramid is moved in the object coordinate system in which the first three-dimensional model is located along the direction of the camera viewpoint corresponding to the perspective image until the face of the quadrangular pyramid is tangent to the boundary of the first three-dimensional model.

[0259] In some embodiments, each camera viewpoint corresponds to a quadrangular pyramid, and the face of the quadrangular pyramid is moved in the direction of the camera viewpoint corresponding to the perspective image until the normal vector of the face of the quadrangular pyramid is parallel to the normal vector of the boundary of the first three-dimensional model, at which time the face of the quadrangular pyramid is tangent to the boundary of the first three-dimensional model.

[0260] For example, the boundary of the first three-dimensional model is obtained (for example, by a rectangular bounding box, and the obtaining of the bounding box can be referred to the description of the bounding box in step 1041 above), the face of the quadrangular pyramid is moved, and the normal vector of the face of the quadrangular pyramid is compared with the normal vector of the model boundary. If the included angle between the two is close to 0 degrees (the included angle is less than or equal to a first preset angle threshold, such as 3 degrees) or close to 180 degrees (the included angle is greater than or equal to a second preset angle threshold, such as 177 degrees) (i.e., they are almost parallel or anti-parallel), it can be considered that the face of the quadrangular pyramid is tangent to the boundary of the first three-dimensional model. Referring to FIG. 16B, which is a second schematic diagram of the principle of obtaining the shading parameters of each perspective image according to an embodiment of the present application, as shown in FIG. 16B, the quadrangular pyramid is moved until the four faces are tangent to the boundary of the first three-dimensional model, and it can also be understood that the quadrangular pyramid is moved (translated) so that the quadrangular pyramid just “sticks” to the first three-dimensional model.

[0261] In step 10432, the intersection of the face of the quadrangular pyramid tangent to the boundary of the first three-dimensional model is determined, and the coordinates of the intersection in the object coordinate system are taken as the viewpoint coordinates.

[0262] In some embodiments, the intersection of the face of the quadrangular pyramid tangent to the boundary of the first three-dimensional model is determined, for example, by least squares solution, the intersection of the four faces shown in FIG. 16B is obtained, and the coordinates of the intersection in the object coordinate system are taken as the viewpoint coordinates. For example, referring to FIG. 16C, which is a third schematic diagram of the principle of obtaining the shading parameters of each perspective image according to an embodiment of the present application, the positional relationship between the viewpoint coordinates and the first three-dimensional model is shown in FIG. 16C.

[0263] Through steps 1041 to 1043, the bounding box of the model object is first obtained to provide an accurate bottom surface reference for subsequent spatial geometry construction, and the combination of preprocessing and object detection algorithm ensures the accuracy of the bounding box; then the bounding box is taken as the bottom surface, a four-pyramid is constructed by combining a preset distance parameter, the view information such as the camera pose is integrated into the geometric structure, and a correlation bridge between the two-dimensional image and the three-dimensional space is established; finally, the coloring parameters of the perspective image are determined through the positional relationship between the surface of the four-pyramid and the first three-dimensional model, and the spatial alignment of the coloring parameters and the three-dimensional model is realized. Through the adaptation of the four-pyramid and the first three-dimensional model, the obtained coloring parameters can accurately reflect the color mapping relationship under each perspective.

[0264] Continuing to refer to FIG. 3A, in step 105, the first color parameter of the first three-dimensional model is updated based on the coloring parameter of each perspective image, to obtain a second three-dimensional model.

[0265] In some embodiments, the first three-dimensional model includes a plurality of second vertices, and the first color parameter of the first three-dimensional model can be set to zero (i.e., the first color parameter of each second vertex of the first three-dimensional model is removed) before the first color parameter of the first three-dimensional model is updated.

[0266] In other embodiments, the first three-dimensional model can not include the first color parameter, for example, the first three-dimensional model is a point cloud model and does not include the first color parameter.

[0267] In some embodiments, the coloring parameter further includes a second color parameter of each pixel point of the model object in the perspective image, and the first three-dimensional model includes a plurality of second vertices, each second vertex corresponding to a first color parameter. Referring to FIG. 3M, step 105 shown in FIG. 3A can be implemented by performing the following steps 1051 to 1052 on each pixel point of the model object in the perspective image, which will be described in detail below.

[0268] In step 1051, a ray with the viewpoint coordinate as the end point and passing through the pixel point is generated in the object coordinate system, and a second vertex on the first three-dimensional model that intersects with the ray is determined as a coloring point.

[0269] In some embodiments, a ray is formed with the viewpoint coordinate as the end point and the pixel point of the model object in the bounding box, and is extended to the first three-dimensional model, and the second vertex on the first three-dimensional model that intersects with the ray is taken as the coloring point.

[0270] In step 1052, the first color parameter of the coloring point in the first three-dimensional model is replaced with the second color parameter of the pixel point, to obtain a second three-dimensional model.

[0271] In some embodiments, the second color parameter of the pixel point is mapped to the coloring point, to obtain a second three-dimensional model.

[0272] For example, referring to FIG. 16D, which is a fourth schematic diagram of a principle of obtaining the shading parameter of each view image according to an embodiment of the present application, as shown in FIG. 16D, the pixel value of the model object in the bounding box of one view image is mapped to the first three-dimensional model. Here, only the shading process of one view is shown for the convenience of understanding.

[0273] In some embodiments, referring to FIG. 3N, the step 105 shown in FIG. 3A can be implemented by performing the following steps 1053 to 1054 on each second vertex in the first three-dimensional model, which are explained in detail as follows.

[0274] In the step 1053, in response to the second vertex being in at least two camera views, the second color parameters of at least two pixel points corresponding to the second vertex in the two camera views are weighted and averaged to obtain a third color parameter.

[0275] In some embodiments, when the second vertex of the first three-dimensional model corresponds to the second color parameters mapped by multiple view images, the weight value is determined according to the included angle between the ray formed by the viewpoint coordinate and the center of the bounding box and the normal of the second vertex (the normal is perpendicular to the tangent plane of the surface of the first three-dimensional model), so as to perform weighted average processing to obtain the third color parameter.

[0276] For example, assuming that the second vertex corresponds to the second color parameters (denoted as A and B) of the pixel points from two view images, the included angles between the rays formed by the viewpoint coordinates corresponding to the two view images and the center of the bounding box and the normal of the second vertex are 60° and 30° respectively, then the third color parameter (C) can be represented as C = cos60°A + cos30°B, wherein cos represents the cosine function.

[0277] In the step 1054, the first color parameter of the second vertex in the first three-dimensional model is replaced by the third color parameter to obtain a second three-dimensional model.

[0278] In some embodiments, referring to FIG. 3O, the step 105 shown in FIG. 3A can be implemented by performing the following steps 1055 to 1057 on each second vertex in the first three-dimensional model, which are explained in detail as follows.

[0279] In the step 1055, in response to the second vertex not being in any camera view, at least two neighbor vertices adjacent to the second vertex within a preset range are obtained.

[0280] In some embodiments, when the second vertex does not correspond to the second color parameter mapped by any view image, at least two neighbor vertices adjacent to the second vertex within a preset range are obtained.

[0281] For example, the first three-dimensional model includes coordinate information of each second vertex, a preset range can be determined by a preset radius value, at least two neighbor vertices in the preset range are obtained according to the coordinate information of each second vertex, or at least two neighbor vertices closest to each second vertex are selected, and the specific obtaining manner of the neighbor vertices is not limited in the embodiments of the present application.

[0282] For example, the preset radius value can be adjusted based on the geometric complexity of the first three-dimensional model. For example, in a vertex dense or complex structure region (for example, the average distance between vertices is less than a preset distance threshold), the radius value is 2 times the distance threshold; in a vertex sparse or flat structure region (for example, the average distance between vertices is greater than or equal to the preset distance threshold), the radius value is 4 times the preset distance threshold. Thus, the problem of insufficient neighbor vertices in the sparse region is solved, and the interpolation calculation of the second color parameter is more in line with the local features of the first three-dimensional model, thereby avoiding the influence of too many neighbor vertices on the interpolation accuracy in the dense region.

[0283] In step 1056, an average value of the second color parameters corresponding to the at least two neighbor vertices is obtained.

[0284] Here, the second color parameters corresponding to the neighbor vertices can be obtained through steps 1051 to 1052, or can be obtained through steps 1053 to 1054, and the specific values of the second color parameters corresponding to the neighbor vertices are different according to the number of camera perspectives at each neighbor vertex.

[0285] In step 1057, the first color parameter of the second vertex in the first three-dimensional model is replaced by the average value, and a second three-dimensional model is obtained.

[0286] Through steps 1051 to 1057, the color parameters of the first three-dimensional model can be accurately and comprehensively updated, and a high-quality second three-dimensional model is obtained. Specifically, first, the ray intersection is determined to determine the coloring point, and the second color parameters of the pixels of the perspective image are directly mapped to the corresponding vertices of the first three-dimensional model, so that the accurate spatial alignment of the two-dimensional image color to the three-dimensional model is realized; second, for the vertices in the multi-perspective range, the color information of different perspectives is fused based on the weighted average processing of the ray and the normal angle, so that the color conflict caused by the single perspective deviation is avoided, and the color transition is more natural; and finally, for the vertices not covered by any perspective, the average value of the colors of the neighbor vertices is used for filling, so that the integrity and continuity of the model color are ensured. This hierarchical processing method not only ensures the accurate restoration of the color of the visible region, but also solves the problem of the perspective blind area through multi-perspective fusion and interpolation completion, so that the color parameters of the second three-dimensional model are significantly improved in accuracy, consistency and integrity, and finally a visual effect closer to the real scene and richer in details is presented.

[0287] The updating of the first color parameter of the first three-dimensional model through step 104 and steps 1051 to 1057 realizes the optimization of the color parameter of the three-dimensional model, so that the color details of the second three-dimensional model are further enhanced, thereby achieving the beneficial effect of improving the quality of the generated three-dimensional model.

[0288] For example, referring to FIGS. 17A, 17B, 17C and 17D, FIG. 17A is a first comparison diagram of a first three-dimensional model and a second three-dimensional model provided by an embodiment of the present application. As can be seen from FIGS. 17A to 17D, after updating the first color parameter of the first three-dimensional model (such as FIG. 17(1), FIG. 17B(1), FIG. 17C(1) and FIG. 17D(1)), the color details of the second three-dimensional model (such as FIG. 17(2), FIG. 17B(2), FIG. 17C(2) and FIG. 17D(2)) are enhanced, and the texture details of the three-dimensional model can be better represented by color (since FIGS. 17A to 17D are gray-scale diagrams, different gray levels in the diagrams can be regarded as different colors).

[0289] For example, the colors (gray levels) of FIG. 17(1), FIG. 17B(1), FIG. 17C(1) and FIG. 17D(1) are relatively “fuzzy and dull”, and the texture details (such as the string button / panel texture of the guitar, the book arrangement of the bookcase, the throw pillow texture of the sofa, and the carving / drawer structure of the desk) are presented in a relatively hazy and weakly layered manner due to the color parameter limitation; the texture details of FIG. 17(2), FIG. 17B(2), FIG. 17C(2) and FIG. 17D(2) are significantly improved, such as the clearer strings of the guitar, the “color (different gray levels) layers” of the books in the bookcase making the bookcase display details more easily recognizable, the throw pillow texture of the sofa being more prominent, and the carving and drawer handle structure of the desk restoring the texture information that the three-dimensional model should have.

[0290] In some embodiments, referring to FIG. 3P, after obtaining the second three-dimensional model, the following steps 106 to 108 can also be performed, which are specifically described as follows.

[0291] In step 106, the second three-dimensional model is subjected to texture optimization to obtain a model texture.

[0292] In some embodiments, the second three-dimensional model is subjected to texture detection processing to obtain a texture detection result, and the model texture is obtained based on the texture detection result.

[0293] For example, in response to the map detection result indicating that the second three-dimensional model does not include a map (without material information), a model map can be generated for the second three-dimensional model by a three-dimensional model rendering tool (such as Blender), for example, by selecting or creating a two-dimensional texture image that will serve as the surface pattern of the model through the three-dimensional model rendering tool, editing the texture image, such as adjusting the size, color correction, applying filters, etc. to meet specific visual effects, and then unfolding the surface of the second three-dimensional model into a two-dimensional plane. This unfolding process can be achieved through UV mapping, which defines how points on the surface of the model correspond to points on the texture image. Adjust the UV coordinates in the UV editor to ensure that the distribution of the texture on the model is uniform and reasonable. Apply the texture image to the UV-unfolded model. Use texture refinement techniques such as seamless texture, texture stitching, and detail mapping to enhance the realism and detail of the texture. Adjust the lighting and shadow effects of the texture according to the lighting model of the model to simulate the lighting conditions in the real world. Set material properties for the model, such as diffuse reflection, glossiness, transparency, etc. Test the texture effect in the rendering view of the three-dimensional model rendering tool, adjust the parameters as needed, and generate a map suitable for the second three-dimensional model to make it have higher realism and visual appeal when rendered.

[0294] For example, in response to the map detection result indicating that the second three-dimensional model includes a map, but the map is not suitable for secondary editing, for example, when the map of the second three-dimensional model cannot be directly edited and used due to the three-dimensional model rendering tool not supporting a specific map format, the map file being damaged due to encoding problems during saving or transmission, etc., the map is regenerated, and the geometric features, materials, and lighting information of the second three-dimensional model are saved as a texture file. Referring to the description above, a model map can be generated for the second three-dimensional model by a three-dimensional model rendering tool.

[0295] Referring to FIG. 18, which is a schematic diagram of the effect of map optimization provided by an embodiment of the present application, as can be seen from FIG. 18, the original map (left) presents a state of being relatively blurred, irregular, and with chaotic details, making it difficult to clearly identify the model structure; while the model map after map optimization (right), through the regular grid lines, clearly outlines the contours, morphology, and structural relationship of the model, not only making the model structure intuitive and identifiable, but also due to this ordered processing, the model map obtained after map optimization can be easily edited, and these effects can be quickly applied in subsequent rendering or real-time display without the need for re-computation, thereby achieving the beneficial effect of improving the rendering efficiency of the three-dimensional model.

[0296] In step 107, the second three-dimensional model is topologically optimized to obtain a third three-dimensional model.

[0297] In some embodiments, the second three-dimensional model can be topologically optimized by a retopology process to obtain a third three-dimensional model.

[0298] For example, the retopology process involves reconstructing the surface geometry of a complex or high-resolution three-dimensional model to create a version with fewer polygons while maintaining the visual appearance of the original model. The second three-dimensional model can be topologically optimized by retopology tools such as QuadRemesh, QuadriFlow, etc. to obtain a third three-dimensional model. Referring to FIG. 19, which is a schematic diagram of the effect of topological optimization according to an embodiment of the present application, it can be seen from FIG. 19 that the second three-dimensional model before optimization (left) has too many faces, and rendering requires a large amount of computing power to process complex faces. The third three-dimensional model after retopology optimization (right) clearly presents the model structure, retains the original visual appearance, greatly simplifies the topology, reduces the rendering computing power requirement, and is more suitable for efficient application on the production side, i.e., the third three-dimensional model has a simple model structure, which can reduce the computing power required for model rendering and is more conducive to application on the production side.

[0299] For example, when the second three-dimensional model cannot be directly topologically optimized (retopology) due to too many faces, the face reduction process can be performed by voxel reconstruction and other optimization methods. After reducing the number of faces of the second three-dimensional model, the topological optimization process is performed. For example, referring to FIG. 20, which is a schematic diagram of the effect of face reduction processing according to an embodiment of the present application, it can be seen from FIG. 20 that the three-dimensional model before face reduction processing has as many as 20W faces, and requires a large amount of computing power and other resources due to the excessive number of faces. After face reduction optimization by voxel reconstruction and other methods, the number of faces is reduced to 4W, and the three-dimensional model presents a more simple structure while maintaining the basic form. This face reduction process effectively reduces the complexity of the model, not only solves the problem of direct topological optimization due to too many faces, but also reduces the computing power burden for subsequent topological optimization and production side application (such as rendering, real-time display, etc.), improves processing efficiency, and makes the three-dimensional model more suitable for actual production processes.

[0300] For example, when the second three-dimensional model cannot be directly topologically optimized due to abnormal structure of part of the model (for example, topological failure caused by non-manifold geometry, which is a three-dimensional shape that cannot be unfolded into a two-dimensional surface with all normal vectors pointing in the same direction, such as multiple faces connecting to the same vertex), the script can be used to edit the isolated or redundant vertices and faces in the second three-dimensional model with the help of a graphics editor (for example, using a script to detect isolated or redundant vertices and faces to delete or merge the isolated vertices or faces), and the faces are flipped to make the normal vectors point in the same direction (for example, selecting all faces, using a script to check the normal direction, flipping the faces with inconsistent normal vectors, and ensuring that the normal vectors point outward). Thus, the second three-dimensional model after repairing the abnormality is converted into a topological optimization process.

[0301] Through topology optimization, a three-dimensional model (for example, a point cloud model, a point shading model, etc.) lacking a grid structure or being irregular in size or a three-dimensional model with too high a number of faces can be reconstructed in model structure through topology optimization, standardized and simplified, and thus better applied in a business environment.

[0302] In step 108, baking is performed based on the model map and the third three-dimensional model to obtain a fourth three-dimensional model.

[0303] In some embodiments, a three-dimensional model rendering tool (such as Blender, etc.) can be used to perform baking based on the model map and the third three-dimensional model to obtain the fourth three-dimensional model.

[0304] For example, in Blender, the type and parameters of baking are set, such as the resolution, sampling rate, margin, etc. of the baked texture, the properties of baking (such as color, light, shadow, AO, etc.) are selected, and the output path and file name of the baked model are set. Start the baking process, Blender will render the information and apply it to the target model (the fourth three-dimensional model), check the baked map and the fourth three-dimensional model to ensure that all details are correctly transferred, export the baked target model to the required format, such as FBX, OBJ or GLTF, etc.

[0305] Through steps 106 to 108, the practicability and quality of the three-dimensional model can be further improved. Specifically, first, the second three-dimensional model is optimized for mapping, and the generated model map not only enhances the realism and visual appeal of the three-dimensional model, but also facilitates secondary editing and improves subsequent rendering efficiency; second, topology optimization simplifies the model structure, reduces the number of faces, and solves the problems of high face number model, large computing power consumption, and non-manifold geometry, etc. through re-topology, face reduction, and structure repair, etc. The model structure is standardized and more suitable for business scenarios; finally, based on the model map and the third three-dimensional model, baking is performed to transfer the texture, lighting, and other detailed information to the fourth three-dimensional model, ensuring that the details are completely retained. This series of processing not only strengthens the visual performance of the three-dimensional model, but also improves the efficiency and compatibility of the three-dimensional model in editing, rendering, and application, allowing the three-dimensional model to maintain a high-quality appearance while better meeting the diverse needs of actual production environments.

[0306] The three-dimensional model generation method provided by the embodiments of the present application can be applied to various scenarios that need to generate three-dimensional models, some of which include: (1) game development: for example, game designers create game characters, environments, and props, etc. through three-dimensional model generation; (2) film and television production: for example, in film and animation production, special effects and animation scenes are created through three-dimensional model generation, etc. to bring visual shock to the audience; (3) industrial design: for example, engineers design product prototypes through three-dimensional model generation, etc. to improve product appearance.

[0307] Referring to FIG. 4, FIG. 4 is a schematic diagram of an application scenario of the three-dimensional model generation method provided in the embodiments of the present application. As shown in FIG. 4, in a game scenario, a game environment and game props (such as the "furniture" shown in FIG. 4) can be created by using the three-dimensional model generation method provided in the embodiments of the present application. The following takes the generation of a three-dimensional model in a game scenario as an example for description.

[0308] Referring to FIG. 5, FIG. 5 is a schematic diagram of the principle framework of the three-dimensional model generation method provided in the embodiments of the present application. The way of obtaining the model material can be to directly obtain an input image as the model material (corresponding to "Option 1: input image" in FIG. 5). The image material generation module is configured to generate a corresponding image from input text data through image material generation processing, so as to obtain the model material (corresponding to "Option 2: input text data" in FIG. 5). The multi-view generation module is configured to generate view images of multiple views based on the model material. The three-dimensional model generation module is configured to obtain a second three-dimensional model based on the multi-view images. The post-processing module is configured to perform texture optimization and topology optimization on the second three-dimensional model, so as to obtain a fourth three-dimensional model.

[0309] For example, referring to FIG. 8, FIG. 8 is a schematic diagram of the architecture of the three-dimensional model generation platform provided in the embodiments of the present application, which includes a first server, a second server and a plurality of third servers.

[0310] The first server (or referred to as a front-end page server) is the front-end part of the user interaction. For example, Nginx can be used as the server, and Vue3 can be used as the front-end framework to build and present the three-dimensional model generation interface. The first server is configured to receive user requests and forward them to the second server for processing.

[0311] The second server (or referred to as a routing server) is the middle layer of the three-dimensional model generation platform system, which includes three modules of model material storage, three-dimensional model storage and remeshing. The second server is responsible for processing the requests from the first server. Specifically, the model material storage (Image storage) is configured to store the images uploaded by the user, the three-dimensional model storage (Obj storage) is configured to store the generated three-dimensional model data, and the remeshing (Remesh) is configured to perform post-processing such as remeshing on the three-dimensional model. Specifically, the second server receives the request from the first server (for example, the first server transmits the request downward through IP proxy), sends the parameter information including the keyword, picture link and the like to the third server, and receives the three-dimensional model download link returned by the third server, such as Uniform Resource Locator (URL).

[0312] The third server (or model server cluster) is used to process specific three-dimensional model generation tasks, receives upper-layer parameter information, and returns a 3D three-dimensional model download URL after processing. Each third server can deploy a model corresponding to a different three-dimensional model generation algorithm. Specifically, a user sends a request through the first server, the request is transmitted to the second server through an IP proxy, the second server transmits parameter information in the request to a specific third server (third server 1, third server 2, third server n) in the back end, and the third server in the back end returns a three-dimensional model download URL to the second server for subsequent processing or storage.

[0313] The three-layer architecture effectively separates the user interface, business logic, and data processing, ensuring system scalability and maintainability.

[0314] Referring to FIG. 6, FIG. 6 is a flowchart of a three-dimensional model generation method of a game scene according to an embodiment of the present application. The following will be described in conjunction with FIG. 6.

[0315] In step 401, game model materials are obtained.

[0316] In some embodiments, the first server obtains game model materials (corresponding to the model materials described above) in response to a triggering operation on a three-dimensional model generation method selection control of the three-dimensional model generation interface.

[0317] For example, referring to FIG. 7A, FIG. 7A is a first schematic diagram of a three-dimensional model generation interface according to an embodiment of the present application. In response to a triggering operation on a three-dimensional model generation method selection control-002 (corresponding to text to 3D), text data is obtained, where the text data is used to describe a three-dimensional model. Based on the text data, image material generation processing is performed to obtain game model materials (see the description of step 1012 above).

[0318] For example, in response to a triggering operation on a three-dimensional model generation method selection control-001 (corresponding to image to 3D), image data is obtained directly as game model materials.

[0319] For example, referring to FIG. 7B, which is a second schematic diagram of a three-dimensional model generation interface provided by an embodiment of the present application, the three-dimensional model generation interface includes five modules, i.e., a three-dimensional model generation based on an image, a three-dimensional model generation based on text data, a personal history record, a queue management, and a re-topology. When the three-dimensional model generation mode selection control-001 (corresponding to image to 3D) in FIG. 7A is triggered, the three-dimensional model generation based on an image module is called to allow a user to upload or select image data, and the user can specify a generation algorithm (for example, InstantMeshes) of the three-dimensional model through “generation algorithm selection”. When the three-dimensional model generation mode selection control-002 (corresponding to text to 3D) in FIG. 7A is triggered, the three-dimensional model generation based on text data module is called to allow the user to input text data. The personal history record module is used to display information of the generated three-dimensional model. The queue management module is used to display a model generation progress. The re-topology module is used to further optimize the generated three-dimensional model to reduce a number of faces of the three-dimensional model or adjust a light intensity of the three-dimensional model obtained through rendering.

[0320] For example, the first server receives a user request and forwards the request to the second server for processing, where the request carries game model materials. The second server stores the game model materials in the request, sends parameter information including keywords (such as text to 3D and image to 3D), picture links (game model material links), and the like to the third server. The third server generates a corresponding three-dimensional model according to the parameter information, returns a three-dimensional model download URL to the second server, and the second server performs subsequent processing or storage.

[0321] In step 402, perspective images of multiple perspectives are generated based on game model materials.

[0322] In some embodiments, the game model materials are two-dimensional game model images. The third server performs feature extraction processing on the game model materials through a pre-trained multi-view generation model to obtain game model image features (corresponding to the model image features in the foregoing description). The third server performs multi-view generation processing based on the game model image features through the pre-trained multi-view generation model to obtain perspective images of multiple perspectives. The implementation of the multi-view generation can refer to the description of steps 1021 to 1022 in the foregoing description, which will not be repeated here.

[0323] For example, the training of the multi-view generation model can be implemented in the following manner: a second data set for a multi-view generation task is constructed, wherein the second data set includes a plurality of single-image samples and a plurality of view image samples of each single-image sample corresponding to a plurality of views (see the description of step 301 above); a feature extraction process is performed on the single-image samples to obtain image feature samples (see the description of step 302 above); a multi-view generation process is performed based on the image feature samples to obtain view prediction images of the plurality of views (see the description of step 303 above); a second loss value of the view prediction images of the plurality of views and the view image samples of the plurality of views is obtained through a pre-set loss function (see the description of step 304 above); and the network parameters of the original multi-view generation model are updated through the second loss value to obtain the trained multi-view generation model (see the description of step 305 above).

[0324] In step 403, a first three-dimensional model is generated based on the view images of the plurality of views, wherein the first three-dimensional model includes a first color parameter.

[0325] In some embodiments, the third server performs feature extraction processing on the view image of each view to obtain image features of each view (see the description of step 1031 above); performs feature encoding processing on the image features of each view to obtain encoded features of each view (see the description of step 1032 above); performs decoding processing based on the encoded features of each view to obtain three-dimensional features (see the description of step 1033 above); and performs three-dimensional reconstruction processing based on the three-dimensional features to obtain the first three-dimensional model (see the description of step 1034 above).

[0326] In other embodiments, the third server can generate the first three-dimensional model based on the view images of the plurality of views through a pre-trained three-dimensional model generation model, such as an Instant Mesh, etc. The present embodiment does not limit the specific three-dimensional model generation model to be used, and the first three-dimensional model can be a point cloud model, a point coloring model, a mesh model (Mesh), etc. The present embodiment does not limit the three-dimensional representation manner of the first three-dimensional model.

[0327] In step 404, a coloring parameter of each view image is obtained.

[0328] In some embodiments, the third server obtains a bounding box of the model object in the view image (see the description of step 1041 above); takes the bounding box as a base surface, determines a first vertex based on a pre-set distance parameter, and constructs a quadrangular pyramid based on the base surface and the first vertex (see the description of step 1042 above); and determines the coloring parameter of the view image through the positional relationship between the surface of the quadrangular pyramid and the first three-dimensional model (see the description of step 1043 above).

[0329] In step 405, the first color parameter of the first three-dimensional model is updated based on the shading parameter of each perspective image, to obtain a second three-dimensional model.

[0330] In some embodiments, the shading parameter further includes a second color parameter of each pixel point of the model object in the perspective image, the first three-dimensional model includes a plurality of second vertices, and each second vertex corresponds to a first color parameter. The third server generates a ray in the object coordinate system, the ray has an end point at the viewpoint coordinate and passes through the pixel point, determines a second vertex on the first three-dimensional model that intersects with the ray as a shading point (see the description of step 1051 above), and replaces the first color parameter of the shading point in the first three-dimensional model with the second color parameter of the pixel point to obtain the second three-dimensional model (see the description of step 1052 above).

[0331] In step 406, the second three-dimensional model is post-processed to obtain a fourth three-dimensional model.

[0332] In some embodiments, the second server performs texture optimization on the second three-dimensional model to obtain a model texture (see the description of step 106 above), performs topology optimization on the second three-dimensional model to obtain a third three-dimensional model (see the description of step 107 above), and performs baking based on the model texture and the third three-dimensional model to obtain the fourth three-dimensional model (see the description of step 108 above).

[0333] For example, referring to FIG. 21A, which is a schematic diagram of the principle of post-processing provided in an embodiment of the present application, topology optimization includes surface reduction processing (for example, a surface reduction method based on voxel reconstruction) on the second three-dimensional model, and re-topology processing on the second three-dimensional model after surface reduction to obtain a third three-dimensional model. Texture optimization can include UV unfolding of a texture image to obtain a model texture, so that the model texture is baked to the third three-dimensional model to obtain a fourth three-dimensional model that can be used in a business environment.

[0334] For example, referring to FIG. 21B, which is a flowchart of post-processing provided in an embodiment of the present application, the following describes the steps in FIG. 21B.

[0335] In step 501, the second three-dimensional model is imported.

[0336] In some embodiments, the second three-dimensional model is imported in a three-dimensional model editing tool (for example, Blender, etc.).

[0337] In step 502, it is determined whether the second three-dimensional model has a material. In response to the second three-dimensional model having a material, the process proceeds to steps 506 to 508. In response to the second three-dimensional model not having a material, the process proceeds to steps 503 to 505.

[0338] In some embodiments, the material information of the second three-dimensional model can be obtained through the material editor of Blender, and whether the second three-dimensional model has a material can also be determined by the category of the second three-dimensional model. For example, when the second three-dimensional model is a point cloud model or a point color model, the second three-dimensional model does not have a material.

[0339] Here, the material (Material) is an attribute parameter used to define how the surface of a three-dimensional model interacts with light. The material determines the appearance of the three-dimensional model when rendered, including color, gloss, transparency, and other characteristics.

[0340] In step 503, a material is added.

[0341] In some embodiments, a new material can be added to the second three-dimensional model through the material editor of Blender.

[0342] In step 504, the second three-dimensional model is copied.

[0343] In some embodiments, the second three-dimensional model with the added material is copied as a copy model.

[0344] In step 505, a color attribute is added to the original second three-dimensional model.

[0345] In some embodiments, a color attribute is added to each vertex of the original model (the original second three-dimensional model without added material) through Blender, so that the vertex color will be used when rendering, which may override the color of the material unless the material uses some specific node to mix these colors. Specifically, if the vertex color is incompatible with the material color (for example, the vertex color completely covers the material color), the rendering effect will only show the vertex color, and in complex material and lighting settings, the vertex color can interact with other attributes of the material (such as texture, transparency, reflectivity, etc.) in a complex way.

[0346] In step 506, the second three-dimensional model is copied.

[0347] In some embodiments, the second three-dimensional model is copied as a copy model.

[0348] In step 507, the map of the original second three-dimensional model is removed.

[0349] In some embodiments, the material information on the original model is removed through the material editor of Blender.

[0350] In step 508, the copy model deletes redundant nodes.

[0351] In some embodiments, redundant nodes in the duplicated model are filtered out, such as through the "Node View" of Blender's Material Editor, viewing the node links of the material, for each unnecessary node, selecting "Delete" or "Remove Node", for example, the node is connected to another node, but its output is not connected anywhere, then it may be a redundant node, before deleting the node, the function of the node needs to be understood to avoid deleting important nodes.

[0352] In step 509, the number of faces that need to be optimized is calculated.

[0353] In some embodiments, the "Properties" panel of Blender can be used to find the "Geometry" property of the model, which will display the number of faces of the model, record the current number of faces, and determine the number of faces that need to be reduced or increased according to the size and detail requirements of the model.

[0354] In step 510, topology optimization.

[0355] In some embodiments, the duplicated model is re-topologized to optimize the geometry of the model.

[0356] In step 511, UV mapping.

[0357] In some embodiments, UV mapping is performed through UV unwrapping algorithms such as automatic unwrapping or manual unwrapping. UV mapping refers to creating a two-dimensional coordinate mapping that converts vertex coordinates on the duplicated model to coordinates on a two-dimensional plane. This process is very important because it allows texture images such as texture maps, material maps, etc. to be correctly applied to three-dimensional models.

[0358] In step 512, specify the baking map.

[0359] In some embodiments, the model map for baking is generated, and the texture type and properties that need to be baked are specified, such as color, concave-convex, or lighting, etc.

[0360] In step 513, superimpose the model.

[0361] In some embodiments, the duplicated model that has been topologically optimized and UV mapped is superimposed onto the original second three-dimensional model.

[0362] In step 514, calculate the extrusion parameters.

[0363] In some embodiments, the extrusion parameters of the duplicated model are calculated to ensure that the duplicated model does not overlap when extruded, thereby avoiding abnormalities in baking textures.

[0364] In step 515, set the baking parameters.

[0365] In some embodiments, parameters required for baking, such as resolution, lighting conditions, etc., are set to ensure that the baking effect meets expectations. The baking process reduces the computational burden of real-time rendering by precomputing lighting and texture information, which means that the setting of baking parameters will directly affect the quality and effect of the final baked texture. By adjusting the baking parameters, the level of detail of the baked texture and the quality of the rendering effect can be controlled. For example, increasing the resolution can increase the detail of the texture, but also increases the baking time and the required data storage. By setting specific baking parameters, specific visual effects can be achieved, such as soft shadows, strong highlights, or specific style lighting effects.

[0366] For example, resolution determines the level of detail of the baked texture, high resolution can provide more detailed texture, but will increase the baking time and file size; lighting conditions include light source type, intensity and direction, these parameters affect the lighting and shadow effect in the baked texture; sampling rate determines the accuracy of the calculation in the baking process, high sampling rate can reduce noise, but also increases the baking time; baking mode selects the type of baked texture, such as diffuse map, normal map, ambient occlusion map, etc. In summary, setting baking parameters is to find a balance between the quality of the model, visual effects and performance, to ensure that the baking result meets the expected artistic effect and practical application requirements.

[0367] In step 516, baking.

[0368] In some embodiments, the generated model map is baked onto the topologically optimized copy model.

[0369] In step 517, a fourth three-dimensional model is exported.

[0370] In some embodiments, the fourth three-dimensional model after retopology and baking is exported in the required format (such as FBX, OBJ or GLTF, etc.), to facilitate subsequent use and application.

[0371] The generation of the first three-dimensional model based on the perspective images of multiple perspectives is realized through steps 401 to 406 and steps 501 to 517, the color parameter of the three-dimensional model is optimized by updating the first color parameter of the first three-dimensional model by obtaining the shading parameter of each perspective image, and the color details of the second three-dimensional model are further enhanced, so as to achieve the beneficial effect of improving the quality of the generated three-dimensional model. Through post-processing of the second three-dimensional model, the model map obtained after map optimization can be conveniently edited again, and these effects can be quickly applied in subsequent rendering or real-time display without the need for re-computation, thereby achieving the beneficial effect of improving the rendering efficiency of the three-dimensional model. Through topology optimization, three-dimensional models that lack available grid structures or are irregular in size (such as point cloud models, point shading models, etc.) or have too many faces can be reconstructed in model structure through topology optimization, so as to realize standardization and simplification of the model structure, and thus better applied in business environment.

[0372] The following continues to illustrate an exemplary structure of the three-dimensional model generation apparatus 133 implemented as a software module. In some embodiments, as shown in FIG. 2, the software module stored in the three-dimensional model generation apparatus 133 in the memory 130 can include:

[0373] A data acquisition module 1331 configured to acquire model materials.

[0374] A multi-view generation module 1332 configured to generate perspective images of multiple perspectives based on the model materials.

[0375] A three-dimensional model generation module 1333 configured to generate a first three-dimensional model based on the perspective images of the multiple perspectives, wherein the first three-dimensional model includes a first color parameter.

[0376] In some embodiments, the three-dimensional model generation module 1333 is further configured to obtain a shading parameter of each of the perspective images.

[0377] In some embodiments, the three-dimensional model generation module 1333 is further configured to update the first color parameter of the first three-dimensional model based on the shading parameter of each of the perspective images to obtain a second three-dimensional model.

[0378] In some embodiments, the three-dimensional model generation module 1333 is further configured to perform the following processing on each of the perspective images: obtaining a bounding box of a model object in the perspective image; taking the bounding box as a base surface, determining a first vertex based on a preset distance parameter, and constructing a quadrangular pyramid based on the base surface and the first vertex; and determining a shading parameter of the perspective image based on a positional relationship between a surface of the quadrangular pyramid and the first three-dimensional model.

[0379] In some embodiments, the shading parameters include a viewpoint coordinate, each of the perspective images corresponds to a camera perspective, and the three-dimensional model generation module 1333 is further configured to move a face of the four-sided pyramid in an object coordinate system in which the first three-dimensional model is located along a direction of the camera perspective corresponding to the perspective image until the face of the four-sided pyramid is tangent to a boundary of the first three-dimensional model, determine an intersection point of the face of the four-sided pyramid tangent to the boundary of the first three-dimensional model, and take a coordinate of the intersection point in the object coordinate system as the viewpoint coordinate.

[0380] In some embodiments, the shading parameters further include a second color parameter of each pixel point of the model object in the perspective image, the first three-dimensional model includes a plurality of second vertices, and the three-dimensional model generation module 1333 is further configured to perform the following processing on each of the pixel points of the model object in the perspective image: generate a ray in the object coordinate system that has the viewpoint coordinate as an end point and passes through the pixel point, determine a second vertex on the first three-dimensional model that intersects with the ray as a shading point, and replace a first color parameter of the shading point in the first three-dimensional model with the second color parameter of the pixel point to obtain a second three-dimensional model.

[0381] In some embodiments, the three-dimensional model generation module 1333 is further configured to perform the following processing on each of the second vertices in the first three-dimensional model: in response to the second vertex being in at least two of the camera perspectives, perform a weighted average processing on second color parameters of at least two pixel points corresponding to the second vertex in the two camera perspectives respectively to obtain a third color parameter, and replace the first color parameter of the second vertex in the first three-dimensional model with the third color parameter to obtain a second three-dimensional model.

[0382] In some embodiments, the three-dimensional model generation module 1333 is further configured to perform the following processing on each of the second vertices in the first three-dimensional model: in response to the second vertex not being in any of the camera perspectives, obtain at least two neighbor vertices adjacent to the second vertex within a preset range, obtain an average value of the second color parameters corresponding to the at least two neighbor vertices, and replace the first color parameter of the second vertex in the first three-dimensional model with the average value to obtain a second three-dimensional model.

[0383] In some embodiments, the data acquisition module 1331 is further configured to acquire text data, where the text data is used to describe a three-dimensional model, and perform image material generation processing based on the text data to obtain the model material.

[0384] In some embodiments, the image material generation process is implemented by a pre-trained image generation model, and the data acquisition module 1331 is further configured to construct a first data set for an image material generation task, wherein the first data set includes a plurality of image samples and a text description of each image sample, and the plurality of image samples have the same style; obtain a pre-trained original image generation model, and perform the following processes by using the original image generation model: performing feature extraction processing on the text description to obtain a text feature; performing image generation processing based on the text feature to obtain a predicted image; obtaining a first loss value of the predicted image and the image sample by using a pre-set loss function; constructing a fine-tuning weight parameter, and adding the fine-tuning weight parameter to network parameters of the original image generation model; updating the fine-tuning weight parameter by using the first loss value to obtain a trained image generation model.

[0385] In some embodiments, the data acquisition module 1331 is further configured to generate a first low-rank matrix initialized based on a Gaussian distribution, and generate a second low-rank matrix initialized based on all zeros; and construct a fine-tuning weight parameter based on the first low-rank matrix and the second low-rank matrix.

[0386] In some embodiments, the data acquisition module 1331 is further configured to keep parameters in the network parameters other than the first low-rank matrix and the second low-rank matrix unchanged, and update the first low-rank matrix and the second low-rank matrix by using the first loss value.

[0387] In some embodiments, the data acquisition module 1331 is further configured to obtain a plurality of three-dimensional model samples, wherein the plurality of three-dimensional model samples have the same style and different shapes; perform rasterization processing on each three-dimensional model sample according to a pre-set fixed view angle to obtain an image sample corresponding to each three-dimensional model sample; perform text generation processing on each image sample to obtain a text description corresponding to each image sample; and combine each image sample and the corresponding text description into the first data set.

[0388] In some embodiments, the multi-view generation module 1332 is further configured to perform feature extraction processing on the model material by using a pre-trained multi-view generation model to obtain a model image feature; and perform multi-view generation processing based on the model image feature by using the pre-trained multi-view generation model to obtain the view images of the plurality of views.

[0389] In some embodiments, the multi-view generation module 1332 is further configured to construct a second data set for a multi-view generation task, where the second data set includes a plurality of single image samples and a plurality of view image samples corresponding to each single image sample; perform the following processing by using a pre-trained original multi-view generation model: performing feature extraction processing on the single image samples to obtain image feature samples; performing multi-view generation processing based on the image feature samples to obtain view prediction images of a plurality of views; obtaining a second loss value of the view prediction images of the plurality of views and the view image samples of the plurality of views by using a pre-set loss function; and updating network parameters of the original multi-view generation model by using the second loss value to obtain a trained multi-view generation model.

[0390] In some embodiments, the multi-view generation module 1332 is further configured to obtain a plurality of three-dimensional model samples, and perform the following processing on each three-dimensional model sample: sampling the three-dimensional model sample according to pre-set view parameters and lighting parameters of a plurality of views to obtain view image samples of the plurality of views of the three-dimensional model sample; sampling the three-dimensional model sample according to the view parameters and the lighting parameters of a single view to obtain a single image sample of the three-dimensional model sample; and combining the view image samples of the plurality of views and the single image sample of each three-dimensional model sample into the second data set.

[0391] In some embodiments, the view parameters of the plurality of views include a camera distance parameter, and the multi-view generation module 1332 is further configured to perform edge detection on each view image sample to obtain an edge detection result; in response to the edge detection result indicating that an edge of a model object in the view image sample exceeds a boundary of the view image sample, increasing the camera distance parameter according to a pre-set parameter, and sampling the three-dimensional model sample according to the increased camera distance parameter to obtain new view image samples of the plurality of views of the three-dimensional model sample.

[0392] In some embodiments, the multi-view generation module 1332 is further configured to perform the following processing for each view: fusing view information of the view with the image feature sample to obtain conditional information features; performing denoising processing on a noise map based on the conditional information features to obtain the view prediction image.

[0393] In some embodiments, the multi-view generation module 1332 is further configured to perform attention encoding on the conditional information features and the noise map to obtain attention encoding features; performing noise prediction based on the attention encoding features to obtain predicted noise; and removing the predicted noise from the noise map to obtain the view prediction image.

[0394] In some embodiments, the three-dimensional model generation module 1333 is further configured to perform feature extraction processing on the view image of each view to obtain image features of each view, perform feature coding processing on the image features of each view to obtain coded features of each view, perform decoding processing based on the coded features of each view to obtain three-dimensional features, and perform three-dimensional reconstruction processing based on the three-dimensional features to obtain the first three-dimensional model.

[0395] In some embodiments, the three-dimensional model generation module 1333 is further configured to perform texture optimization on the second three-dimensional model to obtain a model texture, perform topology optimization on the second three-dimensional model to obtain a third three-dimensional model, and perform baking based on the model texture and the third three-dimensional model to obtain a fourth three-dimensional model.

[0396] The embodiments of the present application provide a computer program product, which includes a computer program or computer executable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions to cause the electronic device to perform the three-dimensional model generation method provided by the embodiments of the present application.

[0397] The embodiments of the present application provide a computer readable storage medium storing computer executable instructions, wherein the computer executable instructions or computer programs are stored in the computer readable storage medium. When the computer executable instructions or computer programs are executed by a processor, the processor will execute the three-dimensional model generation method provided by the embodiments of the present application, for example, the three-dimensional model generation method shown in FIG. 3A.

[0398] In some embodiments, the computer readable storage medium can be a RAM, a ROM, a flash memory, a magnetic surface memory, an optical disc, or a CD-ROM, etc. The computer readable storage medium can also be various devices including one or any combination of the above storage devices.

[0399] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.

[0400] As an example, computer-executable instructions can be, but need not be, related directly to a file. Computer-executable instructions can be, but need not be, directly perceivable as a single file. Computer-executable instructions can be, but need not be, stored in a file that is dedicated to a particular program or utility. Computer-executable instructions can be, but need not be, stored in a file that is stored in a dedicated location on a storage device. Computer-executable instructions can be, but need not be, stored in a file that is stored in a dedicated location on a storage device.

[0401] As an example, computer-executable instructions can be deployed to be executed on one electronic device or alternatively on multiple electronic devices that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0402] To sum up, by the embodiments of the present application, after generating the first three-dimensional model by the model material and generating the perspective images of multiple perspectives based on the model material, the first color parameter of the first three-dimensional model is updated based on the colorizing parameters of each perspective image, which realizes updating and calibrating the first color parameter of the first three-dimensional model under different perspectives, makes the first color parameter more consistent with the real color attribute of the model material after updating, ensures that the color performance of the second three-dimensional model under any observation angle remains highly consistent, avoids local color deviation caused by the limitation of a single perspective image, and thus further enhances the color details of the second three-dimensional model and improves the quality of the generated three-dimensional model.

[0403] The above merely describes the embodiments of the present application, but does not serve to limit the protection scope of the present application. Any modification, equivalent replacement and improvement within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A three-dimensional model generation method applied to an electronic device, the method comprising: obtaining a model material; generating a plurality of perspective images of different perspectives based on the model material; generating a first three-dimensional model based on the plurality of perspective images of different perspectives, wherein the first three-dimensional model comprises a first color parameter; obtaining a coloring parameter of each of the perspective images; updating the first color parameter of the first three-dimensional model based on the coloring parameter of each of the perspective images to obtain a second three-dimensional model.

2. The method of claim 1, wherein, The obtaining of the coloring parameter of each of the perspective images comprises: performing the following processing on each of the perspective images: obtaining a bounding box of a model object in the perspective image; determining a first vertex based on a preset distance parameter with the bounding box as a base surface, and constructing a quadrangular pyramid based on the base surface and the first vertex; determining the coloring parameter of the perspective image based on the positional relationship between the surface of the quadrangular pyramid and the first three-dimensional model.

3. The method of any one of claims 1 to 2, wherein, The coloring parameter comprises a viewpoint coordinate, each of the perspective images corresponds to a camera perspective, and the determination of the coloring parameter of the perspective image based on the positional relationship between the surface of the quadrangular pyramid and the first three-dimensional model comprises: moving the surface of the quadrangular pyramid in the direction of the camera perspective corresponding to the perspective image in the object coordinate system in which the first three-dimensional model is located until the surface of the quadrangular pyramid is tangent to the boundary of the first three-dimensional model; determining the intersection point of the surface of the quadrangular pyramid tangent to the boundary of the first three-dimensional model, and taking the coordinates of the intersection point in the object coordinate system as the viewpoint coordinate.

4. The method according to any one of claims 1 to 3, wherein, The coloring parameter further comprises a second color parameter of each pixel point of the model object in the perspective image, the first three-dimensional model comprises a plurality of second vertices, and the updating of the first color parameter of the first three-dimensional model based on the coloring parameter of each of the perspective images to obtain a second three-dimensional model comprises: performing the following processing on each of the pixel points of the model object in the perspective image: generating a ray passing through the pixel point with the viewpoint coordinate as an end point in the object coordinate system, and determining a second vertex on the first three-dimensional model intersecting with the ray as a coloring point; replacing the first color parameter of the coloring point in the first three-dimensional model with the second color parameter of the pixel point to obtain a second three-dimensional model.

5. The method according to any one of claims 1 to 4, wherein, The updating of the first color parameter of the first three-dimensional model based on the coloring parameter of each of the perspective images to obtain a second three-dimensional model comprises: performing the following processing on each of the second vertices in the first three-dimensional model: in response to the second vertex being in at least two camera perspectives, performing a weighted average processing on the second color parameters of at least two pixel points corresponding to the second vertex in the two camera perspectives respectively to obtain a third color parameter; replacing the first color parameter of the second vertex in the first three-dimensional model with the third color parameter to obtain a second three-dimensional model.

6. The method according to any one of claims 1 to 5, wherein, The first color parameter of the first three-dimensional model is updated based on the shading parameter of each view image to obtain a second three-dimensional model, including: The following processing is performed on each second vertex in the first three-dimensional model: In response to the second vertex not being in any camera view, at least two neighbor vertices adjacent to the second vertex within a preset range are obtained; An average value of the second color parameter corresponding to the at least two neighbor vertices is obtained; The first color parameter of the second vertex in the first three-dimensional model is replaced by the average value to obtain a second three-dimensional model.

7. The method according to any one of claims 1 to 6, wherein, The model material includes: Obtaining text data, wherein the text data is used to describe a three-dimensional model; Based on the text data, image material generation processing is performed to obtain the model material.

8. The method of claim 7, wherein, The image material generation processing is implemented through a pre-trained image generation model, and the image generation model is trained in the following manner: A first data set for an image material generation task is constructed, wherein the first data set includes a plurality of image samples and a text description of each image sample, and the plurality of image samples have the same style; A pre-trained original image generation model is obtained, and the following processing is performed through the original image generation model: Feature extraction processing is performed on the text description to obtain text features; Based on the text features, image generation processing is performed to obtain a predicted image; A first loss value of the predicted image and the image sample is obtained through a pre-set loss function; A fine-tuning weight parameter is constructed, and the fine-tuning weight parameter is added to the network parameters of the original image generation model; The fine-tuning weight parameter is updated through the first loss value to obtain a trained image generation model.

9. The method of any one of claims 7-8, wherein, The construction of the fine-tuning weight parameter includes: A first low-rank matrix based on Gaussian distribution initialization is generated, and a second low-rank matrix based on all-zero initialization is generated; The fine-tuning weight parameter is constructed based on the first low-rank matrix and the second low-rank matrix; The fine-tuning weight parameter is updated through the first loss value, including: The parameters in the network parameters except the first low-rank matrix and the second low-rank matrix remain unchanged, and the first low-rank matrix and the second low-rank matrix are updated through the first loss value.

10. The method of any one of claims 7 to 8, wherein, The construction of the first data set for the image material generation task includes: A plurality of three-dimensional model samples are obtained, wherein the plurality of three-dimensional model samples have the same style and different shapes; Each three-dimensional model sample is rasterized according to a pre-set fixed view angle to obtain an image sample corresponding to each three-dimensional model sample; Text generation processing is performed on each image sample to obtain a text description corresponding to each image sample; Each image sample and the corresponding text description are combined into the first data set.

11. The method according to any one of claims 1 to 10, wherein, The model material is a two-dimensional model image, and the view images of multiple views are generated based on the model material, including: Feature extraction processing is performed on the model material through a pre-trained multi-view generation model to obtain model image features; The pre-trained multi-view generation model is used to perform multi-view generation processing based on the image features, so as to obtain the view images of the multiple views.

12. The method of claim 11, wherein, The multi-view generation model is trained in the following manner: A second data set for a multi-view generation task is constructed, wherein the second data set includes multiple single image samples and view images of multiple views corresponding to each single image sample; The following processing is performed by using a pre-trained original multi-view generation model: Feature extraction processing is performed on the single image sample to obtain an image feature sample; Multi-view generation processing is performed based on the image feature sample to obtain view prediction images of multiple views; A second loss value of the view prediction images of the multiple views and the view image samples of the multiple views is obtained by using a pre-set loss function; The network parameters of the original multi-view generation model are updated based on the second loss value, so as to obtain a trained multi-view generation model.

13. The method according to any one of claims 11 to 12, characterized in that, The construction of the second data set for the multi-view generation task includes: A plurality of three-dimensional model samples are obtained, and the following processing is performed on each three-dimensional model sample: The three-dimensional model sample is sampled according to pre-set view parameters and illumination parameters of multiple views, so as to obtain view image samples of multiple views of the three-dimensional model sample; The three-dimensional model sample is sampled according to the pre-set view parameters and the illumination parameters of a single view, so as to obtain a single image sample of the three-dimensional model sample; The view image samples of multiple views and the single image sample of each three-dimensional model sample are combined into the second data set.

14. The method according to any one of claims 11 to 13, characterized in that, The view parameters of the multiple views include a camera distance parameter, and after the view image samples of multiple views of the three-dimensional model sample are obtained, the method further includes: Edge detection is performed on each view image sample to obtain an edge detection result; In response to the edge detection result representing that the edges of a model object in the view image sample exceed the boundary of the view image sample, the camera distance parameter is increased according to a pre-set parameter, and the three-dimensional model sample is sampled according to the increased camera distance parameter to obtain new view image samples of multiple views of the three-dimensional model sample.

15. The method according to any one of claims 11 to 14, wherein, The multi-view generation processing based on the image feature sample to obtain view prediction images of multiple views includes: The following processing is performed for each view: The view information of the view is fused with the image feature sample to obtain conditional information features; Noise in a noise map is removed based on the conditional information features to obtain the view prediction image.

16. The method according to any one of claims 11 to 15, wherein, The noise removal processing based on the conditional information features to obtain the view prediction image includes: Attention encoding is performed on the conditional information features and the noise map to obtain attention encoding features; Noise prediction is performed based on the attention encoding features to obtain predicted noise; The predicted noise is removed from the noise map to obtain the view prediction image.

17. The method of any one of claims 1 to 16, wherein, The generation of a first three-dimensional model based on the view images of the multiple views includes: perform feature extraction processing on the view image of each of the view angles to obtain image features of each view angle; perform feature coding processing on the image features of each of the view angles to obtain coding features of each view angle; perform decoding processing based on the coding features of each of the view angles to obtain three-dimensional features; perform three-dimensional reconstruction processing based on the three-dimensional features to obtain the first three-dimensional model.

18. The method of any one of claims 1 to 17, wherein, After the second three-dimensional model is obtained, the method further includes: perform texture optimization on the second three-dimensional model to obtain a model texture; perform topology optimization on the second three-dimensional model to obtain a third three-dimensional model; perform baking based on the model texture and the third three-dimensional model to obtain a fourth three-dimensional model.

19. A three-dimensional model generation apparatus, comprising: a data acquisition module configured to acquire model materials; a multi-view generation module configured to generate view images of multiple view angles based on the model materials; a three-dimensional model generation module configured to generate a first three-dimensional model based on the view images of the multiple view angles, wherein the first three-dimensional model comprises a first color parameter; the three-dimensional model generation module is further configured to acquire a coloring parameter of each of the view images; the three-dimensional model generation module is further configured to update the first color parameter of the first three-dimensional model based on the coloring parameter of each of the view images to obtain a second three-dimensional model.

20. An electronic device, comprising: a memory configured to store computer executable instructions or computer programs; a processor configured to execute the computer executable instructions or computer programs stored in the memory to implement the three-dimensional model generation method of any one of claims 1 to 18.

21. A computer readable storage medium storing computer executable instructions or computer programs, wherein the computer executable instructions or computer programs are executed by a processor to implement the three-dimensional model generation method of any one of claims 1 to 18.

22. A computer program product comprising computer executable instructions or computer programs, wherein the computer executable instructions or computer programs are executed by a processor to implement the three-dimensional model generation method of any one of claims 1 to 18.

Citation Information

Patent Citations

  • Method for generating three-dimensional model by text, computer equipment and storage medium thereof

    CN117218281A

  • Three-dimensional model generation method and device and electronic equipment

    CN117372607A

  • Data rendering method, device and equipment and computer readable storage medium

    CN117541703A

  • Image processing apparatus, method and storage medium

    US20190378326A1

Cited By

  • A scene editing and management method, device and storage medium

    CN122156556A

  • Three-dimensional scene generation method, electronic device, and storage medium

    CN122244341A

  • Three-dimensional scene generation method, electronic device, and storage medium

    CN122244341B

  • A method and apparatus for texture remapping of a three-dimensional model and a computing device

    CN122336103A