Virtual object generation method and apparatus, computer device and storage medium
By obtaining object description information and the characteristics of spatial points in the target three-dimensional space, three-dimensional virtual objects are automatically generated, which solves the problem of low efficiency in virtual object development in the existing technology, and achieves efficient and accurate virtual object generation.
Patent Information
- Application Number
- PCT/CN2024/114529
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-01
- Filing Date
- 2024-08-26
- Publication Date
- 2025-05-08
AI Technical Summary
In the process of game development, the development efficiency of virtual objects is low and requires manual development.
A virtual object generation method is provided, by obtaining object description information, obtaining features of multiple spatial points in the target three-dimensional space, and processing these features to generate three-dimensional virtual objects.
The method of automatically generating three-dimensional virtual objects is realized, which improves the efficiency of generating three-dimensional virtual objects and ensures the accuracy of objects.
Smart Images

Figure CN2024114529_08052025_PF_FP_ABST
Abstract
Description
Virtual object generation method, device, computer equipment and storage medium
[0001] This application claims priority to the Chinese patent application filed on November 1, 2023, with application number 202311446956.X and invention name “Virtual object generation method, device, computer equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a method, apparatus, computer device, and storage medium for generating a virtual object. Background Art
[0003] Taking the game scene as an example, with the development of computer technology, games are becoming more and more popular among users. Usually, developers develop games and then release the developed games so that users can experience the games.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, computer device, and storage medium for generating a virtual object, which can improve the efficiency of generating three-dimensional virtual objects. The technical solution is as follows:
[0006] In one aspect, a method for generating a virtual object is provided, the method comprising:
[0007] Obtaining object description information, where the object description information is used to describe the three-dimensional virtual object to be generated;
[0008] acquiring, based on the object description information, first features of a plurality of first spatial points in the target three-dimensional space, the first features of the first spatial points being used to characterize a color of the first spatial points in the target three-dimensional space and a positional relationship with the three-dimensional virtual object, the first spatial points being uniformly distributed in the target three-dimensional space;
[0009] Processing the first feature of each first spatial point by rendering the model to obtain a color and a directed distance of each first spatial point, wherein the directed distance indicates a distance from the first spatial point to the surface of the three-dimensional virtual object in the target three-dimensional space;
[0010] The three-dimensional virtual object is generated in the target three-dimensional space based on the colors and directed distances of the plurality of first spatial points.
[0011] In another aspect, a virtual object generation device is provided, the device comprising:
[0012] An acquisition module, configured to acquire object description information, wherein the object description information is used to describe a three-dimensional virtual object to be generated;
[0013] an updating module, configured to obtain, based on the object description information, first features of a plurality of first spatial points in a target three-dimensional space, the first features of the first spatial points being used to characterize a color of the first spatial points in the target three-dimensional space and a positional relationship with the three-dimensional virtual object, the first spatial points being uniformly distributed in the target three-dimensional space;
[0014] a processing module, configured to process the first feature of each first spatial point by rendering a model to obtain a color and a directed distance of each first spatial point, wherein the directed distance indicates a distance from the first spatial point to the surface of the three-dimensional virtual object in the target three-dimensional space;
[0015] A generating module is configured to generate the three-dimensional virtual object in the target three-dimensional space based on the colors and directed distances of the plurality of first spatial points.
[0016] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the virtual object generation method as described in the above aspects.
[0017] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the virtual object generation method as described in the above aspects.
[0018] On the other hand, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the operations performed by the virtual object generation method as described in the above aspects are implemented.
[0019] In the solution provided in the embodiment of the present application, object description information is used to describe the three-dimensional virtual object to be generated. The object description information is used to obtain the characteristics of each spatial point in the target three-dimensional space, and then the characteristics of each spatial point are processed through the rendering model to obtain the color and directed distance of each spatial point. Based on the color and directed distance of each spatial point, a three-dimensional virtual object can be generated in the target three-dimensional space. In this way, the shape or color of the generated three-dimensional virtual object is the same as the three-dimensional virtual object described by the object description information, thereby ensuring the accuracy of the three-dimensional virtual object, and realizing a method of automatically generating three-dimensional virtual objects using object description information. There is no need to manually develop three-dimensional virtual objects, thereby improving the efficiency of generating three-dimensional virtual objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG1 is a schematic diagram of a structure of an implementation environment provided by an embodiment of the present application;
[0021] FIG2 is a flow chart of a method for generating a virtual object provided in an embodiment of the present application;
[0022] FIG3 is a flowchart of another method for generating a virtual object provided in an embodiment of the present application;
[0023] FIG4 is a schematic diagram of point cloud data provided by an embodiment of the present application;
[0024] FIG5 is a schematic diagram of a directed distance field provided in an embodiment of the present application;
[0025] FIG6 is a schematic diagram of a model of a three-dimensional virtual object provided in an embodiment of the present application;
[0026] FIG7 is a schematic diagram of another model of a three-dimensional virtual object provided in an embodiment of the present application;
[0027] FIG8 is a flowchart of another method for generating a virtual object provided in an embodiment of the present application;
[0028] FIG9 is a schematic diagram of collecting a third spatial point according to an embodiment of the present application;
[0029] FIG10 is a flowchart of obtaining a predicted color and a predicted directed distance of a third spatial point provided by an embodiment of the present application;
[0030] FIG11 is a flowchart of another method for generating a virtual object provided in an embodiment of the present application;
[0031] FIG12 is a schematic diagram of generating a three-dimensional virtual object according to an embodiment of the present application;
[0032] FIG13 is a schematic diagram of another embodiment of the present application for generating a three-dimensional virtual object;
[0033] FIG14 is a schematic diagram of another method for generating a three-dimensional virtual object provided by an embodiment of the present application;
[0034] FIG15 is a schematic diagram of another method for generating a three-dimensional virtual object provided by an embodiment of the present application;
[0035] FIG16 is a schematic structural diagram of a virtual object generation device provided in an embodiment of the present application;
[0036] FIG17 is a schematic structural diagram of another virtual object generation device provided in an embodiment of the present application;
[0037] FIG18 is a schematic structural diagram of a terminal provided in an embodiment of the present application;
[0038] FIG19 is a schematic structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] As used herein, the terms "first," "second," "third," "fourth," "fifth," "sixth," and the like may be used to describe various concepts herein, but unless otherwise specified, these concepts are not limited by these terms. These terms are used solely to distinguish one concept from another. For example, a first spatial point can be referred to as a second spatial point, and similarly, a second spatial point can be referred to as a first spatial point without departing from the scope of this application.
[0040] As used herein, the terms "at least one," "a plurality," "each," and "any" include one, two, or more than two, "a plurality" includes two or more than two, "each" refers to each of the corresponding plurality, and "any" refers to any one of the plurality. For example, the plurality of spatial points includes three spatial points, and "each" refers to each of the three spatial points. "Any" refers to any one of the three spatial points, which can be the first spatial point, the second spatial point, or the third spatial point.
[0041] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the object description information and three-dimensional virtual objects involved in this application are obtained with full authorization.
[0042] In the related art, games include virtual objects. During the game development process, developers usually manually develop the virtual objects in the game, which leads to low development efficiency of the virtual objects.
[0043] The virtual object generation method provided in the embodiments of the present application is performed by a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers. Optionally, the terminal is a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, intelligent voice interaction device, smart home appliance, and vehicle-mounted terminal, etc., but is not limited thereto.
[0044] In some embodiments, the computer device is provided as a server. FIG1 is a schematic diagram of an implementation environment provided by an embodiment of the present application. Referring to FIG1 , the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a wireless or wired network.
[0045] The terminal 101 is used to obtain object description information and send the object description information to the server 102 via a network connection with the server 102. The server 102 is used to receive the object description information sent by the terminal 101 and generate a three-dimensional virtual object described by the object description information based on the object description information.
[0046] Optionally, when generating a three-dimensional virtual object, the server 102 sends the three-dimensional virtual object to the terminal 101 , and the terminal 101 can receive the three-dimensional virtual object and display the three-dimensional virtual object.
[0047] Optionally, an application provided by server 102 is installed on terminal 101, and terminal 101 can use this application to implement functions such as virtual object construction. Optionally, the application is an application in the terminal 101 operating system, or an application provided by a third party. For example, the application is a virtual object construction application, which has the function of constructing virtual objects. Of course, the virtual object construction application can also have other functions, such as review functions, shopping functions, navigation functions, game functions, etc.
[0048] Terminal 101 is used to log in to the application based on the identifier, and can obtain object description information input by the user based on the application, and send the object description information to server 102 through the application. Server 102 is used to receive the object description information sent by terminal 101, generate a three-dimensional virtual object in the target three-dimensional space, and send the three-dimensional virtual object to terminal 101 so that terminal 101 can display the three-dimensional virtual object in the target three-dimensional space through the application.
[0049] FIG2 is a flow chart of a method for generating a virtual object provided in an embodiment of the present application. The method is executed by a computer device. As shown in FIG2 , the method includes:
[0050] 201. A computer device obtains object description information, where the object description information is used to describe a three-dimensional virtual object to be generated.
[0051] In an embodiment of the present application, the three-dimensional virtual object described by the object description information is the three-dimensional virtual object that the user wants to obtain. The user can configure the object description information according to his or her ideas so that the computer device can generate the three-dimensional virtual object described by the object description information based on the object description information, thereby realizing a method of automatically generating three-dimensional virtual objects.
[0052] The object description information can be any type of information, for example, an image, text, or point cloud data. Optionally, the object description information is multimodal information, for example, including at least two of the following: an image, text, or point cloud data. The three-dimensional virtual object can be any type of virtual object, for example, a virtual person, a virtual animal, a virtual item, or a virtual building.
[0053] 202. The computer device obtains first features of multiple first spatial points in the target three-dimensional space based on the object description information. The first features of the first spatial points are used to characterize the color of the first spatial points in the target three-dimensional space and the positional relationship with the three-dimensional virtual object. The first spatial points are evenly distributed in the target three-dimensional space.
[0054] In an embodiment of the present application, the target three-dimensional space is a space used to generate a three-dimensional virtual object. The target three-dimensional space includes multiple first spatial points. The object description information is used to obtain the first feature of each first spatial point in the target three-dimensional space. The first feature of the first spatial point is equivalent to the feature of the first spatial point in the target three-dimensional space when the three-dimensional virtual object is generated in the target three-dimensional space according to the object description information. It can reflect the texture of the first spatial point or the positional relationship with the three-dimensional virtual object. That is, the first feature of the first spatial point is used to characterize the color of the first spatial point in the target three-dimensional space and the positional relationship with the three-dimensional virtual object.
[0055] Among them, the first spatial point is any spatial point in the target three-dimensional space, and multiple first spatial points are evenly distributed in the target three-dimensional space. That is, according to the distribution positions of the multiple first spatial points, the size of the target three-dimensional space can be reflected by the multiple first spatial points, that is, the multiple first spatial points can represent the target three-dimensional space.
[0056] It should be noted that the embodiment of the present application is illustrated by taking the example of obtaining the first feature of the first spatial point based on object description information, while in another embodiment, the process of obtaining the first feature of the first spatial point includes: based on the object description information, updating the second features of multiple first spatial points in the target three-dimensional space to obtain the first features of multiple first spatial points, and the second feature is the feature of the first spatial point in the target three-dimensional space.
[0057] In an embodiment of the present application, the second feature of the first spatial point in the target three-dimensional space is updated using object description information, so that the updated first feature of the first spatial point can reflect the texture of the first spatial point or the positional relationship with the three-dimensional virtual object when a three-dimensional virtual object is generated in the target three-dimensional space according to the object description information.
[0058] The second feature of the first spatial point is used to characterize the texture of the first spatial point in the target three-dimensional space or the positional relationship with the three-dimensional virtual object. The second feature can be represented in any form, for example, the second feature is represented in the form of a vector.
[0059] The second feature is an arbitrary feature. For example, the second feature of the first spatial point is a randomly generated feature or an initialized feature for the first spatial point.
[0060] Optionally, the second feature of the first spatial point is used to characterize the color of the first spatial point in the target three-dimensional space or its positional relationship with any three-dimensional virtual object to be generated before the three-dimensional virtual object is generated based on the object description information.
[0061] In an embodiment of the present application, the second feature of the first spatial point is an initialized feature, which refers to the default feature of the first spatial point in the target three-dimensional space before the three-dimensional virtual object is generated based on the object description information. The second feature of the first spatial point can be applicable to any object description information, that is, for any object description information, it can utilize the existing features of multiple first spatial points in the target three-dimensional space (i.e., the second features), and adopt a feature update method to obtain the features of multiple first spatial points (i.e., the first features) that match the object description information, so that the first features of the multiple first spatial points can reflect the color of the first spatial point or the positional relationship with the three-dimensional virtual object when the three-dimensional virtual object is generated in the target three-dimensional space according to the object description information, so as to ensure that the first features of the multiple first spatial points can be used subsequently to generate a three-dimensional virtual object that matches the object description information in the target three-dimensional space, thereby ensuring the accuracy of the subsequent generation of three-dimensional virtual objects.
[0062] The second feature of the first spatial point being a randomly generated feature means that when generating a three-dimensional virtual object in the target three-dimensional space using each object description information, a feature is first randomly generated for the first spatial point in the target three-dimensional space, and then, according to the object description information, a feature update method is adopted to obtain features (i.e., first features) of multiple first spatial points that match the object description information, so that the first features of the multiple first spatial points can reflect the color of the first spatial point or the positional relationship with the three-dimensional virtual object when the three-dimensional virtual object is generated in the target three-dimensional space according to the object description information, thereby ensuring that the first features of the multiple first spatial points can be used to subsequently generate a three-dimensional virtual object that matches the object description information in the target three-dimensional space, thereby ensuring the accuracy of the subsequent generation of the three-dimensional virtual object. In addition, when the second feature of the first spatial point is a randomly generated feature, the second feature of the first spatial point can be obtained according to the above-mentioned randomly generated method for different object description information; or, after the second feature of the first spatial point is obtained according to the above-mentioned randomly generated method, it can be applied to different object description information, that is, the second feature of the first spatial point is only randomly generated once for different object description information.
[0063] 203. The computer device processes the first feature of each first spatial point by rendering the model to obtain the color and directed distance of each first spatial point, where the directed distance indicates the distance from the first spatial point to the surface of the three-dimensional virtual object in the target three-dimensional space.
[0064] In an embodiment of the present application, a rendering model is used to map the features of a spatial point into the color and directed distance of the spatial point. After the features of each first spatial point in the target three-dimensional space are updated using the object description information, the first feature of the first spatial point is the feature of the first spatial point in the target three-dimensional space when a three-dimensional virtual object is generated in the target three-dimensional space according to the object description information. The first feature of the first spatial point can reflect the texture of the first spatial point or the positional relationship with the three-dimensional virtual object. Therefore, the first feature of each first spatial point is processed separately by the rendering model to obtain the color and directed distance of each first spatial point.
[0065] Among them, the color of the first spatial point represents the color of the first spatial point when the three-dimensional virtual object is generated in the target three-dimensional space according to the object description information. The color of the first spatial point can be represented in any form. For example, the color of the first spatial point is represented in the form of multiple channels, such as the color of the first spatial point is represented in the form of RGB (Red Green Blue). For example, if the color of the first spatial point is (255, 0, 0), it means that the green value and blue value of the first spatial point are 0, and the red value of the first spatial point is 255, then the color of the first spatial point is red; or, the color of the first spatial point is represented in the form of HSB (Hue Saturation Brightness). The directed distance of the first spatial point represents the distance between the first spatial point and the surface of the three-dimensional virtual object when the three-dimensional virtual object is generated in the target three-dimensional space according to the object description information. In an embodiment of the present application, the directed distances of multiple first spatial points can constitute a directed distance field, and the directed distance field can indicate the position and shape of the three-dimensional virtual object in the target three-dimensional space.
[0066] 204. The computer device generates a three-dimensional virtual object in the target three-dimensional space based on the colors and directed distances of the multiple first spatial points.
[0067] In an embodiment of the present application, the directed distances of multiple first spatial points indicate the position and form of the three-dimensional virtual object in the target three-dimensional space. Based on the directed distances of multiple first spatial points, the outline of the three-dimensional virtual object can be outlined in the target three-dimensional space, and the outline of the three-dimensional virtual object can be filled with color according to the color of the first spatial points to obtain a three-dimensional virtual object containing color.
[0068] In the solution provided in the embodiment of the present application, object description information is used to describe the three-dimensional virtual object to be generated. The object description information is used to obtain the characteristics of each spatial point in the target three-dimensional space, and then the characteristics of each spatial point are processed through a rendering model to obtain the color and directed distance of each spatial point. Based on the color and directed distance of each spatial point, a three-dimensional virtual object can be generated in the target three-dimensional space. In this way, the shape or color of the generated three-dimensional virtual object is the same as the three-dimensional virtual object described by the object description information, thereby ensuring the accuracy of the three-dimensional virtual object, and realizing a method of automatically generating three-dimensional virtual objects using object description information. There is no need to manually develop three-dimensional virtual objects, thereby improving the efficiency of generating three-dimensional virtual objects.
[0069] Based on the embodiment shown in Figure 2, the embodiment of the present application can also use object description information to adopt fusion and denoising methods to update the features of spatial points in the target three-dimensional space, and can also use the directed distance of each spatial point in the target three-dimensional space to screen out spatial points located on the surface of the three-dimensional virtual object, and then generate a three-dimensional virtual object in the target three-dimensional space. The specific process is detailed in the following embodiment.
[0070] FIG3 is a flow chart of another method for generating a virtual object provided by an embodiment of the present application. The method is executed by a computer device. As shown in FIG3 , the method includes:
[0071] 301. A computer device obtains object description information, where the object description information is used to describe a three-dimensional virtual object to be generated.
[0072] In one possible implementation, the object description information includes at least one of an image, text, or point cloud data.
[0073] In the embodiments of the present application, the object description information can include any one of images, text, or point cloud data, or can also include any two or three of these. When the object description information includes two or three of these, the object description information is equivalent to multimodal information, that is, the multimodal information is used to describe the 3D virtual object to be generated, so that the multimodal information can be used to control the shape or color of the generated 3D virtual object.
[0074] The point cloud data includes multiple spatial points in three-dimensional space, and the spatial points in the point cloud data can represent the geometric shape of the virtual object. Optionally, the point cloud data also includes the color of each spatial point, so that the point cloud data can not only represent the geometric shape of the virtual object but also the color of the virtual object. In the embodiment of the present application, the point cloud data presents the virtual object in a display expression manner, and the point cloud data is shown in Figure 4.
[0075] In one possible implementation, the generated three-dimensional virtual object can be used in a game as a three-dimensional virtual object in the game. In embodiments of the present application, the three-dimensional virtual object can be a virtual character, virtual item, or virtual scene in the game. In the game, the three-dimensional virtual object is composed of three-dimensional geometric patches and material maps. The three-dimensional geometric patches can form the outline of the three-dimensional virtual object. By adding a texture map to the outline of the three-dimensional virtual object, a colored three-dimensional virtual object can be presented.
[0076] 302. The computer device extracts features from the object description information to obtain object description features.
[0077] The object description feature is used to represent the object description information. The object description feature can be represented in any form. For example, the object description feature is represented in the form of a feature vector.
[0078] 303. The computer device randomly generates a second feature for each first spatial point in the target three-dimensional space, where the second feature is a feature of the first spatial point in the target three-dimensional space, and the first spatial points are evenly distributed in the target three-dimensional space.
[0079] In an embodiment of the present application, a second feature is randomly generated for each first spatial point in the target three-dimensional space, so that multiple first spatial points in the target three-dimensional space have a second feature, and the second features of multiple first spatial points may be the same or different.
[0080] In one possible implementation, the process of randomly generating the second feature for the first spatial point includes: randomly determining a color and a directed distance for the first spatial point, performing feature extraction on the randomly determined color and directed distance of the first spatial point, and obtaining the second feature.
[0081] In the embodiments of the present application, the randomly determined color for a first spatial point refers to the color assigned to the first spatial point. The randomly determined directed distance for the first spatial point refers to the distance from the first spatial point to the surface of the three-dimensional virtual object to be generated in the target three-dimensional space. By randomly determining the color and directed distance for the first spatial point, feature extraction is performed to obtain a second feature, which is the randomly determined feature for the first spatial point.
[0082] In one possible implementation, the second feature includes multiple sub-features, and the resolutions of the multiple sub-features are different. Then step 303 includes: randomly generating sub-features of multiple resolutions for each first spatial point, and forming the second feature of the first spatial point with the sub-features of the multiple resolutions of each first spatial point.
[0083] In an embodiment of the present application, sub-features of different resolutions have different dimensions and contain different information. Using sub-features of different resolutions to represent the features of the first spatial point in the target three-dimensional space can enrich the features of the first spatial point and ensure the accuracy of the features of the first spatial point, so as to ensure the accuracy of the subsequent generation of three-dimensional virtual objects.
[0084] The resolution of the sub-feature indicates the granularity of dividing the target three-dimensional space when obtaining the sub-feature of the first spatial point. When the target three-dimensional space is divided according to different granularities, the sizes of the obtained three-dimensional grids are different.
[0085] For example, when randomly generating multiple sub-features for the first spatial point, the target three-dimensional space is divided into multiple three-dimensional spaces of different granularities, and each three-dimensional space of different granularity is divided into n 3Three-dimensional grids, where n is a positive integer and n is different for three-dimensional spaces of different granularities. The three-dimensional grid to which the first spatial point belongs in the three-dimensional space of each granularity is determined, and a sub-feature of the first spatial point is obtained using the spatial points included in the three-dimensional network to which the first spatial point belongs. Then, multiple sub-features of the first spatial point are obtained using the three-dimensional grids to which the first spatial point belongs in three-dimensional spaces of different granularities, thereby obtaining a second feature of the first spatial point.
[0086] Optionally, if each spatial point in the target three-dimensional space has a default feature, the process of obtaining a sub-feature of the first spatial point includes: for any three-dimensional grid in the three-dimensional space of any granularity, determining the first spatial point contained in the three-dimensional grid, calling any spatial point among the first spatial points contained in the three-dimensional grid the sixth spatial point, and based on the similarity between the default feature of the sixth spatial point and the default feature of each first spatial point contained in the three-dimensional grid, performing a weighted summation on the default feature of each first spatial point contained in the three-dimensional grid to obtain a sub-feature of the sixth spatial point.
[0087] Among them, the default feature refers to an initial feature of a spatial point in the target three-dimensional space, and the default features of different spatial points may be the same or different. The sixth spatial point is a first spatial point in the target three-dimensional space. Optionally, the weighted summation process includes: for each first spatial point contained in the three-dimensional grid, determining the product of the similarity corresponding to the first spatial point and the default feature of the first spatial point, and determining the sum of the products corresponding to multiple first spatial points contained in the three-dimensional grid as a sub-feature of the sixth spatial point. Among them, the similarity corresponding to the first spatial point is the similarity between the default feature of the first spatial point and the default feature of the sixth spatial point. In the above manner, based on the three-dimensional space of each granularity and the default feature of each first spatial point, the second feature of each first spatial point can be obtained.
[0088] In one possible implementation, determining the first spatial point includes uniformly sampling spatial points in the target three-dimensional space to obtain a plurality of first spatial points, wherein the distance between each two adjacent first spatial points in the plurality of first spatial points is equal. In other words, uniformly sampling the target three-dimensional space to obtain the plurality of first spatial points.
[0089] In an embodiment of the present application, due to the excessive number of spatial points in the target three-dimensional space, the spatial points in the target three-dimensional space are uniformly sampled so that the first spatial points obtained by sampling can be evenly distributed in the target three-dimensional space points, and multiple first spatial points can represent the target three-dimensional space points to ensure the efficiency of subsequent three-dimensional virtual objects.
[0090] 304. The computer device fuses the second feature and the object description feature of each first spatial point to obtain a fused feature of each first spatial point.
[0091] In the embodiment of the present application, the second feature of the first spatial point and the object description feature are fused in a manner such as superposition or splicing.
[0092] 305. The computer device denoises the fused features of each first spatial point to obtain a first feature of each first spatial point.
[0093] In an embodiment of the present application, since the second feature of the first spatial point is a feature randomly generated for the first spatial point, the object description feature is used to characterize the object description information. After the second feature of the first spatial point and the object description feature are fused, the fused feature of the first spatial point is denoised so that under the influence of the object description feature, the feature of the first spatial point is guided to change continuously, so that the first feature of the first spatial point matches the object description information, thereby ensuring the accuracy of the first feature.
[0094] In one possible implementation, the second feature of the first spatial point is updated based on the object description information through a diffusion model, that is, the process of obtaining the first feature of the first spatial point includes: extracting features of the object description information through a diffusion model to obtain object description features; randomly generating second features for each first spatial point in the target three-dimensional space through a diffusion model; fusing the second feature and the object description feature of each first spatial point through a diffusion model to obtain a fused feature of each first spatial point; and denoising the fused feature of each first spatial point through a diffusion model to obtain the first feature of each first spatial point.
[0095] The diffusion model is used to update the features of the spatial points using the object description information. The diffusion model can be any model, for example, the diffusion model is a diffusion macro model.
[0096] 306. The computer device processes the first feature of each first spatial point by rendering the model to obtain the color and directed distance of each first spatial point, where the directed distance indicates the distance from the first spatial point to the surface of the three-dimensional virtual object in the target three-dimensional space.
[0097] In one possible implementation, each first feature includes multiple sub-features, and the resolutions of the multiple sub-features are different; step 306 includes: fusing the multiple sub-features of the first spatial point through a rendering model, processing the fused features, and obtaining the color and directed distance of the first spatial point.
[0098] In an embodiment of the present application, sub-features of different resolutions have different dimensions and contain different information. Using sub-features of different resolutions to represent the features of the first spatial point in the target three-dimensional space can enrich the features of the first spatial point and ensure the accuracy of the features of the first spatial point. Then, through the rendering model, multiple sub-features of the first spatial point are processed to ensure the accuracy of the color and directed distance of the obtained first spatial point, thereby ensuring the accuracy of the subsequent generation of three-dimensional virtual objects.
[0099] Optionally, the process of processing the fused features includes: decoding the fused features to obtain the color and directed distance of the first spatial point.
[0100] In an embodiment of the present application, the rendering model has the function of mapping the features of a spatial point into the color and directed distance of the spatial point. After the first feature of the first spatial point is input into the rendering model, the rendering model first fuses the multiple sub-features contained in the first feature, and then is able to decode the fused features so as to decode the fused features into the color and directed distance of the first spatial point.
[0101] Optionally, the process of fusing the multiple sub-features of the first spatial point includes: performing dimensional transformation on the multiple sub-features of the first spatial point so that the dimensions of the multiple sub-features after the transformation are the same, and fusing the multiple sub-features after the transformation.
[0102] In an embodiment of the present application, sub-features of different resolutions have different dimensions. By changing the dimensions of multiple sub-features so that the dimensions of the transformed multiple sub-features are the same, the multiple sub-features can be fused to ensure the accuracy of the fused features.
[0103] 307. The computer device determines a second spatial point from the multiple first spatial points based on the directed distance of the multiple first spatial points, where the second spatial point is located on the surface of the three-dimensional virtual object.
[0104] In an embodiment of the present application, when a three-dimensional virtual object is generated in a target three-dimensional space according to object description information, the directed distance of the first spatial point can reflect whether the first spatial point is located on the surface of the three-dimensional virtual object. Through the directed distances of multiple first spatial points, the spatial point located on the surface of the three-dimensional virtual object can be determined, so that the three-dimensional virtual object can be generated subsequently using the spatial point located on the surface of the three-dimensional virtual object.
[0105] The number of the second spatial points is one or more.
[0106] In an embodiment of the present application, the signed distances of multiple first spatial points can form a signed distance field, which is used to implicitly describe the three-dimensional virtual object to be generated in the target three-dimensional scene. The positive or negative sign of the signed distance of the first spatial point indicates whether the first spatial point is outside or inside the three-dimensional virtual object.
[0107] As shown in Figure 5, each numerical value in Figure 5 corresponds to a first spatial point and is used to represent the distance between the first spatial point and the surface of the 3D virtual object to be generated in the target 3D scene. If the first spatial point is located on the surface of the 3D virtual object, the directed distance of the first spatial point is 0; if the first spatial point is located outside the 3D virtual object, the directed distance of the first spatial point is greater than 0; if the first spatial point is located inside the 3D virtual object, the directed distance of the first spatial point is less than 0. Using the directed distances of multiple first spatial points, a spatial point with a directed distance of 0 can be selected from the multiple first spatial points to serve as the second spatial point.
[0108] 308. The computer device connects the second spatial points in the target three-dimensional space to form a model of the three-dimensional virtual object.
[0109] In an embodiment of the present application, when a three-dimensional virtual object is generated in the target three-dimensional space according to the object description information, the second spatial point in the target three-dimensional space is located on the surface of the three-dimensional virtual object. By connecting the second spatial point, the outline of the three-dimensional virtual object can be outlined, that is, the model of the three-dimensional virtual object.
[0110] In a possible implementation, the model of the three-dimensional virtual object is composed of multiple geometric patches.
[0111] The geometric mesh can be a triangle or a quadrilateral, and each geometric mesh is formed by connecting three or more second space points.
[0112] For example, the model of the three-dimensional virtual object is shown in FIG6 and FIG7 . The model of the three-dimensional virtual object is a three-dimensional structure composed of triangular or quadrilateral geometric facets.
[0113] Optionally, the process of connecting the plurality of second space points in the target three-dimensional space includes: connecting each second space point with an adjacent second space point to obtain a model of the three-dimensional virtual object.
[0114] In the embodiment of the present application, the second space point is connected with adjacent second space points, and then the multiple second space points and the connections between the multiple second space points can constitute a model of a three-dimensional virtual object.
[0115] 309. The computer device renders the model of the three-dimensional virtual object based on the color of the second spatial point to obtain a three-dimensional virtual object including color.
[0116] In an embodiment of the present application, the second spatial point in the target three-dimensional space is located on the surface of the three-dimensional virtual object, and the color of the second spatial point can reflect the color of the surface of the three-dimensional virtual object. Therefore, based on the color of the second spatial point, the model of the three-dimensional virtual object is rendered to obtain a three-dimensional virtual object containing color, so that the obtained three-dimensional virtual object matches the object description information, which not only reflects the shape of the three-dimensional virtual object described by the object description information, but also reflects the color of the three-dimensional virtual object.
[0117] In an embodiment of the present application, the directed distances of multiple first spatial points can constitute a directed distance field, through which the position and shape of the three-dimensional virtual object in the target three-dimensional space can be indicated. The second spatial point in the target three-dimensional space is located on the surface of the three-dimensional virtual object, and the color of the second spatial point can reflect the color of the surface of the three-dimensional virtual object. Therefore, the spatial point located on the surface of the three-dimensional virtual object is determined by the directed distances of multiple first spatial points, so that the three-dimensional virtual object can be subsequently generated using the spatial point located on the surface of the three-dimensional virtual object. By connecting the second spatial points, the outline of the three-dimensional virtual object can be outlined, and then the model of the three-dimensional virtual object is rendered based on the color of the second spatial point to obtain a three-dimensional virtual object containing color, so that the obtained three-dimensional virtual object matches the object description information, which not only reflects the shape of the three-dimensional virtual object described by the object description information, but also reflects the color of the three-dimensional virtual object, thereby ensuring the accuracy of the three-dimensional virtual object.
[0118] In one possible implementation, the model of the three-dimensional virtual object is composed of multiple geometric patches, and step 309 includes: for each geometric patch in the model of the three-dimensional virtual object, based on the color of the second spatial point contained in the geometric patch, determining the texture map corresponding to the geometric patch, and covering each geometric patch contained in the model of the three-dimensional virtual object with the corresponding texture map to obtain a three-dimensional virtual object containing color.
[0119] Among them, the texture map contains the color texture of the corresponding geometric surface. By covering the corresponding texture map on each geometric surface, a three-dimensional virtual object containing color can be obtained to ensure that the obtained three-dimensional virtual object is a closed space, thereby improving the display effect of the three-dimensional virtual object.
[0120] Optionally, the second spatial point contained in the geometric patch is the vertex of the geometric patch; then the method of determining the texture map corresponding to the geometric patch includes: based on the color of the vertices of the geometric patch, using an interpolation algorithm to determine the color of each pixel point in the geometric patch, mapping the color of each pixel point in the geometric patch to the texture space to obtain a texture map.
[0121] In the embodiment of the present application, each geometric patch is formed by connecting three or more second spatial points, that is, the second spatial points that constitute the geometric patch are the vertices of the geometric patch. Texture space is a two-dimensional or three-dimensional coordinate space used to define and process textures. Using an interpolation algorithm, the color of the pixels in the geometric patch can be determined according to the color of the vertices, and the points in the geometric patch can be mapped to the texture space to obtain a texture map, thereby improving the display effect of the texture map and ensuring that the surface details of the subsequent three-dimensional virtual object are displayed more effectively and realistically.
[0122] Optionally, in a game scenario, geometric patches and texture maps can be used as game resources and configured in the game to enable game development.
[0123] It should be noted that the embodiment of the present application is described by taking the example of generating a texture map for each geometric facet. In another embodiment, a texture map can be generated for the model of the three-dimensional virtual object based on the color of the second spatial point, and then the texture map of the model of the three-dimensional virtual object can be overlaid on the model of the three-dimensional virtual object to obtain a three-dimensional virtual object containing color.
[0124] It should be noted that the embodiment of the present application selects a second spatial point from multiple first spatial points, and then uses the color and directed distance of the second spatial point to generate a three-dimensional virtual object. In another embodiment, there is no need to perform the above steps 307-309, but other methods are adopted to generate a three-dimensional virtual object in the target three-dimensional space based on the color and directed distance of multiple first spatial points.
[0125] In the solution provided in the embodiment of the present application, object description information is used to describe the three-dimensional virtual object to be generated. The object description information is used to obtain the characteristics of each spatial point in the target three-dimensional space, and then the characteristics of each spatial point are processed through a rendering model to obtain the color and directed distance of each spatial point. Based on the color and directed distance of each spatial point, a three-dimensional virtual object can be generated in the target three-dimensional space. In this way, the shape or color of the generated three-dimensional virtual object is the same as the three-dimensional virtual object described by the object description information, thereby ensuring the accuracy of the three-dimensional virtual object, and realizing a method of automatically generating three-dimensional virtual objects using object description information. There is no need to manually develop three-dimensional virtual objects, thereby improving the efficiency of generating three-dimensional virtual objects.
[0126] In an embodiment of the present application, the directed distances of multiple first spatial points can form a directed distance field, which is used to describe the three-dimensional virtual object to be generated in the target three-dimensional scene to ensure that the generated three-dimensional virtual object is smooth and complete. In addition, the object description information can be at least one of an image, text, or point cloud data, which can achieve fine control over the generated three-dimensional virtual object and ensure the display effect of the generated three-dimensional virtual object. In addition, the three-dimensional virtual objects and texture maps of the three-dimensional virtual objects provided in the embodiment of the present application can be used in games, which can improve the efficiency of game development.
[0127] It should be noted that the embodiment shown in Figure 3 above adopts a fusion and denoising method to update the second feature of the first spatial point using the object description information. In another embodiment, there is no need to perform the above steps 302-305, but other methods are adopted to update the second features of multiple first spatial points in the target three-dimensional space based on the object description information to obtain the first features of multiple first spatial points.
[0128] In one possible implementation, the process of obtaining the first feature includes: updating the second features of multiple first spatial points in the target three-dimensional space based on object description information through a diffusion model to obtain the first features of the multiple first spatial points.
[0129] In one possible implementation, the object description information includes at least two items of image, text, or point cloud data; then, first features of multiple first spatial points are obtained, including: extracting features of each sub-information in the object description information to obtain features of each sub-information, where the sub-information is any one of image, text, or point cloud data; fusing features of multiple sub-information in the object description information to obtain object description features; and updating second features of multiple first spatial points based on the object description features to obtain first features of multiple first spatial points.
[0130] In an embodiment of the present application, when the object description information includes at least two items of image, text or point cloud data, the object description information is multimodal information. The multimodal information is used to update the features of the first spatial point to obtain the first feature of the first spatial point, and then the first feature of the first spatial point is subsequently used to generate a three-dimensional virtual object, so as to realize the scheme of controlling the generation of three-dimensional virtual objects using multiple types of description information, thereby improving the accuracy of generating three-dimensional virtual objects.
[0131] Optionally, the process of updating the second features of the plurality of first spatial points based on the object description features can be implemented according to the above steps 304 - 305 .
[0132] Optionally, when the object description information is multimodal description information, the features of each sub-information in the object description information can be updated through the self-attention mechanism. That is, the process of obtaining the object description features includes: updating the features of each sub-information through the self-attention mechanism, and fusing the updated features of each sub-information to obtain the object description features.
[0133] In an embodiment of the present application, a self-attention mechanism is adopted to update the features of each sub-information in the object description information, so as to map the features of different types of sub-information into the same feature space, and then the features of different sub-information can be fused into object description features to ensure the accuracy of the object description features, so as to better realize the generation of three-dimensional virtual objects with multimodal information, thereby ensuring the effect of generating three-dimensional virtual objects.
[0134] Optionally, for any sub-information in the object description information, the following relationship is adopted to update the characteristics of the sub-information:
[0135] Among them, Attention(Q,K,V) is used to represent the updated features of the sub-information, Q, K, V are used to represent the features of the sub-information respectively, T is used to represent the transposition of the features, d k It is used to represent the dimension of the features of sub-information, and softmax(·) is used to represent the normalization function.
[0136] Based on the embodiment shown in FIG3 , before using the rendering model to generate a three-dimensional virtual object, the rendering model needs to be trained. The training process is detailed in the following embodiment.
[0137] FIG8 is a flowchart of another method for generating a virtual object provided in an embodiment of the present application. The method is executed by a computer device. As shown in FIG8 , the method includes:
[0138] 801. The computer device photographs a sample virtual object in the sample three-dimensional space based on a virtual camera in the sample three-dimensional space to obtain a sample image.
[0139] In an embodiment of the present application, the sample three-dimensional space includes a sample virtual object and a virtual camera. The virtual camera is used to capture images of the virtual objects in the sample three-dimensional space. Through the virtual camera, the sample virtual objects in the sample three-dimensional space can be captured to obtain a sample image containing the sample virtual object.
[0140] The sample 3D space is a 3D space that is the same as or different from the target 3D space. The sample virtual object is a 3D virtual object in the sample 3D space. The sample virtual object can be any object, such as a virtual person, virtual animal, virtual object, or virtual building. The sample image is an image of the sample virtual object from any perspective, such as a frontal image or a side image of the sample virtual object.
[0141] In one possible implementation, the position of the virtual camera in the sample three-dimensional space can be adjusted arbitrarily. By adjusting the position of the virtual camera in the sample three-dimensional space, the virtual camera can be used to photograph the sample virtual object at any position to obtain a sample image.
[0142] In the embodiment of the present application, the sample images captured by the virtual camera at different positions of the sample image object are different. For example, if the virtual camera captures the sample virtual object from the front, the sample image obtained is the front view of the sample virtual object.
[0143] In one possible implementation, the sample virtual object in the sample three-dimensional space is game data.
[0144] For example, the sample virtual objects in the sample three-dimensional space are virtual objects manually developed by developers. The sample virtual objects are used as samples for training the rendering model so that the subsequent rendering model can learn the ability to map the features of spatial points in the three-dimensional space into the color and directed distance of the spatial points.
[0145] It should be noted that the present embodiment of the present application uses the example of a local device photographing a sample virtual object in a sample three-dimensional space using a virtual camera. However, in another embodiment, step 801 need not be performed, and a sample image may be obtained by other means. The sample image is obtained by photographing the sample virtual object in the sample three-dimensional space using a virtual camera. For example, after another device photographs the sample virtual object in the sample three-dimensional space using a virtual camera to obtain a sample image, it transmits the sample image to the local device, and the local device receives the sample image.
[0146] 802. The computer device determines a third space point corresponding to each pixel point in the sample three-dimensional space based on the position point of the virtual camera in the sample three-dimensional space and the position point of each pixel point in the sample image in the sample three-dimensional space.
[0147] In the embodiments of the present application, based on the imaging principle of a camera, when a virtual camera captures a sample virtual object in a sample 3D space, the resulting sample image is located between the virtual camera and the sample virtual object in the sample 3D space. Each pixel in the sample image is located at a spatial point in the sample 3D space, and the sample image is perpendicular to the shooting angle of the virtual camera. The third spatial point corresponding to a pixel is the spatial point that is mapped to the sample image when the virtual camera captures the sample virtual object.
[0148] A pixel corresponds to one or more third-space points. If a pixel corresponds to a single third-space point, the color of the third-space point is the same as the color of the pixel in the sample image. If a pixel corresponds to multiple third-space points, the color of the fused third-space points is the same as the color of the pixel in the sample image. The position of the virtual camera in the sample 3D space refers to the position of the virtual camera in the sample 3D space and can be represented by coordinates in the sample 3D space.
[0149] In one possible implementation, step 802 includes: determining, in the sample 3D space, a ray starting from the virtual camera's position and passing through the pixel's position; and collecting, along the ray's direction, at least one third-space point. That is, in the sample 3D space, determining, with the virtual camera's position as the starting point and passing through the pixel's position; and collecting, along the ray's direction, at least one third-space point.
[0150] In an embodiment of the present application, a sample image is acquired based on the imaging principle of a camera. According to the imaging principle of a camera, a ray is emitted by a virtual camera, and an image is formed in the sample image at the intersection of the ray and the surface of the sample virtual object. Therefore, in the three-dimensional space of the sample, a ray is determined that starts at the position point of the virtual camera and passes through the position point of the pixel point.
[0151] The pixel position refers to the location of the pixel in the sample 3D space. In the embodiment of the present application, based on the imaging principle of the camera, when a virtual camera captures a sample virtual object in the sample 3D space, the resulting sample image is located between the virtual camera and the sample virtual object in the sample 3D space. Each pixel in the sample image is located at a spatial point in the sample 3D space. The sample pixel position is equivalent to the position of the spatial point where the pixel is located in the sample 3D space.
[0152] In an embodiment of the present application, in the sample three-dimensional space, a ray starting from the position point of the virtual camera and passing through the position point of the pixel intersects with the surface of the sample virtual object. On the ray, the point between the intersection point of the ray with the surface of the sample virtual object and the pixel point can be mapped to the sample image. Then, the color after the intersection point of the ray with the surface of the sample virtual object and the point between the pixel point is fused with the color of the pixel point. Therefore, starting from the position point of the pixel point, at least one third space point is collected on the ray along the direction of the ray, that is, at least one third space point corresponding to the pixel point is obtained, thereby ensuring the accuracy of the collected third space point.
[0153] According to the above method, a plurality of rays can be constructed using the position point of the virtual camera and the position point of each pixel point in the sample three-dimensional space, and the third space point of the corresponding pixel point can be collected from each ray.
[0154] In the embodiment of the present application, based on the imaging principle of the camera, a ray-constructing method is adopted to obtain the third spatial point corresponding to the pixel point in the sample three-dimensional space to ensure the accuracy of the collected third spatial point.
[0155] As shown in Figure 9, a front image 901 and a side image 902 of the sample virtual object are obtained in the sample three-dimensional space, and multiple third space points 905 corresponding to the pixel point 904 in the front image 901 can be collected from the ray 903, and multiple third space points 908 corresponding to the pixel point 907 in the side image 902 can be collected from the ray 906.
[0156] Optionally, the third spatial point corresponding to the pixel point satisfies the following relationship: p(t)=o+tv
[0157] Among them, p(t) is used to represent the third space point of the pixel point, o is used for the position point of the virtual camera in the sample three-dimensional space, t is used to represent an arbitrary distance, t is greater than the distance between the position point of the virtual camera and the position point of the pixel point in the sample three-dimensional space, and v is used to represent the direction of the ray, which is a ray that starts from the position point of the virtual camera and passes through the position point of the pixel point.
[0158] 803. The computer device obtains a sample directed distance of each third space point based on the sample virtual object, where the sample directed distance indicates a distance from the third space point to the surface of the sample virtual object.
[0159] The sample directed distance of the third spatial point is also the directed distance of the third spatial point, so as to indicate the distance from the third spatial point to the surface of the sample virtual object.
[0160] In an embodiment of the present application, the sample three-dimensional space includes a sample virtual object. Based on the sample virtual object in the sample three-dimensional space, the positional relationship between each third space point and the sample virtual object can be determined, that is, whether the third position point is located on the surface of the sample virtual object, or the distance from the surface of the sample virtual object, and then the sample directed distance of each third space point can be determined.
[0161] In one possible implementation, step 803 satisfies the following relationship:
[0162] Among them, f(x) is used to represent the sample directed distance of the third space point x, d is used to represent the scale of the directed distance, and Ω is used to represent the set of spatial points contained in the three-dimensional space occupied by the sample virtual object in the sample three-dimensional space. used to represent the surface of the sample virtual object, It is used to represent the sample directed distance of the third space point x when it belongs to the set Ω; Ω c It is used to represent the set of spatial points in the sample three-dimensional space other than the set Ω, that is, the spatial points other than the sample virtual object in the sample three-dimensional space; Used to indicate that the third space point x belongs to the set Ω c When , the sample of the third space point x has a signed distance. In the embodiment of the present application, the gradient of the signed distance field is The sample virtual object is made to have a smooth surface in the sample three-dimensional space.
[0163] It should be noted that the embodiment of the present application uses pixel points in the sample image captured by the virtual camera to determine the third spatial point from the sample three-dimensional space, and then obtain the sample directed distance of the third spatial point. In another embodiment, there is no need to perform the above steps 801-803, but other methods are adopted to obtain the sample directed distances of multiple third spatial points in the sample three-dimensional space based on the sample virtual object in the sample three-dimensional space.
[0164] 804. The computer device extracts features of each third space point from the sample three-dimensional space.
[0165] The feature of the third spatial point is used to characterize the texture of the third spatial point in the target three-dimensional space or its positional relationship with the sample virtual object. In the embodiment of the present application, any third spatial point may be located outside the sample virtual object or on the surface of the sample virtual object.
[0166] In one possible implementation, the feature of the third spatial point includes multiple sub-features, and the resolutions of the multiple sub-features are different. Then step 804 includes: extracting sub-features of multiple resolutions of the third spatial point from the sample three-dimensional space, and forming the feature of the third spatial point with the sub-features of multiple resolutions of the third spatial point.
[0167] In an embodiment of the present application, sub-features of different resolutions have different dimensions and contain different information. Using sub-features of different resolutions to represent the features of the third spatial point in the three-dimensional space of the sample can enrich the features of the third spatial point and ensure the accuracy of the features of the third spatial point.
[0168] Optionally, when features of multiple third spatial points are obtained, the features of the multiple third spatial points are stored in the form of a hash table, that is, the hash table includes the correspondence between the third spatial points and the features, so that when the rendering model is subsequently trained, the features of the third spatial points can be directly indexed from the hash table.
[0169] Optionally, the hash table includes the coordinates of the third spatial point and the corresponding multiple sub-features. Based on the coordinates of the third spatial point, the corresponding multiple sub-features can be indexed from the hash table to ensure the indexing speed. The indexing speed is fast enough so that the rendering model can directly process the features of the indexed third spatial point to ensure the processing efficiency of the rendering model.
[0170] 805. The computer device processes the features of each third space point through the rendering model to obtain the predicted color and predicted directed distance of each third space point.
[0171] The predicted color of the third spatial point refers to the predicted color of the third spatial point. The color of the third spatial point can be the same as the color of the first spatial point in the above step 203, and will not be repeated here.
[0172] The step 805 is similar to the above step 306 and will not be described again here.
[0173] In the implementation of this application, the process of obtaining the predicted color and predicted directed distance of the third spatial point is shown in Figure 10. Through the coordinates of the third spatial point, the sub-features of multiple resolutions of each third spatial point are extracted from the sample three-dimensional space, and the sub-features of multiple resolutions of each third spatial point are stored in a hash table. Through the rendering model, based on the coordinates of the third spatial point, the sub-features of multiple resolutions of the third spatial point can be quickly indexed from the hash table, and then according to the above step 805, the predicted color and predicted directed distance of the third spatial point are obtained.
[0174] 806. The computer device fuses the predicted colors of the third space points corresponding to the same pixel to obtain the predicted color of each pixel.
[0175] In an embodiment of the present application, the third spatial point corresponding to the pixel point is the spatial point mapped to the sample image that constitutes the pixel point when the sample virtual object is photographed based on the virtual camera. The predicted color of the third spatial point is obtained through the rendering model. The predicted colors of the third spatial points corresponding to the same pixel point are fused as the predicted color of the pixel point, that is, the predicted color of the pixel point can reflect the accuracy of the rendering model.
[0176] In one possible implementation, step 806 includes: determining the transparency of multiple fourth spatial points based on the sample directed distances of multiple fourth spatial points, the transparency is positively correlated with the sample directed distances, the multiple fourth spatial points correspond to the same pixel point, and the fourth spatial point is any one of the third spatial points corresponding to the pixel point; based on the transparency of the multiple fourth spatial points, the predicted colors of the multiple fourth spatial points are fused to obtain the predicted color of the pixel point.
[0177] In an embodiment of the present application, if the fourth spatial point is located on the surface of the three-dimensional virtual object, the directed distance of the fourth spatial point is 0; if the fourth spatial point is located outside the three-dimensional virtual object, the directed distance of the fourth spatial point is greater than 0; if the fourth spatial point is located inside the three-dimensional virtual object, the directed distance of the fourth spatial point is less than 0. The transparency of the fourth spatial point is positively correlated with the sample directed distance, so the transparency of the spatial point located outside the three-dimensional virtual object is large, and the transparency of the spatial point located inside the three-dimensional virtual object is small, so that the spatial point located outside the three-dimensional virtual object will not block the surface of the three-dimensional virtual object. The transparency of the fourth spatial point is used as a weight, and the predicted colors of multiple fourth spatial points corresponding to the same pixel point are fused to obtain the predicted color of the pixel point, so as to ensure that the obtained predicted color is as accurate as possible, avoid the situation where the training rendering model is affected by poor fusion effect, and enable the predicted color to accurately reflect the accuracy of the rendering model.
[0178] Optionally, the process of determining the predicted color of a pixel point includes: determining the weight of each fourth spatial point based on the transparency of multiple fourth spatial points, and fusing the predicted colors of multiple fourth spatial points based on the weights of multiple fourth spatial points to obtain the predicted color of the pixel point.
[0179] In an embodiment of the present application, the transparency of a fourth spatial point can be used as the weight of each fourth spatial point; alternatively, the transparency of multiple fourth spatial points can be normalized, and the resulting value can be used as the weight of the fourth spatial point. The process of fusing the predicted colors of multiple fourth spatial points includes: multiplying the predicted color of each fourth spatial point by the corresponding weight to obtain the product corresponding to each fourth spatial point, and summing the products corresponding to the multiple fourth spatial points to determine the predicted color of the pixel.
[0180] Optionally, the predicted color is represented in the form of multiple channels, and the process of determining the predicted color of the pixel point includes: for each channel of the predicted color, multiplying the value of the predicted color of each fourth spatial point in the channel by the corresponding weight to obtain the product corresponding to each fourth spatial point, and determining the sum of the products corresponding to multiple fourth spatial points as the value of the predicted color of the pixel point in the channel.
[0181] For example, taking the predicted color expressed in RGB form as an example, the values of the predicted colors of multiple fourth spatial points in the R channel are weightedly fused in the above manner to obtain the value of the predicted color of the pixel point in the R channel; the values of the predicted colors of multiple fourth spatial points in the G channel are weightedly fused in the above manner to obtain the value of the predicted color of the pixel point in the G channel; the values of the predicted colors of multiple fourth spatial points in the B channel are weightedly fused in the above manner to obtain the value of the predicted color of the pixel point in the B channel. The predicted color of the pixel point includes the value of the R channel, the value of the G channel and the value of the B channel.
[0182] Optionally, the transparency of the fourth spatial point is positively correlated with the sample directed distance, that is, the opacity of the fourth spatial point is negatively correlated with the sample directed distance. Then, for multiple fourth spatial points corresponding to any pixel point, the predicted color of the pixel point and the predicted colors of the multiple fourth spatial points satisfy the following relationship: w(t)=T(t)ρ(t)
[0183] Among them, C(o,v) is used to represent the predicted color of the pixel point, w(t) is used to represent the weight of the fourth spatial point, c(p(t),v) is used to represent the predicted color of the fourth spatial point, p(t) is used to represent the fourth spatial point of the pixel point, o is used to represent the position of the virtual camera in the sample three-dimensional space, t is used to represent an arbitrary distance, t is greater than the distance between the position of the virtual camera and the position of the pixel point in the sample three-dimensional space, v is used to represent the direction of the ray, which is the ray starting from the position of the virtual camera and passing through the position of the pixel point; T(t) is used to represent the opacity of the spatial point between the pixel point and the fourth spatial point; ρ(t) is used to represent the opacity function (Opaque Density Function), Φ s It is used to represent the set of spatial points in the sample three-dimensional space, and f(p(t)) is used to represent the sample directed distance of the fourth spatial point p(t).
[0184] 807. The computer device trains the rendering model based on the predicted color of each pixel, the color of each pixel in the sample image, the sample directed distances of multiple third space points, and the predicted directed distances.
[0185] In an embodiment of the present application, the predicted color of a pixel is obtained through a rendering model, and the difference between the predicted color of a pixel and the color of the pixel in the sample image can reflect the accuracy of the rendering model. The predicted directed distance of the third space is obtained through a rendering model, and the difference between the sample directed distance of a pixel and the predicted directed distance can reflect the accuracy of the rendering model. Therefore, based on the predicted color of each pixel, the color of each pixel in the sample image, the sample directed distance and the predicted directed distance of multiple third space points, the rendering model is trained so that the predicted color of the pixel obtained by the rendering model is as close as possible to the color of the pixel in the sample image, and the predicted directed distance of the third space point obtained by the rendering model is as close as possible to the sample directed distance, thereby improving the accuracy of the rendering model.
[0186] In one possible implementation, step 807 includes: determining a first loss value based on the predicted color of each pixel and the color of each pixel in the sample image, the first loss value indicating the difference between the color of the same pixel in the sample image and the predicted color; determining a second loss value based on the sample directed distances and predicted directed distances of multiple third space points, the second loss value indicating the difference between the sample directed distances and the predicted directed distances of the same third space point; and training the rendering model based on the first loss value and the second loss value.
[0187] In an embodiment of the present application, a first loss value is determined based on the predicted color of each pixel and the color of each pixel in the sample image, and a second loss value is determined based on the sample directed distances and predicted directed distances of multiple third spatial points. Then, based on the first loss value and the second loss value, the rendering model is trained to reduce the difference between the color of the same pixel in the sample image and the predicted color, and to reduce the difference between the sample directed distance and the predicted directed distance of the same third spatial point, thereby improving the accuracy of the rendering model.
[0188] It should be noted that the embodiment shown in FIG8 is merely an example of training a rendering model based on a sample image of a sample virtual object in a sample three-dimensional space. In another embodiment, the sample virtual object can be photographed from different perspectives based on the virtual object in the sample three-dimensional space to obtain different sample images, and then the rendering model can be trained based on the different sample images. For example, each time a sample image of a perspective is obtained, the rendering model is iterated once based on the sample image, and then a sample image of another perspective is obtained to perform the next iteration on the rendering model. For another example, the rendering model is iterated once based on the sample virtual object in the sample three-dimensional space, and then the rendering model is iterated again based on another sample virtual object in the sample three-dimensional space. In the process of iteratively training the rendering model, if the number of iterations reaches a threshold, or if the sum of the first loss value and the second loss value in the current iteration is less than the loss value threshold, the iterative training of the rendering model is stopped.
[0189] In the solution provided in the embodiment of the present application, when the sample three-dimensional space contains a sample virtual object, the characteristics of the spatial point in the sample three-dimensional space can reflect the texture of the sample virtual object or the positional relationship with the sample virtual object. The rendering model is trained using the sample virtual object in the sample three-dimensional space so that the rendering model can learn to map the characteristics of the spatial point in the three-dimensional space into the color and directed distance of the spatial point, thereby improving the accuracy of the rendering model.
[0190] In addition, a virtual object in the sample three-dimensional space is used to obtain a sample image of the sample virtual object in the sample three-dimensional space, and a plurality of third space points are determined using the position points of the pixel points in the sample image to obtain the sample directed distances of the third space points. The predicted color of each pixel point and the predicted directed distances of the plurality of third space points are obtained through the rendering model. Based on the predicted color of each pixel point, the color of each pixel point in the sample image, the sample directed distances of the plurality of third space points and the predicted directed distances, the rendering model is trained so that the predicted color of the pixel point obtained by the rendering model is as close as possible to the color of the pixel point in the sample image, and the predicted directed distances of the third space point obtained by the rendering model are as close as possible to the sample directed distances, thereby improving the accuracy of the rendering model.
[0191] It should be noted that the embodiment of the present application uses sample images obtained by photographing virtual objects to train the rendering model, while in another embodiment, there is no need to perform the above steps 806-807, but other methods are adopted to train the rendering model based on the predicted colors of multiple third space points, the sample directed distances of multiple third space points and the predicted directed distances.
[0192] In one possible implementation, the process of training the rendering model includes: extracting the sample color of each third space point from the sample three-dimensional space; training the rendering model based on the sample colors and predicted colors of multiple third space points, and the sample directed distances and predicted directed distances of multiple third space points.
[0193] In an embodiment of the present application, the sample three-dimensional space includes a sample virtual object, and the sample three-dimensional space includes multiple spatial points. Then, when the sample three-dimensional space includes the sample virtual object, the color of each spatial point in the sample three-dimensional space can be determined, and the sample color of each third spatial point can be extracted from the sample three-dimensional space. The difference between the sample color and the predicted color of the third spatial point can reflect the accuracy of the rendering model. Therefore, based on the sample colors and predicted colors of multiple third spatial points, and the sample directed distances and predicted directed distances of multiple third spatial points, the rendering model is trained so that the predicted color of the third spatial point obtained by the rendering model is as close as possible to the sample color of the third spatial point, and the predicted directed distance of the third spatial point obtained by the rendering model is as close as possible to the sample directed distance, thereby improving the accuracy of the rendering model.
[0194] Optionally, the process of training the rendering model includes: determining a third loss value based on the sample colors and predicted colors of multiple third space points, the third loss value indicating the difference between the sample color and the predicted color of the same third space point, determining a second loss value based on the sample directed distances and predicted directed distances of multiple third space points, and training the rendering model based on the third loss value and the second loss value.
[0195] In an embodiment of the present application, a third loss value is determined based on the sample colors and predicted colors of multiple third spatial points, and a second loss value is determined based on the sample directed distances and predicted directed distances of multiple third spatial points. Then, the rendering model is trained at the third loss value and the second loss value to reduce the difference between the sample color and the predicted color of the same third spatial point, and to reduce the difference between the sample directed distance and the predicted directed distance of the same third spatial point, thereby improving the accuracy of the rendering model.
[0196] In an embodiment of the present application, the three-dimensional virtual objects in the sample three-dimensional space can be existing three-dimensional game data. The existing three-dimensional game data is used to train the rendering model, and the virtual objects in the three-dimensional space are implicitly expressed in the form of a signed distance field, so that the rendering model can learn to map the features of spatial points in the three-dimensional space into the color and signed distance of the spatial points.
[0197] Based on the embodiment shown in FIG3 , before using the diffusion model to generate a three-dimensional virtual object, the diffusion model needs to be trained. The training process is detailed in the following embodiment.
[0198] FIG11 is a flowchart of a method for generating a virtual object provided in an embodiment of the present application. The method is executed by a computer device. As shown in FIG11 , the method includes:
[0199] 1101. The computer device obtains sample description information and sample features of multiple fifth spatial points in the sample three-dimensional space based on the sample virtual object in the sample three-dimensional space. The sample description information is used to describe the style of the sample virtual object.
[0200] The fifth spatial point is any spatial point in the sample three-dimensional space, and the sample feature of the fifth spatial point is used to characterize the texture of the fifth spatial point in the sample three-dimensional space or the positional relationship with the sample three-dimensional virtual object.
[0201] In a possible implementation, the plurality of fifth spatial points are spatial points located on the surface of the sample virtual object in the sample three-dimensional space, or are the third spatial points determined according to the embodiment shown in FIG. 8 .
[0202] In one possible implementation, the process of obtaining sample description information includes: photographing the sample virtual object based on a virtual camera in the sample three-dimensional space to obtain a sample image, and determining the sample image as the sample description information; or photographing the sample virtual object based on the virtual camera in the sample three-dimensional space to obtain a sample image, converting the sample image to obtain sample text, and determining the sample text as the sample description information; or converting the sample virtual object in the sample three-dimensional space into point cloud data, and determining the point cloud data as the sample description information.
[0203] In an embodiment of the present application, the sample description information can be any type of information, for example, the sample description information is an image, text, or point cloud data. Optionally, the sample description information includes at least one of a sample image, a sample text, or point cloud data. In an embodiment of the present application, different types of information can be used as sample description information so that the diffusion model can be applied to multiple types of sample description information. In addition, when the sample description information includes multiple types of information, the sample description information is multimodal information. The diffusion model is trained based on the multimodal information so that the three-dimensional virtual object can be generated based on the multimodal information to ensure the accuracy of the generated three-dimensional virtual object.
[0204] It should be noted that the above description uses the example of a local device photographing a sample virtual object in a sample three-dimensional space using a virtual camera. In another embodiment, other methods can be used to obtain a sample image. The sample image is obtained by photographing the sample virtual object in the sample three-dimensional space using a virtual camera. For example, after another device photographs the sample virtual object in the sample three-dimensional space using a virtual camera to obtain a sample image, it sends the sample image to the local device, which then receives the sample image.
[0205] In one possible implementation, the process of obtaining the sample features of the fifth spatial point includes: extracting the sample features of the fifth spatial point from the sample three-dimensional space; or, in the case that the fifth spatial point is the third spatial point determined according to the embodiment shown in Figure 8 above, in the process of training the rendering model according to the embodiment shown in Figure 8 above, the features of the third spatial point will be stored in a hash table, and the features of the fifth spatial point will be obtained from the above hash table.
[0206] 1102. The computer device adds noise to the sample feature of each fifth spatial point to obtain a noise feature of each fifth spatial point.
[0207] In the embodiment of the present application, by adding noise to the sample feature of the fifth spatial point, the difference between the obtained noise feature and the sample feature becomes larger, so that the subsequent training of the diffusion model can learn the ability to transform from noise features to sample features.
[0208] In a possible implementation, the process of adding noise includes: adding noise multiple times to the sample feature of the fifth spatial point to obtain the noise feature of the fifth spatial point.
[0209] 1103. The computer device updates the noise feature of each fifth spatial point based on the sample description information through the diffusion model to obtain an updated feature of each fifth spatial point.
[0210] In an embodiment of the present application, the diffusion model is used to update the features of the spatial point using the object description information so that the updated features of the spatial point match the object description information, so that a three-dimensional virtual object can be subsequently generated in the three-dimensional space based on the updated features of the spatial point.
[0211] The step 1103 is similar to the above step 202 and will not be described again here.
[0212] 1104. The computer device trains the diffusion model based on the sample features of each fifth spatial point and the updated features of each fifth spatial point.
[0213] In an embodiment of the present application, the updated features of the fifth spatial point are obtained through a diffusion model. The difference between the sample features of the fifth spatial point and the updated features of the fifth spatial point can reflect the accuracy of the diffusion model. Therefore, based on the sample features of each fifth spatial point and the updated features of each fifth spatial point, the diffusion model is trained to reduce the difference between the sample features of the fifth spatial point and the updated features of the fifth spatial point, thereby improving the accuracy of the diffusion model.
[0214] In one possible implementation, the process of training the diffusion model includes: determining a fourth loss value based on sample features and updated features of multiple fifth spatial points, the fourth loss value indicating the difference between the sample features and the updated features of the same fifth spatial point, and training the diffusion model based on the fourth loss value.
[0215] Optionally, the fourth loss value satisfies the following relationship: L diff =||γ(z s -γ(s), z′0)||2
[0216] Among them, L diff Used to represent the fourth loss value, z s It is used to represent the noise feature obtained by adding multiple noises to the sample feature of the fifth spatial point, γ(s) is used to represent the noise added in the process of adding multiple noises to the sample feature of the fifth spatial point, γ(·) is used to represent the position encoding function, z′0 is used to represent the updated feature of the fifth spatial point, and z s -γ(s) is used to represent the sample feature of the fifth space point, and γ is used to represent a function that determines the difference between features.
[0217] In the solution provided in the embodiment of the present application, based on the sample virtual object in the sample three-dimensional space, sample description information and sample features of multiple fifth spatial points in the sample three-dimensional space are obtained, and by adding noise to the sample features of the fifth spatial point, the difference between the obtained noise features and the sample features is increased, and then the diffusion model is trained based on the sample description information and the noise features of the fifth spatial point, so that the subsequent training of the diffusion model can learn the ability to transform from noise features to sample features, so that the difference between the sample features of the fifth spatial point and the updated features of the fifth spatial point is reduced, thereby improving the accuracy of the diffusion model.
[0218] In the embodiment of the present application, the diffusion model is used to generate new features of the spatial point based on the features of the spatial point and the object description information, and the new features of the spatial point match the object description information.
[0219] Based on the embodiment shown in FIG11 , if the sample features of the plurality of fifth spatial points do not obey a normal distribution, the sample features of the plurality of fifth spatial points can be converted to obey a normal distribution. Then, the diffusion model is trained according to the features obeying the normal distribution. That is, the process of training the diffusion model includes:
[0220] Step 1: Align the sample features of multiple fifth space points to the standard normal distribution through the variational autoencoder.
[0221] Among them, the variational autoencoder is used to align multiple features to a normal distribution. For example, the variational autoencoder is VAE (Variational Auto Encoder).
[0222] In one possible implementation, the method further includes: training a variational autoencoder based on sample features of a fifth space point that obeys a standard normal distribution.
[0223] Optionally, a fifth loss value is determined based on sample features of a fifth spatial point that obeys a standard normal distribution, and the variational autoencoder is trained based on the fifth loss value.
[0224] Optionally, the fifth loss value satisfies the following relationship: L vae =||Φ(x|z)-S||1+λ(D KL (q φ (z|π)||p(z)))
[0225] Among them, L vae is used to represent the fifth loss value, ||Φ(x|z)-S||1 is used to represent the difference between the sample features of the fifth space point that obeys the standard normal distribution and the sample features of multiple fifth space points, S is used to represent the sample features of multiple fifth space points, Φ(x|z) is used to represent the sample features of the fifth space point that obeys the standard normal distribution, x is used to represent the sample features of the fifth space point before transformation, z is used to represent the sample features of the fifth space point that obeys the standard normal distribution after transformation, Φ(·) is used to represent the variational autoencoder; λ is used to represent the weight, q φ (z|π) is used to express that the sample feature z of the fifth spatial point is expected to conform to the Gaussian distribution represented by π, and p(z) is used to express the distribution of the sample features of multiple fifth spatial points after the sample features of multiple fifth spatial points are transformed by the variational autoencoder. D KL Used to represent KL (Kullback-Leibler Divergence, relative entropy) divergence, used to constrain the feature distribution q φ (z|π) is aligned to a target normal distribution p(z).
[0226] Step 2: For the sample features of the fifth spatial point that obey the standard normal distribution, adopt the Markov method to add noise to the sample features of the fifth spatial point multiple times to obtain the noise features of the fifth spatial point that obey the standard normal distribution.
[0227] In one possible implementation, the noise characteristic of the fifth spatial point satisfies the following relationship: α s =1-β s
[0228] Among them, q(x M |x0) is used to represent the noise feature x obtained by adding M noises to the sample feature x0 of the fifth spatial point. M , s is used to represent the number of times the noise is added, s is an integer not less than 1 and not greater than M, M is the total number of times the noise is added, M is an integer greater than 1; q(x s |x s-1 ) is used to represent the noise feature x obtained by adding s-1 noises to the sample feature x0 of the fifth spatial point. s-1 In the case of noise characteristics x s-1 Noise feature x obtained by adding noise s , ∏ is used to represent continuous multiplication; Used to represent the noise characteristics x s-1 Noise feature x obtained by adding noise s It obeys the normal distribution, and the mean of the normal distribution is The variance of the normal distribution is β s I, β s is used to represent the variance coefficient, I is used to represent the identity matrix, α s The coefficient used to represent the fusion noise, α i It is used to represent the coefficient of the noise added for the i-th time, ε is used to represent the noise, and the noise ε obeys the normal distribution, that is, ε~N(0,I).
[0229] Step 3. Use the diffusion model to extract features of the sample description information to obtain sample description features, fuse the sample description features with the noise features of the fifth spatial point that obeys the standard normal distribution to obtain the fused features of the fifth spatial point that obeys the standard normal distribution, denoise the fused features of the fifth spatial point that obeys the standard normal distribution, and obtain the updated features of each fifth spatial point that obeys the standard normal distribution.
[0230] In the embodiment of the present application, the diffusion model denoising process is the opposite process of the Markov process.
[0231] In one possible implementation, the updated features of each fifth spatial point that obeys the standard normal distribution satisfy the following relationship: p θ (x s-1 |x s )=N(x s-1 ;μ θ (x s ,s),∑ θ (x s ,x0)) α s =1-β s
[0232] Among them, p θ (x 0:M ) is used to represent the updated features of the fifth spatial point, x M It is used to represent the noise feature obtained by adding M noises to the sample feature x0 of the fifth spatial point, p θ (x M ) is used to represent the noise characteristics x of the fifth spatial point through the diffusion model M The features obtained by denoising, p θ (x s-1 |x s ) is used to express the noise characteristics x of the fifth space point s In the case of diffusion model, the noise feature x s The noise feature x obtained by denoising s-1 , s is used to represent the number of times the noise is added, s is an integer not less than 1 and not greater than M, M is the total number of times the noise is added, M is an integer greater than 1; N(x s-1 ;μ θ (x s ,s),∑ θ (x s ,x0)) is used to represent the noise feature x through the diffusion model s The noise feature x obtained by denoising s-1 It obeys the normal distribution, and the mean of the normal distribution is μ θ (x s , s), the variance of the normal distribution is ∑ θ (x s , x0), x0 is used to represent the sample features of the fifth space point, α s The coefficient used to represent the fusion noise, α i The coefficient used to represent the noise added for the i-th time, β s Used to represent the coefficient of variance, ε s Used to represent noise. The noise ε obeys the normal distribution, that is, ε~N(0,I), and I is used to represent the unit matrix.
[0233] Step 4: Use a variational autoencoder to restore the updated features of each fifth spatial point that obeys the standard normal distribution, so that the distribution of the restored updated features of the fifth spatial point matches the distribution of the sample features of the fifth spatial point; train the diffusion model based on the sample features of each fifth spatial point and the updated features of each fifth spatial point.
[0234] The embodiment of the present application provides a 3D-AIGC (3Dimensions Artificial Intelligence Generated Content) technology that supports multiple control methods to automatically generate three-dimensional game assets. Users can use pictures, texts or three-dimensional point cloud data individually or in combination as object description information to describe the three-dimensional virtual object they want to generate. Based on the object description information provided by the user, the corresponding three-dimensional geometric facets and corresponding texture maps are automatically generated. The three-dimensional geometric facets and texture maps can be directly imported into the game development engine for game production. It is also possible to generate a three-dimensional virtual object in the target three-dimensional space based on the three-dimensional geometric facets and texture maps, and set the three-dimensional virtual object in the game. This method reduces the workload of developing three-dimensional virtual objects, greatly reduces the development cost of game production, shortens the development cycle, and supports players to independently generate three-dimensional virtual objects and participate in a new game mode of game content creation.
[0235] In an embodiment of the present application, a signed distance field is used to describe the three-dimensional virtual objects to be generated in the target three-dimensional scene to ensure that the generated three-dimensional virtual objects are smooth and complete. The solution provided in the embodiment of the present application supports the control of text, image and point cloud data in three ways, either individually or in combination, and can provide fine control over the generated three-dimensional virtual objects. In addition to generating three-dimensional geometric patches, texture maps based on the physical rendering pipeline are also generated, which can be directly used in the production of game pipelines.
[0236] The method provided in the embodiments of this application can be applied in a variety of scenarios. For example, in a 3D simulation scenario or a gaming scenario. Taking a 3D simulation scenario as an example, based on the solution provided in the embodiments of this application, it is possible to automatically generate a 3D virtual object using the object description information input by the user.
[0237] In the game development scenario, the current game production pipeline mainly relies on manpower to produce game resources. The production of a single game character requires a series of complex processes such as original painting design, high-poly model sculpting, face reduction, baking textures, skinning, and bone binding. Most of this process is completed in different game production software and requires a lot of manpower. The solution provided in the embodiment of the present application provides game producers with a complete set of automatic game resource generation tools. Producers can describe the three-dimensional game resources they want to generate through text description, image input, and three-dimensional point cloud data input, and the system will automatically generate corresponding game resources including geometric faces and texture maps. The generated assets can be directly imported into the game engine, which greatly reduces the cost of game production, speeds up game production, and makes game production more convenient.
[0238] The virtual object generation method provided in the embodiments of the present application can be applied in game development scenarios. Based on the method provided in the embodiments of the present application, developers configure object description information, and the computer device can generate a three-dimensional virtual object by combining the object description information with a diffusion model and a rendering model; or generate geometric facets and texture maps, and developers can configure the three-dimensional virtual object or geometric facets and texture maps in the game to improve the efficiency of game development. In addition, when generating a three-dimensional virtual object, the developer can also adjust the three-dimensional virtual object to adjust the shape or color of the three-dimensional virtual object, thereby generating new geometric facets and new texture maps, and then configuring the new geometric facets and new texture maps in the game.
[0239] For example, the object description information is an image. Based on the method provided in the embodiment of the present application, a three-dimensional virtual object that matches the image can be generated through a diffusion model and a rendering model, as shown in Figures 12 and 13. For another example, the object description information is text. Based on the method provided in the embodiment of the present application, a three-dimensional virtual object that matches the three-dimensional virtual object described by the text can be generated through a diffusion model and a rendering model, as shown in Figure 14. For another example, the object description information is the color point cloud data (Q, 6) corresponding to the game character, where Q is the number of points in the color point cloud data, and 6 is the dimension, where 3 dimensions represent position and 3 dimensions represent color. Based on the method provided in the embodiment of the present application, a corresponding three-dimensional virtual character can be automatically generated through a diffusion model and a rendering model, as shown in Figure 15.
[0240] FIG16 is a schematic diagram of the structure of a virtual object generation device provided in an embodiment of the present application. As shown in FIG16 , the device includes:
[0241] An acquisition module 1601 is used to acquire object description information, where the object description information is used to describe the three-dimensional virtual object to be generated;
[0242] An updating module 1602 is configured to obtain first features of a plurality of first spatial points in the target three-dimensional space based on the object description information, wherein the first features of the first spatial points are used to characterize the color of the first spatial points in the target three-dimensional space and the positional relationship with the three-dimensional virtual object, and the first spatial points are uniformly distributed in the target three-dimensional space.
[0243] a processing module 1603 configured to process the first feature of each first spatial point by rendering the model to obtain a color and a directed distance of each first spatial point, where the directed distance indicates a distance from the first spatial point to the surface of the three-dimensional virtual object in the target three-dimensional space;
[0244] The generating module 1604 is configured to generate a three-dimensional virtual object in the target three-dimensional space based on the colors and directed distances of the plurality of first spatial points.
[0245] In one possible implementation, the updating module 1602 is configured to update the second features of the plurality of first spatial points based on the object description information to obtain the first features of the plurality of first spatial points, where the second features are features of the first spatial points in the target three-dimensional space.
[0246] In another possible implementation, the update module 1602 is used to extract features from the object description information to obtain object description features; randomly generate a second feature for the first spatial point; fuse the second feature of the first spatial point and the object description feature to obtain a fused feature of the first spatial point; and denoise the fused feature of the first spatial point to obtain a first feature of the first spatial point.
[0247] In another possible implementation, generation module 1604 is used to determine multiple second spatial points from multiple first spatial points based on the directed distances of the multiple first spatial points, where the second spatial points are located on the surface of the three-dimensional virtual object; in the target three-dimensional space, the multiple second spatial points are connected to form a model of the three-dimensional virtual object; based on the colors of the multiple second spatial points, the model of the three-dimensional virtual object is rendered to obtain a three-dimensional virtual object containing colors.
[0248] In another possible implementation, each first feature includes multiple sub-features, and the resolutions of the multiple sub-features are different; the processing module 1603 is used to fuse the multiple sub-features of the first spatial point through a rendering model, process the fused features, and obtain the color and directed distance of the first spatial point.
[0249] In another possible implementation, the object description information includes at least one of an image, text, or point cloud data.
[0250] In another possible implementation, the object description information includes at least two items of image, text, or point cloud data; the update module 1602 is used to extract features of each sub-information in the object description information to obtain features of each sub-information, where the sub-information is any one of image, text, or point cloud data; the features of multiple sub-information in the object description information are fused to obtain object description features; based on the object description features, the second features of multiple first spatial points are updated to obtain first features of multiple first spatial points.
[0251] In another possible implementation, as shown in FIG17 , the apparatus further includes:
[0252] Sampling module 1605 is used to uniformly sample the spatial points in the target three-dimensional space to obtain multiple first spatial points, where the distance between each two adjacent first spatial points is equal, that is, it is used to uniformly sample the target three-dimensional space to obtain multiple first spatial points.
[0253] In another possible implementation, as shown in FIG17 , the apparatus further includes:
[0254] The acquisition module 1601 is further configured to acquire sample directed distances of a plurality of third spatial points in the sample three-dimensional space based on the sample virtual object in the sample three-dimensional space, where the sample directed distances indicate distances from the third spatial points to the surface of the sample virtual object;
[0255] Extraction module 1606, for extracting features of each third space point from the sample three-dimensional space;
[0256] The processing module 1603 is further configured to process the features of each third space point using a rendering model to obtain a predicted color and a predicted directed distance of each third space point;
[0257] The training module 1607 is configured to train the rendering model based on the predicted colors of the plurality of third spatial points, the sample directed distances of the plurality of third spatial points, and the predicted directed distances.
[0258] In another possible implementation, acquisition module 1601 is configured to acquire a sample image, where the sample image is obtained by photographing the sample virtual object based on a virtual camera in the sample three-dimensional space; determine, from the sample three-dimensional space, a third spatial point corresponding to each pixel point based on a position of the virtual camera in the sample three-dimensional space and a position of each pixel point in the sample image in the sample three-dimensional space; and acquire a sample directed distance of each third spatial point based on the sample virtual object;
[0259] Training module 1607 is used to fuse the predicted colors of the third space points corresponding to the same pixel point to obtain the predicted color of each pixel point; based on the predicted color of each pixel point, the color of each pixel point in the sample image, the sample directed distances and predicted directed distances of multiple third space points, the rendering model is trained.
[0260] Optionally, the acquisition module 1601 is configured to photograph the sample virtual object based on a virtual camera in the sample three-dimensional space to obtain a sample image.
[0261] In another possible implementation, the acquisition module 1601 is used to determine, in the sample three-dimensional space, a ray that takes the position point of the virtual camera as the starting point and passes through the position point of the pixel point; and takes the position point of the pixel point as the starting point and collects at least one third space point on the ray along the direction of the ray; that is, it is used to determine, in the sample three-dimensional space, a ray that takes the position point of the virtual camera as the starting point and passes through the pixel point; and takes the pixel point as the starting point and collects at least one third space point on the ray along the direction of the ray.
[0262] In another possible implementation, the training module 1607 is used to determine the transparency of multiple fourth spatial points based on the sample directed distances of the multiple fourth spatial points, where the transparency is positively correlated with the sample directed distances, and the multiple fourth spatial points correspond to the same pixel point, and the fourth spatial point is any one of the third spatial points corresponding to the pixel point; based on the transparency of the multiple fourth spatial points, the predicted colors of the multiple fourth spatial points are fused to obtain the predicted color of the pixel point.
[0263] In another possible implementation, the training module 1607 is used to determine a first loss value based on the predicted color of each pixel and the color of each pixel in the sample image, where the first loss value indicates the difference between the color of the same pixel in the sample image and the predicted color; determine a second loss value based on the sample directed distances and predicted directed distances of multiple third space points, where the second loss value indicates the difference between the sample directed distances and the predicted directed distances of the same third space point; and train the rendering model based on the first loss value and the second loss value.
[0264] In another possible implementation, the extraction module 1606 is further configured to extract the sample color of each third spatial point from the sample three-dimensional space;
[0265] The training module 1607 is configured to train the rendering model based on the sample colors and predicted colors of a plurality of third spatial points, and the sample directed distances and predicted directed distances of a plurality of third spatial points.
[0266] In another possible implementation, based on the object description information, the second features of the plurality of first spatial points in the target three-dimensional space are updated to obtain the first features of the plurality of first spatial points. The diffusion model is executed, as shown in FIG17 , and the apparatus further includes:
[0267] The acquisition module 1601 is further configured to acquire sample description information and sample features of multiple fifth spatial points in the sample three-dimensional space based on the sample virtual object in the sample three-dimensional space, where the sample description information is used to describe the style of the sample virtual object;
[0268] A noise adding module 1608 is configured to add noise to the sample feature of each fifth spatial point to obtain a noise feature of each fifth spatial point;
[0269] An updating module 1602 is configured to update the noise feature of each fifth spatial point based on the sample description information by using a diffusion model to obtain an updated feature of each fifth spatial point;
[0270] The training module 1607 is configured to train the diffusion model based on the sample features of each fifth spatial point and the updated features of each fifth spatial point.
[0271] In another possible implementation, the acquisition module 1601 is used to acquire a sample image, where the sample image is obtained by photographing the sample virtual object based on a virtual camera in the sample three-dimensional space, and the sample image is determined as the sample description information; or, the sample virtual object is photographed based on the virtual camera in the sample three-dimensional space to obtain a sample image, the sample image is converted to obtain sample text, and the sample text is determined as the sample description information; or, the sample virtual object in the sample three-dimensional space is converted into point cloud data, and the point cloud data is determined as the sample description information.
[0272] In another possible implementation, the acquisition module 1601 is configured to photograph the sample virtual object based on a virtual camera in the sample three-dimensional space to obtain a sample image.
[0273] It should be noted that the virtual object generation device provided in the above embodiment is merely an example of the division of the aforementioned functional modules. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the virtual object generation device provided in the above embodiment and the virtual object generation method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0274] An embodiment of the present application further provides a computer device, which includes a processor and a memory, wherein the memory stores at least one computer program, which is loaded and executed by the processor to implement the operations performed by the virtual object generation method of the above embodiment.
[0275] Optionally, the computer device is provided as a terminal. FIG18 shows a block diagram of a terminal 1800 provided in an exemplary embodiment of the present application. The terminal 1800 includes: a processor 1801 and a memory 1802.
[0276] The processor 1801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1801 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0277] Memory 1802 may include one or more computer-readable storage media, which may be non-transitory. Memory 1802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 1802 is used to store at least one computer program, which is executed by processor 1801 to implement the virtual object generation method provided in the method embodiment of the present application.
[0278] In some embodiments, terminal 1800 may optionally include a peripheral device interface 1803 and at least one peripheral device. The processor 1801, memory 1802, and peripheral device interface 1803 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1803 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1804, a display screen 1805, a camera assembly 1806, an audio circuit 1807, and a power supply 1808.
[0279] The peripheral device interface 1803 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1801 and the memory 1802. In some embodiments, the processor 1801, the memory 1802, and the peripheral device interface 1803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1801, the memory 1802, and the peripheral device interface 1803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0280] The RF circuit 1804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 1804 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 1804 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1804 may also include circuitry related to Near Field Communication (NFC), although this application does not limit this.
[0281] The display screen 1805 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1805 is a touch screen display, the display screen 1805 also has the ability to collect touch signals on the surface or above the surface of the display screen 1805. The touch signal can be input as a control signal to the processor 1801 for processing. At this time, the display screen 1805 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 1805, which is set on the front panel of the terminal 1800; in other embodiments, there can be at least two display screens 1805, which are respectively set on different surfaces of the terminal 1800 or in a folding design; in other embodiments, the display screen 1805 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 1800. Even more, the display screen 1805 can be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1805 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0282] The camera assembly 1806 is used to capture images or videos. Optionally, the camera assembly 1806 includes a front camera and a rear camera. The front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1806 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0283] The audio circuit 1807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1801 for processing, or input into the radio frequency circuit 1804 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1801 or the radio frequency circuit 1804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as distance measurement. In some embodiments, the audio circuit 1807 may also include a headphone jack.
[0284] Power supply 1808 is used to power various components in terminal 1800. Power supply 1808 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1808 includes a rechargeable battery, the rechargeable battery can be wired or wirelessly rechargeable. A wired rechargeable battery is charged via a wired line, while a wireless rechargeable battery is charged via a wireless coil. The rechargeable battery can also support fast charging technology.
[0285] Those skilled in the art will understand that the structure shown in FIG18 does not constitute a limitation on the terminal 1800 , and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0286] Optionally, the computer device is provided as a server. Figure 19 is a structural diagram of a server provided in an embodiment of the present application. The server 1900 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1901 and one or more memories 1902, wherein at least one computer program is stored in the memory 1902, and at least one computer program is loaded and executed by the processor 1901 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.
[0287] An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the virtual object generation method of the above embodiment.
[0288] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the operations performed by the virtual object generation method of the above embodiment.
[0289] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0290] The above description is only an optional embodiment of the embodiment of the present application and is not intended to limit the embodiment of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiment of the present application should be included in the scope of protection of the present application.
Claims
1. A method for generating a virtual object, executed by a computer device, the method comprising: Obtaining object description information, where the object description information is used to describe the three-dimensional virtual object to be generated; Based on the object description information, first features of a plurality of first spatial points in the target three-dimensional space are acquired, where the first features of the first spatial points are used to characterize the color of the first spatial points in the target three-dimensional space and the positional relationship with the three-dimensional virtual object, and the plurality of first spatial points are evenly distributed in the target three-dimensional space; Processing the first feature of each first spatial point by rendering the model to obtain a color and a directed distance of each first spatial point, wherein the directed distance indicates a distance from the first spatial point to the surface of the three-dimensional virtual object in the target three-dimensional space; The three-dimensional virtual object is generated in the target three-dimensional space based on the colors and the directed distances of the plurality of first spatial points.
2. The method according to claim 1, wherein: The acquiring, based on the object description information, first features of a plurality of first spatial points in the target three-dimensional space includes: Based on the object description information, the second features of the multiple first spatial points are updated to obtain the first features of the multiple first spatial points, where the second features of the first spatial points are features of the first spatial points in the target three-dimensional space.
3. The method according to claim 1 or 2, wherein: The updating of the second features of the plurality of first spatial points based on the object description information to obtain the first features of the plurality of first spatial points includes: Extracting features from the object description information to obtain object description features; randomly generating a second feature for the first spatial point; fusing the second feature of the first spatial point and the object description feature to obtain a fused feature of the first spatial point; The fused features of the first spatial point are denoised to obtain a first feature of the first spatial point.
4. The method according to any one of claims 1 to 3, wherein: The step of generating the three-dimensional virtual object in the target three-dimensional space based on the colors and directed distances of the plurality of first spatial points includes: Based on the directed distances of the plurality of first spatial points, determining a plurality of second spatial points from the plurality of first spatial points, wherein the second spatial points are located on the surface of the three-dimensional virtual object; In the target three-dimensional space, connecting the plurality of second spatial points to form a model of a three-dimensional virtual object; Based on the colors of the plurality of second spatial points, the model of the three-dimensional virtual object is rendered to obtain a three-dimensional virtual object including colors.
5. The method according to any one of claims 1 to 4, wherein: Each first feature includes a plurality of sub-features, and the plurality of sub-features have different resolutions; The first feature of each first spatial point is processed by the rendering model to obtain the color and directed distance of each first spatial point, including: Through the rendering model, multiple sub-features of the first spatial point are fused, and the fused features are processed to obtain the color and directed distance of the first spatial point.
6. The method according to any one of claims 1 to 5, wherein: The object description information includes at least one of an image, text or point cloud data.
7. The method according to any one of claims 1 to 6, wherein: The object description information includes at least two items of image, text or point cloud data; and the updating of the second features of the plurality of first spatial points based on the object description information to obtain the first features of the plurality of first spatial points includes: Performing feature extraction on each sub-information in the object description information to obtain features of each sub-information, wherein the sub-information is any one of the image, text or point cloud data; Merging features of multiple sub-information in the object description information to obtain object description features; Based on the object description feature, the second features of the plurality of first spatial points are updated to obtain the plurality of first spatial points. The first characteristic of a point in space.
8. The method according to any one of claims 1 to 7, wherein: Before acquiring first features of a plurality of first spatial points in the target three-dimensional space based on the object description information, the method further includes: The target three-dimensional space is uniformly sampled to obtain the multiple first space points, wherein the distance between every two adjacent first space points in the multiple first space points is equal.
9. The method according to any one of claims 1 to 8, wherein: The method further comprises: Based on the sample virtual object in the sample three-dimensional space, acquiring sample directed distances of a plurality of third spatial points in the sample three-dimensional space, wherein the sample directed distances indicate distances from the third spatial points to the surface of the sample virtual object; Extracting features of each third space point from the sample three-dimensional space; Processing the features of each third spatial point through the rendering model to obtain a predicted color and a predicted directed distance of each third spatial point; The rendering model is trained based on the predicted colors of the plurality of third spatial points, the sample directed distances and the predicted directed distances of the plurality of third spatial points.
10. The method according to any one of claims 1 to 9, wherein: The acquiring, based on the sample virtual object in the sample three-dimensional space, sample directed distances of a plurality of third spatial points in the sample three-dimensional space comprises: Acquire a sample image, where the sample image is obtained by photographing the sample virtual object based on a virtual camera in the sample three-dimensional space; Based on the position point of the virtual camera in the sample three-dimensional space and the position point of each pixel point in the sample image in the sample three-dimensional space, determining a third space point corresponding to each pixel point in the sample three-dimensional space; Based on the sample virtual object, obtaining a sample directed distance of each third spatial point; The training of the rendering model based on the predicted colors of the plurality of third spatial points, the sample directed distances and the predicted directed distances of the plurality of third spatial points comprises: The predicted colors of the third space points corresponding to the same pixel are merged to obtain the predicted color of each pixel; The rendering model is trained based on the predicted color of each pixel, the color of each pixel in the sample image, the sample directed distances and the predicted directed distances of the plurality of third space points.
11. The method according to any one of claims 1 to 10, wherein: The determining, from the sample three-dimensional space, a third space point corresponding to each pixel point based on the position point of the virtual camera in the sample three-dimensional space and the position point of each pixel point in the sample image in the sample three-dimensional space, comprises: In the sample three-dimensional space, determining a ray starting from the position point of the virtual camera and passing through the pixel point; Taking the pixel point as a starting point and along the direction of the ray, at least one third space point is collected on the ray.
12. The method according to any one of claims 1 to 11, wherein: The step of fusing the predicted colors of the third spatial points corresponding to the same pixel to obtain the predicted color of each pixel includes: Determine the transparency of the plurality of fourth spatial points based on the sample directed distances of the plurality of fourth spatial points, the transparency being positively correlated with the sample directed distances, the plurality of fourth spatial points corresponding to the same pixel point, and the fourth spatial point being any one of the third spatial points corresponding to the pixel point; Based on the transparency of the multiple fourth spatial points, the predicted colors of the multiple fourth spatial points are fused to obtain the predicted color of the pixel point.
13. The method according to any one of claims 1 to 12, wherein: The training of the rendering model based on the predicted color of each pixel, the color of each pixel in the sample image, the sample directed distances and the predicted directed distances of the plurality of third space points comprises: Based on the predicted color of each pixel and the color of each pixel in the sample image, determining a first loss value, wherein the first loss value indicates a difference between the color of the same pixel in the sample image and the predicted color; determining a second loss value based on the sample directed distances and the predicted directed distances of the plurality of third spatial points, the second loss value indicating a difference between the sample directed distance and the predicted directed distance of the same third spatial point; The rendering model is trained based on the first loss value and the second loss value.
14. The method according to any one of claims 1 to 13, wherein: The method further comprises: Extracting the sample color of each third spatial point from the sample three-dimensional space; The training of the rendering model based on the predicted colors of the plurality of third spatial points, the sample directed distances and the predicted directed distances of the plurality of third spatial points comprises: The rendering model is trained based on the sample colors and predicted colors of the plurality of third spatial points, and the sample directed distances and predicted directed distances of the plurality of third spatial points.
15. The method according to any one of claims 1 to 14, wherein: The updating of the second features of the plurality of first spatial points based on the object description information to obtain the first features of the plurality of first spatial points is performed by a diffusion model, and the method further comprises: Based on the sample virtual object in the sample three-dimensional space, acquiring sample description information and sample features of a plurality of fifth spatial points in the sample three-dimensional space, wherein the sample description information is used to describe the style of the sample virtual object; Adding noise to the sample feature of each fifth spatial point to obtain the noise feature of each fifth spatial point; By using a diffusion model and based on the sample description information, the noise feature of each fifth spatial point is updated to obtain an updated feature of each fifth spatial point; The diffusion model is trained based on the sample features of each fifth spatial point and the updated features of each fifth spatial point.
16. The method according to any one of claims 1 to 15, wherein: Based on the sample virtual object in the sample three-dimensional space, sample description information is obtained, including: Acquire a sample image, where the sample image is obtained by photographing the sample virtual object based on a virtual camera in the sample three-dimensional space, and determine the sample image as the sample description information; or, Based on a virtual camera in the sample three-dimensional space, the sample virtual object is photographed to obtain the sample image, the sample image is converted to obtain sample text, and the sample text is determined as the sample description information; or The sample virtual object in the sample three-dimensional space is converted into point cloud data, and the point cloud data is determined as the sample description information.
17. A virtual object generation device, the device comprising: An acquisition module, used to acquire object description information, where the object description information is used to describe the three-dimensional virtual object to be generated; An updating module, configured to obtain first features of a plurality of first spatial points in a target three-dimensional space based on the object description information, wherein the first spatial points are evenly distributed in the target three-dimensional space; a processing module, configured to process the first feature of each first spatial point by rendering a model to obtain a color and a directed distance of each first spatial point, wherein the directed distance indicates a distance from the first spatial point to the surface of the three-dimensional virtual object in the target three-dimensional space; A generating module is used to generate the three-dimensional virtual object in the target three-dimensional space based on the colors and directed distances of the multiple first spatial points.
18. A computer device, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the virtual object generation method according to any one of claims 1 to 16.
19. A computer-readable storage medium, wherein at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the virtual object generation method according to any one of claims 1 to 16.
20. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the operations performed by the virtual object generation method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Three-dimensional reconstruction model training method and device, equipment and storage medium
CN115115805A
Three-dimensional reconstruction model training method and device, equipment and storage medium
CN115222917A
Three-dimensional model reconstruction method and device and computer readable storage medium
CN115294275A
Three-dimensional modeling method and device, storage medium, electronic equipment and product
CN115830227A
Video processing method and device and computer readable storage medium
CN116129011A