Prediction model training method and device, computer equipment, medium and program product

Through the training method of the prediction model, virtual objects that conform to the art design concept are generated, which solves the problem of manual optimization of virtual objects production in the existing technology, and achieves the effect of shortening the development cycle and reducing costs.

CN120219615APending Publication Date: 2025-06-27SHENZHEN TENCENT COMP SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252425.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing game production pipeline relies on manpower to create virtual objects, resulting in a long development cycle and high cost, and the generated virtual objects are evenly cabling and requires manual optimization.

Method used

Provide a training method for predictive models, by obtaining virtual object samples and object description samples after wiring optimization, performing encoding and prediction models training, and generating virtual objects that can be directly used for downstream applications.

Benefits of technology

It shortens the game development cycle and reduces the development cost. The generated virtual objects are properly cabling, which conforms to the art design concept and can be directly used in game applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219615A_ABST
    Figure CN120219615A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a prediction model training method and device, computer equipment, a medium and a program product, and the method comprises the steps: obtaining a virtual object sample and an object description sample used for describing the virtual object sample, and enabling the virtual object sample to be obtained through wiring optimization; the virtual object sample is coded to obtain a sample discrete sequence, and the sample discrete sequence is used for describing the positions of a plurality of patches included in the virtual object sample in the virtual space; predicting the object description sample through the initial prediction model to obtain a prediction discrete sequence; according to the difference between the prediction discrete sequence and the sample discrete sequence, model parameters of the initial prediction model are adjusted, a prediction model is obtained, and the prediction model is used for predicting and obtaining a discrete sequence used for describing the virtual object. Therefore, the accuracy of the prediction model obtained through training is high, wiring of the virtual object is sparse and dense and can be directly used for downstream application, the development period is shortened, and the development cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, computer device, medium and program product for training a prediction model. Background Art

[0002] The current game production pipeline mainly relies on manpower for the production of game assets. The production of a virtual object corresponding to a single game character requires a series of complex processes such as original painting design, high-poly sculpting, decimation, baking textures, skinning, and bone binding. The development of a high-quality virtual object requires a team of dozens of people to spend several months to complete, and most of the time is spent on the sculpting of the details of the virtual object, thus greatly reducing the progress of game development, resulting in a long development cycle and high development costs.

[0003] In the related art, generally, a signed distance field (SDF) is used to generate a field for describing the shape of a virtual object in detail, such as which positions in the field have geometry and the inside and outside of the geometry, etc., and then a mesh model approximately representing the object surface is generated through the marching cubes (MC) algorithm, thereby obtaining the virtual object.

[0004] However, the virtual object obtained by this method still needs to be optimized for wiring manually, so there will also be problems of a long development cycle and high development costs. Summary of the Invention

[0005] To solve the above technical problems, the present application provides a method, device, computer device, medium and program product for training a prediction model, which is used to generate a virtual object that can serve downstream game applications, thereby further shortening the development cycle and reducing the development cost.

[0006] The embodiments of the present application disclose the following technical solutions:

[0007] On the one hand, an embodiment of the present application provides a method for training a prediction model, the method includes:

[0008] Obtain a virtual object sample and an object description sample for describing the virtual object sample, where the virtual object sample is obtained through wiring optimization;

[0009] Encode the virtual object sample to obtain a sample discrete sequence, where the sample discrete sequence is used to describe the positions of multiple patches included in the virtual object sample in the virtual space respectively;

[0010] Predict the object description sample through an initial prediction model to obtain a prediction discrete sequence;

[0011] Adjust the model parameters of the initial prediction model according to the difference between the predicted discrete sequence and the sample discrete sequence, to obtain a prediction model, which is used to predict a discrete sequence for describing a virtual object.

[0012] On the other hand, an embodiment of the present application provides a training device for a prediction model, the device includes: an acquisition unit, an encoding unit, a prediction unit, and a training unit;

[0013] The acquisition unit is used to acquire a virtual object sample and an object description sample for describing the virtual object sample, and the virtual object sample is obtained through wiring optimization;

[0014] The encoding unit is used to encode the virtual object sample to obtain a sample discrete sequence, and the sample discrete sequence is used to describe the positions of multiple patches included in the virtual object sample in the virtual space respectively;

[0015] The prediction unit is used to predict the object description sample through an initial prediction model to obtain a predicted discrete sequence;

[0016] The training unit is used to adjust the model parameters of the initial prediction model according to the difference between the predicted discrete sequence and the sample discrete sequence, to obtain a prediction model, which is used to predict a discrete sequence for describing a virtual object.

[0017] On the other hand, an embodiment of the present application provides a computer device, the computer device includes a processor and a memory:

[0018] The memory is used to store a computer program and transmit the computer program to the processor;

[0019] The processor is used to execute the method described in the above aspect according to the instructions in the computer program.

[0020] On the other hand, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method described in the above aspect.

[0021] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which when running on a computer device, causes the computer device to execute the method described in the above aspect.

[0022] As can be seen from the above technical solutions, a virtual object sample and an object description sample for describing the virtual object sample are obtained. The virtual object sample is obtained by optimizing the wiring of the virtual object, that is, the wiring rule of the virtual object sample is properly sparse and dense, belonging to the wiring and topological structure that conforms to the art design concept, and can be directly applied to downstream applications. The virtual object sample is encoded to obtain a sample discrete sequence for describing the positions of multiple patches included in the virtual object sample in the virtual space respectively. The distribution of the multiple patches described by the sample discrete sequence is also properly sparse and dense. Moreover, the method of describing the multiple patches included in the virtual object sample through the sample discrete sequence is more accurate to improve the accuracy of subsequent training. The object description sample is predicted by an initial prediction model to obtain a prediction discrete sequence, so as to convert the object description sample into a prediction discrete sequence that can be used to construct the virtual object sample. In order to improve the accuracy of the initial prediction model, according to the difference between the prediction discrete sequence and the sample discrete sequence, the model parameters of the initial prediction model are adjusted, so that the difference between the prediction discrete sequence predicted by the initial prediction model and the sample discrete sequence becomes smaller and smaller, that is, the accuracy of the initial prediction model is getting higher and higher, and thus a prediction model is obtained. Thus, the virtual object sample with properly sparse and dense wiring that can be used for downstream applications is used as supervision, that is, the sample discrete sequence corresponding to the virtual object sample is used as the training target, so that the trained prediction model can predict a discrete sequence based on the object description, and based on this discrete sequence, a virtual object corresponding to the object description can be constructed. Moreover, due to the high accuracy of the prediction model, the wiring of the virtual object is properly sparse and dense and can be directly used for downstream applications, thus shortening the development cycle and reducing the development cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 A comparison schematic diagram of wiring optimization provided by an embodiment of the present application;

[0025] Figure 2 A schematic diagram of a computer system for a training method of a prediction model provided by an embodiment of the present application;

[0026] Figure 3 An application scenario schematic diagram of a training method of a prediction model provided by an embodiment of the present application;

[0027] Figure 4Schematic flowchart of a method for training a prediction model provided by an embodiment of the present application;

[0028] Figure 5 Schematic diagram of a patch and a patch block provided by an embodiment of the present application;

[0029] Figure 6 Schematic diagram of a target patch block provided by an embodiment of the present application;

[0030] Figure 7 Schematic diagram of the space occupied by a virtual object sample in a virtual space provided by an embodiment of the present application;

[0031] Figure 8 Schematic diagram of a historical object sample and a virtual object sample provided by an embodiment of the present application;

[0032] Figure 9 Schematic diagram of encoding and decoding provided by an embodiment of the present application;

[0033] Figure 10 Schematic diagram of training a prediction model provided by an embodiment of the present application;

[0034] Figure 11 Schematic diagram of applying a prediction model provided by an embodiment of the present application;

[0035] Figure 12 Schematic diagram of comparison provided by an embodiment of the present application;

[0036] Figure 13 Schematic diagram of the structure of a training device for a prediction model provided by an embodiment of the present application;

[0037] Figure 14 Schematic diagram of the structure of a server provided by an embodiment of the present application;

[0038] Figure 15 Schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0039] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0040] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0041] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0042] It should be understood that although the terms first, second, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0043] It should be noted that before and during the collection of relevant data of the user (such as virtual object samples, object description samples, etc.) in this application, a prompt interface, a pop-up window, or voice prompt information can be displayed. The prompt interface, pop-up window, or voice prompt information is used to prompt the user that their relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation issued by the user for the prompt interface or pop-up window. Otherwise (that is, when the confirmation operation issued by the user for the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. That is to say, all user data collected by this application is collected with the consent and authorization of the user, and the collection, use, and processing of the relevant user data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0044] To facilitate the understanding of the technical solution of this application, some technical terms involved in the embodiments of this application are briefly introduced below.

[0045] (1) Virtual object: It refers to an active object in a virtual space. The active object can be a virtual character, a virtual animal, etc. Optionally, when the virtual space is a three-dimensional virtual environment, the virtual object can be a three-dimensional virtual model, such as represented by a 3D mesh, which includes the three-dimensional information of the virtual object. Each virtual object has its own shape and volume in the three-dimensional virtual environment and occupies a part of the space in the three-dimensional virtual environment. Optionally, the three-dimensional virtual object is a three-dimensional mesh model constructed based on three-dimensional human skeleton technology, and the three-dimensional virtual object realizes different external images by wearing different wearable elements.

[0046] (2) Virtual space: It is the scene displayed (or provided) when the application runs on the terminal. This virtual space can be a simulation of a real scene, a semi-simulated and semi-fictional scene, or a purely fictional scene. The virtual space can be any one of a two-dimensional virtual scene, a 2.5D virtual scene, and a three-dimensional virtual scene, and the present application does not limit this.

[0047] (3) Wiring: It refers to the topological structure of the lines of a virtual object. Simply put, it is the way the lines are connected. Taking a virtual object as a virtual character as an example, through wiring, the muscle trend or the structure of the entity of the virtual object can be reflected, which is the key to building a high-quality model.

[0048] (3) Wiring optimization: It is a process of adjusting the layout of the original line segments on a virtual object, such as adding wiring, deleting wiring, adjusting the shape and position of the wiring, etc., to achieve a better effect.

[0049] In the related art, the virtual object obtained by the signed distance field and the marching cubes algorithm can be as shown in Figure 1 Figure (A). The virtual object obtained in this way is uniformly wired, that is, there is no obvious wiring rule at all positions of the virtual object, and manual wiring optimization is still required to serve downstream game applications. The virtual object obtained by manual wiring optimization is as shown in Figure 1 Figure (B), that is, the wiring rule of the virtual object is properly sparse and dense. For example, the patches are dense at the parts that need to move, such as joints (such as the knees), and sparse on the plane (such as the thighs, etc.).

[0050] That is to say, the wiring of the virtual object obtained by the method in the related art is uniform, and manual participation in adjustment is still required for subsequent use, so there will also be problems of a long development cycle and high development cost.

[0051] Based on this, the embodiment of the present application provides a method for training a prediction model, using a virtual object sample with properly sparse and dense wiring that can be used for downstream applications as supervision, that is, using the sample discrete sequence corresponding to the virtual object sample as the training target, so that the trained prediction model can predict the discrete sequence based on the object description, so that a virtual object corresponding to the object description can be constructed based on the discrete sequence. Moreover, due to the high accuracy of the prediction model, the wiring of the virtual object is properly sparse and dense and can be directly used for downstream applications, thus shortening the development cycle and reducing the development cost.

[0052] To facilitate understanding of the method for training the prediction model provided by the embodiment of the present application, the computer system of the method for training the prediction model will be described below first.

[0053] See Figure 2, This figure is a schematic diagram of a computer system for a method of training a prediction model provided by an embodiment of the present application. The computer system 200 includes a plurality of terminal devices 210 and a server 220, and the terminal devices 210 and the server 220 can communicate through a communication network such as.

[0054] Among them, the communication network uses standard communication technologies and / or protocols, usually the Internet, but can also be any network, including but not limited to any combination of Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, private network or virtual private network. In some embodiments, customized or dedicated data communication technologies can be used to replace or supplement the above data communication technologies.

[0055] The terminal device can be various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner, a smart vehicle-mounted device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. A client for running a target application can be installed in the terminal device. The target application can be an application that supports the training of the prediction model, or can also be other applications that support the production of virtual objects, the storage of virtual objects, and the display of virtual objects. The present application does not make any limitations in this regard. In addition, the present application does not make any limitations on the form of the target application, including but not limited to applications (Apps) installed in the terminal device, applets, etc., and can also be in the form of a web page, etc.

[0056] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and a cloud server that provides basic cloud computing services such as big data. The server can be the background server of the above target application, and is used to provide background services for the client of the target application.

[0057] To facilitate understanding of the method for training the prediction model provided by the embodiment of the present application, the application scenario of the method for training the prediction model will be exemplarily introduced below by taking the execution subject of the method for training the prediction model as the server.

[0058] See Figure 3 , This figure is a schematic diagram of an application scenario of a method for training a prediction model provided by an embodiment of the present application. InFigure 3 in which, the application scenario includes Figure 2 a terminal device 210 and a server 220 in

[0059] The terminal device 210 determines a virtual object sample and its corresponding object description sample, and sends the virtual object sample and its corresponding object description sample to the server 220. Alternatively, the user determines a training sample set through the terminal device 210, and the training sample set includes a virtual object sample and its corresponding object description sample. The terminal device 210 sends the server to the server 220. This application does not make specific limitations on this.

[0060] The server 220 obtains a virtual object sample and the object description sample corresponding to the virtual object sample. The virtual object sample is obtained by optimizing the wiring of the virtual object. That is, the wiring rule of the virtual object sample is properly sparse and dense, and can be directly applied to downstream applications. Encode the virtual object sample to obtain a sample discrete sequence. The sample discrete sequence is used to describe the positions of the multiple patches that make up the virtual object sample in the virtual space, and the distribution of the multiple patches included in the virtual object sample is also properly sparse and dense.

[0061] The server 220 predicts the object description sample through an initial prediction model to obtain a predicted discrete sequence, so as to convert the object description sample into a predicted discrete sequence that can be used to construct a virtual object sample. In order to improve the accuracy of the initial prediction model, according to the difference between the predicted discrete sequence and the sample discrete sequence, adjust the model parameters of the initial prediction model, so that the difference between the predicted discrete sequence predicted by the initial prediction model and the sample discrete sequence becomes smaller and smaller, that is, the accuracy of the initial prediction model becomes higher and higher, and thus a prediction model is obtained.

[0062] The prediction model can predict a target object description to obtain a target discrete sequence, and then decode the target discrete sequence based on the decoding method corresponding to the encoding method to obtain a target virtual object corresponding to the target object description, so as to realize obtaining a corresponding target virtual object based on the target object description, and the wiring of the target virtual object is properly sparse and dense and can be directly applied to downstream applications.

[0063] Thus, using the virtual object sample that can be used in downstream applications and has properly sparse and dense wiring as supervision, that is, using the sample discrete sequence corresponding to the virtual object sample as the training target, so that the trained prediction model can predict a discrete sequence based on the object description, so as to be able to construct a virtual object corresponding to the object description based on the discrete sequence. Moreover, due to the high accuracy of the prediction model, the wiring of the virtual object is properly sparse and dense and can be directly used in downstream applications, thus shortening the development cycle and reducing the development cost.

[0064] The training method of the prediction model provided by the embodiments of the present application can be executed by a server. However, in other embodiments of the present application, the terminal device may also have a similar function to the server, so as to execute the training method of the prediction model provided by the embodiments of the present application, or the training method of the prediction model provided by the embodiments of the present application may be jointly executed by the terminal device and the server. This embodiment does not make any limitations in this regard.

[0065] The following is a detailed introduction to a training method of a prediction model provided by the present application through method embodiments.

[0066] See Figure 4 , which is a schematic flowchart of a training method of a prediction model provided by an embodiment of the present application. For the convenience of description, the following embodiments still take the execution subject of the training method of the prediction model as a server as an example for introduction. As Figure 4 shown, the training method of the prediction model includes S401 - S404.

[0067] S401: Obtain virtual object samples and object description samples for describing the virtual object samples.

[0068] The virtual object samples are obtained by optimizing the wiring of virtual objects, such as those obtained by artists through manual wiring. The virtual object samples have a geometric structure, and the topology and wiring of the geometry are adapted to the artistic design concept. For example, at the parts that need to move, such as joints, the patches are dense, and on the plane, the patches are sparse, and the overall wiring is smooth and beautiful. Different from the uniformly wired geometry obtained by the moving cube algorithm. That is to say, the wiring of the virtual object samples is properly sparse and dense and can be directly applied to downstream applications.

[0069] The object description samples are the content for describing the virtual object samples. For example, the three-dimensional model of the virtual object corresponding to the virtual object sample, the point cloud data describing the virtual object corresponding to the virtual object sample, the picture describing the virtual object sample or its corresponding virtual object, the text describing the virtual object sample or its corresponding virtual object, etc. The present application does not make specific limitations in this regard.

[0070] For example, if the virtual object sample is obtained by an artist through manual wiring optimization and is a three-dimensional model for describing the human body, the object description sample can be the text describing the virtual object sample, the planar image of the virtual object sample, the model before the virtual object sample is wired optimized (which can be called a rough model), etc.

[0071] S402: Encode the virtual object samples to obtain sample discrete sequences.

[0072] The virtual object sample includes multiple patches. Patches are usually used to describe the surface of a three-dimensional object. They can be combined by a large number of adjacent patches to construct complex shapes. Encoding the virtual object sample can be to encode each patch included in the virtual object sample respectively to obtain a sample discrete sequence.

[0073] The sample discrete sequence is the discrete sequence corresponding to the virtual object sample. This sample discrete sequence is used to describe the positions of the multiple patches included in the virtual object sample in the virtual space respectively. The sample discrete sequence includes multiple encoded numbers, so that the positions of the multiple patches included in the virtual object sample can be described by the multiple numbers. That is to say, the virtual object sample corresponding to the positions of the multiple patches described by the sample discrete sequence can be restored. Furthermore, compared with the method of describing the virtual object sample by means of a field, the method of describing by the sample discrete sequence is more accurate.

[0074] The embodiments of the present application do not specifically limit the encoding method, and those skilled in the art can set it according to actual needs. For example, taking a triangular patch as an example, the position of the triangular patch in the virtual space can be determined by the vertex positions of the three vertices. Thus, the vertex positions of the vertices of each patch in the virtual object sample can be arranged in a preset order to obtain the sample discrete sequence. Another example is that encoding can also be based on the normal of the patch. Another example is to encode in units of patch blocks. It will be further illustrated by taking A1 - A2 as an example later and will not be elaborated here.

[0075] S403: Predict the object description sample through the initial prediction model to obtain a predicted discrete sequence.

[0076] The embodiments of the present application do not specifically limit the prediction method. For example, the object description sample can be input into the initial prediction model, and the predicted discrete sequence is output after prediction by the initial prediction model. Another example is that the object description sample can be first input into the feature extraction model, the object description sample features are obtained through feature extraction by the feature extraction model, and then the object description sample features are input into the initial prediction model, and the predicted discrete model is obtained through prediction by the initial prediction model.

[0077] The predicted discrete sequence is the discrete sequence obtained through prediction by the initial prediction model. The predicted discrete sequence is used to describe the positions of the multiple patches included in the virtual object described by the object description sample in the virtual space respectively, in order to restore the virtual object sample described by the object description sample through the predicted discrete sequence. The accuracy of the predicted discrete sequence is related to the accuracy of the initial prediction model.

[0078] The initial prediction model is a prediction model that has not been fully trained. Generally, the two models have the same structure but different model parameters. The prediction accuracy of the initial prediction model is generally lower than that of the prediction model. The initial prediction model can be an autoregressive model or the like, and the present application does not make specific limitations thereto. The feature extraction model is a model used for feature extraction, such as a pre-trained neural network model for matching images and texts (Contrastive Language-Image Pre-Training, CLIP), etc. The present application does not make specific limitations thereto and can be set according to the modality of the object description sample.

[0079] The following takes the initial prediction model being an autoregressive model (AR) as an example for illustration. The autoregressive model is a method for processing time series. The expected value of X is equal to a linear combination of one or several lag periods, plus a constant term, plus a random error. See formula (1) for details.

[0080]

[0081] Among them, X t is the value of the time series at time t; c is the constant term (which can be regarded as the mean); p is the order; i is the order number; is the autoregressive coefficient, which describes the influence of past values on the current value; ε t is the random error, such as the white noise error option, and is usually assumed to follow a Gaussian distribution with a mean of 0 and a variance of ε 2 and σ is assumed to be invariant for any t.

[0082] Therefore, the patches can be used as time series for continuous prediction. For example, the coordinates of the next vertex can be predicted based on the already generated vertices until a predicted discrete sequence is obtained.

[0083] As a possible implementation, a Transformer can also be used to construct an autoregressive model, that is, the decoder part of the Transformer is used as the core of the autoregressive model. Thus, through the self-attention mechanism, the dependencies between any positions in the time series can be captured, which enables it to have powerful modeling capabilities when processing time series. The autoregressive model needs to gradually generate the next element of the time series, which requires the model to accurately understand the context information of the time series, and the Transformer model just meets this requirement. Furthermore, by using a Transformer to construct an autoregressive model, the prediction accuracy can be improved.

[0084] S404: Adjust the model parameters of the initial prediction model according to the difference between the predicted discrete sequence and the sample discrete sequence to obtain the prediction model.

[0085] Improving the prediction accuracy of the initial prediction model means that the accuracy of the predicted discrete sequence obtained by the initial prediction model is getting higher and higher. Based on this, the sample discrete sequence is used to describe the virtual object sample and has relatively high accuracy. Therefore, the sample discrete sequence can be used as the true value, and the model parameters of the initial prediction model can be adjusted according to the difference between the predicted discrete sequence and the sample discrete sequence, so as to obtain the prediction model.

[0086] Among them, the prediction model is used to predict the discrete sequence for describing the virtual object. For example, the user can provide a corresponding object description for the desired virtual object, and then the prediction model predicts the object description to obtain the discrete sequence for describing the virtual object corresponding to the object description, so as to decode the required virtual object based on the discrete sequence.

[0087] The embodiments of this application do not specifically limit the training method. For example, the difference between the predicted discrete sequence and the sample discrete sequence becoming smaller and smaller can be used as the training goal, so that the output predicted discrete sequence is getting closer and closer to the sample discrete sequence, that is, the accuracy of the output predicted discrete sequence is getting higher and higher, and thus a prediction model with relatively high accuracy can be obtained.

[0088] The embodiments of this application also do not specifically limit the form of the training goal, that is, the expression form of the loss function is not limited. For example, the difference between the sample discrete sequence minus the predicted discrete sequence is getting closer and closer to zero. Another example is to represent it through the cross-entropy function, as shown in formula (2).

[0089]

[0090] Among them, CrossEntropy is the loss function; N is the number of predicted discrete sequences obtained by training; y i is the sample discrete sequence; is the predicted discrete sequence.

[0091] The embodiments of this application do not specifically limit the end condition of model training. For example, the predicted discrete sequence outputs an end flag, the output predicted discrete sequence reaches the longest sequence threshold, or the number of iterations reaches the maximum number of iterations, etc. Among them, the end flag is a flag used to identify the end of the predicted discrete sequence, such as "E", etc. The longest sequence threshold refers to the longest length of the predicted discrete sequence, which is used to limit the prediction time. The maximum number of iterations refers to the maximum value of the number of iterations during the training process of the model.

[0092] As can be seen from the above technical solutions, a virtual object sample and an object description sample for describing the virtual object sample are obtained. The virtual object sample is obtained by optimizing the wiring of the virtual object, that is, the wiring rule of the virtual object sample is appropriately sparse and dense, belonging to the wiring and topology structure that conforms to the art design concept and can be directly applied to downstream applications. The virtual object sample is encoded to obtain a sample discrete sequence for describing the positions of multiple patches included in the virtual object sample in the virtual space respectively. The distribution of the multiple patches described by the sample discrete sequence is also appropriately sparse and dense. Moreover, the method of describing the multiple patches included in the virtual object sample through the sample discrete sequence is more accurate to improve the accuracy of subsequent training. The object description sample is predicted through an initial prediction model to obtain a prediction discrete sequence, so as to convert the object description sample into a prediction discrete sequence that can be used to construct the virtual object sample. To improve the accuracy of the initial prediction model, according to the difference between the prediction discrete sequence and the sample discrete sequence, the model parameters of the initial prediction model are adjusted, so that the difference between the prediction discrete sequence predicted by the initial prediction model and the sample discrete sequence becomes smaller and smaller, that is, the accuracy of the initial prediction model becomes higher and higher, and thus a prediction model is obtained. Thus, the virtual object sample with appropriately sparse and dense wiring that can be used in downstream applications is used as supervision, that is, the sample discrete sequence corresponding to the virtual object sample is used as the training target, so that the trained prediction model can predict a discrete sequence based on the object description, and based on this discrete sequence, a virtual object corresponding to the object description can be constructed. Moreover, due to the high accuracy of the prediction model, the wiring of the virtual object is appropriately sparse and dense and can be directly used in downstream applications, thereby shortening the development cycle and reducing the development cost.

[0093] As can be seen from the foregoing, the sample discrete sequence is used to describe the positions of each patch in the virtual space respectively, and the virtual object sample generally includes a large number of patches, resulting in a long length of the sample discrete sequence and a long length of the prediction discrete sequence as well. That is to say, whether it is the predicted discrete sequence (such as the prediction discrete sequence) or the encoded discrete sequence (such as the sample discrete sequence), their lengths will be long, resulting in a long prediction time. Even due to hardware limitations, it is only possible to predict virtual objects including hundreds of patches, thus unable to generate slightly complex geometries and having a small application range, such as being unable to meet the requirements of game production.

[0094] Based on this, an embodiment of the present application provides a specific implementation manner of S402, that is, a specific implementation manner of encoding the virtual object sample to obtain a sample discrete sequence, as shown in A1 - A2.

[0095] A1: Divide multiple patches included in the virtual object sample to obtain multiple patch blocks.

[0096] Divide the M patches included in the virtual object sample to obtain N patch blocks, where each patch block includes multiple patches. Both M and N are positive integers, and M is greater than N. For example, if the virtual object sample includes 10,000 patches, every 4 patches can form a patch block, so 2,500 patch blocks can be obtained.

[0097] A2: Encode the positions of each patch block in the virtual space respectively to obtain a sample discrete sequence.

[0098] For the same virtual object sample, compared with encoding by patches, encoding by patch blocks has a smaller quantity. Therefore, by encoding each of the fewer patch blocks respectively, the length of the obtained sample discrete sequence will be shorter.

[0099] See Figure 5 , which is a schematic diagram of a patch and a patch block provided by an embodiment of the present application. As Figure 5 shown in FIG. (A), the virtual object sample includes 8 patches. Encode each patch. Assuming that the length of the number obtained by encoding a patch is 3, the length of the sample discrete sequence obtained by encoding based on patches is 24. As Figure 5 shown in FIG. (B), the virtual object sample is the same as the virtual object sample shown in Figure 5 FIG. (A). The difference is that Figure 5 in FIG. (B), encoding is based on patch blocks. It can be seen that every 4 patches form a patch block. Assuming the same encoding method, the length of the number obtained by encoding a patch block is also 3. Then the length of the sample discrete sequence obtained by encoding based on patch blocks is 6.

[0100] Therefore, instead of encoding each patch, encoding is performed in units of patch blocks composed of multiple patches, that is, multiple patches are combined to obtain patch blocks. Thus, for the same virtual object sample, the number of patch blocks is less than that of patches. Subsequently, encoding in units of patch blocks is equivalent to encoding after compressing multiple patches, which can reduce the quantity of encoding, and further reduce the length of the predicted discrete sequence and the sample discrete sequence, shortening the prediction time. Moreover, when encoding in units of patch blocks with the same length of the sample discrete sequence, it can make the virtual object sample include more patches, be able to generate complex geometries of nearly ten thousand faces or more, and expand the application scope.

[0101] The embodiment of the present application provides a specific implementation manner of A2, that is, a specific implementation manner of encoding the positions of each patch block in the virtual space respectively to obtain a sample discrete sequence. In this implementation manner, taking the encoding method based on vertices as an example for illustration, specifically see B1 - B4.

[0102] B1: For a target patch block among multiple patch blocks, obtain the vertex positions of the vertices of the multiple patches included in the target patch block in the virtual space.

[0103] For ease of explanation, hereinafter, one patch block among the multiple patch blocks, i.e., the target patch block, will be taken as an example for explanation.

[0104] Obtain the vertex positions of the vertices of the multiple patches included in the target patch block in the virtual space. The vertex positions are used to describe the positions of the vertices of the patches in the virtual space.

[0105] See Figure 6 , which is a schematic diagram of a target patch block provided by an embodiment of the present application. In Figure 6 , the target patch block includes 4 patches, and the vertices of each patch are 645, 425, 412, and 523 respectively. If encoded based on the vertex-based method and in units of patches, 12 vertices need to be encoded. However, if encoded based on the vertex-based method and in units of patch blocks, and the vertices included in the patch block are 412356, only 6 vertices need to be encoded, thus greatly reducing the encoding length, and further reducing the length of the sample discrete sequence, etc.

[0106] The following describes the method of encoding in units of patch blocks.

[0107] B2: Sort the multiple vertex positions according to a preset vertex order to obtain a vertex position sequence corresponding to the target patch block.

[0108] A patch can be described based on multiple vertices. For example, a triangular patch can be described by three vertices. Therefore, in order to be able to restore the patch based on multiple vertices during subsequent decoding, it is necessary to sort the multiple vertex positions according to a preset vertex order to obtain a vertex position sequence corresponding to the target patch, so that the multiple vertex positions included in the vertex position sequence can be restored to a patch based on the preset vertex positions subsequently.

[0109] Among them, the preset vertex order is the arrangement order of the multiple vertices included in the patch block set in advance. Continuing to refer to Figure 6 , the preset vertex order can be 412356, or can be 532146, etc. The present application does not make specific limitations on this. Thus, sort the vertex positions corresponding to the vertices of the multiple patches in the patch block according to the preset vertex order to obtain a vertex position sequence corresponding to the target patch block, so as to determine how the multiple vertices form the multiple patches based on the preset vertex order.

[0110] Taking the preset vertex order of 412356 as an example, the vertex position sequence can be the position of vertex 4 in the virtual space, the position of vertex 1 in the virtual space, the position of vertex 2 in the virtual space, the position of vertex 3 in the virtual space, the position of vertex 5 in the virtual space, and the position of vertex 6 in the virtual space.

[0111] It should be noted that although some vertices in the patch block are shared by multiple patches, still taking Figure 6 as an example, vertex 4 not only belongs to triangle patch 456 but also belongs to triangle patch 412. However, the vertices corresponding to each vertex position in the vertex position sequence do not repeat. For example, vertex 4 in the preset vertex order of 412356 only appears once. That is to say, the preset vertex order not only sorts multiple vertex positions but also can remove the vertices that appear multiple times in the patch block, thereby reducing the length of the vertex position sequence and facilitating the reduction of the length of the subsequent sample discrete sequence.

[0112] B3: Taking multiple patch blocks as target patch blocks respectively to obtain the vertex position sequences corresponding to each patch block.

[0113] Taking the multiple patch blocks included in the virtual object sample as target patch blocks respectively and executing B1 - B3, thereby obtaining the vertex position sequences corresponding to each patch block.

[0114] B4: Encoding the vertex position sequences corresponding to each patch block respectively to obtain the sample discrete sequence.

[0115] The embodiments of the present application do not specifically limit the order between multiple patch blocks. The order can be sorted according to the positional relationship between multiple patch blocks, etc., so as to encode the vertex position sequences corresponding to each patch block respectively to obtain the sample discrete sequence.

[0116] Thus, a patch can be determined by multiple vertices. Sorting the multiple vertices corresponding to multiple patches included in the patch block based on the preset vertex order not only removes the repeatedly appearing vertices within the patch block but also can identify how multiple vertices form multiple patches. That is to say, in the process of encoding with the patch block as the unit, adjacent vertices and patches are compressed, thereby reducing the lengths of the predicted discrete sequence and the sample discrete sequence and shortening the prediction time. Moreover, it is also possible to encode with the patch block as the unit under the same sample discrete sequence length, which can enable the virtual object sample to include more patches, be able to generate complex geometries with nearly ten thousand faces or more, and expand the application scope.

[0117] In addition, the vertex positions of each vertex are generally represented by a three-dimensional Cartesian coordinate system, that is, each vertex position is generally represented by three coordinates, namely the x coordinate, the y coordinate, and the z coordinate. That is to say, the length of the sample discrete sequence is generally K×3×3 = 9K, where K is the number of patches included in the virtual object sample, the first 3 indicates that each patch has 3 vertices, and the second 3 indicates that each vertex has three coordinates x, y, and z. Therefore, factors such as the length and size of the coordinates will affect the length of the sample discrete sequence, that is, limit the encoding length of the virtual object sample, and can only encode relatively simple geometries, such as K < 1000, etc.

[0118] Based on this, the embodiments of the present application provide three specific implementation manners of B1, that is, specific implementation manners for obtaining the vertex positions of the vertices of multiple patches included in the target patch block in the virtual space for the target patch block among multiple patch blocks. For the first manner, refer to C1 - C3 specifically, for the second manner, refer to D1 - D2, and for the third manner, refer to E1 - E6.

[0119] Manner 1: Discretize coordinates.

[0120] C1: Perform coordinate discretization on the vertex positions respectively corresponding to multiple patches included in the virtual object sample to obtain updated vertex coordinates respectively corresponding to each vertex position.

[0121] Discretization means mapping finite individuals in infinite space to a finite space to improve the space-time efficiency of the algorithm. Simply put, it is to centralize data with a large distribution and a small quantity. Coordinate discretization means discretizing the coordinates of the Cartesian coordinate system where the virtual space is located, discretizing infinite coordinate values into finite coordinate values. In fact, it is to make the relatively large virtual object sample become denser, making the entire virtual object sample smaller without changing its structure. For example, if the vertex coordinate of a certain vertex is 3.1415926, through coordinate discretization, it may be converted to 3, thus saving 8 - bit encoding length.

[0122] Therefore, coordinate discretization can be performed on the vertex positions respectively corresponding to multiple patches included in the virtual object sample to obtain updated vertex coordinates respectively corresponding to each vertex position.

[0123] C2: For the target patch block among multiple patch blocks, determine the updated vertex coordinates respectively corresponding to the vertices of multiple patches included in the target patch block.

[0124] C3: Determine the updated vertex coordinates respectively corresponding to the vertices of multiple patches included in the target patch block as the vertex positions of the vertices of multiple patches included in the target patch block in the virtual space.

[0125] Continuing with the example of the target patch block, determine the updated vertex coordinates corresponding to the vertices of the multiple patches included in the target patch block, and determine the updated vertex coordinates corresponding to the vertices of the multiple patches included in the target patch block as the vertex positions of the vertices of the multiple patches included in the target patch block in the virtual space.

[0126] Thus, instead of using the true coordinates of each vertex in the virtual space as the vertex positions, the updated vertex coordinates obtained by discretizing the respective true coordinates are used as the vertex coordinates. Through coordinate discretization, not only can continuous coordinates (generally with a longer length) be changed into discrete coordinates (generally with a shorter length), thereby reducing the coding length by reducing the length of the coordinates, but also the coordinate range of the vertex coordinates can be narrowed. For example, the coordinate range of 0 - 255 is discretized to 0 - 63, thereby reducing the coordinate length by reducing the size of the coordinates, and further reducing the coding length.

[0127] Method 2: Divide the sub - space and encode each sub - space independently.

[0128] D1: Divide the space occupied by the virtual object sample in the virtual space to obtain multiple sub - spaces.

[0129] See Figure 7 , which is a schematic diagram of the space occupied by a virtual object sample in the virtual space provided by an embodiment of the present application. As Figure 7 shown in FIG. (A) therein, the size of the space occupied by the virtual object sample (not shown) in the virtual space is 400×200×400. Among them, the coordinates of vertex D can be expressed as (400, - 200, 400), and 9 digits are required to represent one vertex.

[0130] Based on this, the embodiment of the present application divides the virtual space into multiple sub - spaces, and each space is encoded independently. Continuing to refer to Figure 7 FIG. (B) therein, the space occupied by the virtual object sample in the virtual space is divided into 4 sub - spaces, and the size of each sub - space is 200×200×200.

[0131] D2: For the target patch block among the multiple patch blocks, determine the target sub - space where the target patch block is located from the multiple sub - spaces, and determine the initial vertex positions of the vertices of the multiple patches included in the target patch block in the target sub - space respectively.

[0132] Continuing with the example of the target patch block, determine the sub - space where the target patch block is located, that is, the target sub - space, from the multiple sub - spaces that have been divided, and determine the initial vertex positions of the respective vertices corresponding to the multiple patches included in the target patch block in the target sub - space respectively.

[0133] The initial vertex position is used to describe the position of the vertex in the subspace, and the subspace also has a position in the virtual space. Thus, based on these two positions, the position of the vertex in the virtual space can be obtained. That is, the vertex position of the vertex in the virtual space is determined based on the initial vertex position of the vertex in the target subspace and the position of the target subspace. Among them, the calibration methods for the positions of each subspace in the virtual space are the same. For example, the lower left corner of each subspace is used as the origin, and this origin is determined as the position of the subspace, etc. The present application does not make specific limitations on this.

[0134] Continue to refer to Figure 7 In Figure (B) of [reference], the position of subspace IV in the virtual space can be expressed as (200, 0, 0). The position of vertex D in subspace IV is at the upper right corner, and its initial vertex position can be expressed as (200, -200, 200). The position of vertex D in the virtual space can be the sum of the position of subspace IV in the virtual space and the initial vertex position of vertex D in subspace IV, that is, (400, -200, 200).

[0135] It should be noted that as the multiple subspaces obtained by dividing the virtual space are smaller, for example, if the vertices in each subspace can be represented by a single digit, the length of the encoding will be shorter.

[0136] D3: Sort the multiple initial vertex positions according to the preset vertex order to obtain the vertex position sequence corresponding to the target patch block.

[0137] Thus, instead of using the true positions of each vertex as the basis for sorting, the relative positions of each vertex in each subspace, that is, the initial vertex positions, are used as the basis for sorting to obtain the vertex position sequence. It is equivalent to encoding each subspace separately, thereby avoiding the problem of the long encoding length caused by the large virtual space.

[0138] D4: Use the multiple patch blocks as the target patch blocks respectively to obtain the vertex position sequences corresponding to each patch block respectively.

[0139] D5: Encode the vertex position sequences corresponding to each patch block respectively and the space identifiers of the target subspaces corresponding to each patch block to obtain the sample discrete sequence.

[0140] Since each subspace is encoded independently, in order to ensure accurate decoding later, it is also necessary to encode in combination with the space identifiers of each subspace, so that the subspace where the vertex is located can be determined based on the space identifier later, the position of the subspace can be obtained, and then based on the position of the subspace and the initial vertex position in the vertex position sequence, the positions of each vertex in the virtual space can be obtained.

[0141] Thus, the virtual space is divided into multiple sub-spaces, and independent coding is performed within each sub-space. Each sub-space is not affected by others. Therefore, when performing independent coding within a sub-space, a shorter length can be used for coding, thereby reducing the length of the sample discrete sequence.

[0142] Method 3: Combine Method 1 and Method 2.

[0143] E1: Perform coordinate discretization on the virtual space to obtain updated coordinates corresponding to each coordinate.

[0144] Performing coordinate discretization on the virtual space can be regarded as performing coordinate discretization on the space occupied by the virtual object sample in the virtual space. For example, if the space occupied by the virtual object sample in the virtual space is [0, 255], then each coordinate can be converted into a discrete value between [0, 255] through discretization, and this discrete value is the updated coordinate.

[0145] At this time, it is equivalent to dividing the virtual space into P 3 grids, where P is the upper limit value of xyz, that is, each updated coordinate can be any discrete value between [0, P - 1]. At this time, the coding range of the vertex is P 3 .

[0146] E2: Divide the space occupied by the virtual object sample in the virtual space to obtain multiple sub-spaces.

[0147] Dividing the space occupied by the virtual object sample in the virtual space to obtain q 3 sub-spaces, where q is less than p. At this time, the coding range within a sub-space is [0, p / q - 1]. Thus, the coding range is reduced from P 3 to (p / q) 3 +1. Among them, the last digit is a discrete value between [1, q 3 , representing the sub-space coding where it is located.

[0148] For example, if the space occupied by the virtual object sample in the virtual space is [255, 255, 255], at this time, p is equal to 256, and there are 256 3 grids. Divide this virtual space to obtain 64 3 sub-spaces, that is, q is 64. At this time, the coding range of each sub-space changes from any discrete value between [0, P - 1] to any discrete value between [0, 3]. At this time, the coding range becomes 4 3 +1 = 65.

[0149] E3: For the target patch block among multiple patch blocks, determine the target sub-space where the target patch block is located from multiple sub-spaces, and determine the initial vertex positions of the multiple patches included in the target patch block in the target sub-space.

[0150] The position of the target subspace is used to describe the position of the target subspace in the virtual space, and the initial vertex position is used to describe the position of the vertex in the target subspace. Thus, based on the position of the target subspace and the initial vertex position, the position of the vertex in the virtual space can be determined.

[0151] E4: Sort the multiple initial vertex positions according to a preset vertex order to obtain the vertex position sequence corresponding to the target patch block.

[0152] E5: Take the multiple patch blocks as the target patch blocks respectively to obtain the vertex position sequences corresponding to the respective patch blocks.

[0153] E6: Encode the vertex position sequences corresponding to the respective patch blocks and the space identifiers of the target subspaces corresponding to the respective patch blocks to obtain the sample discrete sequence.

[0154] That is to say, by discretizing the virtual space coordinates, the encoding range can be reduced without changing the structure of the virtual object sample. Then, combined with the independent encoding of each subspace, the encoding range can be further reduced. For example, each subspace is encoded with only one digit, and the multi-digit real position can be obtained by combining the position of the subspace. Thus, if encoding is performed based on the space identifier of the subspace (to obtain the position of the subspace) and the vertex position sequence of the subspace, the sample discrete sequence is obtained, and the length of the sample discrete sequence is reduced.

[0155] Therefore, by combining coordinate discretization and independent encoding of each subspace, the encoding range can be further reduced, thereby reducing the length of the sample discrete sequence, and further increasing the number of patches included in the virtual object sample.

[0156] As a possible implementation manner, the embodiment of the present application provides a specific implementation manner of A1, that is, a specific implementation manner of dividing the multiple patches included in the virtual object sample to obtain multiple patch blocks. For details, refer to F1 - F2.

[0157] F1: Classify the multiple patches included in the virtual object sample to obtain the patches belonging to the first category and the patches belonging to the second category.

[0158] Since the multiple patches are located at different positions of the virtual object sample, the motion amplitude at some positions is relatively large, such as the joint position, and the motion amplitude at some positions is relatively small, such as the plane position, etc. The positions with a relatively large motion amplitude need to minimize the loss of detail information to ensure the encoding quality.

[0159] Based on this, the patches can be first divided into multiple categories, and the movement amplitudes corresponding to the patches in different categories are different. Herein, the movement amplitude corresponding to a patch refers to the movement amplitude corresponding to the position where the virtual object sample is located. For the convenience of description, the following takes two categories as an example for illustration, namely the patches belonging to the first category and the patches belonging to the second category. The movement amplitude corresponding to the patches belonging to the first category is greater than or equal to the amplitude threshold, that is, the movement amplitude corresponding to the patches belonging to the first category is larger, and the movement amplitude of the patches belonging to the second category is less than the amplitude threshold, that is, the movement amplitude corresponding to the patches belonging to the second category is smaller.

[0160] F2: Divide the multiple patches included in the virtual object sample based on the categories of the patches to obtain multiple patch blocks.

[0161] Thus, when dividing the patches into patch blocks, the division is not only based on the quantity but also on the categories of the patches. Continuing with the example of two categories, the patches belonging to the same patch block have the same category. The number of patches included in the patch block composed of the patches belonging to the first category is less than the number of patches included in the patch block composed of the patches belonging to the second category. Thus, the patch block with a larger movement amplitude includes fewer patches, and the patch block with a smaller movement amplitude includes more patches. For example, the patch block at the joint includes 6 patches belonging to the first category, and the patch block on the thigh includes 8 patches belonging to the second category, etc. Thus, while ensuring the quality, the coding length is reduced.

[0162] Therefore, in the process of dividing the patches, not only are the patches belonging to the same category divided into the same patch block, but the patch block where the patch with a larger movement amplitude is located includes a smaller number of patches, and the patch block where the patch with a smaller movement amplitude is located includes a larger number of patches. Thus, while reducing the coding length, the coding quality is ensured, and the accuracy of training the prediction model is improved.

[0163] As a possible implementation manner, in the process of training the prediction model, similar object samples can also be introduced to improve the accuracy of predicting the discrete sequence. For details, refer to G1 - G4.

[0164] G1: Calculate the first similarity between the virtual object sample and multiple historical object samples respectively.

[0165] The historical object sample is a virtual object with wiring and is a virtual object whose wiring has been optimized. For example, a retrieval library can be pre - constructed, and the retrieval library includes multiple historical object samples. Then, the first similarity between the virtual object sample and each historical object sample in the retrieval library can be calculated to obtain multiple first similarities.

[0166] The first similarity is used to describe the similarity between the virtual object sample and the historical object sample. As a possible implementation, the similarity between the whole of the virtual object sample and the whole of the historical object sample can be calculated, or the similarity between the partial structures of the virtual object sample and the partial structures of the historical object sample can be calculated. This application does not make specific limitations in this regard.

[0167] See Figure 8 , which is a schematic diagram of a historical object sample and a virtual object sample provided by an embodiment of this application. In Figure 8 , one of the diagrams (A) and (B) shows the historical object sample, and the other shows the virtual object sample. As can be seen from Figure 8 , even if the historical object sample and the virtual object sample are of different species, the similarity of the wiring rules at the top of the head and the eyes is relatively high.

[0168] The embodiment of this application does not specifically limit the calculation method of the similarity. For example, by separately extracting features from the virtual object sample and the historical object sample, the corresponding feature vectors can be obtained, and the cosine similarity, Euclidean distance, Manhattan distance, 2-norm (Euclidean norm, L2 norm) of the two feature vectors can be calculated, etc. Those skilled in the art can set according to actual needs.

[0169] G2: Determine the historical object samples whose first similarity meets the similarity condition as the first similar object samples.

[0170] The embodiment of this application does not specifically limit the similarity condition, such as the highest first similarity, the top three in the first similarity ranking, the first similarity being greater than the similarity threshold, etc.

[0171] After calculating multiple first similarities, determine the historical object samples whose first similarity meets the similarity condition as the first similar object samples, so as to obtain the historical object samples with a relatively high similarity to the virtual object sample and use them as the first similar object samples.

[0172] G3: Encode the first similar object samples to obtain the first similar discrete sequence.

[0173] The embodiment of this application does not specifically limit the encoding method for the first similar object samples, which can be the same as the encoding method for the virtual object samples in the foregoing S402, etc.

[0174] The first similar discrete sequence is the discrete sequence corresponding to the first similar object sample, and this first similar discrete sequence is used to describe the positions of the multiple patches included in the first similar object sample in the virtual space.

[0175] G4: According to the object description sample and the first similar discrete sequence, perform prediction through the initial prediction model to obtain the predicted discrete sequence.

[0176] The first similar discrete sequence is used to guide the initial prediction model to generate a prediction discrete sequence according to the object description sample during the prediction process, so that the prediction discrete sequence can represent the content that satisfies the object description sample and the first similar discrete sequence.

[0177] For example, input the object description sample and the first similar discrete sequence into the initial prediction model, and make a prediction through the initial prediction model to obtain the prediction discrete sequence. Another example is to extract features from the object description sample and the first similar discrete sequence respectively, and input the feature vectors obtained by feature extraction into the initial prediction model, and make a prediction through the initial prediction model to obtain the prediction discrete sequence.

[0178] For the training process of the initial prediction model, reference can be made to the aforementioned S403 and S404, which will not be elaborated here.

[0179] Therefore, during the process of training the prediction model, a first similar object sample that is relatively similar to the virtual object sample can be introduced, and the first similar discrete sequence obtained based on the first similar object sample is used to assist the prediction model in generating the desired geometry, so that the generated geometry is more reasonable and complete, and the generation quality of the prediction discrete sequence is improved. Moreover, the first similar discrete sequence is obtained based on the historical object sample that has been wired and optimized, so that the generated prediction discrete sequence is also equivalent to being wired and optimized, and the subsequent manual wiring and optimization steps can be omitted based on the prediction model, improving the efficiency of wiring and optimization.

[0180] As a possible implementation manner, the embodiment of the present application also provides a way to further increase the number of patches included in the virtual object sample. For details, refer to H1-H4.

[0181] H1: Divide the virtual object sample according to parts to obtain multiple sub-virtual object samples.

[0182] When the number of patches included in the virtual object sample is greater than the number threshold, or in other words, when the size of the virtual object sample is large, the object description sample corresponding to the virtual object sample includes multiple sub-object description samples. Different sub-object description samples describe different parts of the virtual object sample. Taking the virtual object sample as a human three-dimensional model as an example, the head can be described to obtain a sub-object description sample, the legs can be described to obtain a sub-object description sample, and so on. Thus, after describing each part, an object description sample composed of multiple sub-object description samples is obtained.

[0183] The virtual object samples are divided according to each part, each part corresponds to a sub-virtual object sample, and each sub-virtual object sample corresponds to a sub-object description sample. That is to say, when the number of patches included in the virtual object sample is greater than the number threshold, or in other words, when the size of the virtual object sample is large, the virtual object sample can be divided according to the part to obtain multiple sub-virtual object samples. Different sub-virtual object samples describe different parts, and the same part corresponds to a sub-virtual object sample and a sub-object description sample.

[0184] H2: Encode each sub-virtual object sample respectively to obtain the sub-sample discrete sequences corresponding to each sub-virtual object sample respectively.

[0185] The encoding method can refer to the aforementioned S402 and will not be elaborated here.

[0186] Encode each sub-virtual object sample respectively to obtain the sub-sample discrete sequences corresponding to each sub-virtual object sample respectively. The multiple sub-sample discrete sequences constitute the sample discrete sequence.

[0187] H3: Use the initial prediction model to predict each of the multiple sub-object description samples respectively to obtain the sub-prediction discrete sequences corresponding to each sub-object description sample respectively.

[0188] The initial prediction model predicts each sub-object description sample respectively to obtain the sub-prediction discrete sequences corresponding to each sub-object description sample respectively. The multiple sub-prediction discrete sequences constitute the prediction discrete sequence.

[0189] As a possible implementation, where i is a positive integer, the i-th sub-virtual object sample can be encoded first to obtain the sub-sample discrete sequence corresponding to the i-th sub-virtual object sample. Then, use the initial prediction model to predict the i-th sub-object description sample corresponding to the i-th sub-virtual object sample to obtain the sub-prediction discrete sequence corresponding to the i-th sub-object description sample, and store the sub-prediction discrete sequence corresponding to the i-th sub-object description sample in the memory. Before encoding the (i + 1)-th sub-virtual object sample, delete the content related to the i-th sub-virtual object sample from the video memory, such as the sub-sample discrete sequence, sub-prediction discrete sequence, etc., so as to reduce the video memory consumption.

[0190] Thus, by predicting only one part of the virtual object sample each time, it is similar to the sliding window method to continuously predict the virtual object sample. Therefore, only some patches are encoded, which can reduce the video memory consumption, improve the prediction efficiency, and accelerate the training speed of the prediction model.

[0191] H4: Adjust the model parameters of the initial prediction model according to the differences between each sub-prediction discrete sequence and the corresponding sub-sample discrete sequence to obtain the prediction model.

[0192] Adjust the model parameters of the initial prediction model according to the sum of the differences between each sub-prediction discrete sequence and its corresponding sub-sample discrete sequence, so as to obtain a prediction model.

[0193] Thus, when the number of patches included in the virtual object sample is large, the virtual object sample can be divided into multiple sub-virtual object samples, and a corresponding sub-object description sample can be constructed for each sub-virtual object sample, so as to encode and predict each part of the virtual object sample separately, thereby reducing the consumption of video memory and accelerating the training of the prediction model.

[0194] After training to obtain a prediction model, the usage process of the prediction model will be described below. For details, see I1-I3.

[0195] I1: Obtain a target object description.

[0196] The target object description is the description of the virtual object that the user wants to obtain, which can be a 3D model, point cloud data, picture, text, etc. The present application does not make specific limitations on this.

[0197] I2: Predict the target object description according to the prediction model to obtain a target discrete sequence.

[0198] The prediction model is used to predict a discrete sequence for describing a virtual object. For example, the target object description can be input into the prediction model, and through the prediction of the prediction model, a target discrete sequence can be obtained.

[0199] The target discrete sequence is used to describe the positions of multiple patches included in the target virtual object corresponding to the target object description in the virtual space.

[0200] I3: Construct according to the target discrete sequence to obtain the target virtual object corresponding to the target object description.

[0201] As can be seen from the foregoing, the target discrete sequence is used to describe the positions of multiple patches included in the target virtual object corresponding to the target object description in the virtual space. If you want to obtain the target virtual object described by the target discrete sequence, it is also necessary to construct based on the patches described by the target discrete sequence to obtain the target virtual object.

[0202] The embodiments of the present application do not specifically limit the method for obtaining the target virtual object based on the target discrete sequence. For example, multiple patches are obtained from the target discrete sequence based on the decoding method corresponding to the encoding method used in training, so as to construct the target virtual object based on the multiple patches. For details, see J1-J2.

[0203] J1: Obtain a preset decoding method.

[0204] The preset decoding method is used to decode the preset encoding method, and the preset encoding method is the encoding method used to encode the virtual object sample.

[0205] Continue to refer to the foregoing Figure 6 , if the preset encoding method is to encode multiple vertex positions in the order of 412356, the preset decoding method is to construct a patch with the 1st, 2nd, and 3rd bits, a patch with the 3rd, 4th, and 5th bits, a patch with the 1st, 3rd, and 5th bits, and a patch with the 1st, 5th, and 6th bits.

[0206] J2: Decode the target discrete sequence according to the preset decoding method to obtain the target virtual object corresponding to the target object description.

[0207] The target discrete sequence can be an indefinite-length one-dimensional discrete integer sequence. The target discrete sequence can be segmented based on the number of patches included in the patch block to obtain subsequences corresponding to each patch block. For each patch block, use the preset decoding method to decode its corresponding subsequence, that is, obtain vertex positions from the one-dimensional discrete integer sequence in the same way and construct corresponding patches to obtain the target virtual object.

[0208] Taking the decoding process of a patch block as an example, continue to refer to Figure 6 , the preset decoding method is that the 1st, 2nd, and 3rd bits construct a patch, that is, obtain the triangular patch 412, the 3rd, 4th, and 5th bits construct a patch, that is, obtain the triangular patch 235, the 1st, 3rd, and 5th bits construct a patch, that is, obtain the triangular patch 425, and the 1st, 5th, and 6th bits construct a patch, that is, obtain the triangular patch 456. Thus, the triangular patches 412, 235, 425, and 456 form a patch block.

[0209] In addition, for the convenience of encoding and decoding, some special encodings can be added during the encoding and decoding processes, such as S representing start, E representing end, A1 representing the first subspace, A2 representing the second subspace, etc. This application does not make specific limitations, and those skilled in the art can set them according to needs, so as to facilitate the segmentation of the target discrete sequence.

[0210] Refer to Figure 9 , this figure is a schematic diagram of an encoding and decoding method provided by an embodiment of this application. In Figure 9 , taking the virtual object sample as an example, encode the virtual object sample to obtain a sample discrete sequence such as S5372063218...537352E, and decoding based on this sample discrete sequence will also obtain this virtual object sample.

[0211] Thus, by using the preset encoding method and the preset decoding method for encoding and decoding, there is no need to separately train the corresponding encoder and decoder, which is convenient and fast and has high accuracy.

[0212] Thus, after the prediction model is trained, the prediction model can predict the target object description to obtain a target discrete sequence, which includes multiple patches for constructing the target virtual object. Decoding the target discrete sequence yields the target virtual object corresponding to the target object description. That is to say, based on the prediction model and the target object description, the target virtual object corresponding to the target object description can be directly generated, and this target virtual object is optimized in wiring, that is, the target virtual object has appropriate wiring density and can be used in downstream applications, thus shortening the development cycle and reducing the development cost.

[0213] As a possible implementation manner, the embodiment of the present application provides a specific implementation manner of I2, that is, a specific implementation manner of predicting the target object description according to the prediction model to obtain the target discrete sequence. For details, refer to K1-K4.

[0214] K1: Calculate the second similarity between the target object description and multiple historical object samples respectively.

[0215] Calculate the similarity between the target object description and each historical object sample respectively to obtain multiple second similarities, that is, each second similarity is used to describe the similarity between the target object description and a historical object sample.

[0216] The embodiment of the present application does not specifically limit the calculation method of the similarity. For related parts, refer to the description of G1 above and will not be elaborated here.

[0217] K2: Determine the historical object samples with second similarities meeting the similarity condition as the second similar object samples.

[0218] After calculating multiple second similarities, determine the historical object samples with second similarities meeting the similarity condition as the second similar object samples, so as to obtain the historical object samples with relatively high similarity to the virtual object samples and use them as the second similar object samples.

[0219] K3: Encode the second similar object samples to obtain the second similar discrete sequence.

[0220] The embodiment of the present application does not specifically limit the encoding method for the second similar object samples, which can be the same as the encoding method for the virtual object samples in S402 above, etc.

[0221] The second similar discrete sequence is the discrete sequence corresponding to the second similar object sample, and this second similar discrete sequence is used to describe the positions of multiple patches included in the second similar object sample in the virtual space respectively.

[0222] K4: According to the target object description and the second similar discrete sequence, perform prediction through the prediction model to obtain the target discrete sequence.

[0223] As can be seen from the foregoing, in the process of training the prediction model, relatively similar first similar object samples can be introduced, so that in the process of using the prediction model trained in this way, relatively similar second similar object samples can also be introduced. The training process of the prediction model can refer to the foregoing G1-G4 and will not be elaborated here.

[0224] The second similar discrete sequence is used to guide the prediction model to generate the target discrete sequence according to the target object description during the prediction process, so that the target discrete sequence can represent the content that satisfies the target object description and the second similar discrete sequence.

[0225] For example, input the target object description and the second similar discrete sequence into the prediction model, and perform prediction through the prediction model to obtain the target discrete sequence. Another example is to perform feature extraction on the target object description and the second similar discrete sequence respectively, and input the feature vectors obtained by feature extraction into the prediction model, and perform prediction through the prediction model to obtain the target discrete sequence.

[0226] Thus, in the process of using the prediction model, a second similar object sample relatively similar to the target object description can be introduced, and the second similar discrete sequence obtained based on the second similar object sample is used to assist the prediction model in generating the desired geometry, making the generated geometry more reasonable and complete, not only improving the generation quality of the target virtual object described by the target discrete sequence, but also improving the efficiency of wiring optimization.

[0227] As a possible implementation manner, the embodiments of the present application provide a specific implementation manner of I2, that is, a specific implementation manner of obtaining the target discrete sequence by predicting the target object description according to the prediction model. For details, refer to L1-L2.

[0228] As can be seen from the foregoing, the target object description can be a 3D model, point cloud data, picture, text, etc., that is, the target object description can be in multiple modalities. Therefore, in order to ensure that the prediction model can perform predictions on target object descriptions of different modalities, it is necessary to ensure that target object descriptions of different modalities can be mapped to the same feature space.

[0229] For the sake of convenience of description, the following takes two different modalities, namely the first modality and the second modality, as examples for illustration.

[0230] L1: If the target object description is in the first modality, then feature extraction is performed on the target object description according to the first modality encoding model to obtain the first object description feature, and the target discrete sequence is obtained by predicting the first object description feature according to the prediction model.

[0231] The first modality encoding model is used to perform feature extraction on the target object description belonging to the first modality, and the first object description feature is the feature vector obtained by performing feature extraction on the target object description belonging to the first modality through the first modality encoding model.

[0232] L2: If the target object description is in the second modality, then feature extraction is performed on the target object description according to the second modality encoding model to obtain the second object description feature, and the target discrete sequence is obtained by predicting the second object description feature according to the prediction model.

[0233] The first modality encoding model and the second modality encoding model are different models. The second modality encoding model is used to perform feature extraction on the target object description belonging to the second modality, and the second object description feature is the feature vector obtained by performing feature extraction on the target object description belonging to the second modality through the second modality encoding model.

[0234] It should be noted that even though the first object description feature and the second object description feature are obtained by performing feature extraction through different encoding models, the first object description feature and the second object description feature are in the same feature space.

[0235] The embodiments of the present application do not specifically limit the manner in which the target object description features of multiple modalities are obtained through different encoding models and the resulting feature vectors can still be in the same feature space. Taking the pre-trained neural network model for matching images and texts (Constrastive Language-Image Pre-training, CLIP) as an example, the CLIP model uses an image encoder and a text encoder to encode images and texts respectively, and through a contrastive learning method, it can map the image features corresponding to the images and the text mappings corresponding to the texts to the same feature space.

[0236] Based on this, if the target object description can be a 3D model, point cloud data, picture, text, etc., then three modality encoding models can be trained, namely a modality encoding model corresponding to the 3D model and point cloud data (wherein, by sampling the surface of the 3D model, point cloud data can be obtained), a modality encoding model corresponding to the picture, and a modality encoding model corresponding to the text.

[0237] Therefore, in order to improve the user experience, various modal representation methods are provided for the target object description, such as 3D models, point cloud data, pictures, texts, etc. Regardless of the modality of the target object description, the corresponding modal encoding model provided by the embodiments of the present application can be used for feature extraction, so as to map the target object descriptions of various modalities to the same feature space, which is convenient for the prediction model to understand and then make predictions, improving the accuracy of predictions.

[0238] To facilitate further understanding of the technical solution provided by the embodiments of the present application, the following takes the execution subject of the training method of the prediction model provided by the embodiments of the present application as a server as an example to give an overall exemplary introduction to the training method of the prediction model.

[0239] The following first describes the training process of the prediction model.

[0240] See Figure 10 , this figure is a training schematic diagram of a prediction model provided by the embodiments of the present application.

[0241] S1: Obtain virtual object samples and object description samples.

[0242] Among them, the virtual object samples are obtained through wiring optimization, and the object description samples are samples describing the virtual object samples, which can be 3D models, point cloud data, pictures, texts, etc.

[0243] S2: Encode the virtual object samples to obtain sample discrete sequences.

[0244] The sample discrete sequences are used to describe the positions of multiple patches included in the virtual object samples in the virtual space respectively.

[0245] S3: Sample the virtual object samples to obtain the point cloud data of the virtual object samples.

[0246] S4: Retrieve in the retrieval library based on the point cloud data of the virtual object samples to obtain the first similar object samples.

[0247] In order to quickly match the first similar object samples, a retrieval library is pre-constructed in this embodiment. The retrieval library includes the point cloud data corresponding to multiple historical object samples. Then, a pre-trained three-dimensional neural network can be used to convert each point cloud data to the representation in the latent space, that is, the corresponding feature vector is obtained by feature extraction. Thus, the first similarity between the point cloud data of the virtual object samples and the point cloud data of multiple historical object samples is calculated based on the feature vectors. The historical object samples whose first similarity meets the similarity condition are determined as the first similar object samples.

[0248] Assume that S is the retrieval library, that is, the set of all virtual object samples of art handicrafts. The pseudo-code for constructing the retrieval library is:

[0249] for mesh in S:

[0250] point_cloud = sample(mesh)

[0251] latent_code = f(point_cloud)

[0252] Among them, mesh is a historical object sample included in the retrieval library. sample is a sampling operation that samples a certain number of point clouds on the surface of the historical object sample, and point_cloud is the point cloud data corresponding to the historical object sample. f is a pre-trained three-dimensional neural network. The input is the point cloud data, and the output is the corresponding feature vector, that is, latent code.

[0253] S5: Encode the first similar object sample to obtain the first similar discrete sequence.

[0254] Using the aforementioned encoding method, the encodable patch length can be expanded from several hundred patches to tens of thousands of patches, meeting the needs of generating most game assets.

[0255] S6: According to the object description sample and the first similar discrete sequence, perform prediction through the initial prediction model to obtain the predicted discrete sequence.

[0256] S7: Adjust the model parameters of the initial prediction model according to the difference between the predicted discrete sequence and the sample discrete sequence to obtain the prediction model.

[0257] Among them, the first similar discrete sequence can be added to the initial prediction model (such as an autoregressive model) through the attention mechanism to control the initial prediction model to generate the predicted discrete sequence, and it can be an indefinite-length sequence. When the end coding "E" appears in the sequence or the longest sequence limit is reached, it exits.

[0258] Thus, the trained prediction model can predict the discrete sequence corresponding to the virtual object for the object description based on the object description.

[0259] After training the prediction model, the usage process of the prediction model will be described below.

[0260] See Figure 11 , this figure is an application schematic diagram of a prediction model provided by an embodiment of the present application.

[0261] S8: Obtain the target object description.

[0262] S9: Retrieve in the retrieval library based on the target object description to obtain the second similar object sample.

[0263] Specifically, calculate the second similarity between the target object description and multiple historical object samples, and determine the historical object samples whose second similarity meets the similarity condition as the second similar object samples.

[0264] S10: Encode the second similar object samples to obtain a second similar discrete sequence.

[0265] S11: According to the target object description and the second similar discrete sequence, perform prediction through a prediction model to obtain a target discrete sequence.

[0266] S12: Construct according to the target discrete sequence to obtain a target virtual object corresponding to the target object description.

[0267] Specifically, obtain a preset decoding method, which is used to decode a preset encoding method, and the preset encoding method is the encoding method used to encode virtual object samples. Thus, decode the target discrete sequence according to the preset decoding method to obtain a target virtual object corresponding to the target object description.

[0268] See Figure 12 , this figure is a comparison schematic diagram provided by an embodiment of the present application. The target object description is a rough model, as shown in Figure 12 Figure (A) therein, and this rough model can be obtained through the marching cubes algorithm. The target virtual object obtained based on the prediction model is shown in Figure 12 Figure (B) therein, with clearer wiring and proper density, and can be directly applied to downstream applications.

[0269] Thus, the embodiment of the present application can directly generate a corresponding target virtual object that conforms to the game art design style according to target object descriptions in different modalities such as rough models, texts, pictures, and point clouds, greatly reducing the game production threshold and cost, and shortening the game production time. Moreover, it can also be mixed with other marching cubes algorithms, that is, use the marching cubes algorithm to generate a rough model and then use this embodiment to convert it into a target virtual object that meets the requirements of art wiring. Thereby reducing the game production cost, accelerating the game production speed, and making the game production more convenient.

[0270] Regarding the training method of the prediction model described above, the present application also provides a corresponding training device for the prediction model to enable the above-mentioned training method of the prediction model to be applied and implemented in practice.

[0271] See Figure 13 , this figure is a structural schematic diagram of a training device for a prediction model provided by an embodiment of the present application. As shown in Figure 13 , the training device 1300 for the prediction model includes: an acquisition unit 1301, an encoding unit 1302, a prediction unit 1303, and a training unit 1304;

[0272] The obtaining unit 1301 is configured to obtain a virtual object sample and an object description sample for describing the virtual object sample, where the virtual object sample is obtained through wiring optimization;

[0273] The encoding unit 1302 is configured to encode the virtual object sample to obtain a sample discrete sequence, where the sample discrete sequence is used to describe the positions of multiple patches included in the virtual object sample in the virtual space respectively;

[0274] The prediction unit 1303 is configured to predict the object description sample through an initial prediction model to obtain a prediction discrete sequence;

[0275] The training unit 1304 is configured to adjust the model parameters of the initial prediction model according to the difference between the prediction discrete sequence and the sample discrete sequence to obtain a prediction model, where the prediction model is used to predict a discrete sequence for describing a virtual object.

[0276] It can be seen from the above technical solution that a virtual object sample and an object description sample for describing the virtual object sample are obtained. The virtual object sample is obtained by optimizing the wiring of the virtual object, that is, the wiring rule of the virtual object sample is appropriately sparse and dense, belonging to the wiring and topological structure that conforms to the art design concept and can be directly applied to downstream applications. The virtual object sample is encoded to obtain a sample discrete sequence for describing the positions of multiple patches included in the virtual object sample in the virtual space respectively. The distribution of the multiple patches described by the sample discrete sequence is also appropriately sparse and dense. Moreover, the method of describing the multiple patches included in the virtual object sample through the sample discrete sequence is more accurate to improve the accuracy of subsequent training. The object description sample is predicted through the initial prediction model to obtain a prediction discrete sequence, so as to convert the object description sample into a prediction discrete sequence that can be used to construct the virtual object sample. In order to improve the accuracy of the initial prediction model, the model parameters of the initial prediction model are adjusted according to the difference between the prediction discrete sequence and the sample discrete sequence, so that the difference between the prediction discrete sequence predicted by the initial prediction model and the sample discrete sequence becomes smaller and smaller, that is, the accuracy of the initial prediction model is getting higher and higher, and thus a prediction model is obtained. Thus, the virtual object sample with appropriately sparse and dense wiring that can be used for downstream applications is used as supervision, that is, the sample discrete sequence corresponding to the virtual object sample is used as the training target, so that the trained prediction model can predict a discrete sequence based on the object description, so as to construct a virtual object corresponding to the object description based on the discrete sequence. Moreover, due to the high accuracy of the prediction model, the wiring of the virtual object is appropriately sparse and dense and can be directly used for downstream applications, thereby shortening the development cycle and reducing the development cost.

[0277] As a possible implementation manner, the encoding unit 1302 is specifically configured to:

[0278] Divide the multiple patches included in the virtual object sample to obtain multiple patch blocks, where each patch block includes multiple patches;

[0279] Encode the positions of each patch block in the virtual space respectively to obtain the sample discrete sequence.

[0280] As a possible implementation manner, the encoding unit 1302 is specifically configured to:

[0281] For a target patch block among the multiple patch blocks, obtain the vertex positions of the multiple patches included in the target patch block in the virtual space;

[0282] Sort the multiple vertex positions according to a preset vertex order to obtain the vertex position sequence corresponding to the target patch block, where the vertices corresponding to the respective vertex positions in the vertex position sequence do not repeat;

[0283] Take the multiple patch blocks as the target patch block respectively to obtain the vertex position sequences corresponding to each patch block respectively;

[0284] Encode the vertex position sequences corresponding to each patch block respectively to obtain the sample discrete sequence.

[0285] As a possible implementation manner, the encoding unit 1302 is specifically configured to:

[0286] Perform coordinate discretization on the vertex positions corresponding to the multiple patches included in the virtual object sample to obtain the updated vertex coordinates corresponding to the respective vertex positions;

[0287] For a target patch block among the multiple patch blocks, determine the updated vertex coordinates corresponding to the vertices of the multiple patches included in the target patch block;

[0288] Determine the updated vertex coordinates corresponding to the vertices of the multiple patches included in the target patch block as the vertex positions of the multiple patches included in the target patch block in the virtual space.

[0289] As a possible implementation manner, the encoding unit 1302 is specifically configured to:

[0290] Divide the space occupied by the virtual object sample in the virtual space to obtain multiple subspaces;

[0291] For a target patch block among the multiple patch blocks, determine a target subspace where the target patch block is located from the multiple subspaces, and determine the initial vertex positions of the vertices of the multiple patches included in the target patch block in the target subspace respectively. The vertex positions of the vertices in the virtual space are determined based on the initial vertex positions of the vertices in the target subspace and the position of the target subspace in the virtual space;

[0292] Sort the multiple initial vertex positions according to the preset vertex order to obtain a vertex position sequence corresponding to the target patch block;

[0293] Encode the vertex position sequences respectively corresponding to each patch block and the space identifiers of the target subspaces corresponding to each patch block to obtain the sample discrete sequence.

[0294] As a possible implementation manner, the encoding unit 1302 is specifically configured to:

[0295] Classify the multiple patches included in the virtual object sample to obtain patches belonging to the first category and patches belonging to the second category. The movement amplitude of the patches belonging to the first category is greater than or equal to the amplitude threshold, and the movement amplitude of the patches belonging to the second category is less than the amplitude threshold;

[0296] Divide the multiple patches included in the virtual object sample based on the categories of the patches to obtain multiple patch blocks. The categories of the patches belonging to the same patch block are the same, and the number of patches included in the patch block composed of the patches belonging to the first category is less than the number of patches included in the patch block composed of the patches belonging to the second category.

[0297] As a possible implementation manner, the device further includes a similarity calculation unit, configured to:

[0298] Calculate a first similarity between the virtual object sample and multiple historical object samples respectively;

[0299] Determine the historical object samples whose first similarity meets the similarity condition as the first similar object samples;

[0300] Encode the first similar object samples to obtain a first similar discrete sequence;

[0301] The prediction unit 1303 is specifically configured to:

[0302] Predict through an initial prediction model according to the object description sample and the first similar discrete sequence to obtain a prediction discrete sequence.

[0303] As a possible implementation, if the number of patches included in the virtual object sample is greater than the number threshold, the object description sample includes a plurality of sub-object description samples, and different sub-object description samples describe different parts of the virtual object sample;

[0304] The encoding unit 1302 is specifically configured to:

[0305] Divide the virtual object sample according to the parts to obtain a plurality of sub-virtual object samples;

[0306] Encode each of the sub-virtual object samples respectively to obtain a sub-sample discrete sequence corresponding to each of the sub-virtual object samples, and the sample discrete sequence includes sub-sample discrete sequences corresponding to each of the virtual sub-objects;

[0307] The prediction unit 1303 is specifically configured to predict the plurality of sub-object description samples respectively through the initial prediction model to obtain sub-prediction discrete sequences corresponding to the plurality of sub-object description samples;

[0308] The training unit 1304 is specifically configured to adjust the model parameters of the initial prediction model according to the difference between each sub-prediction discrete sequence and the corresponding sub-sample discrete sequence to obtain the prediction model.

[0309] As a possible implementation, the device further includes an application unit, configured to:

[0310] Obtain a target object description;

[0311] Predict the target object description according to the prediction model to obtain a target discrete sequence, and the target discrete sequence is used to describe the positions of a plurality of patches included in the target virtual object corresponding to the target object description in the virtual space;

[0312] Construct according to the target discrete sequence to obtain the target virtual object corresponding to the target object description.

[0313] As a possible implementation, the application unit is specifically configured to:

[0314] Calculate a second similarity between the target object description and a plurality of historical object samples respectively;

[0315] Determine the historical object samples whose second similarity meets the similarity condition as second similar object samples;

[0316] Encode the second similar object samples to obtain a second similar discrete sequence;

[0317] Predict according to the target object description and the second similar discrete sequence through the prediction model to obtain the target discrete sequence.

[0318] As a possible implementation, the application unit is specifically configured to:

[0319] If the target object description is in the first modality, extract features from the target object description according to the first modality encoding model to obtain the first object description feature, and predict the first object description feature according to the prediction model to obtain the target discrete sequence;

[0320] If the target object description is in the second modality, extract features from the target object description according to the second modality encoding model to obtain the second object description feature, and predict the second object description feature according to the prediction model to obtain the target discrete sequence. The first modality and the second modality are different modalities, the first modality encoding model and the second modality encoding model are different models, and the first object description feature and the second object description feature are in the same feature space.

[0321] As a possible implementation, the application unit is specifically configured to:

[0322] Obtain a preset decoding method, where the preset decoding method is used to decode a preset encoding method, and the preset encoding method is the encoding method used to encode the virtual object sample;

[0323] Decode the target discrete sequence according to the preset decoding method to obtain the target virtual object corresponding to the target object description.

[0324] An embodiment of the present application also provides a computer device, which can be a server or a terminal device. Below, the computer device provided by the embodiment of the present application will be introduced from the perspective of hardware implementation. Among them, Figure 14 The structure diagram of the server is shown, Figure 15 The structure diagram of the terminal device is shown.

[0325] See Figure 14, This figure is a schematic diagram of a server structure provided by an embodiment of the present application. The server 1400 may vary significantly due to configuration or performance differences, and may include one or more processors 1422, such as Central Processing Units (CPUs), a memory 1432, and a storage medium 1430 (e.g., one or more mass storage devices) for storing one or more application programs 1442 or data 1444. Among them, the memory 1432 and the storage medium 1430 may be transient storage or persistent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the processor 1422 may be configured to communicate with the storage medium 1430 and execute a series of instruction operations in the storage medium 1430 on the server 1400.

[0326] The server 1400 may further include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0327] The steps performed by the server in the above embodiments may be based on the Figure 14 server structure shown.

[0328] Among them, the processor 1422 is used to perform the following steps:

[0329] Obtain a virtual object sample and an object description sample for describing the virtual object sample, where the virtual object sample is obtained through wiring optimization;

[0330] Encode the virtual object sample to obtain a sample discrete sequence, where the sample discrete sequence is used to describe the positions of multiple patches included in the virtual object sample in the virtual space;

[0331] Predict the object description sample through an initial prediction model to obtain a prediction discrete sequence;

[0332] Adjust the model parameters of the initial prediction model according to the difference between the prediction discrete sequence and the sample discrete sequence to obtain a prediction model, where the prediction model is used to predict a discrete sequence for describing a virtual object.

[0333] Optionally, the processor 1422 may also execute the method steps of any specific implementation of the training method of the prediction model in the embodiments of the present application.

[0334] Refer to Figure 15 , which is a schematic structural diagram of a terminal device provided in an embodiment of the present application. Taking the terminal device as a smart phone as an example for illustration, Figure 15 The block diagram of a part of the structure of the smart phone is shown. The smart phone includes: a Radio Frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a Wireless Fidelity (WiFi) module 1570, a processor 1580, and a power supply 1590 and other components. Those skilled in the art can understand that Figure 15 The smart phone structure shown in

[0335] does not limit the smart phone, and may include more or fewer components than shown in the figure, or combine some components, or different component arrangements. Figure 15 The following specifically introduces each component of the smart phone:

[0336] The RF circuit 1510 can be used to receive and send information or signals during a call. Specifically, after receiving the downlink information of the base station, it is given to the processor 1580 for processing; in addition, the uplink data designed is sent to the base station.

[0337] The memory 1520 can be used to store software programs and modules. The processor 1580 realizes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 1520.

[0338] The input unit 1530 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function control of the smart phone. Specifically, the input unit 1530 may include a touch panel 1531 and other input devices 1532. The touch panel 1531, also known as a touch screen, can collect touch operations of the user on or near it, and drive the corresponding connection device according to a pre-set program. In addition to the touch panel 1531, the input unit 1530 may further include other input devices 1532. Specifically, the other input devices 1532 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.

[0339] The display unit 1540 can be used to display information input by the user or information provided to the user, as well as various menus of the smart phone. The display unit 1540 may include a display panel 1541. Optionally, the display panel 1541 can be configured in the form of, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0340] The smart phone may further include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. As for other sensors that the smart phone may also be configured with, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., they will not be elaborated here.

[0341] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the smart phone. The audio circuit 1560 can transmit the electrical signal converted from the received audio data to the speaker 1561, and the speaker 1561 converts it into a sound signal for output; on the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and then converted into audio data. After the audio data is output to the processor 1580 for processing, it is sent through the RF circuit 1510 to, for example, another smart phone, or the audio data is output to the memory 1520 for further processing.

[0342] The processor 1580 is the control center of the smart phone. It connects various parts of the entire smart phone using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1520, and by calling the data stored in the memory 1520, it executes various functions of the smart phone and processes data. Optionally, the processor 1580 may include one or more processing units.

[0343] The smart phone also includes a power supply 1590 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 1580 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.

[0344] Although not shown, the smart phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0345] In the embodiment of the present application, the memory 1520 included in the smart phone can store a computer program and transmit the computer program to the processor.

[0346] The processor 1580 included in the smart phone can execute the training method of the prediction model provided in the above embodiment according to the instructions in the computer program.

[0347] An embodiment of the present application further provides a computer-readable storage medium for storing a computer program, which is used to execute the training method of the prediction model provided in the above embodiment.

[0348] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which, when running on a computer device, causes the computer device to execute the training method of the prediction model provided in various optional implementation manners of the above aspect.

[0349] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium can be at least one of the following media: read-only memory (abbreviation: ROM), RAM, magnetic disk, or optical disc, etc., which can store computer programs.

[0350] In the embodiments of the present application, the terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0351] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit including the function of that module or unit.

[0352] It should be noted that the embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0353] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Based on the implementation manners provided in the above aspects, the present application can also be further combined to provide more implementation manners. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A prediction model training method, characterized in that: The method comprises: Acquire a virtual object sample and an object description sample for describing the virtual object sample, wherein the virtual object sample is obtained through wiring optimization; Encoding the virtual object sample to obtain a sample discrete sequence, where the sample discrete sequence is used to describe positions of a plurality of face pieces included in the virtual object sample in a virtual space; Predicting the object description sample by using an initial prediction model to obtain a predicted discrete sequence; According to the difference between the predicted discrete sequence and the sample discrete sequence, the model parameters of the initial prediction model are adjusted to obtain a prediction model, and the prediction model is used to predict a discrete sequence for describing a virtual object.

2. The method according to claim 1, characterized in that: The step of encoding the virtual object samples to obtain a discrete sequence of samples includes: Dividing the multiple facets included in the virtual object sample to obtain multiple facet blocks, each of the facet blocks including multiple facets; The positions of the respective surface blocks in the virtual space are encoded respectively to obtain the sample discrete sequence.

3. The method according to claim 2, characterized in that: The step of respectively encoding the position of each of the face blocks in the virtual space to obtain the sample discrete sequence includes: For a target patch block among the multiple patch blocks, obtaining vertex positions of multiple patches included in the target patch block in the virtual space; Sorting the plurality of vertex positions according to a preset vertex order to obtain a vertex position sequence corresponding to the target face block, wherein the vertices corresponding to the vertex positions in the vertex position sequence are not repeated; Taking the plurality of surface blocks as the target surface blocks respectively, and obtaining vertex position sequences corresponding to the surface blocks respectively; The vertex position sequences corresponding to the respective facet blocks are encoded to obtain the sample discrete sequence.

4. The method according to claim 3, characterized in that The step of obtaining, for a target patch block among the plurality of patch blocks, vertex positions of vertices of a plurality of patches included in the target patch block in the virtual space comprises: Discretize the coordinates of the vertex positions respectively corresponding to the multiple facets included in the virtual object sample to obtain updated vertex coordinates respectively corresponding to the vertex positions; For a target patch block among the multiple patch blocks, determining update vertex coordinates corresponding to vertices of multiple patches included in the target patch block; The updated vertex coordinates corresponding to the vertices of the multiple patches included in the target patch block are determined as the vertex positions of the vertices of the multiple patches included in the target patch block in the virtual space.

5. The method according to claim 3, characterized in that: The method further comprises: Dividing the space occupied by the virtual object sample in the virtual space to obtain a plurality of subspaces; The step of obtaining, for a target patch block among the plurality of patch blocks, vertex positions of vertices of a plurality of patches included in the target patch block in the virtual space comprises: For a target patch block among the multiple patch blocks, determine a target subspace where the target patch block is located from the multiple subspaces, and determine initial vertex positions of vertices of multiple patches included in the target patch block in the target subspace, respectively, wherein the vertex positions of the vertices in the virtual space are determined based on the initial vertex positions of the vertices in the target subspace and the positions of the target subspace in the virtual space; The step of sorting the plurality of vertex positions according to a preset vertex order to obtain a vertex position sequence corresponding to the target face block includes: Sorting the plurality of initial vertex positions according to the preset vertex order to obtain a vertex position sequence corresponding to the target face block; The step of encoding the vertex position sequences corresponding to the respective facet blocks to obtain the sample discrete sequence comprises: The vertex position sequences corresponding to the respective patch blocks and the spatial identifiers of the target subspaces corresponding to the respective patch blocks are encoded to obtain the sample discrete sequence.

6. The method according to claim 2, characterized in that The step of dividing the plurality of facets included in the virtual object sample to obtain a plurality of facet blocks comprises: Classifying a plurality of patches included in the virtual object sample to obtain patches belonging to a first category and patches belonging to a second category, wherein the motion amplitude corresponding to the patches belonging to the first category is greater than or equal to an amplitude threshold, and the motion amplitude of the patches belonging to the second category is less than the amplitude threshold; Based on the categories of the patches, the multiple patches included in the virtual object sample are divided to obtain multiple patch blocks, the patches belonging to the same patch block have the same category, and the number of patches included in the patch block composed of the patches belonging to the first category is less than the number of patches included in the patch block composed of the patches belonging to the second category.

7. The method according to claim 1, characterized in that The method further comprises: Calculating first similarities between the virtual object sample and a plurality of historical object samples respectively; Determine the historical object sample whose first similarity meets the similarity condition as a first similar object sample; Encoding the first similar object samples to obtain a first similar discrete sequence; The method of predicting the object description sample by using the initial prediction model to obtain a predicted discrete sequence includes: According to the object description sample and the first similar discrete sequence, prediction is performed using an initial prediction model to obtain a predicted discrete sequence.

8. The method according to claim 1, characterized in that: If the number of facets included in the virtual object sample is greater than the number threshold, the object description sample includes a plurality of sub-object description samples, and different sub-object description samples describe different parts of the virtual object sample; The step of encoding the virtual object samples to obtain a discrete sequence of samples includes: Dividing the virtual object sample according to the parts to obtain a plurality of sub-virtual object samples; Encode each of the sub-virtual object samples respectively to obtain a sub-sample discrete sequence corresponding to each of the sub-virtual object samples, wherein the sample discrete sequence includes a sub-sample discrete sequence corresponding to each of the virtual sub-objects; The step of predicting the description sample by using the initial prediction model to obtain a predicted discrete sequence includes: Predicting the multiple sub-object description samples respectively by using the initial prediction model to obtain sub-prediction discrete sequences corresponding to the multiple sub-object description samples respectively; The step of adjusting the model parameters of the initial prediction model according to the difference between the prediction discrete sequence and the sample discrete sequence to obtain the prediction model comprises: According to the difference between each of the sub-prediction discrete sequences and the corresponding sub-sample discrete sequences, the model parameters of the initial prediction model are adjusted to obtain the prediction model.

9. The method according to claim 1, characterized in that: The method further comprises: Get the target object description; Predicting the target object description according to the prediction model to obtain a target discrete sequence, wherein the target discrete sequence is used to describe positions of a plurality of face pieces included in the target virtual object corresponding to the target object description in the virtual space; The target discrete sequence is constructed to obtain a target virtual object corresponding to the target object description.

10. The method according to claim 9, characterized in that The step of predicting the target object description according to the prediction model to obtain a target discrete sequence includes: Calculating second similarities between the target object description and a plurality of historical object samples respectively; Determine the historical object sample whose second similarity meets the similarity condition as a second similar object sample; Encoding the second similar object sample to obtain a second similar discrete sequence; According to the target object description and the second similar discrete sequence, prediction is performed using the prediction model to obtain the target discrete sequence.

11. The method according to claim 9, characterized in that The step of predicting the target object description according to the prediction model to obtain a target discrete sequence includes: If the target object description is a first modality, extracting features of the target object description according to the first modality coding model to obtain first object description features, and predicting the first object description features according to the prediction model to obtain the target discrete sequence; If the target object description is the second modality, features are extracted from the target object description according to the second modality coding model to obtain second object description features, and the second object description features are predicted according to the prediction model to obtain the target discrete sequence, the first modality and the second modality are different modalities, the first modality coding model and the second modality coding model are different models, and the first object description features and the second object description features are in the same feature space.

12. The method according to claim 9, characterized in that The step of constructing according to the target discrete sequence to obtain a target virtual object corresponding to the target object description includes: Obtaining a preset decoding method, where the preset decoding method is used to decode a preset encoding method, where the preset encoding method is an encoding method used to encode the virtual object sample; The target discrete sequence is decoded according to the preset decoding method to obtain a target virtual object corresponding to the target object description.

13. A prediction model training device, characterized in that: The device comprises: an acquisition unit, an encoding unit, a prediction unit and a training unit; The acquisition unit is used to acquire a virtual object sample and an object description sample used to describe the virtual object sample, wherein the virtual object sample is obtained through wiring optimization; The encoding unit is used to encode the virtual object sample to obtain a sample discrete sequence, where the sample discrete sequence is used to describe the positions of a plurality of facets included in the virtual object sample in the virtual space; The prediction unit is used to predict the object description sample by using an initial prediction model to obtain a predicted discrete sequence; The training unit is used to adjust the model parameters of the initial prediction model according to the difference between the predicted discrete sequence and the sample discrete sequence to obtain a prediction model, and the prediction model is used to predict a discrete sequence for describing a virtual object.

14. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 12 according to the computer program.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 12.

16. A computer program product comprising a computer program, characterized in that When the method is executed on a computer device, the computer device is enabled to execute the method according to any one of claims 1 to 12.