Virtual object data processing method and device, electronic equipment, computer readable storage medium and computer program product

By sampling the interaction process of virtual scenes and analyzing it using deep learning models, the performance index values ​​of virtual objects are determined, which solves the problem of insufficient accuracy in existing technologies and achieves more accurate performance evaluation and improved user experience.

CN121513451APending Publication Date: 2026-02-13SHENZHEN TENCENT NETWORK INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411109119.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies cannot accurately reflect the dynamism and complexity of complex interaction processes when determining the performance index values ​​of virtual objects in virtual scenes, resulting in inaccurate performance index values ​​and affecting user experience.

Method used

By sampling the interaction process of the virtual scene, state data and actual operation data are obtained. A pre-trained deep learning model is used to determine the reference operation probability distribution of each sampling point. Sub-performance index values ​​are calculated by combining the actual operation data, and finally the performance index value of the virtual object is determined.

Benefits of technology

It improves the accuracy of virtual object performance metrics, enabling more sensitive capture of dynamic changes during interaction, thereby enhancing user experience and interaction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121513451A_ABST
    Figure CN121513451A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device of a virtual object, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps that the interaction process of a virtual scene is sampled, a sampling data set is obtained, and the sampling data set comprises state data of a virtual object at multiple sampling points and actual operation data of each sampling point; for each sampling point, determining probability distribution of the sampling point corresponding to the multiple pieces of reference operation data according to the state data; determining sub-performance index values of the virtual object at the sampling points according to the probability corresponding to the reference operation data in the probability distribution in response to the actual operation data queried in the multiple reference operation data; and determining the performance index value of the virtual object in the interaction process according to the sub-performance index values of the virtual object at the plurality of sampling points. Through the method and the device, the accuracy of determining the performance index value of the virtual object in the interaction process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to artificial intelligence technology, and in particular to a data processing method and device for a virtual object, an electronic device, a computer readable storage medium, and a computer program product. BACKGROUND

[0002] In the process in which a user controls a virtual object to interact with other virtual objects in a virtual scene, a performance indicator value of the virtual object in the interaction process is an important factor for the user to consider when making a strategy selection in the virtual scene. Taking a game scene as an example, by determining the performance indicator value of a virtual object controlled by a player in the interaction process, other players can more accurately evaluate the skill level of the player, and thus make different cooperation strategies.

[0003] The related technology determines the performance indicator value of the virtual object in the interaction process depending on a final win or loss result statistical indicator, but this method easily ignores the complexity and dynamics in the virtual scene, resulting in that the determined performance indicator value of the virtual object is not accurate enough to reflect the performance of the virtual object. SUMMARY

[0004] Embodiments of the present application provide a data processing method and device for a virtual object, an electronic device, a computer readable storage medium, and a computer program product, which can improve the accuracy of determining the performance indicator value of the virtual object in the interaction process.

[0005] The technical solutions of the embodiments of the present application are implemented as follows:

[0006] The embodiments of the present application provide a data processing method for a virtual object, which comprises:

[0007] sampling an interaction process of a virtual scene to obtain a sampling data set, wherein the sampling data set comprises state data of the virtual object at a plurality of sampling points and actual operation data of each sampling point;

[0008] for each sampling point, determining a probability distribution of a plurality of reference operation data corresponding to the sampling point according to the state data;

[0009] in response to querying the actual operation data in the plurality of reference operation data, determining a sub-performance indicator value of the virtual object at the sampling point according to a probability corresponding to the reference operation data in the probability distribution;

[0010] determining a performance indicator value of the virtual object in the interaction process according to the sub-performance indicator values of the virtual object at a plurality of sampling points.

[0011] The embodiments of the present application provide a data processing device for a virtual object, which comprises:

[0012] a data sampling module, configured to sample an interaction process of a virtual scene to obtain a sampling data set, wherein the sampling data set comprises state data of a virtual object at a plurality of sampling points and actual operation data of each sampling point;

[0013] a first determining module, configured to determine, for each sampling point, a probability distribution of a plurality of reference operation data corresponding to the sampling point according to the state data;

[0014] a second determining module, configured to, in response to querying the actual operation data in the plurality of reference operation data, determine a sub-performance index value of the virtual object at the sampling point according to a probability corresponding to the reference operation data in the probability distribution;

[0015] a third determining module, configured to determine a performance index value of the virtual object in the interaction process according to sub-performance index values of the virtual object at a plurality of sampling points.

[0016] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:

[0017] a memory, configured to store computer executable instructions;

[0018] a processor, configured to execute the computer executable instructions stored in the memory to implement the data processing method for a virtual object provided in the embodiments of the present application.

[0019] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores a computer program or computer executable instructions, and is configured to implement the data processing method for a virtual object provided in the embodiments of the present application when executed by a processor.

[0020] A computer program product is provided in an embodiment of the present application, and the computer program product comprises a computer program or computer executable instructions, and is configured to implement the data processing method for a virtual object provided in the embodiments of the present application when executed by a processor.

[0021] The embodiments of the present application have the following beneficial effects:

[0022] By acquiring the state data and the actual operation data of the virtual object in the interaction process at a plurality of sampling points, compared with the related art which only relies on the final win or lose result, the data granularity processed is more detailed, more detailed information in the interaction process can be acquired, and therefore more comprehensive reference can be provided for subsequent prediction of the performance indicator value; by the manner of determining the performance indicator value by the sub performance indicator value corresponding to the actual operation data in each sampling point probability distribution and in combination with the sub performance indicator values of the plurality of sampling points, compared with the related art which only determines the performance indicator value of the virtual object based on a simple statistical indicator, different dynamic changes of different situations between the plurality of sampling points can be more sensitively captured, and the determination of the performance indicator value of the virtual object is more accurate. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is an architecture schematic diagram of a data processing system 100 of a virtual object provided by an embodiment of the present application;

[0024] Figure 2 is a structure schematic diagram of a terminal 400 provided by an embodiment of the present application;

[0025] Figure 3A is a first flow schematic diagram of a data processing method of a virtual object provided by an embodiment of the present application;

[0026] Figure 3B is a flow schematic diagram of a training method of a deep learning model provided by an embodiment of the present application;

[0027] Figure 3C is a flow schematic diagram of prediction of an evaluation parameter provided by an embodiment of the present application;

[0028] Figure 3D is a flow schematic diagram of parameter updating of a deep learning model provided by an embodiment of the present application;

[0029] Figure 3E is a flow schematic diagram of determination of a sub performance indicator value provided by an embodiment of the present application;

[0030] Figure 3F is a flow schematic diagram of determination of a performance indicator value provided by an embodiment of the present application;

[0031] Figure 4 is a flow schematic diagram of player performance scoring provided by an embodiment of the present application;

[0032] Figure 5 is a sampling data schematic diagram provided by an embodiment of the present application;

[0033] Figure 6 is a first schematic diagram of a performance indicator value of a virtual object provided by an embodiment of the present application;

[0034] Figure 7 is a second schematic diagram of performance index values of virtual objects provided by an embodiment of the present application.

[0035] It should be noted that the above-mentioned "first", "second" are only used to distinguish different schemes, and do not represent the advantages or disadvantages of the schemes or the priority in the implementation process. DETAILED DESCRIPTION

[0036] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0037] In the following description, "some embodiments" are related to a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0038] In the following description, the terms "first / second / third" are only used to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first / second / third" can interchange the specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.

[0040] Unless otherwise specified, at least one described below refers to one or more cases, and "a plurality of" can refer to two or more cases.

[0041] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0042] The related data collection and processing in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behavior within the scope of authorization of laws and regulations and the personal information subject.

[0043] Before the embodiments of the present application are further described in detail, the terms and phrases involved in the embodiments of the present application are explained, which are applicable to the following explanations.

[0044] 1) Virtual scene: a scene displayed (or provided) by an application when running on a terminal device. The scene can be a simulation environment of the real world, or a virtual environment that is half-simulated and half-fictitious, or a purely fictitious virtual environment. For example, it can be a game scene in a user's game session, which can include player characters, non-player characters, sky, land, sea, etc., and the land can include desert, city, etc. environment elements, and the user can control the player character to interact with other player characters or non-player characters in the virtual scene.

[0045] 2) Virtual object: an image of various people and things that can interact in a virtual scene, or a movable object in a virtual scene. The movable object can be a virtual character, a virtual animal, an animation character, etc., such as a person, an animal, etc. displayed in a virtual scene. The virtual object can be a virtual image in a virtual scene that represents a user. The virtual scene can include multiple virtual objects, each virtual object having its own shape and volume in the virtual scene, occupying a part of the space in the virtual scene. For example, in a game scene, a player character controlled by a user or a non-player character controlled by a machine that interacts with other player characters or non-player characters in the virtual scene.

[0046] 3) State data: feature data of a virtual object in a virtual scene at a sampling point, such as the position, blood volume, skill achievement, etc. of a virtual object and other virtual objects in a virtual scene of a game session at 5 minutes and 10 seconds when the interaction process of the virtual object and other virtual objects is carried out.

[0047] 4) Actual operation data: data of an operation actually performed by a virtual object at a specified sampling point in the process of interacting in a virtual scene. For example, at a certain sampling point in a game scene, an enemy releases a skill directly in front of a virtual object, and the virtual object moves three steps backward.

[0048] 5) Reference operation data: the data of the candidate operation output by the pre-trained deep learning model according to the state data of the sampling point in the process of the virtual object interacting in the virtual scene at a specified sampling point. For example, in a game scene, the enemy releases a skill directly in front of the virtual object, and the reference operation data of the virtual object can be moving two steps to the left, moving two steps to the right, moving two steps backward, or moving one step forward.

[0049] 6) Sub-performance indicator value: an indicator value used to represent the performance of the virtual object in the interaction process of the virtual scene at a sampling point.

[0050] 7) Performance indicator value: an indicator value used to represent the comprehensive performance of the virtual object in the interaction process of the virtual scene in a stage. For example, in a game scene, it can be an indicator value representing the comprehensive performance of the virtual object in one or more interaction processes.

[0051] 8) Actual evaluation parameter: the feedback obtained by the virtual object after performing an operation in a state in the interaction process of the virtual scene. For example, in a game scene, the actual evaluation parameter can be a feedback value, a game score, a survival time, a resource collection amount, etc. For example, when the virtual object hits the enemy, it will get an experience value feedback of 100.

[0052] 9) Predicted evaluation parameter: the feedback predicted by the deep learning model after the virtual object performs an operation in a state in the interaction process of the virtual scene. For example, in a game scene, when the virtual object hits the enemy, the deep learning model predicts that the virtual object will get an experience value feedback of 80.

[0053] 10) Response to: used to represent the conditions or states on which the performed operations depend. When the dependent conditions or states are met, the performed operation(s) can be real-time or have a set delay. In the absence of specific instructions, there is no restriction on the execution order of multiple operations performed.

[0054] The related art relies on simple statistical indicators such as the final victory or defeat result and the final score of the virtual object to determine the performance indicator value of the virtual object in the interaction process. Although this method can quickly obtain the performance indicator value of the virtual object, it easily ignores the dynamics and complexity of the virtual scene interaction process, such as ignoring the impact of individual performance and team cooperation on the final result in the interaction process, and cannot consider the different levels of different virtual objects and the dynamic changes of the same virtual object in the interaction process, resulting in low accuracy of the performance indicator value of the virtual object in a complex scene, thereby reducing the user experience of controlling the virtual object.

[0055] Based on the above analysis, the applicant finds that the prior art cannot accurately evaluate the performance indicator value of the virtual object in the interaction process under a complex scenario. To solve the above problem, the present embodiment provides a data processing method of a virtual object, which can improve the accuracy of determining the performance indicator value of the virtual object in the interaction process.

[0056] To solve the above problem, the present embodiment provides a data processing method and device of a virtual object, an electronic device, a computer readable storage medium and a computer program product, which can improve the accuracy of determining the performance indicator value of the virtual object in the interaction process. The following describes an exemplary application of the electronic device provided by the present embodiment. The electronic device provided by the present embodiment can be implemented as a notebook computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart television, a vehicle-mounted terminal, and various types of terminals. The electronic device can also be implemented as a server. The following describes an exemplary application when the electronic device is implemented as a terminal.

[0057] Referring to Figure 1 , Figure 1 is an architecture diagram of a data processing system 100 of a virtual object provided by the present embodiment. To implement a data processing application of a virtual object, the terminal 400 (an exemplary graphical interface 411 is shown) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0058] Taking scoring the performance of a player in a game scenario as an example, the terminal 400 is configured to send game state data of a plurality of sampling points sampled in the game process of the player and actual operation data of the player at each sampling point to the server 200. The server 200 is configured to determine a probability distribution of a plurality of reference operation data corresponding to the state according to the game state data of each sampling point. If the actual operation data of the player is queried in the plurality of reference operation data, the sub-performance indicator value of the player at each sampling point is determined according to the probability of the corresponding reference operation data in the probability distribution. Then, the performance indicator value of the player in the interaction process is determined according to the sub-performance indicator values of the player at the plurality of sampling points, and the performance indicator value of the player is sent to the terminal 400 through the network 300 and displayed on the graphical interface 411.

[0059] In some embodiments, the terminal 400 is used to acquire game state data and actual operation data of the player at multiple sampling points during the game; determine the probability distribution of multiple reference operation data corresponding to the state based on the game state data at each sampling point; if the player's actual operation data is found in the multiple reference operation data, determine the sub-performance index value of the player at each sampling point based on the probability of the corresponding reference operation data in the probability distribution; and then determine the performance index value of the player during the interaction process based on the sub-performance index value of the player at multiple sampling points, and display it on the graphical interface 411.

[0060] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0061] The data processing method for virtual objects provided in this application is applicable to scenarios where the performance of virtual objects during interaction is evaluated. By determining the performance index values ​​of virtual objects during interaction and scoring their performance, the method can assess the interaction quality, user experience, and usability of virtual objects (such as virtual characters, virtual assistants, etc.), thereby providing feedback and improvement directions for users or developers. This helps developers optimize the design and functionality of virtual objects to provide a more natural, smoother, and more effective interactive experience. Applicable application scenarios include:

[0062] 1) Multiplayer competitive games: In competitive games, users control player characters to interact with player characters or non-player characters in the virtual environment. By determining the performance index values ​​of player characters or non-player characters in the game, users can adjust their individual or team cooperative game strategies in a timely manner, thereby improving the user's gaming experience.

[0063] 2) Virtual Reality (VR) Games: In VR games, players interact with objects in a virtual environment, such as shooting, talking, or solving puzzles. By determining the performance metrics of players during the interaction process, developers can understand the quality of the game experience and thus improve the design.

[0064] 3) Chatbot: In the case of a chatbot having a natural language conversation with a user, by determining the performance indicator value of the chatbot during the interaction, the developer can improve the chatbot's conversation quality in providing conversation services based on the chatbot's performance, thereby improving user satisfaction.

[0065] 4) Virtual Customer Service: The interaction quality of virtual customer service when providing services is very important, and scoring it can help optimize customer service processes and improve user experience.

[0066] 5) Virtual Assistant: In the process of interacting with a virtual assistant in a smart home appliance, such as interacting with a virtual assistant in a smart speaker, by determining the performance indicator value of the virtual assistant during the interaction, the developer can improve the service of the smart home appliance based on the performance of the virtual assistant, which helps to improve the accuracy and responsiveness of the service.

[0067] 6) Online Education: In an online education platform, by determining the performance indicator value of a virtual teacher or tutor during the interaction, corresponding improvement measures can be developed based on the performance of the virtual teacher or tutor to improve teaching quality and learning experience.

[0068] 7) Telemedicine: In the interaction process between a virtual doctor and a patient, by determining the performance indicator value of the virtual doctor during the interaction, the effectiveness of the telemedicine service and patient satisfaction can be measured based on the performance of the virtual doctor.

[0069] 8) Social Media: On social media platforms, the interaction quality of virtual characters with users has an important impact on user experience. By determining the performance indicator value of the virtual character during the interaction, the developer can improve the user interaction quality of the social media platform based on the performance of the virtual character.

[0070] 9) Smart Home: In the interaction process between a virtual housekeeper in a smart home system and a user, by determining the performance indicator value of the virtual housekeeper during the interaction, the user experience and practicality of the smart home system can be improved based on the performance of the virtual housekeeper.

[0071] 10) Autonomous Driving: In the interaction process between a virtual driver in an autonomous driving system and a passenger, the interaction performance score of the virtual driver can timely feedback the opinions and suggestions of the passenger, which helps to improve the user acceptance and safety of the autonomous driving system.

[0072] Reference Figure 2 , Figure 2 is a structural schematic diagram of a terminal 400 provided by an embodiment of the present application, Figure 2The illustrated terminal 400 includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components of terminal 400 are coupled together by a bus system 440, which is configured to permit communication between the components. The bus system 440 includes a power bus, a control bus, and a status bus, among others. For purposes of illustration, the bus system 440 is shown in the illustrated embodiment as a single bus system; however, it should be understood that the bus system 440 can include a combination of multiple buses that function as a bus system. Figure 2

[0073] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general purpose processor, a Digital Signal Processor (DSP), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or the like. The general purpose processor can be a microprocessor, or any conventional processor, or the like.

[0074] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432 that facilitate user input, such as user interface components that facilitate input of a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls, and the like.

[0075] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. The memory 450 optionally includes one or more storage devices remotely located from the processor 410.

[0076] The memory 450 includes volatile memory or non-volatile memory, or both. Non-volatile memory can be read only memory (ROM), volatile memory can be random access memory (RAM). The memory 450 described herein is intended to include any suitable type of memory.

[0077] In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or superset thereof, which are described below.

[0078] The operating system 451 includes a system program for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks. ​

[0079] a network communication module 452 for communicating to other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;

[0080] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., display screens, speakers, etc.) associated with the user interface 430 (e.g., user interfaces for operating the peripheral device and displaying content and information);

[0081] an input processing module 454 for detecting and interpreting one or more user inputs or interactions from one or more input devices 432.

[0082] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software, Figure 2 A data processing apparatus 455 for virtual objects stored in the memory 450 is shown, which can be software in the form of programs and plug-ins, etc., including the following software modules: a data sampling module 4551, a first determining module 4552, a second determining module 4553, and a third determining module 4554. These modules are logical, and thus can be combined or further split according to the implemented functions. The functions of each module will be described below.

[0083] In some embodiments, the terminal or server can implement the virtual object data processing method provided by the embodiments of the present application by running various computer executable instructions or computer programs. For example, the computer executable instructions can be microprogram level commands, machine instructions or software instructions. The computer program can be a native program or a software module in the operating system; can be a native application (APP), i.e. a program that needs to be installed in the operating system to run, such as a game APP; or can be a small program that can be embedded into any APP, i.e. a program that only needs to be downloaded into a browser environment to run. In summary, the above computer executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules or plug-ins.

[0084] Taking a computer program as an application program, in actual implementation, the terminal device 400 is installed and runs an application program supporting a virtual scene. The application program can be any one of a first-person shooting game (FPS), a third-person shooting game, a virtual reality application program, a three-dimensional map program, a card strategy game, a sports game, or a multi-player gun battle survival game. A user uses the terminal device 400 to operate a virtual object located in a virtual scene to perform an activity, which includes but is not limited to at least one of adjusting a body posture, crawling, walking, running, riding, jumping, driving, picking up, shooting, attacking, throwing, and building a virtual building. Illustratively, the virtual character can be a virtual person, such as a simulated person character or an animation character.

[0085] The exemplary application and implementation of the terminal provided by the embodiments of the present application will be described in combination with the data processing method of the virtual object provided by the embodiments of the present application.

[0086] In the following, the data processing method of the virtual object provided by the embodiments of the present application will be described. As described above, the electronic device implementing the data processing method of the virtual object of the embodiments of the present application can be a terminal or a server, or a combination of the two. Therefore, the execution subject of each step will not be repeated in the following description.

[0087] Referring to Figure 3A , Figure 3A is a flowchart of the data processing method of the virtual object provided by the embodiments of the present application. Taking a terminal as an execution subject, the steps shown in Figure 3A will be described in combination.

[0088] In step 101, the interaction process of the virtual scene is sampled to obtain a sampling data set, wherein the sampling data set includes state data of the virtual object at a plurality of sampling points and actual operation data of each sampling point.

[0089] In some embodiments, the virtual scene can be an environment for interaction of a virtual object (for example, a player character or a non-player character in a game), for example, an environment for a game character to fight in a virtual scene. The interaction process of the virtual scene is a process of realizing interaction of a player character in a virtual scene with other player characters or non-player characters by a user controlling actions of the player character in the game.

[0090] In some embodiments, the interactive process of the virtual scene can be sampled at a set time interval to obtain a sampling data set. Here, the state data represents the attribute features of the virtual objects in the virtual scene at the sampling points, for example, in a game scene, the attribute features of the virtual objects can include the features of attributes such as position, speed, health value, energy value, etc.; the actual operation data refers to the data of the actual operation of the virtual object at the sampling point in the process of interaction between the virtual object and other virtual objects in the virtual scene, which can include the type, target, intensity, etc. of the operation, for example, the specific behavior instruction data of the virtual object in the game scene, such as moving, attacking, using items, etc.

[0091] As an example, in a game scene, the pre-set time interval can be one second, that is, the interactive process of the virtual scene is sampled at a frequency of once per second, and the game state data and the actual operation data of the player at each sampling point are collected as a sampling data set. Among them, the game state data is the slice data at the sampling time point, for example, in the virtual scene of a game match, the positions, blood volumes, skill achievements, etc. of the virtual objects and other virtual objects when the interactive process of the virtual objects and other virtual objects reaches 5 minutes and 10 seconds; and the actual operation data is the operation set data of the virtual object within one second from the sampling time point.

[0092] As an example, refer to Figure 5 , Figure 5 is a sampling data diagram provided by the embodiments of the present application, taking the performance evaluation of the virtual objects of two camps (red and blue) in a virtual scene as an example, each of the red and blue camps includes 4 virtual objects, after the start of the interactive process of the virtual scene, the terminal samples each virtual object of the two camps, and sends the obtained sampling data set to the server. As shown in Figure 5 , the sampling data set includes state data 501 and actual operation data 502, wherein the state data 501 includes the position, blood volume, skill achievement data of each virtual object of the red and blue camps at the sampling point, and the actual operation data 502 is the actual operation data of each virtual object of the red and blue camps at the sampling point.

[0093] In step 102, for each sampling point, the probability distribution of the plurality of reference operation data corresponding to the sampling point is determined according to the state data.

[0094] In some embodiments, the plurality of reference operation data corresponding to the sampling point can be discrete actions or continuous actions. Here, the discrete action refers to an action that can be clearly distinguished and is not continuous, for example, the virtual object moves in different directions (up, down, left, right) in the game; the continuous action refers to the value range of the action of the virtual object being a continuous interval, for example, the action of the virtual object lifting the arm.

[0095] As an example, when a virtual object in a virtual scene drives an airplane, it is necessary to adjust the flight height and select the flight mode at the same time, where the flight height is a continuous action, for example, it can be set from the ground to the maximum flight height, and the flight mode is a discrete action, for example, it can be set to manual mode or automatic cruise mode, etc.

[0096] In some embodiments, for the state data of each sampling point, key features can be extracted from the state data, and the probability distribution of the virtual object taking different reference operation data is predicted based on the extracted key features. The probability distribution of multiple reference operation data based on the state data of the sampling point can be achieved by training a probability prediction model using a machine learning algorithm.

[0097] In some embodiments, the "determining, for each sampling point, the probability distribution of the sampling point corresponding to multiple reference operation data according to the state data" in step 102 can be achieved by performing the following processing for each sampling point: determining the probability of the virtual object executing multiple reference operation data respectively when the virtual object is in the state corresponding to the state data, and combining the probabilities of executing multiple reference operation data into a probability distribution.

[0098] In some embodiments, for the state data of each sampling point, the virtual object candidate has multiple operations, and when the virtual object is in the state corresponding to the state data of the sampling point, the probability of the virtual object executing multiple reference operation data respectively can be determined, and then the probabilities of the virtual object executing multiple reference operation data are combined into a probability distribution.

[0099] As an example, in a game scene, it is assumed that an enemy located in front of a player character releases a skill to the player character. At this time, if the player character moves backward, it can avoid the skill released by the enemy. It is determined that when the reference operation data is the player character moving backward, the probability of executing the reference operation data is P(backward) = 0.6. If the player character moves left or right, there is a probability of avoiding the skill released by the enemy. It is determined that when the reference operation data is the player character moving left, the probability of executing the reference operation data is P(left) = 0.20, and when the reference operation data is the player character moving right, the probability of executing the reference operation data is P(right) = 0.15. If the player character moves forward or does not move, it will be injured. It is determined that when the reference operation data is the player character moving forward or not moving, the probability of executing the reference operation data is P(forward or not moving) = 0.05. Therefore, the probability distribution of the sampling point corresponding to multiple reference operation data can be represented as P = {0.60(backward), 0.20(left), 0.15(right), 0.05(forward or not moving)}.

[0100] In some embodiments, the probability distribution of the plurality of reference operation data is determined by a pre-trained deep learning model.

[0101] As an example, the pre-trained deep learning model can be a Multilayer Perceptron (MLP), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Variational Autoencoder (VAE), a Transformer, a Long Short-Term Memory (LSTM), a Proximal Policy Optimization (PPO), or the like.

[0102] Referring to Figure 3B , Figure 3B is a flowchart of a method for training a deep learning model according to an embodiment of the present application. The deep learning model can be trained by performing steps 201 to 206 of the method. Figure 3B The following will be described in detail.

[0103] In step 201, the interaction process of the sample virtual scene is sampled to obtain a first sample state of the sample virtual object at a first sampling point.

[0104] In some embodiments, the sample virtual scene can be an environment for the sample virtual object (e.g., a player character or a non-player character in a game) to interact, for example, an environment for a game character to fight in the sample virtual scene. The interaction process of the sample virtual scene is a process of realizing the interaction of the player character with other player characters or non-player characters in the sample virtual scene by the user controlling the actions of the player character in the game.

[0105] In some embodiments, the interaction process of the sample virtual scene can be sampled at a pre-set time interval. Here, the first sample state of the sample virtual object at the first sampling point represents the values of various attributes of the virtual object in the virtual scene at the first sampling point.

[0106] As an example, in a game scenario, the preset time interval can be one second, that is, the interaction process of the sample virtual scene is sampled at a frequency of once per second to obtain a first sample state of the sample virtual object at a first sampling point. The first sample state of the game is the state of the virtual object in the virtual scene at the sampling time point, for example, in the sample virtual scene of the game match, the interaction process of the sample virtual object and other sample virtual objects is at 5 minutes and 10 seconds, the attribute values of the sample virtual object and other sample virtual objects such as position, blood volume, skill achievement, etc.

[0107] In step 202, the probability distribution of the sample virtual object performing the plurality of sample reference operations respectively in the first sample state is determined by the initialized deep learning model.

[0108] In some embodiments, the probability of the sample virtual object performing the plurality of sample reference operations respectively in the first sample state can be determined by the initialized deep learning model, and the probabilities of performing the plurality of sample reference operations are combined into a probability distribution.

[0109] As an example, the initialized deep learning model can be a fully connected neural network, a convolutional neural network, a recurrent neural network, a variational autoencoder, a Transformer, a long short-term memory network, and a proximal policy optimization, etc. These models are designed for reinforcement learning tasks and can directly output the probability distribution of actions.

[0110] As an example, the actor model in the actor-critic (AC) model can be used as the initialized deep learning model to determine the probabilities of the sample virtual object performing the plurality of sample reference operations respectively in the first sample state, that is, p1=P(operation 1), p2=P(operation 2), p3=P(operation 3)…pn=P(operation n), and the probabilities of performing the plurality of sample reference operations are combined into a probability distribution, which can be represented as P={p1, p2, p3…pn}.

[0111] In some embodiments, the probability of the sample virtual object performing the plurality of sample reference operations respectively in the first sample state is determined by the initialized deep learning model, and the probabilities of performing the plurality of sample reference operations are combined into a probability distribution.

[0112] In some embodiments, when the sample virtual object is in the first sample state, firstly, specific features and attributes of the virtual object in the first sample state of the sample virtual object are obtained, such as the position, speed, resource possession, life value, etc. of the sample virtual object; then, according to the interaction rules and strategies of the virtual scene, a plurality of candidate sample reference operations are determined, such as moving, attacking, defending or collecting resources, etc.; finally, according to the specific features and attributes of the virtual object in the first sample state, the probabilities of the sample virtual object executing the plurality of sample reference operations respectively when in the first sample state are determined, and the probabilities of executing the plurality of sample reference operations are combined into a probability distribution.

[0113] As an example, when the sample virtual object is in the first sample state, the probabilities of the sample virtual object executing the plurality of sample reference operations at the first sampling point are determined by the initialized deep learning model as follows: the probability of executing the sample reference operation 1 is P(sample reference operation 1) = 0.3, the probability of executing the sample reference operation 2 is P(sample reference operation 2) = 0.4, the probability of executing the sample reference operation 3 is P(sample reference operation 3) = 0.2, and the probability of executing the sample reference operation 4 is P(sample reference operation 4) = 0.1. Therefore, the probability distribution of the sample virtual object executing the plurality of sample reference operations respectively in the first sample state can be expressed as P = {0.3(sample reference operation 1), 0.4(sample reference operation 2), 0.2(sample reference operation 3), 0.1(sample reference operation 4)}.

[0114] In step 203, one sample reference operation is selected from the plurality of sample reference operations as the first sample operation, the first sample operation is executed, and the second sample state of the sample virtual object at the second sampling point is obtained, wherein the second sample state is the state updated after the first sample operation is executed in the first sample state.

[0115] In some embodiments, when the sample virtual object is in the first sample state, one sample reference operation is selected from the plurality of sample reference operations predicted by the initialized deep learning model, as the first sample operation and executed, and the first sample state of the first sampling point of the sample virtual scene is updated to the second sample state of the second sampling point.

[0116] In some embodiments, the "selecting one sample reference operation from the plurality of sample reference operations as the first sample operation" in step 203 can be implemented by executing any one of the following processes: selecting any one sample reference operation from the plurality of sample reference operations as the first sample operation; selecting one sample reference operation from the plurality of sample reference operations as the first sample operation according to a preset condition, wherein the preset condition includes the sample reference operation corresponding to the maximum probability in the probability distribution.

[0117] In other embodiments, any one of the plurality of sample reference operations can be selected as the first sample operation.

[0118] As an example, any one of the plurality of sample reference operations of the above example, i.e., from sample reference operation 1, sample reference operation 2, sample reference operation 3, and sample reference operation 4, can be selected as the first sample operation. For example, sample reference operation 3 can be selected as the first sample operation, or sample reference operation 4 can be selected as the first sample operation.

[0119] In some embodiments, one of the plurality of sample reference operations can be selected as the first sample operation according to a preset condition, wherein the preset condition includes a sample reference operation corresponding to a maximum probability in a probability distribution.

[0120] As an example, in the four sample reference operations of the above example, the probability of executing sample reference operation 1 is P(sample reference operation 1) = 0.3, the probability of executing sample reference operation 2 is P(sample reference operation 2) = 0.4, the probability of executing sample reference operation 3 is P(sample reference operation 3) = 0.2, and the probability of executing sample reference operation 4 is P(sample reference operation 4) = 0.1. When the preset condition is a sample reference operation corresponding to a maximum probability in a probability distribution, sample reference operation 2 is selected as the first sample operation because the probability value of sample reference operation 2 is greater than that of the other three sample reference operations.

[0121] The embodiments of the present application can select one of the plurality of sample reference operations as the first sample operation, or select one of the plurality of sample reference operations as the first sample operation according to a preset condition, so that the developer can formulate different operation selection strategies according to the needs, and the operation selection is suitable for more situations, so that the subsequent analysis of the virtual object is more personalized, thereby improving the accuracy of the evaluation of the performance index value of the virtual object.

[0122] In step 204, an actual evaluation parameter of executing the first sample operation in the first sample state is obtained.

[0123] In some embodiments, when the sample virtual object in the sample virtual scene executes the first sample operation in the first sample state, corresponding environmental feedback will appear in the sample virtual scene. Here, the actual evaluation parameter refers to the feedback obtained after the virtual object executes an operation in a state during the interaction process in the virtual scene. For example, in a game scene, the evaluation parameter can be a feedback value, a game score, a survival time, a resource collection amount, etc. After executing the first sample operation in the first sample state, changes in feedback, the position of the virtual object, the life value, resources, etc. will appear in the game environment.

[0124] In step 205, the parameter prediction is performed according to the second sample state to obtain a predicted evaluation parameter.

[0125] In some embodiments, the parameter prediction can be performed according to the second sample state by the initialized deep learning model to obtain the predicted evaluation parameter.

[0126] As an example, the parameter prediction can be performed according to the second sample state by a Critic model in the actor-critic model as the initialized deep learning model to obtain the predicted evaluation parameter. In the game scenario, the predicted evaluation parameter can be an expected feedback of the sample virtual object, an expected game score, an expected survival time, an expected resource collection amount, etc.

[0127] In some embodiments, referring to Figure 3C , Figure 3C is a flowchart of the predicted evaluation parameter provided by the embodiments of the present application. Figure 3B Step 205 of the above embodiment can be implemented by steps 2051-2052 of the above embodiment, which are specifically described as follows. Figure 3C

[0128] In step 2051, the state data in the second sample state is subjected to feature extraction by the initialized deep learning model to obtain state features.

[0129] In some embodiments, the state data of the sample virtual scene in the second sample state is first obtained, and then the state data in the second sample state is subjected to feature extraction by the initialized deep learning model to obtain state features.

[0130] As an example, in the game scenario, the state data in the second sample state in the game virtual scene can include the position of the player character, the health value, the resource amount, the game time, the enemy state, etc.

[0131] In some embodiments, before the state data in the second sample state is subjected to feature extraction by the initialized deep learning model to obtain state features, the state data of the sample virtual scene in the second sample state can be preprocessed to be better used for feature extraction.

[0132] ​In some embodiments, the preprocessing of the state data of the sample virtual scene in the second sample state includes the following processing: first, data cleaning is performed on the state data of the sample virtual scene in the second sample state, to check whether there are errors or abnormal values in the data, such as missing values, abnormally large or small values, etc., and methods such as interpolation, mean filling, removal or replacement of abnormal values are used to process these problems; then, data standardization or normalization processing is performed to reduce the dimension difference between different features, where standardization is to scale the data to a fixed range based on a normal distribution (for example, mean value is 0 and standard deviation is 1), and normalization is to scale the data to another given range (for example, between 0 and 1); finally, the category features are converted into one-hot encoding or label encoding form.

[0133] In step 2052, a parameter prediction is performed based on the state features, to obtain a predicted evaluation parameter.

[0134] In some embodiments, the initialized deep learning model is called based on the state features to perform the parameter prediction, to obtain the predicted evaluation parameter.

[0135] Continuing to refer to Figure 3B , the step 205 is continued to be described.

[0136] In step 206, the parameters of the initialized deep learning model are updated according to the difference between the actual evaluation parameter and the predicted evaluation parameter, to obtain a pre-trained deep learning model.

[0137] In some embodiments, according to the difference between the actual evaluation parameter and the predicted evaluation parameter, a loss value can be determined by using a loss function, and according to the loss value, the parameters of the initialized deep learning model are updated by using a back propagation algorithm, to obtain the pre-trained deep learning model.

[0138] In some embodiments, referring to Figure 3D , Figure 3D is a flowchart of parameter updating of a deep learning model provided by an embodiment of the present application. Figure 3B The step 206 of Figure 3D may be implemented by steps 2061 to 2062 of

[0139] In step 2061, a loss value is determined according to the difference between the actual evaluation parameter and the predicted evaluation parameter.

[0140] In some embodiments, the difference between the actual evaluation parameter and the predicted evaluation parameter can be taken as the difference between the actual evaluation parameter and the predicted evaluation parameter. According to the difference between the actual evaluation parameter and the predicted evaluation parameter, a loss value can be determined by using a loss function.

[0141] As an example, the loss function used to determine the loss value can be a Mean Squared Error (MSE), a Root Mean Squared Error (RMSE), a Mean Absolute Error (MAE), a Cross-Entropy Loss, or the like.

[0142] In step 2062, according to the loss value, the parameters of the initialized deep learning model are updated to obtain a pre-trained deep learning model.

[0143] In some embodiments, the loss value determined according to the difference between the actual evaluation parameter and the predicted evaluation parameter can be used to determine the gradient of the deep learning model using a backpropagation algorithm, and the parameters of the initialized deep learning model are updated according to the principle of gradient descent to obtain a pre-trained deep learning model.

[0144] The embodiments of the present application can minimize the loss value by adjusting the parameters of the deep learning model, so that the prediction result of the deep learning model for the evaluation parameter is closer and closer to the true value, thereby improving the accuracy of the deep learning model in predicting the state. At the same time, by adjusting the parameters of the deep learning model, the loss value is more effectively reduced, which can reduce the time required for training the deep learning model, thereby saving training resources.

[0145] Continuing to refer to Figure 3A , the step 102 is explained.

[0146] In step 103, in response to querying the actual operation data in the plurality of reference operation data, the probability of the corresponding reference operation data in the probability distribution is determined as the probability of the actual operation data, and the sub-performance indicator value of the virtual object at the sampling point is determined according to the probability of the actual operation data.

[0147] In some embodiments, by traversing the plurality of reference operation data corresponding to the state data of the sampling point, it is determined whether there is actual operation data of the virtual object in the plurality of reference operation data. If the actual operation data is queried in the plurality of reference operation data, the probability of the corresponding reference operation data is taken as the belonging probability of the actual operation data, and the sub-performance indicator value of the virtual object at the sampling point is determined according to the belonging probability of the actual operation data.

[0148] As an example, if the belonging probability of the queried actual operation data of the virtual object in the probability distribution of the plurality of reference operation data is p a , and the maximum probability in the probability distribution of the discrete action is p max , then the sub-performance indicator value of the virtual object at the sampling point is score = p a / p max, wherein the score ranges from 0 to 1. For example, the maximum probability in the probability distribution P of the plurality of reference operation data is 0.60, and the probability value of the probability p a of the actual operation action of the virtual object in the probability distribution of the plurality of reference operation data is 0.20, the sub-performance indicator value of the virtual object at the sampling point is 0.20 / 0.60=1 / 3.

[0149] In some embodiments, referring to Figure 3E , Figure 3E is a flowchart of determining a sub-performance indicator value provided by the embodiments of the present application. Figure 3A Step 103 of the method can be implemented by steps 1031 to 1032 of the method. Figure 3E

[0150] In step 1031, the probability corresponding to the reference operation data in the probability distribution is taken as the target probability of the virtual object.

[0151] As an example, it is assumed that an enemy located in front of the player character releases a skill to the player character, at this time, if the player character moves backward, the skill released by the enemy can be avoided, and the probability of the player character moving backward is 0.6; if the player character moves leftward or rightward, there is a probability to avoid the skill released by the enemy, the probability of the player character moving leftward is 0.2, and the probability of the player character moving rightward is 0.15; if the player character moves forward, it will be damaged, and the probability of the player character moving forward is 0.05; according to the game state data at this time, the probability distribution of the plurality of reference operation data is P={0.60(backward), 0.20(leftward), 0.15(rightward), 0.05(forward)}; if the player's operation at this time is to control the player character to move leftward, the probability value of the probability p a of the actual operation action of the player belongs to is 0.20.

[0152] In step 1032, the ratio of the target probability to the maximum probability in the probability distribution is taken as the sub-performance indicator value of the virtual object at the sampling point.

[0153] As an example, in the above example, since the maximum probability in the probability distribution P of the plurality of reference operation data is 0.60, and the probability value of the probability p a of the actual operation action of the player belongs to is 0.20, therefore, in this state, the sub-performance indicator value of the player is 0.20 / 0.60=1 / 3.

[0154] ​The embodiment of the application determines the probability to which the actual operation data of the virtual object belongs by traversing the probability distribution of the plurality of reference operation data and taking the probability of the plurality of reference operation data in the probability distribution as a reference, determines the sub-performance index value of the virtual object at the sampling point according to the probability of the corresponding reference operation data in the probability distribution when the actual operation data is queried in the plurality of reference operation data, and can quickly locate the probability corresponding to the actual operation data of the virtual object, thereby improving the speed of evaluating the sub-performance index value of the virtual object at the sampling point.

[0155] In some embodiments, in response to the actual operation data not being queried in the plurality of reference operation data, a preset index value is determined as the sub-performance index value of the virtual object at the sampling point.

[0156] In some embodiments, when the plurality of reference operation data corresponding to the state data of the sampling point is traversed, if the actual operation data is not queried in the plurality of reference operation data, a preset index value is determined as the sub-performance index value of the virtual object at the sampling point.

[0157] As an example, the probability distribution of the plurality of reference operation data determined according to the state data in the above example is P={0.60(backward), 0.20(left), 0.15(right), 0.05(forward)}; if the player does not move in the state at the sampling point, it indicates that the actual operation data of the player is not queried in the plurality of reference operation data, and if the preset index value is 0, it can be determined that the sub-performance index value of the player at the sampling point is the preset index value, that is, 0.

[0158] The embodiment of the application determines the probability to which the actual operation data of the virtual object belongs by traversing the probability distribution of the plurality of reference operation data and taking the probability of the plurality of reference operation data in the probability distribution as a reference, determines the sub-performance index value of the virtual object at the sampling point according to the probability of the corresponding reference operation data in the probability distribution when the actual operation data is queried in the plurality of reference operation data, and can quickly locate the probability corresponding to the actual operation data of the virtual object, thereby improving the speed of evaluating the sub-performance index value of the virtual object at the sampling point.

[0159] In step 104, the performance index value of the virtual object in the interaction process is determined according to the sub-performance index value of the virtual object at the plurality of sampling points.

[0160] In some embodiments, the performance index value of the virtual object in different periods of the interaction process can be determined according to the sub-performance index value of the virtual object at the plurality of sampling points corresponding to the different periods.

[0161] The interactive process of the virtual scene can include a pre-stage and a post-stage, both of which include multiple stages (i.e., at least two stages), and the multiple stages included in the pre-stage precede the multiple stages included in the post-stage.

[0162] Taking an example in which the interactive process includes M stages, the pre-stage can include the first N stages of the M stages, i.e., the first stage to the Nth stage, and the post-stage can include the N+1th stage to the Mth stage of the M stages, where M and N are positive integers, and M is greater than N.

[0163] As a special case, when M = N+1, the post-stage includes the N+1th stage to the Mth stage of the M stages, for example, when M = N+1 = 3, the post-stage includes the third stage to the third stage of the three stages, i.e., the post-stage has only one stage. When M is greater than N+1, it can be ensured that the post-stage includes at least two stages, for example, when M = 4 and N+1 = 3, the post-stage includes the third stage to the fourth stage of the four stages, i.e., the post-stage includes two stages.

[0164] As an example, the interactive process can occur in a continuous time period, for example, in the case of a player playing a game continuously for two hours, the total duration of the interactive process is two hours, and the player plays a total of 10 games in two hours, then the first 5 games of the 10 games can be regarded as the pre-stage, and the last 5 games can be regarded as the post-stage, and the performance indicator value of the player in the interactive process of the 10 games reflects the change of the user's game skills in two hours.

[0165] As an example, the interactive process can also occur in multiple discrete time periods, for example, a player participates in a game at least once a week in the past year, a total of 500 games, then the first 300 games of the 500 games can be regarded as the pre-stage, and the last 200 games can be regarded as the post-stage, and the performance indicator value of the player in the interactive process of the 500 games reflects the change of the user's game skills in a year.

[0166] In some embodiments, in the case where the interactive process includes M stages, and the multiple sampling points are in the first N stages of the M stages, i.e., for the pre-stage of the interactive process, the "determining the performance indicator value of the virtual object in the interactive process according to the sub-performance indicator values of the virtual object at the multiple sampling points" in the above step 104 can be implemented by performing the following processing for each of the first N stages: determining a first average value of the sub-performance indicator values of the virtual object at the multiple sampling points in the stage, and taking the first average value as a first performance indicator value of the virtual object in each stage, where M and N are positive integers, and M is greater than N.

[0167] In some embodiments, the first N stages among the M stages that the sampling points are located in can be artificially set, for example, the first 5 games among 10 games can be set as the first N stages among the M stages; or can be autonomously set by the system according to a preset rule, where the preset rule can be that the time difference between the time length of the first N stages in the early stage and the time length of the N+1th to Mth stages in the later stage is less than a time threshold, or that the difference between the number of stages in the early stage and the number of stages in the later stage is less than or equal to 1, for example, the first 6 games among 11 games can be set as the first N stages.

[0168] Taking games as an example, the stages described above can be a round (also referred to as a game) in a round-based game. For non-round-based games, the stages described above can also be divided according to a set time length (for example, 5 minutes), in which case the time lengths of different stages are the same; the stages can also be divided according to the process of interaction between a virtual object and different virtual objects in a game scene, for example, the process of interaction between a virtual object A controlled by a player and a virtual object B is the first stage, the process of interaction between the virtual object and a virtual object B is the second stage, and the stages can also be the process of executing all commands before the player achieves a preset target task, in which case the time lengths of different stages can be different.

[0169] As an example, in the game scene, when M is 10 and N is 5, the interaction process of the player includes the interaction process of 10 games, and in the case that the plurality of sampling points are located in the first 5 games among the 10 games, for each game among the first 5 games, a first average value of the sub-performance indicator values of the player at the plurality of sampling points corresponding to each game is determined, and the first average value is taken as a first performance indicator value of the virtual object in the interaction process of a game.

[0170] In some embodiments, the interaction process includes M stages, and in the case that the plurality of sampling points are located in the first N stages among the M stages, the step 104 of “determining a performance indicator value of the virtual object in the interaction process according to the sub-performance indicator values of the virtual object at the plurality of sampling points” can be implemented by the following manner: determining a second average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the first N stages, and taking the second average value as a second performance indicator value of the virtual object in the first N stages, the second performance indicator value being used to represent the overall performance reflected by the virtual object in the interaction process of the first N stages, where M and N are positive integers, and M is greater than N.

[0171] As an example, in the game scenario, when M is 10 and N is 5, the interactive process of the player includes the interactive process of 10 games, and in the case that the plurality of sampling points are in the first 5 games of the 10 games, the sub-performance indicator values of the player at the plurality of sampling points corresponding to the first 5 games can be determined, and the average of the sub-performance indicator values of the player at the plurality of sampling points corresponding to the first 5 games is taken as the second performance indicator value of the virtual object in the interactive process of the first 5 games.

[0172] According to the sub-performance indicator values of the virtual object at the plurality of sampling points, the embodiments of the present application determine the performance indicator values of the virtual object in the interactive process at different stages, and by using the average of multiple game scores, the performance indicator values of the virtual object can be quickly determined, thereby avoiding the problem that the performance indicator values of different virtual objects are determined by using a unified preset initial value in the related art, and the accuracy of determining the performance indicator values of the virtual object is improved.

[0173] In some embodiments, the interactive process includes M stages, and in the case that the plurality of sampling points are in the (N+1)th to Mth stages of the M stages, referring to Figure 3F , Figure 3F is a flowchart of determining a performance indicator value provided by the embodiments of the present application, Figure 3A Step 104 in the flowchart can be implemented by executing steps 1041 to 1044 in Figure 3F The following is a specific description.

[0174] In step 1041, a third average of the sub-performance indicator values of the virtual object at the plurality of sampling points in the (N+1)th to Mth stages is determined.

[0175] In some embodiments, the interactive process includes M stages, and in the case that the plurality of sampling points are in the (N+1)th to Mth stages of the M stages, a third average of the sub-performance indicator values of the virtual object at the plurality of sampling points in the (N+1)th to Mth stages can be determined according to the sub-performance indicator values of the virtual object at the plurality of sampling points in each stage, and the third average is used to represent the overall performance reflected by the virtual object in the interactive process of the (N+1)th to Mth stages.

[0176] As an example, in the game scenario, when M is 10 and N is 5, the interactive process of the player includes the interactive process of 10 games, and in the case that the plurality of sampling points are in the 6th to 10th games of the 10 games, a third average of the sub-performance indicator values of the virtual object at the plurality of sampling points in the 6th to 10th games can be determined according to the sub-performance indicator values of the virtual object at the plurality of sampling points in the 6th to 10th games.

[0177] In step 1042, a fourth average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the Mth stage is determined.

[0178] In some embodiments, the fourth average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the Mth stage can be determined according to the sub-performance indicator values of the virtual object at the plurality of sampling points in the Mth stage.

[0179] As an example, the fourth average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the 10th game can be determined according to the sub-performance indicator values of the player at the plurality of sampling points in the 10th game.

[0180] In step 1043, the product of the difference between the third average value and the fourth average value and the learning rate parameter is determined as a performance indicator increment value.

[0181] Here, the learning rate is a scalar that controls the step size of parameter updates in the gradient descent algorithm. If the learning rate is too large, the parameter updates in the training process may be too fast, skipping the optimal solution, and the model may not learn the data well; if the learning rate is too small, the training process may be too slow, and a large number of iterations may be required to find the optimal solution. The setting method of the learning rate parameter includes manual setting according to empirical values, adjusting the learning rate using a pre-trained model, using a learning rate decay strategy, etc.

[0182] In some embodiments, the value of the learning rate is related to the virtual object, that is, the stronger the preset learning ability of the virtual object, the larger the value of the corresponding learning rate, and the pre-trained model can adjust the learning rate according to the characteristic parameters of the virtual object. Taking the characteristic parameter of the confrontation duration of the virtual object in the previous plurality of stages as an example, if the virtual object encounters the same virtual object in the previous plurality of stages, when the virtual object faces the attack of the same virtual object again in the current stage, the ability of the virtual object to resist the attack is improved, and the value of the corresponding learning rate is increased.

[0183] As an example, the performance indicator increment value can be represented as lr(score n -score old ), wherein score n is the fourth average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the Mth stage, score old is the third average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the N+1th to Mth stages, and lr is the learning rate parameter, which can be set to 0.1, for example.

[0184] In step 1044, the sum of the performance indicator increment value and the third average value is taken as the third performance indicator value of the virtual object in the N+1th to Mth stages, wherein M and N are positive integers, and M is greater than N.

[0185] As an example, the third performance indicator value of the virtual object in the N+1th to Mth stage can be represented as score new = score old + lr(score n - score old ).

[0186] The embodiments of the present application determine the performance indicator value of the virtual object in the interactive process at different stages according to the sub-performance indicator values of the virtual object at multiple sampling points. In the early stage, the performance indicator value of different virtual objects is determined by using the average of multiple local scores, which can consider the difference in initial strength of different virtual objects, quickly determine the performance indicator value of different virtual objects in the early stage, and avoid the problem that the performance indicator value of the virtual object is not accurate enough due to the use of a unified preset initial value for different virtual objects in the related art, thereby improving the accuracy of determining the performance indicator value of the virtual object. In the later stage, by setting the learning rate parameter, the performance indicator value of the same virtual object at different stages in the later stage can be dynamically updated. Compared with the way of updating the performance indicator value of the virtual object at different stages in the later stage in the related art, the data processing amount in the process of updating the performance indicator value at different stages in the later stage is reduced, and the efficiency of determining the performance indicator value of the virtual object at different stages in the later stage is improved.

[0187] In some embodiments, after performing the above step 104 of “determining the performance indicator value of the virtual object in the interactive process according to the sub-performance indicator values of the virtual object at multiple sampling points”, the performance indicator value of the virtual object in the interactive process can be displayed on the graphical interface of the terminal.

[0188] In some embodiments, in the game scene, the performance indicator value of the virtual object in the interactive process at each time period can be displayed. Assuming that the player ends the 2-hour interactive process from 11:00 to 13:00, and the player has played a total of 10 games, the first 5 games are regarded as the early stage, and the last 5 games are regarded as the later stage, the early performance indicator value of the virtual object in the early interactive process and the later performance indicator value of the virtual object in the later interactive process can be automatically displayed, respectively, wherein the early performance indicator value can be the average of the sub-performance indicator values of the multiple sampling points of the first 5 games, and the later performance indicator value can be determined according to the above steps 1041 to 1044.

[0189] As an example, refer to Figure 6 , Figure 6is a first schematic diagram of performance index values of virtual objects provided by an embodiment of the present application, which shows the performance index values of virtual object 1, virtual object 2, virtual object 3 and virtual object 4 in the early interaction process and the performance index values of virtual object 1, virtual object 2, virtual object 3 and virtual object 4 in the later interaction process respectively, and sorts the performances of the four virtual objects according to the sizes of the performance index values.

[0190] In some embodiments, the performance index values of the virtual object and the other virtual objects in the interaction process of the corresponding historical tracing time period can also be displayed respectively in response to the selection operation of the user on the historical tracing time period of the interaction process at the current time, or after the interaction process of the virtual object ends, a list of virtual objects in the interaction process is displayed in the graphical interface, and the performance index values of the selected virtual object in the interaction process are displayed in response to the selection of the user on the virtual object.

[0191] As an example, if the current time is 13:00, referring to Figure 7 , Figure 7 is a second schematic diagram of performance index values of virtual objects provided by an embodiment of the present application, as shown in Figure 7 , in the display interface 700, in response to the selection operation on the time selection control 701, it is determined that the historical tracing time period is 2 hours before the current time, that is, from 11:00 to 13:00, after the end of the interaction process, the virtual object list 702 and the object selection list 703 are shown in the display interface 700, wherein the virtual object list 702 includes virtual object 1, virtual object 2, virtual object 3 and virtual object 4, and the object selection list 703 includes selection boxes corresponding to virtual object 1, virtual object 2, virtual object 3 and virtual object 4 respectively. In response to the operation of the user on the selection box 704, it is determined to display the performance index values corresponding to virtual object 2, at this time, the performance index values of virtual object 2 in the early interaction process and the performance index values of virtual object 2 in the later interaction process are displayed in the index value display area 705 in the display interface 710.

[0192] Compared with the related art which only relies on the final win or loss result, the data granularity processed by the embodiments of the present application is more detailed, and more detailed information in the interactive process can be obtained, so that more comprehensive reference can be provided for subsequent prediction of performance indicator values. Compared with the related art which only determines the performance indicator values of the virtual object based on simple statistical indicators, the embodiments of the present application determine the performance indicator values of the virtual object in the interactive process by taking the sub-performance indicator values of the virtual object at each sampling point as reference and combining the sub-performance indicator values of the multiple sampling points, which can more sensitively capture the dynamic changes of different situations among the multiple sampling points, so that the determination of the performance indicator values of the virtual object is more accurate. By traversing the probability distributions of the multiple reference operation data, the probability of the multiple reference operation data in the probability distribution is taken as reference to determine the probability to which the actual operation data of the virtual object belongs. When the actual operation data is not found in the multiple reference operation data, the preset indicator value is determined as the sub-performance indicator value of the virtual object at the sampling point, which formulates a scheme for determining the sub-performance indicator value of the virtual object at the sampling point under different situations, avoids the problem of inaccurate analysis caused by the fact that the actual operation data cannot be found in the multiple reference operation data, and can improve the accuracy evaluation of the sub-performance indicator value of the virtual object at the sampling point.

[0193] In the following, an exemplary application of the embodiments of the present application to scoring the performance of a player in a game scenario will be described.

[0194] In the process of a multiplayer online game, the accuracy of the performance score of a player is crucial to the game experience of the player, which can not only motivate the player to improve skills, but also help game designers optimize matching algorithms. When scoring the performance of a game player, the related art mainly uses a scoring algorithm to rely on simple statistical indicators such as win or loss results and scores of player roles for scoring. Although this method can simply and quickly score the performance of a player in a game, it is easy to ignore the dynamics and complexity of the game in the case of complex game mechanisms and player strategies, which reduces the accuracy of the performance score of the player and thus reduces the game experience of the player.

[0195] For example, the Elo Rating System (ELO), the system is based on a simple assumption: if a player wins a game against another player, the player rating of the winning player should be increased, and vice versa, the adjustment of the rating depends on the difference between the ratings of the two players and the result of the game. However, this method has the following technical problems: long convergence time, the ELO rating system takes a long time to rate the game players, and a large amount of game data is needed to obtain a stable rating result; oversimplification: the ELO rating algorithm assumes that the skill level of all players is static, and the magnitude of the settlement update will become smaller and smaller, ignoring the dynamic changes of the player's skill; lack of individualization: ELO rating mainly depends on the win or lose result of the game, and ignores the individual performance and team cooperation in the game, so that the mediocre players of the winning side will also get high ratings, and the excellent players of the losing side will get low ratings.

[0196] To solve the above problems, the embodiment of the present application provides a data processing method of a virtual object, which scores the performance of a player in a game based on a reinforcement learning model. As an example, the reinforcement learning model can be an AC model, which is composed of an Actor model and a Critic model. These two parts work together to optimize the behavior strategy of an agent (Agent). The embodiment of the present application regards the game as a reinforcement learning environment, and the game state, action and feedback correspond to the state of the environment, the operation of the player and the feedback in the game. By training the Actor model and the Critic model in the AC model, the strategy function and the evaluation function of the agent can be obtained. In the application stage, the strategy probability distribution (probability distribution of multiple predicted operations) output by the Actor model in the pre-trained AC model is taken as a reference to score the actual operation of the player.

[0197] For the same game state, comparing the strategy probability distribution output by the Actor model with the actual operation of the player, if the actual operation of the player has a larger probability value in the strategy probability distribution, it corresponds to a higher performance score; if the actual operation of the player has a smaller probability value in the strategy probability distribution, it corresponds to a lower performance score. The data processing method of the virtual object provided by the embodiment of the present application is faster and more accurate than the related art processing, and can directly obtain the performance score of the player based on single operation data without the need for a large amount of game data for training; by using the average score of the last n games and the learning rate update mechanism, the performance score of the player can be dynamically updated; and the win or lose of the game is not strongly related to the game, but is strongly related to the operation data of the player. For the performance of two players who are both winners, individualized scoring can also be made.

[0198] The embodiment of the application provides a training process of an AC model and an application process of scoring the performance of a player by using the pre-trained AC model.

[0199] The Actor model is a part of the AC model responsible for learning the strategy. The Actor model receives the state of the environment as input and outputs the action that should be taken under the given state. By learning a strategy function π(s|a), the maximum of long-term cumulative feedback is achieved, where s is the state of the game session and a is the action corresponding to the state. The Actor model can be composed of a deep neural network. The input of the network is the state of the game, and the output is the corresponding strategy probability distribution. The strategy probability distribution can be discrete or continuous, depending on the specific application scenario. The Actor model learns the optimal strategy by exploring the environment and trying different actions. According to the feedback provided by the Critic model, the action selection under the corresponding state is adjusted so that the action corresponding to higher feedback can be selected.

[0200] The Critic model is a part of the AC algorithm responsible for evaluating the strategy. It can evaluate whether the action selected by the Actor model has received high enough feedback and provide a signal to the Actor model to improve the strategy. By observing the action of the Actor and the feedback obtained, the evaluation function V(s) is learned. This function can estimate the expected cumulative feedback under a given state s. The Critic can also be composed of a deep neural network. The input of the network is the game state after the Actor model performs the action, and the output is the corresponding state value. Here, each game of the player in the game can be regarded as a reinforcement learning process. Each operation of the player in the game is recorded as an action, and the win or lose result of the game and the score of the player during the game can be used as a feedback signal. Through training, the AC model can learn the optimal strategy and value estimation of the player under different game states.

[0201] The training process of the AC model is described below. First, the network parameters of the Actor model and the Critic model are randomly initialized. Second, the state data in the game match is sampled, such as the enemy position, teammate position, blood volume, etc. at a preset historical time. Then, the sampled state data is input into the Actor model to output the probability distribution of multiple actions under the corresponding state, and a random action is executed to generate the corresponding feedback, and the state at the historical time is updated to the next state. Next, the Critic model predicts the value of the state after the Actor model executes the action, and updates the evaluation function V(s) according to the difference between the predicted value and the feedback. The Actor model determines the policy gradient using the reinforcement learning policy gradient method according to the value feedback of the state after executing the action by the Critic model, and updates the parameters of the policy function π(s|a). Finally, the above training process is repeated until the preset training number or convergence condition is reached, and the network parameters of the Actor model and the Critic model in the AC model are stopped updating. Through the above training process, the Actor model and the Critic model cooperate with each other, constantly learn and improve, and finally make the agent learn to make optimal decisions in a complex environment.

[0202] The pre-trained Actor model can output the probability distribution of various optional actions for a given state. Since the goal of reinforcement learning is to obtain the maximum feedback, that is, to win the game, a larger probability in the action probability distribution means a higher probability of winning. Therefore, the probability value corresponding to the action can be used as a reference to score the player's performance. The process of scoring the player's performance by the pre-trained AC model is described below.

[0203] Referring to Figure 4 , Figure 4 is a flowchart of scoring the player's performance provided by the embodiments of the present application. As shown in Figure 4 , first, the state data is sampled. In order to score the player's performance based on the Actor model, the model and the match data need to be prepared in advance. Since the game performance score is generally performed during the settlement period after a match, the complete data of a match can be obtained after the player's match ends, including the state data in the match, such as the enemy position, teammate position, blood volume, etc. at a certain time, and the player's operation data, such as triggering the move key or skill key, etc. The game match data can be sampled at a preset sampling frequency, such as once per second, to collect the state data and the player's operation data at each sampling point. Here, the state data of the game is the slice data at the sampling time point, and the player's operation data is the set data of the player's operation within one second duration after the sampling time point.

[0204] Secondly, the game state data obtained by each sampling point is input into the Actor model, and the Actor model determines the probability distribution of the corresponding discrete action according to the input state data, which is used to represent the probability of taking each action in the state.

[0205] Then, the actual action of the player is scored with reference to the probability distribution of the action output by the Actor model. The probability p a of the actual action of the player in the probability distribution of the action output by the Actor model is queried. a If there is no actual action of the player in the action output by the Actor model, the probability value of p max is 0; and if the maximum probability in the probability distribution of the discrete action is p a , then the score of the actual action of the player is score=p max / p a , and the value range of the score result score is [0, 1].

[0206] Finally, the state data samples of other sampling points are continuously traversed, and the performance of the player in the corresponding state is scored according to the above processing method. When the traversal ends, the average value of the scores of all sampling points is taken as the overall performance score of the player.

[0207] As an example, it is assumed that an enemy located in front of the player character releases a skill to the player character. At this time, if the player character moves backward, the skill released by the enemy can be avoided, and the probability of the player character moving backward is 0.6; if the player character moves left or right, there is a probability of avoiding the skill released by the enemy, and the probability of the player character moving left is 0.2 and the probability of the player character moving right is 0.15; if the player character moves forward or does not move, it will be damaged, and the probability of the player character moving forward is 0.05; the probability distribution of the discrete action output by the Actor model according to the game state data at this time is P={0.60(backward), 0.20(left), 0.15(right), 0.05(forward)}; if the operation of the player at this time is to control the player character to move left, the probability p a of the actual action of the player is 0.20, and since the maximum probability in the probability distribution P is 0.60, the performance score of the player in this state is 0.20 / 0.60=1 / 3; if the player does not operate, the score is 0.

[0208] The data processing method of the virtual object provided in the embodiments of the present application only needs one game data, and can score the performance of the player in the game, which is faster than the related art. However, considering the volatility of the performance of the player, the performance scores of the same player in two games may be quite different, so the embodiments of the present application increase the amplitude updating mechanism, and configure the learning rate to make the score result of the performance of the player more stable.

[0209] In the related art, when scoring the performance of the player, a uniform initial value score is usually set for the player, and the initial value score is updated in the settlement of the subsequent game, but this way ignores the difference in the initial strength of different players, so the score of the player in the early stage of the game is particularly inaccurate. In order to solve this problem, the embodiments of the present application use the average score of the last n games as the comprehensive performance score of the player for the game data of the early stage of the player, and use the learning rate updating mechanism to dynamically update the performance score of the player for the game data of the later stage.

[0210] As an example, the first 10 games of the player are taken as the game data of the early stage, and the game data after the 10th game is taken as the game data of the later stage.

[0211] After the nth game in the early stage, the update result of the player score is shown in formula (1).

[0212]

[0213] score i is the single game score of the i th game in the early stage. By using the average score of multiple games, the game strength of the player can be quickly evaluated, and the way of using the preset initial value in the related art is avoided, which needs to process dozens of games before the player score converges to the stable stage.

[0214] After the nth game in the later stage, the update result of the player score is shown in formula (2).

[0215] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score old is the score of the player before updating, and score n is the score of the player after updating.

[0001] score n -score old ) (2)

[0216] score oldThe single-game score of the nth game is lr, and the learning rate can be set to 0.1. Here, the learning rate can also be configured according to application needs. A larger learning rate will make the player's score update more, and a smaller learning rate will make the opposite. By using the learning rate update mechanism, the player's performance score can be dynamically updated, and by adjusting the learning rate, the effect of stably and smoothly updating the player's performance score can be achieved.

[0217] The embodiment of the application only needs to sample the state data and operation data of one game to evaluate the player's strength, which is faster than the related art that needs a large number of game win-lose structures to score the player's performance. In the early stage, the average of the player's performance scores of the last n games is used as the player's comprehensive score, and in the later stage, the learning rate update mechanism is used. Compared with the related art that uses a preset initial value, the early stage score accuracy of the player is low, and after a large number of games in the later stage, it will tend to be stable, and the update amplitude will become smaller and smaller. The virtual object processing method provided by the embodiment of the application can quickly evaluate the player's game strength in the early stage, dynamically update the player's performance score in the later stage, and by adjusting the learning rate, the effect of stably and smoothly updating the player's performance score can be achieved. When scoring the player's performance, the embodiment of the application is strongly related to the player's operation data, compared with the related art that is strongly related to the final win-lose of the game. The embodiment of the application can score different players who are on the winning side individually. Especially in the multiplayer team competition game scene, when the strength of the players in the team is different, the virtual object processing method provided by the embodiment of the application can score the players' respective operations more carefully, rather than just looking at the final win-lose result, improving the accuracy of the performance score of different players and improving the players' experience.

[0218] The following continues to illustrate an exemplary structure of the virtual object data processing apparatus 455 provided by the embodiment of the application as a software module. In some embodiments, as shown in Figure 2 The software modules stored in the virtual object data processing apparatus 455 of the memory 450 can include:

[0219] The data sampling module 4551 is configured to sample the interactive process of the virtual scene to obtain a sampling data set, wherein the sampling data set includes state data of the virtual object at a plurality of sampling points and actual operation data of each sampling point.

[0220] The first determination module 4552 is configured to determine, for each sampling point, a probability distribution of a plurality of reference operation data corresponding to the sampling point according to the state data.

[0221] The second determining module 4553 is configured to, in response to that the actual operation data is queried in the plurality of reference operation data, determine the sub-performance indicator value of the virtual object at the sampling point according to the probability of the corresponding reference operation data in the probability distribution.

[0222] The third determining module 4554 is configured to determine the performance indicator value of the virtual object in the interaction process according to the sub-performance indicator values of the virtual object at the plurality of sampling points.

[0223] In some embodiments, the second determining module 4553 is further configured to take the probability of the corresponding reference operation data in the probability distribution as a target probability of the virtual object, and take a ratio of the target probability to a maximum probability in the probability distribution as the sub-performance indicator value of the virtual object at the sampling point.

[0224] In some embodiments, the second determining module 4553 is further configured to, in response to that the actual operation data is not queried in the plurality of reference operation data, determine a preset indicator value as the sub-performance indicator value of the virtual object at the sampling point.

[0225] In some embodiments, the interaction process includes M stages, and in a case where the plurality of sampling points are in the first N stages of the M stages, the third determining module 4554 is further configured to, for each of the first N stages, perform the following processing: determine a first average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the stage, and take the first average value as a first performance indicator value of the virtual object in each stage, where M and N are both positive integers, and M is greater than N.

[0226] In some embodiments, the interaction process includes M stages, and in a case where the plurality of sampling points are in the first N stages of the M stages, the third determining module 4554 is further configured to determine a second average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the first N stages, and take the second average value as a second performance indicator value of the virtual object in the first N stages, where M and N are both positive integers, and M is greater than N.

[0227] In some embodiments, the interaction process includes M stages, and in a case where the plurality of sampling points are in the (N+1)th to M stages of the M stages, the third determining module 4554 is further configured to determine a third average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the (N+1)th to M stages, determine a fourth average value of the sub-performance indicator values of the virtual object at the plurality of sampling points in the M stage, determine a product of a difference between the third average value and the fourth average value and a learning rate parameter as a performance indicator increment value, and take a sum of the performance indicator increment value and the third average value as a third performance indicator value of the virtual object in the (N+1)th to M stages, where M and N are both positive integers, and M is greater than N.

[0228] In some embodiments, the first determining module 4552 is further configured to, for each sampling point, determine probabilities of the virtual object performing the plurality of reference operation data respectively when the virtual object is in the state corresponding to the state data, and combine the probabilities of performing the plurality of reference operation data into a probability distribution.

[0229] In some embodiments, the probability distribution is determined by a pre-trained deep learning model, and the first determining module 4552 is further configured to sample an interaction process of the sample virtual scene to obtain a first sample state of the sample virtual object at a first sampling point; determine, by the initialized deep learning model, a probability distribution of the sample virtual object performing a plurality of sample reference operations respectively in the first sample state; select one of the plurality of sample reference operations as the first sample operation, and perform the first sample operation to obtain a second sample state of the sample virtual object at a second sampling point, where the second sample state is a state updated after the first sample operation is performed in the first sample state; obtain an actual evaluation parameter of performing the first sample operation in the first sample state; perform parameter prediction based on the second sample state to obtain a predicted evaluation parameter; and update parameters of the initialized deep learning model based on a difference between the actual evaluation parameter and the predicted evaluation parameter to obtain the pre-trained deep learning model.

[0230] In some embodiments, the first determining module 4552 is further configured to determine probabilities of the sample virtual object performing the plurality of sample reference operations respectively when the sample virtual object is in the first sample state, and combine the probabilities of performing the plurality of sample reference operations into a probability distribution.

[0231] In some embodiments, the first determining module 4552 is further configured to perform any one of the following processes: select any one of the plurality of sample reference operations as the first sample operation; or select one of the plurality of sample reference operations as the first sample operation based on a preset condition, where the preset condition includes that the sample reference operation corresponding to the maximum probability in the probability distribution.

[0232] In some embodiments, the first determining module 4552 is further configured to perform feature extraction on the state data in the second sample state by the initialized deep learning model to obtain state features, and perform parameter prediction based on the state features to obtain a predicted evaluation parameter.

[0233] In some embodiments, the first determining module 4552 is further configured to determine a loss value based on a difference between the actual evaluation parameter and the predicted evaluation parameter, and update parameters of the initialized deep learning model based on the loss value to obtain the pre-trained deep learning model.

[0234] The embodiment of the present application provides a computer program product, which comprises a computer program or computer executable instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the data processing method of the virtual object provided in the embodiment of the present application.

[0235] The embodiment of the present application provides a computer readable storage medium, which stores computer executable instructions or computer programs. When the computer executable instructions or computer programs are executed by a processor, the processor executes the data processing method of the virtual object provided in the embodiment of the present application, for example, as shown in the data processing method of the virtual object. Figure 3A The data processing method of the virtual object.

[0236] In some embodiments, the computer readable storage medium can be RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM memory, etc. It can also be various devices including one or any combination of the above storage.

[0237] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.

[0238] As an example, the computer executable instructions can but not necessarily correspond to files in a file system, can be stored in part of a file storing other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperative files (for example, files storing one or more modules, subroutines or code portions).

[0239] As an example, the computer executable instructions can be deployed to be executed on one electronic device, or executed on multiple electronic devices located in one place, or executed on multiple electronic devices distributed in multiple places and interconnected through a communication network.

[0240] In summary, by obtaining state data and actual operation data of the virtual object in the interaction process at multiple sampling points through the embodiments of the present application, compared with the related art which only relies on the final win or lose result, the data granularity processed by the embodiments of the present application is more detailed, more detailed information in the interaction process can be obtained, the subsequent processing is more comprehensive, so as to provide a more comprehensive reference for the prediction of the subsequent performance indicator value; for each sampling point, the probability distribution of the plurality of reference operation data is determined according to the state data, the probability corresponding to the actual operation data in the probability distribution is inquired according to the probability distribution of the reference operation data, and the sub-performance indicator value of the virtual object at the sampling point is determined according to the probability to which the actual operation data belongs; finally, the performance indicator value of the virtual object in the interaction process is determined according to the sub-performance indicator value of the virtual object at the plurality of sampling points, the sub-performance indicator value corresponding to the actual operation data in each sampling point probability distribution is determined, and the performance indicator value is determined by combining the sub-performance indicator values of the plurality of sampling points, compared with the related art which only determines the performance indicator value of the virtual object based on a simple statistical indicator, the embodiments of the present application determine the sub-performance indicator value of the virtual object at each sampling point by referring to the probability distribution results of the plurality of sampling points, and finally determine the performance indicator value in the interaction process, which can more sensitively capture the dynamic changes of different situations among the plurality of sampling points, so that the determination of the performance indicator value of the virtual object is more accurate. By traversing the probability distribution of the plurality of reference operation data, the probability of the plurality of reference operation data in the probability distribution is referred to, the probability to which the actual operation data of the virtual object belongs is determined, when the actual operation data is inquired in the plurality of reference operation data, the sub-performance indicator value of the virtual object at the sampling point is determined according to the probability of the corresponding reference operation data in the probability distribution, the probability corresponding to the actual operation data of the virtual object can be quickly located, and the speed of evaluating the sub-performance indicator value of the virtual object at the sampling point is improved. By adjusting the parameters of the deep learning model to minimize the loss value, the prediction result of the evaluation parameter of the deep learning model can be closer and closer to the true value, so as to improve the accuracy of the deep learning model in predicting the state, at the same time, by adjusting the parameters of the deep learning model, the loss value is more effectively reduced, the time required for training the deep learning model can be reduced, and the training resources are saved.

[0241] The above merely describes the embodiments of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement and improvement made within the spirit and scope of the present application shall be included in the protection scope of the present application.

Claims

1. A data processing method for virtual objects, characterized in that, The method includes: The interaction process of the virtual scene is sampled to obtain a sampled dataset, wherein the sampled dataset includes the state data of the virtual object at multiple sampling points and the actual operation data of each sampling point; For each of the sampling points, the probability distribution of multiple reference operation data corresponding to the sampling point is determined based on the state data; In response to finding the actual operation data in the plurality of reference operation data, the sub-performance index value of the virtual object at the sampling point is determined according to the probability of the reference operation data in the probability distribution; The performance index value of the virtual object in the interaction process is determined based on the sub-performance index values ​​of the virtual object at multiple sampling points.

2. The method according to claim 1, characterized in that, Determining the sub-performance index value of the virtual object at the sampling point based on the probability corresponding to the reference operation data in the probability distribution includes: The probability corresponding to the reference operation data in the probability distribution is taken as the target probability of the virtual object; The ratio of the target probability to the maximum probability in the probability distribution is used as the sub-performance index value of the virtual object at the sampling point.

3. The method according to claim 1, characterized in that, The method further includes: In response to the fact that the actual operation data is not found in the plurality of reference operation data, the preset index value is determined as the sub-performance index value of the virtual object at the sampling point.

4. The method according to claim 1, characterized in that, The interaction process includes M stages. When the plurality of sampling points are in the first N stages of the M stages, determining the performance index value of the virtual object in the interaction process based on the sub-performance index values ​​of the virtual object at the plurality of sampling points includes: For each of the first N stages, the following processing is performed: A first average value of the sub-performance index values ​​of the virtual object at the plurality of sampling points in the stage is determined, and the first average value is used as the first performance index value of the virtual object in each stage, wherein M and N are both positive integers, and M is greater than N.

5. The method according to claim 1, characterized in that, The interaction process includes M stages. When the plurality of sampling points are in the first N stages of the M stages, determining the performance index value of the virtual object in the interaction process based on the sub-performance index values ​​of the virtual object at the plurality of sampling points includes: Determine the second average value of the sub-performance index values ​​of the virtual object at the plurality of sampling points in the first N stages, and use the second average value as the second performance index value of the virtual object in the first N stages, wherein M and N are both positive integers, and M is greater than N.

6. The method according to any one of claims 1 to 5, characterized in that, The interaction process includes M stages. When the plurality of sampling points are in the (N+1)th to the Mth stage of the M stages, determining the performance index value of the virtual object in the interaction process based on the sub-performance index values ​​of the virtual object at the plurality of sampling points includes: Determine the third average value of the sub-performance index values ​​of the virtual object at the plurality of sampling points in the N+1 to Mth stages; Determine the fourth average value of the sub-performance index values ​​of the virtual object at the plurality of sampling points in the Mth stage; The product of the difference between the third average and the fourth average and the learning rate parameter is determined as the performance metric increment; The sum of the performance index increment and the third average value is used as the third performance index value of the virtual object in the N+1 to M stages, where M and N are both positive integers, and M is greater than N.

7. The method according to claim 1, characterized in that, For each sampling point, determining the probability distribution of multiple reference operation data corresponding to that sampling point based on the state data includes: Perform the following processing for each sampling point: Determine the probability of the virtual object executing multiple reference operation data when it is in the state corresponding to the state data, and combine the probabilities of executing multiple reference operation data into a probability distribution.

8. The method according to claim 7, characterized in that, The probability distribution is determined by a pre-trained deep learning model, which is trained in the following manner: The interaction process of the sample virtual scene is sampled to obtain the first sample state of the sample virtual object at the first sampling point; The probability distribution of the sample virtual object performing multiple sample reference operations in the first sample state is determined by the initialized deep learning model. Select one sample reference operation from the plurality of sample reference operations as the first sample operation, execute the first sample operation, and obtain the second sample state of the sample virtual object at the second sampling point, wherein the second sample state is the state updated after executing the first sample operation in the first sample state; Obtain the actual evaluation parameters for performing the first sample operation under the first sample state; Based on the state of the second sample, parameter prediction is performed to obtain the predicted evaluation parameters; Based on the difference between the actual evaluation parameters and the predicted evaluation parameters, the parameters of the initialized deep learning model are updated to obtain the pre-trained deep learning model.

9. The method according to claim 8, characterized in that, The step of determining the probability distribution of the sample virtual object performing multiple sample reference operations in the first sample state through an initialized deep learning model includes: Determine the probability that the sample virtual object performs multiple sample reference operations when it is in the first sample state, and combine the probabilities of performing multiple sample reference operations into a probability distribution.

10. The method according to claim 8, characterized in that, Selecting a sample reference operation as the first sample operation from the plurality of sample reference operations includes: Perform any of the following processes: Select any one of the plurality of sample reference operations as the first sample operation; According to preset conditions, a sample reference operation is selected as the first sample operation from the plurality of sample reference operations. The preset conditions include the sample reference operation corresponding to the highest probability in the probability distribution.

11. The method according to claim 8, characterized in that, The step of predicting parameters based on the second sample state to obtain prediction evaluation parameters includes: The state features are obtained by extracting features from the state data of the second sample state through an initialized deep learning model. Based on the state characteristics, parameter prediction is performed to obtain the prediction evaluation parameters.

12. The method according to any one of claims 8 to 11, characterized in that, The step of updating the parameters of the initialized deep learning model based on the difference between the actual evaluation parameters and the predicted evaluation parameters to obtain the pre-trained deep learning model includes: The loss value is determined based on the difference between the actual evaluation parameters and the predicted evaluation parameters; Based on the loss value, the parameters of the initialized deep learning model are updated to obtain the pre-trained deep learning model.

13. A data processing device for virtual objects, characterized in that, The device includes: The data sampling module is used to sample the interaction process of the virtual scene to obtain a sampling dataset, wherein the sampling dataset includes the state data of the virtual object at multiple sampling points and the actual operation data of each sampling point; The first determining module is used to determine, for each sampling point, the probability distribution of multiple reference operation data corresponding to the sampling point based on the state data; The second determining module is used to determine the sub-performance index value of the virtual object at the sampling point based on the probability of the actual operation data corresponding to the reference operation data in the probability distribution when the actual operation data is found in the plurality of reference operation data. The third determining module is used to determine the performance index value of the virtual object in the interaction process based on the sub-performance index values ​​of the virtual object at multiple sampling points.

14. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the data processing method for the virtual object according to any one of claims 1 to 12.

15. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the data processing method for the virtual object according to any one of claims 1 to 12.

16. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the data processing method for the virtual object according to any one of claims 1 to 12.