A method, device, computer equipment and storage medium for intelligent agent interaction
By extracting the interaction state characteristics in the virtual interactive scenario and determining the target operation, the agent can interact with the virtual account more accurately in the virtual interactive scenario, solving the problem of low accuracy in the existing technology of the agent's interaction and achieving more efficient virtual account interaction.
Patent Information
- Application Number
- CN202110828710.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-22
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-07-22
AI Technical Summary
In the prior art, the interaction accuracy of the agent is low and cannot flexibly interact with the virtual account in the virtual interactive scenario.
By loading the target agent, the target interactive scene image corresponding to the control operation of the virtual account is obtained, the interaction status characteristics are extracted, the target scheduling operation and the target interactive operation are determined, and the associated virtual controlled elements are controlled in the virtual interactive scene.
It improves the interaction accuracy of the agent, allowing it to accurately imitate the real interaction between virtual accounts, and enhances macro decision-making and interactive capabilities.
Smart Images

Figure CN113813592B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an intelligent agent interaction method, device, computer equipment and storage medium. Background Art
[0002] With the continuous development of technology, more and more devices can not only provide virtual interactive scenes for multiple virtual accounts, but also provide intelligent agents for a single virtual account to interact in virtual interactive scenes. For example, in the game scene, a game account can interact with an intelligent agent when it is not matched with other game accounts; for another example, a game account can improve its own interactive ability by interacting with an intelligent agent.
[0003] Usually, after a virtual account performs a control operation on a virtual controlled element associated with the virtual account, the agent can only determine the corresponding feedback operation for the virtual controlled element associated with the agent. However, after the virtual account performs a control operation, not only a single element of the virtual interactive scene is affected, but the influence of the control operation is diverse. The traditional agent interaction method does not consider the real interaction process in the virtual interactive scene, making it impossible for the agent to flexibly interact with the virtual account in the virtual interactive scene.
[0004] It can be seen that under the existing technology, the interaction accuracy of intelligent agents is low. Summary of the invention
[0005] The embodiments of the present application provide an intelligent agent interaction method, apparatus, computer equipment and storage medium for solving the problem of low interaction accuracy of intelligent agents.
[0006] In a first aspect, a method for intelligent agent interaction is provided, comprising:
[0007] In response to the interactive request instruction triggered by the virtual account, loading the target agent;
[0008] In response to a control operation triggered by the virtual account on a first target virtual controlled element associated with the virtual account in a target virtual interactive scene, acquiring a target interactive scene image corresponding to the control operation;
[0009] Extracting target interaction state features from the target interaction scene image;
[0010] Determine the target scheduling operation and the target interaction operation corresponding to the target agent based on the target interaction state characteristics;
[0011] In response to the target scheduling operation and the target interaction operation, a second target virtual controlled element associated with the target agent is controlled in the target virtual interaction scene.
[0012] In a second aspect, an intelligent agent interaction device is provided, comprising:
[0013] Loading module: used to load the target agent in response to the interactive request instruction triggered by the virtual account;
[0014] A processing module: configured to obtain a target interactive scene image corresponding to the control operation triggered by the virtual account on a first target virtual controlled element associated with the virtual account in a target virtual interactive scene;
[0015] The processing module is also used to: extract target interaction state features from the target interaction scene image;
[0016] The processing module is also used to: determine the target scheduling operation and target interaction operation corresponding to the target agent based on the target interaction state characteristics;
[0017] The processing module is also used to: in response to the target scheduling operation and the target interaction operation, control a second target virtual controlled element associated with the target agent in the target virtual interaction scene.
[0018] Optionally, the target agent is trained in the following manner:
[0019] The processing module is further used to: perform multiple rounds of iterative training on the intelligent agent to be trained based on the interaction process between the intelligent agent to be trained and the preset reference intelligent agent in the sample virtual interactive scene, until the preset training target is met, and output the intelligent agent to be trained as the target intelligent agent, wherein in one round of iterative training, the processing module is specifically used to:
[0020] Based on a first sample interactive state feature corresponding to a first sample interactive scene image in the sample virtual interactive scene, predicting a sample scheduling operation performed by the to-be-trained agent on a sample virtual controlled element associated with the to-be-trained agent in the sample virtual interactive scene, and predicting a sample interactive operation performed by the to-be-trained agent on the sample virtual controlled element after executing the sample scheduling operation;
[0021] Based on the second sample interaction state characteristics corresponding to the second sample interaction scene image generated after executing the sample scheduling operation, and the third sample interaction state characteristics corresponding to the third sample interaction scene image generated after executing the sample interaction operation, the model parameters of the intelligent agent to be trained are adjusted.
[0022] Optionally, the processing module is specifically used for:
[0023] Based on the second sample interaction state feature, according to the preset scheduling incentive strategy, determining the scheduling incentive data of the sample scheduling operation, wherein the scheduling incentive data is used to characterize the degree of completion of the sample scheduling operation and the degree of influence of the sample scheduling operation on the sample interaction result;
[0024] Based on the third sample interaction state feature, and in accordance with a preset interaction incentive strategy, determining interaction incentive data of the sample interaction operation, wherein the interaction incentive data is used to characterize the degree of influence of the sample interaction operation on the sample interaction result;
[0025] The error values between the scheduling incentive data and the interaction incentive data and the preset target incentive data are determined respectively, and the model parameters of the to-be-trained intelligent agent are adjusted based on the obtained error values.
[0026] Optionally, the processing module is further used for:
[0027] After adjusting the model parameters of the intelligent agent to be trained based on the obtained error values, determining the evaluation value of the intelligent agent to be trained according to a preset scoring strategy based on the scheduling incentive data and the interaction incentive data obtained through multiple rounds of iterative training, wherein the evaluation value is used to characterize the training degree of the intelligent agent to be trained;
[0028] If the evaluation value converges, the agent to be trained is output as the target agent.
[0029] Optionally, the processing module is specifically used for:
[0030] Based on the selection probability corresponding to each reference agent in the preset reference agent set, randomly select a reference agent from each reference agent;
[0031] Based on the interaction process between the intelligent agent to be trained and the extracted reference intelligent agent in the sample virtual interaction scene, performing multiple rounds of iterative training on the intelligent agent to be trained;
[0032] If the subject to be trained does not meet the training objective when the sample interaction results between the subject to be trained and the extracted reference subject are obtained, then a reference subject is re-extracted from each of the reference subjects, and the subject to be trained is continuously trained for multiple rounds of iterations;
[0033] If the intelligent agent to be trained meets the training objective, the intelligent agent to be trained is output as the target intelligent agent.
[0034] Optionally, the processing module is further used for:
[0035] Before outputting the intelligent agent to be trained as the target intelligent agent, counting the number of times the intelligent agent to be trained is iteratively trained;
[0036] If the counted number of training times reaches a preset specified number of times, the agent to be trained is output as a reference agent and added to the reference agent set;
[0037] The training times are reset to zero, the iterative training of the agent to be trained is continued, and the reference agent set is updated based on the re-counted training times.
[0038] Optionally, the processing module is further used for:
[0039] Based on a first sample interaction state feature corresponding to a first sample interaction scene image in the sample virtual interaction scene, predict the sample scheduling operation performed by the to-be-trained agent on a sample virtual controlled element associated with the to-be-trained agent in the sample virtual interaction scene, and predict that after the to-be-trained agent performs the sample scheduling operation and before the sample interaction operation performed on the sample virtual controlled element, perform region recognition processing on the first sample interaction scene image to obtain a first interaction result region, a first global viewing region, and a first local viewing region;
[0040] Performing image feature extraction processing on the first interaction result area, the first global viewing area, and the first local viewing area respectively, and obtaining corresponding first feature vectors, first global viewing feature matrices, and first local viewing feature matrices respectively, wherein the first feature vector is used to characterize interaction information related to the sample interaction result, the first global viewing feature matrix is used to characterize position information of the sample virtual controlled element, position information of the reference virtual controlled element associated with the reference agent, and position information of the scene elements included in the sample virtual interaction scene, and the first local viewing feature matrix is used to characterize position information of the sample virtual controlled element included in the first local viewing area, position information of the reference virtual controlled element included in the first local viewing area, and position information of the scene elements included in the first local viewing area;
[0041] The first feature vector, the first global perspective feature matrix and the first local perspective feature matrix are used as first sample interaction state features corresponding to the first sample interaction scene image.
[0042] Optionally, the processing module is specifically used for:
[0043] Based on the first feature vector and the first global perspective feature matrix, predicting a sample scheduling operation performed by the to-be-trained agent on the sample virtual controlled element;
[0044] Based on the sample scheduling operation, the first global perspective feature matrix and the first local perspective feature matrix, predict the predicted feature vector, the predicted global perspective feature matrix and the predicted local perspective feature matrix corresponding to the predicted interactive scene image generated after executing the sample scheduling operation;
[0045] Based on the predicted feature vector and the predicted local view feature matrix, a sample interactive operation performed by the to-be-trained agent on the sample virtual controlled element is predicted.
[0046] Optionally, the processing module is specifically used for:
[0047] Dividing the first global viewing area into a plurality of sub-areas;
[0048] Predicting a target sub-region corresponding to the sample virtual controlled element based on the first feature vector and the first global viewing angle feature matrix;
[0049] The sample scheduling operation is obtained based on the sub-region where the sample virtual controlled element is currently located and the target sub-region corresponding to the sample virtual controlled element.
[0050] Optionally, the processing module is further used for:
[0051] After obtaining the sample scheduling operation based on the sub-region where the sample virtual controlled element is currently located and the target sub-region corresponding to the sample virtual controlled element, based on the sample scheduling operation, controlling the sample virtual controlled element to move to the corresponding target sub-region to obtain a second sample interactive scene image generated by the to-be-trained intelligent agent;
[0052] The second sample interaction state feature corresponding to the second sample interaction scene image is extracted to obtain a second feature vector, a second global perspective feature matrix and a second local perspective feature matrix.
[0053] Optionally, the processing module is further used for:
[0054] After extracting the second sample interaction state feature corresponding to the second sample interaction scene image, determining the first scheduling sub-stimulus of the sample scheduling operation based on whether the sub-region where the sample virtual controlled element is currently located matches the target sub-region corresponding to the sample virtual controlled element indicated by the sample scheduling operation;
[0055] Determining a second scheduling sub-stimulus of the sample scheduling operation based on a change value between the first eigenvector and the second eigenvector corresponding to the sample virtual controlled element;
[0056] Based on a weighted sum of the first scheduling sub-stimulus and the second scheduling sub-stimulus, scheduling stimulus data for the sample scheduling operation is determined.
[0057] Optionally, the intelligent agent to be trained includes a quantitative information extraction module, wherein the quantitative information extraction module is used to extract sample interaction state features corresponding to each sample interaction scene image;
[0058] The intelligent agent to be trained also includes a training module, wherein the training module is used to obtain each scheduling incentive data and each interaction incentive data based on each sample interaction state feature, and adjust the model parameters of the intelligent agent to be trained based on the obtained each scheduling incentive data and each interaction incentive data.
[0059] Optionally, the processing module is specifically used for:
[0060] If the training module includes a scheduling model and an interactive model, and the target incentive data includes scheduling target incentive data and interactive target incentive data, determining a scheduling error value between the scheduling incentive data and the scheduling target incentive data, and adjusting a model parameter of the scheduling model based on the obtained scheduling error value;
[0061] An interaction error value between the interaction incentive data and the interaction target incentive data is determined, and a model parameter of the interaction model is adjusted based on the obtained interaction error value.
[0062] According to a third aspect, a computer device is provided, comprising:
[0063] A memory for storing program instructions;
[0064] The processor is used to call the program instructions stored in the memory and execute the method as described in the first aspect according to the obtained program instructions.
[0065] According to a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to the first aspect.
[0066] In the embodiment of the present application, in response to the control operation triggered by the virtual account for the first target virtual controlled element, the target scheduling operation and the target interactive operation corresponding to the target intelligent agent can be determined, so that the second target virtual controlled element can be controlled in response to the target scheduling operation and the target interactive operation. The target scheduling operation obtained in response to the control operation can control the second target virtual controlled element to perform the scheduling action, reflecting the interactive strategy obtained by the target intelligent agent in response to the control operation from a macro perspective, rather than simply responding to the control operation in interactive actions, thereby improving the interactive accuracy of the target intelligent agent.
[0067] At the same time, when there are multiple second target virtual controlled elements, it is not a single control of a second target virtual controlled element to respond to the control operation of the virtual account, but the target scheduling operation is obtained by considering the overall situation of all second target virtual controlled elements, thereby controlling one second target virtual controlled element or multiple second target virtual controlled elements based on the target scheduling operation, thereby further improving the interaction accuracy of the target intelligent body.
[0068] The target interactive operation obtained in response to the control operation can control the second target virtual controlled element to perform interactive actions, which reflects the interactive ability of the target intelligent body from a microscopic perspective, and can maximize the interactive ability on the basis of the correct interactive strategy, thereby improving the interactive accuracy of the target intelligent body. From both macroscopic and microscopic perspectives, the macroscopic decision-making ability of the target intelligent body is improved without reducing the interactive ability, so that the target intelligent body can provide accurate and diversified feedback to the virtual account. In the embodiment of the present application, the interaction between the target intelligent body and the virtual account can accurately imitate the real interaction between the virtual accounts, thereby improving the interactive accuracy of the target intelligent body. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1a A schematic diagram of a principle of an intelligent agent interaction method provided for related technologies;
[0070] Figure 1b A schematic diagram of the principle of the intelligent agent interaction method provided in the embodiment of the present application Figure 2 ;
[0071] Figure 2 An application scenario of the intelligent agent interaction method provided in the embodiment of the present application;
[0072] Figure 3a A third schematic diagram of a principle of an intelligent agent interaction method provided in an embodiment of the present application;
[0073] Figure 3b A flowchart of an intelligent agent interaction method provided in an embodiment of the present application is shown in FIG1;
[0074] Figure 4a A fourth schematic diagram of a principle of an intelligent agent interaction method provided in an embodiment of the present application;
[0075] Figure 4b A schematic diagram 5 of a principle of an intelligent agent interaction method provided in an embodiment of the present application;
[0076] Figure 4c A schematic diagram of the principle of the intelligent agent interaction method provided in the embodiment of the present application Figure 6 ;
[0077] Figure 4d A schematic diagram of the principle of the intelligent agent interaction method provided in the embodiment of the present application Figure 7 ;
[0078] Figure 5a A schematic diagram of the principle of the intelligent agent interaction method provided in the embodiment of the present application Figure 8 ;
[0079] Figure 5b A schematic diagram of the principle of the intelligent agent interaction method provided in the embodiment of the present application Figure 9 ;
[0080] Figure 6 A schematic diagram of a principle of an intelligent agent interaction method provided in an embodiment of the present application is shown in Figure 10;
[0081] Figure 7 A schematic diagram of a process of an intelligent agent interaction method provided in an embodiment of the present application Figure 2 ;
[0082] Figure 8 A structural schematic diagram 1 of an intelligent agent interaction device provided in an embodiment of the present application;
[0083] Fig. 9 A schematic diagram of a structure of an intelligent agent interaction device provided in an embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0084] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0085] Some of the terms used in the embodiments of the present application are explained below to facilitate understanding by those skilled in the art.
[0086] (1) Multiplayer online battle arena (MOBA):
[0087] In the competition, multiple people are usually divided into two teams, which compete with each other in scattered game maps. Everyone controls the selected character through an interface, and usually does not need to operate organizational units such as buildings, resources, and training arms in the game.
[0088] (2) Agent:
[0089] An intelligent agent is a computational entity that resides in a certain environment, can function continuously and autonomously, and has characteristics such as residency, responsiveness, sociality, and initiative. Intelligent agent is a very important concept in the field of artificial intelligence. Any entity that has independent thoughts and can interact with the environment can be abstracted as an intelligent agent.
[0090] (3) Reinforcement learning (RL):
[0091] Reinforcement learning can be used to describe and solve the problem of how an agent learns strategies to maximize rewards or achieve specific goals during its interaction with the environment. If a behavior of an agent leads to positive incentives from the environment, the tendency of the agent to behave in the future will be strengthened. The goal of the agent is to find the optimal strategy in each discrete state to maximize the expected incentives.
[0092] The embodiments of the present application involve cloud technology and artificial intelligence technology (AI), and are designed based on cloud computing and cloud storage in cloud technology, as well as computer vision technology (CV) and machine learning (ML) in artificial intelligence technology.
[0093] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.
[0094] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". From the user's perspective, the resources in the "cloud" are infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for as needed.
[0095] As a provider of basic cloud computing capabilities, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an Infrastructure as a Service (IaaS) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0096] According to the logical function division, the Platform as a Service (PaaS) layer can be deployed on the IaaS layer, and the Software as a Service (SaaS) layer can be deployed on the PaaS layer, or SaaS can be directly deployed on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0097] Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0098] At present, the storage method of the storage system is: create a logical volume, and when creating a logical volume, allocate physical storage space for each logical volume. The physical storage space may be composed of disks of a storage device or several storage devices. The client stores data on a logical volume, that is, stores the data on the file system. The file system divides the data into many parts, each of which is an object. The object contains not only data but also additional information such as data identification (ID entity, ID). The file system writes each object into the physical storage space of the logical volume, and the file system records the storage location information of each object, so that when the client requests to access the data, the file system can allow the client to access the data according to the storage location information of each object.
[0099] The process of the storage system allocating physical storage space to a logical volume is as follows: based on the estimated capacity of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of the Redundant Array of Independent Disks (RAID), the physical storage space is divided into stripes in advance. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.
[0100] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0101] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0102] Computer vision technology Computer vision is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify, track and measure targets, and further perform graphic processing so that the computer processing becomes an image that is more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0103] Machine learning is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0104] The following is a brief introduction to the application areas of the intelligent agent interaction method provided in the embodiments of the present application.
[0105] With the continuous development of technology, more and more devices can not only provide virtual interaction scenes for multiple virtual accounts, but also provide a single virtual account with an intelligent agent that can interact virtually in the virtual interaction scene. Taking the game scene as an example, for example, a game account can interact virtually with an intelligent agent when it is not matched with other game accounts; for another example, a game account can improve its own interaction ability or interaction level by interacting with an intelligent agent.
[0106] Taking the simulated driving practice scenario as an example, for example, a virtual account can simulate driving on the client by operating the steering wheel or pedals, the intelligent agent can simulate other vehicles on the road, and the virtual account can interact virtually with the intelligent agent, such as turning on the turn signal to signal other vehicles, turning the steering wheel to change lanes, honking the horn to alert other vehicles, etc., to achieve the purpose of driving practice.
[0107] Taking a simulated interview scenario as an example, for example, a virtual account can control a virtual character through a client to find an interview location in a designated area, control the virtual character to make preparations for the interview at the interview location, and answer questions from the virtual interviewer simulated by the intelligent agent through voice, thereby achieving the purpose of improving interview skills.
[0108] Usually, after a virtual account performs a control operation on a virtual controlled element associated with the virtual account, the agent can only determine a corresponding feedback operation for the virtual controlled element associated with the agent. Take a game scenario as an example, for example, if the virtual account controls a hero associated with the virtual account to perform an attack action, the agent can only control the virtual controlled element associated with the agent that is attacked to perform an evasion operation.
[0109] However, after the virtual account performs the control operation, it is not only the single element of the virtual interactive scene that is affected, but the influence of the control operation is diverse. Take the game scene as an example. For example, when a hero selected by the agent is attacked, another hero can be dispatched to replenish the blood of the hero. The hero under attack can directly attack instead of evading, and use skills to kill the opponent hero.
[0110] Furthermore, the agent is trained based on the virtual interactive scene. The traditional method of training the agent is to train the agent based on the attribute information of the virtual interactive scene. Please refer to Figure 1a Based on the control rules of virtual controlled elements, the location of scene elements and the duration of interaction contained in the attribute information of virtual interactive scenes, the intelligent agent is trained so that the trained intelligent agent can predict the triggered virtual interaction when facing a virtual account. Taking the game scene as an example, for example, based on the skill attributes of the heroes in the game, the location of obstacles in the game map and the duration of each skill possessed by the heroes, the intelligent agent is trained so that the trained intelligent agent can predict the triggered evasive action when facing the attack action of the virtual account.
[0111] However, since the attribute information of the virtual interactive scene is pre-set, the interactive ability of the trained intelligent agent tends to be single. When the intelligent agent faces an action performed by a virtual controlled element associated with a virtual account, it can only control the virtual controlled element associated with the intelligent agent that interacts virtually with the virtual controlled element to perform the corresponding action. Taking the game scene as an example, for example, when the intelligent agent faces an attack action issued by any hero selected by the virtual account, it usually controls the hero selected by the intelligent agent to trigger an evasive action. However, when a hero selected by the intelligent agent is attacked, another hero can be dispatched to replenish the blood of the hero. The hero under attack can not evade, but directly attack, use skills to kill the opponent hero, etc.
[0112] For another example, when the intelligent agent faces a situation where the hero selected by the virtual account enters the attack range of the hero selected by the intelligent agent, the intelligent agent usually controls the hero selected by the intelligent agent to attack the hero selected by the virtual account. However, when the hero selected by the virtual account enters the attack range of the hero selected by the intelligent agent, the hero selected by the intelligent agent can perform evasive actions in the grass and wait for other heroes selected by the intelligent agent to arrive nearby before attacking the hero selected by the virtual account, thereby achieving tactical coordination between teammates.
[0113] In the actual interaction process, the interaction methods and approaches are important, and the interaction process is rich and varied. Traditional intelligent agent interaction methods and traditional methods of training intelligent agents do not consider the real interaction process in virtual interaction scenarios, which makes it possible for trained intelligent agents to be unable to conduct virtual interactions from a macro decision-making perspective and unable to flexibly conduct virtual interactions with virtual accounts in virtual interaction scenarios. As a result, when virtual accounts interact virtually with intelligent agents, they cannot accurately restore the real interaction experience when interacting virtually with other virtual accounts. It can be seen that under existing technologies, the interaction accuracy of intelligent agents is low.
[0114] In order to solve the problem of low accuracy of intelligent agent interaction, this application proposes an intelligent agent interaction method. Figure 1b The method loads a target agent in response to an interaction request instruction triggered by a virtual account. In response to a control operation triggered by the virtual account on a first target virtual controlled element associated with the virtual account in a target virtual interactive scene, a target interactive scene image corresponding to the control operation is obtained. Target interactive state features are extracted from the target interactive scene image. Target scheduling operations and target interactive operations corresponding to the target agent are determined based on the target interactive state features. In response to the target scheduling operations and target interactive operations, a second target virtual controlled element associated with the target agent is controlled in the target virtual interactive scene.
[0115] It should be noted that the target scheduling operation and the target interactive operation are only two angles of division of the operations that the target agent can perform, and are not limited to a certain operation. The target scheduling operation is used to control the second target virtual controlled element to perform the target scheduling action, and the target interactive operation is used to control the second target virtual controlled element to perform the target interactive action. When the target agent is triggered to perform the target scheduling operation and the target interactive operation, it can be that when the target scheduling operation is not scheduled, only the target interactive operation is performed, or when the target interactive operation is not interactive, only the target scheduling operation is performed, or when the target scheduling operation is not scheduled and the target interactive operation is not interactive, no operation is performed, etc. When there are multiple target virtual controlled elements, the target scheduling operation performed by the target agent may include the scheduling action performed by controlling each target virtual controlled element, and the target interactive operation performed by the target agent may include the interactive action performed by controlling each target virtual controlled element. The target agent can control all target virtual controlled elements at the same time, or only one or more target virtual controlled elements, etc., which can be set according to the actual scenario, and will not be specifically introduced here.
[0116] In the embodiment of the present application, in response to the control operation triggered by the virtual account for the first target virtual controlled element, the target scheduling operation and the target interactive operation corresponding to the target intelligent agent can be determined, so that the second target virtual controlled element can be controlled in response to the target scheduling operation and the target interactive operation. The target scheduling operation obtained in response to the control operation can control the second target virtual controlled element to perform the scheduling action, reflecting the interactive strategy obtained by the target intelligent agent in response to the control operation from a macro perspective, rather than simply responding to the control operation in interactive actions, thereby improving the interactive accuracy of the target intelligent agent.
[0117] At the same time, when there are multiple second target virtual controlled elements, it is not a single control of a second target virtual controlled element to respond to the control operation of the virtual account, but the target scheduling operation is obtained by considering the overall situation of all second target virtual controlled elements, thereby controlling one second target virtual controlled element or multiple second target virtual controlled elements based on the target scheduling operation, thereby further improving the interaction accuracy of the target intelligent body.
[0118] The target interactive operation obtained in response to the control operation can control the second target virtual controlled element to perform interactive actions, which reflects the interactive ability of the target intelligent body from a microscopic perspective, and can maximize the interactive ability on the basis of the correct interactive strategy, thereby improving the interactive accuracy of the target intelligent body. From both macroscopic and microscopic perspectives, the macroscopic decision-making ability of the target intelligent body is improved without reducing the interactive ability, so that the target intelligent body can provide accurate and diversified feedback to the virtual account. In the embodiment of the present application, the interaction between the target intelligent body and the virtual account can accurately imitate the real interaction between the virtual accounts, thereby improving the interactive accuracy of the target intelligent body.
[0119] The application scenarios of the intelligent agent interaction method provided in this application are explained below.
[0120] Please refer to Figure 2 , which is an application scenario of the intelligent agent interaction method provided in the embodiment of the present application. The application scenario includes a client 101, an intelligent agent interaction terminal 102, and an intelligent agent training terminal 103. The client 101 and the intelligent agent interaction terminal 102 can communicate with each other, and the intelligent agent interaction terminal 102 and the intelligent agent training terminal 103 can communicate with each other. The communication method can be to communicate using wired communication technology, such as communicating by connecting a network cable or a serial port cable; or to communicate using wireless communication technology, such as communicating through Bluetooth or wireless fidelity (wireless fidelity, WIFI) and other technologies, without specific limitation.
[0121] Client 101 generally refers to a device that can log in to a virtual account, such as a terminal device, a third-party application that a terminal device can access, or a web page that a terminal device can access. For example, terminal devices include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc. Agent interaction terminal 102 generally refers to a device that can virtually interact with a virtual account or agent, such as a terminal device or a server. For example, a server includes a cloud server, a local server, or an associated third-party server. Agent training terminal 103 generally refers to a device that can train an agent, such as a terminal device or a server. Client 101, agent interaction terminal 102, and agent training terminal 103 can all use cloud computing to reduce the occupancy of local computing resources; cloud storage can also be used to reduce the occupancy of local storage resources.
[0122] As an embodiment, the client 101 and the agent interaction terminal 102 can be the same device, the agent interaction terminal 102 and the agent training terminal 103 can be the same device, the client 101 and the agent training terminal 103 can be the same device, and the client 101, the agent interaction terminal 102 and the agent training terminal 103 can be the same device, without specific limitation. In the embodiment of the present application, the client 101, the agent interaction terminal 102 and the agent training terminal 103 are introduced as examples of different devices.
[0123] The following is based on Figure 2 , the intelligent agent interaction method provided in the embodiment of the present application is specifically introduced.
[0124] Before the client 101 performs virtual interaction with the target agent of the agent interaction terminal 102, the agent interaction terminal 102 may first obtain the target agent. After the agent training terminal 103 performs iterative training on the agent to be trained, the target agent is obtained, and the agent training terminal 103 sends the target agent to the agent interaction terminal 102, and the agent interaction terminal 102 receives the target agent sent by the agent training terminal 103.
[0125] The following first introduces the process of training the to-be-trained intelligent agent by the intelligent agent training terminal 103.
[0126] Please refer to Figure 3aBased on the interaction process between the intelligent agent to be trained and the preset reference intelligent agent in the sample virtual interaction scene, the intelligent agent to be trained is trained for multiple rounds of iterations until the preset training goal is met, and the intelligent agent to be trained is output as the target intelligent agent. A virtual interaction scene simulated by two intelligent agents is used to integrate the characteristics of the real interaction process into the training process for the target intelligent agent. Based on the real interaction training intelligent agent in the simulated virtual interaction scene, the trained target intelligent agent can flexibly interact with the virtual account in the real virtual interaction scene, which improves the interaction accuracy of the trained intelligent agent.
[0127] As an embodiment, based on the interaction process between the intelligent agent to be trained and the preset reference intelligent agent in the sample virtual interactive scene, before the intelligent agent to be trained is subjected to multiple rounds of iterative training, the preset reference intelligent agent can be obtained first. There are many methods for obtaining the reference intelligent agent, for example, other devices send the reference intelligent agent to the intelligent agent training terminal 103, and for another example, the reference intelligent agent can be obtained by the intelligent agent training terminal 103 based on the preset reference intelligent agent set. The reference intelligent agent set can include various reference intelligent agents with the same interactive ability as the intelligent agent to be trained, so that the interactive method for obtaining the expected sample interactive result can be learned from multiple interactions with the same strength, thereby improving the interactive accuracy of the trained intelligent agent to be trained. The reference intelligent agent set can also include various reference intelligent agents with higher interactive ability than the intelligent agent to be trained, so that the high-ability interactive method can be learned from the interactive method of the high-ability reference intelligent agent, thereby improving the interactive accuracy of the trained intelligent agent to be trained. The reference agent set may also include reference agents whose interaction capabilities are lower than those of the agent to be trained, so that the interaction methods of the low-ability reference agents can be used as negative teaching materials to learn the interaction methods of the high-ability agents and improve the interaction accuracy of the trained agent to be trained. The reference agent set may also include various types of reference agents, and the reference agents may be randomly selected or each type of reference agent may be extracted with a preset probability, etc., without specific limitation.
[0128] The following is an introduction to the process in which the agent training terminal 103 obtains a reference agent based on a preset reference agent set and trains the agent to be trained.
[0129] Based on the selection probability corresponding to each reference agent in the preset reference agent set, a reference agent is randomly selected from each reference agent. The selection probability corresponding to each reference agent can be determined based on the frequency of each reference agent being selected. The higher the frequency of selection, the lower the selection probability, and the lower the frequency of selection, the higher the selection probability. The selection probability corresponding to each reference agent can also be determined based on the acquisition time of each reference agent. The shorter the time between the acquisition time and the current time, the higher the selection probability, and the longer the time, the lower the selection probability. There is no specific limitation on the selection probability corresponding to each reference agent.
[0130] After extracting the reference agent, the agent to be trained is trained for multiple rounds of iterations based on the interaction process between the agent to be trained and the extracted reference agent in the sample virtual interaction scene. If the agent to be trained does not currently meet the training objectives when the sample interaction results between the agent to be trained and the extracted reference agent are obtained, then a reference agent can be re-extracted from each reference agent, and the agent to be trained is trained for multiple rounds of iterations. After multiple rounds of iterations of training the agent to be trained, if the agent to be trained currently meets the training objectives, the agent to be trained is output as the target agent.
[0131] As an embodiment, if the subject to be trained does not meet the training goal when the sample interaction results between the subject to be trained and the extracted reference subject are obtained, then a reference subject is re-extracted from each reference subject, and the subject to be trained is continuously trained for multiple rounds of iterations. If the subject to be trained meets the training goal, the subject to be trained is output as the target subject.
[0132] As an embodiment, there are multiple methods for obtaining a reference agent set. For example, other devices send a preset reference agent set to the agent training terminal 103. For another example, the reference agent set is obtained during the training of the agent to be trained, etc., without specific limitation.
[0133] The following is an example of the process of obtaining a reference agent set during the training of an agent to be trained.
[0134] When the intelligent agent to be trained has not been trained, the preset reference intelligent agent set may only include the intelligent agent to be trained itself. When the intelligent agent to be trained is subjected to multiple rounds of iterative training based on the interaction process between the intelligent agent to be trained and the reference intelligent agent in the sample virtual interactive scene, the number of training times is accumulated for each round of iterative training, and the number of training times of the iterative training of the intelligent agent to be trained is counted.
[0135] If the counted number of training times does not reach the preset specified number of times, then iterative training continues. If the counted number of training times reaches the specified number of times, then the current agent to be trained is used as a reference agent and added to the reference agent set. After adding the reference agent to the reference agent set, the number of training times can be reset to zero, and the agent to be trained can continue to be iteratively trained, and the new number of training times can be recounted. Before the trained target agent is obtained, the reference agents included in the reference agent set are constantly updated, and new reference agents are continuously added. Reference agents whose multiple sample interaction results do not meet the preset indicators can be removed from the reference agent set.
[0136] Taking the game scenario as an example, if a reference agent in the reference agent set fails to fight against the agent to be trained for many consecutive times, it means that the reference agent has a weak fighting ability and cannot be used to train the agent to be trained. Therefore, the reference agent can be removed from the reference agent set to ensure that each reference agent in the reference agent set maintains the same interaction level, thereby improving the accuracy of training the agent to be trained based on the reference agent. In the process of training the agent to be trained, the virtual interaction data is obtained based on the virtual interaction between agents, and does not require the participation of any virtual account, which reduces the difficulty of obtaining virtual interaction data.
[0137] The following is an example of an iterative training of the training agent. Please refer to Figure 3b , which is a flow chart of the intelligent agent interaction method provided in an embodiment of the present application.
[0138] S301, based on the first sample interaction state feature corresponding to the first sample interaction scene image in the sample virtual interaction scene, predict the sample scheduling operation performed by the to-be-trained intelligent agent on the sample virtual controlled element associated with the to-be-trained intelligent agent in the sample virtual interaction scene, and predict the sample interaction operation performed by the to-be-trained intelligent agent on the sample virtual controlled element after executing the sample scheduling operation.
[0139] Before making a prediction based on the first sample interaction state feature, the first sample interaction state feature corresponding to the first sample interaction scene image in the sample virtual interaction scene can be obtained. There are many methods to obtain the first sample interaction state feature, for example, receiving the first sample interaction state feature corresponding to the first sample interaction scene image sent by other devices, for example, when obtaining the first sample interaction scene image, calculating the first sample interaction state feature in real time, etc. The following is an example of a method for determining the first sample interaction state feature corresponding to the first sample interaction scene image in the sample virtual interaction scene.
[0140] When the intelligent agent to be trained conducts virtual interaction with the reference intelligent agent, each frame of the sample virtual interaction scene can be used as a sample interaction scene image, or certain frames of the sample virtual interaction scene can be used as individual sample interaction scene images, etc., without specific limitation.
[0141] For the first sample interactive scene image, the intelligent agent to be trained can perform region recognition processing on the first sample interactive scene image to obtain a first interactive result region, a first global perspective region, and a first local perspective region. Among them, the first sample interactive scene image can be the first frame image when the intelligent agent to be trained and the reference intelligent agent enter the sample virtual interactive scene, or it can be any frame image in the sample virtual interactive scene, without specific limitation. The first interactive result region is used to characterize the current interaction between the intelligent agent to be trained and the reference intelligent agent, the first global perspective region is used to characterize the position information of the sample virtual controlled elements associated with the intelligent agent to be trained and the reference virtual controlled elements associated with the reference intelligent agent in the virtual interactive scene, and the first local perspective region is used to characterize the position information of the sample virtual controlled elements and the reference virtual controlled elements contained in the sample virtual interactive scene under the perspective of a certain sample virtual controlled element.
[0142] Take the game scenario as an example, for example, please refer to Figure 4a , is a possible interface diagram of the first sample interactive scene image. Please refer to Figure 4b , is the first interactive result area in the first sample interactive scene image, which may include the survival status of each hero associated with the to-be-trained agent, the number of heroes killed by the to-be-trained agent and the reference agent, the number of heroes killed by the opponent, and the duration of the battle. Please refer to Figure 4c , is the first global viewing area in the first sample interactive scene image. The first global viewing area may include the position information of each hero associated with the to-be-trained agent and each hero associated with the reference agent in the sample virtual interactive scene, etc., without any specific limitation. Please refer to Figure 4d , is the first local viewing area in the first sample interactive scene image. The first local viewing area may include the position information of the hero associated with the to-be-trained intelligent agent and the hero associated with the reference intelligent agent contained in the perspective of a certain hero in the sample virtual interactive scene, as well as the skill information that can be used by the hero corresponding to the hero's perspective, etc., without specific limitation.
[0143] After obtaining the first interactive result area, the first global perspective area, and the first local perspective area, image feature extraction processing can be performed on the first interactive result area, the first global perspective area, and the first local perspective area, respectively, to obtain the corresponding first feature vector, the first global perspective feature matrix, and the first local perspective feature matrix, respectively. The first feature vector, the first global perspective feature matrix, and the first local perspective feature matrix are used as the first sample interactive state feature corresponding to the first sample interactive scene image.
[0144] The first eigenvector is used to characterize the interaction information related to the sample interaction results, the first global perspective feature matrix is used to characterize the position information of the sample virtual controlled elements, the position information of the reference virtual controlled elements associated with the reference agent, and the position information of the scene elements contained in the sample virtual interaction scene, and the first local perspective feature matrix is used to characterize the position information of the sample virtual controlled elements contained in the first local perspective area, the position information of the reference virtual controlled elements contained in the first local perspective area, and the position information of the scene elements contained in the first local perspective area.
[0145] As an embodiment, the intelligent agent to be trained may include a quantitative information extraction module, wherein the quantitative information extraction module is used to extract sample interaction state features corresponding to each sample interaction scene image generated based on virtual interaction in each preset sample virtual interaction scene.
[0146] After obtaining the first sample interaction state feature corresponding to the first sample interaction scene image, it is possible to predict the sample scheduling operation performed by the to-be-trained intelligent agent on the sample virtual controlled element associated with the to-be-trained intelligent agent in the sample virtual interaction scene based on the first sample interaction state feature corresponding to the first sample interaction scene image in the sample virtual interaction scene, and to predict the sample interaction operation performed by the to-be-trained intelligent agent on the sample virtual controlled element after executing the sample scheduling operation. The following is an example introduction to the process of predicting the sample scheduling operation and the sample interaction operation, respectively.
[0147] Forecast sample scheduling operation:
[0148] Based on the first eigenvector and the first global perspective feature matrix, the sample scheduling operation performed by the agent to be trained on the sample virtual controlled element can be predicted. Based on the first eigenvector, the current interaction situation can be obtained, and the gap between the current interaction situation and the expected sample interaction result can be obtained; based on the first global perspective feature matrix, the position information of the sample virtual controlled element associated with the agent to be trained and the position information of the reference virtual controlled element associated with the reference agent can also be obtained from a macro perspective. Therefore, when predicting the sample scheduling operation performed by the agent to be trained on the sample virtual controlled element, the favorable position information of the sample virtual controlled element can be analyzed in the direction of narrowing the gap between the current interaction situation and the expected sample interaction result, and the sample scheduling operation can be obtained. The sample scheduling operation is predicted based on the characteristics of the virtual account operating the virtual controlled element in the real interaction process, thereby improving the interaction accuracy of the trained agent. The sample scheduling operation can include the scheduling direction, please refer to Figure 5a , a specified number of direction angles can be divided according to the current position of the sample virtual controlled element. The sample scheduling operation can also include a scheduling distance, so that the sample virtual controlled element can be controlled to move a corresponding scheduling distance in the scheduling direction through the sample scheduling operation, so that the sample virtual controlled element moves from the current position to the target position.
[0149] The agent to be trained can also divide the first global view area into multiple sub-areas. Based on the first feature vector and the first global view feature matrix, predict the target sub-area corresponding to the sample virtual controlled element. Based on the sub-area where the sample virtual controlled element is currently located and the target sub-area corresponding to the sample virtual controlled element, obtain the sample scheduling operation. Taking the game scene as an example, please refer to Figure 5b , which is the first global viewing area divided into multiple sub-areas.
[0150] After obtaining the sample scheduling operation, based on the sample scheduling operation, the sample virtual controlled element is controlled to move to the corresponding target sub-area to obtain a second sample interactive scene image generated by the to-be-trained intelligent agent. The second sample interactive state feature corresponding to the second sample interactive scene image is extracted to obtain a second feature vector, a second global perspective feature matrix, and a second local perspective feature matrix.
[0151] Based on the sample scheduling operation, there are multiple processes for controlling the sample virtual controlled element to move to the corresponding target sub-area to obtain the second sample interactive scene image generated by the intelligent agent to be trained, for example, controlling the sample virtual controlled element to move from the current sub-area to the corresponding target sub-area, and starting the timing. Based on the preset moving speed for the sample virtual controlled element, determine the reference time for the sample virtual controlled element to move from the current sub-area to the corresponding target sub-area. If the timing time reaches the reference time, the second sample interactive scene image generated by the intelligent agent to be trained is obtained.
[0152] For another example, the sample virtual controlled element is controlled to move from the current sub-region to the corresponding target sub-region, and the moving distance is recorded. If the recorded moving distance reaches the reference distance between the current sub-region where the sample virtual controlled element is located and the target sub-region, a second sample interactive scene image generated by the intelligent agent to be trained is obtained.
[0153] Predict sample interactive operations:
[0154] Based on the sample scheduling operation, the first global perspective feature matrix and the first local perspective feature matrix, the predicted feature vector, the predicted global perspective feature matrix and the predicted local perspective feature matrix corresponding to the predicted interactive scene image generated by the to-be-trained intelligent agent after the sample scheduling operation is performed are predicted. The predicted interactive scene image is a scene image when the sample virtual controlled element reaches the target sub-area after the to-be-trained intelligent agent performs the sample scheduling operation.
[0155] After obtaining the predicted feature vector, predicted global perspective feature matrix and predicted local perspective feature matrix corresponding to the predicted interactive scene image, the sample interactive operation performed by the to-be-trained intelligent agent on the sample virtual controlled element is predicted based on the predicted feature vector and the predicted local perspective feature matrix. The predicted feature vector can characterize the possible interactive situation after the sample scheduling operation is performed, thereby further characterizing the gap between the possible interactive situation after the sample scheduling operation is performed and the expected sample interactive result. The predicted global perspective feature matrix can characterize the position information of the sample virtual controlled element and the position information of the reference virtual controlled element after the sample virtual controlled element reaches the target sub-region, thereby further characterizing whether the relative position between the sample virtual controlled element and the reference virtual controlled element is favorable. The predicted local perspective feature matrix can characterize the position information of the sample virtual controlled element and the position information of the reference virtual controlled element in the predicted local perspective region where each sample virtual controlled element is located, thereby further characterizing whether the relative position between the sample virtual controlled element and the reference virtual controlled element is favorable. Thus, the next sample interactive operation can be predicted based on the predicted situation after the sample scheduling operation is performed.
[0156] After obtaining the predicted sample scheduling operation and sample interaction operation, the model parameters of the intelligent agent to be trained are adjusted based on the second sample interaction state feature corresponding to the second sample interaction scene image generated after executing the sample scheduling operation, and the third sample interaction state feature corresponding to the third sample interaction scene image generated after executing the sample interaction operation. There are many methods for adjusting the model parameters of the intelligent agent to be trained based on the second sample interaction state feature and the third sample interaction state feature, for example, adjusting the model parameters of the intelligent agent to be trained based on the error value between the second sample interaction state feature and the third sample interaction state feature and the preset interaction state feature. For another example, the model parameters of the intelligent agent to be trained are adjusted based on the scheduling incentive data and interaction incentive data corresponding to the second sample interaction state feature and the third sample interaction state feature. Steps S302 to S304 are introduced as an example of a method for adjusting the model parameters of the intelligent agent to be trained based on the scheduling incentive data and interaction incentive data corresponding to the second sample interaction state feature and the third sample interaction state feature.
[0157] S302, based on the second sample interaction state feature corresponding to the second sample interaction scene image generated by the to-be-trained intelligent agent after executing the sample scheduling operation, determine the scheduling incentive data of the sample scheduling operation according to the preset scheduling incentive strategy.
[0158] The first scheduling sub-stimulus of the sample scheduling operation is determined based on whether the current position of the sample virtual controlled element matches the target position corresponding to the sample virtual controlled element indicated by the sample scheduling operation, or based on whether the sub-region where the sample virtual controlled element is currently located matches the target sub-region corresponding to the sample virtual controlled element indicated by the sample scheduling operation.
[0159] If it matches, then the first scheduling sub-excitation is a positive excitation, and if it does not match, then the first scheduling sub-excitation is a negative excitation. For example, based on the distance difference between the current position of the sample virtual controlled element and the target position corresponding to the sample virtual controlled element indicated by the sample scheduling operation, if it is determined that the distance difference is less than the first specified distance threshold, then it is determined that the current position of the sample virtual controlled element matches the target position corresponding to the sample virtual controlled element indicated by the sample scheduling operation, and a preset excitation value is given. If the distance difference is greater than the first specified distance threshold and less than the second specified distance threshold, then it is determined that the current position of the sample virtual controlled element does not completely match the target position corresponding to the sample virtual controlled element indicated by the sample scheduling operation, and a value corresponding to a specified percentage of the preset excitation value is given. If the distance difference is greater than the third specified distance threshold, then it is determined that the current position of the sample virtual controlled element does not match the target position corresponding to the sample virtual controlled element indicated by the sample scheduling operation, and a preset negative excitation value is given.
[0160] Based on the change value between the first eigenvector and the second eigenvector corresponding to the sample virtual controlled element, the second scheduling sub-incentive of the sample scheduling operation is determined. Taking the game scene as an example, the second scheduling sub-incentive of the sample scheduling operation can be determined according to the change of hero experience value, gold coin change, health change, number of kills, number of kills, and health change of the main buildings.
[0161] After obtaining the first scheduling sub-incentive and the second scheduling sub-incentive, the scheduling incentive data of the sample scheduling operation can be determined based on the weighted sum of the first scheduling sub-incentive and the second scheduling sub-incentive, please refer to formula (1). The weight can be a pre-set value or a value learned during the training of the intelligent agent to be trained, and there is no specific limitation.
[0162] R t =w d *R d +w e *R e (1)
[0163] Among them, R t is the first scheduling sub-stimulus R d and the second scheduling sub-stimulus R e The weighted sum of d is the first scheduling sub-stimulus R d The weight, w e The second scheduling sub-stimulus R e The weight of .
[0164] After obtaining the weighted sum of the first scheduling sub-incentive and the second scheduling sub-incentive, the cumulative sum of the weighted sums of all first scheduling sub-incentives and second scheduling sub-incentives obtained from the current process to the end of the interaction between the training agent and the reference agent can be determined based on the Bellman Equation. Please refer to formulas (2) and (3).
[0165]
[0166] V(S t )=E[R t+1 +λV(S t+1 )|S t =s] (3)
[0167] Among them, the scheduling incentive data V(S t ) can be the expected value of the weighted sum of all first scheduling sub-incentives and second scheduling sub-incentives obtained by the training agent and the reference agent from the current to the end of the interaction, λ k is the attenuation coefficient, and s represents the virtual scene at the corresponding moment of the first sample interactive scene image.
[0168] The scheduling incentive data includes not only data used to characterize the degree of influence of sample scheduling operations on sample interaction results, but also data used to characterize the degree of completion of sample scheduling operations, that is, dense incentive data and sparse incentive data are combined to obtain scheduling incentive data, which avoids the situation where local optimal solutions appear when the intelligent agent predicts virtual interactive operations, reduces the accuracy of interaction, and improves the decision-making ability of the trained target intelligent agent.
[0169] S303, based on the third sample interaction state feature corresponding to the third sample interaction scene image generated by the to-be-trained intelligent agent after executing the sample interaction operation, determine the interaction incentive data of the sample interaction operation according to the preset interaction incentive strategy.
[0170] After performing the sample interaction operation, a third sample interaction scene image can be obtained. The quantitative information extraction module can be used to extract the third sample interaction state features corresponding to the third sample interaction scene image to obtain a third feature vector, a third global perspective feature matrix and a third local perspective feature matrix.
[0171] Based on the change value between the second eigenvector and the third eigenvector corresponding to the sample virtual controlled element, the interactive incentive of the sample interactive operation is determined. Based on the interactive incentive, the scheduling incentive data of the sample scheduling operation is determined. Taking the game scene as an example, the interactive incentive of the sample interactive operation can be determined based on the changes in the hero's experience value, gold coins, health, kills, kills, and health of the main buildings. Similarly, the cumulative sum or the expectation of the cumulative sum of all possible interactive incentives that can be obtained from the current process to the end of the interaction between the trained agent and the reference agent can be determined based on the Bellman Equation.
[0172] As an embodiment, the intelligent agent to be trained may also include a training module, wherein the training module is used to obtain each scheduling incentive data and each interaction incentive data based on each sample interaction state feature, and adjust the model parameters of the intelligent agent to be trained based on the obtained each scheduling incentive data and each interaction incentive data.
[0173] S304, respectively determining the error values between the scheduling incentive data and the interaction incentive data and the preset target incentive data, and adjusting the model parameters of the intelligent agent to be trained based on the obtained error values.
[0174] If the training module includes a scheduling model and an interaction model, the model parameters of the scheduling model and the model parameters of the interaction model can be adjusted based on the scheduling incentive data and the interaction incentive data, and reinforcement learning can be performed on the scheduling model and the interaction model. For example, the error values between the scheduling incentive data and the interaction incentive data and the preset target incentive data are determined respectively. After obtaining each error value, the model parameters of the scheduling model and the model parameters of the interaction model can be adjusted according to each error value obtained.
[0175] For another example, if the target incentive data includes scheduling target incentive data and interactive target incentive data, then the scheduling error value between the scheduling incentive data and the scheduling target incentive data can be determined, and the model parameters of the scheduling model can be adjusted based on the obtained scheduling error value. At the same time, the interactive error value between the interactive incentive data and the interactive target incentive data can also be determined, and the model parameters of the interactive model can be adjusted based on the obtained interactive error value.
[0176] By using a hierarchical reinforcement learning method from both macro and micro perspectives, complex prediction problems can be simplified into two simple prediction problems. Without lowering the training standard for the agent's interactive ability, the training of the agent's macro decision-making ability is added, so that the trained target agent can accurately achieve the real interactive effect of the virtual account during virtual interaction, thereby improving the interaction accuracy of the target agent.
[0177] As an embodiment, in the process of training the intelligent agent to be trained, after each round of iterative training, the trained intelligent agent to be trained can be evaluated, and the evaluation value of the intelligent agent to be trained is determined based on the scheduling incentive data and interaction incentive data obtained each time before, and the scheduling incentive data and interaction incentive data obtained this time, and the evaluation value is used to characterize the training degree of the intelligent agent to be trained. For example, the training degree of the intelligent agent to be trained is evaluated through the ELO evaluation mechanism.
[0178] In the process of training the intelligent agent to be trained, the trained intelligent agent to be trained can also be evaluated after multiple rounds of iterative training, and the specific evaluation timing is not limited. When the evaluation values obtained multiple times tend to converge, the intelligent agent to be trained is output as the target intelligent agent.
[0179] During the training process of the intelligent agent to be trained, the number of iterations can also be counted. If the number of iterations reaches a preset maximum number, the intelligent agent to be trained is output as the target intelligent agent.
[0180] As an embodiment, the training process of the agent to be trained can be easily and quickly expanded to multiple machines in parallel through multi-container Docker images according to the available machine capacity, so please refer to Figure 6, you can train the to-be-trained intelligent agent and different reference intelligent agents on multiple machines at the same time, or you can train multiple to-be-trained intelligent agents on multiple machines at the same time, etc., which greatly improves the efficiency of AI battle data generation.
[0181] After obtaining the target agent, the agent training terminal 103 may send the target agent to the agent interaction terminal 102 , so that the agent interaction terminal 102 uses the target agent after receiving the target agent sent by the agent training terminal 103 .
[0182] The following is an introduction to the process of interacting with the target agent.
[0183] Please refer to Figure 7 , is a flowchart of the interaction process using the target agent.
[0184] S701, in response to an interaction request instruction triggered by a virtual account, loading a target agent.
[0185] The virtual account can trigger an interaction request instruction through the client, and load the target agent in response to the interaction request instruction triggered by the virtual account. Thus, the virtual account can perform virtual interaction with the target agent in the target virtual interaction scene.
[0186] S702, in response to a control operation triggered by a virtual account on a first target virtual controlled element associated with the virtual account in a target virtual interactive scene, obtaining a target interactive scene image corresponding to the control operation.
[0187] The virtual account can trigger a control operation on the first target virtual controlled element associated with the virtual account in the target virtual interactive scene through the client, and obtain the target interactive scene image corresponding to the control operation in response to the control operation triggered by the virtual account on the first target virtual controlled element associated with the virtual account in the target virtual interactive scene.
[0188] S703, extracting target interaction state features from the target interaction scene image.
[0189] After obtaining the target interaction scene image corresponding to the control operation, the target interaction state features of the target interaction scene image can be extracted. The process of extracting the target interaction state features can refer to the process of extracting the sample interaction state features of the sample interaction scene image introduced above, and will not be repeated here.
[0190] S704, determining the target scheduling operation and target interaction operation corresponding to the target agent based on the target interaction state characteristics.
[0191] After obtaining the target interaction state characteristics, the target scheduling operation and target interaction operation corresponding to the target agent can be determined based on the target interaction state characteristics. The process of determining the target scheduling operation and the target interaction operation can refer to the process of determining the sample scheduling operation and the sample interaction operation introduced in the previous article, which will not be repeated here.
[0192] S705, in response to the target scheduling operation and the target interaction operation, controlling a second target virtual controlled element associated with the target agent in the target virtual interaction scene.
[0193] After determining the target scheduling operation and the target interaction operation, the target intelligent agent executes the target scheduling operation and the target interaction operation, and in response to the target scheduling operation and the target interaction operation, controls the second target virtual controlled element associated with the target intelligent agent in the target virtual interaction scene to realize the interaction between the target intelligent agent and the virtual account.
[0194] The target agent and the virtual account can interact multiple times until an interaction end instruction is received, ending the interaction between the target agent and the virtual account. This application will not repeat each interaction. If an interaction end instruction is received, then based on the interaction end instruction, a target interaction result between the virtual account and the target agent in the target virtual interaction scenario is generated. The interaction end instruction can be triggered by the virtual account through the client, or by the target agent, or it can be automatically generated at the end of the virtual interaction process, without specific restrictions.
[0195] Among them, the target intelligent agent is obtained by training based on sample virtual interaction data. The specific training process can refer to the previous introduction to the training process of the intelligent agent to be trained.
[0196] The following uses a game scenario as an example to illustrate the intelligent agent interaction method provided in the embodiment of the present application.
[0197] Reference agents are sequentially extracted from the reference agent set to perform reinforcement learning training on the training agent. In the first training, the training agent can play a game with itself. As the reference agent set expands, other reference agents can be extracted from the reference agent set to perform reinforcement learning training on the training agent in subsequent training until the evaluation value of the training agent converges or the number of training times reaches the upper limit, then the training agent is output as the target agent.
[0198] In a training process, the trained agent can control multiple sample virtual controlled elements, i.e., sample heroes, and the reference agent can control multiple reference virtual controlled elements, i.e., multiple reference heroes. The virtual interactive scene also includes multiple environmental elements, i.e., shields and NPCs in the battle game. After the trained agent and the reference agent enter the battle game, the corresponding first sample interactive scene image can be obtained from the perspective of each sample hero.
[0199] Based on the first sample interactive state feature corresponding to the first sample interactive scene image, the sample scheduling operation performed by the to-be-trained agent on each sample hero in the battle game can be predicted, and the sample scheduling operation can include the scheduling action corresponding to each sample hero. The sample interactive operation performed by the to-be-trained agent on the sample hero after performing the sample scheduling operation can also be predicted, and the sample interactive operation can include the skill action corresponding to each sample hero.
[0200] After executing the sample scheduling operation for each sample hero, a second sample interactive scene image is obtained. Based on the second sample interactive state feature corresponding to the second sample interactive scene image, the scheduling incentive data of the sample scheduling operation can be determined. After executing the sample interactive operation for each sample hero, a third sample interactive scene image is obtained. Based on the third sample interactive state feature corresponding to the third sample interactive scene image, the interactive incentive data of the sample interactive operation can be determined.
[0201] Reinforcement learning training can be performed on the training agent based on the obtained scheduling incentive data and interaction incentive data. For example, hierarchical reinforcement learning training can be performed based on the scheduling incentive data and interaction incentive data, respectively.
[0202] After multiple rounds of training to obtain the target agent, the virtual account can click the "Smart Battle" button in the battle game to load the target agent. The virtual account can control the first target virtual controlled element, that is, the first target hero to move to the camp where the target agent is located, and perform skill attacks on the second target virtual controlled element that the target agent can control, that is, the second target hero or NPC, and avoid the skill attacks performed by the second target hero or NPC on the first target hero.
[0203] If the virtual account kills all the second target heroes or NPCs, or the battle game time is over, an interaction end instruction can be generated. After obtaining the interaction end instruction, a target interaction result can be generated. The target interaction result can indicate whether the virtual account wins or the target intelligent entity wins. It can also indicate how many second target heroes and NPCs each first target hero in the virtual account has killed. It can also indicate how many first target heroes and NPCs each second target hero in the target intelligent entity has killed, etc. There is no specific restriction.
[0204] Based on the same inventive concept, the embodiment of the present application provides an intelligent agent interaction device, which is equivalent to the target intelligent agent discussed above and can realize the functions corresponding to the aforementioned intelligent agent interaction method. Figure 8 , the device includes a loading module 801 and a processing module 802, wherein:
[0205] Loading module 801: used to load the target agent in response to the interactive request instruction triggered by the virtual account;
[0206] Processing module 802: for acquiring a target interactive scene image corresponding to the control operation in response to a control operation triggered by a virtual account on a first target virtual controlled element associated with the virtual account in a target virtual interactive scene;
[0207] The processing module 802 is also used to: extract target interaction state features from the target interaction scene image;
[0208] The processing module 802 is also used to: determine the target scheduling operation and the target interaction operation corresponding to the target agent based on the target interaction state characteristics;
[0209] The processing module 802 is also used to: in response to the target scheduling operation and the target interaction operation, control the second target virtual controlled element associated with the target agent in the target virtual interaction scene.
[0210] In one possible embodiment, the target agent is trained in the following manner:
[0211] The processing module 802 is further used to: perform multiple rounds of iterative training on the intelligent agent to be trained based on the interaction process between the intelligent agent to be trained and the preset reference intelligent agent in the sample virtual interactive scene, until the preset training target is met, and output the intelligent agent to be trained as the target intelligent agent, wherein in one round of iterative training, the processing module 802 is specifically used to:
[0212] Based on the first sample interactive state feature corresponding to the first sample interactive scene image in the sample virtual interactive scene, predict the sample scheduling operation performed by the to-be-trained intelligent agent on the sample virtual controlled element associated with the to-be-trained intelligent agent in the sample virtual interactive scene, and predict the sample interactive operation performed by the to-be-trained intelligent agent on the sample virtual controlled element after performing the sample scheduling operation;
[0213] Based on the second sample interaction state characteristics corresponding to the second sample interaction scene image generated after executing the sample scheduling operation, and the third sample interaction state characteristics corresponding to the third sample interaction scene image generated after executing the sample interaction operation, the model parameters of the intelligent agent to be trained are adjusted.
[0214] In a possible embodiment, the processing module 802 is specifically configured to:
[0215] Based on the second sample interaction state feature, according to the preset scheduling incentive strategy, determining the scheduling incentive data of the sample scheduling operation, wherein the scheduling incentive data is used to characterize the degree of completion of the sample scheduling operation and the degree of influence of the sample scheduling operation on the sample interaction result;
[0216] Based on the third sample interaction state feature, and in accordance with a preset interaction incentive strategy, determining interaction incentive data of the sample interaction operation, wherein the interaction incentive data is used to characterize the degree of influence of the sample interaction operation on the sample interaction result;
[0217] The error values between the scheduling incentive data and the interaction incentive data and the preset target incentive data are determined respectively, and the model parameters of the intelligent agent to be trained are adjusted based on the obtained error values.
[0218] In a possible embodiment, the processing module 802 is further configured to:
[0219] After adjusting the model parameters of the intelligent agent to be trained based on the obtained error values, the evaluation value of the intelligent agent to be trained is determined according to a preset scoring strategy based on the scheduling incentive data and the interaction incentive data obtained through multiple rounds of iterative training, wherein the evaluation value is used to characterize the training degree of the intelligent agent to be trained;
[0220] If the evaluation value converges, the output of the agent to be trained is used as the target agent.
[0221] In a possible embodiment, the processing module 802 is specifically configured to:
[0222] Based on the selection probability corresponding to each reference agent in the preset reference agent set, a reference agent is randomly selected from each reference agent;
[0223] Based on the interaction process between the intelligent agent to be trained and the extracted reference intelligent agent in the sample virtual interaction scene, the intelligent agent to be trained is trained for multiple rounds of iterations;
[0224] If the subject to be trained does not meet the training objectives when the sample interaction results between the subject to be trained and the extracted reference subject are obtained, a reference subject is re-extracted from each reference subject, and multiple rounds of iterative training are continued for the subject to be trained;
[0225] If the intelligent agent to be trained meets the training objectives, the intelligent agent to be trained is output as the target intelligent agent.
[0226] In a possible embodiment, the processing module 802 is further configured to:
[0227] Before outputting the intelligent agent to be trained as the target intelligent agent, counting the number of iterative trainings of the intelligent agent to be trained;
[0228] If the statistical training times reach the preset specified times, the output of the agent to be trained is used as the reference agent and added to the reference agent set;
[0229] Reset the number of training times to zero, continue iterative training of the agent to be trained, and update the reference agent set based on the re-counted number of training times.
[0230] In a possible embodiment, the processing module 802 is further configured to:
[0231] Based on a first sample interaction state feature corresponding to a first sample interaction scene image in a sample virtual interaction scene, predict the sample scheduling operation performed by the to-be-trained intelligent agent on a sample virtual controlled element associated with the to-be-trained intelligent agent in the sample virtual interaction scene, and predict the to-be-trained intelligent agent after performing the sample scheduling operation and before performing the sample interaction operation on the sample virtual controlled element, perform region recognition processing on the first sample interaction scene image to obtain a first interaction result region, a first global perspective region, and a first local perspective region;
[0232] Performing image feature extraction processing on the first interaction result area, the first global perspective area, and the first local perspective area respectively, and obtaining corresponding first eigenvectors, first global perspective feature matrices, and first local perspective feature matrices respectively, wherein the first eigenvector is used to characterize the interaction information related to the sample interaction result, the first global perspective feature matrix is used to characterize the position information of the sample virtual controlled elements, the position information of the reference virtual controlled elements associated with the reference agent, and the position information of the scene elements included in the sample virtual interaction scene, and the first local perspective feature matrix is used to characterize the position information of the sample virtual controlled elements included in the first local perspective area, the position information of the reference virtual controlled elements included in the first local perspective area, and the position information of the scene elements included in the first local perspective area;
[0233] The first eigenvector, the first global perspective feature matrix and the first local perspective feature matrix are used as first sample interaction state features corresponding to the first sample interaction scene image.
[0234] In a possible embodiment, the processing module 802 is specifically configured to:
[0235] Based on the first feature vector and the first global perspective feature matrix, predicting a sample scheduling operation performed by the to-be-trained agent on the sample virtual controlled element;
[0236] Based on the sample scheduling operation, the first global perspective feature matrix and the first local perspective feature matrix, predict the predicted feature vector, the predicted global perspective feature matrix and the predicted local perspective feature matrix corresponding to the predicted interactive scene image generated after the sample scheduling operation is performed;
[0237] Based on the predicted feature vector and the predicted local view feature matrix, the sample interactive operations performed by the to-be-trained agent on the sample virtual controlled elements are predicted.
[0238] In a possible embodiment, the processing module 802 is specifically configured to:
[0239] Dividing the first global viewing area into a plurality of sub-areas;
[0240] Based on the first eigenvector and the first global perspective feature matrix, predict the target sub-region corresponding to the sample virtual controlled element;
[0241] Based on the sub-region where the sample virtual controlled element is currently located and the target sub-region corresponding to the sample virtual controlled element, a sample scheduling operation is obtained.
[0242] In a possible embodiment, the processing module 802 is further configured to:
[0243] After obtaining the sample scheduling operation based on the sub-region where the sample virtual controlled element is currently located and the target sub-region corresponding to the sample virtual controlled element, the sample virtual controlled element is controlled to move to the corresponding target sub-region based on the sample scheduling operation to obtain a second sample interactive scene image generated by the to-be-trained intelligent agent;
[0244] The second sample interaction state feature corresponding to the second sample interaction scene image is extracted to obtain a second feature vector, a second global perspective feature matrix, and a second local perspective feature matrix.
[0245] In a possible embodiment, the processing module 802 is further configured to:
[0246] After extracting the second sample interaction state feature corresponding to the second sample interaction scene image, determining the first scheduling sub-stimulus of the sample scheduling operation based on whether the sub-region where the sample virtual controlled element is currently located matches the target sub-region corresponding to the sample virtual controlled element indicated by the sample scheduling operation;
[0247] Determine a second scheduling sub-stimulus of the sample scheduling operation based on a change value between a first eigenvector and a second eigenvector corresponding to the sample virtual controlled element;
[0248] Based on a weighted sum of the first scheduling sub-stimulus and the second scheduling sub-stimulus, scheduling stimulus data for the sample scheduling operation is determined.
[0249] In a possible embodiment, the intelligent agent to be trained includes a quantitative information extraction module, wherein the quantitative information extraction module is used to extract sample interaction state features corresponding to each sample interaction scene image;
[0250] The intelligent agent to be trained also includes a training module, wherein the training module is used to obtain each scheduling incentive data and each interaction incentive data based on each sample interaction state feature, and adjust the model parameters of the intelligent agent to be trained based on the obtained each scheduling incentive data and each interaction incentive data.
[0251] In a possible embodiment, the processing module 802 is specifically configured to:
[0252] If the training module includes a scheduling model and an interactive model, and the target incentive data includes scheduling target incentive data and interactive target incentive data, a scheduling error value between the scheduling incentive data and the scheduling target incentive data is determined, and a model parameter of the scheduling model is adjusted based on the obtained scheduling error value;
[0253] Determine an interaction error value between the interaction incentive data and the interaction target incentive data, and adjust a model parameter of the interaction model based on the obtained interaction error value.
[0254] Based on the same inventive concept, an embodiment of the present application provides a computer device, and the computer device 900 is introduced below.
[0255] Please refer to Fig. 9 The above-mentioned intelligent agent interaction device can run on a computer device 900, and the current version and historical version of the data storage program and the application software corresponding to the data storage program can be installed on the computer device 900. The computer device 900 includes a display unit 940, a processor 980 and a memory 920, wherein the display unit 940 includes a display panel 941 for displaying a user interactive operation interface, etc.
[0256] In a possible embodiment, the display panel 941 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).
[0257] The processor 980 is used to read the computer program and then execute the method defined by the computer program. For example, the processor 980 reads the data storage program or file, so as to run the data storage program on the computer device 900 and display the corresponding interface on the display unit 940. The processor 980 may include one or more general-purpose processors and may also include one or more DSPs (Digital Signal Processors) to perform related operations to implement the technical solutions provided in the embodiments of the present application.
[0258] The memory 920 generally includes internal memory and external memory, and the internal memory can be a random access memory (RAM), a read-only memory (ROM), and a cache (CACHE), etc. The external memory can be a hard disk, an optical disk, a USB disk, a floppy disk, or a tape drive, etc. The memory 920 is used to store computer programs and other data, and the computer program includes applications corresponding to each client, etc. Other data may include data generated after the operating system or application is run, and the data includes system data (such as configuration parameters of the operating system) and user data. In the embodiment of the present application, program instructions are stored in the memory 920, and the processor 980 executes the program instructions stored in 920 to implement any of the intelligent agent interaction methods discussed in the previous figure.
[0259] The display unit 940 is used to receive input digital information, character information or contact touch operation / contactless gesture, and generate signal input related to user settings and function control of the computer device 900. Specifically, in the embodiment of the present application, the display unit 940 may include a display panel 941. The display panel 941 is, for example, a touch screen, which can collect user touch operations on or near it (such as operations performed by the user using fingers, stylus, or any other suitable object or accessory on or on the display panel 941), and drive corresponding connection devices according to a pre-set program.
[0260] In a possible embodiment, the display panel 941 may include two parts: a touch detection device and a touch controller. The touch detection device detects the touch position of the player, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 980, and can receive and execute the command sent by the processor 980.
[0261] The display panel 941 may be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 940, the computer device 900 may further include an input unit 930, which may include a graphic input device 931 and other input devices 932, wherein the other input devices may include, but are not limited to, one or more of a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, and a joystick.
[0262] In addition to the above, the computer device 900 may also include a power supply 990 for supplying power to other modules, an audio circuit 960, a near field communication module 970, and an RF circuit 910. The computer device 900 may also include one or more sensors 950, such as an acceleration sensor, a light sensor, a pressure sensor, etc. The audio circuit 960 specifically includes a speaker 961 and a microphone 962, etc. For example, the computer device 900 can collect the user's voice through the microphone 962 to perform corresponding operations, etc.
[0263] As an embodiment, the number of processors 980 may be one or more, and the processor 980 and the memory 920 may be coupled or relatively independently configured.
[0264] As an example, Fig. 9 The processor 980 in the embodiment can be used to implement the following Figure 8 The functions of the loading module 801 and the processing module 802.
[0265] As an example, Fig. 9 The processor 980 in can be used to implement the corresponding functions of the server 102 discussed above.
[0266] A person skilled in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.
[0267] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0268] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An agent interaction method, characterized in that: include: In response to the interactive request instruction triggered by the virtual account, the target intelligent agent is loaded; wherein the target intelligent agent is obtained by performing multiple rounds of iterative training on the intelligent agent to be trained; each round of training is performed: based on the scheduling incentive data and interactive incentive data determined in this round of training, the model parameters of the intelligent agent to be trained are adjusted; the scheduling incentive data represents: the degree of completion of the sample scheduling operation predicted in this round of training, and the degree of influence of the sample scheduling operation on the sample interaction result generated in this round of training; the interactive incentive data represents: the degree of influence of the sample interaction operation predicted in this round of training on the sample interaction result; In response to a control operation triggered by the virtual account on a first target virtual controlled element associated with the virtual account in a target virtual interactive scene, acquiring a target interactive scene image corresponding to the control operation; Extracting target interaction state features from the target interaction scene image; Determine the target scheduling operation and the target interaction operation corresponding to the target agent based on the target interaction state characteristics; In response to the target scheduling operation and the target interaction operation, a second target virtual controlled element associated with the target agent is controlled in the target virtual interaction scene.
2. The method according to claim 1, characterized in that: The target agent is trained in the following way: Based on the interaction process between the intelligent agent to be trained and the preset reference intelligent agent in the sample virtual interactive scene, the intelligent agent to be trained is trained for multiple rounds until the preset training target is met, and the intelligent agent to be trained is output as the target intelligent agent, wherein in one round of iterative training, the following operations are performed: Based on a first sample interactive state feature corresponding to a first sample interactive scene image in the sample virtual interactive scene, predicting a sample scheduling operation performed by the to-be-trained agent on a sample virtual controlled element associated with the to-be-trained agent in the sample virtual interactive scene, and predicting a sample interactive operation performed by the to-be-trained agent on the sample virtual controlled element after executing the sample scheduling operation; Based on the second sample interaction state characteristics corresponding to the second sample interaction scene image generated after executing the sample scheduling operation, and the third sample interaction state characteristics corresponding to the third sample interaction scene image generated after executing the sample interaction operation, the model parameters of the intelligent agent to be trained are adjusted.
3. The method according to claim 2, characterized in that Based on the second sample interaction state feature corresponding to the second sample interaction scene image generated after executing the sample scheduling operation, and the third sample interaction state feature corresponding to the third sample interaction scene image generated after executing the sample interaction operation, adjusting the model parameters of the to-be-trained intelligent agent includes: Based on the second sample interaction state feature, according to the preset scheduling incentive strategy, determining the scheduling incentive data of the sample scheduling operation, wherein the scheduling incentive data is used to characterize the degree of completion of the sample scheduling operation and the degree of influence of the sample scheduling operation on the sample interaction result; Based on the third sample interaction state feature, and in accordance with a preset interaction incentive strategy, determining interaction incentive data of the sample interaction operation, wherein the interaction incentive data is used to characterize the degree of influence of the sample interaction operation on the sample interaction result; The error values between the scheduling incentive data and the interaction incentive data and the preset target incentive data are determined respectively, and the model parameters of the to-be-trained intelligent agent are adjusted based on the obtained error values.
4. The method according to claim 2, characterized in that: Based on the interaction process between the intelligent agent to be trained and the preset reference intelligent agent in the sample virtual interactive scene, the intelligent agent to be trained is trained for multiple rounds until the preset training target is met, and the intelligent agent to be trained is output as the target intelligent agent, including: Based on the selection probability corresponding to each reference agent in the preset reference agent set, randomly select a reference agent from each reference agent; Based on the interaction process between the intelligent agent to be trained and the extracted reference intelligent agent in the sample virtual interaction scene, performing multiple rounds of iterative training on the intelligent agent to be trained; If the subject to be trained does not meet the training objective when the sample interaction results between the subject to be trained and the extracted reference subject are obtained, then a reference subject is re-extracted from each of the reference subjects, and the subject to be trained is continuously trained for multiple rounds of iterations; If the intelligent agent to be trained meets the training objective, the intelligent agent to be trained is output as the target intelligent agent.
5. The method according to claim 4, characterized in that Before outputting the to-be-trained intelligent agent as the target intelligent agent, the method further includes: Counting the number of times the iterative training is performed on the intelligent agent to be trained; If the counted number of training times reaches a preset specified number of times, the agent to be trained is output as a reference agent and added to the reference agent set; The training times are reset to zero, the iterative training of the agent to be trained is continued, and the reference agent set is updated based on the re-counted training times.
6. The method according to claim 2, characterized in that Based on a first sample interactive state feature corresponding to a first sample interactive scene image in the sample virtual interactive scene, predicting a sample scheduling operation performed by the to-be-trained agent on a sample virtual controlled element associated with the to-be-trained agent in the sample virtual interactive scene, and predicting a sample interactive operation performed by the to-be-trained agent on the sample virtual controlled element after performing the sample scheduling operation, further comprising: Performing region recognition processing on the first sample interactive scene image to obtain a first interactive result region, a first global viewing area, and a first local viewing area; Performing image feature extraction processing on the first interaction result area, the first global perspective area, and the first local perspective area respectively, and obtaining corresponding first feature vectors, first global perspective feature matrices, and first local perspective feature matrices respectively, wherein the first feature vector is used to characterize interaction information related to the sample interaction result, the first global perspective feature matrix is used to characterize position information of the sample virtual controlled element, position information of the reference virtual controlled element associated with the reference agent, and position information of the scene elements included in the sample virtual interaction scene, and the first local perspective feature matrix is used to characterize position information of the sample virtual controlled element included in the first local perspective area, position information of the reference virtual controlled element included in the first local perspective area, and position information of the scene elements included in the first local perspective area; The first feature vector, the first global perspective feature matrix and the first local perspective feature matrix are used as first sample interaction state features corresponding to the first sample interaction scene image.
7. The method according to claim 6, characterized in that The method predicts a sample scheduling operation performed by the to-be-trained agent on a sample virtual controlled element associated with the to-be-trained agent in the sample virtual interactive scene based on a first sample interactive state feature corresponding to a first sample interactive scene image in the sample virtual interactive scene, and predicts a sample interactive operation performed by the to-be-trained agent on the sample virtual controlled element after the to-be-trained agent performs the sample scheduling operation, including: Based on the first feature vector and the first global perspective feature matrix, predicting a sample scheduling operation performed by the to-be-trained agent on the sample virtual controlled element; Based on the sample scheduling operation, the first global perspective feature matrix and the first local perspective feature matrix, predict the predicted feature vector, the predicted global perspective feature matrix and the predicted local perspective feature matrix corresponding to the predicted interactive scene image generated after executing the sample scheduling operation; Based on the predicted feature vector and the predicted local view feature matrix, a sample interactive operation performed by the to-be-trained agent on the sample virtual controlled element is predicted.
8. The method according to claim 7, characterized in that Predicting a sample scheduling operation performed by the to-be-trained agent on the sample virtual controlled element based on the first feature vector and the first global view feature matrix includes: Dividing the first global viewing area into a plurality of sub-areas; Predicting a target sub-region corresponding to the sample virtual controlled element based on the first feature vector and the first global viewing angle feature matrix; The sample scheduling operation is obtained based on the sub-region where the sample virtual controlled element is currently located and the target sub-region corresponding to the sample virtual controlled element.
9. The method according to claim 8, characterized in that After obtaining the sample scheduling operation based on the sub-region where the sample virtual controlled element is currently located and the target sub-region corresponding to the sample virtual controlled element, the method further includes: Based on the sample scheduling operation, control the sample virtual controlled element to move from the current sub-region to the corresponding target sub-region, and start timing; Based on a preset moving speed for the sample virtual controlled element, determining a reference time length for the sample virtual controlled element to move from a current sub-region to a corresponding target sub-region; If the timing duration reaches the reference duration, obtaining the second sample interactive scene image; The second sample interaction state feature corresponding to the second sample interaction scene image is extracted to obtain a second feature vector, a second global perspective feature matrix and a second local perspective feature matrix.
10. The method according to claim 9, characterized in that After extracting the second sample interaction state feature corresponding to the second sample interaction scene image, the method further includes: Determine a first scheduling sub-stimulus of the sample scheduling operation based on whether the sub-region where the sample virtual controlled element is currently located matches the target sub-region corresponding to the sample virtual controlled element indicated by the sample scheduling operation; Determining a second scheduling sub-stimulus of the sample scheduling operation based on a change value between the first eigenvector and the second eigenvector corresponding to the sample virtual controlled element; Based on a weighted sum of the first scheduling sub-stimulus and the second scheduling sub-stimulus, scheduling stimulus data for the sample scheduling operation is determined.
11. The method according to any one of claims 2 to 10, characterized in that: The intelligent agent to be trained includes a quantitative information extraction module, wherein the quantitative information extraction module is used to extract sample interaction state features corresponding to each sample interaction scene image; The intelligent agent to be trained also includes a training module, wherein the training module is used to obtain each scheduling incentive data and each interaction incentive data based on each sample interaction state feature, and adjust the model parameters of the intelligent agent to be trained based on the obtained each scheduling incentive data and each interaction incentive data.
12. The method according to claim 11, characterized in that Determining the error values between the scheduling incentive data and the interaction incentive data and the preset target incentive data respectively, and adjusting the model parameters of the to-be-trained agent based on the obtained error values, including: If the training module includes a scheduling model and an interactive model, and the target incentive data includes scheduling target incentive data and interactive target incentive data, determining a scheduling error value between the scheduling incentive data and the scheduling target incentive data, and adjusting a model parameter of the scheduling model based on the obtained scheduling error value; An interaction error value between the interaction incentive data and the interaction target incentive data is determined, and a model parameter of the interaction model is adjusted based on the obtained interaction error value.
13. An intelligent agent interactive device, characterized in that: include: A loading module is used to load a target agent in response to an interaction request instruction triggered by a virtual account; wherein the target agent is obtained by performing multiple rounds of iterative training on the agent to be trained; each round of training is performed by adjusting the model parameters of the agent to be trained based on the scheduling incentive data and interaction incentive data determined in this round of training; the scheduling incentive data represents: the degree of completion of the sample scheduling operation predicted in this round of training, and the degree of influence of the sample scheduling operation on the sample interaction result generated in this round of training; the interaction incentive data represents: the degree of influence of the sample interaction operation predicted in this round of training on the sample interaction result; A processing module: configured to obtain a target interactive scene image corresponding to the control operation triggered by the virtual account on a first target virtual controlled element associated with the virtual account in a target virtual interactive scene; The processing module is also used to: extract target interaction state features from the target interaction scene image; The processing module is also used to: determine the target scheduling operation and the target interaction operation corresponding to the target agent based on the target interaction state characteristics; The processing module is also used to: in response to the target scheduling operation and the target interaction operation, control a second target virtual controlled element associated with the target agent in the target virtual interaction scene.
14. A computer device, characterized in that: include: A memory for storing program instructions; A processor, configured to call the program instructions stored in the memory, and execute the method according to any one of claims 1 to 12 according to the obtained program instructions.
15. A computer-readable storage medium, characterized in that: The storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Intelligent body sense and touch sense interaction scene simulating method based on VR technology
CN107670272A
Training learning method and system for realizing individuation of learnable ability model of intelligent virtual digital animal
CN110866588A