Model training method and related device

By providing the state and perception information of the target object in model training, and combining the behavioral task achievement adjustment parameters, the problems of high training difficulty and low simulation quality in the existing technology are solved, and efficient and anthropomorphic behavior control is achieved.

CN120285569APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410046779.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the model training method relies on player data, which leads to high training difficulty and poor generalization, and the generated instruction generation model simulation quality is low, making it difficult to effectively control game objects in new scenarios.

Method used

By providing the state information and perception information of the target object, generating behavior control instructions, and adjusting model parameters in combination with the achievement degree of behavioral tasks, it realizes efficient training without player data.

Benefits of technology

It improves the authenticity and rationality of behavioral control instructions, reduces the difficulty of training, and enhances the applicability of the model in new scenarios and anthropomorphic interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120285569A_ABST
    Figure CN120285569A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training method and a related device, when an instruction generation model used for generating a behavior control instruction for a target object is trained, target perception information used for simulating information actually perceived by a program user through controlling the target object can be provided for the model. Therefore, the model can generate the behavior control instruction based on the perception information similar to the program user, so that the model can know the logic of various behavior control instructions made by the program user, and the behavior control of the target object through the behavior control instruction is more personified; and better interaction experience can be brought during interaction with the target object. Meanwhile, the reasonable degree of the behavior control instruction output by the model can be measured through the achievement degree of the behavior task, so that the model parameters can be effectively adjusted, the behavior control instruction actually made by a program user does not need to be collected, and the model training difficulty is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and in particular to a model training method and related devices. Background Art

[0002] In order to meet diverse interaction needs, in addition to the ability for players to interact with each other, game programs usually also provide functions for players to interact with game objects controlled by the game program itself (commonly known as human-computer interaction). For example, the game program may provide a human-computer game function, in which players can control their own game objects in the game and play against game objects automatically controlled by the game program.

[0003] In the related art, in order to make the game program have better automatic control functions, an instruction generation model for generating action instructions based on the state of the game object can be obtained through model training. During the game running process, the game program can collect the object state of the automatically controlled game object in real time, input it into the instruction generation model to obtain action control instructions, and realize automatic control of the game object based on the action control instructions.

[0004] However, the model training method in the related art requires a large amount of data on actual action control instructions made by players during the game for various states of the controlled game objects, so as to generate sample pairs between object states and action control instructions for model learning. On the one hand, this training method is overly dependent on player data. While data collection is difficult, it is difficult to achieve efficient training in new game scenarios, and it is also difficult to effectively control object states that have never appeared. On the other hand, the instruction generation model trained in this way can only barely simulate the operations of real players, and the simulation quality is poor, which gives players a low sense of reality. Summary of the invention

[0005] In order to solve the above technical problems, the present application provides a model training method, which, when generating behavior control instructions for a target object through a model, can provide target perception information of the information actually perceived by the model program user by controlling the target object, so that the model can generate behavior control instructions based on perception information similar to that of the program user, making the behavior control of the target object through behavior control instructions more humanized and more logical, thereby bringing a better interactive experience when interacting with the target object.

[0006] The embodiments of the present application disclose the following technical solutions:

[0007] In a first aspect, an embodiment of the present application discloses a model training method, the method comprising:

[0008] Determine the target object state information and target perception information corresponding to the target object at the target moment. The target object state information is used to characterize the object state of the target object in the target program, and the perception information is used to simulate the information that can be perceived by the user of the target program by controlling the target object at the target moment;

[0009] Generate a target behavior control instruction corresponding to the target moment through an initial instruction generation model according to the target object state information and the target perception information. The target behavior control instruction is used to control the target object to perform a target object behavior at the target moment;

[0010] Adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task by the target object performing the target object behavior, and obtain an instruction generation model. The achievement degree determined by the instruction generation model reaches the target achievement degree. The achievement degree is used to characterize the rational degree of the object behavior of the target object, and the rational degree characterized by the target achievement degree is greater than the first preset threshold. The instruction generation model is used to determine a behavior control instruction for controlling the target object according to the object state information and perception information corresponding to the target object.

[0011] In a possible implementation manner, the target program is a shooting game program, the first initial object behavior is a shooting behavior, and the target behavior deviation includes a horizontal shooting direction deviation and / or a vertical shooting direction deviation.

[0012] In a possible implementation manner, the generating a target behavior control instruction corresponding to the target moment according to the target object state information and the target perception information includes:

[0013] Generate a second initial behavior control instruction corresponding to the target moment according to the target object state information and the target perception information. The second initial behavior control instruction is used to control the target object to perform the target object behavior at the target moment with a first behavior amplitude;

[0014] Based on that the first behavior amplitude does not exceed the behavior amplitude threshold corresponding to the target object behavior, determine the second initial behavior control instruction as the target behavior control instruction. The behavior amplitude threshold is the maximum amplitude for controlling the target object to perform the target object behavior through the target program interface, and the target program interface is used to control the target object;

[0015] Based on the amplitude of the first behavior exceeding the behavior amplitude threshold, adjust the second initial behavior control instruction to obtain the target behavior control instruction, where the target behavior control instruction is used to control the target object to execute the target object behavior at the target moment with the behavior amplitude threshold.

[0016] In a possible implementation, the target object behavior is an object behavior for a target interaction object in the target program. Adjusting the model parameters of the initial instruction generation model according to the achievement degree of the behavior task by the target object through executing the target object behavior to obtain an instruction generation model includes:

[0017] Based on the time interval between the target moment and the object appearance moment reaching the reaction duration, control the target object to execute the target object behavior through the target behavior control instruction. The object appearance moment is the moment when the target interaction object appears in the program interface corresponding to the target program. The reaction duration is positively correlated with the perception difficulty corresponding to the target interaction object, and the perception difficulty is the difficulty of perceiving the target interaction object through the program interface at the object appearance moment.

[0018] Adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task by the target object through executing the target object behavior to obtain an instruction generation model.

[0019] In a possible implementation, the perception difficulty is positively correlated with the distance between the target interaction object and the target object at the object appearance moment, and is positively correlated with the position difference corresponding to the display position of the target interaction object in the program interface at the object appearance moment. The position difference is the difference between the display position and the center position of the program interface.

[0020] In a possible implementation, the target program is a shooting game program, and the target object behavior is to aim at the target interaction object.

[0021] In a second aspect, an embodiment of the present application discloses a model training device, and the device includes a determination unit, a generation unit, and an adjustment unit:

[0022] The determination unit is configured to determine the target object state information and target perception information corresponding to the target object at the target moment. The target object state information is used to characterize the object state corresponding to the target object in the target program, and the perception information is used to simulate the information that can be perceived by the user of the target program by controlling the target object at the target moment.

[0023] The generating unit is configured to generate a target behavior control instruction corresponding to the target moment through an initial instruction generation model according to the target object state information and the target perception information, where the target behavior control instruction is used to control the target object to perform a target object behavior at the target moment;

[0024] The adjusting unit is configured to adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task by the target object performing the target object behavior, so as to obtain an instruction generation model. The achievement degree determined by the instruction generation model reaches a target achievement degree. The achievement degree is used to characterize the rational degree of the object behavior of the target object, and the rational degree characterized by the target achievement degree is greater than a first preset threshold. The instruction generation model is used to determine a behavior control instruction for controlling the target object according to the object state information and perception information corresponding to the target object.

[0025] In a possible implementation manner, the target perception information includes scene perception information, and the scene perception information is used to characterize the positions of scene objects that can be perceived by the program user at the target moment. The scene objects are used to form the target scene where the target object is located at the target moment.

[0026] In a possible implementation manner, the scene perception information is determined in the following way:

[0027] Determine the ray starting point of the detection ray according to the position of the target object at the target moment;

[0028] Project the detection ray from the ray starting point to the target scene where the target object is located, and determine the object point that the detection ray first passes through. The object point is used to form the scene object;

[0029] Determine the position information corresponding to the object point as the scene perception information.

[0030] In a possible implementation manner, the detection ray includes a horizontal detection ray. Projecting the detection ray from the ray starting point to the target scene where the target object is located and determining the object point that the detection ray first passes through includes:

[0031] Project the horizontal detection ray horizontally from the ray starting point to the target scene, and determine the object point that the horizontal detection ray first passes through.

[0032] In a possible implementation manner, the horizontal detection ray includes a plane detection ray. Determining the ray starting point of the detection ray according to the position of the target object at the target moment includes:

[0033] Based on the position of the target object at the target moment, determine the camera plane corresponding to the target object, where the camera plane is used to render the target program interface corresponding to the target moment in the target program, the target program interface is used to control the target object in the target program, and the pixel values of the pixel points included in the target program interface are determined based on the projection of the scene object onto the camera plane;

[0034] Determine the starting point of the ray located on the camera plane;

[0035] Projecting the horizontal detection ray horizontally from the starting point of the ray towards the target scene to determine the object point where the horizontal detection ray first passes through includes:

[0036] Project the plane detection ray from the starting point of the ray towards the target scene in a first direction perpendicular to the camera plane to determine the object point where the plane detection ray first passes through, where the first direction is the direction of observing the target scene through the target program interface.

[0037] In a possible implementation, the horizontal detection ray includes a surrounding detection ray. Determining the starting point of the detection ray according to the position of the target object at the target moment includes:

[0038] According to the position of the target object at the target moment, determine the starting point of the ray located on the target object, where the starting point of the ray is the object point on the target object where the probability of interacting with the scene object is greater than a second preset threshold;

[0039] Projecting the horizontal detection ray horizontally from the starting point of the ray towards the target scene to determine the object point where the horizontal detection ray first passes through includes:

[0040] Project multiple surrounding detection rays horizontally from the starting point of the ray towards the target scene to determine the object points where the multiple surrounding detection rays first pass through respectively, where the ray directions corresponding to the multiple surrounding detection rays are different.

[0041] In a possible implementation, the detection ray includes a vertical detection ray. Determining the starting point of the detection ray according to the position of the target object at the target moment includes:

[0042] According to the position of the target object at the target moment, determine the detection plane located above the target object, where the detection plane is a horizontal plane;

[0043] Determine the starting point of the ray located on the detection plane;

[0044] Projecting the detection ray from the starting point of the ray towards the target scene where the target object is located, and determining the object point where the detection ray first passes through, includes:

[0045] Projecting the vertical detection ray vertically from the starting point of the ray towards the target scene, and determining the object point where the vertical detection ray first passes through.

[0046] In a possible implementation manner, the target perception information includes interaction object perception information, which is used to represent the position of the interaction object that can be perceived by the program user at the target moment. The interaction object is used to interact with the target object in the target scene by performing object behaviors. The target scene is the scene where the target object is located at the target moment, and the interaction object is not a scene object that constitutes the target scene.

[0047] In a possible implementation manner, the interaction object perception information is determined by the following method:

[0048] Determining whether there is such a scene object between the target object and the interaction object;

[0049] Based on the fact that there is no such scene object between the target object and the interaction object, determining the position information corresponding to the interaction object as the interaction object perception information, where the position information is used to identify the position of the interaction object in the target scene;

[0050] Based on the fact that there is such a scene object between the target object and the interaction object, and the interaction object meets the perceivable condition, determining the area information corresponding to the interaction object as the interaction object perception information, where the area information is used to identify the position area of the interaction object in the target scene. The position area includes multiple positions, and the multiple positions include the position of the interaction object in the target scene.

[0051] In a possible implementation manner, the perceivable condition includes that the interaction object has sound effect information in a playing state at the target moment, and the target distance between the interaction object and the target object is less than a distance threshold;

[0052] The number of positions included in the position area is inversely correlated with the volume of the sound effect information and inversely correlated with the target distance.

[0053] In a possible implementation, the perceivable condition includes that in a first historical period, there is a first historical moment satisfying that at the first historical moment, there is no such scenario object between the interaction object and the target object. The end moment of the first historical period is the target moment, and the period length of the first historical period is less than a first length threshold;

[0054] The number of positions included in the position area is positively correlated with the period length of a first target period, and the first target period is the period from the first historical moment to the target moment.

[0055] In a possible implementation, the perceivable condition includes that in a second historical period, there is a second historical moment satisfying that at the second historical moment, the perception information corresponding to the perception sharing object includes the target interaction object perception information corresponding to the interaction object, and the perception sharing object is an object in the target program that shares perception information with the target object;

[0056] The number of positions included in the position area is positively correlated with the period length of a second target period and is positively correlated with the number of positions identified by the target interaction object perception information, and the second target period is the period from the second historical moment to the target moment.

[0057] In a possible implementation, the target perception information further includes scenario path information, and the scenario path information is used to identify the actionable path corresponding to the target object in the target scenario, and the target scenario is the scenario where the target object is located at the target moment.

[0058] In a possible implementation, the scenario path information is determined by the following method:

[0059] Determine multiple reachable positions corresponding to the target scenario, and the multiple reachable positions are positions that the target scenario supports the object to reach;

[0060] Generate initial path information corresponding to the target scenario according to the multiple reachable positions, and the initial path information is used to identify multiple initial paths between the multiple reachable positions;

[0061] Determine multiple key positions among the multiple reachable positions. When the program user controls the target object to move at the target moment, the probability of reaching the key positions is greater than the probability of reaching the reachable positions other than the key positions among the reachable positions;

[0062] Determine the scenario path information corresponding to the target scenario according to the initial path information and the multiple key positions, and the actionable path is the initial path among the multiple initial paths that includes the key positions.

[0063] In a possible implementation, the generating unit is specifically configured to:

[0064] Generate a plurality of pending behavior control instructions according to the target object state information and the target perception information;

[0065] Determine a plurality of sub-behavior control instructions corresponding to the target pending behavior control instruction, where the plurality of sub-behavior control instructions have corresponding sub-object behaviors and control behaviors respectively. The sub-object behaviors corresponding to the plurality of sub-behavior control instructions are used to constitute the object behavior corresponding to the target pending behavior control instruction. The target sub-behavior control instruction corresponds to a target control behavior, and the target control behavior executed through the target program interface is used to generate the target sub-behavior control instruction. The target sub-behavior control instruction is any one of the plurality of sub-behavior control instructions, the target pending behavior control is any one of the plurality of pending behavior control instructions, and the target program interface is used to control the target object;

[0066] Based on the control behaviors corresponding to the plurality of sub-behavior control instructions being a plurality of control behaviors that can be executed in parallel through the target program interface, determine the target pending behavior control instruction as the target behavior control instruction.

[0067] In a possible implementation, the generating unit is specifically configured to:

[0068] Generate a first initial behavior control instruction corresponding to the target time according to the target object state information and the target perception information, where the first initial behavior control instruction is used to control the target object to perform a first initial object behavior at the target time;

[0069] Generate the target object control instruction according to the behavior deviation parameter and the first initial behavior control instruction, where the behavior deviation parameter corresponds to the target behavior deviation, and the target object behavior is the first initial object behavior that generates the target behavior deviation when executed.

[0070] In a possible implementation, the target program is a shooting game program, the first initial object behavior is a shooting behavior, and the target behavior deviation includes a horizontal shooting direction deviation and / or a vertical shooting direction deviation.

[0071] In a possible implementation, the generating unit is specifically configured to:

[0072] Generate a second initial behavior control instruction corresponding to the target time according to the target object state information and the target perception information, where the second initial behavior control instruction is used to control the target object to perform the target object behavior at the target time with a first behavior amplitude;

[0073] Based on the fact that the amplitude of the first behavior does not exceed the behavior amplitude threshold corresponding to the target object behavior, determine the second initial behavior control instruction as the target behavior control instruction, where the behavior amplitude threshold is the maximum amplitude for controlling the target object to perform the target object behavior through the target program interface, and the target program interface is used to control the target object;

[0074] Based on the fact that the amplitude of the first behavior exceeds the behavior amplitude threshold, adjust the second initial behavior control instruction to obtain the target behavior control instruction, where the target behavior control instruction is used to control the target object to perform the target object behavior at the target moment with the behavior amplitude threshold.

[0075] In a possible implementation, the target object behavior is an object behavior for a target interaction object in the target program, and the adjustment unit is specifically configured to:

[0076] Based on the fact that the time interval between the target moment and the object appearance moment reaches the reaction duration, control the target object to perform the target object behavior through the target behavior control instruction, where the object appearance moment is the moment when the target interaction object appears in the program interface corresponding to the target program, and the reaction duration is positively correlated with the perception difficulty corresponding to the target interaction object, and the perception difficulty is the difficulty of perceiving the target interaction object at the object appearance moment through the program interface;

[0077] According to the achievement degree of the behavior task by the target object through performing the target object behavior, adjust the model parameters of the initial instruction generation model to obtain the instruction generation model.

[0078] In a possible implementation, the perception difficulty is positively correlated with the distance between the target interaction object and the target object at the object appearance moment, and is positively correlated with the position difference corresponding to the display position of the target interaction object in the program interface at the object appearance moment, where the position difference is the difference between the display position and the center position of the program interface.

[0079] In a possible implementation, the target program is a shooting game program, and the target object behavior is to aim at the target interaction object.

[0080] In a third aspect, an embodiment of the present application discloses a computer device, which includes a processor and a memory:

[0081] The memory is used to store a computer program and transmit the computer program to the processor;

[0082] The processor is configured to execute the model training method according to any one of the first aspects based on the instructions in the computer program;

[0083] In a fourth aspect, an embodiment of the present application discloses a computer-readable storage medium for storing a computer program for executing the model training method according to any one of the first aspects;

[0084] In a fifth aspect, an embodiment of the present application discloses a computer program product including a computer program, which, when running on a computer device, causes the computer device to execute the model training method according to any one of the first aspects.

[0085] It can be seen from the above technical solutions that when training an instruction generation model for automatically controlling a target object to execute various object behaviors, target object state information for characterizing the object state of the target object at the target moment and target perception information for simulating the information that can be perceived by the program user through controlling the target object at the target moment can be provided to the initial instruction generation model, so that the initial instruction generation model can generate a target behavior control instruction based on information similar to the information referred to by the program user when making a behavior control decision. Furthermore, during the generation process of the behavior control instruction, the object behavior control logic of the program user can be learned more effectively, which helps to improve the authenticity and rationality of the determined target behavior control instruction, and makes the target object behavior controlled by the target behavior control instruction more conform to the control logic of the program user. During the training process, a behavior task can be set for the target object, and the degree of achievement of the behavior task by the target object through executing the target object behavior can characterize the rationality of the object behavior. Therefore, the model parameters of the initial instruction generation model can be adjusted based on this degree of achievement to obtain an instruction generation model, so that the degree of achievement determined by the instruction generation model can reach a target degree of achievement with a relatively large characterized rationality. During the model parameter adjustment process, the model can learn how to determine a more reasonable behavior control instruction. Thus, on the one hand, the present application improves the authenticity and rationality of the behavior control instruction in the data dimension input to the model, and on the other hand, supervises the rationality of the behavior control instruction in the output dimension of the behavior control instruction and feeds back the model parameters based on the rationality, enabling it to learn how to reasonably generate behavior control instructions. At the same time, during the above training process, it is not necessary to collect the actual behavior control data of the program user controlling the target object, so the training difficulty is low, and effective training can also be achieved for novel programs lacking usage data, improving the scenario versatility of object automatic control. Description of the Drawings

[0086] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0087] Figure 1 Schematic diagram of a model training method in an actual application scenario provided by an embodiment of the present application;

[0088] Figure 2 Flowchart of a model training method provided by an embodiment of the present application;

[0089] Figure 3 Schematic diagram of a model training method provided by an embodiment of the present application;

[0090] Figure 4 Schematic diagram of a model training method provided by an embodiment of the present application;

[0091] Figure 5 Schematic diagram of a model training method provided by an embodiment of the present application;

[0092] Figure 6 Schematic diagram of a model training method provided by an embodiment of the present application;

[0093] Figure 7 Schematic diagram of a model training method provided by an embodiment of the present application;

[0094] Figure 8 Schematic diagram of a model training method provided by an embodiment of the present application;

[0095] Figure 9 Schematic diagram of a model training method provided by an embodiment of the present application;

[0096] Figure 10 Schematic diagram of a model training method provided by an embodiment of the present application;

[0097] Figure 11 Schematic diagram of a model training method provided by an embodiment of the present application;

[0098] Figure 12 Schematic diagram of a model training method provided by an embodiment of the present application;

[0099] Figure 13 Schematic diagram of a model training method provided by an embodiment of the present application;

[0100] Figure 14Flowchart of a model training method in an actual application scenario provided by an embodiment of the present application;

[0101] Figure 15 Structural block diagram of a model training device provided by an embodiment of the present application;

[0102] Figure 16 Structural diagram of a terminal provided by an embodiment of the present application;

[0103] Figure 17 Structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0104] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0105] In order to achieve automatic control of the object behavior of an object in a program, an instruction generation model can be obtained through model training. The instruction generation model can automatically generate behavior control instructions for controlling the object to execute the object behavior.

[0106] In the related art, when training the instruction generation model, the model input is the object state information corresponding to the object in the program. For example, in a game program, the object state information such as the health value, movement speed, and movement direction corresponding to the object can be input. The instruction generation model will output the corresponding behavior control instructions. The object behavior corresponding to the behavior control instructions is the object behavior that the instruction generation model believes the object is most likely to execute in the object state identified by the object state information. Relevant personnel can collect the behavior control instructions actually made by the program user who controls the object in each object state. Thus, through the difference between the behavior control instructions in these two dimensions, the accuracy of the instruction generation model in generating behavior control instructions can be characterized. Furthermore, the model parameters can be adjusted based on this difference, so that the behavior control instructions determined by the instruction generation model are close to the behavior control instructions actually made by the program user.

[0107] However, in the model training process in the related art, only the behavior control instructions determined by the model are forced to fit the actual behavior control instructions made by the program user. The model cannot know the basis for the program user to make these behavior control instructions, so it cannot learn the logic of the program user to make these behavior control instructions, and can only roughly learn the mapping relationship between the object state and the behavior control instructions. On the one hand, the model cannot perform sufficient behavioral logic analysis on the object states that have appeared during the training process, resulting in poor authenticity of the determined behavior control instructions, and it is difficult for the object controlled by the behavior control instructions to bring a high-quality object interaction experience. On the other hand, this method requires a large amount of actual behavior control instructions of the program user in various object states to be collected. The workload of the preliminary preparation work is large, the model training is difficult, and it is difficult to determine accurate and effective behavior control instructions for object states that have not appeared before.

[0108] To solve the above technical problems, the present application provides a model training method. When training an instruction generation model for generating behavior control instructions for a target object, on the basis of providing the model with target object state information for characterizing the target object, it is also possible to provide the model with target perception information for simulating the information actually perceived by the program user through controlling the target object, so that the model can generate behavior control instructions based on perception information similar to that of the program user, enabling the model to understand the logic of the program user to make various behavior control instructions, and further making the behavior control of the target object by the behavior control instructions more anthropomorphic and more logical, and being able to bring a better interaction experience when interacting with the target object. At the same time, the present application can measure the rationality of the behavior control instructions output by the model through the achievement degree of the behavior task, so as to effectively adjust the model parameters, without collecting the actual behavior control instructions made by the program user, and the model training is less difficult.

[0109] It can be understood that this method can be applied to a computer device, which is a computer device capable of model training, such as a terminal device or a server. This method can be independently executed by a terminal device or a server, or can be applied to a network scenario where a terminal device and a server communicate, and is executed in cooperation with the terminal device and the server. Among them, the terminal device can be a device such as a mobile phone, a tablet computer, a laptop computer, or a desktop computer. The terminal device can also include a variety of virtual reality devices, such as an augmented reality (AR) device, such as an AR glasses, an AR screen, etc., and can also include a virtual reality (VR) device, such as a head-mounted VR glasses, etc. The server can be understood as an application server or a web server. In actual deployment, the server can be an independent server, a cluster server, or a cloud server, etc.

[0110] To facilitate understanding of the technical solution provided by this application, next, a model training method provided by this application will be introduced in combination with an actual application scenario.

[0111] See Figure 1 , Figure 1 is a schematic diagram of a model training method in an actual application scenario provided by an embodiment of this application. In this actual application scenario, the computer device is a server 101 with a model training function, the target program is a shooting game program, the target object is the game character controlled by the player, and the program user is the player.

[0112] As Figure 1 shown, the server 101 can obtain the target object state information and the target perception information corresponding to the target object at the target moment, where the target object state information is used to characterize the object state corresponding to the target object at the target moment, such as the health value, movement speed, skill release times, etc. of the game character. The target perception information is used to simulate the information that the program user can perceive by controlling the target object at the target moment. For example, through Figure 1 the game interface corresponding to the target moment shown in, the player can perceive information such as buildings, corridors, obstacles (such as railings), and whether there are hostile characters in the game scene. The perceived information of the object state is two main bases for the program user to make behavior control instructions.

[0113] Server 101 can input the target object status information and target perception information into the initial instruction generation model, enabling the initial instruction generation model to simulate the logic of a program user making instructions to generate target behavior control instructions. This can improve the rationality and logic of the model in generating behavior control instructions at the input end, making the target behavior control instructions more in line with the actual behavior control instructions made by the program user at the target moment. At the output end, in order to enable the model to learn how to improve the rationality of behavior control instructions, a behavior task can be set. The target behavior control instruction can control the target object to execute the target object behavior, and the degree of achievement of the behavior task by the target object through executing the target object behavior can represent the rationality of the target object executing the target object behavior. For example, the behavior task can be to defeat a target number of enemy objects, and the degree of achievement is the number of enemy objects actually defeated by the game object through executing the target object behavior. The higher the degree of achievement, the higher the rationality of the object behavior executed by the target object.

[0114] Based on this, Server 101 can adjust the model parameters of the initial instruction generation model according to this degree of achievement to obtain an instruction generation model. When the target behavior control instruction output by this instruction generation model performs the behavior task, it can reach the target degree of achievement, and this target degree of achievement corresponds to a relatively high degree of rationality. Furthermore, this parameter adjustment process can enable the initial instruction generation model to learn how to determine a behavior control instruction with a higher degree of rationality, that is, this instruction generation model can have the ability to accurately generate a behavior control instruction with a relatively high degree of rationality based on the object status information and perception information corresponding to the target object. Thus, the behavior control instruction determined based on the model can reasonably control the target object, enabling the target object to execute object behaviors with a relatively high degree of rationality, strong authenticity, and in line with the control logic of the program user. For example, in Figure 1 the shooting game shown, it can automatically control the game character to reasonably execute object behaviors such as shooting, moving, and turning.

[0115] It can be seen that this application can improve the rationality and authenticity of the behavior control instructions determined by the model from two dimensions of input information and output supervision. Without collecting data of the program user, it realizes the effective training of the instruction generation model, thereby improving the effectiveness and versatility of the model while reducing the model training difficulty, enabling the model to have a clearer understanding of the execution logic of object behaviors, and being able to make more reasonable responses to object status and perception information that have not appeared before.

[0116] Next, the model training method provided by this application will be introduced in detail in conjunction with the accompanying drawings.

[0117] See Figure 2 , Figure 2The flowchart of a model training method provided by an embodiment of this application. In this embodiment, the computer device can be any of the above computer devices with model training functions. The method includes:

[0118] S201: Determine the target object state information and target perception information corresponding to the target object at the target time.

[0119] Among them, the target object can be any object in the target program that can control the behavior of the execution object. The target program can be any program including the target object. For example, the target program can be a game program, and the target object is a controllable game character in the game program, etc. The program user is the party that controls the object through the target program. For example, it can be the player corresponding to the game program, etc., and is not the target program itself.

[0120] First, this application first carefully analyzes the control logic of the program user to control the object to execute the object behavior when using the program. Usually, the program user controls the object based on the following two dimensions of information: object state information and perceivable information. The object state information is used to represent the object state of the controlled object. The object state refers to the state of the object itself. For example, in a game program, the object state information can include the health value information, movement speed information, facing information, etc. of the game character. When the health value information is low, the player tends to control the game character to perform object behaviors such as defense and escape. When the health value information is high, the player tends to control the game character to perform offensive object behaviors; the perceivable information is the information perceived by the program user by controlling the target object. For example, the program scene where the target object is located, whether there are other objects near the target object, etc. When there are hostile game characters around the game character and the health value information of the controlled character is high, the player tends to control the game character to attack the hostile game character; when there are hostile game characters around the game character and the health value information of the controlled character is low, the player tends to control the game character to hide.

[0121] Based on this, in the present application, in order to enable the instruction generation model to determine a more reasonable behavior control instruction that conforms to the behavior control logic of the program user, the information in the above two dimensions can be used as the basis for the model to determine the behavior control instruction. During the model training process, taking the target object as an example, the computer device can determine the target object state information and target perception information corresponding to the target object at the target moment. The target object state information is used to represent the object state of the target object in the target program, that is, the object state of the target object at the target moment. The perception information is used to simulate the information that can be perceived by the user of the target program by controlling the target object at the target moment. Thus, it can provide the information basis for the initial instruction generation model to control the target object to execute the object behavior at the target moment, enabling the initial instruction generation model to simulate the control logic of the program user to generate the behavior control instruction, improving the rationality and logic of the model in analyzing the behavior control instruction, and further enabling the determined behavior control instruction to be more in line with the behavior control logic of the program user.

[0122] S202: Through the initial instruction generation model, according to the target object state information and the target perception information, generate the target behavior control instruction corresponding to the target moment.

[0123] The instruction generation model in the present application can be any model that supports the technical solution of the present application. The behavior control instruction is a control instruction for controlling the object to execute the object behavior. Through the target object state information and the target perception information, the initial instruction generation model can start from the state of the target object itself and the information that can be perceived through the target object, analyze which object behavior is more reasonable for the target object to execute and is more in line with the behavior control logic of the program user, and obtain the target behavior control instruction. Among them, the target behavior control instruction is used to control the target object to execute the target object behavior at the target moment.

[0124] S203: According to the degree of achievement of the behavior task by the target object through executing the target object behavior, adjust the model parameters of the initial instruction generation model to obtain the instruction generation model.

[0125] In this application, in order to get rid of the dependence of the model training process on the usage data of the program user, the computer device can set behavior tasks based on the requirements of model training. The degree of achievement of the behavior task by the target object through performing object behaviors can be used to measure the reasonableness of the object behaviors. Generally, the higher the reasonableness of the object behaviors, the higher the efficiency of performing the behavior task, and the higher the corresponding degree of achievement. For example, for a battle game program, the behavior task can be to defeat 100 game characters. The more game characters the target object defeats through performing object behaviors, the higher the degree of achievement of the behavior task, indicating that the object behaviors of the target object are more reasonable. The behavior task can be determined only based on the object behavior requirements of the target object in the target program, without collecting the behavior control instructions generated by the program user during actual use. Therefore, the preliminary preparation work required for model training is less, the model training difficulty is lower, and the efficiency is higher.

[0126] Based on this, the computer device can control the target object to perform the target object behaviors through the target behavior control instructions, and determine the degree of achievement of the behavior task by the target object through performing the target object behaviors. The computer device can adjust the model parameters of the initial instruction generation model based on this degree of achievement, so that the degree of achievement determined by the initial instruction generation model gradually approaches the target degree of achievement. This degree of achievement is used to characterize the reasonableness of the object behaviors of the target object, and the reasonableness characterized by the target degree of achievement is greater than the first preset threshold, that is, the target degree of achievement corresponds to a relatively high reasonableness. Thus, the initial instruction generation model can learn how to determine more reasonable behavior control instructions, and obtain the trained instruction generation model. The degree of achievement determined by the instruction generation model can reach the target degree of achievement. The instruction generation model can be used to determine the behavior control instructions for controlling the target object according to the object state information and perception information corresponding to the target object. For example, in a shooting game program, the instruction generation model can be used to generate behavior control instructions according to the object state information such as the health value information and the number of bullets of the game character, and according to the perception information such as whether there are hostile game characters around, the distance and position of the hostile game characters, to control the game character to perform object behaviors such as shooting, aiming, and moving.

[0127] As can be seen from the above technical solution, when training an instruction generation model for automatically controlling a target object to perform various object behaviors, target object state information for characterizing the object state of the target object corresponding to the target moment, and target perception information for simulating the information that can be perceived by the program user through controlling the target object at the target moment can be provided to the initial instruction generation model, so that the initial instruction generation model can generate a target behavior control instruction based on information similar to the information referred to by the program user when making a behavior control decision. Furthermore, during the generation process of the behavior control instruction, the object behavior control logic of the program user can be effectively learned, which helps to improve the authenticity and rationality of the determined target behavior control instruction, and makes the target object behavior controlled based on the target behavior control instruction more conform to the control logic of the program user. On the one hand, this application improves the authenticity and rationality of the behavior control instruction in the data dimension input to the model, and on the other hand, supervises the rationality of the behavior in the dimension of the output behavior control instruction, and feeds back the model parameters based on the rationality, enabling it to learn how to reasonably generate behavior control instructions. At the same time, during the above training process, it is not necessary to collect the actual behavior control data of the program user controlling the target object, so the training difficulty is low, and effective training can also be achieved for novel programs lacking usage data, improving the scenario versatility of object automatic control.

[0128] Next, the specific content of one of the core improvements of this application - perception information - will be introduced in detail.

[0129] It can be understood that the information that can be perceived by the program user in the program mainly includes two dimensions: scene perception information and object perception information. Scene perception information is the information used to perceive the program scene where the object is located, such as perceiving the scene structure, obstacle positions, passage positions, etc. in a game scene. Object perception information is the information used to perceive other objects included in the scene, such as perceiving the positions, distances, etc. of other game characters in a game scene.

[0130] First, the scene perception information will be introduced in detail. In a possible implementation, the target perception information may include scene perception information, and the scene perception information is used to characterize the positions of scene objects that can be perceived by the program user at the target moment, and the scene objects are used to form the target scene where the target object is located at the target moment. For example, in Figure 1 the scene objects can be railings, corridors, etc. in a game scene, and these scene objects can form the game scene where the game character is located. Through the scene perception information, the instruction generation model can analyze how to control the target object to perform reasonable object behaviors in the target scene without conflicting with the target scene, such as how to move to avoid obstacles in the target scene, how to move without colliding with the scene edge (such as a wall) in the target scene, etc.

[0131] Among them, there are various ways to determine the scene perception information, and each determination method will be introduced in detail below.

[0132] It can be understood that the scene perception information simulates the scene perceived by the program user. Since the program user usually perceives the scene through the visual dimension, the computer device can simulate the line of sight of the program user observing the scene to determine the scene perception information. In one possible implementation, the scene perception information can be determined in the following way:

[0133] The computer device can determine the ray origin of the detection ray according to the position of the target object at the target moment. The detection ray is used to simulate the program user's perception of the target scene in the visual dimension. The computer device can project the detection ray from the ray origin to the target scene where the target object is located, and determine the object point that the detection ray first passes through. The object point is used to form the scene object. Since the position of the ray origin changes with the position of the target object, the detection method of the detection ray also changes with the position of the target object. Thus, the detection ray can more accurately and realistically simulate the program user's perception of the target scene in the visual dimension by controlling the target object. This object point is most likely the object point that the program user can perceive at the target moment. Based on this, the computer device can determine the position information corresponding to the object point as the scene perception information, so that the instruction generation model can obtain a scene perception similar to that of the program user at the target moment. That is, when the program user is a user, the scene perception information can simulate the human eye's perception of the scene terrain.

[0134] Specifically, there are various forms of scene perception through the detection ray, and each form of the detection ray will be introduced in detail below.

[0135] First of all, it can be understood that the information that the program user can perceive during the process of controlling the target object is usually determined based on the way the target object observes the target scene. The way the target object observes the target scene is mainly visual observation in the horizontal direction. Therefore, usually, the information that the program user can perceive is mainly visual perception information in the horizontal direction. For example, in Figure 1 , the information that the program user can perceive through the game interface is the scene directly in front of the game character.

[0136] Based on this, in a possible implementation, the computer device can detect the scene in the horizontal direction by detecting rays. In this implementation, the detection rays can include horizontal detection rays. When projecting the detection rays from the ray origin to the target scene where the target object is located and determining the object point where the detection ray first passes, the computer device can project the horizontal detection rays horizontally from the ray origin to the target scene and determine the object point where the horizontal detection ray first passes. These object points are the object points mainly perceived by the program user in the visual dimension.

[0137] Among them, based on different scene perception requirements, there can also be multiple types of horizontal detection rays. Next, two types of horizontal detection rays will be mainly introduced in detail.

[0138] The first type: Plane detection rays

[0139] It can be understood that when using the target program, the program user mainly perceives information through the program interface of the target program. For example, when using a game program, the player observes the game scene and understands the status of the game characters through the game program interface. Based on this, in a possible implementation, the computer device can generate plane detection rays based on the program interface to simulate the perception of the program scene by the program user in the visual dimension through the program interface.

[0140] In this implementation, the horizontal detection rays can include plane detection rays. When determining the ray origin of the detection ray according to the position of the target object at the target moment, the computer device can first locate the plane position corresponding to the program interface. The computer device can determine the camera plane corresponding to the target object according to the position of the target object at the target moment. The camera plane is used to render the target program interface corresponding to the target program at the target moment. Among them, the target program interface is used to control the target object in the target program, and the pixel values of the pixel points included in the target program interface are determined based on the projection of the scene object onto the camera plane. That is, the target program interface is obtained by projecting onto the camera plane. Therefore, the content projected by the target scene onto the camera plane is the scene mainly perceived through the program interface. For example, Figure 3 in, the camera plane is located at the axis position of the target object, and the slope in the target scene can be displayed in the program interface of the target program in the form of projection onto the camera plane. When the position of the target object changes, it will drive the change of the camera plane position, thereby changing the content projected by the target scene onto the camera plane, and further changing the scene content perceived by the program user through the program interface.

[0141] Based on this, in order to simulate the scene information perceived through the target program interface, the computer device can determine the starting point of the ray located on the camera plane. Starting from the starting point of the ray and horizontally projecting a horizontal detection ray towards the target scene, when determining the object point that the horizontal detection ray first passes through, the computer device can project a plane detection ray from the starting point of the ray towards the target scene in a first direction perpendicular to the camera plane, and determine the object point that the plane detection ray first passes through. The first direction is the direction of observing the target scene through the target program interface. This object point is the object point that can be projected onto the camera plane through projection and affects the display effect of the target program interface. Therefore, it is also the object point that the program user can perceive through the target program interface.

[0142] As Figure 4 shown, 5 plane detection rays can be emitted from the camera plane, and 5 object points, namely object point 1 to object point 5 on the slope, can be detected, and the corresponding position information can be obtained. This position information can be the depth information corresponding to the object point, that is, the distance from the camera plane. The first direction is the direction that the target object faces, that is, the direction that the program user observes the target scene through the target program interface. Through the information of these 5 object points, the model can perceive that there is a slope in front of the target object, so as to control the target object to perform object behaviors such as detouring and climbing.

[0143] The second type: surrounding detection ray

[0144] In a possible implementation, in order to be able to perceive the scene objects distributed around the target object with a finer granularity, the computer device can also use surrounding detection rays for scene perception. In this implementation, the horizontal detection ray can include the surrounding detection ray. The surrounding detection ray refers to the ray that starts from the position where the target object is located and horizontally detects around the target object. Therefore, compared with the plane detection ray that horizontally detects in a single direction, the perception effect on the scene objects around the target object is better.

[0145] When determining the starting point of the detection ray according to the position of the target object at the target moment, the computer device can determine the starting point of the ray located on the target object according to the position of the target object at the target moment. This starting point of the ray is the object point on the target object where the probability of interacting with the scene object is greater than the second preset threshold. Starting from the starting point of the ray and horizontally projecting a horizontal detection ray towards the target scene, when determining the object point that the horizontal detection ray first passes through, the computer device can project multiple surrounding detection rays from the starting point of the ray towards the target scene horizontally, and determine the object points that the multiple surrounding detection rays first pass through respectively. The ray directions corresponding to the multiple surrounding detection rays are different, so as to accurately detect the object points distributed around the target object and achieve a finer granularity of scene perception.

[0146] That is, the computer device can focus on the parts of the target object that are likely to interact with the surrounding scene objects to determine the scene perception information, so that the behavior control instructions determined based on the scene perception information can enable the target object to interact reasonably with the surrounding scene objects. For example, as Figure 5 shown, the head and feet of the target object are parts with a relatively high probability of colliding with obstacles in the scene. Therefore, the computer device can place the ray origin at the head and feet of the target object, so that the surrounding detection ray can accurately detect the obstacles near the head and feet of the target object. Furthermore, when generating behavior control instructions based on the scene perception information, the probability of the head and feet of the target object colliding with the scene objects can be reduced, and the rationality of the behavior control instructions can be further improved.

[0147] It can be understood that when the program user performs scene perception through vision, usually not only can the directly visible parts be perceived, but also the parts that may not be directly visible can be automatically supplemented in the mind based on the attributes of the scene objects. For example, for Figure 4 the slope shown in, when seeing the slope, the directly visible part is the uphill part. However, after the program user knows that there is a slope ahead, it can naturally perceive that there is a flat part above the slope.

[0148] In addition, in a three-dimensional scene, the position information of an object point is usually generated by the position information in the horizontal dimension and the vertical dimension in space. When the program user perceives the position of an object point in the target scene, usually the position information in the horizontal dimension and the vertical dimension is combined for perception. The above horizontal detection ray mainly perceives the position information in the horizontal dimension, such as the distance in the horizontal dimension between the target object and the like. Therefore, through the scene detection in the vertical dimension, the position information dimension of the object point can be further enriched. Furthermore, when the instruction generation model generates behavior control instructions, the accuracy and authenticity of the perception of the object point position can be improved, thereby further improving the generation accuracy and logic of the behavior control instructions.

[0149] In a possible implementation, the detection ray may include a vertical detection ray. When determining the ray origin of the detection ray based on the position of the target object at the target moment, the computer device may determine a detection plane located above the target object according to the position of the target object at the target moment. The detection plane is a horizontal plane, and this detection plane is used for detecting the position of object points in the vertical dimension. The computer device may determine the ray origin located on the detection plane. When projecting the detection ray from the ray origin to the target scene where the target object is located and determining the object point that the detection ray first passes through, the computer device may project a vertical detection ray vertically from the ray origin to the target scene to determine the object point that the vertical detection ray first passes through. Thus, the position of the object point determined in this way can, on the one hand, be used for the program user to indirectly perceive object points with relatively difficult direct perception in the visual dimension, such as the top surface of a slope, the top surface of a box, the top surface of a ladder, etc. On the other hand, it can enrich the expression dimension of the scene point position, enabling the instruction generation model to more accurately and realistically perceive the scene points in the scene through the scene perception information, thereby further improving the accuracy of the generated behavior control instructions.

[0150] As Figure 6 shown, when the program user observes the target scene, object points 1 to 4 are directly visible object points, that is, object points that can also be detected by horizontal detection rays. For this part of object points, the vertical detection ray can enrich the dimension of the position information, so as to more accurately express the position of the object points; object points 5 and 6 are object points that cannot be directly observed in the visual dimension, but can be indirectly perceived through the attributes of the slope scene object. Therefore, the vertical detection ray can perceive this part of object points, improving the authenticity of the simulation of the actual scene perception method of the scene perception information for the program user.

[0151] It should be emphasized that the above-mentioned multiple detection rays can be used alone or in combination with multiple detection devices. For example, in an actual application scenario, the horizontal detection ray can integrate the plane detection ray and the surrounding detection ray, and can further combine with the vertical detection ray to achieve the most accurate and realistic perception of the object points in the target scene.

[0152] The above content mainly introduces the technology of this application for scene perception. Next, the technology for object perception will be introduced in detail.

[0153] In a possible implementation, the target perception information may include interaction object perception information, which is used to represent the position of the interaction object that the program user can perceive at the target moment. Here, the interaction object is used to interact with the target object by performing object behaviors in the target scenario. For example, in a game program, the interaction object may be a hostile game character in the game scene, and the target object may be the controlled game character. By controlling the game character to perform object behaviors, interactions such as attacking the hostile game character can be carried out. The target scenario is the scenario where the target object is located at the target moment. It should be emphasized that the interaction object is not a scene object that constitutes the target scenario, and it is an object in a different dimension from the scene object. That is, the above method for determining the scene perception information may not be used to perceive the interaction object.

[0154] Through the interaction object perception information, the instruction generation model can perceive information similar to the position information of the interaction object perceived by the program user. Therefore, after generating the behavior control instruction based on this interaction object perception information, it can more reasonably control the target object to interact with the interaction object by performing object behaviors. For example, in a game program, through the interaction object perception information, the position of the hostile game character can be perceived, so that it can be reasonably controlled whether the game character attacks the hostile game character or moves towards the hostile game character.

[0155] Among them, similar to the perception of the scene, the perception of the interaction object can also include two forms: direct perception and indirect perception.

[0156] First, in a possible implementation, the interaction object perception information can be determined in the following way:

[0157] The computer device can determine whether there is a scene object between the target object and the interaction object. The scene object is the object that constitutes the target scenario mentioned above. Based on the fact that there is no scene object between the target object and the interaction object, it means that when looking at the interaction object from the side of the target object, the line of sight will not be blocked by the scene object. Therefore, when the program user controls the target object, it is very likely that the interaction object can be directly perceived, and the perception of the interaction object is relatively clear and accurate. Therefore, in order to simulate the perception of the interaction object by the program user, the computer device can directly determine the position information corresponding to the interaction object as the interaction object perception information. The position information is used to identify the position of the interaction object in the target scenario, so that the instruction generation model can accurately perceive the interaction object, and then can control the target object to interact reasonably with the interaction object through the determined behavior control instruction. For example, in a shooting game program, when the interaction object perception information is used to identify the accurate position of the interaction object, the generated behavior control instruction can be used to control the target object to accurately aim at the interaction object and shoot.

[0158] Based on the existence of a scene object between the target object and the interaction object, and the interaction object meeting the perceivability condition, on the one hand, the line of sight from the target object to observe the interaction object will be blocked by the scene object, so it is very likely that the interaction object cannot be directly perceived. On the other hand, the perceivability condition is used to judge whether the interaction object can be perceived. If the interaction object meets this perceivability condition, it means that although it cannot be directly perceived from the visual dimension, it can be indirectly perceived from other dimensions (such as the auditory dimension), etc. However, for the program user, the perception accuracy of this perception method is lower than that of directly perceiving through the visual dimension. Based on this, in order to simulate the actual perception method of the program user, the computer device can determine the area information corresponding to the interaction object as the interaction object perception information. This area information is used to identify the position area where the interaction object is located in the target scene. The position area includes multiple positions, and among the multiple positions, there is the position where the interaction object is located in the target scene. That is, when the program user can only indirectly perceive the interaction object, the provided position information is relatively vague, so that the instruction generation model can only perceive the general position and cannot perceive the specific position, thus being able to distinguish the perception accuracy of direct perception and indirect perception, simulate the actual perception method of the program user, and improve the authenticity of the interaction object perception information.

[0159] It should be emphasized that the above analysis of the interaction object and the target object is all carried out for the object positions at the target moment, so the interaction object perception information corresponding to the target moment can be obtained.

[0160] Among them, the perceivability conditions used to judge whether indirect perception is formed can include various types, and each perceivability condition will be introduced in detail next.

[0161] The first type: Sound effect perception

[0162] In a possible implementation, the perceivability condition includes that the interaction object has sound effect information in a playing state at the target moment, and the target distance between the interaction object and the target object is less than the distance threshold. The sound effect information in a playing state means that at the target moment, in the target program, the interaction object is emitting this sound effect information. This distance threshold is used to measure whether the sound effect information can be perceived by the program user controlling the target object. When the target distance between the interaction object and the target object is less than the distance threshold, it means that the distance between the interaction object and the target object is relatively close. Therefore, the program user controlling the target object can probably perceive the sound effect information from the auditory dimension, and thus can probably perceive the general orientation of the interaction object even in the case of being invisible. Therefore, it meets the perceivability condition.

[0163] Among them, the number of positions included in the position area is inversely correlated with the volume of the sound effect information and inversely correlated with the target distance. Since the position corresponding to the interaction object must be included in the positions included in the position area, the fewer the number of positions included in the position area, the more accurate the representation of the position where the interaction object is located. That is, in the case of indirectly perceiving through sound effect information, the perception accuracy of the position where the interaction object is located is positively correlated with the volume and inversely correlated with the distance between objects, which conforms to the characteristics of perceiving through sound effect information in the actual scenario.

[0164] The second type: self-historical perception

[0165] It can be understood that the program user usually has a certain position memory of the interaction objects perceived in the historical period. For example, if the program user has just encountered an interaction object, then in the short term, the memory of the position where the interaction object is located can be retained, so that a relatively vague perception of the position where the interaction object is currently located can be obtained. If the program user saw the interaction object on the right side of their controlled object 10 seconds ago, then at the current moment, even if the interaction object cannot be seen, it can be roughly perceived that the interaction object is located in the right position.

[0166] Based on this, in order to simulate the perception of the program user based on the information in the historical period, the perceivable condition may include that in the first historical period, there is a first historical moment that satisfies that at the first historical moment, there is no scene object between the interaction object and the target object, that is, at the first historical moment, the interaction object can be directly perceived by the program user controlling the target object in the visual dimension, so the specific position of the interaction object can be perceived at the first historical moment. The end moment of the first historical period is the target moment, and the length of the first historical period is less than the first length threshold, that is, the first historical period is used to simulate the period when the program user retains the position perception memory. Beyond this period, the probability that the program user forgets the position of the interaction object is relatively high, or the position of the interaction object has changed sufficiently, and the historical memory fails.

[0167] The number of positions included in the position area is positively correlated with the length of the first target period, and the first target period is the period from the first historical moment to the target moment. That is, the farther away from the moment when the position of the interaction object can be directly perceived, the lower the identification accuracy of the regional information for the actual position of the interaction object, so as to simulate the characteristic that the position perception of the interaction objects that have appeared in the historical period by the program user gradually fails over time.

[0168] The third type: historical perception of perceiving shared objects

[0169] In a target program, there may be some objects that can share perception information with a target object. For example, in a game program, a game character controlled by a program user can form a team with other game characters controlled by program users. Other program users can use voice or other means to inform this program user of the interaction object information (such as the position of a hostile character) they perceive through controlling the game character. Thus, this program user can generally understand the location of the interaction object without directly perceiving its position, and can then control the game character to perform corresponding object behaviors, such as approaching or surrounding the interaction object.

[0170] Based on this, in order to simulate the historical perception through other objects, the perceivable condition may include that in a second historical period, there is a second historical moment that satisfies the condition that the perception information corresponding to the perception sharing object includes the target interaction object perception information corresponding to the interaction object at the second historical moment, that is, at the second historical moment, the interaction object can be perceived by the program user who controls the perception sharing object. The perception sharing object is an object in the target program that shares perception information with the target object, and the sharable perception information may include the above object perception information, scene perception information, etc., which is not limited here.

[0171] The number of positions included in the position area is positively correlated with the length of the second target period and positively correlated with the number of positions identified by the target interaction object perception information. The second target period is the period from the second historical moment to the target moment. Similar to the way of its own historical perception, the position perception of the interaction object by the program user corresponding to the shared perception object will gradually become invalid over time, so the position perception shared with the program user corresponding to the target object will also gradually become invalid as time passes. At the same time, since the interaction object perception is carried out in a shared manner, the accuracy of the position perception of the interaction object depends on the source of the shared information, that is, the accuracy of the interaction object perception through the perception sharing object. Thus, the accuracy of the position identification of the interaction object by this area information is inversely correlated with the length of the second target period and positively correlated with the perception accuracy of the shared perception information source.

[0172] Overall, the perception of the interaction object can be as Figure 7As shown, the computer device can first determine whether the interaction object is exposed, that is, whether it can be directly perceived in the visual dimension. This determination can be made by connecting the object points on the target object to the object points on the interaction object, and judging whether there are object points of the scene object on the connection line to determine whether there is a scene object in the middle, so as to judge whether it can be directly perceived. Otherwise, it can be judged whether it is indirectly perceived through the perceivable conditions, including the above three indirect perception methods. Through direct perception, the corresponding position of the interaction object can be obtained, that is, accurate position information. Through indirect perception, the area where the interaction object is located can be obtained, that is, fuzzy position information. Taking sound effect perception as an example, the object perception information can be the area information where the sound effect information is located, which can represent the orientation where the sound effect information is located, rather than accurate position information. Based on the perception accuracy of the sound effect information satisfying this perception accuracy curve, the greater the distance between objects, the smaller the volume, and the lower the perception accuracy.

[0173] Through the above method, the instruction generation model can simulate the true perception of the target scene and the interaction object by the program user, so that the determined behavior control instruction can further conform to the control logic of the program user, avoiding the situation that the target object makes unrealistic object behaviors due to the amount and accuracy of the information perceived by the model far exceeding that of the program user. For example, if too much perception information is provided to the instruction generation model, in a shooting game program, it may cause the target object to make abnormal object behaviors that the program user cannot perform, such as accurately shooting through walls and accurately dodging through walls.

[0174] In addition to the perception information in the above two dimensions, this application also includes the perception information in the third dimension: path perception. It can be understood that usually, the paths that the target object can move in the target scene are limited. For example, there can be multiple channels in the target scene, and the target object can only move in the channels. When the program user controls the movement of the target object, it is usually determined by combining the nearby selectable paths and the movement requirements. Therefore, by providing the path information corresponding to the target object, it can help the instruction generation model more reasonably determine the behavior control instruction.

[0175] In a possible implementation manner, the target perception information further includes scene path information, and the scene path information is used to identify the actionable path corresponding to the target object in the target scene. The target scene is the scene where the target object is located at the target moment, and the target object can act in the actionable path. For example, when the instruction generation model perceives that there is a hostile object near the target object, it can determine a reasonable path to avoid / attack the hostile object through the scene path information.

[0176] Among them, it can be understood that since the scene composition of the target scene may be relatively complex, at each position, the target object may have multiple paths to move, and the rationality of moving along different paths may vary. For example, in Figure 8 the path of the target scene may include three positions: position 1, position 2, and position 3. The target object is at position 1. If the target object needs to move upward, it can choose to move to position 2 or position 3. However, position 2 is at the intersection of multiple paths, so there are more subsequent path options. Compared with position 3, it is more convenient for subsequent execution of object behavior. Therefore, the rationality of position 2 is better than that of position 3.

[0177] Based on this, in a possible implementation, in order to further improve the rationality of the behavior control instruction, the computer device can first remove the part of the path information that is less helpful for determining the behavior control instruction. The scene path information can be determined in the following way:

[0178] The computer device can first determine multiple reachable positions corresponding to the target scene. The multiple reachable positions are the positions that the target scene supports the object to reach. For example Figure 8 the position 1, position 2, and position 3 in. The computer device can generate the initial path information corresponding to the target scene based on the multiple reachable positions. The initial path information is used to identify multiple initial paths between the multiple reachable positions. As Figure 9 shown, according to position 1, position 2, and position 3, two initial paths, path 1 and path 2, can be determined. Among them, path 1 is used to reach position 2 from position 1, and path 2 is used to reach position 3 from position 1.

[0179] Then, the computer device can determine multiple key positions among the multiple reachable positions. The key position is the position with a relatively high selection probability when the program user actually controls the target object. That is, when the program user controls the target object to move at the target moment, the probability of reaching the key position is greater than the probability of reaching the reachable positions other than the key position among the reachable positions. The computer device can determine the scene path information corresponding to the target scene based on the initial path information and the multiple key positions. The actionable path is the initial path that includes the key position among the multiple initial paths. That is, the computer device will delete the paths that do not include the key position in the initial paths. These paths are the paths with a relatively low selection probability when the program user actually controls the target object. Therefore, when automatically controlling the target object to move, moving through these paths does not conform to the actual control logic of the program user. Thus, the behavior control instruction determined based on the processed scene path information can control the target object to select a more reasonable path to move in the target scene, further improving the rationality and logic of the object behavior.

[0180] For example, under normal circumstances, the program user will usually choose to move to positions that have been selected more times and have more subsequent optional paths. Therefore, they will tend to choose paths that can reach these positions. Thus, these positions can be determined as key positions. Also, for the position where the target object is currently located, the program user usually tends to move to nearby positions rather than positions farther away in the target scenario. Therefore, when determining the scene path information corresponding to the target moment, positions closer to the position of the target object at the target moment can be determined as key positions. The method for determining key positions can include various ways and is not limited here. As Figure 10 shown, since the subsequent path selection for position 2 is much more than that for position 3, when the program user controls the movement of the target object, the probability of reaching position 2 is much higher than that of reaching position 3. Position 2 can be determined as a key position. When determining the scene path information, path 2 leading to position 3 can be removed, and only path 1 is retained to prevent the instruction generation model from selecting path 2 with a lower degree of rationality for movement.

[0181] The above content mainly introduced the information content for the input instruction generation model in detail. To improve the rationality and authenticity of the behavior control instruction, the present application can also perform fine-grained control on the determination process of the behavior control instruction. Next, the technical content in this regard will be introduced in detail.

[0182] First of all, it can be understood that under normal circumstances, the program user controls the object in the target program through the program interface displayed by the target program. Therefore, the object behaviors that the object can perform should be the object behaviors corresponding to the control behaviors that the program user can make through the program interface. For example, as Figure 11 shown, in the game interface corresponding to the target moment, the left hand controllable area of the player includes a shooting control for controlling the shooting of the game character, a movement control for controlling the movement of the game character, and the right hand controllable area includes a shooting control for controlling the shooting of the game character, a squat control for controlling the squatting of the game character, a jump control for controlling the jumping of the game character, a skill release control for controlling the release of skills of the game character, etc. Under normal circumstances, for each controllable area, the player can only select one control to trigger. For example, when the left hand triggers the shooting control, it cannot move, etc. Thus, it can be seen that the control behaviors that can be made through the program interface are limited. Therefore, not any behavior control instruction is the behavior control instruction that the program user can actually generate when controlling the target object. For example, it is impossible to perform object behaviors such as moving, squatting, and shooting simultaneously.

[0183] Based on this, in a possible implementation, in order to further improve the rationality and authenticity of the behavior control instruction, the computer device can restrict the object behaviors that the behavior control instruction can control and execute. When executing step S202, the computer device can execute steps S2021 - S2023 (not shown in the figure), and steps S2021 - S2023 are a possible implementation of step S202, including:

[0184] S2021: Generate a plurality of pending behavior control instructions according to the target object state information and the target perception information.

[0185] Among them, the plurality of pending behavior control instructions are the plurality of pending behavior control instructions that the initial instruction control model believes can be used to control the target object to execute the object behavior under the object state and perception information identified by the information. If no subsequent processing is performed, the initial instruction generation model will only output the pending behavior control instruction with the highest execution probability as the target behavior control instruction.

[0186] S2022: Determine a plurality of sub - behavior control instructions corresponding to the target pending behavior control instruction.

[0187] The plurality of sub - behavior control instructions have respectively corresponding sub - object behaviors and control behaviors. The sub - object behaviors respectively corresponding to the plurality of sub - behavior control instructions are used to constitute the object behavior corresponding to the target pending behavior control instruction. For example, the object behavior can be shooting while moving, then the plurality of sub - object behaviors can be the moving behavior and the shooting behavior.

[0188] In this application, the behavior control instruction can be generated by the control behavior of the program user on the program interface. Each control behavior can generate a sub - behavior control instruction. The program user can execute multiple control behaviors simultaneously. For example, multiple controls can be triggered simultaneously, so as to generate a behavior control instruction including a plurality of sub - behavior control instructions at the same moment.

[0189] The target sub - behavior control instruction corresponds to the target control behavior, that is, the target control behavior executed through the target program interface is used to generate the target sub - behavior control instruction. The target sub - behavior control instruction is any one of the plurality of sub - behavior control instructions, the target pending behavior control is any one of the plurality of pending behavior control instructions, and the target program interface is used to control the target object. In this step, the computer device can split the pending behavior control instruction, so as to determine the control behavior combination that can generate the pending behavior control instruction, and further can judge whether the pending behavior control instruction is a reasonable behavior control instruction that the program user can generate through the control behavior.

[0190] S2023: Based on the control behaviors corresponding to multiple sub-behavior control instructions being multiple control behaviors that can be executed in parallel through the target program interface, determine the target pending behavior control instruction as the target behavior control instruction.

[0191] If the control behaviors corresponding to multiple sub-behavior control instructions are multiple control behaviors that can be executed in parallel through the target program interface, it indicates that the target pending behavior control instruction is a behavior control instruction that the program user can generate through control behaviors, with a relatively high degree of reasonableness. Therefore, it can be output as the target behavior control instruction. Conversely, if they are not multiple control behaviors that can be executed in parallel through the target program interface, it means that the program user cannot generate this behavior control instruction through control behaviors, that is, the program user cannot control the target object to execute the object behavior corresponding to this behavior control instruction. Therefore, the degree of reasonableness is relatively low, and the computer device can control the model not to output this behavior control instruction.

[0192] As Figure 12 shown, in the initial instruction generation model, it includes a sub-behavior control instruction generation part corresponding to the left hand and a sub-behavior control instruction generation part corresponding to the right hand, which are respectively used to generate sub-behavior control instructions corresponding to the left hand and sub-behavior control instructions corresponding to the right hand. These sub-behavior control instructions will be screened according to the control behaviors that can be executed in parallel. Among them, the control behaviors that can be executed in parallel include one of the control behaviors that the left hand can execute and one of each of the control behaviors that the right hand can execute. Thus, among the screened sub-behavior control instructions, there can be at most one sub-behavior control instruction corresponding to the control behavior executed by the left hand and one sub-behavior control instruction corresponding to the control behavior executed by the right hand. Furthermore, the target behavior control instruction for output can be combined, and this target behavior control instruction can be used to control the target object not to execute the object behavior that the program user cannot control to execute, such as not squatting and jumping at the same time.

[0193] In addition, the parallelizable control behaviors can also be restricted by quantity. For example, when a player controls through the program interface, usually at most three fingers can be used for operation simultaneously. Therefore, it can be restricted that at most three sub-behavior control instructions can be included in the same behavior control instruction. The specific setting of the parallelizable control behaviors can be based on the actual control method corresponding to the target program, which is not limited here.

[0194] In addition to restricting based on the parallel execution of control behaviors, in one possible implementation, the computer device can also restrict the behavior control instructions from the execution amplitude of the object behavior. It can be understood that during the use of the target program, the amplitude of the object behavior that the program user can control the target object to execute instantaneously is limited. For example, when controlling the target object to turn through the program interface, due to the limitation of the program user's control ability or the size limitation of the program interface, the program user may be able to control the target object to turn up to 30 degrees at a single moment. However, when the computer device actively controls the target object, since the processing speed of the computer device is relatively fast, it can control the target object to turn 90 degrees or even more at a single moment. Based on this, in order to further improve the authenticity of the object behavior executed by the automatically controlled target object, the computer device can restrict the amplitude of the object behavior executed by the target object at a single moment.

[0195] When executing step S202, the computer device can execute steps S2026 - S2028 (not shown in the figure). Steps S2026 - S2028 are a possible implementation of step S202 and include:

[0196] S2026: Generate a second initial behavior control instruction corresponding to the target moment according to the target object state information and the target perception information.

[0197] The second initial behavior control instruction is used to control the target object to execute the target object behavior at the first behavior amplitude at the target moment.

[0198] S2027: Based on the fact that the first behavior amplitude does not exceed the behavior amplitude threshold corresponding to the target object behavior, determine the second initial behavior control instruction as the target behavior control instruction.

[0199] The computer device can preset the behavior amplitude threshold, which is the maximum amplitude for controlling the target object to execute the target object behavior through the target program interface, that is, to simulate the maximum amplitude when the program user controls the target object to execute the target object behavior through the target program interface at the target moment. For example, when the target object behavior is turning, the behavior amplitude threshold can be 30 degrees. The target program interface is used to control the target object, that is, the program interface displayed by the target program at the target moment.

[0200] S2028: Based on the fact that the first behavior amplitude exceeds the behavior amplitude threshold, adjust the second initial behavior control instruction to obtain the target behavior control instruction.

[0201] If the amplitude of the first line exceeds the behavior amplitude threshold, it indicates that directly executing the second initial behavior control instruction at this time will cause the amplitude of the target object's execution of the target object behavior to be too large, resulting in distorted object behavior. Therefore, the computer device can adjust the second initial behavior control instruction to reduce the behavior amplitude corresponding to the second initial behavior control instruction, and obtain a target behavior control instruction, which is used to control the target object to execute the target object behavior at the target moment with the behavior amplitude threshold. For example, if the second initial behavior control instruction is used to control the target object to turn 90 degrees at the target moment, the target behavior control instruction obtained through adjustment can be used to control the target object to turn 30 degrees at the target moment.

[0202] In addition to being able to simulate the program user in terms of the control method, the computer device can also simulate the program user in terms of the control speed and accuracy.

[0203] It can be understood that, generally, due to the limited perception ability and limb control ability of people, when controlling the target object, it is usually impossible to perform object behaviors with too high precision. For example, when controlling the target object to shoot, the shooting angle may deviate and it is impossible to hit the target 100%. Therefore, in a possible implementation manner, in order to make the automatically controlled target object closer to the actual control characteristics of the program user, the computer device can control the target object to execute object behaviors with a certain deviation through the generated behavior control instructions.

[0204] When executing step S202, the computer device can execute steps S2024 - S2025 (not shown in the figure), and steps S2024 - S2025 are a possible implementation manner of step S202, including:

[0205] S2024: Generate a first initial behavior control instruction corresponding to the target moment according to the target object state information and the target perception information.

[0206] The first initial behavior control instruction is used to control the target object to execute the first initial object behavior at the target moment.

[0207] S2025: Generate a target object control instruction according to the behavior deviation parameter and the first initial behavior control instruction.

[0208] The behavior deviation parameter corresponds to the target behavior deviation. The target object behavior is the first initial object behavior that generates the target behavior deviation when executed, that is, through this behavior deviation parameter, the target object can generate this target behavior deviation when executing this first initial object behavior, so as to simulate the behavior deviation of the program user when controlling the target object to execute this first initial object behavior, and improve the authenticity of the execution of the object behavior.

[0209] Among them, the behavioral deviation can include various forms. For example, in a possible implementation, the target program can be a shooting game program, the first initial object behavior is a shooting behavior, and the target behavioral deviation includes a horizontal shooting direction deviation and / or a vertical shooting direction deviation, so as to simulate the phenomenon that when a player controls a target object to shoot, due to reasons such as low aiming accuracy or hand jitter, a shooting deviation occurs.

[0210] Among them, the shooting direction deviation is usually also affected by the object state. For example, when the target object is in a moving state, the shooting direction deviation is usually larger, and it is more difficult to accurately shoot the enemy in the moving state. The horizontal shooting direction deviation mainly affects the hit rate of shooting, and the vertical shooting direction deviation not only affects the hit rate, but also affects the distribution of the shooting hit parts.

[0211] In the stage of aiming at and shooting the enemy, the computer device can count the player data to obtain the mapping function between turning and turning accuracy. The formula is as follows, where Arc is the angle difference between the character controlled by the player and the hostile character, and a and b are different constant coefficients, and different parameters can be set according to the strength of the automatically controlled character required to achieve intelligent control of the character at different game levels.

[0212] noise = a * log(Arc) + b

[0213] It can be seen from the above formula that the shooting direction deviation noise is positively correlated with the angle, that is, the larger the angle between the character controlled by the player and the hostile character, the greater the deviation in aiming and shooting, which is in line with the accuracy of the player's actual control of the game character for aiming and shooting.

[0214] Next, the simulation of the control speed dimension will be introduced. It can be understood that when the program user controls the target object, it is usually controlled based on the perceived information. When new perceivable information appears, there is usually a time interval between perceiving the information and actually controlling the target object to execute the object behavior, that is, the reaction time of the program user. However, for the computer device, the generation of the behavior control instruction can be carried out instantly when new information appears. Therefore, if the time for executing the object behavior is not controlled, the situation of distortion may occur due to the too fast reaction speed of the automatically controlled target object. For example, the character controlled by the computer device can aim and shoot instantly when discovering the hostile character, while the player usually needs a period of time to react.

[0215] Based on this, in a possible implementation, in order to further improve the authenticity of the object behavior executed by the automatically controlled target object, the computer device can simulate the reaction time of the program user. In this implementation, the target object behavior is the object behavior for the target interaction object in the target program. For example, shooting at the hostile game character in the game program, etc. When executing step S203, the computer device can execute steps S2031 - S2032 (not shown in the figure). Steps S2031 - S2032 are a possible implementation of step S203 and include:

[0216] S2031: Based on the time interval between the target moment and the object appearance moment reaching the reaction duration, control the target object to execute the target object behavior through the target behavior control instruction.

[0217] Among them, the object appearance moment is the moment when the target interaction object appears in the program interface corresponding to the target program, that is, the moment when the program user who controls the target object through the program interface can directly perceive the target interaction object. The reaction duration is positively correlated with the perception difficulty corresponding to the target interaction object. The perception difficulty is the difficulty of perceiving the target interaction object at the object appearance moment through the program interface. That is, the easier it is for the program user to perceive the target interaction object, the shorter the reaction duration, and vice versa, so as to simulate the actual situation of the program user perceiving the target interaction object.

[0218] Based on the time interval between the target moment and the object appearance moment reaching the reaction duration, it means that the object behavior starts to be executed at this time, which conforms to the actual reaction situation of the program user. Therefore, the computer device can control the target object to execute the target object behavior through the target behavior control instruction. On the contrary, if the time interval between the target moment and the object appearance moment does not reach the reaction duration, it means that the time interval from the moment when the program user can perceive the target interaction object is too short. Under normal circumstances, the program user has not had time to react to control the target object to execute the object behavior. Therefore, at this time, the computer device can not execute the target behavior control instruction, so as to simulate a more realistic reaction duration.

[0219] S2032: Adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task by the target object through executing the target object behavior, and obtain the instruction generation model.

[0220] Among them, there are various ways to measure the perception difficulty. Usually, the degree of prominence of the target interaction object displayed in the program interface determines the perception reaction speed of the program user. That is, the more prominent the display, the lower the perception difficulty of the program user, and the faster the reaction speed. Based on this, in a possible implementation, the perception difficulty can be positively correlated with the distance between the target interaction object and the target object at the moment the object appears, and positively correlated with the position difference corresponding to the display position of the target interaction object in the program interface. The position difference is the difference between the display position and the center position of the program interface. That is, the closer the distance between the target interaction object and the target object, usually the larger the display area in the program interface, so the perception difficulty is lower; the program user mainly perceives through the center position in the program interface, so the closer the display position of the target interaction object is to the center position, usually the easier it is for the program user to perceive.

[0221] Taking a shooting game program as an example, in a possible implementation, the target program is a shooting game program, and the target object behavior is to aim at the target interaction object. That is, through this method, the reaction duration required for the player to go from discovering the target interaction object to aiming at the target interaction object can be simulated. Among them, the determination method of the reaction duration Delay can be shown by the following formula:

[0222] Delay∝(Angle,Distance)

[0223] Among them, Angle is the position difference corresponding to the display position of the target interaction object. In other words, the closer the target interaction object is to the center position, the shorter the simulated reaction time. The reaction time is also related to the distance Distance between the target interaction object and the target object. The closer the distance, the larger the display area of the target interaction object in the program interface, the lower the perception difficulty, and the shorter the reaction time.

[0224] To sum up, for a shooting game program, as Figure 13 shown, the computer device can simulate the player's actual operations in three stages: the stage of perceiving the enemy, the stage of aiming at the enemy, and the stage of shooting the enemy. In the stage of perceiving the enemy, the reaction time of the player can be aligned, and after reaching the reaction duration, the target object is controlled to perform the aiming behavior; in the stage of aiming at the enemy, the turning amplitude at a single moment can be restricted to align the turning distribution; in the stage of shooting the enemy, the shooting deviation can be adjusted to align the shooting ability of the player, so as to realistically simulate the player's operations in the shooting game program as a whole.

[0225] To facilitate the understanding of the technical solution provided by this application, next, a model training method provided by an embodiment of this application will be introduced in combination with an actual application scenario.

[0226] See Figure 14 , Figure 14 Figure 14 is a flowchart of a model training method in an actual application scenario provided by an embodiment of this application. In this actual application scenario, the computer device may be a server of a shooting game program, the target program may be a shooting game program, and the target object is a game character that can be controlled in the game program. The method includes:

[0227] S1401: Determine the object state information corresponding to the target object at the target moment.

[0228] The object state information may include character health information, character position information, character skill quantity information, etc.

[0229] S1402: Determine the perception information corresponding to the target object at the target moment.

[0230] Among them, the perception information may include the following dimensions:

[0231] S14021: Determine the scene perception information.

[0232] The scene perception information can be obtained by combining three methods: planar detection, vertical detection, and surrounding detection. Respectively, use planar detection rays, vertical detection rays, and surrounding detection rays to detect scene points in the target scene, and determine the position information of the scene points as the scene perception information.

[0233] S14022: Determine the scene path information.

[0234] The server can obtain the map resource file of the game scene. Then, screen out the positions that the character can reach from the map resource file. These positions usually include surface positions such as the ground, stairs, platforms, etc. that can be walked on. Subsequently, based on the reachable positions, generate a path structure mesh. A mesh is a mesh structure used to represent the path that a character can walk on the map. Here, the navigation mesh (NavMesh) generation algorithm can be used to create the path structure mesh. Finally, perform clipping and pruning operations on the generated path structure mesh to remove the paths that do not include key positions, so as to simplify the map representation and reduce the computational complexity, thereby obtaining the scene path information. The scene path information can be composed of vertices and the edges between vertices. The vertices represent key positions on the map (such as corners, intersections, etc.), and the edges represent the paths that the character can follow.

[0235] S14023: Determine the interaction object perception information.

[0236] The server can determine the interaction object perception information from two dimensions: direct perception and indirect perception, as Figure 7 shown.

[0237] S1403: Determine the initial behavior control instruction through the initial instruction generation model.

[0238] The initial instruction generation model may include various model architectures, which are not limited here.

[0239] S1404: Adjust the initial behavior control instruction to obtain the target behavior control instruction.

[0240] Among them, the adjustment of the initial behavior control instruction may include the following dimensions:

[0241] S14041: Adjust based on the behavior amplitude.

[0242] The server can adjust to make the behavior amplitude corresponding to the target behavior control instruction not exceed the behavior amplitude threshold of the corresponding target object behavior.

[0243] S14042: Adjust based on the sub-behavior control instructions included in the target behavior control instruction.

[0244] Through adjustment, the control behavior combination corresponding to the target behavior control instruction can be made into multiple control behaviors that can be executed in parallel.

[0245] S14043: Adjust based on the behavior deviation.

[0246] Through adjustment, when controlling the target object to execute the target object behavior through the target behavior control instruction, the behavior deviation of the player in object behaviors such as aiming and shooting during operation can be simulated.

[0247] S1405: Based on the time interval between the target moment and the object appearance moment reaching the reaction duration, control the target object to execute the target object behavior through the target behavior control instruction.

[0248] Through this step, the reaction duration of the player between perceiving information and making control can be simulated to further improve the authenticity of automatic control.

[0249] S1406: According to the achievement degree of the behavior task by the target object through executing the target object behavior, adjust the model parameters of the initial instruction generation model to obtain the instruction generation model.

[0250] The behavior task can be to defeat a specified number of hostile game characters within a specified time. If the achievement degree reaches the preset target achievement degree, or the number of training iterations reaches the upper limit, the model training can be stopped to obtain the final instruction generation model. Otherwise, parameter adjustment can continue. Among them, since this application does not refer to the actual operation data of players, in order to strengthen the behavior expression of the automatically controlled game characters in various game scenarios, the server can enrich the scene locations involved in the target object during the training process. In the early stage of training, the server can randomly place the target object in various corners of the game scene to enable it to fully explore. Thus, it can be avoided that when the automatically controlled game character enters some game scenes not involved in the training process, the problem of abnormal object behavior occurs.

[0251] S1407: Control the object in the shooting game program to automatically execute the object behavior through the instruction generation model.

[0252] It can be seen from the above technical solutions that the model training method provided by this application has the following technical improvements:

[0253] 1. Through perfect modeling, the training process can produce good anthropomorphism, logic and realism without relying on the actual data of the program user, reducing the model training difficulty and improving the model training efficiency.

[0254] 2. By simulating the program user from two dimensions of information perception and behavior operation, the instruction generation model can understand the object control logic of the program user from the root, rather than simply imitating the behavior of the program user. Therefore, it has a wide range of applications and can be widely used in various scenarios, getting rid of the limitations on the behavior control instructions that the model can generate when training based on the collected data.

[0255] 3. Since the dependence on the data of the program user is small, for programs with less data of the program user, the instruction generation model trained by this application can also make the controlled object maintain good anthropomorphism and has a wider range of applications.

[0256] Based on the model training method provided in the above embodiments, this application also provides a model training device. See Figure 15 , Figure 15 is the structural block diagram of a model training device provided in an embodiment of this application. The device 1500 includes a determination unit 1501, a generation unit 1502 and an adjustment unit 1503:

[0257] The determining unit 1501 is configured to determine the target object state information and the target perception information corresponding to the target object at the target moment. The target object state information is used to represent the object state of the target object in the target program, and the perception information is used to simulate the information that can be perceived by the user of the target program by controlling the target object at the target moment;

[0258] The generating unit 1502 is configured to generate a target behavior control instruction corresponding to the target moment through an initial instruction generation model according to the target object state information and the target perception information. The target behavior control instruction is used to control the target object to perform a target object behavior at the target moment;

[0259] The adjusting unit 1503 is configured to adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task achieved by the target object by performing the target object behavior, so as to obtain an instruction generation model. The achievement degree determined by the instruction generation model reaches the target achievement degree. The achievement degree is used to represent the reasonable degree of the object behavior of the target object. The reasonable degree represented by the target achievement degree is greater than a first preset threshold. The instruction generation model is used to determine a behavior control instruction for controlling the target object according to the object state information and perception information corresponding to the target object.

[0260] In a possible implementation manner, the target perception information includes scene perception information, and the scene perception information is used to represent the position of the scene object that can be perceived by the program user at the target moment. The scene object is used to constitute the target scene where the target object is located at the target moment.

[0261] In a possible implementation manner, the scene perception information is determined by the following method:

[0262] Determine the ray starting point of the detection ray according to the position of the target object at the target moment;

[0263] Project the detection ray from the ray starting point to the target scene where the target object is located, and determine the object point where the detection ray first passes through. The object point is used to constitute the scene object;

[0264] Determine the position information corresponding to the object point as the scene perception information.

[0265] In a possible implementation manner, the detection ray includes a horizontal detection ray. Projecting the detection ray from the ray starting point to the target scene where the target object is located and determining the object point where the detection ray first passes through includes:

[0266] Starting from the ray origin, horizontally project the horizontal detection ray towards the target scene, and determine the object point where the horizontal detection ray first passes through.

[0267] In a possible implementation manner, the horizontal detection ray includes a planar detection ray. The determining the ray origin of the detection ray according to the position of the target object at the target moment includes:

[0268] According to the position of the target object at the target moment, determine the camera plane corresponding to the target object. The camera plane is used to render the target program interface corresponding to the target moment in the target program. The target program interface is used to control the target object in the target program. The pixel values of the pixel points included in the target program interface are determined based on the projection of the scene object onto the camera plane;

[0269] Determine the ray origin located on the camera plane;

[0270] The starting from the ray origin, horizontally project the horizontal detection ray towards the target scene, and determine the object point where the horizontal detection ray first passes through, includes:

[0271] Starting from the ray origin, project the planar detection ray towards the target scene in a first direction perpendicular to the camera plane, and determine the object point where the planar detection ray first passes through. The first direction is the direction of observing the target scene through the target program interface.

[0272] In a possible implementation manner, the horizontal detection ray includes a surrounding detection ray. The determining the ray origin of the detection ray according to the position of the target object at the target moment includes:

[0273] According to the position of the target object at the target moment, determine the ray origin located on the target object. The ray origin is the object point on the target object where the probability of interacting with the scene object is greater than a second preset threshold;

[0274] The starting from the ray origin, horizontally project the horizontal detection ray towards the target scene, and determine the object point where the horizontal detection ray first passes through, includes:

[0275] Starting from the ray origin, horizontally project multiple surrounding detection rays towards the target scene, and determine the object points where the multiple surrounding detection rays first pass through respectively. The ray directions corresponding to the multiple surrounding detection rays are different.

[0276] In a possible implementation, the detection ray includes a vertical detection ray. Determining the ray starting point of the detection ray according to the position of the target object at the target moment includes:

[0277] According to the position of the target object at the target moment, determine a detection plane located above the target object, and the detection plane is a horizontal plane;

[0278] Determine the ray starting point located on the detection plane;

[0279] Projecting the detection ray from the ray starting point to the target scene where the target object is located, and determining the object point first passed by the detection ray includes:

[0280] Project the vertical detection ray vertically from the ray starting point to the target scene, and determine the object point first passed by the vertical detection ray.

[0281] In a possible implementation, the target perception information includes interaction object perception information, and the interaction object perception information is used to represent the position of the interaction object that can be perceived by the program user at the target moment. The interaction object is used to interact with the target object in the target scene by performing object behaviors. The target scene is the scene where the target object is located at the target moment, and the interaction object is not a scene object that constitutes the target scene.

[0282] In a possible implementation, the interaction object perception information is determined in the following manner:

[0283] Determine whether there is the scene object between the target object and the interaction object;

[0284] Based on the fact that there is no such scene object between the target object and the interaction object, determine the position information corresponding to the interaction object as the interaction object perception information, and the position information is used to identify the position where the interaction object is located in the target scene;

[0285] Based on the fact that there is the scene object between the target object and the interaction object, and the interaction object meets the perceivable condition, determine the area information corresponding to the interaction object as the interaction object perception information, and the area information is used to identify the position area where the interaction object is located in the target scene. The position area includes multiple positions, and the multiple positions include the position where the interaction object is located in the target scene.

[0286] In a possible implementation, the perceivable condition includes that the interaction object has sound effect information in a playing state at the target moment, and the target distance between the interaction object and the target object is less than a distance threshold;

[0287] The number of positions included in the position area is inversely correlated with the volume of the sound effect information and inversely correlated with the target distance.

[0288] In a possible implementation, the perceivable condition includes that in a first historical period, there is a first historical moment satisfying that at the first historical moment, there is no such scene object between the interaction object and the target object, the end moment of the first historical period is the target moment, and the period length of the first historical period is less than a first length threshold;

[0289] The number of positions included in the position area is positively correlated with the period length of a first target period, and the first target period is the period from the first historical moment to the target moment.

[0290] In a possible implementation, the perceivable condition includes that in a second historical period, there is a second historical moment satisfying that at the second historical moment, the target interaction object perception information corresponding to the interaction object is included in the perception information corresponding to the perceived shared object, and the perceived shared object is an object in the target program that shares perception information with the target object;

[0291] The number of positions included in the position area is positively correlated with the period length of a second target period and positively correlated with the number of positions identified by the target interaction object perception information, and the second target period is the period from the second historical moment to the target moment.

[0292] In a possible implementation, the target perception information further includes scene path information, and the scene path information is used to identify the actionable path corresponding to the target object in the target scene, and the target scene is the scene where the target object is located at the target moment.

[0293] In a possible implementation, the scene path information is determined in the following manner:

[0294] Determine a plurality of reachable positions corresponding to the target scene, and the plurality of reachable positions are positions that the target scene supports the object to reach;

[0295] Generate initial path information corresponding to the target scene according to the plurality of reachable positions, and the initial path information is used to identify multiple initial paths between the plurality of reachable positions;

[0296] Determine multiple critical positions among the multiple reachable positions. When the program user controls the target object to move at the target moment, the probability of reaching the critical positions is greater than the probability of reaching the reachable positions other than the critical positions among the reachable positions;

[0297] According to the initial path information and the multiple critical positions, determine the scenario path information corresponding to the target scenario. The actionable path is the initial path among the multiple initial paths that includes the critical positions.

[0298] In a possible implementation manner, the generating unit 1502 is specifically configured to:

[0299] Generate multiple pending behavior control instructions according to the target object state information and the target perception information;

[0300] Determine multiple sub-behavior control instructions corresponding to the target pending behavior control instruction. The multiple sub-behavior control instructions have corresponding sub-object behaviors and control behaviors respectively. The sub-object behaviors corresponding to the multiple sub-behavior control instructions are used to constitute the object behavior corresponding to the target pending behavior control instruction. The target sub-behavior control instruction corresponds to the target control behavior. The target control behavior executed through the target program interface is used to generate the target sub-behavior control instruction. The target sub-behavior control instruction is any one of the multiple sub-behavior control instructions. The target pending behavior control is any one of the multiple pending behavior control instructions. The target program interface is used to control the target object;

[0301] Based on the control behaviors corresponding to the multiple sub-behavior control instructions being multiple control behaviors that can be executed in parallel through the target program interface, determine the target pending behavior control instruction as the target behavior control instruction.

[0302] In a possible implementation manner, the generating unit 1502 is specifically configured to:

[0303] Generate a first initial behavior control instruction corresponding to the target moment according to the target object state information and the target perception information. The first initial behavior control instruction is used to control the target object to perform a first initial object behavior at the target moment;

[0304] Generate the target object control instruction according to the behavior deviation parameter and the first initial behavior control instruction. The behavior deviation parameter corresponds to the target behavior deviation. The target object behavior is the first initial object behavior that generates the target behavior deviation when executed.

[0305] In a possible implementation, the target program is a shooting game program, the first initial object behavior is a shooting behavior, and the target behavior deviation includes a horizontal shooting direction deviation and / or a vertical shooting direction deviation.

[0306] In a possible implementation, the generating unit 1502 is specifically configured to:

[0307] Generate a second initial behavior control instruction corresponding to the target moment according to the target object state information and the target perception information, where the second initial behavior control instruction is used to control the target object to execute the target object behavior with a first behavior amplitude at the target moment;

[0308] Based on that the first behavior amplitude does not exceed the behavior amplitude threshold corresponding to the target object behavior, determine the second initial behavior control instruction as the target behavior control instruction, where the behavior amplitude threshold is the maximum amplitude for controlling the target object to execute the target object behavior through the target program interface, and the target program interface is used to control the target object;

[0309] Based on that the first behavior amplitude exceeds the behavior amplitude threshold, adjust the second initial behavior control instruction to obtain the target behavior control instruction, where the target behavior control instruction is used to control the target object to execute the target object behavior with the behavior amplitude threshold at the target moment.

[0310] In a possible implementation, the target object behavior is an object behavior for a target interaction object in the target program, and the adjusting unit 1503 is specifically configured to:

[0311] Based on that the time interval between the target moment and the object appearance moment reaches the reaction duration, control the target object to execute the target object behavior through the target behavior control instruction, where the object appearance moment is the moment when the target interaction object appears in the program interface corresponding to the target program, and the reaction duration is positively correlated with the perception difficulty corresponding to the target interaction object, and the perception difficulty is the difficulty of perceiving the target interaction object at the object appearance moment through the program interface;

[0312] Adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task performed by the target object through executing the target object behavior to obtain an instruction generation model.

[0313] In a possible implementation, the perception difficulty is positively correlated with the distance between the target interaction object and the target object at the moment when the object appears, and is positively correlated with the position difference corresponding to the display position of the target interaction object in the program interface at the moment when the object appears. The position difference is the difference between the display position and the center position of the program interface.

[0314] In a possible implementation, the target program is a shooting game program, and the target object behavior is to aim at the target interaction object.

[0315] The embodiment of the present application further provides a computer device. Please refer to Figure 16 As shown, this computer device may be a terminal device. Taking the terminal device as a mobile phone as an example:

[0316] Figure 16 The block diagram of a part of the structure of the mobile phone related to the terminal device provided by the embodiment of the present application is shown. Refer to Figure 16 , the mobile phone includes: a radio frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a wireless fidelity (WiFi) module 770, a processor 780, and a power supply 790, etc. Those skilled in the art can understand that Figure 16 the mobile phone structure shown in

[0317] does not constitute a limitation on the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange different components. Figure 16 The following specifically introduces each component of the mobile phone:

[0318] The RF circuit 710 can be used for receiving and sending information or signals during a call. Specifically, after receiving the downlink information from the base station, it is sent to the processor 780 for processing. Additionally, the uplink data designed is sent to the base station. Generally, the RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 710 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0319] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 720 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0320] The input unit 730 can be used to receive input numeric or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 730 can include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel 731), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 731 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 780, and can receive and execute the commands sent by the processor 780. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 731. In addition to the touch panel 731, the input unit 730 can also include other input devices 732. Specifically, the other input devices 732 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0321] The display unit 740 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 740 can include a display panel 741. Optionally, the display panel 741 can be configured in the form of a liquid crystal display (LCD for short), an organic light-emitting diode (OLED for short), etc. Further, the touch panel 731 can cover the display panel 741. When the touch panel 731 detects a touch operation thereon or nearby, it transmits it to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides corresponding visual output on the display panel 741 according to the type of touch event. Although in Figure 16 it, the touch panel 731 and the display panel 741 are implemented as two independent components to realize the input and output functions of the mobile phone, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.

[0322] The mobile phone may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 741 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.

[0323] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the mobile phone. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and then converted into audio data. After the audio data is output to the processor 780 for processing, it is sent through the RF circuit 710 to, for example, another mobile phone, or the audio data is output to the memory 720 for further processing.

[0324] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 770, which provides users with wireless broadband Internet access. Although Figure 16 the WiFi module 770 is shown, it can be understood that it does not belong to the essential composition of the mobile phone and can be completely omitted within the scope of not changing the essence of the invention according to needs.

[0325] The processor 780 is the control center of the mobile phone, connecting various parts of the entire mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 720, and by calling the data stored in the memory 720, it executes various functions of the mobile phone and processes data, thereby performing an overall detection of the mobile phone. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 780 either.

[0326] The mobile phone further includes a power supply 790 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 780 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.

[0327] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0328] In this embodiment, the processor 780 included in the terminal device further has the following functions:

[0329] Determine the target object state information and target perception information corresponding to the target object at the target moment, where the target object state information is used to represent the object state of the target object corresponding to the target program, and the perception information is used to simulate the information that can be perceived by the user of the target program by controlling the target object at the target moment;

[0330] Generate a target behavior control instruction corresponding to the target moment through an initial instruction generation model according to the target object state information and the target perception information, where the target behavior control instruction is used to control the target object to perform a target object behavior at the target moment;

[0331] Adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task by the target object through performing the target object behavior, to obtain an instruction generation model. The achievement degree determined by the instruction generation model reaches a target achievement degree. The achievement degree is used to represent the reasonable degree of the object behavior of the target object, and the reasonable degree represented by the target achievement degree is greater than a first preset threshold. The instruction generation model is used to determine a behavior control instruction for controlling the target object according to the object state information and perception information corresponding to the target object.

[0332] The embodiment of the present application further provides a server. Please refer to Figure 17 as shown Figure 17The structural diagram of server 800 provided by an embodiment of this application. Server 800 may vary significantly due to configuration or performance differences, and may include one or more central processing units (CPUs) 822 (for example, one or more processors) and a memory 832, and one or more storage media 830 (for example, one or more mass storage devices) that store application programs 842 or data 844. Among them, the memory 832 and the storage media 830 may be transient storage or persistent storage. The programs stored in the storage media 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 822 may be configured to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the server 800.

[0333] Server 800 may further include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0334] The steps performed by the server in the above embodiments may be based on Figure 17 the server structure shown.

[0335] An embodiment of this application also provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute any one of the model training methods described in the foregoing embodiments.

[0336] An embodiment of this application also provides a computer program product including a computer program, and when it runs on a computer device, it causes the computer device to execute the model training method described in any one of the above embodiments.

[0337] It can be understood that in the specific implementation of this application, when it comes to data related to user information (such as player data), etc., when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0338] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes of various kinds.

[0339] It should be noted that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0340] As described above, it is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that, The method includes: Determining the target object state information and target perception information corresponding to the target object at the target moment, where the target object state information is used to characterize the object state of the target object corresponding in the target program, and the perception information is used to simulate the information that the user of the target program can perceive by controlling the target object at the target moment; Generating, by an initial instruction generation model, a target behavior control instruction corresponding to the target moment according to the target object state information and the target perception information, where the target behavior control instruction is used to control the target object to perform a target object behavior at the target moment; Adjusting the model parameters of the initial instruction generation model according to the achievement degree of the behavior task by the target object performing the target object behavior, to obtain an instruction generation model, where the achievement degree determined by the instruction generation model reaches a target achievement degree. The achievement degree is used to characterize the rational degree of the object behavior of the target object, and the rational degree characterized by the target achievement degree is greater than a first preset threshold. The instruction generation model is used to determine a behavior control instruction for controlling the target object according to the object state information and perception information corresponding to the target object.

2. The method according to claim 1, wherein The target perception information includes scene perception information, where the scene perception information is used to characterize the position of the scene object that the program user can perceive at the target moment, and the scene object is used to constitute the target scene where the target object is located at the target moment.

3. The method according to claim 2, wherein The scene perception information is determined by the following method: Determining the ray origin of the detection ray according to the position where the target object is located at the target moment; Projecting the detection ray from the ray origin to the target scene where the target object is located, and determining the object point that the detection ray first passes through, where the object point is used to constitute the scene object; Determining the position information corresponding to the object point as the scene perception information.

4. The method according to claim 3, wherein The detection ray includes a horizontal detection ray. Projecting the detection ray from the ray origin to the target scene where the target object is located and determining the object point that the detection ray first passes through includes: Projecting the horizontal detection ray horizontally from the ray origin to the target scene, and determining the object point that the horizontal detection ray first passes through.

5. The method according to claim 4, wherein The horizontal detection ray includes a planar detection ray. Determining the ray origin of the detection ray according to the position where the target object is located at the target moment includes: Determining the camera plane corresponding to the target object according to the position where the target object is located at the target moment, where the camera plane is used to render the target program interface corresponding to the target moment of the target program. The target program interface is used to control the target object in the target program, and the pixel values of the pixel points included in the target program interface are determined based on the projection of the scene object onto the camera plane; Determining the ray origin located on the camera plane; Projecting the horizontal detection ray horizontally from the ray starting point towards the target scene and determining the object point first passed by the horizontal detection ray includes: Projecting the plane detection ray from the ray starting point towards the target scene in a first direction perpendicular to the camera plane, and determining the object point first passed by the plane detection ray, where the first direction is the direction of observing the target scene through the target program interface.

6. The method according to claim 4, characterized in that, The horizontal detection ray includes a surrounding detection ray. Determining the ray starting point of the detection ray according to the position of the target object at the target moment includes: According to the position of the target object at the target moment, determining the ray starting point located on the target object, where the ray starting point is the object point on the target object with a probability of interacting with the scene object greater than a second preset threshold; Projecting the horizontal detection ray horizontally from the ray starting point towards the target scene and determining the object point first passed by the horizontal detection ray includes: Projecting multiple surrounding detection rays horizontally from the ray starting point towards the target scene, and determining the object points first passed by the multiple surrounding detection rays respectively, where the ray directions corresponding to the multiple surrounding detection rays are different.

7. The method according to claim 3, wherein The detection ray includes a vertical detection ray. Determining the ray starting point of the detection ray according to the position of the target object at the target moment includes: According to the position of the target object at the target moment, determining the detection plane located above the target object, where the detection plane is a horizontal plane; Determining the ray starting point located on the detection plane; Projecting the detection ray from the ray starting point towards the target scene where the target object is located and determining the object point first passed by the detection ray includes: Projecting the vertical detection ray vertically from the ray starting point towards the target scene, and determining the object point first passed by the vertical detection ray.

8. The method according to claim 1, wherein The target perception information includes interaction object perception information, which is used to represent the position of the interaction object that can be perceived by the program user at the target moment. The interaction object is used to interact with the target object in the target scene by performing object behaviors. The target scene is the scene where the target object is located at the target moment, and the interaction object is not the scene object used to constitute the target scene.

9. The method according to claim 8, wherein The interaction object perception information is determined by the following method: Determining whether there is the scene object between the target object and the interaction object; Based on the absence of the scene object between the target object and the interaction object, determining the position information corresponding to the interaction object as the interaction object perception information, where the position information is used to identify the position of the interaction object in the target scene; Based on the existence of the scenario object between the target object and the interaction object, and the interaction object satisfying the perceivable condition, the area information corresponding to the interaction object is determined as the interaction object perception information. The area information is used to identify the position area where the interaction object is located in the target scenario. The position area includes multiple positions, and the multiple positions include the position where the interaction object is located in the target scenario.

10. The method according to claim 9, wherein The perceivable condition includes that the interaction object has sound effect information in a playing state at the target moment, and the target distance between the interaction object and the target object is less than the distance threshold; The number of positions included in the position area is inversely correlated with the volume of the sound effect information and inversely correlated with the target distance.

11. The method according to claim 9, wherein The perceivable condition includes that in the first historical period, there is a first historical moment satisfying that at the first historical moment, there is no such scenario object between the interaction object and the target object. The end moment of the first historical period is the target moment, and the length of the first historical period is less than the first length threshold; The number of positions included in the position area is positively correlated with the length of the first target period. The first target period is the period from the first historical moment to the target moment.

12. The method according to claim 9, characterized in that The perceivable condition includes that in the second historical period, there is a second historical moment satisfying that at the second historical moment, the perception information corresponding to the perception sharing object includes the target interaction object perception information corresponding to the interaction object. The perception sharing object is the object in the target program that shares perception information with the target object; The number of positions included in the position area is positively correlated with the length of the second target period and positively correlated with the number of positions identified by the target interaction object perception information. The second target period is the period from the second historical moment to the target moment.

13. The method according to claim 1, characterized in that, The target perception information further includes scene path information. The scene path information is used to identify the actionable path corresponding to the target object in the target scenario. The target scenario is the scenario where the target object is located at the target moment.

14. The method according to claim 13, wherein The scene path information is determined by the following method: Determine multiple reachable positions corresponding to the target scenario. The multiple reachable positions are the positions that the target scenario support object can reach; Generate initial path information corresponding to the target scenario according to the multiple reachable positions. The initial path information is used to identify multiple initial paths between the multiple reachable positions; Determine multiple key positions among the multiple reachable positions. When the program user controls the movement of the target object at the target moment, the probability of reaching the key positions is greater than the probability of reaching the reachable positions other than the key positions among the reachable positions; According to the initial path information and the multiple key positions, determine the scene path information corresponding to the target scenario. The actionable path is the initial path among the multiple initial paths that includes the key positions.

15. The method according to claim 1, characterized in that, Generating the target behavior control instruction corresponding to the target moment according to the target object state information and the target perception information includes: Generating a plurality of pending behavior control instructions according to the target object state information and the target perception information; Determining a plurality of sub-behavior control instructions corresponding to the target pending behavior control instruction, where the plurality of sub-behavior control instructions have respectively corresponding sub-object behaviors and control behaviors, and the sub-object behaviors corresponding to the plurality of sub-behavior control instructions are used to constitute the object behavior corresponding to the target pending behavior control instruction. The target sub-behavior control instruction corresponds to the target control behavior, and the target control behavior executed through the target program interface is used to generate the target sub-behavior control instruction. The target sub-behavior control instruction is any one of the plurality of sub-behavior control instructions, the target pending behavior control is any one of the plurality of pending behavior control instructions, and the target program interface is used to control the target object; Based on that the control behaviors corresponding to the plurality of sub-behavior control instructions are a plurality of control behaviors that can be executed in parallel through the target program interface, determining the target pending behavior control instruction as the target behavior control instruction.

16. The method according to claim 1, characterized in that Generating the target behavior control instruction corresponding to the target moment according to the target object state information and the target perception information includes: Generating a first initial behavior control instruction corresponding to the target moment according to the target object state information and the target perception information, where the first initial behavior control instruction is used to control the target object to perform a first initial object behavior at the target moment; Generating the target object control instruction according to the behavior deviation parameter and the first initial behavior control instruction, where the behavior deviation parameter corresponds to the target behavior deviation, and the target object behavior is the first initial object behavior that generates the target behavior deviation during execution.

17. A model training device, characterized in that, The device includes a determination unit, a generation unit, and an adjustment unit: The determination unit is used to determine the target object state information and the target perception information corresponding to the target object at the target moment. The target object state information is used to characterize the object state corresponding to the target object in the target program, and the perception information is used to simulate the information that the user of the target program can perceive by controlling the target object at the target moment; The generation unit is used to generate, through the initial instruction generation model, the target behavior control instruction corresponding to the target moment according to the target object state information and the target perception information, where the target behavior control instruction is used to control the target object to perform the target object behavior at the target moment; The adjustment unit is configured to adjust the model parameters of the initial instruction generation model according to the achievement degree of the behavior task achieved by the target object through performing the target object behavior, so as to obtain an instruction generation model. The achievement degree determined by the instruction generation model reaches a target achievement degree. The achievement degree is used to characterize the rationality of the object behavior of the target object. The rationality characterized by the target achievement degree is greater than a first preset threshold. The instruction generation model is configured to determine a behavior control instruction for controlling the target object according to the object state information and perception information corresponding to the target object.

18. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is configured to store a computer program and transmit the computer program to the processor; The processor is configured to execute the model training method according to any one of claims 1-16 based on the instructions in the computer program.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store a computer program, and the computer program is used to execute the model training method according to any one of claims 1-16.

20. A computer program product including a computer program, when running on a computer device, causes the computer device to execute the model training method according to any one of claims 1-16.