Data determination method, device, electronic device and computer storage medium
By obtaining the team value data of the target user's operational object team while executing and not executing team tasks, combining machine learning and non-monotonic evaluation functions, the flexibility and universality of team value evaluation are solved, and a more accurate team value allocation is achieved.
Patent Information
- Application Number
- CN202211060317.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The team value evaluation methods of team members in the existing technology lack flexibility and universality, and cannot accurately reflect the team contribution of team members.
By determining the target user's target operation object and other operating objects within the observable range to perform the target team tasks, the first team value data is obtained, and the second team value data is obtained when the target operation object does not participate in the target team task. Combined with the greedy marginal contribution idea, machine learning and non-monotonic evaluation functions are used to allocate team value.
There is no need to develop specific measurement standards, which improves the accuracy and universality of team value allocation.
Smart Images

Figure CN115624762B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and specifically to a data determination method, device, electronic device, and computer storage medium. Background Art
[0002] With the development of computer technology, more and more game products are entering the public's field of vision. As a result, people's requirements for the actual experience of game products are also constantly improving.
[0003] In team collaboration game products, after the team collaboration task is completed, the team value of the members who participated in the team collaboration task is often evaluated. However, in the existing technical field, the evaluation method of team members' team value is generally implemented through some established measurement standards. This evaluation method lacks flexibility and universality in distribution. More importantly, this evaluation method is limited by the interpretability of the measurement standards and cannot accurately link the team value of team members with their team contributions.
[0004] Therefore, how to conduct team evaluation on each member of the team has become a technical problem that technical personnel in this field urgently need to solve. Summary of the Invention
[0005] The present application provides a data determination method, device, electronic device and computer storage medium for accurately and flexibly determining the team value of team members in various games.
[0006] This application provides a data determination method, comprising:
[0007] Determining first team value data of an operation object team consisting of a target operation object of a target user and other operation objects within an observable range of the target user, when the operation object team performs a target team task;
[0008] Determining second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task;
[0009] The team value data of the target user to the operation object team through the target operation object is determined according to the first team value data and the second team value data.
[0010] This application also provides a data determination device, including:
[0011] a first determining unit, configured to determine first team value data of an operation object team, when an operation object team consisting of a target operation object of a target user and other operation objects within an observable range of the target user performs a target team task;
[0012] A second determining unit is configured to determine second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task;
[0013] A third determining unit is configured to determine, based on the first team value data and the second team value data, the team value of the target user to the operation object team through the target operation object.
[0014] The present application also provides an electronic device, including:
[0015] processor;
[0016] The memory is used to store a method program, and when the program is read and executed by the processor, the data determination method is executed.
[0017] The present application also provides a computer storage medium, which stores a computer program, and when the program is executed, the data determination method is implemented.
[0018] Compared with the prior art, this application has the following advantages:
[0019] The data determination method provided in the present application includes: determining the first team value data of the operation object team when the operation object team composed of the target operation object of the target user and other operation objects within the observable range of the target user performs the target team task; determining the second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task; and determining the team value data of the target user to the operation object team through the target operation object based on the first team value data and the second team value data.
[0020] This method combines the idea of greedy marginal contribution and allocates value to the target user by using an operation object team consisting of the target user's target operation object and other operation objects within the target user's observable range. The first team value data of the operation object team when the target team task is performed, and the second team value data when the target operation object does not participate in the operation object team and the operation object team fails to complete the target team task, thereby completing the value allocation of the target user. The team value allocation method provided in this application does not require the formulation of specific measurement standards, thereby improving the accuracy and universality of team value allocation. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flow chart of a data determination method provided in an embodiment of the present application;
[0022] Figure 2 A schematic diagram of an operation object team within the observable range of a target user provided by another embodiment of the present application;
[0023] Figure 3 A schematic diagram of the structure of a motion encoder provided in another embodiment of the present application;
[0024] Figure 4 A schematic diagram of a function curve of a data classification function provided in another embodiment of the present application;
[0025] Figure 5 A schematic diagram of training an environment forward prediction model provided in another embodiment of the present application;
[0026] Figure 6 A flowchart for obtaining team value data provided in another embodiment of the present application;
[0027] Figure 7 A schematic structural diagram of a data determination device provided in another embodiment of the present application;
[0028] Figure 8 A schematic structural diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0029] The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0030] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. Descriptive terms such as "a," "a," "a first," and "a second," etc., used in this application and the appended claims, are not intended to limit quantity or sequence, but are used to distinguish information of the same type from one another.
[0031] With the development of computer technology, more and more game products are entering the public's field of vision. As a result, people's requirements for the actual experience of game products are also constantly improving.
[0032] In team collaboration game products, after the team collaboration task is completed, the team value of the members who participated in the team collaboration task is often evaluated. However, in the existing technical field, the evaluation method of team members' team value is generally implemented through established measurement standards. This evaluation method lacks flexibility and universality in allocation. More importantly, this evaluation method is limited by the interpretability of the measurement standards and cannot accurately link the team value of team members with their team contributions.
[0033] The embodiment of the present application provides a data determination method that combines the idea of greedy marginal contribution and completes the value allocation to the target user by using the first team value data of the target operation object team when the target operation object of the target user and other operation objects within the target user's observable range perform the target team task, and the second team value data when the target operation object does not participate in the operation object team and the operation object team does not complete the target team task. The team value allocation method provided by the present application does not require the formulation of specific measurement standards, thereby improving the accuracy and universality of team value allocation.
[0034] The implementation subject of the method can be various types of user terminals such as laptops, tablet computers, desktop computers, set-top boxes, mobile devices (for example, mobile phones, personal digital messaging devices, game consoles), or a combination of any two or more of these data processing devices, or a server.
[0035] Please refer to Figure 1 , Figure 1 This is a flow chart of a data determination method provided in an embodiment of the present application, the method comprising the following steps S101 to S103:
[0036] Step S101 : determining first team value data of an operation object team composed of a target operation object of a target user and other operation objects within an observable range of the target user, when the operation object team performs a target team task.
[0037] In an optional embodiment of the present application, the target user can be understood as a game player, and the target operation object of the target user can be further understood as a virtual character or virtual object operated by the game player in the game. For example, in a multiplayer online tactical competitive game, the target operation object of the target user can be the corresponding virtual game character, and in some other types of game scenes, the target operation object can also be a virtual object such as an airplane or tank operated by the game player. This application does not impose any restrictions on this.
[0038] The first team value data can be understood as the overall value of the team when the target user operates the target operation object to participate in the team task.
[0039] The observable range can be understood as the game scene displayed in the graphical user interface of the target user when the target user operates the target operation object. This can be the entire game scene displayed in the graphical user interface, or a portion of the game scene displayed in the graphical user interface. For example, in a third-person game, the game scene displayed in the graphical user interface is the game scene observed from the "head" or "eye" direction of the target operation object, that is, the game scene displayed in the graphical user interface is the game scene within the field of view of the target operation object. In this case, the observable range of the target user is the entire game scene displayed in the graphical user interface. For another example, in a massively multiplayer online game, the game scene displayed in the graphical user interface is the game scene centered on the target operation object, that is, the game scene displayed in the graphical user interface includes both the game scene within the field of view of the target operation object and the game scene outside the field of view of the target operation object. In this case, the observable range of the target user can be the entire game scene displayed in the graphical user interface, the game scene within the field of view of the target operation object, or the portion of the game scene centered on the target operation object.
[0040] The specifics of the observable range can vary depending on the specific game situation. For example, if the game scene displayed within the graphical user interface (GUI) includes too many operational objects during the current game session, then to improve the accuracy of team value allocation, the observable range can be narrowed to the portion of the game scene centered on the target operational object. Another example: if the game scene displayed within the GUI includes too few operational objects during the current game session, then to improve the comprehensiveness of team value allocation, the observable range can be expanded to the entire game scene displayed within the GUI. Once the target user's observable range is determined, all operational objects within the observable range should be included in the operational object team. If the observable range is the game scene within the target operational object's field of view, then the operational objects included in the operational object team are the target operational object and other operational objects within the target operational object's field of view. If the observable range is the entire or partial game scene centered on the target operational object, then the operational objects included in the operational object team are the target operational object and other operational objects surrounding the target operational object.
[0041] For example, please refer to Figure 2 , Figure 2 A schematic diagram of a team of operating objects within the observable range of a target user is provided in accordance with another embodiment of the present application.
[0042] The operation object team includes: a first operation object 201, a second operation object 202, a third operation object 203, and a fourth object 204, wherein the first operation object 201 can be understood as the target operation object in step S102.
[0043] Further, Figure 2 A circular area centered on the first operation object 201 is also shown in FIG. 4 , and the circular area can be considered as the observable range of the target user.
[0044] Furthermore, the first operation object 201 , the second operation object 202 and the third operation object 203 are in the first camp, and the fourth object 204 is in the second camp. In the game scene, the first camp is opposite to the second camp.
[0045] In an optional embodiment of the present application, the above-mentioned first operation object 201 to the third operation object 203, and the fourth object 204 can all be regarded as virtual game characters operated by the game player. In this case, when the above-mentioned fourth object 204 is defeated by any one or more of the first operation object 201, the second operation object 202 and the third operation object 203, the built-in scoring system of the game can score the above-mentioned four objects based on the performance of the first operation object 201, the second operation object 202 and the third operation object 203 in the process of defeating the fourth object 204, and the performance of the fourth object 204 when fighting against the above-mentioned first to third operation objects (also called contribution distribution).
[0046] Please refer to Table 1, which shows a scoring system built into the game based on the method provided by this application. After the first to third operating objects jointly defeat the fourth object, the scoring of the above four objects is shown as follows:
[0047] Table 1:
[0048] First Camp First camp score Second Camp Second camp score First operation object +1 Fourth Object -1 Second operand +1 The third operation object +1
[0049] In other words, after the first through third operating subjects simultaneously attacked and defeated fourth subject 204, it can be considered that the first through third operating subjects simultaneously attacked fourth subject 204 during this confrontation, demonstrating a high degree of teamwork. Therefore, the first through third operating subjects' team contributions in this confrontation are all positive 1. However, since fourth subject 204 was significantly separated from its team in the second camp, its team contribution in this confrontation is considered negative 1.
[0050] It should be understood that the above description of the operational object teams and the scoring of each object within the teams is intended only to facilitate understanding of the application scenarios of the method described herein and is not intended to limit the scenarios or specific scoring methods. For example, the fourth object in this application could be a building within the game scene that is in opposition to the first camp. In this case, the fourth object is not controlled by a human and does not need to be scored. For another example, the first and second camps described above could be understood as two teams in a football match, where the first operational object is responsible for shooting, and the second and third operational objects are responsible for passing the ball to the first operational object. In this case, the second and third operational objects only score if the pass is successful, while the first operational object only scores if the shot is successful. Otherwise, points are deducted from the operational objects.
[0051] Furthermore, in order to obtain the first team value data in the above step S101, the method further includes: obtaining observation information of the target user within the observable range.
[0052] The observation information of the target user within the observable range can be understood as a set of features that the target user can discover in the current game scene, that is, the scene features of the target user within the observable range of the game scene.
[0053] In an optional implementation manner of the present application, the observation information may include object features of the target operation object and environmental features within the observable range of the target user.
[0054] Specifically, the object features of the target operation object may include: remaining health information of the game character operated by the target user, remaining skill cooldown time information of the game character, and game props information available to the game character.
[0055] In an optional implementation manner of the present application, the object features of the target operation object may be expressed in the form of a vector or a matrix to facilitate the subsequent output and transmission of the observation value.
[0056] Please refer to Table 2, which is a schematic table of object features that may appear in the game scene of the target operation object provided in the embodiment of the present application. In order to distinguish different object features, different user features are represented by F1 and F2 respectively:
[0057] Table 2:
[0058]
[0059] F1 in Table 2 can be understood as the object feature F1 of the target operation object = [100, 20, 0, 30, 4], that is, the remaining health of the target operation object at this time is 100, the cooling time of skill 1 is 20 seconds, the cooling time of skill 2 is 0 seconds, the cooling time of skill 3 is 30 seconds, and the remaining number of game props that can be used by the target object is 4; F2 in Table 2 can be understood as the object feature F2 of the target operation object = [70, 10, 5, 0, 10], that is, the remaining health of the target operation object at this time is 70, the cooling time of skill 1 is 10 seconds, the cooling time of skill 2 is 5 seconds, the cooling time of skill 3 is 0 seconds, and the remaining number of game props that can be used is 10.
[0060] It can be understood that the above introduction to the object characteristics of the target operation object provided in the embodiment of the present application is only for the purpose of facilitating the understanding of the present application, and is not used to limit the specific expression form of the object characteristics of the target operation object. The object characteristics of the target operation object may have different expression forms in different application scenarios, and the present application does not impose any restrictions on this.
[0061] Furthermore, the scene features within the observable range can be understood as the geographical location features of the target operation object in the game scene. For example, the geographical location is that the target operation object is in a dangerous area preset in the game, the target operation object is located in a mountain or a river in the game, etc., the distance between the target object and the attackable target, etc.
[0062] Please refer to Table 3, which shows two different forms of expression structures of the target operation object provided in an embodiment of the present application, wherein S1 and S2 represent two different scene features respectively.
[0063] Table 3:
[0064]
[0065]
[0066] S1 in Table 3 can be understood as the scene feature of the target operation object S1 = [D0, M1, WO], that is, the position of the target operation object at this time is not in the danger zone, but in the mountain structure and not in a lake or water; S2 in Table 3 can be understood as the scene feature of the target operation object S2 = [D1, M0, W1], that is, the position of the target operation object at this time is in the danger zone, not in the mountain structure, but in a lake or water.
[0067] It should be noted that the above explanation of the observation information of the target user in the target scene is only an optional implementation method of the observation information provided by this application. For different types of games, the observation information also has different composition methods. Therefore, the above explanation of the observation information is only used to understand the observation information, and is not used to limit the specific meaning of the observation information. For example, in some real-time strategy game scenarios, the target user may need to operate multiple different target operation objects, and achieve common game expectations through these target operation objects together with the operation objects of other users. At this time, the object characteristics of the target operation object may include the object characteristics of the multiple target operation objects of the target user; the scene characteristics within the observable range may be the environmental characteristics of the environment in which the multiple target operation objects are located.
[0068] Furthermore, the first team value data can be obtained by following steps S1 to S3:
[0069] Step S1, obtaining first action data of the target operation object;
[0070] The first action data may be understood as action data performed by the target user manipulating the target operation object. For example, the action data may include movement data, attack data, and the like.
[0071] Please refer to Table 4, which is a table of two sets of action data structures of target operation objects provided in an embodiment of the present application.
[0072] Table 4:
[0073] Motion data Mobile data Posture data Attack data A1 quit squat 0 A2 move clamber 1
[0074] A1 in Table 4 can be understood as the action data A1 of the target operation object = [quit, squat, 0], where quit indicates that the target operation object is in a stopped state, squat indicates that the target operation object is in a crouching state, and 0 indicates that the target operation object is not attacking, that is, A1 indicates that the target operation object is in a crouching and stationary state; A2 in Table 4 can be understood as the action data A2 of the target operation object = [move, clamber, 1], where move indicates that the target operation object is in a moving state, clamber indicates that the target operation object is in a climbing state, and 1 indicates that the target operation object is performing an attack action, that is, A2 indicates that the target operation object is in a climbing, moving and attacking state.
[0075] It should be noted that the above introduction to the action data provided in the embodiment of the present application is only for the purpose of facilitating the understanding of the meaning of the action data, and is not used to limit the specific form of the action data. For example, in some optional embodiments, the action data can also be expressed in the form of a digital vector or matrix, and this application does not impose any restrictions on this.
[0076] Step S2, obtaining an action data set of second action data of other operation objects within the observable range of the target user;
[0077] In an optional implementation of the present application, the second action data of the other operation objects is similar to the first action data, and can be understood as a data set of movement data, attack data, and other data of the operation objects of other users.
[0078] It can be understood that since the acquisition of the action data set of the second action data (i.e., the data set of the movement data and attack data of the operation object) needs to rely on the real-time collection of the actions of other operation objects within the observable range of the target user, and in game scenarios or other similar scoring scenarios, since the confrontation between opposing camps or players is random, the collection of action data generated within the observable range of the target operation object at all times will undoubtedly greatly reduce work efficiency and also cause data collection errors.
[0079] Therefore, in order to avoid the above problems, in another optional embodiment of the present application, in order to weaken the influence of other operation objects in the process of determining the target user's team value based on the impact of the target user's operation of the target operation object on the overall value of the team, the present application limits the second action data based on a centralized idea (that is, only relying on the collected observation values of the target user and the action data of the target operation object itself to complete the method of obtaining the team value).
[0080] In this manner, the action data set of the second action data is an action data set in which the action data of the other operation objects are set to the globally optimal action data by default.
[0081] In another optional embodiment of the present application, the action data of the other operation objects default to the globally optimal action data, which can be understood as follows: when the action data of the target operation object is guaranteed to be optimal, the action data corresponding to the other operation objects are the globally optimal action data. The globally optimal action data of the other operation objects can be determined by the first non-monotonic evaluation function described in step S3. Specifically, since the first non-monotonic evaluation function is a function that is fitted with the observation information of the target user, the first action data of the target operation object, and the action data set of the second action data of other operation objects within the observable range as input items, and the estimated team value data of the target user's operation object team when performing the task as the output item, then when the output item data is the largest, the action data set of the second action data of other operation objects within the corresponding observable range is the globally optimal action data of the other operation objects.
[0082] Step S3: Obtain the first team value data based on the action data set of the first action data and the second action data and the observation information in the target scene.
[0083] In an optional embodiment of the present application, the above-mentioned step S3 can be understood as bringing the action data set of the above-mentioned first action data, second action data and the observation information under the target scene into a preset first non-monotonic evaluation function, thereby obtaining the first team value data.
[0084] The first non-monotonic evaluation function can be understood as a utility function. In the above step S2, in order to determine the global optimal action data of the other operation objects, all possible action data sets of the other operation objects can be brought into the non-monotonic evaluation function in turn, and then the first team value data can be determined while determining the global optimal action data.
[0085] In another optional embodiment of the present application, the method for obtaining the first non-monotonicity evaluation function may include the following steps:
[0086] First, each user in the user set who participates in the target team task is taken as the target user, and the observation information of each target user, the first action data of the target operation object of each target user, and the action data set of the second action data of other operation objects within the observable range of each target user are obtained.
[0087] The user set may be understood as all game players who participate in a certain game at the same time. In this step, each user in the user set who participates in the target team task is defined as a target user.
[0088] The action data set of obtaining the observation information of each target user, the first action data of the target operation object of each target user, and the second action data of other operation objects within the observable range of each target user is actually obtaining training data samples for training the first non-monotonic evaluation function.
[0089] Second, the observation information of each target user, the first action data of the target operation object of each target user, and the action data set of the second action data of other operation objects within the observable range of each target user are taken as input data, and the first non-monotonic evaluation function to be trained is input to obtain the estimated first team value data of the operation object team of each target user when performing the target task output by the first non-monotonic evaluation function to be trained.
[0090] The first non-monotonic evaluation function is generated through repeated iterative training. The first non-monotonic evaluation function before or during training can be understood as the first non-monotonic evaluation function to be trained.
[0091] The acquired observation information, first action data, and second action data are combined as input data and fed into the first monotonicity evaluation function to be trained. Output data can be obtained, which can be understood as first team value data. However, since the first monotonicity evaluation function to be trained is still in the iterative training process, the first team value data it outputs is not accurate data but rather estimated data. This data is the estimated first team value data and is not used to evaluate the first team value, but only to optimize the first monotonicity evaluation function to be trained.
[0092] Third, the estimated first team value data of the operation object team of each target user when performing the target task is mixed to obtain the estimated first overall value data of the user set.
[0093] In an optional embodiment, the estimated first team value data of each target user's operation object team when performing the target task can be mixed and processed through a hybrid neural network, and the estimated first overall value data of the user set can be obtained after processing.
[0094] The overall value data can be understood as an overall score for all game players participating in a certain game, such as the total score, the total number of times a skill is used, etc.
[0095] The estimated first overall value data can be understood as: data obtained by mixing the estimated first team value data obtained by the first monotonic evaluation function to be trained. This data is not accurate data, but estimated data. It is not used to evaluate the overall value, but only used to optimize the training of the first monotonic evaluation function to be trained.
[0096] Fourth, obtain the real overall value data of the user set.
[0097] The real overall value data of the user set is the score automatically calculated by the game after the game ends, which can be understood as the true value of the overall value data.
[0098] Fifth, the estimated first overall value data of the user set is used to fit the real overall value data of the user set, and the first non-monotonic evaluation function to be trained is iteratively trained to obtain the first non-monotonic evaluation function.
[0099] With the goal of ensuring that the estimated first overall value data is as close as possible to the actual overall value data, the first non-monotonic evaluation function to be trained is iteratively trained. When the difference between the estimated first overall value data obtained by the first non-monotonic evaluation function to be trained and the actual overall value data is within a preset error range (e.g., ±0.5%), the first non-monotonic evaluation function to be trained is considered to have been trained and can be used to evaluate the first team's value.
[0100] In another optional embodiment of the present application, the first non-monotonic evaluation function to be trained is a function that is fitted with the action data set of the target user's observation information, the first action data of the target operation object, and the second action data of other operation objects within the observable range as input items, and the estimated team value data of the target user's operation object team when performing the task as the output item.
[0101] The first non-monotonic evaluation function can be obtained by iteratively training the first non-monotonic evaluation function to be trained. In practical applications, since the data dimension of the action data that can be taken by different operation objects is uncertain, and the computational complexity of the process of traversing all action data combinations to determine the first team value data is too large. Therefore, in an optional embodiment of the present application, different action data can be mapped to the potential category space through an action encoder to reduce the category dimension of the action data.
[0102] In the specific application process, this application uses machine learning (ML) training to obtain the action encoder. Machine learning (a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines) is a branch of artificial intelligence (AI) technology that specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning generally includes artificial neural networks, reset networks, reinforcement learning, transfer learning, inductive learning and other technologies.
[0103] Please refer to Figure 3 , Figure 3 A schematic structural diagram of a motion encoder provided in another embodiment of the present application.
[0104] like Figure 3 As shown, Figure 3 It includes: a first fully connected layer 301, a second fully connected layer 302, a third fully connected layer 303 and a self-attention network 304.
[0105] exist Figure 3 In the action encoder shown, O i represents the observation value of the target user in the target scene in step S101 above, U i An action data set representing the second action data in step S302 above;
[0106] In an embodiment of the present application, the first fully connected layer 301 is used to integrate the observation values of the target user in the target scene, and the second fully connected layer 302 is used to integrate the action data set of the second action data, so as to adjust the action data sets of the first action data and the second action data so that they can complete the classification while reducing the influence of each action data on the classification.
[0107] Furthermore, the self-attention network 303 is used to assign weights to the output data of the first fully connected layer 301 and the second fully connected layer 402 after receiving the output data, so as to classify the action data set of the second action data in combination with the observation value of the target user in the target scene, thereby reducing the data dimension of the action data.
[0108] Furthermore, the third fully connected layer 304 is used to classify the action data set of the second action data according to the output data of the self-attention network 303, and output the classification result Zi.
[0109] Furthermore, the classification of the action data set is specifically performed in conjunction with the corresponding data classification function. Figure 4 , Figure 4 A schematic diagram of a function curve of a data classification function provided in another embodiment of the present application.
[0110] The horizontal axis of the function curve of the data classification function represents different classification results, that is, Z1 represents classification result 1, Z2 represents classification result 2, and Z3 represents classification result 3.
[0111] The vertical axis of the function curve represents the probabilities corresponding to different classification results.
[0112] In the embodiment of the present application, the classification result corresponding to the highest point of the classification function is taken as the classification result of the action data set of the second action data, such as Figure 4 As shown, Figure 4 Shown is a function curve corresponding to the classification result Z2.
[0113] As described above, the action encoder is specifically used to classify the action data set of the second action data. Specifically, in order to obtain an accurate action encoder, the action encoder is trained together with an environment forward prediction model to optimize the action data compressed and output by the action encoder while compressing the action data set of the second action data.
[0114] Please refer to Figure 5 , Figure 5 A schematic diagram of training an environment forward prediction model provided in another embodiment of the present application.
[0115] The environment forward prediction model is a model for evaluating the target user based on the observation value of the target user and the action data of the operable object of the target user.
[0116] like Figure 5As shown, in the actual application process, the input data of the environmental forward prediction model 501 are Oi and Ai, that is, the observation value within the observable range of the target user and the first action data of the operable object of the target user mentioned in the embodiment of the present application; the output data of the environmental forward prediction model 501 is the score of the target user action data, denoted as Rt.
[0117] In the actual application process, in order to determine whether the action data output by the action encoder meets the requirements for the subsequent determination of the team value of the target user, this application uses the difference between a constant value data Q1 and the output data Rt of the environment forward prediction model as the loss value L of the action encoder. e , so that when the loss value is greater than or equal to a preset loss threshold, the internal parameters of the motion encoder are adjusted until the loss value is less than the preset loss threshold.
[0118] Furthermore, in order to facilitate understanding of this application Figure 3 The working principle of the action encoder is shown in the figure. The classification process of the action encoder is illustrated by examples in the following specific application scenarios.
[0119] Assume that the observation value of the target user 1 in the target scene is O1 = [100, 20, 0, 1], that is, the remaining blood volume of the target operation object K1 of the target user 1 is 100, the first skill cooling time of the target operation object is 20 seconds remaining, the second skill cooling time is 0 seconds remaining, and the target operation object is in the danger zone, that is, the object feature of the target operation object K1 is F1 = [100, 20, 0], and the scene feature of the target operation object is S1 = [D1].
[0120] Furthermore, the target user's observable range includes the first operation object, the second operation object, and a target object that needs to be attacked jointly by the first operation object and the second operation object to complete a team task;
[0121] The optional action data of the first operation object include: moving, attacking;
[0122] The optional action data of the second operation object also includes: move, attack;
[0123] At this time, if A10 indicates that the first operation object does not move, A11 indicates that the first operation object moves; A20 indicates that the first operation object does not attack, A21 indicates that the first operation object attacks; B10 indicates that the second operation object does not move, B11 indicates that the second operation object moves; B20 indicates that the second operation object does not attack, B21 indicates that the second operation object attacks, the optional second action data is as shown in Table 5 below.
[0124] Table 5:
[0125]
[0126] After determining the action data sets of the second action data of the different combinations, the action data sets may be sequentially input into the action encoder while inputting the observation value of the target user in the target scene.
[0127] After fitting the action data set and the observation value through the action encoder, corresponding weights are assigned to the observation value, attack action and movement action respectively. It can be understood that in the embodiment of the present application, the first operation object and the second operation object need to attack the target object together to complete the team task. Therefore, the weight assigned to the action of "attack" here should be higher than the weight of the action of "move" and the observation value.
[0128] After the action encoder assigns corresponding weights to the observation value and the action data set, the weighted action data set is classified through the self-attention network to determine the action category of the action data set.
[0129] It can be understood that when a higher weight is given to the action of "attack", when classifying the action data set, it is more inclined to classify the action combination in which the first operation object and the second operation object both have the action data of "attack" into the same category, that is, when the weight of the action of "attack" is higher than the preset threshold, the action set can be divided into two categories according to whether the first operation object and the second operation object in the action set both have the action data of "attack", that is, the action data set that has both A21 and B21 in Table 1 is divided into one category (which can be recorded as Z1), and the other action data sets are divided into another category (which can be recorded as Z0).
[0130] It can be seen from this that the reason why this application classifies the action data set of the second action data mentioned in the above step S3 is to weaken the action categories of different action data, reduce the data dimensions of the action data, and then classify the second action data, thereby reducing the complexity of determining the value of the first team and improving processing efficiency.
[0131] Step S102: determining the second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task.
[0132] Similar to step S101 of the present application, the second team value data can be obtained through the following steps S4 and S5:
[0133] Step S4, obtaining third action data when the target operation object does not participate in the target team task;
[0134] Step S5: Obtain the second team value data based on the third action data and the observation information.
[0135] Different from the first action data described in step S1 of the present application, the third action data is not the action data generated by the target user operating the target operation object to perform a certain action, but is based on the assumption that the target operation object has not participated in the operation object team. The target operation object may generate an action. It can be understood that when the target operation object does not participate in the operation object team, the overall value of the target operation object to the team task can be ignored. Therefore, when the target operation object does not participate in the operation object team and acts alone, the specific nature of the third action data has no effect on the value of the second team.
[0136] In an optional implementation of the present application, the above-mentioned step S5 can be understood as bringing the above-mentioned third action data and the observation information under the target scene into a preset second non-monotonic evaluation function to obtain the second team value data.
[0137] The second non-monotonic evaluation function, like the first non-monotonic evaluation function described in step S3 of the present application, can also be understood as a utility function. The difference between the second non-monotonic evaluation function and the first non-monotonic evaluation function is that the input item of the second non-monotonic evaluation function does not include the action data set of the second action data of other operation objects within the observable range of the target user.
[0138] Specifically, the method for obtaining the second non-monotonicity evaluation function may include the following steps:
[0139] First, each user in the user set who has not participated in the target team task is taken as the target user, and observation information of each target user and third action data of the target operation object of each target user are obtained.
[0140] The user set can be understood as all game players who participate in a certain game at the same time. In this step, each user in the user set who does not participate in the target team task is defined as a target user.
[0141] The target user is not participating in the target team task. Therefore, the target user does not have any other operation objects within the observable range, nor does he or she have any action data sets for the second action data of other operation objects. Therefore, the training data samples for training the second non-monotonic evaluation function only include the observation information of each target user and the third action data of each target user's target operation object.
[0142] Second, the observation information of each target user and the third action data of the target operation object of each target user are obtained as input data, and the second non-monotonic evaluation function to be trained is input to obtain the estimated second team value data of the operation object team of each target user when performing the target task output by the second non-monotonic evaluation function to be trained.
[0143] The second non-monotonic evaluation function is generated through repeated iterative training. The second non-monotonic evaluation function before or during training can be understood as the second non-monotonic evaluation function to be trained.
[0144] The acquired observation information and third action data are input into the second monotonicity evaluation function to be trained, resulting in output data, which can be understood as the second team value data. However, since the second monotonicity evaluation function to be trained is still in the iterative training process, the second team value data it outputs is not accurate data, but rather an estimate. This data is the estimated second team value data and is not used to evaluate the second team value, but only to optimize the second monotonicity evaluation function to be trained.
[0145] Third, the estimated second team value data of the operation object team of each target user when performing the target task is mixed to obtain the estimated second overall value data of the user set.
[0146] In an optional implementation, a hybrid neural network can be used to combine the estimated second team value data for each target user's target team when performing the target task. This data can be used to generate the estimated second overall value data for the user set. This data is not accurate data, but rather an estimate. It is not used to evaluate the overall value, but only to optimize the second monotonic evaluation function to be trained.
[0147] Fourth, obtain the real overall value data of the user set.
[0148] The real overall value data of the user set is the score automatically calculated by the game after the game ends, which can be understood as the true value of the overall value data.
[0149] Fifth, the estimated second overall value data of the user set is used to fit the real overall value data of the user set, and the second non-monotonic evaluation function to be trained is iteratively trained to obtain the second non-monotonic evaluation function.
[0150] The second non-monotonic evaluation function to be trained is iteratively trained with the goal of ensuring that the estimated second overall value data is as close as possible to the true overall value data. When the difference between the estimated second overall value data obtained by the second non-monotonic evaluation function to be trained and the true overall value data is within a preset error range (e.g., ±0.5%), the second non-monotonic evaluation function to be trained is considered to have been trained and can be used to evaluate the second team value.
[0151] In another optional embodiment of the present application, the second non-monotonic evaluation function to be trained is a function fitted with the observation information of the target user and the third action data of the target operation object as input items, and the estimated team value data of the target user's operation object team when performing the task as output items.
[0152] The second non-monotonic evaluation function can be obtained by iteratively training the second non-monotonic evaluation function to be trained.
[0153] In another optional embodiment of the present application, the training of the first non-monotonic evaluation function in step S3 and the training of the second non-monotonic evaluation function in step S5 can be combined into the training of the non-monotonic evaluation function. The non-monotonic evaluation function can be understood as a utility function comprising two parts for calculating the overall value of the team. When training the non-monotonic evaluation function, the relevant data of the users in the user set who participated in the target team task can be input into the first part, and the relevant data of the users in the user set who did not participate in the target team task can be input into the second part. The non-monotonic evaluation function can obtain the first team value data based on the action data set of the first action data, the second action data and the observation information, and can also obtain the second team value data based on the third action data and the observation information.
[0154] Step S103: determining the team value data of the target user to the operation object team through the target operation object according to the first team value data and the second team value data.
[0155] In the embodiment of the present application, the above step S103 determines the team value data of the target operation object to the operation object team by subtracting the second team value data from the first team value data to obtain the team value data of the target user to the operation team through the target operation object, that is, if the first team value data is recorded as Q coll, let the second team value data be Q alone , the team value data of the target operation object to the operation object team is recorded as Q p , then Q p =Q coll -Q alone .
[0156] Further, in order to facilitate understanding of the method for obtaining team value provided in the embodiment of the present application, the following is combined with Figure 6 This method is further elaborated.
[0157] Please refer to Figure 6 , Figure 6 A flowchart for obtaining team value data is provided in another embodiment of the present application.
[0158] Figure 6 It includes an action encoder 601, a first MLP network 602, and a second MLP network 603.
[0159] The input data of the action encoder 601 is U1, that is, U1 is the global optimal action data of other operation objects within the observable range of the target user mentioned in step S2 of this application;
[0160] In this embodiment, it is assumed that after U1 is classified based on the action encoder 601, the classification result obtained is Z1, that is, other operating objects within the observation range of the target user jointly attack the same target within the visual range.
[0161] The input data of the first MLP network 602 are Z1, O1 and A1, where O1 is the observation value of the target user in the target scenario mentioned in step S101 of this application; A1 is the first action data of the target operation object in step S1 of this application, that is, the action data occurring in the target operation object.
[0162] The above data is processed by the first MLP network 602 to obtain the team value Q coll , that is, the first team value data mentioned in step S3 of this application.
[0163] The input data of the second MLP network 603 are O1 and A2, where O1 is the same as O1 input to the first MLP network 602, that is, the observation value of the target user in the target scenario; A2 is the third action data of the target operation object mentioned in step S4 of this application when it does not participate in the target team task.
[0164] The above data is processed by the second MLP network 604 to obtain the team value Q alone , that is, the second team value data mentioned in step S5 of this application.
[0165] Finally, through Q coll -Q alone The subtraction operation is performed to obtain the team value data Q of the target operation object of the target user to the operation team. p .
[0166] In summary, the method for obtaining the team value data provided in the embodiment of the present application, combined with the concept of greedy marginal contribution, completes the value distribution to the target user by using the operation object team composed of the target operation object of the target user and other operation objects within the observable range of the target user, the first team value data of the operation object team when the target team task is performed, and the second team value data when the target operation object does not participate in the operation object team and the operation object team does not complete the target team task. The team value distribution method provided in the present application does not require the formulation of specific measurement standards, thereby improving the accuracy and universality of team value distribution.
[0167] Another embodiment of the present application provides a data determination device. Since the device embodiment is basically similar to the above-mentioned method embodiment, the description is relatively simple. For relevant points, please refer to the corresponding description of the above-mentioned method embodiment. The following description of the device embodiment is merely illustrative.
[0168] Please refer to Figure 7 , Figure 7 A schematic diagram of the structure of a data determination device provided in another embodiment of the present application, the device comprising:
[0169] A first determining unit 701 is configured to determine first team value data of an operation object team, when the operation object team, which is composed of a target operation object of a target user and other operation objects within an observable range of the target user, performs a target team task;
[0170] A second determining unit 702 is configured to determine second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task;
[0171] The third determining unit 703 is configured to determine the team value data of the target user to the operation object team through the target operation object according to the first team value data and the second team value data.
[0172] In an optional embodiment, the apparatus further comprises: a first obtaining unit;
[0173] The first obtaining unit is configured to obtain observation information of the target user within the observable range, where the observation information includes object features of the target operation object of the target user and environmental features within the observable range of the target user.
[0174] Specifically, the first obtaining unit is basically the same as the method for obtaining the observation information of the target user within the observable range mentioned in step S101 of the data determination method embodiment. For relevant matters, please refer to the description of step S101 in the method embodiment of this application, and no further details will be given here.
[0175] In an optional embodiment, the object characteristics of the target operation object include at least one of the following information: usable prop information of the target operation object, skill cooling time information of the target operation object, and remaining health information of the target operation object.
[0176] In an optional implementation, the environmental features within the observable range of the target user include: regional information of the game scene where the target operation object is located.
[0177] In an optional implementation, the first determining unit 701 includes:
[0178] A first data obtaining unit, configured to obtain first action data of the target operation object;
[0179] A second data obtaining unit, configured to obtain an action data set of second action data of other operation objects within an observable range of the target user;
[0180] The first value data obtaining unit is configured to obtain the first team value data according to the first action data, the action data set of the second action data, and the observation information.
[0181] Specifically, the process by which the first data acquisition unit obtains the first action data is substantially the same as step S1 in the method embodiment provided herein, the process by which the second data acquisition unit obtains the action data set of the second action data is substantially the same as step S2 in the method embodiment provided herein, and the process by which the first value data acquisition unit obtains the first team value data is substantially the same as step S3 in the method embodiment provided herein. For relevant details, reference may be made to the description of steps S1 to S3 in the method embodiment of this application, which will not be repeated here.
[0182] In an optional implementation, obtaining the first action data of the target object includes:
[0183] In response to the operation instruction information sent by the target user to the target operation object, first action data of the target operation object corresponding to the operation instruction information is obtained.
[0184] In an optional implementation, the action data set of the second action data includes:
[0185] The action data set of the other operation objects when the action data of the other operation objects is optimal is used by default.
[0186] In an optional embodiment, before the step of obtaining the first team value data based on the first action data, the action data set of the second action data, and the observation information, the method further includes:
[0187] Inputting the motion data set of the second motion data into a preset motion encoder to obtain fourth motion data output by the preset motion encoder;
[0188] The action encoder is used to classify the action data set of the second action data based on a preset action category.
[0189] In an optional implementation, obtaining the first team value data according to the first action data, the action data set of the second action data, and the observation information includes:
[0190] The first action data, the fourth action data and the observation information are input into a preset first non-monotonic evaluation function to obtain the first team value data output by the first non-monotonic evaluation function.
[0191] In an optional embodiment, the method for obtaining the first non-monotonic evaluation function includes:
[0192] Taking each user in the user set who participates in the target team task as the target user, obtaining the observation information of each target user, the first action data of the target operation object of each target user, and the action data set of the second action data of other operation objects within the observable range of each target user;
[0193] Using the acquired observation information of each target user, the first action data of the target operation object of each target user, and the action data set of the second action data of other operation objects within the observable range of each target user as input data, inputting the first non-monotonic evaluation function to be trained, and obtaining the estimated first team value data of the operation object team of each target user when performing the target task output by the first non-monotonic evaluation function to be trained;
[0194] Mixing the estimated first team value data of the operation object team of each target user when performing the target task to obtain the estimated first overall value data of the user set;
[0195] Obtaining real overall value data of the user set;
[0196] The estimated first overall value data of the user set is used to fit the real overall value data of the user set, and the first non-monotonic evaluation function to be trained is iteratively trained to obtain the first non-monotonic evaluation function.
[0197] In an optional embodiment, the first non-monotonic evaluation function to be trained is a function that is fitted with the action data set of the target user's observation information, the first action data of the target operation object, and the second action data of other operation objects within the observable range as input items, and the estimated team value data of the target user's operation object team when performing the task as the output item.
[0198] In an optional implementation, the second determining unit 702 includes:
[0199] A third data obtaining unit is configured to obtain third action data when the target operation object does not participate in the target team task;
[0200] The second value data obtaining unit is used to obtain the second team value data according to the third action data and the observation information.
[0201] Specifically, the process of the third action data obtaining unit obtaining the third action data is basically the same as step S4 in the embodiment of the method of the present application, and the process of the second value data obtaining unit obtaining the second team value data is basically the same as step S5 in the embodiment of the method of the present application. For relevant matters, please refer to the relevant description of step S4 and step S5 in the embodiment of the method of the present application, and no further details will be given here.
[0202] In an optional implementation, obtaining the third action data when the target operation object does not participate in the target team task includes:
[0203] determining an action data set that can be executed by the target operation object when the target operation object does not participate in the target team task;
[0204] At least one action data is selected from the action data set executable by the target operation object as the third action data.
[0205] In an optional implementation, obtaining the second team value data based on the third action data and the observation information includes:
[0206] The third action data and the observation information are input into a preset second non-monotonic evaluation function to obtain second team value data output by the second non-monotonic evaluation function.
[0207] In an optional embodiment, the method for obtaining the second non-monotonic evaluation function includes:
[0208] Taking each user in the user set who has not participated in the target team task as the target user, obtaining observation information of each target user and third action data of the target operation object of each target user;
[0209] Using the obtained observation information of each target user and the third action data of the target operation object of each target user as input data, inputting the second non-monotonic evaluation function to be trained, and obtaining estimated second team value data of the operation object team of each target user when performing the target task output by the second non-monotonic evaluation function to be trained;
[0210] Mixing the estimated second team value data of the operation object team of each target user when performing the target task to obtain the estimated second overall value data of the user set;
[0211] Obtaining real overall value data of the user set;
[0212] The estimated second overall value data of the user set is used to fit the real overall value data of the user set, and the second non-monotonic evaluation function to be trained is iteratively trained to obtain the second non-monotonic evaluation function.
[0213] In an optional embodiment, the second non-monotonic function to be trained is a function that is fitted with the observation information of the target user and the third action data of the target operation object as input items, and the estimated team value data of the target user's operation object team when performing the task as output items.
[0214] In an optional embodiment, determining, based on the first team value data and the second team value data, the team value data of the target user to the operation object team through the target operation object includes:
[0215] A difference between the first team value data and the second team value data is obtained, and the difference is used as the team value data of the target user on the operation object team through the target operation object.
[0216] In an optional embodiment, the observable range includes: a game scene displayed in a graphical user interface of the target user.
[0217] Another embodiment of the present application also provides an electronic device, please refer to Figure 8 , Figure 8 A schematic structural diagram of an electronic device provided in another embodiment of the present application.
[0218] The electronic device includes:
[0219] At least one processor 801, at least one memory 803, at least one communication interface 802, and at least one communication bus 804;
[0220] Optionally, the communication interface 802 may be an interface of a communication module, such as an interface of a GSM module;
[0221] The processor 801 may be a CPU, or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application;
[0222] The memory 803 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0223] Among them, the memory 803 is a program for storing the method, and when the program is read and executed by the processor, the method provided by the method embodiment of the present application is executed.
[0224] A sixth embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program, and when the program is executed, the method provided in the above method embodiment is implemented.
[0225] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0226] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0227] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0228] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0229] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as systems or electronic devices. Therefore, the present application may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A data determination method, characterized in that: include: Determining first team value data of an operation object team consisting of a target operation object of a target user and other operation objects within an observable range of the target user, when the operation object team performs a target team task; Determining second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task; The team value data of the target user to the operation object team through the target operation object is determined according to the first team value data and the second team value data.
2. The method according to claim 1, characterized in that The method further comprises: Observation information of the target user within the observable range is obtained, where the observation information includes object features of a target operation object of the target user and environmental features within the observable range of the target user.
3. The method according to claim 2, characterized in that The object feature of the target operation object includes at least one of the following information: usable item information of the target operation object, skill cooldown time information of the target operation object, and remaining health information of the target operation object.
4. The method according to claim 2, characterized in that The environmental features within the observable range of the target user include: regional information of the game scene where the target operation object is located.
5. The method according to claim 2, characterized in that In the case of determining that an operation object team consisting of a target operation object of a target user and other operation objects within an observable range of the target user performs a target team task, the first team value data of the operation object team includes: Obtaining first action data of the target operation object; Obtain an action data set of second action data of other operation objects within an observable range of the target user; The first team value data is obtained according to the first action data, the action data set of the second action data and the observation information.
6. The method according to claim 5, characterized in that The obtaining of the first action data of the target object includes: In response to the operation instruction information sent by the target user to the target operation object, first action data of the target operation object corresponding to the operation instruction information is obtained.
7. The method according to claim 5, characterized in that The action data set of the second action data includes: The action data set of the other operation objects under the condition that the action data of the other operation objects are optimal is used by default.
8. The method according to claim 5, characterized in that Before the step of obtaining the first team value data based on the first action data, the action data set of the second action data, and the observation information, the method further includes: Inputting the motion data set of the second motion data into a preset motion encoder to obtain fourth motion data output by the preset motion encoder; The action encoder is used to classify the action data set of the second action data based on a preset action category.
9. The method according to claim 8, characterized in that The obtaining of the first team value data according to the first action data, the action data set of the second action data, and the observation information includes: The first action data, the fourth action data and the observation information are input into a preset first non-monotonic evaluation function to obtain the first team value data output by the first non-monotonic evaluation function.
10. The method according to claim 9, characterized in that The method for obtaining the first non-monotonic evaluation function includes: Taking each user in the user set who participates in the target team task as the target user, obtaining the observation information of each target user, the first action data of the target operation object of each target user, and the action data set of the second action data of other operation objects within the observable range of each target user; Using the acquired observation information of each target user, the first action data of the target operation object of each target user, and the action data set of the second action data of other operation objects within the observable range of each target user as input data, inputting the first non-monotonic evaluation function to be trained, and obtaining the estimated first team value data of the operation object team of each target user when performing the target task output by the first non-monotonic evaluation function to be trained; Mixing the estimated first team value data of the operation object team of each target user when performing the target task to obtain the estimated first overall value data of the user set; Obtaining real overall value data of the user set; The estimated first overall value data of the user set is used to fit the real overall value data of the user set, and the first non-monotonic evaluation function to be trained is iteratively trained to obtain the first non-monotonic evaluation function.
11. The method according to claim 10, characterized in that The first non-monotonic evaluation function to be trained is a function that is fitted with the action data set of the target user's observation information, the first action data of the target operation object, and the second action data of other operation objects within the observable range as input items, and the estimated team value data of the target user's operation object team when performing the task as the output item.
12. The method according to claim 2, characterized in that The determining of the second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task includes: Obtaining third action data when the target operation object does not participate in the target team task; The second team value data is obtained based on the third action data and the observation information.
13. The method according to claim 12, characterized in that The obtaining of the third action data when the target operation object does not participate in the target team task includes: determining an action data set that can be executed by the target operation object when the target operation object does not participate in the target team task; At least one action data is selected from the action data set executable by the target operation object as the third action data.
14. The method according to claim 12, characterized in that The obtaining of the second team value data according to the third action data and the observation information includes: The third action data and the observation information are input into a preset second non-monotonic evaluation function to obtain second team value data output by the second non-monotonic evaluation function.
15. The method according to claim 14, characterized in that The method for obtaining the second non-monotonic evaluation function includes: Taking each user in the user set who has not participated in the target team task as the target user, obtaining observation information of each target user and third action data of the target operation object of each target user; Using the obtained observation information of each target user and the third action data of the target operation object of each target user as input data, inputting the second non-monotonic evaluation function to be trained, and obtaining estimated second team value data of the operation object team of each target user when performing the target task output by the second non-monotonic evaluation function to be trained; Mixing the estimated second team value data of the operation object team of each target user when performing the target task to obtain the estimated second overall value data of the user set; Obtaining real overall value data of the user set; The estimated second overall value data of the user set is used to fit the real overall value data of the user set, and the second non-monotonic evaluation function to be trained is iteratively trained to obtain the second non-monotonic evaluation function.
16. The method according to claim 15, characterized in that The second non-monotonic evaluation function to be trained is a function that is fitted with the observation information of the target user and the third action data of the target operation object as input items, and the estimated team value data of the target user's operation object team when performing the task as output items.
17. The method according to claim 1, wherein The determining, based on the first team value data and the second team value data, the team value data of the target user to the operation object team through the target operation object includes: A difference between the first team value data and the second team value data is obtained, and the difference is used as the team value data of the target user on the operation object team through the target operation object.
18. The method according to any one of claims 1 to 17, characterized in that: The observable range includes: the game scene displayed in the graphical user interface of the target user.
19. A data determination device, characterized in that: include: a first determining unit, configured to determine first team value data of an operation object team, when an operation object team consisting of a target operation object of a target user and other operation objects within an observable range of the target user performs a target team task; A second determining unit is configured to determine second team value data of the operation object team when the target operation object does not participate in the target team task and the other operation objects do not complete the target team task; A third determining unit is configured to determine, based on the first team value data and the second team value data, the team value data of the target user to the operation object team through the target operation object.
20. An electronic device, characterized in that: include: processor; A memory for storing a method program, wherein when the program is read and executed by the processor, the method according to any one of claims 1 to 18 is executed.
21. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the program is executed, the method according to any one of claims 1 to 18 is implemented.
Citation Information
Patent Citations
Interactive task control method and device, intelligent decision model training method and device and medium
CN111589166A
Multi-agent reinforcement learning method and system based on sparse attention mechanism
CN112949856A