Business data processing method, apparatus, device, and storage medium

CN121570816BActive Publication Date: 2026-09-22HANGZHOU BULLET FINGER UNIVERSE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511725970.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-09-22
Estimated Expiration
2045-11-21

AI Technical Summary

Technical Problem

[0005]本公开提供一种业务数据处理方法、装置、设备及存储介质,以至少解决相关技术中难以准确筛选出优质对战阵容的问题

Benefits of technology

本公开获取预设业务对应的随机阵容;将所述随机阵容输入阵容强化模型进行阵容强化调整,得到初始阵容;其中,所述阵容强化模型为采用样本业务的样本随机阵容对预设模型进行调整得到样本调整阵容,以及基于所述样本随机阵容与所述样本调整阵容之间的对战结果对所述预设模型进行训练得到;基于所述初始阵容以及所述阵容强化模型,构建所述随机阵容对应的阵容序列;所述阵容序列包括至少两个按照确定时间排序的对抗阵容,从而通过阵容强化模型快速构建得到一系列的优化阵容;所述阵容序列中排在首位的对抗阵容为所述初始阵容,所述阵容序列中排序靠后的对抗阵容基于对排序靠前的对抗阵容进行调整得到;将所述阵容序列中排序靠后的预设数量个对抗阵容确定为头部阵容,从而保证了提取的头部阵容的对战能力强于其他阵容。本公开实现了快速、准确地构建对战能力强大的阵容。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121570816B_ABST
    Figure CN121570816B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a business data processing method, device, equipment and storage medium. The method comprises: obtaining a random lineup corresponding to a preset business; inputting the random lineup into a lineup strengthening model for lineup strengthening adjustment to obtain an initial lineup; constructing a lineup sequence corresponding to the random lineup based on the initial lineup and the lineup strengthening model; the lineup sequence comprises at least two confrontation lineups sorted according to a determined time; determining a preset number of confrontation lineups at the back of the lineup sequence as head lineups; wherein the lineup strengthening model is obtained by adjusting a preset model using a sample random lineup of a sample business to obtain a sample adjusted lineup, and training the preset model based on the battle result between the sample random lineup and the sample adjusted lineup. The present disclosure realizes rapid and accurate construction of a lineup with strong confrontation ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a business data processing method, apparatus, device, and storage medium. Background Technology

[0002] In the field of video game development, game balance testing is a crucial step in ensuring game playability and long-term operational success. With the diversification of game genres and the increasing complexity of game mechanics, traditional balance testing methods face numerous challenges. Currently, the mainstream testing methods mainly include three categories: manual testing, A / B testing, and automated simulation testing.

[0003] Manual testing methods rely primarily on the experience and judgment of game designers, assessing balance by manually adjusting parameters and simulating gameplay. While intuitive, this method is inefficient and struggles to cover all possible strategy combinations in the game. A / B testing adjusts game balance by collecting behavioral data from real players; while reliable, it is time-consuming and can negatively impact player experience. Automated simulation testing uses statistical simulation methods, improving efficiency, but lacks adaptability to the specific mechanics of different game types.

[0004] The generation methods for test cases for related technologies are relatively fixed, making it difficult to fully cover all possible strategies in games. Secondly, the testing system lacks adaptability to the special mechanisms of different game types. Thirdly, there is a trade-off between testing accuracy and computational efficiency, making it difficult to control computational costs while ensuring test quality. Especially when facing complex games with multiple characters, skills, and resource systems, it is difficult to complete comprehensive testing with limited computing power, thus making it difficult to accurately select high-quality battle lineups. Summary of the Invention

[0005] This disclosure provides a business data processing method, apparatus, device, and storage medium to at least solve the problem of accurately selecting high-quality match lineups in related technologies. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a business data processing method is provided, comprising: Get the random lineup corresponding to the preset service; The random lineup is input into the lineup enhancement model for lineup enhancement and adjustment to obtain the initial lineup; the lineup enhancement model is obtained by adjusting the preset model with the sample random lineup of the sample business to obtain the sample adjusted lineup, and by training the preset model based on the battle results between the sample random lineup and the sample adjusted lineup. Based on the initial lineup and the lineup enhancement model, a lineup sequence corresponding to the random lineup is constructed; the lineup sequence includes at least two adversarial lineups ordered according to a certain time; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier; A predetermined number of opposing lineups that are ranked later in the lineup sequence are identified as the top lineups.

[0006] In one exemplary implementation, the random lineup is at least two, and the step of constructing the lineup sequence corresponding to the random lineup based on the initial lineup and the lineup enhancement model includes: Input the initial lineup corresponding to each random lineup into the lineup enhancement model to enhance and adjust the lineup, and output the current adjusted lineup corresponding to each random lineup. Input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model to enhance and adjust the lineup, and output the current updated lineup corresponding to each random lineup. The current updated lineup sequence corresponding to each random lineup is used as the current adjusted lineup corresponding to each random lineup, and the current adjusted lineup corresponding to each random lineup is repeatedly input into the lineup enhancement model for lineup enhancement and adjustment, and the current updated lineup corresponding to each random lineup is output until the preset conditions are met. Based on the initial lineup, the currently adjusted lineup, and the currently updated lineup corresponding to each random lineup, a lineup sequence corresponding to each random lineup is constructed.

[0007] In an exemplary embodiment, the steps of reusing the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup, repeatedly inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and outputting the current updated lineup corresponding to each random lineup until a preset condition is met include: For each random lineup, determine the cumulative number of lineups including the initial lineup, the currently adjusted lineup, and the currently updated lineup; If the cumulative number of lineups does not reach the target number, the current updated lineup sequence corresponding to each random lineup is used as the current adjusted lineup corresponding to each random lineup again, and the process of inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment is repeated until the current updated lineup corresponding to each random lineup is output until the preset conditions are met.

[0008] In one exemplary implementation, determining a preset number of opposing lineups that rank lower in the lineup sequence as top-tier lineups includes: Get the preset number of opposing lineups that are sorted at the end of the lineup sequence corresponding to each random lineup, and obtain the enhanced lineup corresponding to each random lineup. The enhanced lineup corresponding to each random lineup is determined as the top lineup.

[0009] In one exemplary embodiment, the training method for the lineup enhancement model includes: Obtain the random sample lineup of the sample service; The sample random lineup is input into the preset model for lineup adjustment to obtain the sample adjusted lineup; The sample random lineup and the sample adjusted lineup are used to simulate a battle in the simulated scenario of the sample business, and the sample battle results are obtained. The reward data corresponding to the preset model is calculated based on the sample battle results; if the sample battle results indicate that the sample adjusted lineup wins, the reward data is a positive number; if the sample battle results indicate that the sample random lineup wins, the reward data is a negative number. The model parameters of the preset model are adjusted according to the reward data until the training end condition is met, and the preset model at the end of training is determined as the lineup enhancement model.

[0010] In one exemplary implementation, adjusting the model parameters of the preset model based on the reward data until the training termination condition is met includes: Obtain the action sequence during the lineup adjustment process; If the reward data is positive, adjust the model parameters of the preset model to increase the probability of using the action sequence until the training termination condition is met; If the reward data is negative, adjust the model parameters of the preset model to reduce the probability of using the action sequence until the training termination condition is met.

[0011] In one exemplary implementation, calculating the reward data corresponding to the preset model based on the sample battle results includes: If there are at least two sample battle results, determine the average of the at least two sample battle results to obtain the sample average battle result; The difference between the adjusted lineup and the random lineup in the sample is determined based on the average battle results of the sample. The reward data is determined based on the lineup difference results; the lineup difference results are positively correlated with the absolute value of the reward data.

[0012] In one exemplary embodiment, the step of inputting the random sample lineup into the preset model for lineup adjustment to obtain the adjusted sample lineup includes: Obtain the set of sample fields from the random sample lineup; The sample fields in the sample field set are converted into fixed-length sample vectors using one-hot encoding. The individual sample vectors are concatenated to obtain a concatenated sample vector, which is then input into the preset model for lineup adjustment to obtain the adjusted sample lineup.

[0013] In one exemplary embodiment, the step of inputting the sample concatenation vector into the preset model for lineup adjustment to obtain the adjusted sample lineup includes: Obtain the number of fields in the sample fields set, and determine the number of output heads in the neural network of the preset model based on the number of fields; the number of output heads is equal to the number of fields. The sample concatenation vector is input into the preset model, and a multi-head masking mechanism is used to process one sample vector in the sample concatenation vector based on each output head to obtain the sample output result corresponding to each sample vector; The sample adjustment lineup is obtained based on the sample output results corresponding to each sample vector.

[0014] In one exemplary embodiment, after determining a preset number of opposing lineups that rank lower in the lineup sequence as top lineups, the method further includes: Obtain the lineup attribute information corresponding to each top lineup; Based on the lineup attribute information corresponding to each top lineup, a top lineup curve is drawn and visualized.

[0015] In one exemplary implementation, the sample random lineup is the lineup of the sample service under the current version, the lineup enhancement model is the model corresponding to the current version, and the method further includes: If the sample service has an updated version, obtain the updated random lineup of samples under the updated version; The lineup enhancement model is trained based on the updated random lineup of the sample to obtain the updated lineup enhancement model; the updated lineup enhancement model is used to adjust the random lineup under the updated version.

[0016] According to a second aspect of the present disclosure, a business data processing apparatus is provided, comprising: The random lineup acquisition module is configured to acquire a random lineup corresponding to a preset business. The initial lineup acquisition module is configured to input the random lineup into the lineup enhancement model for lineup enhancement and adjustment to obtain the initial lineup; the lineup enhancement model is obtained by adjusting the preset model with the sample random lineup of the sample business to obtain the sample adjusted lineup, and by training the preset model based on the battle results between the sample random lineup and the sample adjusted lineup. The lineup sequence construction module is configured to construct a lineup sequence corresponding to the random lineup based on the initial lineup and the lineup enhancement model; the lineup sequence includes at least two adversarial lineups ordered according to a determined time; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier; The top lineup determination module is configured to determine a preset number of opposing lineups that are ranked later in the lineup sequence as the top lineup.

[0017] In one exemplary implementation, the random lineup is at least two, and the lineup sequence construction module includes: The current lineup adjustment unit is configured to input the initial lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement adjustment, and output the current adjusted lineup corresponding to each random lineup; The current lineup update unit is configured to input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model to perform lineup enhancement and adjustment, and output the current updated lineup corresponding to each random lineup. The repeated adjustment unit is configured to perform the following steps: re-enable the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup, repeatedly input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement adjustment, and output the current updated lineup corresponding to each random lineup until the preset conditions are met. The sequence construction unit is configured to construct a lineup sequence corresponding to each random lineup based on the initial lineup, the currently adjusted lineup, and the currently updated lineup corresponding to each random lineup.

[0018] In one exemplary embodiment, the repetition adjustment unit includes: The initial lineup determination subunit is configured to perform, for each random lineup, the cumulative number of lineups: the initial lineup, the current adjusted lineup, and the current updated lineup; The lineup adjustment subunit is configured to perform the following steps if the cumulative number of lineups does not reach the target number: re-use the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup, and repeatedly input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and output the current updated lineup corresponding to each random lineup until the preset conditions are met.

[0019] In one exemplary implementation, the head lineup determination module includes: The enhanced lineup determination unit is configured to retrieve a preset number of opposing lineups from the lineup sequence corresponding to each random lineup, and obtain the enhanced lineup corresponding to each random lineup. The head lineup determination unit is configured to determine the enhanced lineup corresponding to each random lineup as the head lineup.

[0020] In one exemplary embodiment, the apparatus further includes: The sample lineup acquisition module is configured to execute the sample random lineup acquisition service; The sample lineup adjustment module is configured to input the random lineup of the sample into the preset model to adjust the lineup and obtain the adjusted lineup of the sample. The sample result determination module is configured to perform a simulated battle between the sample random lineup and the sample adjusted lineup in the simulated scenario of the sample business, and obtain the sample battle result. The reward data determination module is configured to calculate the reward data corresponding to the preset model based on the sample battle results; if the sample battle results indicate that the sample adjusted lineup wins, the reward data is a positive number; if the sample battle results indicate that the sample random lineup wins, the reward data is a negative number. The model training module is configured to adjust the model parameters of the preset model according to the reward data until the training termination condition is met, and to determine the preset model at the end of training as the lineup enhancement model.

[0021] In one exemplary embodiment, the model training module includes: The action sequence acquisition unit is configured to acquire the action sequence during the lineup adjustment process; The first adjustment unit is configured to perform the following actions: if the reward data is positive, adjust the model parameters of the preset model to increase the probability of using the action sequence until the training termination condition is met. The second adjustment unit is configured to adjust the model parameters of the preset model to reduce the probability of using the action sequence until the training termination condition is met if the reward data is negative.

[0022] In one exemplary embodiment, the reward data determination module includes: The sample average result determination unit is configured to perform the following: if there are at least two sample battle results, determine the average of the at least two sample battle results to obtain the sample average battle result; The sample difference determination unit is configured to determine the lineup difference result between the sample adjusted lineup and the sample random lineup based on the average battle result of the sample; The reward data determination unit is configured to determine the reward data based on the lineup difference result; the lineup difference result is positively correlated with the absolute value of the reward data.

[0023] In one exemplary embodiment, the sample lineup adjustment module includes: The sample field acquisition unit is configured to acquire the set of sample fields in the random sample lineup. The sample vector conversion unit is configured to perform one-hot encoding of the sample fields in the sample field set and convert them into a fixed-length sample vector. The sample lineup adjustment unit is configured to concatenate the individual sample vectors to obtain a concatenated sample vector, and then input the concatenated sample vector into the preset model for lineup adjustment to obtain the adjusted sample lineup.

[0024] In one exemplary embodiment, the sample lineup adjustment unit includes: The output head determination subunit is configured to perform the following operations: obtain the number of fields in the sample fields set, and determine the number of output heads in the neural network of the preset model based on the number of fields; the number of output heads is equal to the number of fields. The sample result determination subunit is configured to input the sample concatenation vector into the preset model, and use a multi-head masking mechanism to process one sample vector in the sample concatenation vector based on each output head to obtain the sample output result corresponding to each sample vector; The sample adjustment subunit is configured to execute the sample output results based on each sample vector to obtain the sample adjustment lineup.

[0025] In one exemplary embodiment, the apparatus further includes: The lineup attribute acquisition module is configured to retrieve the lineup attribute information corresponding to each top lineup. The display module is configured to draw a curve chart of the top lineups based on the lineup attribute information corresponding to each top lineup and display it visually.

[0026] In one exemplary embodiment, the sample random lineup is the lineup of the sample service under the current version, the lineup enhancement model is the model corresponding to the current version, and the device further includes: The updated lineup acquisition module is configured to acquire the sample updated random lineup under the updated version if the sample service has an updated version. The updated model training module is configured to perform lineup enhancement training on the lineup enhancement model based on the updated random lineup of the sample, so as to obtain an updated lineup enhancement model; the updated lineup enhancement model is used to perform lineup enhancement adjustments on the random lineup under the updated version.

[0027] According to a third aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the business data processing method described above.

[0028] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by an electronic device processor, enables the electronic device to perform the business data processing method as described above.

[0029] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the business data processing method described above.

[0030] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: This disclosure obtains a random lineup corresponding to a preset service; inputs the random lineup into a lineup enhancement model for lineup enhancement and adjustment to obtain an initial lineup; wherein, the lineup enhancement model is obtained by adjusting the preset model with a sample random lineup of a sample service to obtain a sample adjusted lineup, and by training the preset model based on the battle results between the sample random lineup and the sample adjusted lineup; based on the initial lineup and the lineup enhancement model, a lineup sequence corresponding to the random lineup is constructed; the lineup sequence includes at least two adversarial lineups ordered according to a certain time, thereby quickly constructing a series of optimized lineups through the lineup enhancement model; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier; a preset number of adversarial lineups ranked later in the lineup sequence are determined as the top lineups, thereby ensuring that the extracted top lineups have stronger combat capabilities than other lineups. This disclosure achieves the rapid and accurate construction of lineups with strong combat capabilities.

[0031] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0033] Figure 1 This is an application environment diagram illustrating a business data processing method according to an exemplary embodiment.

[0034] Figure 2 This is a flowchart illustrating a business data processing method according to an exemplary embodiment.

[0035] Figure 3 This is a schematic diagram illustrating a training method for a lineup enhancement model according to an exemplary embodiment.

[0036] Figure 4 This is a schematic diagram illustrating the cumulative reward during a model training process according to an exemplary embodiment.

[0037] Figure 5 This is a schematic diagram illustrating the structure of a training system for a lineup enhancement model according to an exemplary embodiment.

[0038] Figure 6 This is a flowchart illustrating a method for constructing a lineup sequence corresponding to the random lineup based on the initial lineup and the lineup enhancement model, according to an exemplary embodiment.

[0039] Figure 7 This is a flowchart illustrating a method, according to an exemplary embodiment, of repeatedly inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and outputting the current updated lineup corresponding to each random lineup until a preset condition is met.

[0040] Figure 8 This is a block diagram illustrating a business data processing apparatus according to an exemplary embodiment.

[0041] Figure 9 This is a block diagram illustrating a server according to an exemplary embodiment.

[0042] Figure 10 This is a block diagram illustrating an electronic device for business data processing according to an exemplary embodiment. Detailed Implementation

[0043] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0044] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0045] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure require user authorization or full authorization from all parties when the embodiments of this disclosure are applied to specific products or technologies. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0046] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained below: Game Balance Testing (GBT) is an evaluation method for game design. It verifies the fairness of game characters, skills, economic systems, and other elements through methods such as simulating player behavior, data analysis, and adversarial experiments, ensuring that different strategies or gameplay styles have reasonable competitiveness.

[0047] Genetic Algorithm (GA) is an optimization algorithm that simulates the natural evolutionary process. It iteratively generates solutions through operations such as selection, crossover, and mutation, and is suitable for global search of complex problems and parameter optimization in machine learning.

[0048] Heuristic search (HS) is a search strategy based on experience or intuition. It guides the search direction through an evaluation function, significantly reducing computational load, and is often used in scenarios such as path planning and artificial intelligence decision-making.

[0049] Currently, the main technical solutions in the field of game balance testing include the following: (1) Manual testing method This method relies on the experience of game designers or testing teams, assessing balance by manually adjusting game parameters (such as character attributes, skill strength, resource output, etc.) and conducting simulated battles. Testers typically design specific battle scenarios, observe the performance of different strategies or character combinations, and make adjustments based on subjective judgment or simple data analysis. This method is usually optimized in conjunction with player feedback from the game's beta testing phase.

[0050] (2) A / B test A / B testing assesses game balance by dividing players into different groups, each experiencing different game parameter configurations (such as skill damage, economy system, etc.), and collecting real player behavior data (such as win rate, usage rate, game duration, etc.). This method relies on big data analysis to determine the optimal parameter configuration by comparing the performance of players in different groups. It is typically used for post-launch balance adjustments, such as hero strength optimization in MOBA games.

[0051] (3) Automated simulation testing Using rule-driven simulators, a large number of battle simulations are conducted in a virtual environment to statistically analyze indicators such as win rate and strength of different strategies or character combinations. For example, Monte Carlo simulation generates a large amount of battle data through random sampling to evaluate balance.

[0052] Manual testing methods are highly subjective: test results heavily rely on the experience of the planners, making it difficult to guarantee objectivity and consistency. They also suffer from low coverage: due to limited manpower, it's impossible to exhaustively list all possible strategy combinations or parameter configurations, easily overlooking extreme cases. Furthermore, they are inefficient: manual adjustments and testing are time-consuming, making them unsuitable for the demands of rapid iterative development.

[0053] A / B testing has several drawbacks: It has a long cycle, requiring a large amount of player data to draw reliable conclusions, making it unsuitable for rapid validation during the development phase. It can also negatively impact player experience, as unbalanced test parameters may lead to poor experiences for some players and even trigger negative feedback. Furthermore, it has limited applicability, only applicable to games already released and unsuitable for effective testing in the early stages of development.

[0054] Automated simulation testing suffers from poor generalization ability: existing automated simulation testing methods struggle to adapt to the specific mechanics of different game genres. High computational cost: high-precision simulations (such as deep reinforcement learning) require substantial computational resources, making them difficult to run efficiently in typical development environments.

[0055] This embodiment constructs an agent capable of generating adversarial lineups in real time through model training. It discovers globally advantageous lineups through continuous self-play, designs an incremental learning mechanism, and utilizes a historical policy network to quickly adapt to new versions. It also combines a policy network with a visualization analysis module to output interpretable schematic diagrams.

[0056] Please see Figure 1 The diagram illustrates an application environment for a business data processing method according to an exemplary embodiment. The application environment may include a server 01 and a client 02.

[0057] Specifically, in the embodiments of this specification, server 01 may include a standalone server, a distributed server, or a server cluster composed of multiple servers. It may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 01 may include network communication units, processors, and memory, etc. Specifically, server 01 can train a lineup enhancement model, input a random lineup into the lineup enhancement model for lineup enhancement and adjustment, and obtain an initial lineup; based on the initial lineup and the lineup enhancement model, construct a lineup sequence corresponding to the random lineup; and determine a preset number of adversarial lineups ranked later in the lineup sequence as the head lineup; and send the head lineup to client 02.

[0058] Specifically, in the embodiments of this specification, the client 02 may include physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, smart wearable devices, and in-vehicle terminals, and may also include software running on the physical device, such as web pages provided to users by some service providers, or applications provided to users by these service providers. Specifically, the client 02 can be used to display the top lineup.

[0059] Figure 2 This is a flowchart illustrating a business data processing method according to an exemplary embodiment, such as... Figure 2 As shown, this method can be applied to Figure 1 The server 01 shown includes the following steps.

[0060] In step S201, a random lineup corresponding to a preset service is obtained.

[0061] In this embodiment of the disclosure, the preset service can be a game to be tested for balance. The preset service can include multiple service attribute types, and the random lineup is determined based on at least some of the multiple service attribute types. The service attribute types can include, but are not limited to, virtual character types, skill types, etc. in the game. For example, the service attribute types can include unit types, hero types, etc. in the game. Each service attribute type can include at least two service attribute information. For example, if the service attribute type is a unit, its corresponding service attribute information can include archers, infantry, cavalry, etc. In different games, virtual character types can include different service attribute information. The random lineup also includes the service attribute information corresponding to the service attribute types.

[0062] In step S203, the random lineup is input into the lineup enhancement model for lineup enhancement and adjustment to obtain the initial lineup; the lineup enhancement model is obtained by adjusting the preset model with the sample random lineup of the sample business to obtain the sample adjusted lineup, and by training the preset model based on the battle results between the sample random lineup and the sample adjusted lineup.

[0063] In this embodiment, a model can be pre-trained to obtain a lineup enhancement model. For example, a sample random lineup from a sample service can be used to adjust the preset model to obtain a sample adjusted lineup. Then, based on the battle results between the sample random lineup and the sample adjusted lineup, the preset model is trained to obtain the lineup enhancement model. The lineup enhancement model is used to adjust the input lineup and output an output lineup capable of defeating the input lineup. The sample service and the preset service can be of the same type, and the sample random lineup and the random lineup of the preset service can be of the same type. The random lineup can be input into the lineup enhancement model for lineup enhancement adjustment to obtain an initial lineup; in a scenario where the initial lineup battles against the random lineup, the initial lineup has a higher probability of winning than the random lineup.

[0064] The preset model can be a reinforcement learning model or other types of models. Reinforcement learning (RL) is a machine learning method whose basic framework is the Markov decision process. It allows an agent to learn the optimal policy through trial and error in its interaction with the environment. The agent performs actions in the environment and receives feedback, i.e., rewards, based on the results of the actions. For example, the preset model can adopt the PPO algorithm, which is a reinforcement learning algorithm based on policy gradients.

[0065] In step S205, based on the initial lineup and the lineup enhancement model, a lineup sequence corresponding to the random lineup is constructed; the lineup sequence includes at least two adversarial lineups ordered according to a determined time; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier.

[0066] In this embodiment, the initial lineup can be input into the lineup enhancement model to obtain an output lineup capable of defeating the initial lineup; then, the output lineup can be re-input into the lineup enhancement model to obtain a new enhanced lineup, and so on, thereby obtaining multiple opposing lineups ordered sequentially by time, with later opposing lineups being obtained by adjusting earlier opposing lineups. The number of lineups in the lineup sequence is at least two, and this number can be set and adjusted according to actual needs, without being specifically limited here.

[0067] In step S207, a preset number of opposing lineups that are ranked later in the lineup sequence are determined as the top lineups.

[0068] In this embodiment, since the later a lineup appears in the lineup sequence, its combat capability is stronger, a preset number of opposing lineups ranked later in the lineup sequence can be determined as the top lineups. The preset number can be set according to actual circumstances. The top lineups selected in this embodiment can be used for recommendation in preset business scenarios.

[0069] This disclosure provides an embodiment for obtaining a random lineup corresponding to a preset service; inputting the random lineup into a lineup enhancement model for lineup enhancement and adjustment to obtain an initial lineup; wherein, the lineup enhancement model is obtained by adjusting the preset model using a sample random lineup of a sample service to obtain a sample adjusted lineup, and by training the preset model based on the battle results between the sample random lineup and the sample adjusted lineup; based on the initial lineup and the lineup enhancement model, a lineup sequence corresponding to the random lineup is constructed; the lineup sequence includes at least two adversarial lineups ordered according to a predetermined time, thereby quickly constructing a series of optimized lineups through the lineup enhancement model; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier; a preset number of adversarial lineups ranked later in the lineup sequence are determined as the top lineups, thereby ensuring that the extracted top lineups have stronger combat capabilities than other lineups. This disclosure achieves the rapid and accurate construction of lineups with strong combat capabilities.

[0070] In some embodiments, after obtaining multiple top lineups, the multiple top lineups can be grouped into pairs and a balance test can be performed to select the optimal target lineup and recommend it to the user.

[0071] In some embodiments, analogous to a professional esports coach, after seeing the opponent's lineup, runes, and other choices, the coach can quickly devise a counter-strategy based on their understanding of the champion pool. This counter-strategy is more effective in terms of skill and synergy against the opponent's lineup. Given these two lineups, if an AI with identical behavior trees were to play against the player, the latter lineup would most likely be stronger and have a higher win rate.

[0072] This embodiment accomplishes a similar task using reinforcement learning artificial intelligence (AI). Given a random lineup, the AI ​​is required to output a counter-lineup. The quality of the counter-lineup depends on the battle results of these two lineups in a combat simulator. The AI's goal is to eliminate as many enemy soldiers as possible (considering absolute values) using the given counter-lineup. After training based on reinforcement learning, this AI model, for a given input lineup A, can consistently output a stronger lineup B with a high probability. Based on this, a lineup enhancement model can be obtained through reinforcement learning training.

[0073] In some embodiments, such as Figure 3 As shown, Figure 3 A schematic diagram illustrating a training method for a lineup enhancement model, including: S301: Obtain the random sample lineup of the sample service; S303: Input the random sample lineup into the preset model to adjust the lineup, and obtain the adjusted sample lineup; S305: Obtain the sample random lineup and the sample adjusted lineup to simulate a battle in the simulated scenario of the sample business, and obtain the sample battle result; S307: Calculate the reward data corresponding to the preset model based on the sample battle results; if the sample battle results indicate that the sample adjusted lineup wins, the reward data is a positive number; if the sample battle results indicate that the sample random lineup wins, the reward data is a negative number. S309: Adjust the model parameters of the preset model according to the reward data until the training end condition is met, and determine the preset model at the end of training as the lineup enhancement model.

[0074] In this embodiment, the sample service and the preset service can be of the same type, and the sample random lineup and the preset service's random lineup are of the same type. The sample battle results can include at least one battle data point; in the game scenario, the battle data can include, but is not limited to, real-time data used to calculate the battle win rate, battle survival rate, and battle attrition ratio; at least two of the battle win rate, battle survival rate, and battle attrition ratio can be calculated from the battle data as evaluation index data, thereby calculating the reward data corresponding to the preset model based on the evaluation index data. If the sample battle result indicates that the sample adjusted lineup wins, the reward data is positive; if the sample battle result indicates that the sample random lineup wins, the reward data is negative; the absolute value of the reward data can be determined based on the magnitude of the evaluation index data; the two are positively correlated. Then, the model parameters of the preset model can be adjusted based on the reward data until the training termination condition is met, and the preset model at the end of training is determined as the lineup enhancement model.

[0075] In some embodiments, the step of inputting the random sample lineup into the preset model for lineup adjustment to obtain the adjusted sample lineup includes: Obtain the set of sample fields from the random sample lineup; The sample fields in the sample field set are converted into fixed-length sample vectors using one-hot encoding. The individual sample vectors are concatenated to obtain a concatenated sample vector, which is then input into the preset model for lineup adjustment to obtain the adjusted sample lineup.

[0076] In this embodiment, the random sample lineup may include multiple sample fields to construct a sample field set, which represents the state space of the random sample lineup. Then, the sample fields in the sample field set are one-hot encoded into fixed-length sample vectors so that the preset model can process the sample vectors. The sample vectors are then concatenated to obtain a concatenated sample vector, which is then input into the preset model for lineup adjustment to obtain the adjusted sample lineup. The adjusted sample lineup can also be composed of fields, which represent the action space of the adjusted sample lineup. Thus, the adjusted sample lineup can be quickly generated using the sample fields in the random sample lineup.

[0077] In some embodiments, the step of inputting the sample concatenation vector into the preset model for lineup adjustment to obtain the adjusted sample lineup includes: Obtain the number of fields in the sample fields set, and determine the number of output heads in the neural network of the preset model based on the number of fields; the number of output heads is equal to the number of fields. The sample concatenation vector is input into the preset model, and a multi-head masking mechanism is used to process one sample vector in the sample concatenation vector based on each output head to obtain the sample output result corresponding to each sample vector; The sample adjustment lineup is obtained based on the sample output results corresponding to each sample vector.

[0078] In this embodiment of the disclosure, the preset model may include multiple output heads for processing each sample vector in the sample concatenation vector, and a multi-head masking mechanism may be used for data processing, with the masking of the later output head depending on the result of the previous output head; the sample output result may be a sample output field, thereby obtaining a sample adjustment lineup constructed from multiple sample output fields.

[0079] In some embodiments, within a game scenario, a lineup includes several heroes, each possessing several skills. Heroes are arranged in front and back rows, and the order of their skills has an impact, creating a cyclical counter-system. The state space represents the opponent's lineup observed by the model, i.e., the model's output. Specific lineup details are as follows: General level, star rating, and equipment skill level are not considered at this time. (Generals are all level 50 and tier 5; battle tactics are all level 10 and tier 5.) We are not considering situations where the opponent's forces have fewer than 3 members and are not equipped with all their skills.

[0080] The requirement is that the IDs of the three opposing generals must be different.

[0081] The requirement is that the skill IDs of the six opponents must be different.

[0082] Default positioning: General 1 is the front row general, General 2 is the back row general, and General 3 is the back row general.

[0083] Its corresponding state space size is:

[0084] In this embodiment of the disclosure, as shown in Table 1, multiple sample fields are displayed. A random sample lineup is equivalent to the set of all fields in Table 1. The vectors transformed from each sample field are concatenated in sequence and used as the input of the neural network of the preset model.

[0085] Table 1

[0086] In this embodiment of the disclosure, the action space represents the lineup that the AI ​​proposes to play after observing the opponent's lineup.

[0087] The specific lineup details are as follows: General level, star rating, and equipment skill level are not considered at this time. (Generals are all level 50 and tier 5; battle tactics are all level 10 and tier 5.) We will not consider situations where the deployed troops have fewer than 3 members or do not have all their skills equipped.

[0088] The requirement is that the IDs of the three generals in battle must be different from each other, but they can be the same as the ID of the general in the current state (opponent).

[0089] The requirement is that the six skill IDs used in battle must be different from each other, but they can be the same as the skill IDs of the opponent in the current state.

[0090] Default positioning: General 1 is the front row general, General 2 is the back row general, and General 3 is the back row general.

[0091] Its corresponding motion space size is:

[0092] In this embodiment, the neural network has 10 output heads for action output, corresponding sequentially to each field in Table 2. For example, the value range of the first head is [0, 1, 2, 3], corresponding to the four types of units. Furthermore, a multi-head mask mechanism is implemented in the output to support the game's requirements for unit lineups. The "multi-head mask mechanism" refers to applying the same mask to each attention head during the multi-head attention calculation process. For example, if the game requires that three generals cannot be repeated, and General 1's output head selects General 'a' from 84 generals, then the multi-head mask mechanism allows General 2's output head to mask the already selected General 'a' (e.g., selecting General 'b'), and General 3's output head to mask both selected Generals 'a' and 'b'. This progressive, ordered mechanism, where the masking of each subsequent output head depends on the result of the previous output head, is called multi-head masking.

[0093] Table 2

[0094] The reward design for the preset model is as follows: Given the opponent's lineup (random lineup, lineup A) as follows:

[0095] in, This refers to the collection of generals in lineup A. This refers to three different general characters in lineup A. This represents the set of skills used in lineup A. This indicates that lineup A has multiple different skills in its skill set. This represents the set of soldiers used in lineup A. This indicates the unit type of the soldiers in lineup A.

[0096] The strategy lineup for dealing with its output is as follows:

[0097] in, This refers to the collection of generals in lineup B. This refers to three different general characters within the generals of lineup B. This represents the set of skills used in lineup B. This indicates that lineup B has multiple different skills in its skill set. This represents the set of soldiers used in lineup B. This indicates the unit type of the soldiers in lineup B. In this example, the unit type is the same as that of the soldiers in lineup A.

[0098] Considering that game developers often focus on "win rate" and "unit losses," the reward can be defined as:

[0099] In the design of rewards, This refers to the average troop loss of lineups A and B over 100 battles. Here, 30,000 is used as the normalization parameter, which can be understood as both sides having a maximum troop strength of 30,000. This reward will dynamically change from large to small, and from positive to negative, depending on whether the battle is a complete victory, a narrow victory, a draw, a close defeat, or a complete defeat.

[0100] The AI ​​generates a random opponent lineup A, outputs a counter lineup B, and collects battle data. The data includes: the win rate of lineup B in this battle (the battle simulator at this point includes a version that includes draws, making it difficult to win; this has been optimized in later versions), the difference in casualties (how many more troops lineup A lost compared to lineup B), and the generals, skills, and troop types of lineup B.

[0101] For example, such as Figure 4 As shown, Figure 4 This diagram illustrates the cumulative reward during model training. The horizontal axis represents the size of the training data, and the vertical axis represents the cumulative reward data. As the training process progresses, the AI's cumulative reward increases, meaning that the AI ​​model (team lineup enhancement model) outputting lineup B is more likely to defeat the given lineup A by a significant margin.

[0102] After obtaining the AI ​​model described above, several random lineups are used as seeds. The AI ​​is continuously fed in, and the results are fed back in, similar to a "left foot stepping on right foot" process, to obtain a lineup sequence. After repeatedly trying this process and accumulating multiple different lineup sequences, the stronger parts of each sequence (generally lineups from later sequences) are selected to form the top lineup. While there is no transitivity in strength between lineups (A lineup > B lineup, B lineup > C lineup, but this does not imply A lineup > C lineup), there are "generally strong" lineups that are strong against most other lineups (e.g., those with higher numerical values ​​or synergistic mechanics).

[0103] Finally, by using the visualization module to draw the top lineups, you can obtain a series of lineups that are "generally strong".

[0104] In some embodiments, calculating the reward data corresponding to the preset model based on the sample battle results includes: If there are at least two sample battle results, determine the average of the at least two sample battle results to obtain the sample average battle result; The difference between the adjusted lineup and the random lineup in the sample is determined based on the average battle results of the sample. The reward data is determined based on the lineup difference results; the lineup difference results are positively correlated with the absolute value of the reward data.

[0105] In this embodiment, at least two battles can be conducted in a sample service scenario using a random sample lineup and a modified sample lineup, and the sample battle results of each battle can be obtained, thus obtaining at least two sample battle results. Then, the average of the at least two sample battle results can be calculated to obtain the sample average battle result. The lineup difference result between the modified sample lineup and the random sample lineup can be determined based on the sample average battle result. And the reward data can be determined based on the lineup difference result. The lineup difference result is positively correlated with the absolute value of the reward data. For example, the lineup difference result is proportional to the absolute value of the reward data. This embodiment avoids the randomness of a single sample battle result by calculating the sample average battle result of multiple battles, and further determines the reward data, thereby improving the accuracy of the reward data.

[0106] In some embodiments, adjusting the model parameters of the preset model based on the reward data until the training termination condition is met includes: Obtain the action sequence during the lineup adjustment process; If the reward data is positive, adjust the model parameters of the preset model to increase the probability of using the action sequence until the training termination condition is met; If the reward data is negative, adjust the model parameters of the preset model to reduce the probability of using the action sequence until the training termination condition is met.

[0107] In this embodiment of the disclosure, during model training, the action sequence during each training phase of lineup adjustment can be acquired; then, the probability trend of the action sequence in subsequent model training is determined based on reward data, thereby adjusting the model parameters; for example, after acquiring the action sequence during lineup adjustment, if the sample battle result indicates that the sample adjusted lineup defeats the sample random lineup, the model parameters of the preset model are adjusted to increase the probability of the action sequence being used; if the sample battle result indicates that the sample random lineup defeats the sample adjusted lineup, the model parameters of the preset model are adjusted to decrease the probability of the action sequence being used; until the training termination condition is met, thereby ensuring that the model converges quickly and can accurately predict the output lineup that can defeat the input lineup based on the input lineup.

[0108] For example, in a game scenario, such as Figure 5 As shown, Figure 5 This is a schematic diagram of a training system for a lineup enhancement model. The system includes a Unity editor for the game environment, a reinforcement learning module, and a game-side combat simulator. Unity, as the game's runtime environment, provides the following key functions: Virtual environment construction: Developers can create game scenes, character models, etc., within Unity as a training environment for machine learning. Real-time data interaction: The Unity engine enables efficient interaction between the agent and the environment, including perceiving environmental states and executing actions. Multi-platform compatibility: Supports game development needs on PC and mobile platforms, facilitating cross-platform training and deployment. During training, Unity continuously generates sample random lineups (Lineup A) and sends them to the reinforcement learning module. Through neural network inference using a preset model, Lineup B is obtained and returned to Unity. Unity then sends Lineup A and Lineup B to the server of the game-side combat simulator. The server uses Lineup A and Lineup B to initiate in-game combat gameplay (the specific combat gameplay is not limited; only win / loss and damage statistics are required), and sends the battle results back to Unity. This section can simulate multiple matches to calculate average win rate, battle losses, etc., to avoid randomness. The Unity part sends the results of the battles between lineup A and lineup B back to the reinforcement learning module, which can calculate rewards based on the battle results and update the parameters of the preset model, thereby training a lineup enhancement model.

[0109] In some embodiments, such as Figure 6 As shown, the random lineup has at least two elements. The process of constructing the lineup sequence corresponding to the random lineup based on the initial lineup and the lineup enhancement model includes: S601: Input the initial lineup corresponding to each random lineup into the lineup enhancement model to perform lineup enhancement and adjustment, and output the current adjusted lineup corresponding to each random lineup; S603: Input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model to enhance and adjust the lineup, and output the current updated lineup corresponding to each random lineup; S605: Take the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup again, and repeat the steps of inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and outputting the current updated lineup corresponding to each random lineup until the preset conditions are met. S607: Based on the initial lineup, the currently adjusted lineup, and the currently updated lineup corresponding to each random lineup, construct the lineup sequence corresponding to each random lineup.

[0110] In this embodiment of the disclosure, during application, at least two random lineups can be obtained to construct a seed lineup pool. Then, for each random lineup, multiple lineup adjustments are performed using a lineup enhancement model, with each subsequent lineup adjustment using the output lineup from the previous adjustment as the input lineup. For example, the initial lineup corresponding to each random lineup is input into the lineup enhancement model for lineup enhancement adjustment, outputting the current adjusted lineup for each random lineup. The current adjusted lineup corresponding to each random lineup is then input into the lineup enhancement model for lineup enhancement adjustment, outputting the current updated lineup for each random lineup. The current updated lineup sequence for each random lineup is then used again as the current adjusted lineup for each random lineup. This process of inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement adjustment and outputting the current updated lineup for each random lineup continues until a preset condition is met. The preset condition can be determined by the number of repetitions or the number of lineups in the lineup sequence. Finally, the initial lineup corresponding to each random lineup, all currently adjusted lineups, and all currently updated lineups are aggregated, allowing for the rapid construction of a lineup sequence corresponding to each random lineup based on actual needs.

[0111] In some embodiments, such as Figure 7 As shown, the steps of reusing the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup, repeatedly inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and outputting the current updated lineup corresponding to each random lineup until the preset conditions are met include: S6051: For each random lineup, determine the cumulative number of the initial lineup, the current adjusted lineup, and the current updated lineup; S6053: If the cumulative number of lineups does not reach the target number, the current updated lineup sequence corresponding to each random lineup is used as the current adjusted lineup corresponding to each random lineup again, and the steps of inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and outputting the current updated lineup corresponding to each random lineup are repeated until the preset conditions are met.

[0112] In this embodiment, the target number of lineups in each lineup sequence can be preset according to actual needs, and then the number of times the model is used repeatedly is determined based on the target number. If the cumulative number of lineups does not reach the target number, the current updated lineup sequence corresponding to each random lineup is used again as the current adjusted lineup corresponding to each random lineup, and the process of inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and outputting the current updated lineup corresponding to each random lineup is repeated until a preset condition is met. The preset condition can be set to the cumulative number of lineups reaching the target number. This allows for flexible control of the number of lineups in the lineup sequence. If the cumulative number of lineups reaches the target number, then the lineup sequence corresponding to each random lineup is constructed based on the initial lineup, the current adjusted lineup, and the current updated lineup.

[0113] In some embodiments, determining a preset number of opposing lineups that are ranked later in the lineup sequence as top lineups includes: Get the preset number of opposing lineups that are sorted at the end of the lineup sequence corresponding to each random lineup, and obtain the enhanced lineup corresponding to each random lineup. The enhanced lineup corresponding to each random lineup is determined as the top lineup.

[0114] In this embodiment, for multiple random lineups, a lineup enhancement model can be used to obtain the lineup sequence corresponding to each random lineup. Then, a fixed number of enhanced lineups are selected from each lineup sequence. Finally, the enhanced lineups in each lineup sequence are aggregated to obtain multiple top lineups, thereby ensuring the diversity of top lineups. In application, teams can be formed based on multiple top lineups for battle, thereby further filtering out lineups with strong combat capabilities.

[0115] In some embodiments, constructing the lineup sequence corresponding to the random lineup based on the initial lineup and the lineup enhancement model includes: inputting the initial lineup into the lineup enhancement model to construct the initial lineup sequence corresponding to the random lineup; using the initial lineup sequence as the current lineup sequence and inputting the current lineup sequence into the lineup enhancement model to construct the current updated lineup sequence; using the current updated lineup sequence as the current lineup sequence again, and repeating the step of inputting the current lineup sequence into the lineup enhancement model to construct the current updated lineup sequence until a preset condition is met; and constructing the lineup sequence corresponding to the random lineup based on the initial lineup sequence, the current lineup sequence, and the current updated lineup sequence.

[0116] For example, the step of reusing the current updated lineup sequence as the current lineup sequence and repeating the step of inputting the current lineup sequence into the lineup enhancement model to construct the current updated lineup sequence until a preset condition is met includes: determining the current cumulative sequence number based on the initial lineup sequence, the current lineup sequence, and the current updated lineup sequence; if the current cumulative sequence number is less than a preset number, it is determined that the preset condition is not met, and the step of inputting the current lineup sequence into the lineup enhancement model to construct the current updated lineup sequence is repeated until the current cumulative sequence number is greater than or equal to the preset number.

[0117] For example, the step of inputting the initial lineup into the lineup enhancement model to construct the initial lineup sequence corresponding to the random lineup includes: inputting the initial lineup into the lineup enhancement model for lineup enhancement adjustment and outputting the current adjusted lineup; inputting the current adjusted lineup into the lineup enhancement model for lineup enhancement adjustment and outputting the current updated lineup; determining the cumulative number of lineups of the initial lineup, the current adjusted lineup, and the current updated lineup; if the cumulative number of lineups does not reach the target number, using the current updated lineup as the current adjusted lineup again, and repeating the step of inputting the current adjusted lineup into the lineup enhancement model for lineup enhancement adjustment and outputting the current updated lineup; if the cumulative number of lineups reaches the target number, constructing the initial lineup sequence corresponding to the random lineup based on the initial lineup, the current adjusted lineup, and the current updated lineup.

[0118] In some embodiments, after determining a preset number of opposing lineups that are ranked later in the lineup sequence as top lineups, the method further includes: Obtain the lineup attribute information corresponding to each top lineup; Based on the lineup attribute information corresponding to each top lineup, a top lineup curve is drawn and visualized.

[0119] In this embodiment, each top-tier lineup can consist of one or more lineup attribute information. In a game scenario, the lineup attribute information may include specific troop types, such as archers, infantry, and cavalry, as well as specific hero character types. Then, based on the lineup attribute information corresponding to each top-tier lineup, a top-tier lineup curve can be plotted and visualized, allowing users to intuitively view the lineup attribute information of the top-tier lineup.

[0120] In some embodiments, the sample random lineup is the lineup of the sample service in the current version, the lineup enhancement model is the model corresponding to the current version, and the method further includes: If the sample service has an updated version, obtain the updated random lineup of samples under the updated version; The lineup enhancement model is trained based on the updated random lineup of the sample to obtain the updated lineup enhancement model; the updated lineup enhancement model is used to adjust the random lineup under the updated version.

[0121] In this embodiment, if an updated version of the sample service exists, an updated random lineup of samples under the updated version is obtained; then, the lineup enhancement model is trained based on the updated random lineup of samples to obtain an updated lineup enhancement model; its training method is similar to that of the lineup enhancement model, essentially fine-tuning the lineup enhancement model to obtain the updated lineup enhancement model; the updated lineup enhancement model is used to perform lineup enhancement adjustments on the random lineup under the updated version. For example, reinforcement learning AI, during its design and training process, learns and fixes relevant knowledge of lineup combination and matching into its neural network parameters. Its input and action space are friendly to game changes (often numerical and mechanic changes). When the game iterates frequently, it can calculate wins, losses, and damage based on the previously trained version in the current new game version (corresponding to a new "battle simulation server"), thus quickly migrating to the new game environment while retaining previously learned knowledge.

[0122] The model training method in this embodiment reduces invalid combination testing by 70% compared to genetic algorithms through self-game-based directional exploration; enhanced dynamic adaptability and historical strategy migration shorten retraining time after version iteration by 50%; interpretable complex relationships and clear reactivity graph reveal actionable insights such as "the strength overflow of the combination of hero X and skill Y in the infantry system".

[0123] For example, this embodiment can change different unit pools into a competitive relationship (such as the cavalry pool and infantry pool evolving alternately), replacing the static baseline lineup with adversarial evaluation. This may improve the efficiency of finding the global optimal solution, but it will increase the consumption of computational resources.

[0124] This embodiment can enhance generalization ability through transfer learning, building a game type knowledge base (such as a MOBA / RTS parameter mapping table) on top of the abstraction layer, and accelerating the adaptation to new games through pre-trained models. However, it requires additional labeled datasets, resulting in high initial implementation costs.

[0125] In this embodiment, lineup decisions can be made directly based on reinforcement learning AI. Compared to lineup search based on genetic algorithms, reinforcement learning AI can make decisions more efficiently and assemble a lineup stronger than the given lineup. It adopts a hierarchical action space design: a cascaded decision structure of unit type → hero → skill, which is compatible with legality constraints (such as hero duplicate detection); a reward shaping mechanism; a dual-objective reward that integrates damage difference and win rate to avoid local optima; and the strength relationship is visualized based on the strength relationship of lineup confrontation data.

[0126] Figure 8 This is a block diagram of a business data processing apparatus according to an exemplary embodiment. (Refer to...) Figure 8 The device includes: The random lineup acquisition module 810 is configured to acquire the random lineup corresponding to the preset business. The initial lineup acquisition module 820 is configured to input the random lineup into the lineup enhancement model for lineup enhancement and adjustment to obtain the initial lineup; the lineup enhancement model is obtained by adjusting the preset model with the sample random lineup of the sample business to obtain the sample adjusted lineup, and by training the preset model based on the battle results between the sample random lineup and the sample adjusted lineup. The lineup sequence construction module 830 is configured to construct a lineup sequence corresponding to the random lineup based on the initial lineup and the lineup enhancement model; the lineup sequence includes at least two adversarial lineups ordered according to a determined time; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier. The head lineup determination module 840 is configured to determine a preset number of opposing lineups that are ranked later in the lineup sequence as the head lineup.

[0127] In one exemplary implementation, the random lineup is at least two, and the lineup sequence construction module includes: The current lineup adjustment unit is configured to input the initial lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement adjustment, and output the current adjusted lineup corresponding to each random lineup; The current lineup update unit is configured to input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model to perform lineup enhancement and adjustment, and output the current updated lineup corresponding to each random lineup. The repeated adjustment unit is configured to perform the following steps: re-enable the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup, repeatedly input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement adjustment, and output the current updated lineup corresponding to each random lineup until the preset conditions are met. The sequence construction unit is configured to construct a lineup sequence corresponding to each random lineup based on the initial lineup, the currently adjusted lineup, and the currently updated lineup corresponding to each random lineup.

[0128] In one exemplary embodiment, the repetition adjustment unit includes: The initial lineup determination subunit is configured to perform, for each random lineup, the cumulative number of lineups: the initial lineup, the current adjusted lineup, and the current updated lineup; The lineup adjustment subunit is configured to perform the following steps if the cumulative number of lineups does not reach the target number: re-use the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup, and repeatedly input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and output the current updated lineup corresponding to each random lineup until the preset conditions are met.

[0129] In one exemplary implementation, the head lineup determination module includes: The enhanced lineup determination unit is configured to retrieve a preset number of opposing lineups from the lineup sequence corresponding to each random lineup, and obtain the enhanced lineup corresponding to each random lineup. The head lineup determination unit is configured to determine the enhanced lineup corresponding to each random lineup as the head lineup.

[0130] In one exemplary embodiment, the apparatus further includes: The sample lineup acquisition module is configured to execute the sample random lineup acquisition service; The sample lineup adjustment module is configured to input the random lineup of the sample into the preset model to adjust the lineup and obtain the adjusted lineup of the sample. The sample result determination module is configured to perform a simulated battle between the sample random lineup and the sample adjusted lineup in the simulated scenario of the sample business, and obtain the sample battle result. The reward data determination module is configured to calculate the reward data corresponding to the preset model based on the sample battle results; if the sample battle results indicate that the sample adjusted lineup wins, the reward data is a positive number; if the sample battle results indicate that the sample random lineup wins, the reward data is a negative number. The model training module is configured to adjust the model parameters of the preset model according to the reward data until the training termination condition is met, and to determine the preset model at the end of training as the lineup enhancement model.

[0131] In one exemplary embodiment, the model training module includes: The action sequence acquisition unit is configured to acquire the action sequence during the lineup adjustment process; The first adjustment unit is configured to perform the following actions: if the reward data is positive, adjust the model parameters of the preset model to increase the probability of using the action sequence until the training termination condition is met. The second adjustment unit is configured to adjust the model parameters of the preset model to reduce the probability of using the action sequence until the training termination condition is met if the reward data is negative.

[0132] In one exemplary embodiment, the reward data determination module includes: The sample average result determination unit is configured to perform the following: if there are at least two sample battle results, determine the average of the at least two sample battle results to obtain the sample average battle result; The sample difference determination unit is configured to determine the lineup difference result between the sample adjusted lineup and the sample random lineup based on the average battle result of the sample; The reward data determination unit is configured to determine the reward data based on the lineup difference result; the lineup difference result is positively correlated with the absolute value of the reward data.

[0133] In one exemplary embodiment, the sample lineup adjustment module includes: The sample field acquisition unit is configured to acquire the set of sample fields in the random sample lineup. The sample vector conversion unit is configured to perform one-hot encoding of the sample fields in the sample field set and convert them into a fixed-length sample vector. The sample lineup adjustment unit is configured to concatenate the individual sample vectors to obtain a concatenated sample vector, and then input the concatenated sample vector into the preset model for lineup adjustment to obtain the adjusted sample lineup.

[0134] In one exemplary embodiment, the sample lineup adjustment unit includes: The output head determination subunit is configured to perform the following operations: obtain the number of fields in the sample fields set, and determine the number of output heads in the neural network of the preset model based on the number of fields; the number of output heads is equal to the number of fields. The sample result determination subunit is configured to input the sample concatenation vector into the preset model, and use a multi-head masking mechanism to process one sample vector in the sample concatenation vector based on each output head to obtain the sample output result corresponding to each sample vector; The sample adjustment subunit is configured to execute the sample output results based on each sample vector to obtain the sample adjustment lineup.

[0135] In one exemplary embodiment, the apparatus further includes: The lineup attribute acquisition module is configured to retrieve the lineup attribute information corresponding to each top lineup. The display module is configured to draw a curve chart of the top lineups based on the lineup attribute information corresponding to each top lineup and display it visually.

[0136] In one exemplary embodiment, the sample random lineup is the lineup of the sample service under the current version, the lineup enhancement model is the model corresponding to the current version, and the device further includes: The updated lineup acquisition module is configured to acquire the sample updated random lineup under the updated version if the sample service has an updated version. The updated model training module is configured to perform lineup enhancement training on the lineup enhancement model based on the updated random lineup of the sample, so as to obtain an updated lineup enhancement model; the updated lineup enhancement model is used to perform lineup enhancement adjustments on the random lineup under the updated version.

[0137] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0138] In one exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored in the memory, it implements the business data processing method provided in any of the above embodiments.

[0139] The electronic device can be a terminal, a server, or a similar computing device. Taking a server as an example... Figure 9 This is a block diagram of a server according to an exemplary embodiment, such as... Figure 9 As shown, the server 900 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 910 (CPUs 910 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 930 for storing data, and one or more storage media 920 (e.g., one or more mass storage devices) for storing application programs 923 or data 922. The memory 930 and storage media 920 may be temporary or persistent storage. The program stored in the storage media 920 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 910 may be configured to communicate with the storage media 920 and execute the series of instruction operations stored in the storage media 920 on the server 900. Server 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows Server™, Mac OSX™, Unix™, Linux™, FreeBSD™, etc.

[0140] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 940 may be a radio frequency (RF) module used for wireless communication with the Internet.

[0141] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 900 may also include... Figure 9 The more or fewer components shown, or having the same Figure 9 The different configurations shown.

[0142] In one exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 930 including instructions, which can be executed by a processor 910 of a server 900 to perform the above-described method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0143] In one exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the business data processing method provided in any of the above embodiments.

[0144] Figure 10 This is a block diagram illustrating an electronic device for business data processing according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a business data processing method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse. Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0145] This disclosure obtains a random lineup corresponding to a preset service; inputs the random lineup into a lineup enhancement model for lineup enhancement and adjustment to obtain an initial lineup; wherein, the lineup enhancement model is obtained by adjusting the preset model with a sample random lineup of a sample service to obtain a sample adjusted lineup, and by training the preset model based on the battle results between the sample random lineup and the sample adjusted lineup; based on the initial lineup and the lineup enhancement model, a lineup sequence corresponding to the random lineup is constructed; the lineup sequence includes at least two adversarial lineups ordered according to a certain time, thereby quickly constructing a series of optimized lineups through the lineup enhancement model; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier; a preset number of adversarial lineups ranked later in the lineup sequence are determined as the top lineups, thereby ensuring that the extracted top lineups have stronger combat capabilities than other lineups. This disclosure achieves the rapid and accurate construction of lineups with strong combat capabilities.

[0146] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0147] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0148] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A business data processing method, characterized in that, include: Get the random lineup corresponding to the preset service; The random lineup is input into the lineup enhancement model for lineup enhancement and adjustment to obtain the initial lineup; The lineup enhancement model is obtained by training a preset model based on the battle results between a sample random lineup and a sample adjusted lineup from the sample business. The sample adjusted lineup is obtained by adjusting the preset model using the sample random lineup. The random lineup consists of at least two. The initial lineup corresponding to each random lineup is input into the lineup enhancement model for lineup enhancement and adjustment, and the current adjusted lineup corresponding to each random lineup is output. Input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model to enhance and adjust the lineup, and output the current updated lineup corresponding to each random lineup. The current updated lineup sequence corresponding to each random lineup is used as the current adjusted lineup corresponding to each random lineup, and the current adjusted lineup corresponding to each random lineup is repeatedly input into the lineup enhancement model for lineup enhancement and adjustment, and the current updated lineup corresponding to each random lineup is output until the preset conditions are met. Based on the initial lineup, the currently adjusted lineup, and the currently updated lineup corresponding to each random lineup, a lineup sequence corresponding to each random lineup is constructed; the lineup sequence includes at least two adversarial lineups ordered according to a determined time; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier; A predetermined number of opposing lineups that are ranked later in the lineup sequence are identified as the top lineups.

2. The method according to claim 1, characterized in that, The steps of reusing the current updated lineup sequence corresponding to each random lineup as the current adjusted lineup corresponding to each random lineup, repeatedly inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment, and outputting the current updated lineup corresponding to each random lineup until the preset conditions are met include: For each random lineup, determine the cumulative number of lineups including the initial lineup, the currently adjusted lineup, and the currently updated lineup; If the cumulative number of lineups does not reach the target number, the current updated lineup sequence corresponding to each random lineup is used as the current adjusted lineup corresponding to each random lineup again, and the process of inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment is repeated until the current updated lineup corresponding to each random lineup is output until the preset conditions are met.

3. The method according to claim 1, characterized in that, The step of determining a preset number of opposing lineups that are ranked lower in the lineup sequence as top lineups includes: Get the preset number of opposing lineups that are sorted at the end of the lineup sequence corresponding to each random lineup, and obtain the enhanced lineup corresponding to each random lineup. The enhanced lineup corresponding to each random lineup is determined as the top lineup.

4. The method according to any one of claims 1-3, characterized in that, The training method for the lineup enhancement model includes: Obtain the random sample lineup of the sample service; The sample random lineup is input into the preset model for lineup adjustment to obtain the sample adjusted lineup; The sample random lineup and the sample adjusted lineup are used to simulate a battle in the simulated scenario of the sample business, and the sample battle results are obtained. The reward data corresponding to the preset model is calculated based on the sample battle results; if the sample battle results indicate that the sample adjusted lineup wins, the reward data is a positive number; if the sample battle results indicate that the sample random lineup wins, the reward data is a negative number. The model parameters of the preset model are adjusted according to the reward data until the training end condition is met, and the preset model at the end of training is determined as the lineup enhancement model.

5. The method according to claim 4, characterized in that, The step of adjusting the model parameters of the preset model based on the reward data until the training termination condition is met includes: Obtain the action sequence during the lineup adjustment process; If the reward data is positive, adjust the model parameters of the preset model to increase the probability of using the action sequence until the training termination condition is met; If the reward data is negative, adjust the model parameters of the preset model to reduce the probability of using the action sequence until the training termination condition is met.

6. The method according to claim 4, characterized in that, The calculation of reward data corresponding to the preset model based on the sample battle results includes: If there are at least two sample battle results, determine the average of the at least two sample battle results to obtain the sample average battle result; The difference between the adjusted lineup and the random lineup in the sample is determined based on the average battle results of the sample. The reward data is determined based on the lineup difference results; the lineup difference results are positively correlated with the absolute value of the reward data.

7. The method according to claim 4, characterized in that, The step of inputting the random sample lineup into the preset model for lineup adjustment to obtain the adjusted sample lineup includes: Obtain the set of sample fields from the random sample lineup; The sample fields in the sample field set are converted into fixed-length sample vectors using one-hot encoding. The individual sample vectors are concatenated to obtain a concatenated sample vector, which is then input into the preset model for lineup adjustment to obtain the adjusted sample lineup.

8. The method according to claim 7, characterized in that, The step of inputting the sample concatenation vector into the preset model for lineup adjustment to obtain the sample adjusted lineup includes: Obtain the number of fields in the sample fields set, and determine the number of output heads in the neural network of the preset model based on the number of fields; the number of output heads is equal to the number of fields. The sample concatenation vector is input into the preset model, and a multi-head masking mechanism is used to process one sample vector in the sample concatenation vector based on each output head to obtain the sample output result corresponding to each sample vector; The sample adjustment lineup is obtained based on the sample output results corresponding to each sample vector.

9. The method according to claim 1, characterized in that, After determining the predetermined number of opposing lineups that are ranked lower in the lineup sequence as the top lineups, the method further includes: Obtain the lineup attribute information corresponding to each top lineup; Based on the lineup attribute information corresponding to each top lineup, a top lineup curve is drawn and visualized.

10. The method according to claim 1, characterized in that, The sample random lineup is the lineup of the sample service under the current version, the lineup enhancement model is the model corresponding to the current version, and the method further includes: If the sample service has an updated version, obtain the updated random lineup of samples under the updated version; The lineup enhancement model is trained based on the updated random lineup of the sample to obtain the updated lineup enhancement model; the updated lineup enhancement model is used to adjust the random lineup under the updated version.

11. A business data processing apparatus, characterized in that, include: The random lineup acquisition module is configured to acquire a random lineup corresponding to a preset business. The initial lineup acquisition module is configured to input the random lineup into the lineup enhancement model for lineup enhancement and adjustment to obtain the initial lineup; the lineup enhancement model is obtained by training a preset model based on the battle results between sample random lineups and sample adjusted lineups in sample business; the sample adjusted lineup is obtained by adjusting the preset model using the sample random lineup; the random lineup is at least two. The lineup sequence construction module is configured to input the initial lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement adjustment, and output the current adjusted lineup corresponding to each random lineup. Input the current adjusted lineup corresponding to each random lineup into the lineup enhancement model to enhance and adjust the lineup, and output the current updated lineup corresponding to each random lineup. The current updated lineup sequence corresponding to each random lineup is used as the current adjusted lineup corresponding to each random lineup, and the process of inputting the current adjusted lineup corresponding to each random lineup into the lineup enhancement model for lineup enhancement and adjustment is repeated until the current updated lineup corresponding to each random lineup is output until the preset condition is met; based on the initial lineup, the current adjusted lineup, and the current updated lineup corresponding to each random lineup, a lineup sequence corresponding to each random lineup is constructed; the lineup sequence includes at least two adversarial lineups ordered according to a determined time; the adversarial lineup ranked first in the lineup sequence is the initial lineup, and the adversarial lineups ranked later in the lineup sequence are obtained by adjusting the adversarial lineups ranked earlier; The top lineup determination module is configured to determine a preset number of opposing lineups that are ranked later in the lineup sequence as the top lineup.

12. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the business data processing method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device is able to perform the business data processing method as described in any one of claims 1-10.

14. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the business data processing method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Game data processing method and device, equipment and storage medium

    CN115212576A

  • Data processing method and device, electronic equipment and readable storage medium

    CN116764233A