A Multi-Level Deduction Interaction Data Processing Method and System Based on Intelligent Game
By designing an intelligent game system that includes multi-level deduction components and agent service components, the problems of unclear logic and unclear multi-level deduction processes in the existing technology are solved, and more efficient use of intelligent game capabilities and adversarial analysis experience is achieved.
Patent Information
- Application Number
- CN202411817720.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The logic in the existing intelligent game technology is unclear, the multi-level deduction process is unclear, and the problem of insufficient utilization of existing adversarial analysis experience.
A multi-level deduction interactive data processing method and system based on intelligent game is designed, including a control unit and an intelligent game unit. The control unit includes user development components and third-party environment access components, and the intelligent game unit includes agent service components, multi-level deduction components, database components and system background components. The system realizes the encapsulation of intelligent game information and the processing of multi-level deduction data through the upper, intermediate and lower deduction action sets of multi-level deduction components, combining the agent service components and database components.
This system improves the logical clarity of the intelligent game process and the transparency of multi-level deduction, enhances the utilization of existing adversarial analysis experience, and improves the overall ability of intelligent game.
Smart Images

Figure CN119514697B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent agent intelligent game, and particularly relates to a multi-level deduction interaction data processing method and system based on intelligent game. Background Art
[0002] The development of artificial intelligence technology has gone through three stages: computational intelligence, perceptual intelligence, and cognitive intelligence. Computational intelligence mainly involves rule-based operations such as scientific computing, logical processing, and statistical queries; Perceptual intelligence, represented by image understanding, speech recognition, and machine translation, is based on deep learning models; Cognitive intelligence, represented by understanding, reasoning, thinking, and decision-making, emphasizes cognitive reasoning and autonomous learning capabilities. Intelligent game, as a frontier paradigm of decision-making intelligence, is a high-order manifestation of cognitive intelligence. It is mainly supported by game theory, emphasizing incorporating other intelligent agents (teammates and opponents) into one's own decision-making for rule self-learning, game confrontation evolution, interpretable strategy recommendation, etc. The development of artificial intelligence has created new possibilities for the game field. At the theoretical level, artificial intelligence provides new solutions for games, such as computational complexity problems in Nash equilibrium and equilibrium in cooperative games. In practical applications, artificial intelligence has also found new applications in game scenarios such as Go and games. For example, AlphaGo, AlphaStar, etc. have obtained excellent strategies through the application of deep reinforcement learning methods and defeated top human players in human-machine game confrontations.
[0003] Chinese invention patent CN112295229B, an intelligent game confrontation platform, includes a simulation deduction platform and an AI confrontation training environment management subsystem. From the aspects of system function composition, interaction relationship, and work process, a system design method for the intelligent game confrontation platform is proposed, aiming to solve the problems of environmental simulation and test system construction in intelligent game confrontation training experiments. The simulation deduction platform provides functions such as battlefield environment simulation, intelligence simulation, weapon platform simulation, damage simulation, and real-time intelligent instruction intervention; The AI confrontation training environment management subsystem creates a training confrontation platform by mirroring the simulation deduction platform using container technology through configuring operation environment parameters and operation condition parameters, and dynamically forms a prototype of the confrontation training test environment. The main focus of this invention is to enhance environmental adaptability and can be used as a plugin for other intelligent game platforms, but the main logic is low and the performance improvement is small.
[0004] Chinese Invention Patent CN113379054A, an open intelligent game ecosystem platform, where the platform includes a central site and at least one adversarial system sub-site, and each adversarial system sub-site operates independently to achieve adversarial functions; the central site interacts with the sub-site data interfaces of each adversarial system sub-site through the central data interface; when the sub-site data interface of any adversarial system sub-site is opened, the central site transmits the user information of the transferred user to the any adversarial system sub-site, and the any adversarial system sub-site transmits the sub-site adversarial data to the central site; wherein, the transferred user is the user who logs in to the any adversarial system sub-site through the central site. The idea of this invention patent is to use the distributed computing function to improve the overall operation efficiency, which is an idea to improve the operation speed, but the improvement in the real intelligent game logic is still very small.
[0005] None of the methods and ways provided in the above invention patent specifically and deeply explain the problem of how to improve the ability of the game itself in the current intelligent game process, nor do they involve the way of reusing the existing experience in multiple intelligent games. Therefore, the logical architecture of the existing intelligent game is not ideal enough. Summary of the Invention
[0006] The present invention provides a multi-level deduction interactive data processing method and system based on intelligent games to solve the technical problems of unclear logic, unclear multi-level deduction process, and insufficient utilization of existing adversarial analysis experience in the existing intelligent game process.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] The present invention provides a multi-level deduction interactive data processing method and system based on intelligent games, including:
[0009] A parameter control unit and an intelligent game unit, the parameter control unit includes a user development component and a third-party environment access component, and the intelligent game unit includes an intelligent agent service component, a multi-level deduction component, a database component, and a system background component;
[0010] The user development component is used for the basic parameter configuration of the system, controlling the game process, and providing a visual operation interface;
[0011] The third-party environment access component is used to provide an information interface for accessing a third-party environment. The user development component selects the accessed third-party environment, and the third-party environment access component provides third-party environment access component information to the database component;
[0012] The agent service component includes a single-agent service set and a multi-agent service set. The agent service component is used to encapsulate agent game information, and the agent game information is encapsulated as an experience information set and stored in the database component. The agent service component receives multi-level deduction data from the multi-level deduction component;
[0013] The multi-level deduction component includes a superior deduction action set, a middle-level deduction action set, and a lower-level deduction action set. The superior deduction action set, the middle-level deduction action set, and the lower-level deduction action set complete multi-level deduction through deduction actions. The multi-level deduction component generates the multi-level deduction data and sends it to the agent service component;
[0014] The database component receives third-party environment access component information from the third-party environment access component, and the database component receives and provides the experience information set from the agent service component;
[0015] The system background component provides interfaces and routing services to network-connect the user development component, the third-party environment access component, the agent service component, the multi-level deduction component, and the database component.
[0016] Further, the user development component is deployed on a single control computer;
[0017] The user development component can adjust the configuration parameters of the agent service component, so as to select the use of the single-agent service set and the multi-agent service set;
[0018] The user development component selects the third-party environment accessed in the third-party environment access component;
[0019] The user development component views the operation and maintenance status information through a visual operation interface.
[0020] Further, the third-party environment access component can connect to multiple third-party environments;
[0021] The third-party environment is an existing simulation operation platform;
[0022] The third-party environment access component stores the used existing simulation operation platform in the database component.
[0023] Further, the agent service component includes a single-agent service set and a multi-agent service set;
[0024] The single-agent service set copies the experience information set in the database component into experience information 1 and experience information 2, optimizes experience information 1 using policy gradient optimization to obtain a policy network, optimizes experience information 2 using value gradient optimization to obtain a value network, and the policy network obtains deterministic actions and sends them to the value network to predict the single-agent action set;
[0025] The single-agent training process is as follows: control the game engine to start, the single agent acts, obtains and sends the current state, evaluates the situation to calculate the decision. On the one hand, continue to obtain the situation, and on the other hand, use the current state to fill the experience information, infer the state change for parameter update, judge whether the expected state change converges. If it does not converge, return to the step of inferring the state change for parameter update. If it converges, store the experience information, store the experience information in the database together with the information obtained on the one hand, and judge whether the game result meets the expectation;
[0026] The multi-agent service set shares the experience information set in the database component with all agents in the multi-agent service set to implement parallel execution of the experience information set strategy;
[0027] The single-agent service set and the multi-agent service set process the experience information to obtain game information, and after encapsulation, store it in the database component.
[0028] Further, the multi-level deduction component obtains observation information by observing the agent service component;
[0029] The observation information includes campaign-level observation information and tactical-level observation information.
[0030] Further, the multi-level deduction process of the multi-level deduction component is as follows:
[0031] In multi-level deduction process one, the superior deduction action set generates campaign-level decisions through a pulsed neural network based on the monotonic value function decomposition algorithm, and obtains the superior deduction action strategy and sends it to the intermediate deduction action set;
[0032] In multi-level deduction process two, the intermediate deduction action level generates tactical-level decisions through a pulsed neural network based on the monotonic value function decomposition algorithm, and obtains the intermediate deduction action strategy and sends it to the lower-level deduction action set;
[0033] In multi-level deduction process three, the lower-level deduction action level generates tactical-level decisions through a pulsed neural network based on the proximal policy optimization algorithm.
[0034] Further, when the database component receives third-party environment access component information, it receives 1 type of third-party environment at a time;
[0035] The experience information set is a training experience information set and a test experience information set.
[0036] Furthermore, the generation process of the training information set and the test experience information set is as follows:
[0037] When the database component first uses a third-party environment, the agent service component receives the experience information set. At this time, the experience information set is the test experience information set. The agent service component processes and packages the test experience information set and sends it to the database component to obtain the training experience information set;
[0038] When the database component continues to use the third-party environment, the agent service component obtains the training experience information set from the database component, processes and packages it to obtain a new training experience information set, and sends it to the database component.
[0039] Furthermore, it also includes a viewing component that can record and store the multi-level deduction process and the experience information, and implement playback and speed playback according to the user's development component.
[0040] Furthermore, the system background component simultaneously controls the user access of the parameter control unit, generates operation and maintenance status information, and maintains the stable operation of this system.
[0041] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0042] The present invention provides a multi-level deduction interaction data processing method and system based on intelligent game, including: a parameter control unit and an intelligent game unit. The parameter control unit includes a user development component and a third-party environment access component, and the intelligent game unit includes an agent service component, a multi-level deduction component, a database component, and a system background component. The user development component is used for basic parameter configuration of the system, controlling the game process, and providing a visual operation interface. The third-party environment access component is used to provide an information interface for accessing a third-party environment. The user development component selects the third-party environment to be accessed, and the third-party environment access component provides third-party environment access component information to the database component. The agent service component includes a single-agent service set and a multi-agent service set. The agent service component is used to encapsulate agent game information, and the agent game information is encapsulated as an experience information set and stored in the database component. The agent service component receives multi-level deduction data from the multi-level deduction component. The multi-level deduction component includes an upper-level deduction action set, a middle-level deduction action set, and a lower-level deduction action set. The upper-level deduction action set, the middle-level deduction action set, and the lower-level deduction action set complete multi-level deduction through deduction actions. The multi-level deduction component generates the multi-level deduction data and sends it to the agent service component. The database component receives third-party environment access component information from the third-party environment access component, and the database component receives and provides the experience information set from the agent service component. The system background component provides interfaces and routing services to network-connect the user development component, the third-party environment access component, the agent service component, the multi-level deduction component, and the database component. The components of the present invention are clear and modularly divided, facilitating implementation.
[0043] In the present invention, the user development component can adjust the configuration parameters of the agent service component, so as to select to use the single-agent service set and the multi-agent service set, and can switch between the single-agent and multi-agent cases, reducing resource waste.
[0044] The multi-level deduction component in the present invention includes a three-level deduction process, and can perform targeted intelligent games using three strategies for high and low levels.
[0045] The present invention distinguishes the generation of training information sets and test experience information sets, and can freely switch under the same and different environment conditions, improving the analysis efficiency. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 It is a schematic structural diagram of a multi-level deduction interaction data processing method and system based on intelligent game provided by the present invention;
[0048] Figure 2 It is a flowchart of the single-agent training process of the present invention;
[0049] Figure 3 It is a schematic structural diagram of the multi-level deduction component of the present invention. Detailed implementation manners
[0050] To make the objectives, technical solutions and advantages of the present invention clearer, the following will further describe the implementation manners of the present invention in detail with reference to the accompanying drawings. Embodiment
[0051] This embodiment provides a multi-level deduction interaction data processing method and system based on intelligent game, including: a parameter control unit and an intelligent game unit. The parameter control unit includes a user development component and a third-party environment access component. The intelligent game unit includes an agent service component, a multi-level deduction component, a database component, and a system background component;
[0052] The user development component is used for the basic parameter configuration of this system, controlling the game process, and providing a visual operation interface;
[0053] The third-party environment access component is used to provide an information interface for accessing a third-party environment. The user development component selects the accessed third-party environment, and the third-party environment access component provides third-party environment access component information to the database component;
[0054] The agent service component includes a single-agent service set and a multi-agent service set. The agent service component is used to encapsulate agent game information. The agent game information is encapsulated as an experience information set and stored in the database component. The agent service component receives multi-level deduction data from the multi-level deduction component;
[0055] The multi-level deduction component includes an upper-level deduction action set, a middle-level deduction action set, and a lower-level deduction action set. The upper-level deduction action set, the middle-level deduction action set, and the lower-level deduction action set complete multi-level deduction through deduction actions. The multi-level deduction component generates the multi-level deduction data and sends it to the agent service component;
[0056] The database component receives third-party environment access component information from the third-party environment access component, and the database component receives and provides the experience information set from the agent service component;
[0057] The system background component provides interfaces and routing services to network-connect the user development component, the third-party environment access component, the agent service component, the multi-level deduction component, and the database component.
[0058] Specifically, the user development component is deployed on a single control computer;
[0059] The user development component can adjust the configuration parameters of the agent service component, so as to select to use a single-agent service set and a multi-agent service set;
[0060] The user development component selects the third-party environment accessed in the third-party environment access component;
[0061] The user development component views the operation and maintenance status information through a visual operation interface.
[0062] Please refer to Figure 1 the structural schematic diagram of the multi-level deduction interactive data processing method and system provided by the present invention as shown.
[0063] Specifically, the third-party environment access component can connect to multiple third-party environments;
[0064] The third-party environment is an existing simulation operation platform;
[0065] The third-party environment access component stores the used existing simulation operation platform into the database component.
[0066] Specifically, the agent service component includes a single-agent service set and a multi-agent service set;
[0067] The single-agent service set copies the experience information set in the database component into experience information 1 and experience information 2, optimizes experience information 1 using policy gradient optimization to obtain a policy network, optimizes experience information 2 using value gradient optimization to obtain a value network, and the policy network obtains a deterministic action and sends it into the value network to predict a single-agent action set;
[0068] Please refer to Figure 2 the flowchart of the single-agent training process of the present invention as shown.
[0069] The single-agent training process is as follows: start the control game engine, the single agent acts, obtain and send the current state, evaluate the situation and calculate the decision. On the one hand, continue to obtain the situation, on the other hand, use the current state to fill the experience information, infer the state change and update the parameters, judge whether the expected state change converges. If it does not converge, return to the step of inferring the state change and updating the parameters. If it converges, store the experience information, store the experience information into the database together with the situation on the one hand, and judge whether the game result meets the expectation;
[0070] The multi-agent service set shares the experience information set in the database component with all agents in the multi-agent service set, enabling parallel execution of the experience information set strategy;
[0071] The single-agent service set and the multi-agent service set process the experience information to obtain game information, which is encapsulated and stored in the database component.
[0072] Specifically, the multi-level deduction component obtains observation information by observing the agent service component;
[0073] The observation information includes campaign-level observation information and tactical-level observation information.
[0074] Specifically, the multi-level deduction process of the multi-level deduction component is as follows:
[0075] In the first multi-level deduction process, the upper-level deduction action set generates campaign-level decisions through a pulsed neural network based on the monotonic value function decomposition algorithm, and obtains the upper-level deduction action strategy and sends it to the middle-level deduction action set;
[0076] In the second multi-level deduction process, the middle-level deduction action level generates tactical-level decisions through a pulsed neural network based on the monotonic value function decomposition algorithm, and obtains the middle-level deduction action strategy and sends it to the lower-level deduction action set;
[0077] In the third multi-level deduction process, the lower-level deduction action level generates tactical-level decisions through a pulsed neural network based on the proximal policy optimization algorithm.
[0078] Please refer to Figure 3 the structural schematic diagram of the multi-level deduction component of the present invention shown.
[0079] Specifically, when the database component receives third-party environment access component information, it receives 1 type of third-party environment at a time;
[0080] The experience information set is a training experience information set and a test experience information set.
[0081] Specifically, the generation process of the training information set and the test experience information set is as follows:
[0082] When the database component first uses a third-party environment, the agent service component receives the experience information set. At this time, the experience information set is the test experience information set. The agent service component processes and encapsulates the test experience information set and sends it to the database component to obtain the training experience information set;
[0083] When the database component continues to use the third-party environment, the agent service component obtains the training experience information set from the database component, processes and encapsulates it to obtain a new training experience information set, and sends it to the database component.
[0084] Specifically, the system further includes a viewing component that can record and store the multi-level deduction process and the experience information, and implement playback and speed-doubling playback according to the user development component.
[0085] Specifically, the system background component simultaneously controls the user access of the parameter control unit, generates operation and maintenance status information, and maintains the stable operation of the system. Embodiment
[0086] This embodiment provides a method and system for processing multi-level deduction interaction data based on intelligent gaming. The system is specifically as follows:
[0087] In this embodiment, the Warcraft III engine produced by Blizzard Entertainment is used, version 1.32.10, and the AI-enhanced AMAI 3.3.1 version. Warcraft III is a classic real-time strategy game. Its battle mode and AI design have been tested for more than 20 years, and countless artificial intelligence testers have tested through this game.
[0088] AMAI (Advanced Melee AI) is a well-known open-source AI-enhanced MOD in Warcraft III. It completely rewrites the game's AI script, enabling the AI to have a behavior pattern closer to that of human players, including aspects such as tactics, strategies, and operations.
[0089] The artificial development component in this embodiment is actually a computer working platform, which sends an environment selection command to the third-party environment access component, selects Warcraft III, version 1.32.10, AI-enhanced AMAI 3.3.1 version in the third-party environment component, and selects the regular map TurtleRock as the third-party environment.
[0090] At this time, the third-party environment access component provides the third-party environment access component information to the database component. In this embodiment, it is to store the existing built-in enhanced version AMAI 3.3.1 AI in the database component through the third-party environment component as experience information.
[0091] At this time, for a single agent, for a single unit in the battle map, when giving a single operation instruction or the AI default instruction, the single-agent service set is selected. For a single agent in Warcraft III, there are 12 operation instructions and 6 item instructions for regular units.
[0092] The single-agent service set copies the AMAI 3.3.1 experience information set in the database component into experience information 1 and experience information 2, optimizes experience information 1 using policy gradient optimization to obtain a policy network, optimizes experience information 2 using value gradient optimization to obtain a value network, and the policy network feeds the deterministic actions into the value network to predict the single-agent action set, realizing the behavior operation of a single agent in simple actions.
[0093] The single-agent training process is as follows: Start the control game engine, the single agent acts, obtains and sends the current state, evaluates the situation to calculate the decision. On the one hand, continue to obtain the situation, and on the other hand, use the current state to fill the experience information, reason about the state change for parameter update, judge whether the expected state change converges. If it does not converge, return to the step of reasoning about the state change for parameter update. If it converges, store the experience information, store the experience information together with the former to the database, and judge whether the game result meets the expectation.
[0094] For multiple agents, use the centralized training method and the behavior method of multiple executions. The training experience of each unit is shared with all units. When the number of units is large, the above operation instructions can be executed in parallel, thus accelerating the learning process.
[0095] The single-agent service set and the multi-agent service set process the experience information to obtain game information, and after encapsulation, store it in the database component.
[0096] In this embodiment, the main deduction process of the multi-level deduction component is divided into 3 levels.
[0097] Multi-level deduction process one: The superior deduction action set generates campaign-level decisions through a pulsed neural network based on the monotonic value function decomposition algorithm, and obtains the superior deduction action strategy and sends it to the intermediate deduction action set;
[0098] Multi-level deduction process two: The intermediate deduction action level generates tactical-level decisions through a pulsed neural network based on the monotonic value function decomposition algorithm, and obtains the intermediate deduction action strategy and sends it to the lower deduction action set;
[0099] Multi-level deduction process three: The lower deduction action level generates tactical-level decisions through a pulsed neural network based on the proximal policy optimization algorithm.
[0100] The superior deduction action set is used to analyze and implement the global strategy in the battle process. For existing models, including but not limited to:
[0101] 1. Fast attack tactic: This strategy involves sacrificing economic and technological progress to quickly build an army in the early stage of the game to defeat the opponent. To accelerate this process, players often need to secretly build military buildings near the enemy's base and may mobilize all farmers to participate in the attack.
[0102] 2. One-wave Tactics: By predicting the moments when the enemy's forces are weak, adopting specific operation strategies, and assembling as powerful an army as possible to launch a general attack at these critical moments, which is usually accompanied by the timely upgrade of key technologies.
[0103] 3. Suppression Tactics: Dispatch troops to advance forward, forcing the opponent into an unfavorable position or restricting their expansion, while taking the opportunity to expand oneself.
[0104] 4. Harassment Tactics: Utilize highly mobile units, air units, or through transport units to avoid the enemy's main force and directly attack the enemy's peasants, thereby weakening the enemy's economic foundation.
[0105] 5. Map Control: In a situation where the two sides are evenly matched, the player moves flexibly on the battlefield, conducts extensive reconnaissance to understand the position and composition of the enemy's main force, and waits for the right opportunity to attack or harass.
[0106] 6. Steal Economy and Steal Technology: When it is considered that the enemy will not conduct detailed reconnaissance, take the risk of opening a satellite mine in a remote area to gain an economic advantage, or sacrifice troops or economy to research key technologies in advance, so as to gain the initiative in the attack.
[0107] 7. Late-game Tactics: In the early and mid-stages of the game, focus on defense and expansion, build a large number of defensive buildings to stabilize the front line. With the support of economy and technology, finally produce a large number of high-level units and gradually consume and eliminate the opponent.
[0108] 8. Base Trading Tactics: When the frontal battle is inferior to the opponent, adopt guerrilla tactics, avoid the enemy's main force, and circle around to their base to cause mutual destruction when the enemy is attacking.
[0109] The intermediate deduction action set mainly aims at the strategy generation in a smaller-scale battle, such as the strategy generation in the early-stage jungle clearing process with fewer units.
[0110] The monotonic function decomposition method Qmix is a classic algorithm in multi-agent reinforcement learning (MARL), which makes some improvements on the basis of VDN (Value-Decomposition Networks). The main advantage of Qmix is that it performs better in an environment where there are large differences among individual agents, can better handle the interaction and cooperation problems among agents, and can realize the strategy of optimizing the global result by changing exploration relying on the reward mechanism under the condition of less information interaction.
[0111] The Spiking Neural Network (SNN) is a commonly used network model in fields such as military command and strategy optimization. The SNN processes information by simulating the spiking behavior of biological neurons. Neurons in the SNN will emit spikes when they receive sufficient input, rather than continuously outputting activation values, and can have higher efficiency and energy efficiency when dealing with emergencies and urgent tasks such as games and military operations.
[0112] The lower-level deduction action level generates tactical-level decisions through a spiking neural network based on the Proximal Policy Optimization (PPO) algorithm. The lower-level deduction action in this embodiment generates specific actions of a specific unit in the scenario. The core idea of the PPO algorithm is to limit the difference between the new policy and the old policy during each policy update to maintain the stability of the policy. Specifically, PPO introduces a "surrogate" objective function that takes into account the difference between the new and old policies during each update to ensure that the new policy does not deviate too far from the old policy.
[0113] In addition, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0114] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1the functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 one process or more processes and / or boxes Figure 1 steps for implementing the functions specified in one or more boxes
[0116] It should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element
[0117] Finally, it should be noted that the above description is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once the basic creative concept of the present invention is known, several improvements and refinements can be made without departing from the principle of the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention
Claims
1. A multi-level deduction interactive data processing method and system based on intelligent game, characterized in that: include: A control unit and an intelligent game unit, wherein the control unit includes a user development component and a third-party environment access component, and the intelligent game unit includes an agent service component, a multi-level deduction component, a database component, and a system background component; The user development component is used to configure the basic parameters of the system, control the game process, and provide a visual operation interface; The third-party environment access component is used to provide an information interface for accessing a third-party environment, the user development component selects a third-party environment to access, and the third-party environment access component provides the database component with third-party environment access component information; The agent service component includes a single agent service set and a multi-agent service set, the agent service component is used to encapsulate agent game information, the agent game information is encapsulated as an experience information set and stored in the database component, and the agent service component receives multi-level deduction data from the multi-level deduction component; The multi-level deduction component includes an upper-level deduction action set, an intermediate deduction action set and a lower-level deduction action set, and the upper-level deduction action set, the intermediate deduction action set and the lower-level deduction action set complete multi-level deduction through deduction actions, and the multi-level deduction component generates the multi-level deduction data and sends it to the intelligent agent service component; The database component receives the third-party environment access component information from the third-party environment access component, and the database component receives and provides the experience information set from the agent service component; The system background component provides interface and routing services to connect the user development component, third-party environment access component, agent service component, multi-level deduction component, and database component to the network; The agent service component includes a single-agent service set and a multi-agent service set; The single-agent service set copies the experience information set in the database component into experience information 1 and experience information 2, optimizes experience information 1 using policy gradient optimization to obtain a policy network, optimizes experience information 2 using value gradient optimization to obtain a value network, and obtains deterministic actions from the policy network and sends them to the value network to predict the single-agent action set; The single agent training process is as follows: the control game engine starts, the single agent takes action, obtains and sends the current state, evaluates the situation and calculates the decision, on the one hand, continues to obtain the situation, on the other hand, uses the current state to fill in the experience information, infers the state change to update the parameters, and determines whether the state change expectation converges. If not, return to the inference state change to perform the parameter update step. If converged, store the experience information, and store the experience information in the database together with the other hand to determine whether the game result meets the expectation. The multi-agent service set shares the experience information set in the database component with all agents in the multi-agent service set to implement parallel execution of the experience information set strategy; The single-agent service set and the multi-agent service set process the experience information to obtain game information, and encapsulate and store it in the database component; The system background component also controls user access to the control unit, generates operation and maintenance status information, and maintains the smooth operation of the system.
2. According to claim 1, a multi-level deduction interactive data processing method and system based on intelligent game, characterized in that: The user developed component is deployed on a single control computer; The user development component can adjust the configuration parameters of the agent service component to select and use a single-agent service set and a multi-agent service set; The user development component selects a third-party environment accessed in the third-party environment access component; The user development component views the operation and maintenance status information through a visual operation interface.
3. According to claim 1, a multi-level deduction interactive data processing method and system based on intelligent game, characterized in that: The third-party environment access component can connect to multiple third-party environments; The third-party environment is an existing simulation operation platform; The third-party environment access component stores the used existing simulation operation platform in the database component.
4. According to claim 1, a multi-level deduction interactive data processing method and system based on intelligent game, characterized in that: The multi-level deduction component obtains observation information by observing the intelligent agent service component; The observation information includes campaign-level observation information and tactical-level observation information.
5. According to claim 1, a multi-level deduction interactive data processing method and system based on intelligent game, characterized in that: The multi-level deduction process of the multi-level deduction component is as follows: In the multi-level deduction process 1, the upper-level deduction action set generates a campaign-level decision through a pulse neural network based on a monotone value function decomposition algorithm, and obtains an upper-level deduction action strategy and sends it to the intermediate deduction action set; In the second multi-level deduction process, the intermediate deduction action level generates tactical-level decisions through a pulse neural network based on a monotone value function decomposition algorithm, and obtains an intermediate deduction action strategy which is sent to a lower-level deduction action set; In multi-level deduction process three, the lower-level deduction action level generates tactical-level decisions through a pulse neural network based on a proximal strategy optimization algorithm.
6. According to claim 1, a multi-level deduction interactive data processing method and system based on intelligent game, characterized in that: The database component receives one third-party environment at a time when receiving the third-party environment access component information; The experience information set includes a training experience information set and a testing experience information set.
7. A multi-level deduction interactive data processing method and system based on intelligent game according to claim 6, characterized in that: The training experience information set and the test experience information set generation process are as follows: When the database component currently uses the third-party environment for the first time, the intelligent agent service component receives the experience information set, which is a test experience information set. The intelligent agent service component processes the test experience information set, encapsulates it, and sends it to the database component to obtain a training experience information set. When the database component continues to use the third-party environment, the intelligent agent service component obtains the training experience information set from the database component, processes and encapsulates it to obtain a new training experience information set and sends it to the database component.
8. The method and system for multi-level deduction interactive data processing based on intelligent game according to claim 1, characterized in that: It also includes a viewing component that can record and store the multi-level deduction process and the experience information, and realize playback and double-speed playback according to the user-developed component.
Citation Information
Patent Citations
A smart game adversarial platform
CN112295229B
Open type intelligent game ecological platform
CN113379054A
Man-machine intelligent game system
CN113633994A