A method and system for generating complex game scenarios using artificial intelligence algorithms

By decoupling participant strategy adversarial records from environmental mutation parameters, a non-steady-state policy dependency network is generated. Dynamic preference calculus and environmental constraint conflict variable injection are then performed, solving the coupling distortion problem caused by the separation of environment and policy modeling in existing technologies. The generated game scenario reflects the nonlinear conflict transmission characteristics in real games, improving the credibility of complex game deduction.

CN120975826BActive Publication Date: 2026-04-03BEIJING XINYAN HECHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

The existing technology's separation of environment and strategy modeling weakens the coupling relationship between environmental mutation parameters and strategy decision chains, failing to reflect the dynamic locking effect of environmental disturbances on strategy selection nodes in real games. Path bifurcation lacks the interactive correlation of multi-participant decision chains, and the adversarial branches of generated scenarios lack the interactive correlation of multi-participant decision chains. The inference results have a systematic deviation from the actual game complexity.

Method used

By acquiring a historical game dataset containing multiple participants' strategy adversarial records and environmental mutation parameters, a strategy topology parsing engine is used to decouple the coupling relationship between the participants' strategy adversarial records and environmental mutation parameters, generating a non-steady-state strategy dependency network. Based on this network, dynamic preference calculation is performed to iteratively correct the participants' behavioral preference values, and environmental constraint conflict variables are injected synchronously to generate a complex game scenario with adversarial path bifurcation.

Benefits of technology

It achieves dynamic interlocking of environmental disturbances and strategy decision chains, and the generated game scenario reflects the nonlinear conflict transmission characteristics in real games, which improves the credibility and decision support value of complex game deduction and solves the coupling distortion problem caused by the separation of environment and strategy modeling in traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975826B_ABST
    Figure CN120975826B_ABST
Patent Text Reader

Abstract

This application provides a method and system for generating complex game scenarios using artificial intelligence algorithms. The method first acquires a historical game dataset containing records of multiple participants' policy adversarial actions and environmental mutation parameters. Then, it processes the historical game dataset using a policy topology parsing engine, outputting a non-stationary policy dependency network. Next, it iteratively corrects the participants' behavioral preference values, simultaneously injecting environmental constraint conflict variables from the environmental mutation parameters into the game nodes. Finally, based on the corrected participant behavioral preference values ​​and the injected environmental constraint conflict variables into the game nodes, a complex game scenario with adversarial path bifurcation is generated. The technical solution provided by this application achieves dynamic interlocking between environmental disturbances and policy decision chains, solving the coupling distortion problem caused by the separation of environment and policy modeling in traditional solutions, thus enhancing the credibility and decision support value of complex game deduction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and system for generating complex game scenarios using artificial intelligence algorithms. Background Technology

[0002] In complex game scenarios such as strategy simulation and emergency decision-making, the behavior of multiple participants is influenced by the interplay of dynamic environmental changes (such as market fluctuations) and strategic adversarial forces. This necessitates the real-time generation of highly adversarial decision-making environments with multi-branch evolutionary paths. Traditional manually pre-set scenarios struggle to cover nonlinear conflict nodes and dynamic coupling relationships.

[0003] Current mainstream solutions employ a modeling mechanism that separates environmental parameters from policy behavior. First, a policy decision network is trained using a reinforcement learning model. Then, environmental interference is simulated based on a Bayesian probability model. Finally, environmental parameters are added as independent weights to the policy path generation scenario. This solution triggers path bifurcation by setting a preset environmental impact threshold, thus constructing a basic adversarial scenario.

[0004] However, this scheme has three key flaws. First, due to the separate modeling of environment and strategy, the coupling relationship between environmental mutation parameters and the strategy decision chain is weakened to a linear superposition, failing to reflect the dynamic locking effect of environmental disturbances on strategy selection nodes in real games (such as mutations at specific times directly changing tactical dependencies). Second, path branching relies solely on environmental threshold triggers, without incorporating conflict preference correction mechanisms for strategy behavior. Third, the adversarial branches of the generated scenario lack the interactive correlation of multi-participant decision chains, resulting in a systematic deviation between the inferred results and the actual game complexity. Summary of the Invention

[0005] This application provides a method and system for generating complex game scenarios using artificial intelligence algorithms to solve the problems of environmental mutation and decoupling distortion of the strategy decision chain, as well as the lack of behavioral preference coordination in path forking in the prior art.

[0006] Firstly, this application provides a method for generating complex game scenarios optimized using artificial intelligence algorithms, including:

[0007] Obtain a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label;

[0008] The historical game dataset is processed by a policy topology parsing engine to decouple the coupling relationship between participant policy adversarial records and environmental mutation parameters, so as to output a non-steady-state policy dependency network.

[0009] Based on the non-steady-state policy dependency network, dynamic preference calculus is performed to iteratively correct the participant behavior preference values ​​in the game nodes to obtain the corrected participant behavior preference values. At the same time, the environmental constraint conflict variables in the environmental mutation parameters are injected into the game nodes.

[0010] Based on the modified participant behavior preference values ​​and the conflict variables of the environmental constraints injected into the game nodes, a complex game scenario with adversarial path bifurcation is generated.

[0011] Optionally, a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters is obtained, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label, including:

[0012] Extract raw interaction records from real-world game event logs, where the raw interaction records contain sequences of strategic exchanges between participants and environmental interference events;

[0013] Parse the adversarial behavior identifiers in the strategy confrontation sequence to generate a participant strategy confrontation record. Each participant strategy confrontation record contains continuous decision actions and corresponding decision consequence identifiers.

[0014] The environmental disturbance event is decomposed into discrete environmental mutation parameters, wherein the environmental mutation parameters include the action time window and the influence intensity value;

[0015] For each participant's strategy adversarial record, an associated dynamic decision chain label is created, and the associated participant strategy adversarial record is bound to the environmental mutation parameter to form a historical game dataset.

[0016] Optionally, the historical game dataset is processed by a policy topology parsing engine to decouple the coupling relationship between participant policy adversarial records and environmental mutation parameters, in order to output a non-stationary policy dependency network, including:

[0017] Separate the participant strategy adversarial records and environmental mutation parameters from the historical game dataset, and identify the strategy selection nodes in the participant strategy adversarial records;

[0018] Analyze the time interval of the environmental mutation parameter and determine the overlapping area between the time interval of the parameter and the effective time period of the strategy selection node.

[0019] Within the overlapping region, calculate the interference correlation strength between the strategy selection node and the environmental mutation parameters;

[0020] Eliminate the coupling connections whose interference correlation strength exceeds a preset threshold, and generate a strategy dependency network composed of multiple game nodes;

[0021] In the policy dependency network, the decision dependencies between policy selection nodes are marked to form a non-steady-state policy dependency network.

[0022] Optionally, dynamic preference calculus is performed based on the non-steady-state policy dependency network to iteratively correct the participant's behavioral preference values ​​in the game nodes, resulting in corrected participant behavioral preference values, including:

[0023] Extract policy dependency paths containing multiple decision segments from the non-steady-state policy dependency network, wherein each decision segment consists of the decision dependency relationship between two consecutive game nodes;

[0024] Identify the policy evolution stage markers in the policy dependency path, and divide the decision segment sequence according to the policy evolution stage markers;

[0025] Obtain the behavior records of each decision segment within the strategy evolution stage, and compare them with preset dependency rules to generate strategy selection bias;

[0026] Calculate the decision inertia offset for each decision segment within each of the strategy evolution stages;

[0027] Based on the decision inertia offset, the participants' behavioral preference values ​​are adjusted at the corresponding game nodes to form intermediate modified preference values;

[0028] By aggregating the intermediate modified preference values ​​at multiple strategy evolution stages, the modified participant behavior preference values ​​at the game nodes are obtained.

[0029] Optionally, a decision inertia offset is calculated for each decision segment within each strategy evolution stage, including:

[0030] The behavioral preference values ​​of participants in the preceding game nodes in the decision segment are extracted as the inertial transmission benchmark.

[0031] Obtain the strategy selection deviation value corresponding to the current decision segment;

[0032] The preference correction coefficient is determined based on the strength of the decision dependency relationship in the decision segment;

[0033] The inertial transfer benchmark and the strategy selection deviation value are fused through a conflict correlation function, and then the preference correction coefficient is superimposed to generate the decision inertial offset.

[0034] Optionally, the environmental constraint conflict variables in the environmental mutation parameters are synchronously injected into the game node, including:

[0035] Extract environmental mutation parameters from the historical game dataset, and separate dynamic constraint relationship quantification terms from the environmental mutation parameters;

[0036] The dynamic constraint relationship quantification term is converted into an environmental constraint conflict variable, which includes the target identifier and the conflict intensity level.

[0037] Locate the target game node in the non-steady-state policy dependency network, wherein the decision dependency relationship of the target game node is matched with the action object identifier;

[0038] The environmental constraint conflict variables are merged into the state descriptor of the target game node.

[0039] Optionally, based on the modified participant behavioral preference values ​​and the environmental constraint conflict variables injected into the game nodes, a complex game scenario with adversarial path bifurcation is generated, including:

[0040] The initial decision evolution path is extracted based on the aforementioned non-steady-state policy dependency network;

[0041] The corrected participant behavioral preference values ​​are mapped to the state descriptors of the corresponding game nodes to form preference reinforcement decision trajectories;

[0042] Detect the conflict activation values ​​of environmental constraint conflict variables in the game nodes;

[0043] When the conflict activation value exceeds the preset scenario fork threshold, a conflict interference event is triggered at the corresponding game node.

[0044] In response to the conflict interference event, the decision dependency relationship is reconstructed based on the conflict type identifier of the environmental constraint conflict variable, and a conflict interference path is generated;

[0045] At the game node location that triggers the conflict interference event, the preference-enhancing decision trajectory is combined with the conflict interference path to construct a complex game scenario that includes adversarial path bifurcation.

[0046] Secondly, this application provides a complex game scenario generation system optimized using artificial intelligence algorithms, comprising:

[0047] The acquisition module is used to acquire a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label;

[0048] The decoupling module is used to process the historical game dataset through the policy topology parsing engine, decouple the coupling relationship between the participants' policy adversarial records and the environmental mutation parameters, so as to output a non-steady-state policy dependency network.

[0049] The correction module is used to perform dynamic preference calculation based on the non-steady-state policy dependency network, iteratively correct the participant behavior preference values ​​in the game node, obtain the corrected participant behavior preference values, and simultaneously inject the environmental constraint conflict variables in the environmental mutation parameters into the game node.

[0050] The generation module is used to generate a complex game scenario with adversarial path bifurcation based on the corrected participant behavior preference values ​​and the environmental constraint conflict variables injected into the game nodes.

[0051] Thirdly, this application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to realize a method for generating complex game scenarios using artificial intelligence algorithms as described in the first aspect above.

[0052] Fourthly, this application provides a computer storage medium storing a computer program, which, when executed by a computer, implements a method for generating complex game scenarios using artificial intelligence algorithms as described in the first aspect.

[0053] This application employs a strategy topology analysis engine to deeply decouple historical game data, accurately removing spurious correlations between environmental mutation parameters and participant strategy adversarial records, and constructing a non-stationary network that truly reflects the dynamics of strategy dependencies. Based on this, dynamic preference calculation and environmental constraint conflict variable injection are performed simultaneously, enabling game nodes to simultaneously bear the results of strategy behavior preference corrections and the intensity of environmental disturbance conflicts. Finally, a complex game scenario containing adversarial path bifurcation is generated based on a dual-source collaborative driving mechanism. This technical solution achieves dynamic interlocking of environmental disturbances and strategy decision chains for the first time, solving the coupling distortion problem caused by the separation of environment and strategy modeling in traditional solutions. The generated adversarial scenario possesses the nonlinear conflict transmission characteristics of real games.

[0054] Furthermore, regarding the scenario bifurcation construction stage, an innovative conflict interference event triggering mechanism is adopted. When the conflict activation value exceeds a preset threshold, the decision dependency relationship is dynamically reconstructed, generating a native interference path that strictly matches the conflict type. By physically combining the preference-enhanced decision trajectory with the reconstructed conflict interference path at a designated game node, an adversarial path bifurcation with dual imprints of behavioral preferences and environmental conflict is formed. This breakthrough solves the defect in existing solutions where path bifurcation is disconnected from the participant's decision chain. The generated game scenario retains the historical inertia characteristics of strategy evolution while realistically reflecting the fission effect of decision dependency relationships caused by environmental mutations, significantly improving the credibility and decision support value of complex game deduction.

[0055] These or other aspects of this application will become more apparent from the description of the following embodiments. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 A flowchart of a method for generating complex game scenarios using artificial intelligence algorithms, as provided in this application, is shown.

[0058] Figure 2 A schematic diagram of the structure of a complex game scenario generation system optimized by artificial intelligence algorithms provided in this application is shown;

[0059] Figure 3 A schematic diagram of the structure of a computing device provided in this application is shown. Detailed Implementation

[0060] To enable those skilled in the art to better understand the present application, the technical solution of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0061] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.

[0062] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0063] Figure 1 This application provides a flowchart of a complex game scenario optimized using artificial intelligence algorithms, such as... Figure 1 As shown, the method includes:

[0064] Step 101: Obtain a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label.

[0065] Optionally, step 101 may specifically include the following steps:

[0066] Step 1011: Extract the original interaction records from the real-world game event log, wherein the original interaction records contain the strategy confrontation sequence between the participants and environmental interference events;

[0067] Step 1012: Parse the adversarial behavior identifiers in the strategy confrontation sequence to generate participant strategy confrontation records. Each participant strategy confrontation record contains continuous decision actions and corresponding decision consequence identifiers.

[0068] Step 1013: Decompose the environmental disturbance event into discrete environmental mutation parameters, wherein the environmental mutation parameters include the action time window and the influence intensity value;

[0069] Step 1014: Associate dynamic decision chain labels with each participant's strategy adversarial record, and bind the associated participant strategy adversarial records with the environmental mutation parameters to form a historical game dataset.

[0070] In the above scheme, the original interaction records are unprocessed data extracted from a business competition event database (such as e-commerce platform business war logs), containing strategy confrontation sequences (market competition actions between enterprises) and environmental disturbance events (sudden policies / market changes). Participant strategy confrontation records are formatted records generated by parsing the strategy confrontation sequences, containing continuous decision actions (such as price reductions / new product launches) and decision consequence identifiers (market share rise / fall markers), reconstructing enterprise competitive strategies. Environmental mutation parameters are structured data decomposed from environmental disturbance events, containing the time window of effect (policy effective period / market fluctuation period) and impact intensity values ​​(policy strength / fluctuation amplitude level). Dynamic decision chain labels are chain-like labels linking decision actions and consequences (such as "price reduction - market share increase → price increase - customer churn"), identifying the causal logic of business decisions. The historical game dataset is an integrated dataset of strategy records and environmental parameters bound to decision chain labels.

[0071] In this embodiment, firstly, the business competition database (such as retail platform business war logs) is read through the data interface in step 1011 to extract the original interaction records. This includes strategy confrontation sequences (e.g., "Platform A lowers prices on Monday → Platform B promotes on Tuesday") and environmental interference events (e.g., "New advertising law is released on Wednesday"). Secondly, in step 1012, the behavior parsing module identifies competitive behavior identifiers (e.g., the "price reduction" identifier "PRC_DOWN") in the strategy confrontation sequences and converts them into structured strategy confrontation records. Each record contains continuous decision actions (e.g., ["price reduction", "increase advertising"]) and consequence identifiers (e.g., ["market share +5%", "customer -3%"]). Next, in step 1013, the environmental interference events (e.g., the new advertising law) are processed by the event decomposer and broken down into environmental mutation parameters. The time window of action is the policy's effective period (e.g., May 1st - December 31st), and the impact intensity value is mapped through the rule base (e.g., "strict restriction" → intensity value 0.9). Finally, in step 1014, dynamic decision chain labels (such as "PRC_DOWN-SHARE_UP→AD_ADD-CUST_DOWN") are generated for each policy record in chronological order and bound to mutation parameters to form a historical dataset.

[0072] Following the specific implementation of the previous step, in the cross-border logistics competition case: Step 1011 extracts records from the logistics company database: strategy confrontation sequence ("Company X lowers prices by 10% on Monday → Company Y launches same-day delivery service on Tuesday"), environmental interference event ("Fuel tax increase on Wednesday"). Step 1012 parses and generates Company X's strategy record: decision action = ["price reduction", "service upgrade"], consequence identifier = ["orders +15%", "cost +8%"]. Step 1013 decomposes the fuel tax event: impact time window = June 1st - December 31st, impact intensity value = 0.85 (high cost shock). Step 1014 associates the dynamic decision chain label "price reduction - order increase → service upgrade - cost increase", binds it with the fuel tax parameter, and stores it in the dataset. E-commerce platform competition case: Step 1011 extracts platform business war logs: strategy sequence ("Platform A offers limited-time discounts on Sundays → Platform B provides member subsidies on Mondays"), interference event ("New data security regulations are introduced on Tuesday"). Step 1012 Generate a record for Platform A: Action = ["Limited-time discount", "Increased benefits"], Consequence = ["Traffic +20%", "Retention -5%"]. Step 1013 Decompose the new regulation event: Time window = Immediate effect, Intensity value = 0.75 (moderate compliance cost).

[0073] Step 1014 binds the tags "Discount-Traffic Increase → Benefits-Retention Decrease" with the new regulations parameters.

[0074] This step automates the analysis of business competition data, transforming raw business battle records into structured datasets that carry causal chains of decision-making. It accurately quantifies the scope and intensity of environmental interference such as policies and markets, fully preserving the action-outcome relationships of corporate competitive strategies, and providing a high-fidelity data foundation for business game theory simulations.

[0075] Step 102: Process the historical game dataset through the policy topology parsing engine to decouple the coupling relationship between the participant policy adversarial records and the environmental mutation parameters, so as to output a non-steady-state policy dependency network.

[0076] Optionally, step 102 may specifically include the following steps:

[0077] Step 1021: Separate the participant strategy adversarial records and environmental mutation parameters in the historical game dataset, and identify the strategy selection nodes in the participant strategy adversarial records;

[0078] Step 1022: Analyze the time interval of the environmental mutation parameter and determine the overlapping area between the time interval of the parameter and the effective time period of the strategy selection node.

[0079] Step 1023: Calculate the interference correlation strength between the strategy selection node and the environmental mutation parameter within the overlapping region;

[0080] Step 1024: Eliminate the coupling connections whose interference correlation strength exceeds a preset threshold, and generate a strategy dependency network composed of multiple game nodes;

[0081] Step 1025: Mark the decision dependencies between policy selection nodes in the policy dependency network to form a non-steady-state policy dependency network.

[0082] In the above scheme, the strategy selection node is a key decision point in the enterprise's strategy confrontation record (such as the critical moment of launching a new product / adjusting pricing), used to locate the position of environmental disturbances. The overlapping area is the interval where the time period of the environmental change parameter and the effective time period of the enterprise's strategy node coincide (such as the overlap between the implementation period of new regulations and the promotional activity period), used to lock in the actual scope of impact. The interference correlation strength quantifies the degree of impact of environmental changes on the strategy selection node (such as the impact of tax rate adjustments on pricing decisions). The strategy dependency network is a network of multiple game nodes (representing the enterprise's decision state) connected by competitive relationships, forming the basic framework of business game theory. The non-steady-state strategy dependency network is a network where the competitive relationship between enterprises changes dynamically with the market (such as the transformation of cooperation into competition), reflecting the uncertainty of the business environment.

[0083] In this embodiment, firstly, step 1021 separates the enterprise strategy records and environmental parameters in the historical dataset, and uses decision point identification technology to locate the strategy selection node (e.g., extracting the "price reduction decision time" as a node from the "Company X price reduction → Company Y subsidy" record). Secondly, step 1022 analyzes the time interval of the environmental change parameter (e.g., January-June for the new tax policy) and compares it with the effective time period of the strategy node (e.g., February-April for Company X's price reduction activity) to determine the overlapping area (February-April). Next, step 1023 calculates the interference correlation strength of the environmental change on the strategy node within the overlapping area using an impact assessment model (e.g., tax rate increase × overlap ratio of activity periods). Then, step 1024, when the interference correlation strength exceeds a preset threshold (e.g., strength > 0.7), severs the spurious correlation between the node and the environmental parameter, generating a decoupled pure strategy dependency network (retaining only the competitive relationship between enterprises). Finally, through step 1025, the decision dependencies between nodes in the strategy dependency network (such as the competitive linkage of "Company X pricing → Company Y promotion") are marked to form a non-stationary network that reflects market changes.

[0084] Following the specific implementation of the previous step, the e-commerce platform competition case: Step 1021: Separate the dataset: Platform X's strategy records (including the "March promotion" node) and the new advertising regulation parameters (intensity value 0.8). Step 1022: Analyze the regulation's effective period (April-September) and the promotion node time (March-May), with the overlapping area being April-May. Step 1023: Calculate the interference strength: Regulation strength 0.8 × time overlap ratio 60% = intensity value 0.48. Step 1024: Since 0.48 < the critical value 0.7, retain the association. Step 1025: Mark the competitive dependency relationship "Platform X promotion → Platform Y subsidy". Logistics company competition case: Step 1021: Separate the data: Company A's strategy records (including the "June acceleration service" node) and the fuel tax parameters (intensity 0.9). Step 1022: Analyze the fuel tax period (June-December) and the acceleration service period (June-August), with the overlapping area being June-August. Step 1023: Calculate the intensity: Tax rate intensity 0.9 × Time overlap rate 100% = Intensity value 0.9. Step 1024: Since 0.9 > 0.7, eliminate the spurious association between "Fuel Tax → Speed-up Service". Step 1025: Mark the market adversarial relationship between "Company A's Speed-up Service → Company B's Price Reduction".

[0085] This step precisely identifies the overlap between policy impacts and corporate decision-making, stripping away non-critical environmental interference while preserving genuine competitive relationships. The generated non-stationary network avoids the distortion caused by the mechanical superposition of environment and strategy in traditional solutions, and realistically reflects the uncertainty of market games through dynamic dependency labeling, providing a precise underlying model for corporate strategic deduction.

[0086] Step 103: Perform dynamic preference calculus based on the non-steady-state policy dependency network, iteratively correct the participant behavior preference values ​​in the game node, obtain the corrected participant behavior preference values, and simultaneously inject the environmental constraint conflict variables in the environmental mutation parameters into the game node.

[0087] Optionally, step 103 may specifically include the following steps:

[0088] Step 1031: Extract policy dependency paths containing multiple decision segments from the non-steady-state policy dependency network, wherein each decision segment consists of the decision dependency relationship between two consecutive game nodes;

[0089] Step 1032: Identify the policy evolution stage markers in the policy dependency path, and divide the decision segment sequence according to the policy evolution stage markers;

[0090] Step 1033: Obtain the behavior record values ​​of each decision segment within the strategy evolution stage, and compare them with the preset dependency rules to generate strategy selection bias;

[0091] Step 1034: Calculate the decision inertia offset for each decision segment within each strategy evolution stage;

[0092] Optionally, step 1034 may specifically include the following steps:

[0093] The participant behavior preference values ​​of the preceding game nodes in the decision segment are extracted as the inertia transmission benchmark; the strategy selection deviation value corresponding to the current decision segment is obtained; the preference correction coefficient is determined according to the decision dependency strength of the decision segment; the inertia transmission benchmark and the strategy selection deviation value are fused through a conflict correlation function, and then the preference correction coefficient is superimposed to generate the decision inertia offset.

[0094] Extract environmental mutation parameters from the historical game dataset, and separate dynamic constraint relationship quantification terms from the environmental mutation parameters; convert the dynamic constraint relationship quantification terms into environmental constraint conflict variables, which include an object identifier and a conflict intensity level; locate the target game node in the non-steady-state policy dependency network, where the decision dependency of the target game node matches the object identifier; merge the environmental constraint conflict variables into the state descriptor of the target game node. Step 1035: Adjust the participant's behavioral preference values ​​at the corresponding game node according to the decision inertia offset to form intermediate corrected preference values;

[0095] Step 1036: Aggregate the intermediate modified preference values ​​of the multiple strategy evolution stages to obtain the modified participant behavior preference values ​​of the game node.

[0096] In the above scheme, the decision segment is a competitive relationship fragment (e.g., "price reduction → promotion") between two consecutive firm decision nodes in a non-steady-state strategy dependency network, constituting the basic unit of preference calculus. The strategy evolution stage marker is a key point in the strategy path that identifies the transition of business state (e.g., the transition from "market expansion → profit defense"), used to divide the competitive stages. The strategy selection bias is the difference between the actual behavior value of the decision segment (e.g., the actual price reduction) and the preset rule (e.g., "maximum price reduction of 10% under cost constraints"). The decision inertia offset is a correction amount that integrates the behavioral preferences of previous nodes (e.g., historical promotional intensity) with the current strategy bias. The environmental constraint conflict variable is a structured conflict carrier for the transformation of environmental mutation parameters, containing an identifier of the object of action (pointing to a specific firm node) and a conflict intensity level (e.g., policy shock level).

[0097] In this embodiment, firstly, step 1031 extracts the policy dependency path (e.g., "X platform price reduction → Y platform subsidy") from the non-steady-state policy dependency network and decomposes it into decision segments ("price reduction-subsidy" relationship segments). Secondly, step 1032 identifies the policy evolution stage markers in the path (e.g., "subsidy war begins" marker) and divides the decision segment sequence according to the markers (e.g., stage 1: price war segment; stage 2: service upgrade segment). Next, step 1033: obtain the behavior record values ​​of each decision segment (e.g., the actual price reduction of 15% on platform X), and compare them with the preset dependency rules (e.g., "the rule requires a price reduction of ≤10%)" to generate a strategy selection bias (deviation value = +5%). Then, step 1034: calculate the decision inertia offset—① extract the preference value of the preceding node (e.g., the preference value of 0.8 during the market expansion period) as the inertia benchmark; ② obtain the current strategy selection bias value (+5%); ③ determine the correction coefficient (0.05) based on the decision dependency strength of this segment (e.g., the competitive correlation degree of 0.7); ④ fuse the inertia benchmark and the deviation value (0.8 × 5%) through the conflict correlation function, and superimpose the correction coefficient to generate the offset (4% + 0.05 = 4.05%). Environmental variable injection sub-process: a) extract environmental mutation parameters The process involves: a) separating the dynamic constraint relationship quantification item ("compliance cost increase = 0.6") from the data (e.g., the new advertising law); b) converting it into an environmental constraint conflict variable (object of action = "advertising placement node", conflict intensity = "medium"); c) locating the target game node matching the identifier (e.g., the advertising placement node on platform Y); d) adding "compliance cost = medium" to the node state description and conflict variable. Then, in step 1035, the node preference value is adjusted using the decision inertia offset (e.g., original preference value 0.7 → corrected to 0.7 + 4.05% = 0.745) to form an intermediate correction value. Finally, in step 1036, multiple stage intermediate values ​​are aggregated (e.g., price war stage value 0.745 + service stage value 0.68) to obtain the final corrected node preference value (0.712).

[0098] Following the specific implementation of the previous step, in the case of a chain supermarket competition: Step 1031: Extract the path "Supermarket A price reduction → Supermarket B member day", and break it down into the "price reduction - member day" decision segment. Step 1032: Identify the "Member Day" node and mark it as the "customer protection stage". Step 1033: Obtain the price reduction behavior value of Supermarket A (reduction of 12%), and compare it with the preset rule "cost constraint reduction ≤ 10%" to generate a deviation of +2%. Step 1034: Calculate the offset: ① Preceding expansion period preference value 0.75; ② Current deviation +2%; ③ Dependency strength 0.6 → correction coefficient 0.03; ④ Fusion calculation: 0.75 × 2% = 1.5% → offset after superposition coefficient = 1.53%.

[0099] Environmental Injection: a) Extract new food safety regulations and separate the quantifiable item "Increase in testing costs = 0.7"; b) Transform conflict variables (object of action = "fresh produce node", intensity = "high"); c) Locate the fresh produce section node in supermarket B; d) Add "Food safety compliance cost = high" to the node status. Step 1035: Adjust the preference value of supermarket B's member day node: original value 0.8 → corrected to 0.8 + 1.53% = 0.815. Step 1036: Aggregate the two-stage values ​​(price war 0.815 + customer protection 0.78) to obtain the final preference value of 0.797. Logistics Company Case: Step 1031: Extract the path "Company C speeds up → Company D lowers prices". Step 1033: Obtain the speed-up behavior value of Company C (timeliness improvement of 20%), compare it with the rule "maximum speed-up of 15%" to generate a deviation of +5%. Environmental injection: a) Extract the fuel tax policy and separate the quantitative item "fuel cost increase = 0.8"; b) Transform the conflict variable (object of action = "transportation node", intensity = "high"); c) Locate the variable to be injected into the transportation node of company D.

[0100] This step quantifies strategy execution deviations in stages and integrates historical preference inertia to generate dynamic correction quantities, enabling corporate decision-making preferences to adaptively adjust with the competitive process. Simultaneously injected environmental conflict variables precisely pinpoint policy-sensitive nodes, forming a dual-track driving mechanism of "competitive strategy preference correction + policy impact labeling," realistically recreating the dynamic interlocking effect between market behavior and the policy environment in business games.

[0101] Step 104: Generate a complex game scenario with adversarial path bifurcation based on the corrected participant behavior preference values ​​and the environmental constraint conflict variables injected into the game nodes.

[0102] Optionally, step 104 may specifically include the following steps:

[0103] Step 1041: Extract the initial decision evolution path based on the non-steady-state policy dependency network;

[0104] Step 1042: Map the corrected participant behavior preference values ​​to the state descriptors of the corresponding game nodes to form preference reinforcement decision trajectories;

[0105] Step 1043: Detect the conflict activation value of the environmental constraint conflict variable in the game node;

[0106] Step 1044: When the conflict activation value exceeds the preset scenario fork threshold, a conflict interference event is triggered at the corresponding game node.

[0107] Step 1045: In response to the conflict interference event, reconstruct the decision dependency relationship based on the conflict type identifier of the environmental constraint conflict variable, and generate a conflict interference path;

[0108] Step 1046: At the game node position where the conflict interference event is triggered, the preference enhancement decision trajectory is combined with the conflict interference path to construct a complex game scenario that includes adversarial path bifurcation.

[0109] In the above scheme, the initial decision evolution path is the basic competitive route of enterprises extracted from the non-steady-state strategy dependency network (such as the node chain of "price reduction → promotion → membership growth"), which constitutes the scenario generation baseline. The preference reinforcement decision trajectory is an enhanced path (such as "high promotion intensity trajectory") formed by embedding the modified behavioral preference value (such as risk preference coefficient) into the node state descriptor. The conflict activation value is the threshold indicator for triggering path bifurcation in the environmental constraint conflict variables (such as the policy compliance cost value), used to detect the critical point of environmental mutation impact. The conflict interference event is the path reconstruction instruction generated when the conflict activation value exceeds the threshold (such as "ad traffic restriction event"), driving the scenario bifurcation. The conflict interference path is a new path branch generated based on the conflict type identifier (such as "cost constraint type") to reconstruct the decision dependency relationship (such as "shift to offline promotion to replace online advertising").

[0110] In this embodiment, firstly, step 1041 extracts the initial decision evolution path (e.g., the node sequence "Supermarket A price reduction → Supermarket B member day → customer growth") from the non-steady-state policy dependency network. Secondly, step 1042 writes the corrected participant behavior preference value (e.g., the preference value of Supermarket B member day node is 0.8) into the corresponding node state descriptor, forming a preference-reinforced decision trajectory (e.g., the path marked "high member benefit investment"). Next, step 1043 detects the conflict activation value of environmental constraint conflict variables in the game nodes (e.g., the "new advertising law compliance cost value 0.85" of the Supermarket B advertising placement node). Then, step 1044: when the conflict activation value exceeds the preset scenario bifurcation threshold (e.g., threshold 0.8), a conflict interference event is triggered at the corresponding node (generating an "advertising compliance risk event"). Then, step 1045 responds to the event, reconstructs the decision dependency relationship according to the conflict type identifier (e.g., "cost constraint type") (e.g., changing "online advertising dependency" to "community promotion dependency"), and generates a conflict interference path (e.g., a new branch "price reduction → community ground promotion"). Finally, in step 1046, at the node position of the triggered event (ad placement node), the preference reinforcement decision trajectory (original price reduction → member day path) and the conflict interference path (new community promotion path) are physically combined to construct a game scenario with branching (forming a dual-path scenario).

[0111] Following the specific implementation of the previous step, in the case of chain supermarket competition: Step 1041 extracts the initial path "Supermarket A price reduction → Supermarket B member day → customer growth". Step 1042 maps the modified preference value of Supermarket B member day node 0.8 to the state descriptor, forming the "high equity investment" trajectory. Step 1043 detects the new advertising law compliance cost value of 0.85 at the advertising node. Step 1044 triggers the "advertising compliance risk event" because 0.85 > threshold 0.8. Step 1045 reconstructs the dependency relationship according to the conflict type "cost constraint": the original "online advertising → customer growth" is changed to "community promotion → customer growth", generating a new path "price reduction → community promotion". Step 1046 combines the original trajectory and the new path at the advertising node position: Main path: Supermarket A price reduction → Supermarket B high equity member day (original trajectory) Branch path: Supermarket A price reduction → Supermarket B community promotion (new path) E-commerce platform case: Step 1041 extracts the path "Platform X subsidy → Platform Y live streaming traffic". Step 1043 detected a data security compliance value of 0.9 for the live streaming node (as required by the new regulations). Step 1044 triggered a "data compliance event". Step 1045, based on the conflict type "technical limitation", refactored the dependency to "short video traffic generation", generating a new path "subsidy → short video traffic generation". Step 1046 formed a fork scenario: two parallel paths: subsidy + live streaming and subsidy + short video.

[0112] This step precisely captures the impact points of policy mutations by using conflict activation thresholds, and reconstructs new competitive paths that conform to regulatory characteristics based on conflict types. By physically combining the original strategy trajectory with the new path, an adversarial bifurcation scenario is formed that retains the characteristics of corporate decision-making preferences while reflecting environmental constraints, realistically recreating the dynamic process of "policy mutations triggering a shift in competitive paths" in market games.

[0113] Figure 2 This application provides a schematic diagram of the structure of a complex game scenario generation system optimized using artificial intelligence algorithms, as shown below. Figure 2 As shown, the system includes:

[0114] The acquisition module 21 is used to acquire a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label.

[0115] The decoupling module 22 is used to process the historical game dataset through the policy topology parsing engine, decouple the coupling relationship between the participant's policy adversarial records and the environmental mutation parameters, so as to output a non-steady-state policy dependency network.

[0116] The correction module 23 is used to perform dynamic preference calculation based on the non-steady-state policy dependency network, iteratively correct the participant behavior preference values ​​in the game node, obtain the corrected participant behavior preference values, and simultaneously inject the environmental constraint conflict variables in the environmental mutation parameters into the game node.

[0117] The generation module 24 is used to generate a complex game scenario with adversarial path bifurcation based on the modified participant behavior preference values ​​and the environmental constraint conflict variables injected into the game nodes.

[0118] Figure 2 The aforementioned complex game scenario generation system optimized using artificial intelligence algorithms can execute... Figure 1 The implementation principle and technical effects of the complex game scenario generation method optimized by artificial intelligence algorithms described in the illustrated embodiment will not be repeated here. The specific methods by which each module and unit performs operations in the complex game scenario generation system optimized by artificial intelligence algorithms in the above embodiments have been described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0119] In one possible design, Figure 2 The complex game scenario generation system optimized using artificial intelligence algorithms in the illustrated embodiment can be implemented as a computing device, such as... Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0120] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 32.

[0121] The processing component 32 is used for the above Figure 1 The embodiment describes a method for generating complex game scenarios using artificial intelligence algorithms.

[0122] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0123] Storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0124] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.

[0125] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.

[0126] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.

[0127] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.

[0128] This application also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 The embodiment shown illustrates a method for generating complex game scenarios using artificial intelligence algorithms.

[0129] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0130] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for generating complex game scenarios optimized using artificial intelligence algorithms, characterized in that, include: Obtain a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label; The historical game dataset is processed by a policy topology parsing engine to decouple the coupling relationship between participant policy adversarial records and environmental mutation parameters, so as to output a non-steady-state policy dependency network. Based on the non-steady-state policy dependency network, dynamic preference calculus is performed to iteratively correct the participant behavior preference values ​​in the game nodes to obtain the corrected participant behavior preference values. At the same time, the environmental constraint conflict variables in the environmental mutation parameters are injected into the game nodes. Based on the modified participant behavior preference values ​​and the conflict variables of the environmental constraints injected into the game nodes, a complex game scenario with adversarial path bifurcation is generated. The step of generating a complex game scenario with adversarial path bifurcation based on the modified participant behavioral preference values ​​and the environmental constraint conflict variables injected into the game nodes includes: The initial decision evolution path is extracted based on the aforementioned non-steady-state policy dependency network; The corrected participant behavioral preference values ​​are mapped to the state descriptors of the corresponding game nodes to form preference reinforcement decision trajectories; Detect the conflict activation values ​​of environmental constraint conflict variables in the game nodes; When the conflict activation value exceeds the preset scenario fork threshold, a conflict interference event is triggered at the corresponding game node. In response to the conflict interference event, the decision dependency relationship is reconstructed based on the conflict type identifier of the environmental constraint conflict variable, and a conflict interference path is generated; At the game node location that triggers the conflict interference event, the preference-enhancing decision trajectory is combined with the conflict interference path to construct a complex game scenario that includes adversarial path bifurcation.

2. The method according to claim 1, characterized in that, Obtain a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label, including: Extract raw interaction records from real-world game event logs, where the raw interaction records contain sequences of strategic exchanges between participants and environmental interference events; Parse the adversarial behavior identifiers in the strategy confrontation sequence to generate a participant strategy confrontation record. Each participant strategy confrontation record contains continuous decision actions and corresponding decision consequence identifiers. The environmental disturbance event is decomposed into discrete environmental mutation parameters, wherein the environmental mutation parameters include the action time window and the influence intensity value; For each participant's strategy adversarial record, an associated dynamic decision chain label is created, and the associated participant strategy adversarial record is bound to the environmental mutation parameter to form a historical game dataset.

3. The method according to claim 1, characterized in that, The historical game dataset is processed by a policy topology parsing engine to decouple the coupling relationship between participant policy adversarial records and environmental mutation parameters, in order to output a non-stationary policy dependency network, including: Separate the participant strategy adversarial records and environmental mutation parameters from the historical game dataset, and identify the strategy selection nodes in the participant strategy adversarial records; Analyze the time interval of the environmental mutation parameter and determine the overlapping area between the time interval of the parameter and the effective time period of the strategy selection node. Within the overlapping region, calculate the interference correlation strength between the strategy selection node and the environmental mutation parameters; Eliminate the coupling connections whose interference correlation strength exceeds a preset threshold, and generate a strategy dependency network composed of multiple game nodes; In the policy dependency network, the decision dependencies between policy selection nodes are marked to form a non-steady-state policy dependency network.

4. The method according to claim 1, characterized in that, Based on the aforementioned non-steady-state policy dependency network, dynamic preference calculus is performed to iteratively correct the participant's behavioral preference values ​​at the game nodes, resulting in corrected participant behavioral preference values, including: Extract policy dependency paths containing multiple decision segments from the non-steady-state policy dependency network, wherein each decision segment consists of the decision dependency relationship between two consecutive game nodes; Identify the policy evolution stage markers in the policy dependency path, and divide the decision segment sequence according to the policy evolution stage markers; Obtain the behavior records of each decision segment within the strategy evolution stage, and compare them with preset dependency rules to generate strategy selection bias; Calculate the decision inertia offset for each decision segment within each of the strategy evolution stages; Based on the decision inertia offset, the participants' behavioral preference values ​​are adjusted at the corresponding game nodes to form intermediate modified preference values; By aggregating the intermediate modified preference values ​​at multiple strategy evolution stages, the modified participant behavior preference values ​​at the game nodes are obtained.

5. The method according to claim 4, characterized in that, Calculate the decision inertia offset for each decision segment within each of the aforementioned strategy evolution stages, including: The behavioral preference values ​​of participants in the preceding game nodes in the decision segment are extracted as the inertial transmission benchmark. Obtain the strategy selection deviation value corresponding to the current decision segment; The preference correction coefficient is determined based on the strength of the decision dependency relationship in the decision segment; The inertial transfer benchmark and the strategy selection deviation value are fused through a conflict correlation function, and then the preference correction coefficient is superimposed to generate the decision inertial offset.

6. The method according to claim 1, characterized in that, Synchronously injecting environmental constraint conflict variables from the environmental mutation parameters into the game node includes: Extract environmental mutation parameters from the historical game dataset, and separate dynamic constraint relationship quantification terms from the environmental mutation parameters; The dynamic constraint relationship quantification term is converted into an environmental constraint conflict variable, which includes the target identifier and the conflict intensity level. Locate the target game node in the non-steady-state policy dependency network, wherein the decision dependency relationship of the target game node is matched with the action object identifier; The environmental constraint conflict variables are merged into the state descriptor of the target game node.

7. A complex game scenario generation system optimized using artificial intelligence algorithms, applied to the complex game scenario generation method optimized using artificial intelligence algorithms according to any one of claims 1-6, characterized in that, include: The acquisition module is used to acquire a historical game dataset containing multiple participant strategy adversarial records and environmental mutation parameters, wherein each participant strategy adversarial record in the historical game dataset is associated with a dynamic decision chain label; The decoupling module is used to process the historical game dataset through the policy topology parsing engine, decouple the coupling relationship between the participants' policy adversarial records and the environmental mutation parameters, so as to output a non-steady-state policy dependency network. The correction module is used to perform dynamic preference calculation based on the non-steady-state policy dependency network, iteratively correct the participant behavior preference values ​​in the game node, obtain the corrected participant behavior preference values, and simultaneously inject the environmental constraint conflict variables in the environmental mutation parameters into the game node. The generation module is used to generate a complex game scenario with adversarial path bifurcation based on the corrected participant behavior preference values ​​and the environmental constraint conflict variables injected into the game nodes.

8. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a method for generating complex game scenarios using artificial intelligence algorithms as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that, The system contains a computer program that, when executed by a computer, implements a method for generating complex game scenarios using artificial intelligence algorithms as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Hierarchical multi-agent game confrontation and collaborative decision-making algorithm based on federated learning

    CN119443312A

  • Big data analysis and prediction-based strategy dynamic optimization system and method

    CN120337978A