Key action extraction method and device based on state transition diagram and storage medium
By constructing a state transition graph and using frequency-weighted edge betweenness to evaluate action importance, the problem of low efficiency in extracting key actions in existing technologies is solved, and more accurate action extraction and decision logic reflection are achieved.
Patent Information
- Application Number
- CN202511684801.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies are inefficient and unable to accurately assess the impact of actions on the upstream and downstream of the action sequence when extracting key actions from relevant data generated by intelligent game models. This results in incomplete action strategies that fail to accurately reflect the core decision-making logic of the intelligent game model.
By constructing a state transition graph, the frequency-weighted betweenness number of each edge is determined. Actions are sorted according to importance information, and redundant edge removal operations are iteratively executed until a preset stopping condition is met, thus identifying key actions.
It improves the efficiency of extracting key actions, ensuring that the extracted actions are more complete and can more accurately reflect the core decision-making logic of the intelligent game model.
Smart Images

Figure CN121705693A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interpretable game model technology, and in particular to a method, device and storage medium for extracting key actions based on state transition diagrams. Background Technology
[0002] With the development of artificial intelligence technology, intelligent game theory models have been applied in scenarios such as game competition and decision-making. During the competition, these models generate relevant data, which typically contains a large number of actions. However, not all actions are equally important to the final game outcome. Therefore, accurately extracting key actions from the data generated by intelligent game theory models has become a pressing technical problem. Summary of the Invention
[0003] This application provides a method, device, and storage medium for extracting key actions based on state transition diagrams, in order to solve the technical problem of how to accurately extract key actions from relevant data generated by intelligent game models.
[0004] In a first aspect, embodiments of this application provide a method for extracting key actions based on a state transition diagram, including: Step S10: Obtain the sequence sample set generated by the intelligent game model during the adversarial process; Step S20: Construct a state transition graph based on the sequence sample set. The state transition graph includes nodes composed of states and edges composed of actions. Step S30: Determine the frequency-weighted edge betweenness of the action corresponding to each edge in the state transition graph; Step S40: Determine the importance information of the action corresponding to each edge based on the frequency-weighted edge betweenness of the action corresponding to each edge; Step S50: Sort the actions in the state transition graph according to the importance information of the action corresponding to each edge to obtain the action importance order; Step S60: Remove redundant edges from the state transition graph based on the order of action importance; Step S70: Iterate through steps S30 to S60 until the preset stopping condition is met, then determine the actions corresponding to the remaining edges in the state transition graph as key actions.
[0005] In conjunction with the first aspect, in some possible implementations, step S20 includes: Extract the state set and action set from the sequence sample set; Construct an initial state transition diagram based on the set of states and the set of actions; In the initial state transition graph, nodes with only in-degree 0 are marked as start nodes, nodes with only out-degree 0 are marked as end nodes, and weights are assigned according to the frequency of the action corresponding to each edge in the sequence sample set, resulting in a state transition graph containing start and end nodes.
[0006] In combination with the first aspect and the above implementation methods, in some possible implementation methods, step S30 includes: Traverse all edges in the state transition graph, and for the current edge encountered, perform the following steps: Obtain the target ratio of each node pair in the state transition graph for the current edge. The target ratio is used to represent the ratio of the product of the second quantity and the weight corresponding to the current edge to the first quantity. The first quantity represents the total number of shortest paths between node pairs. The second quantity represents the number of shortest paths in the node pair that pass through the current edge. The weight corresponding to the current edge is used to represent the frequency of the action corresponding to the current edge in the sequence sample set. The frequency-weighted edge betweenness of the current edge is obtained by summing the target ratios of each node in the state transition graph for the current edge.
[0007] Combining the first aspect and the above implementation methods, in some possible implementation methods, the target ratio of each node in the state transition graph relative to the current edge is calculated using the following formula: ; in, For the current edge, Weight the edge betweenness of the current edge according to its frequency. and For any pair of nodes in the state transition graph, As the first quantity, For the second quantity, This represents the weight corresponding to the current edge.
[0008] In combination with the first aspect and the above implementation methods, in some possible implementation methods, step S60 includes: The candidate edge set is determined based on at least one edge with the smallest frequency-weighted betweenness in the state transition graph. If the candidate edge set contains only one candidate edge, then the candidate edge is determined as the edge to be processed; If the candidate edge set contains multiple candidate edges, then the edge to be processed is determined from among the multiple candidate edges; The judgment result is obtained based on the edge to be processed and the state transition graph. The judgment result is used to indicate whether a new starting node or a new ending node is added to the state transition graph after the edge to be processed is removed. The new starting node is a node in the state transition graph whose out-degree becomes 0 due to the removal of the edge to be processed, excluding the already marked starting node. The new ending node is a node in the state transition graph whose in-degree becomes 0 due to the removal of the edge to be processed, excluding the already marked ending node. If the judgment result indicates that no new starting node or new ending node is added to the state transition graph after removing the edge to be processed, then the edge to be processed is identified as a redundant edge, and the redundant edge is removed from the state transition graph. If the judgment result indicates that after removing the edge to be processed, a new starting node or a new ending node is added to the state transition graph, then the edge to be processed is retained.
[0009] Combining the first aspect and the above implementation methods, in some possible implementation methods, the edge to be processed is determined from multiple candidate edges, including: Multiple candidate edges are compared using the weights corresponding to each candidate edge to obtain the comparison result. The weights corresponding to the candidate edges are used to characterize the frequency of the action corresponding to the candidate edge in the sequence sample set. If the comparison results indicate that the weights of the candidate edges are not the same, then the candidate edge with the smallest weight among the multiple candidate edges is determined as the edge to be processed. If the comparison results indicate that the weights of the candidate edges are the same, then the edge to be processed is randomly selected from the multiple candidate edges.
[0010] In combination with the first aspect and the above implementation methods, in some possible implementations, the method further includes: If the judgment result indicates that after removing the edge to be processed, a new starting node or a new ending node is added to the state transition graph, then the stopping condition is satisfied.
[0011] Combining the first aspect and the above implementation methods, in some possible implementation methods, if the judgment result indicates that no new starting node or new ending node is added to the state transition graph after removing the edge to be processed, then the edge to be processed is identified as a redundant edge, and after removing the redundant edge in the state transition graph, the following is also included: The weights of the remaining edges in the state transition graph are normalized based on the frequency of occurrence of the actions corresponding to the remaining edges in the sequence sample set.
[0012] Secondly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the key action extraction method based on the state transition diagram of the first aspect.
[0013] Thirdly, embodiments of this application provide a non-transitory computer-readable storage medium storing a computer program thereon, wherein when the computer program is executed by a processor, it implements the steps of the key action extraction method based on the state transition diagram of the first aspect.
[0014] The key action extraction method, device, and storage medium based on state transition graphs provided in this application first acquire a set of sequence samples generated by an intelligent game model during the adversarial process, and then construct a state transition graph based on this set, with states forming nodes and actions forming edges. Next, the frequency-weighted edge betweenness factor of each action corresponding to an edge in the state transition graph is determined, and the importance information of each action corresponding to an edge is determined based on this frequency-weighted edge betweenness factor. Subsequently, all actions are sorted according to the importance information to obtain an action importance order, and redundant edge removal is performed on the state transition graph based on this action importance order. Finally, the steps of determining importance information, sorting, and removing redundant edges are iteratively executed until a preset stopping condition is met, and the actions corresponding to the remaining edges in the final graph are identified as key actions. Because it performs offline analysis on the generated sequence sample set, it avoids reliance on the online operation of the intelligent game model, thereby improving the efficiency of key action extraction. Furthermore, by constructing a state transition graph and using frequency-weighted edge betweenness as a metric, the connectivity effect of actions in the global path can be evaluated, rather than assessing their value in isolation. This overcomes the limitation of related technologies that cannot measure the degree of influence of actions between upstream and downstream parts of the action sequence, resulting in more complete extracted key actions that can more accurately reflect the core decision-making logic of the intelligent game model. In summary, the proposed solution can effectively improve the accuracy of key action extraction. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating the key action extraction method based on a state transition diagram provided in this application embodiment; Figure 2 This is a schematic diagram of the process for constructing a state transition diagram provided in an embodiment of this application; Figure 3 This is an example schematic diagram of the state transition diagram provided in the embodiments of this application; Figure 4 This is an example schematic diagram of the frequency-weighted side betweenness provided in an embodiment of this application; Figure 5This is a schematic diagram of the redundant edge removal process provided in an embodiment of this application; Figure 6 This is a schematic diagram of the process for determining the edge to be processed from multiple candidate edges, provided in an embodiment of this application. Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0018] With the development of artificial intelligence technology, intelligent game theory models have been applied in scenarios such as game competition and decision-making. Specifically, an intelligent game theory model is an AI model that can interact with the external environment or other intelligent agents, and select and execute specific actions from a pre-set set of actions based on its current state information to achieve a predetermined game objective. During the competition, intelligent game theory models generate relevant data, which typically contains a large number of actions, but not all actions are equally important to the final game outcome. Therefore, there is a need to extract key actions from the relevant data generated by intelligent game theory models during competition.
[0019] For example, in some related technologies, the approach adopted is to perform online value assessment on each action in the relevant data generated by the intelligent game model. Specifically, one approach is to input the action sequence in the relevant data back into the intelligent game model, calculate the value assessment result of each action through the online operation of the intelligent game model, and select key actions based on the value assessment results. Another approach is to aggregate multiple states in the relevant data and perform aggregate judgment on the actions associated with the aggregated states to determine the importance of the actions.
[0020] It is evident that the aforementioned technologies have shortcomings: they rely heavily on the online operation of the intelligent game model, resulting in low efficiency in extracting key actions; at the same time, such solutions cannot adequately measure the impact of a specific action on the upstream and downstream states of the complete action sequence, making it difficult to accurately assess the importance of actions at the strategy level, ultimately leading to incomplete extracted action strategies that fail to accurately reflect the core decision-making logic of the intelligent game model.
[0021] Therefore, how to accurately extract key actions from the relevant data generated by intelligent game models has become an urgent technical problem to be solved.
[0022] To address the aforementioned issues, the solution provided in this application primarily includes: first, acquiring a sequence sample set generated by the intelligent game model during the adversarial process, and constructing a state transition graph based on this set, with states constituting nodes and actions constituting edges; next, determining the frequency-weighted edge betweenness factor of each action corresponding to an edge in the state transition graph, and then determining the importance information of each action corresponding to an edge based on this frequency-weighted edge betweenness factor; subsequently, sorting all actions according to the importance information to obtain an action importance order, and performing redundant edge removal operations on the state transition graph based on this action importance order; finally, iteratively executing the steps of determining importance information, sorting, and redundant edge removal until a preset stopping condition is met, and identifying the actions corresponding to the remaining edges in the final graph as key actions. Because it performs offline analysis on the generated sequence sample set, it avoids reliance on the online operation of the intelligent game model, thereby improving the efficiency of key action extraction. Furthermore, by constructing a state transition graph and using frequency-weighted edge betweenness as a metric, the connectivity effect of actions in the global path can be evaluated, rather than assessing their value in isolation. This overcomes the limitation of related technologies that cannot measure the degree of influence of actions between upstream and downstream parts of the action sequence, resulting in more complete extracted key actions that can more accurately reflect the core decision-making logic of the intelligent game model. In summary, the proposed solution can effectively improve the accuracy of key action extraction.
[0023] The following will provide a detailed description of the key action extraction method based on state transition diagrams provided in the embodiments of this application.
[0024] Please see Figure 1 , Figure 1 This is a flowchart illustrating a key action extraction method based on a state transition diagram, provided as an embodiment of this application. Figure 1 As shown, the method in this application embodiment may include the following steps S10-S70.
[0025] Step S10: Obtain the sequence sample set generated by the intelligent game model during the adversarial process.
[0026] Specifically, the first step is to obtain the sequence sample set generated by the intelligent game model during the adversarial process. Here, the intelligent game model refers to an artificial intelligence model that can interact with the environment or other intelligent agents and select and execute actions based on the current state to achieve a preset game objective; the sequence sample set refers to the information set containing continuous states and actions generated by the intelligent game model during the adversarial process.
[0027] For example, the architecture of an intelligent game model can be represented as follows: it includes a perception module, a decision-making module, and an execution module. The perception module receives environmental state information, the decision-making module outputs action instructions based on the state information, and the execution module applies the action instructions to the environment. The continuous state-action-state data generated by this cyclical interaction process constitutes the basis of the sequence sample set. For example, the training process of the intelligent game model can be represented as follows: first, an initial game model is constructed; then, through reinforcement learning algorithms, it continuously trials and errors in interaction with the environment, adjusting the parameters of the initial game model based on reward signals to maximize cumulative rewards; finally, the initial game model is trained into an intelligent game model capable of generating a sequence sample set. For example, the specific implementation of the adversarial process of the intelligent game model can be represented as follows: the intelligent game model interacts with one or more adversaries in a preset game environment, selecting the optimal action based on its own and the adversaries' states to achieve victory or a specific goal. The state transition data generated during this adversarial process is recorded to form a sequence sample set.
[0028] Regarding the steps for obtaining the sequence sample set generated by the intelligent game model during the adversarial process, some possible implementations include extracting the sequence sample set from the historical data of the intelligent game model through a pre-defined data interface. Other possible implementations involve real-time monitoring of the intelligent game model's adversarial process, capturing and recording the generated data stream, thereby constructing the sequence sample set.
[0029] Step S20: Construct a state transition graph based on the sequence sample set. The state transition graph includes nodes consisting of states and edges consisting of actions.
[0030] Specifically, in order to structure abstract sequence data into a graph structure for analyzing relationships, it is necessary to construct a state transition graph based on the sequence sample set. The state transition graph includes nodes composed of states and edges composed of actions. Here, the state transition graph can be a directed graph; a state refers to a complete description of the environment in which the intelligent game model is located at a certain moment; and an action refers to the decision or operation that the intelligent game model can perform in a specific state.
[0031] Regarding this step, in some possible implementations, the sequence sample set can be traversed, with the states used as nodes in the state transition graph and the actions used as edges connecting the corresponding nodes, thus constructing the state transition graph. In some possible implementations, based on the constructed state transition graph, a weight can be assigned to each edge in the graph; this weight is used to characterize the statistical frequency of the corresponding action in the sequence sample set.
[0032] For example, the sequence sample set and state transition diagram in steps S10 and S20 above can be represented in the following form: Each sample in the sequence sample set can be represented as: ; in, and These represent the states at time t and time t+1, respectively. Represents the action at time t; the set of all states is The set of all actions is And there are , Based on this, the state transition diagram It can be represented as: ; in, As a state transition diagram The set of points, As a state transition diagram The set of edges. In the state transition graph In this process, nodes with only in-degree of 0 are marked as starting nodes, nodes with only out-degree of 0 are marked as ending nodes, and weights are assigned to each edge based on the frequency of the action corresponding to each edge in the sequence sample set.
[0033] Step S30: Determine the frequency-weighted edge betweenness of the action corresponding to each edge in the state transition graph.
[0034] Specifically, to quantitatively assess the criticality of the action corresponding to each edge in global path connectivity, it is necessary to determine the frequency-weighted edge betweenness factor of the action corresponding to each edge in the state transition graph. The frequency-weighted edge betweenness factor of the action corresponding to an edge is a comprehensive index that combines the action's pass rate in the global shortest path with its frequency of occurrence in the samples, used to evaluate the criticality of the action.
[0035] Regarding this step, in some possible implementations, for each edge in the state transition graph, a preset frequency-weighted edge betweenness formula can be applied to perform mathematical operations on the state transition graph to determine the frequency-weighted edge betweenness of the action corresponding to each edge. In some possible implementations, the standard edge betweenness of each edge in the unweighted case can be calculated first; then, the standard edge betweenness of each edge can be multiplied by the frequency of occurrence of the action corresponding to that edge in the sequence sample set to obtain the frequency-weighted edge betweenness of the action corresponding to each edge.
[0036] Step S40: Determine the importance information of the action corresponding to each edge based on the frequency-weighted edge betweenness of the action corresponding to each edge.
[0037] Specifically, in order to transform the frequency-weighted edge betweenness number (FIN) into ranking importance information, it is necessary to determine the importance information of the action corresponding to each edge based on the FIN of the action corresponding to each edge. Here, the importance information of the action corresponding to the edge refers to the numerical value quantified by the FIN, reflecting the criticality of the action in the state transition network. Regarding this step, in some possible implementations, the frequency-weighted edge betweenness numbers of the actions corresponding to each edge can be directly used as the importance information of the actions corresponding to that edge. In other possible implementations, all frequency-weighted edge betweenness numbers can be normalized, and the processed result can be used as the importance information of the actions corresponding to each edge.
[0038] Step S50: Sort the actions in the state transition graph according to the importance information of the action corresponding to each edge to obtain the action importance order.
[0039] Specifically, to determine the priority of removing redundant actions, the actions in the state transition graph need to be sorted according to the importance information of the actions corresponding to each edge, resulting in an action importance order. Sorting refers to arranging all actions in the state transition graph according to their importance information, for example, in descending order, where actions with higher importance information values are ranked higher. The action importance order refers to the sequence generated after sorting, reflecting the relative importance of each action.
[0040] Regarding this step, some possible implementations include using a sorting algorithm to rank all actions in the state transition diagram based on importance information, thus obtaining the action importance order. Other possible implementations involve dividing all actions into high-importance and low-importance categories based on a preset importance threshold, then ranking the actions within each category, and finally integrating the results to obtain the action importance order.
[0041] Step S60: Remove redundant edges from the state transition graph based on the order of action importance.
[0042] Specifically, to simplify the state transition graph and remove non-critical actions, redundant edge removal is required based on the order of action importance. Redundant edge removal refers to the process of deleting edges of lower importance from the state transition graph according to certain rules, while maintaining the connectivity of the graph structure.
[0043] Regarding this step, some possible implementations include: starting with the least important action and attempting to remove corresponding edges from the state transition graph sequentially based on action importance, and performing connectivity checks. If removal does not disrupt the graph's pre-defined structure, the removal is confirmed, completing one round of redundant edge removal. Another possible implementation is to set a removal ratio threshold, removing a fixed proportion of the least important edges each time based on action importance, then evaluating the graph as a whole. If the evaluation results meet the requirements, the current round of redundant edge removal is complete.
[0044] Step S70: Iterate through steps S30 to S60 until the preset stopping condition is met, then determine the actions corresponding to the remaining edges in the state transition graph as key actions.
[0045] Specifically, in this embodiment, the stopping condition refers to the condition triggered when any removal attempt in the redundant edge removal operation would disrupt the structural integrity of the state transition graph. For example, the stopping condition could be manifested as the generation of a new starting node or a new ending node in the graph after the removal of a candidate edge; or as the breaking of at least one path from the marked starting node to the marked ending node in the state transition graph after the removal of a candidate edge; or as the finding that the minimum value of the frequency-weighted edge betweenness factor corresponding to all remaining edges is still higher than a preset threshold after importance evaluation of all remaining edges. There are many other possible implementations of the stopping condition, which will not be listed here. The key action refers to the action corresponding to the edge that is ultimately retained in the state transition graph after the stopping condition is met and iteration stops.
[0046] In order to find the core action set that maintains the integrity of the strategy through cyclic reduction, it is necessary to iteratively execute steps S30 to S60 until the preset stopping condition is met, and then determine the actions corresponding to the remaining edges in the state transition graph as key actions.
[0047] Regarding this step, in some possible implementations, after each successful redundant edge removal operation, the updated state transition graph can be used as the input for the next iteration. Based on this updated state transition graph, the following steps can be repeated: determining the frequency-weighted edge betweenness constant of each action corresponding to each edge, determining the importance information of each action corresponding to each edge, sorting the actions in the state transition graph to obtain the action importance order, and performing the next redundant edge removal operation, until the stopping condition is met. In some possible implementations, after each successful redundant edge removal operation, the weights of the edges in the state transition graph can be updated only based on the statistical frequency of the remaining edges in the sequence sample set. Then, based on the updated weight information, the following steps can be repeated: determining the frequency-weighted edge betweenness constant of each action corresponding to each edge, determining the importance information of each action corresponding to each edge, sorting the actions in the state transition graph to obtain the action importance order, and determining the next redundant edge removal operation, until the stopping condition is met. In some possible implementations, after the stopping condition is met and the iteration stops, the state transition graph before the last successful execution of the redundant edge removal operation can be used as the final graph structure, and the set of actions corresponding to all edges in the final graph structure can be determined as the final key action and output.
[0048] In this embodiment, firstly, a sequence sample set generated by the intelligent game model during the adversarial process is obtained, and based on this, a state transition graph is constructed, with states forming nodes and actions forming edges. Next, the frequency-weighted edge betweenness factor of each action corresponding to an edge in the state transition graph is determined, and then the importance information of each action corresponding to an edge is determined based on this frequency-weighted edge betweenness factor. Subsequently, all actions are sorted according to the importance information to obtain an action importance order, and redundant edge removal is performed on the state transition graph based on this action importance order. Finally, the steps of determining importance information, sorting, and removing redundant edges are iteratively executed until a preset stopping condition is met, and the actions corresponding to the remaining edges in the final graph are identified as key actions. Because it performs offline analysis on the generated sequence sample set, it avoids relying on the online operation of the intelligent game model, thereby improving the efficiency of key action extraction. Furthermore, by constructing a state transition graph and using frequency-weighted edge betweenness as a metric, the connectivity effect of actions in the global path can be evaluated, rather than assessing their value in isolation. This overcomes the limitation of related technologies that cannot measure the degree of influence of actions between upstream and downstream parts of the action sequence, resulting in more complete extracted key actions that can more accurately reflect the core decision-making logic of the intelligent game model. In summary, the proposed solution can effectively improve the accuracy of key action extraction.
[0049] Please see Figure 2 This application provides a schematic flowchart for constructing a state transition diagram, as shown in the embodiments below. Figure 2As shown, the method in this application embodiment may include the following steps S201-S203, which can be used as a further refinement of "step S20: constructing a state transition diagram based on the sequence sample set".
[0050] S201, extract the state set and action set from the sequence sample set; S202, Construct the initial state transition diagram based on the state set and action set; S203, in the initial state transition graph, the node with only in-degree 0 is marked as the starting node, the node with only out-degree 0 is marked as the ending node, and the corresponding weight is assigned according to the frequency of the action corresponding to each edge in the sequence sample set, so as to obtain the state transition graph containing the starting node and the ending node.
[0051] Specifically, considering that the key action extraction scheme relies on the online operation of the intelligent game model, which has problems such as low generation efficiency and insufficient consideration of the importance of actions at the strategy level, this embodiment proposes a key action extraction scheme based on state transition diagrams.
[0052] First, it is necessary to extract the state set and action set from the sequence sample set. The state set refers to the set of all states in the action sequence sample set; the action set refers to the set of all actions in the action sequence sample set.
[0053] Regarding this step, in some possible implementations, each sample in the sequence sample set can be traversed, the state information contained in the sample can be summarized to form a state set, and the action information contained in the sample can be summarized to form an action set.
[0054] After completing the above extraction, an initial state transition diagram needs to be constructed based on the state set and action set.
[0055] Regarding this step, in some possible implementations, each state in the state set can be used as a node, and each action in the action set can be used as a directed edge connecting the corresponding node, thereby constructing the initial state transition graph.
[0056] Furthermore, to evaluate the importance of the state transition graph and perform subsequent redundant edge removal, nodes with only an in-degree of 0 are marked as start nodes, and nodes with only an out-degree of 0 are marked as end nodes in the initial state transition graph. Weights are then assigned based on the frequency of the action corresponding to each edge in the sequence sample set, resulting in a state transition graph containing start and end nodes. More specifically, in-degree refers to the number of edges pointing to a node in the graph; an in-degree of 0 means there are no other states to transition to that node. The start node is the node in the state transition graph that represents the starting state of the game. The out-degree refers to the number of edges originating from a node in the graph; an out-degree of 0 means that the node cannot transition to any other state. The end node is the node in the state transition graph that represents the ending state of the game.
[0057] Regarding this step, in some possible implementations, all nodes in the initial state transition graph can be traversed, the in-degree and out-degree of each node can be calculated, and the start node and end node can be marked according to the calculation results. At the same time, the number of times each action appears in the sequence sample set is counted, the frequency is calculated according to the statistical results, and the calculated frequency is used as the weight of the corresponding edge, thereby generating a state transition graph containing the start node, end node, and weights.
[0058] For a better understanding of the state transition diagram in this embodiment, please refer to [link to relevant documentation]. Figure 3 , Figure 3 This is an example schematic diagram of a state transition diagram provided in an embodiment of this application. Specifically, Figure 3 The left side shows a sequence sample set, which is represented by a multi-layered structure, with each layer representing a sample sequence. Figure 3 The middle section illustrates the correspondence between samples and the frequencies of their corresponding actions in the sequence sample set, where the sample column contains... , , The frequency columns of the corresponding actions in the sequence sample set are 0.3, 0.4, and 0.2, respectively. Figure 3 The right side shows the state transition diagram, which contains... , , , , , Nodes, among which As the starting node, For the termination node, , , , As an intermediate node, nodes are connected by directed edges. The values on these edges (0.3, 0.4, 0.2, 0.25, 0.15, 0.2, 0.3, 0.2) represent the frequencies of the corresponding actions in the sequence sample set, and these frequencies serve as the weights of the corresponding edges in the state transition graph. In the state transition graph, Node to The edge weight of the node is 0.3. Node to The edge weight of the node is 0.4. Node to The edge weight of the node is 0.2. Node to The edge weight of the node is 0.25. Node to The edge weight of the node is 0.15. Node to The edge weight of the node is 0.2. Node to The edge weight of the node is 0.3. Node to The edge weight of the node is 0.2. In this state transition graph, The only node has an out-degree of 0. The node has only one in-degree of 0. , , , The in-degree and out-degree of a node are both non-zero.
[0059] In this embodiment, the state set and action set are first extracted from the sequence sample set, and then an initial state transition graph is constructed. Finally, a state transition graph containing the start node, end node, and weights is determined. This process transforms the continuous game sequence samples in the time dimension into a graph structure composed of state nodes and action edges. This transformation provides the data structure foundation for subsequent applications of graph theory-based algorithms, such as frequency-weighted edge betweenness calculation and redundant edge removal operations, to analyze the state transition graph. Simultaneously, by structurally representing the decision-making process of the intelligent game model in the form of a state transition graph, support is provided for tracing the internal state transitions and action selection logic of the model, thereby helping to improve the interpretability of the intelligent game model.
[0060] In one embodiment, the above "step S30: determining the frequency-weighted edge betweenness number of the action corresponding to each edge in the state transition graph" can be further refined and may include the following steps: Traverse all edges in the state transition graph, and for the current edge encountered, perform the following steps: Obtain the target ratio of each node pair in the state transition graph for the current edge. The target ratio is used to represent the ratio of the product of the second quantity and the weight corresponding to the current edge to the first quantity. The first quantity represents the total number of shortest paths between node pairs. The second quantity represents the number of shortest paths in the node pair that pass through the current edge. The weight corresponding to the current edge is used to represent the frequency of the action corresponding to the current edge in the sequence sample set. The frequency-weighted edge betweenness of the current edge is obtained by summing the target ratios of each node in the state transition graph for the current edge.
[0061] Specifically, considering that the value of a single action cannot be fully reflected by simply evaluating its statistical frequency or in isolation in related technologies, this embodiment proposes a frequency-weighted edge betweenness calculation scheme that combines global path connectivity with local statistical characteristics.
[0062] Traversal refers to the process of checking and processing all edges in the state transition graph one by one, ensuring that each edge is processed once as the current edge.
[0063] To accurately quantify the criticality of the action corresponding to each edge in the state transition graph within the global path, it is necessary to traverse all edges in the state transition graph. For the current edge encountered, the following steps are performed: Obtain the target ratio of each node pair in the state transition graph for the current edge. The target ratio is used to represent the ratio of the product of the second quantity and the weight corresponding to the current edge to the first quantity. The first quantity represents the total number of shortest paths between node pairs. The second quantity represents the number of shortest paths in the node pair that pass through the current edge. The weight corresponding to the current edge is used to represent the frequency of the action corresponding to the current edge in the sequence sample set.
[0064] More specifically, the total number of shortest paths between node pairs refers to the total number of all non-repeating paths in the state transition graph that start from one node in a node pair and reach another node with the shortest path length; the number of shortest paths in a node pair that pass through the current edge refers to the number of paths that include the current edge among all the shortest paths between node pairs.
[0065] Regarding this step, some possible implementations can employ shortest path algorithms from graph theory, such as breadth-first search or Dijkstra's algorithm, to calculate the shortest path between all pairs of nodes in the state transition graph and to count the first quantity and the second quantity for each pair of nodes relative to the current edge. Simultaneously, the weight corresponding to the current edge is read from the state transition graph; this weight represents the frequency of the action corresponding to the current edge in the sequence sample set. Finally, the second quantity is multiplied by the weight corresponding to the current edge, and then divided by the first quantity to obtain the target ratio of the pair of nodes relative to the current edge.
[0066] Furthermore, the target ratios of each node in the state transition graph for the current edge are summed to obtain the frequency-weighted edge betweenness of the current edge. Here, summation refers to the mathematical summation of the target ratios calculated for the current edge by each node in the state transition graph.
[0067] Regarding this step, in some possible implementations, an accumulator can be initialized with an initial value of zero; then, each node pair in the state transition graph is traversed, and the target ratio calculated for each node pair for the current edge is added to the accumulator; when the target ratios of all node pairs have been accumulated, the final value in the accumulator is the frequency-weighted edge betweenness of the current edge.
[0068] It should be noted that the above traversal process applies to all edges in the state transition graph. That is, for each edge in the edge set of the state transition graph, the process of obtaining the target ratio of each corresponding node pair and accumulating it is performed once to obtain the frequency-weighted edge betweenness of the edge itself.
[0069] For a better understanding of the state transition diagram in this embodiment, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is an example diagram illustrating the frequency-weighted side betweenness numbers provided in an embodiment of this application. Specifically, Figure 4 The table on the left lists the actions and their corresponding frequency-weighted betweenness numbers, where the actions... The frequency-weighted edge betweenness factor is 0.05, and the action... The frequency-weighted betweenness number is 0.067, and the action... The frequency-weighted edge betweenness number is 0.027, and the action... The frequency-weighted edge betweenness number is 0.033, and the action... The frequency-weighted edge betweenness factor is 0.02, and the action... The frequency-weighted edge betweenness number is 0.027, and the action... The frequency-weighted edge betweenness factor is 0.04, and the action... The frequency-weighted edge betweenness is 0.027. In the state transition diagram on the right, the node... and The edges between them correspond to actions The weight of this edge is 0.05, and the calculated frequency-weighted edge betweenness is 0.05; node and The edges between them correspond to actions The weight of this edge is 0.067, and the calculated frequency-weighted edge betweenness is 0.067; node and The edges between them correspond to actions The weight of this edge is 0.027, and the calculated frequency-weighted edge betweenness is 0.027; node and The edges between them correspond to actions The weight of this edge is 0.033, and the calculated frequency-weighted edge betweenness is 0.033; node and The edges between them correspond to actions The weight of this edge is 0.02, and the calculated frequency-weighted edge betweenness is 0.02; node and The edges between them correspond to actions The weight of this edge is 0.027, and the calculated frequency-weighted edge betweenness is 0.027; node and The edges between them correspond to actions The weight of this edge is 0.04, and the calculated frequency-weighted edge betweenness is 0.04; node and The edges between them correspond to actions The weight of this edge is 0.027, and the calculated frequency-weighted edge betweenness is 0.027.
[0070] In this embodiment, the role of actions in the global state transition network is quantified based on the number of shortest paths between node pairs. A weighted factor representing the frequency of an action in the sequence sample set is introduced to combine the shortest path metric with the statistical frequency of the action. Finally, a frequency-weighted edge betweenness factor is generated for each edge in the state transition graph. This value serves as the quantification basis for subsequent redundant action screening steps, ensuring that the judgment of action importance considers both the action's passage through multiple state transition paths and its frequency of occurrence in the sequence sample set.
[0071] In one embodiment, the above step "obtaining the target ratio of each node in the state transition graph for the current edge" is calculated using the following formula: ; in, For the current edge, Weight the edge betweenness of the current edge according to its frequency. and For any pair of nodes in the state transition graph, As the first quantity, For the second quantity, This represents the weight corresponding to the current edge.
[0072] Specifically, in order to calculate the current edge Frequency-weighted border betweenness This requires traversing and processing all node pairs in the state transition graph. First, for the current edge, an accumulator is initialized with an initial value of zero to store the final frequency-weighted edge betweenness. Next, all node pairs in the state transition graph are traversed, where... and For any two distinct pairs of nodes.
[0073] For example, for each pair of nodes, the following calculation is performed: First, calculate the first quantity, which is the total number of shortest paths between node pairs. For example, the shortest path algorithm in graph theory, such as breadth-first search, can be applied to count the number of nodes. To the node The number of all shortest paths.
[0074] Second, calculate the second quantity, which is the number of shortest paths passing through the current edge in each node pair. In the process of calculating the shortest path, the number of shortest paths that include the current edge is further counted.
[0075] Third, obtain the weight corresponding to the current edge. This weight is determined when constructing the state transition graph, and its value represents the frequency of the action corresponding to the current edge in the sequence sample set. .
[0076] Fourth, calculate the target ratio of the node relative to the current edge. According to the formula in this embodiment, the second quantity... Weight corresponding to the current edge Multiply, then divide by the first quantity This yields the target ratio of the node's contribution to the current edge.
[0077] Fifth, accumulate the target ratio. Add the target ratio calculated in step four to the initialized accumulator.
[0078] After all node pairs in the state transition graph have undergone the above steps, the final value in the accumulator is the value of the current edge. Frequency-weighted border betweenness .
[0079] In this embodiment, the abstract frequency-weighted edge betweenness formula is transformed into a specific calculation process through the above steps. This method not only considers the statistical frequency of actions in the sequence sample set, i.e., the weight corresponding to the current edge, but also... It also fully considers the pivotal role of this action in the global state transition network, namely its passage through all shortest paths, determined by the first quantity. Second quantity Common characterization. Therefore, the calculated frequency-weighted side betweenness numbers are... It can more accurately quantify the importance of each action corresponding to each edge in the state transition graph, providing a reliable data foundation for subsequent action sorting and redundant edge removal operations.
[0080] Please see Figure 5 This document provides a flowchart illustrating the process of redundant edge removal in an embodiment of this application. Figure 5 As shown, the method in this application embodiment may include the following steps S601-S606. Steps S601-S606 can be used as a further refinement of "Step S60: Perform redundant edge removal operation on the state transition graph based on the order of action importance".
[0081] S601, determine the candidate edge set based on at least one edge with the smallest frequency-weighted edge betweenness in the state transition diagram; S602, If the candidate edge set contains only one candidate edge, then the candidate edge is determined as the edge to be processed; S603, If the candidate edge set contains multiple candidate edges, then determine the edge to be processed from among the multiple candidate edges; S604. Obtain the judgment result based on the edge to be processed and the state transition graph. The judgment result is used to indicate whether a new starting node or a new ending node is added to the state transition graph after removing the edge to be processed. The new starting node refers to the node in the state transition graph, other than the already marked starting node, whose out-degree becomes 0 due to the removal of the edge to be processed. The new ending node refers to the node in the state transition graph, other than the already marked ending node, whose in-degree becomes 0 due to the removal of the edge to be processed. S605 If the judgment result indicates that no new starting node or new ending node is added to the state transition graph after removing the edge to be processed, then the edge to be processed is identified as a redundant edge, and the redundant edge is removed from the state transition graph. S606 If the judgment result indicates that after removing the edge to be processed, a new starting node or a new ending node is added to the state transition graph, then the edge to be processed is retained.
[0082] Specifically, the first step is to determine the candidate edge set based on at least one edge with the smallest frequency-weighted betweenness number in the state transition graph. It is understandable that, since there may be multiple edges with the same frequency-weighted betweenness number, the candidate edge set may contain one or more candidate edges.
[0083] Regarding this step, in some possible implementations, all edges in the state transition graph can be traversed, the frequency-weighted edge betweenness number corresponding to each edge can be calculated, and the edge with the smallest frequency-weighted edge betweenness number can be added to the candidate edge set.
[0084] In one scenario, the candidate edge set contains only one candidate edge, meaning that the edge with the smallest frequency-weighted betweenness factor is unique. In this case, the candidate edge can be identified as the edge to be processed.
[0085] In some cases, the candidate edge set contains multiple candidate edges, meaning that the edge with the smallest frequency-weighted edge betweenness is not unique. In this situation, the edge to be processed can be determined from among the multiple candidate edges. Regarding the process of determining the edge to be processed from multiple candidate edges, some possible implementations can compare the weights corresponding to the multiple candidate edges and determine the candidate edge with the smallest weight as the edge to be processed; in some possible implementations, a candidate edge can be randomly selected as the edge to be processed. There are also many other ways to determine the edge to be processed from multiple candidate edges, which will not be listed here.
[0086] After identifying the edge to be processed, to ensure the integrity of the strategy in the state transition graph, a judgment result needs to be obtained based on the edge to be processed and the state transition graph. The judgment result is used to indicate whether a new starting node or a new ending node is added to the state transition graph after removing the edge to be processed. Here, a new starting node refers to a node in the state transition graph, other than the already marked starting node, whose out-degree becomes 0 due to the removal of the edge to be processed; a new ending node refers to a node in the state transition graph, other than the already marked ending node, whose in-degree becomes 0 due to the removal of the edge to be processed.
[0087] Regarding the process of obtaining the judgment result based on the edge to be processed and the state transition graph, some possible implementations can simulate removing the edge to be processed, and then checking the in-degree and out-degree of other nodes in the state transition graph besides the marked start and end nodes to determine whether to add a new start node or a new end node. It should be noted that the process of obtaining the judgment result does not actually remove the edges in the state transition graph.
[0088] Depending on the different outcomes of the judgment, at least two of the following situations and corresponding steps exist: In one scenario, the judgment result indicates that removing the edge to be processed does not add a new starting node or a new ending node to the state transition graph. This means that removing the edge to be processed will not disrupt the connectivity of the state transition graph. In this case, the edge to be processed can be identified as a redundant edge, and the redundant edge can be removed from the state transition graph. Regarding this step, in some possible implementations, the structure of the state transition graph can be updated, and the weights of the remaining edges can be renormalized based on the sequence sample set.
[0089] In one scenario, the decision indicates that removing the edge to be processed would add a new starting node or a new ending node to the state transition graph. This means that removing the edge would disrupt the policy integrity of the state transition graph. In this case, the edge to be processed can be retained. Regarding this step, in some possible implementations, the edge to be processed can be recorded as a retained edge, and subsequent removal operations can be stopped.
[0090] In addition, under certain possible circumstances, the "judgment result indicates that after removing the edge to be processed, a new starting node or a new ending node is added to the state transition graph" will also trigger the stopping condition. Specifically, when it is detected that removing any edge to be processed will result in the appearance of a new starting node or a new ending node, the stopping condition is determined to be met.
[0091] In this embodiment, redundant edges are removed based on frequency-weighted edge betweenness and node degree to reduce the state transition graph. The scheme determines the candidate edge set based on frequency-weighted edge betweenness, ensuring that the removal operation targets actions with low importance in terms of global path connectivity and sample occurrence frequency. When multiple edges of equal importance exist, the edge to be processed is determined by comparing weights or random selection, providing a basis for the removal operation. Before each removal, the scheme simulates whether removing the edge will generate a new starting node or a new ending node. This criterion is related to the integrity of the policy path represented by the state transition graph, ensuring that the removal operation does not break the effective policy chain from the starting state to the ending state. By iteratively executing evaluation, judgment, and removal operations, redundant edges can be gradually eliminated while maintaining graph connectivity and policy integrity until the stopping condition is met. The edges ultimately retained in the state transition graph are the key actions that maintain the decision-making logic of the intelligent game model.
[0092] Please see Figure 6 This application provides a flowchart illustrating the process of determining the edge to be processed from multiple candidate edges, as shown in the embodiments of this application. Figure 6 As shown, the method of this application embodiment may include the following steps S6031-S6033, which can be used as a further refinement of the step "determining the edge to be processed among multiple candidate edges".
[0093] S6031, compare multiple candidate edges using the weights corresponding to each candidate edge to obtain the comparison result, wherein the weights corresponding to the candidate edges are used to characterize the frequency of the action corresponding to the candidate edge in the sequence sample set; S6032, if the comparison result indicates that the weights of each candidate edge are different, then the candidate edge with the smallest weight among the multiple candidate edges is determined as the edge to be processed. S6033, if the comparison result indicates that the weights of each candidate edge are the same, then the edge to be processed is randomly determined from multiple candidate edges.
[0094] Specifically, the first step is to compare the weights of each candidate edge among multiple candidate edges to obtain the comparison results. The weights of the candidate edges are used to characterize the frequency of the actions corresponding to the candidate edges in the sequence sample set. The comparison refers to the operation of comparing the numerical values of the weights of each candidate edge among multiple candidate edges to determine whether the weights of each candidate edge are the same.
[0095] Based on the different outcomes of the comparison, at least two scenarios and corresponding steps exist: In one scenario, the comparison result indicates that the weights corresponding to the candidate edges are different. This means that among multiple candidate edges, at least two candidate edges have different weight values, allowing for priority establishment based on weight differences. In this case, the candidate edge with the smallest weight among the multiple candidate edges can be identified as the edge to be processed. Regarding this step, in some possible implementations, multiple candidate edges can be traversed, the weight corresponding to each candidate edge can be obtained, and the candidate edge with the smallest weight value can be identified as the edge to be processed. Here, the weight is used to characterize the frequency of the action corresponding to the candidate edge in the sequence sample set.
[0096] In one scenario, the comparison result indicates that all candidate edges have the same weight. This means that among multiple candidate edges, all corresponding to the same weight value, making it impossible to establish priority based on weight differences. In this case, the edge to be processed can be randomly selected from the multiple candidate edges. Regarding this step, in some possible implementations, one candidate edge can be randomly selected as the edge to be processed.
[0097] In this embodiment, for the case where multiple candidate edges correspond to the same frequency-weighted edge betweenness, a comparison mechanism based on the weights of each candidate edge is introduced as the first priority criterion. This ensures that when weights differ, a unique edge to be processed can be determined based on the magnitude of the weight values. Simultaneously, when weights are also the same, random selection is used as the second priority criterion, providing an unbiased method for determining the edge to be processed for multiple candidate edges with completely identical weights. This embodiment, through hierarchical judgment criteria, transforms the selection uncertainty caused by the same frequency-weighted edge betweenness into a deterministic or random but structurally complete decision-making process. This ensures that the edge to be processed can be uniquely determined among multiple candidate edges, providing a clear execution target for subsequent redundant edge removal operations.
[0098] In one embodiment, the key action extraction method based on state transition diagram of this application further includes the following steps: If the judgment result indicates that after removing the edge to be processed, a new starting node or a new ending node is added to the state transition graph, then the stopping condition is satisfied.
[0099] Specifically, considering the need to ensure the structural integrity and policy continuity of the state transition graph during redundant edge removal, this embodiment proposes a scheme to set stopping conditions by detecting changes in node degree.
[0100] If the judgment result indicates that removing the edge to be processed adds a new starting node or a new ending node to the state transition graph, it means that removing the edge to be processed will destroy the integrity of the policy chain represented by the state transition graph. At this time, it can be determined that the stopping condition is met. Here, a new starting node refers to a node in the state transition graph, other than the already marked starting node, whose out-degree becomes 0 due to the removal of the edge to be processed; a new ending node refers to a node in the state transition graph, other than the already marked ending node, whose in-degree becomes 0 due to the removal of the edge to be processed.
[0101] Regarding this step, in some possible implementations, after determining that the stopping condition is met, the structural information of the current state transition graph can be recorded, and the actions corresponding to the remaining edges in the current state transition graph can be determined as the final set of critical actions. In some possible implementations, after determining that the stopping condition is met, a notification message to stop iteration can be output, and the state transition graph before the last successful execution of the redundant edge removal operation can be used as the final state transition graph, where the actions corresponding to the remaining edges in the final state transition graph are the critical actions.
[0102] In this embodiment, the criterion for satisfying the stopping condition is whether a new starting node or a new ending node is added after removing the edge to be processed. This ensures that during the redundant edge removal operation, all nodes in the state transition graph, except for the marked starting and ending nodes, have non-zero in-degree and out-degree. This ensures that at least one path from the marked starting node to the marked ending node remains connected after the removal operation, thus maintaining the integrity of the policy chain represented by the state transition graph. Therefore, the actions corresponding to the remaining edges in the state transition graph are components of this policy chain, reflecting the core decision-making logic of the intelligent game model.
[0103] In one embodiment, after the above step "if the judgment result indicates that no new starting node or new ending node is added to the state transition graph after removing the edge to be processed, then the edge to be processed is identified as a redundant edge, and the redundant edge is removed from the state transition graph", the following steps may also be included: The weights of the remaining edges in the state transition graph are normalized based on the frequency of occurrence of the actions corresponding to the remaining edges in the sequence sample set.
[0104] Specifically, to ensure that the weights corresponding to the remaining edges in the state transition graph can reconstruct an effective probability distribution after removing redundant edges, and to adapt to the frequency-weighted edge betweenness calculation in the next iteration, the weights of the remaining edges in the state transition graph need to be normalized based on the frequency of the actions corresponding to the remaining edges in the sequence sample set. Normalization refers to a data processing procedure that scales a set of values to a specific interval, such as [0, 1], to make the processed values comparable. For example, in the embodiments of this application, normalization can refer to summing and normalizing the weights of all remaining edges in the state transition graph; that is, dividing the weight of each remaining edge by the sum of the weights of all remaining edges, so that the sum of the normalized weights of all remaining edges is 1.
[0105] Regarding this step, in some possible implementations, one approach is to first traverse the remaining edges in the state transition graph. For each remaining edge, based on its corresponding action, query the frequency of that action in the sequence sample set and determine that frequency as the initial weight of the remaining edge. Then, calculate the sum of the initial weights of all remaining edges. Finally, for each remaining edge, divide its initial weight by the sum of the initial weights to obtain the normalized weight of the remaining edge, and update the normalized weight as the weight of the remaining edges in the state transition graph.
[0106] In this embodiment, the weights of the remaining edges in the state transition graph are normalized to reflect the state transition graph structure after removing redundant edges. Therefore, in subsequent iterations, the calculation of the frequency-weighted edge betweenness is based on the updated relative weights, avoiding weight distortion caused by changes in the graph structure.
[0107] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 1301, a communication interface 1302, a memory 1303, and a communication bus 1304, wherein the processor 1301, the communication interface 1302, and the memory 1303 communicate with each other via the communication bus 1304. The processor 1301 can call a computer program in the memory 1303 to execute steps of a key action extraction method based on a state transition diagram, such as including: Step S10: Obtain the sequence sample set generated by the intelligent game model during the adversarial process; Step S20: Construct a state transition graph based on the sequence sample set. The state transition graph includes nodes composed of states and edges composed of actions. Step S30: Determine the frequency-weighted edge betweenness of the action corresponding to each edge in the state transition graph; Step S40: Determine the importance information of the action corresponding to each edge based on the frequency-weighted edge betweenness of the action corresponding to each edge; Step S50: Sort the actions in the state transition graph according to the importance information of the action corresponding to each edge to obtain the action importance order; Step S60: Remove redundant edges from the state transition graph based on the order of action importance; Step S70: Iterate through steps S30 to S60 until the preset stopping condition is met, then determine the actions corresponding to the remaining edges in the state transition graph as key actions.
[0108] Furthermore, when the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0109] On the other hand, embodiments of this application also provide a non-transitory computer-readable storage medium storing a computer program. The computer program is used to cause a processor to execute the steps of the methods provided in the above embodiments, including, for example: Step S10: Obtain the sequence sample set generated by the intelligent game model during the adversarial process; Step S20: Construct a state transition graph based on the sequence sample set. The state transition graph includes nodes composed of states and edges composed of actions. Step S30: Determine the frequency-weighted edge betweenness of the action corresponding to each edge in the state transition graph; Step S40: Determine the importance information of the action corresponding to each edge based on the frequency-weighted edge betweenness of the action corresponding to each edge; Step S50: Sort the actions in the state transition graph according to the importance information of the action corresponding to each edge to obtain the action importance order; Step S60: Remove redundant edges from the state transition graph based on the order of action importance; Step S70: Iterate through steps S30 to S60 until the preset stopping condition is met, then determine the actions corresponding to the remaining edges in the state transition graph as key actions.
[0110] Non-transitory computer-readable storage media can be any available medium or data storage device that can be accessed by a processor, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)).
[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for extracting key actions based on state transition diagrams, characterized in that, include: Step S10: Obtain the sequence sample set generated by the intelligent game model during the adversarial process; Step S20: Construct a state transition graph based on the sequence sample set, the state transition graph including nodes composed of states and edges composed of actions; Step S30: Determine the frequency-weighted edge betweenness number of the action corresponding to each edge in the state transition graph; Step S40: Determine the importance information of the action corresponding to each edge based on the frequency-weighted edge betweenness of the action corresponding to each edge; Step S50: Sort the actions in the state transition graph according to the importance information of the action corresponding to each edge to obtain the action importance order; Step S60: Perform redundant edge removal operation on the state transition graph based on the order of action importance; Step S70: Iteratively execute steps S30 to S60 until a preset stopping condition is met, and then determine the actions corresponding to the remaining edges in the state transition graph as key actions.
2. The method according to claim 1, characterized in that, Step S20 includes: Extract the state set and action set from the sequence sample set; Construct an initial state transition diagram based on the state set and the action set; In the initial state transition graph, nodes with only in-degree of 0 are marked as starting nodes, nodes with only out-degree of 0 are marked as ending nodes, and weights are assigned according to the frequency of the action corresponding to each edge in the sequence sample set, resulting in a state transition graph containing the starting node and the ending node.
3. The method according to claim 1, characterized in that, Step S30 includes: Traverse all edges in the state transition graph, and for the current edge encountered, perform the following steps: Obtain the target ratio of each node pair in the state transition graph for the current edge, wherein the target ratio is used to characterize the ratio of the product of the second quantity and the weight corresponding to the current edge to the first quantity, the first quantity characterizes the total number of shortest paths between the node pairs, the second quantity characterizes the number of shortest paths in the node pairs that pass through the current edge, and the weight corresponding to the current edge is used to characterize the frequency of the action corresponding to the current edge in the sequence sample set. The frequency-weighted edge betweenness of the current edge is obtained by summing the target ratios of each node in the state transition graph for the current edge.
4. The method according to claim 3, characterized in that, The target ratio of each node in the state transition graph for the current edge is calculated using the following formula: ; in, For the current edge, The frequency-weighted edge betweenness of the current edge, and For any pair of nodes in the state transition diagram For the first quantity, For the second quantity, The weight is the weight corresponding to the current edge.
5. The method according to claim 1, characterized in that, Step S60 includes: The candidate edge set is determined based on at least one edge with the smallest frequency-weighted edge betweenness in the state transition diagram. If the candidate edge set contains only one candidate edge, then the candidate edge is determined as the edge to be processed; If the candidate edge set contains multiple candidate edges, then the edge to be processed is determined from among the multiple candidate edges; A judgment result is obtained based on the edge to be processed and the state transition graph. The judgment result is used to indicate whether a new starting node or a new ending node is added to the state transition graph after the edge to be processed is removed. The new starting node refers to a node in the state transition graph, other than the already marked starting node, whose out-degree becomes 0 due to the removal of the edge to be processed. The new ending node refers to a node in the state transition graph, other than the already marked ending node, whose in-degree becomes 0 due to the removal of the edge to be processed. If the judgment result indicates that after removing the edge to be processed, no new starting node or new ending node is added to the state transition graph, then the edge to be processed is identified as a redundant edge, and the redundant edge is removed from the state transition graph. If the judgment result indicates that after removing the edge to be processed, a new starting node or a new ending node is added to the state transition graph, then the edge to be processed is retained.
6. The method according to claim 5, characterized in that, The step of determining the edge to be processed from the plurality of candidate edges includes: The multiple candidate edges are compared using the weights corresponding to each candidate edge to obtain a comparison result, wherein the weights corresponding to the candidate edges are used to characterize the frequency of the action corresponding to the candidate edge in the sequence sample set; If the comparison result indicates that the weights corresponding to each candidate edge are not the same, then the candidate edge with the smallest weight among the multiple candidate edges is determined as the edge to be processed. If the comparison result indicates that the weights corresponding to each of the candidate edges are the same, then the edge to be processed is randomly determined from the plurality of candidate edges.
7. The method according to claim 5, characterized in that, The method further includes: If the judgment result indicates that after removing the edge to be processed, a new starting node or a new ending node is added to the state transition graph, then the stopping condition is determined to be met.
8. The method according to claim 5, characterized in that, If the determination result indicates that after removing the edge to be processed, no new starting node or new ending node is added to the state transition graph, then the edge to be processed is identified as a redundant edge, and after removing the redundant edge from the state transition graph, the method further includes: The weights of the remaining edges in the state transition graph are normalized based on the frequency of occurrence of the actions corresponding to the remaining edges in the sequence sample set.
9. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the key action extraction method based on the state transition diagram as described in any one of claims 1 to 8.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the key action extraction method based on a state transition diagram as described in any one of claims 1 to 8.