Player character and BOSS fight real-time analysis method, system and equipment and medium
Through real-time acquisition and cross-verification of game data, combined with advanced model identification and modeling technology, dynamic battlefield analysis is generated, and the problems of low manual recording efficiency, single analysis and lagging feedback in the existing technology are solved, achieving efficient and accurate real-time battle strategy generation and adjustment.
Patent Information
- Application Number
- CN202510452362.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
AI Technical Summary
Existing game battle analysis technology relies on manual data recording inefficient and poor accuracy, single analysis dimensions, rigid model, and lagging feedback, and unable to provide real-time strategy adjustment suggestions.
Through real-time synchronous acquisition of screen pixel streams and game engine API data, cross-verification is adopted for inter-frame differential verification algorithm, combined with the improved YOLOv5 model to identify obstacles, use the three-dimensional spatial interpolation algorithm to generate environmental perception data, build a dynamic attribute relationship map, use the two-way graph attention network modeling skills to restrain relationships, use the reinforcement learning decision model to generate real-time battle strategies, and update model parameters online through the meta-gradient descent algorithm.
It realizes efficient and accurate analysis of player character battles, provides real-time strategy adjustment suggestions, improves analysis efficiency and accuracy, adapts to different combat styles and character characteristics, reduces manual intervention, and enhances the accuracy and reliability of data collection.
Smart Images

Figure CN120285577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a real-time analysis method, system, device and medium for the battle between a player character and a BOSS. Background Art
[0002] With the rapid development of e-sports and online games, the battle analysis between a player character and a BOSS (Boss monster) has become a key link in enhancing the game experience and competitive level. However, the existing game battle analysis technologies generally have the following defects:
[0003] 1. Strong dependence on manual work: Traditional battle analysis methods mainly rely on players to manually record battle data, such as key indicators like skill release time, damage output, and health value changes. This method is not only inefficient, as players need to be distracted from the intense battle to record data, but also extremely prone to inaccurate data recording due to human negligence or subjective judgment, which seriously affects the reliability and effectiveness of subsequent analysis results.
[0004] 2. Single analysis dimension: Most current analysis methods only focus on basic attributes of players such as health value and magic value, while ignoring deep factors such as environmental interaction, character skill restraint relationships, and the impact of dynamic obstacles. This single analysis dimension limits the comprehensive understanding of the battle process and makes it difficult to discover deep battle strategies.
[0005] 3. Rigid model: Existing battle analysis models usually adopt fixed analysis templates and parameter settings, lacking adaptability to different battle styles and character characteristics. In actual battles, different players may have completely different operation habits, skill release sequences, and tactical preferences, and BOSSes also often have unique attack patterns, skill combinations, and weakness distributions. The rigid analysis model is difficult to accurately capture these differences, resulting in a deviation between the analysis result and the actual situation and being unable to provide effective strategic guidance for players.
[0006] 4. Feedback lag: Traditional battle analysis methods usually conduct analysis after the battle ends, through methods such as replay of videos and statistical data for post-event analysis. Although this method can provide a certain summary of the battle and lessons learned, due to the slow generation of analysis results, it cannot provide real-time strategic adjustment suggestions for players. In fast-paced game battles, this feedback lag may cause players to miss the best response opportunity and be unable to adjust tactics and strategies in a timely manner, thus affecting the battle result and game experience. Summary of the Invention
[0007] The object of the present invention is to provide a real-time analysis method, system, device and medium for the battle between a player character and a BOSS, which improves the analysis efficiency and accuracy of the battle between the player character and the BOSS, and can help the player quickly understand the battle dynamics, improve the combat strategy and operation skills, so as to solve at least one of the above-mentioned prior art problems.
[0008] In a first aspect, the present invention provides a real-time analysis method for the battle between a player character and a BOSS, and the method specifically includes:
[0009] Collect the screen pixel stream and game engine API data in real-time synchronization, and use the inter-frame difference verification algorithm to cross-verify the screen pixel stream and game engine API data to obtain the verified combat data;
[0010] Take the verified combat data as the input, identify the dynamic obstacles on the battlefield through an improved YOLOv5 model, and generate the environmental perception data by using the three-dimensional space interpolation algorithm;
[0011] Construct a dynamic attribute relationship graph according to the environmental perception data and the pre-set character status data, and use a bi-directional graph attention network to model the time-varying restraint relationship between character skills, and update the edge weight matrix through a time sliding window mechanism;
[0012] Input the dynamic attribute relationship graph and the player's historical operation data into the reinforcement learning decision model together to generate a battle strategy tree, and dynamically adjust the search depth of the battle strategy tree through the real-time calculated battle complexity index to generate the optimal battle strategy;
[0013] According to the strategy execution effect data of the battle strategy tree, online update the decision model parameters through the meta-gradient descent algorithm.
[0014] In a second aspect, the present invention provides a real-time analysis system for the battle between a player character and a BOSS, and the system specifically includes:
[0015] A first real-time analysis module, which is used to collect the screen pixel stream and game engine API data in real-time synchronization, and use the inter-frame difference verification algorithm to cross-verify the screen pixel stream and game engine API data to obtain the verified combat data;
[0016] A second real-time analysis module, which is used to take the verified combat data as the input, identify the dynamic obstacles on the battlefield through an improved YOLOv5 model, and generate the environmental perception data by using the three-dimensional space interpolation algorithm;
[0017] A third real-time analysis module, which is used to construct a dynamic attribute relationship graph according to the environmental perception data and the pre-set character status data, and use a bi-directional graph attention network to model the time-varying restraint relationship between character skills, and update the edge weight matrix through a time sliding window mechanism;
[0018] The fourth real-time analysis module is used to jointly input the dynamic attribute relationship graph and the player's historical operation data into the reinforcement learning decision model to generate a battle strategy tree, dynamically adjust the search depth of the battle strategy tree through the real-time calculated combat complexity index, and generate the optimal battle strategy;
[0019] The fifth real-time analysis module is used to online update the decision model parameters according to the strategy execution effect data of the battle strategy tree through the meta-gradient descent algorithm.
[0020] In a third aspect, the present invention provides a computer device, including: a memory, a processor, and a computer program stored on the memory. When the computer program is executed on the processor, it implements the real-time analysis method for the player character to fight against the BOSS as described in any one of the above methods.
[0021] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the real-time analysis method for the player character to fight against the BOSS as described in any one of the above methods.
[0022] Compared with the prior art, the present invention has at least one of the following technical effects:
[0023] 1. The present invention improves the analysis efficiency and accuracy of the player character fighting against the BOSS, and can help the player quickly understand the battle dynamics, improve the battle strategy and operation skills.
[0024] 2. The present invention realizes the automatic analysis of the battle process by synchronously collecting the screen pixel stream and the game engine API data in real time, reduces the manual intervention, and improves the analysis efficiency.
[0025] 3. The present invention not only focuses on the player's basic attributes, but also fully considers deep factors such as environmental interaction, character skill restraint relationship, and dynamic obstacle influence, and realizes the comprehensive analysis of the battle process.
[0026] 4. The present invention adopts advanced technologies such as dynamic attribute relationship graph and bidirectional graph attention network, can adapt to different battle styles and character characteristics, and improves the accuracy and reliability of the analysis results.
[0027] 5. The present invention realizes the real-time adjustment and optimization of the battle strategy through the reinforcement learning decision model and the meta-gradient descent algorithm, provides timely strategy adjustment suggestions for the player, and helps to improve the battle result.
[0028] 6. The present invention realizes the efficient and accurate data collection of the game interface through adaptive region segmentation and differential sampling, and improves the data collection efficiency and quality.
[0029] 7. The present invention realizes the modal consistency verification of the screen pixel stream and the game engine API data through the inter-frame difference verification algorithm, and triggers the repair process in case of anomalies to ensure the accuracy and reliability of the data.
[0030] 8. The present invention realizes dynamic obstacle detection by improving the YOLOv5 model, and combines the cross-modal attention mechanism to generate joint feature representations, improving the accuracy of obstacle detection and the richness of feature representations.
[0031] 9. The present invention constructs an environmental heat map using a three-dimensional adaptive kernel density estimation algorithm to generate environmental perception data, enhancing the spatial perception ability of the battlefield environment.
[0032] 10. The present invention encodes the skill attributes into multi-dimensional feature vectors and calculates the synergy effect coefficients, providing basic data support for constructing a dynamic attribute relationship graph. Further, based on the skill embedding representation and the synergy effect coefficients, a dynamic attribute relationship graph is constructed to reveal the complex relationships between skills, providing strong support for strategy formulation.
[0033] 11. The present invention uses a spatio-temporal graph attention network and a bidirectional message passing mechanism to dynamically reason about the skill restraint relationship, realizing the real-time update and accurate modeling of the skill restraint relationship.
[0034] 12. The present invention dynamically adjusts the edge weight parameters of the dynamic attribute relationship graph based on real-time combat log data, improving the accuracy and adaptability of the skill restraint relationship.
[0035] 13. The present invention automatically expands the nodes of the dynamic attribute relationship graph and initializes the connection relationships, realizing the rapid integration of new skill combinations and strategy updates.
[0036] 14. The present invention fuses multi-source data into a unified state representation, and uses an improved Monte Carlo tree search algorithm to construct an initial battle strategy tree, providing a basis for strategy generation. Further, the battle strategy tree is optimized through a node value evaluation function and a dynamic search depth adjustment formula, and the optimal battle strategy is output, improving the effectiveness and adaptability of the strategy. When a low-reward value branch is detected, a pruning operation is triggered to improve the search efficiency and strategy quality of the strategy tree. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] Figure 1It is a schematic flowchart of a method for real-time analysis of the battle between a player character and a BOSS provided by an embodiment of the present invention;
[0039] Figure 2 It is a schematic structural diagram of a system for real-time analysis of the battle between a player character and a BOSS provided by an embodiment of the present invention;
[0040] Figure 3 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0041] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0042] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0043] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0044] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.
[0045] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0046] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0047] In the embodiments of this application, the execution subject of the process includes a terminal device. The terminal device includes, but is not limited to: devices such as servers, computers, smartphones, and tablet computers that can execute the methods disclosed in this application. Figure 1 The flowchart showing the method for real-time analysis of the battle between a player character and a BOSS disclosed in the first embodiment of the present invention is shown as follows and will be described in detail:
[0048] S101, collect the screen pixel stream and game engine API data in real-time synchronization, and use the inter-frame difference verification algorithm to cross-verify the screen pixel stream and game engine API data to obtain the verified battle data.
[0049] In this embodiment, in the game scenario of the battle between a player character and a BOSS, accurately obtaining battle data is crucial for analyzing the player's operations and formulating battle strategies. Traditional methods may rely only on a single data source (such as the screen pixel stream or game engine API data), but due to the complexity of the game running environment (such as screen rendering delay, API data transmission error, etc.), the single data source may have problems of inaccurate or missing data, affecting subsequent analysis and decision-making.
[0050] Specifically, by using game development tools or third-party software interfaces, the pixel information of the game screen can be captured in real-time on the player's device. A fixed acquisition frequency (such as 30 frames per second) can be set, and the screen pixel data of each frame can be saved in an image format (such as an RGB image), and the acquisition timestamp is recorded. Through the API interface provided by the game engine, relevant data inside the game can be obtained, such as the positions, health points, skill states, etc. of the player character and the BOSS. Similarly, the acquisition timestamp is recorded for each data item to ensure the time synchronization with the acquisition of the screen pixel stream.
[0051] For the screen pixel stream, calculate the pixel difference between two adjacent frames of images. Specifically, subtract the pixel value of each pixel in the current frame from the pixel value at the corresponding position in the previous frame to obtain a difference matrix. If the differences of most pixels in the difference matrix are less than a preset threshold (such as 10), it is considered that the two frames of images are basically the same in terms of the picture content; otherwise, it is considered that the picture has changed.
[0052] According to the acquisition timestamps, align the screen pixel stream and the game engine API data in time to ensure that the data at the same moment can be cross-validated. For each time-aligned data point, associate the change in the picture content in the screen pixel stream with the state change in the game engine API data. For example, when the screen pixel stream shows that the player character has released a skill, check whether the state of the skill in the game engine API data has changed from "not released" to "released".
[0053] If the changes in the screen pixel stream and the game engine API data are consistent at the same moment, it is considered that this data point is valid and mark it as passed the verification. If there is an inconsistency between the two, such as the screen pixel stream shows that the BOSS has been attacked, but the health of the BOSS in the game engine API data has not changed, it is considered that this data point may be abnormal and further processing is required (such as marking it as suspicious data, discarding this data point, or performing data repair).
[0054] In this embodiment, by synchronously collecting the screen pixel stream and the game engine API data in real time and using the inter-frame difference verification algorithm for cross-validation, it is possible to effectively avoid the problems of inaccurate or missing data that may exist in a single data source. For example, in the case where the screen pixel stream is out of sync with the game engine API data due to rendering latency, cross-validation can promptly detect this inconsistency and improve the data accuracy through reasonable processing methods (such as data repair or discarding suspicious data).
[0055] The inter-frame difference verification algorithm can detect abnormal changes in the screen pixel stream and the game engine API data, such as data mutations and data losses. By processing these abnormal data, the reliability of the data can be improved, providing a more accurate data basis for subsequent steps such as battlefield dynamic obstacle recognition and environmental perception data generation.
[0056] Accurately verifying the combat data is a prerequisite for subsequent steps such as battlefield dynamic obstacle recognition, modeling of character skill restraint relationships, and generation of combat strategies. Through the technical solution of this embodiment, it is possible to provide reliable data support for the entire real-time analysis method of the player character's battle with the BOSS, improving the performance and accuracy of the entire system.
[0057] S102. Take the verified combat data as input, identify dynamic obstacles on the battlefield through an improved YOLOv5 model, and use a three-dimensional spatial interpolation algorithm to generate environmental perception data.
[0058] In this embodiment, in the game scenario where the player character battles against the BOSS, accurately identifying dynamic obstacles on the battlefield and generating environmental perception data are crucial for formulating effective battle strategies. Traditional methods are difficult to cope with the complex and ever-changing battlefield environment, have a low recognition accuracy for dynamic obstacles, and cannot comprehensively and accurately generate environmental perception data, affecting subsequent decision-making and analysis.
[0059] Specifically, introduce an attention mechanism, such as the SE (Squeeze-and-Excitation) attention module, into the backbone network and detection head of the YOLOv5 model. The SE module adaptively adjusts the weights of each channel by learning the relationships between channels, enabling the model to pay more attention to the feature channels related to obstacles and improving the accuracy of obstacle recognition. For the small dynamic obstacles that may appear in the game, add a small target detection layer to the detection head of the YOLOv5 model. This detection layer has a smaller receptive field and higher resolution, enabling it to better detect small target obstacles. Perform data augmentation on the image data in the verified combat data, such as random rotation, flipping, scaling, adding noise, etc. Through data augmentation, increase the diversity of training data and improve the generalization ability of the model.
[0060] Collect verified combat data from a large number of game battle scenarios, label the dynamic obstacles in the images, and construct a training dataset. The labeling content includes the category of the obstacle (such as monster, trap, obstacle, etc.) and the bounding box coordinates. Use the constructed training dataset to train the improved YOLOv5 model. Optimize the model using common object detection loss functions (such as cross-entropy loss and IoU loss). Through multiple iterations of training, enable the model to accurately identify dynamic obstacles on the battlefield. Input the verified combat data into the trained improved YOLOv5 model, and the model outputs the category and bounding box coordinate information of each dynamic obstacle.
[0061] According to the actual situation of the game scene, a three-dimensional space coordinate system is constructed with the player character as the origin, and the directions and unit lengths of the three coordinate axes X, Y, and Z are determined. The coordinate information of the bounding boxes of the dynamic obstacles recognized by the improved YOLOv5 model is mapped into the three-dimensional space coordinate system. Since the image data in the verification combat data is usually two-dimensional, it is necessary to determine the position of the obstacle on the Z-axis through game engine API data or other prior knowledge (such as the height information of the obstacle). For the areas that may exist in the battlefield and are not directly detected, a three-dimensional space interpolation algorithm (such as Kriging interpolation algorithm) is used to interpolate and estimate the position and attributes of the obstacles. Through interpolation, a continuous three-dimensional space environmental perception data is generated, which contains information such as the positions, categories, and possible movement trajectories of all dynamic obstacles in the battlefield.
[0062] In this embodiment, through improvement measures such as introducing an attention mechanism, adding a small target detection layer, and performing data augmentation, the improved YOLOv5 model can more accurately identify dynamic obstacles in the battlefield. Especially in the complex and changeable battlefield environment, the recognition accuracy of smaller, blurred, or occluded dynamic obstacles has been significantly improved.
[0063] The environmental perception data generated by using the three-dimensional space interpolation algorithm contains detailed information of all dynamic obstacles in the battlefield, including positions, categories, and possible movement trajectories, etc. These data provide a more comprehensive and accurate basis for formulating subsequent battle strategies, enabling players to better understand the battlefield environment and make more reasonable decisions.
[0064] The improved YOLOv5 model has a high inference speed and can process the verification combat data in real time to quickly identify dynamic obstacles. At the same time, the three-dimensional space interpolation algorithm can be dynamically updated according to the obstacle information obtained in real time, so that the environmental perception data always remains up-to-date, enhancing the real-time performance and adaptability of the system.
[0065] S103. Construct a dynamic attribute relationship graph based on the environmental perception data and the pre-set character state data, and use a bi-directional graph attention network to model the time-varying restraint relationship between character skills, and update the edge weight matrix through a time sliding window mechanism.
[0066] In this embodiment, in the game scene of the player character fighting against the BOSS, the restraint relationship between character skills is not static, but is affected by environmental perception data (such as battlefield terrain, obstacle distribution, etc.) and character state data (such as health value, energy value, skill cooldown time, etc.). Traditional methods usually use a fixed restraint relationship matrix to describe the restraint relationship between skills, which cannot reflect the changes in the skill restraint relationship in real time, resulting in inaccurate and inflexible formulation of battle strategies.
[0067] Specifically, collect the generated environmental perception data, including information such as the positions, categories, and movement trajectories of dynamic obstacles in the battlefield, as well as the terrain features of the battlefield (such as flat areas, areas with dense obstacles, etc.). Obtain the pre-set character status data of the player character and the BOSS, including health points, energy values, skill cooldown times, skill attributes (such as attack power, defense power, skill range, etc.). Clean and standardize the collected environmental perception data and character status data, and convert different types of data into a unified format for subsequent graph construction.
[0068] Define the player character, the BOSS, and the key obstacles in the battlefield (such as obstacles that have a significant impact on the battle) as the nodes of the graph. Each node contains corresponding attribute information, such as the type of the node (player character, BOSS, obstacle), position, status, etc. Define the edges between the nodes according to the potential restraint relationships between the character skills, the influence of the environmental perception data on the skill effects, and the influence of the character status data on the skill release. For example, if a certain skill of the player character has a restraining effect on a certain skill of the BOSS under specific terrain conditions, establish an edge between the player character node and the BOSS node, and mark the attributes of the edge, such as terrain conditions, skill restraint relationships, etc. Store the constructed dynamic attribute relationship graph in a graph database for subsequent query and analysis.
[0069] The bidirectional graph attention network includes an input layer, a bidirectional graph attention layer, and an output layer. In the input layer, use the dynamic attribute relationship graph as the input. The node features in the graph include character status data and obstacle attribute data, and the edge features include skill restraint relationships and environmental condition data. In the bidirectional graph attention layer, adopt the bidirectional graph attention mechanism to propagate information forward and backward respectively, so that each node can simultaneously consider the influence of its neighbor nodes on it and its influence on the neighbor nodes. For example, during forward propagation, consider the influence of a certain skill of the BOSS on the player character's skills; during backward propagation, consider the restraining effect of the player character's skills on the BOSS's skills. In the output layer, output the time-varying restraint relationships of each skill on other skills, represented in the form of an edge weight matrix.
[0070] Design a comprehensive loss function, including the accuracy loss of the skill restraint relationship and the smoothness loss of the edge weight matrix. The accuracy loss of the skill restraint relationship is calculated by comparing the restraint relationship output by the network with the restraint relationship in the actual combat data; the smoothness loss of the edge weight matrix is calculated by restricting the change range of the edge weight matrix at adjacent time steps. Collect a large amount of game battle data, including the order of character skill releases, battle results, environmental perception data, etc., to construct a training data set. Use the training data set to train the bidirectional graph attention network, and optimize the network parameters through the backpropagation algorithm, so that the network can accurately model the time-varying restraint relationships between character skills.
[0071] Set a time-sliding window of a fixed size, such as the battle data in the last 10 seconds. Integrate the environmental perception data, character status data, and skill release data within the time-sliding window as the input data at the current moment. Use the input data at the current moment to recalculate the edge weight matrix through the trained bidirectional graph attention network to reflect the latest time-varying restraint relationship between character skills.
[0072] In this embodiment, through the time-sliding window mechanism, the edge weight matrix can be updated in real time to accurately reflect the time-varying restraint relationship between character skills affected by environmental perception data and character status data, enabling players to adjust their battle strategies according to the latest skill restraint relationship. The bidirectional graph attention network can comprehensively consider the bidirectional influence between nodes, enabling players to better understand the interaction between skills, thereby formulating more accurate and flexible battle strategies. The combination of the dynamic attribute relationship graph and the bidirectional graph attention network enables the system to adapt to different battlefield environments and character statuses, improving the robustness of the system.
[0073] S104, input the dynamic attribute relationship graph and the player's historical operation data into the reinforcement learning decision model together to generate a battle strategy tree, and dynamically adjust the search depth of the battle strategy tree through the battle complexity index calculated in real time to generate the optimal battle strategy.
[0074] In this embodiment, in complex multiplayer online battle games, accurately generating effective battle strategies is crucial for improving players' gaming experience and win rate. Traditional battle strategy generation methods often rely only on a single data source, such as the player's historical operation data, while ignoring the rich dynamic attribute relationship information in the game world. In addition, the fixed search depth of the battle strategy tree is difficult to adapt to battle scenarios of different complexities, which may lead to low strategy generation efficiency or poor strategy quality. Therefore, a method that comprehensively utilizes multiple data sources and can dynamically adjust the search depth is needed to generate the optimal battle strategy.
[0075] Specifically, clean and organize the player's historical operation data, and extract key operation information, such as operation type (attack, defense, use of items, etc.), operation target, operation time, etc. Convert the player's historical operation data into a numerical feature vector, for example, count the frequency and success rate of the player using various operations in different time periods.
[0076] Adopt deep reinforcement learning models, such as variants of the Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) algorithms. The input of the model is the feature representation of the dynamic attribute relationship graph and the feature vector of the player's historical operation data, and the output is the Q-values or action probability distributions of all possible actions. Use historical battle data to train the reinforcement learning decision model. During the training process, take the dynamic attribute relationship graph and the player's historical operation data as inputs, take the win-loss result of the battle as the reward signal, and use appropriate loss functions (such as mean squared error loss function) and optimization algorithms (such as Adam optimization algorithm) to update the model parameters.
[0077] During the battle, according to the current game state (including the current state of the dynamic attribute relationship graph and the latest information of the player's historical operation data), use the trained reinforcement learning decision model to output the Q-values or action probability distributions of all possible actions, and construct a battle strategy tree. Each node of the strategy tree represents a possible action, and each edge represents the transition from one action to another. Set an initial search depth for searching in the battle strategy tree.
[0078] Define multiple battle complexity metrics, such as the number of characters, the difference degree of character attributes, the frequency of item use, the environmental change rate, etc. According to the current game state, calculate the values of each complexity metric in real time, and use the method of weighted summation to calculate the battle complexity index.
[0079] According to the calculated battle complexity index, dynamically adjust the search depth of the battle strategy tree. When the battle complexity index is relatively high, increase the search depth to explore more possible strategies; when the battle complexity index is relatively low, decrease the search depth to improve the efficiency of strategy generation. Set the upper and lower limits of the search depth, such as the minimum search depth is 2 and the maximum search depth is 5. When the battle complexity index is greater than the high complexity threshold, increase the search depth by 1; when the battle complexity index is less than the low complexity threshold, decrease the search depth by 1.
[0080] In the battle strategy tree, use methods such as Monte Carlo Tree Search (MCTS) to search within the adjusted search depth range, and select the action sequence with the highest evaluation value as the optimal battle strategy. Send the generated optimal battle strategy to the game client, and let the player or game character execute the corresponding operations.
[0081] In this embodiment, by comprehensively utilizing the dynamic attribute relationship graph and the player's historical operation data, the reinforcement learning decision model can learn more comprehensive and accurate policy knowledge, thereby generating higher-quality battle strategies. According to the battle complexity index calculated in real time, the search depth of the battle strategy tree is dynamically adjusted, enabling the method to adapt to battle scenarios of different complexities and improving the policy generation efficiency while ensuring the policy quality. The generated optimal battle strategy can better handle various complex situations in the game, enhancing the strategic and interesting aspects of the game and bringing a better gaming experience to players.
[0082] S105. According to the policy execution effect data of the battle strategy tree, online update the decision model parameters through the meta-gradient descent algorithm.
[0083] In this embodiment, in a complex and ever-changing battle scenario, the pre-trained battle decision model may not be able to well adapt to the new battle environment and opponent strategies. Traditional model update methods usually require collecting a large amount of offline data for re-training. This method is not only time-consuming but also unable to respond promptly to the dynamic changes during the battle. Therefore, a method for online updating the decision model parameters is needed to improve the real-time adaptability and decision-making quality of the model.
[0084] Specifically, during the battle, the execution results of each decision action are recorded in real time, including the victory or defeat of the battle, changes in character status (such as health value, energy value, etc.), and resource consumption (such as skill cooldown time, number of item uses, etc.). The collected policy execution effect data is preprocessed and converted into a format suitable for model update. For example, the victory or defeat of the battle is converted into a binary label (1 for victory, 0 for defeat), and the changes in character status and resource consumption are converted into numerical features. Using the current decision model parameters as the initial parameters, an inner loss function is constructed using the policy execution effect data. The inner loss function can be a weighted sum of the classification loss function for the victory or defeat of the battle and the regression loss function for the changes in character status and resource consumption. By performing inner optimization steps (such as using the mini-batch stochastic gradient descent algorithm), an intermediate model parameter is obtained. Calculate the gradient of the intermediate model parameter with respect to the initial model parameter, and this gradient is the meta-gradient. The meta-gradient reflects the impact of the policy execution effect data on the initial model parameters. According to the calculated meta-gradient, an appropriate optimization algorithm (such as the Adam optimization algorithm) is used to online update the initial decision model parameters. After each parameter update, a small portion of the battle data is used to online evaluate the updated model, and the evaluation metrics of the model (such as win rate, average score, etc.) are calculated. According to the online evaluation results, if the performance of the model is improved, continue to update the parameters according to the current update strategy; if the performance of the model decreases, adjust the relevant parameters of the meta-gradient descent algorithm (such as the learning rate, number of inner optimization steps, etc.), or re-collect the policy execution effect data for model update.
[0085] In this embodiment, by online updating the decision model parameters, the model can promptly respond to dynamic changes during the battle process, such as the adjustment of the opponent's strategy, the change of the environment, etc., thereby improving the adaptability of the model in different battle scenarios. Updating the model based on the policy execution effect data enables the model to learn more effective battle strategies, improving the battle victory rate and decision-making quality. Compared with the traditional offline retraining method, the online updating method does not require collecting a large amount of offline data, reducing the time and cost of data collection and model training.
[0086] In some embodiments, in the above step S101, the real-time synchronous acquisition of the screen pixel stream and the game engine API data specifically includes:
[0087] Using an adaptive region segmentation algorithm to divide the game interface into multiple game regions, where the game regions include a skill release area, an environment interaction area, and a status display area;
[0088] Setting different sampling frequencies for each game region, and using the corresponding sampling frequency to collect different screen pixel streams in each game region in real time;
[0089] Constructing a state-event mapping table, and converting the player character attribute vector, BOSS state vector, and environment event vector output by the game engine into a continuous state vector according to the state-event mapping table to form game engine API data.
[0090] In this embodiment, the original image of the game interface is obtained, and preprocessing operations such as grayscale conversion and denoising are performed on it to improve the accuracy of subsequent region segmentation. Edge features, color features, texture features, etc. of the game interface image are extracted. For example, the Canny edge detection algorithm is used to extract the edge information of the image, and the color distribution of the image is analyzed using a color histogram. An adaptive region segmentation algorithm, such as a region segmentation algorithm based on clustering (such as the K-means clustering algorithm) or a region segmentation algorithm based on graph theory (such as the normalized cut algorithm), is used to divide the game interface into multiple game regions according to the extracted features. In this embodiment, the game interface is divided into a skill release area, an environment interaction area, and a status display area. The skill release area contains relevant buttons and display areas for the player and enemy characters to release skills; the environment interaction area includes elements for interacting with the game environment, such as pickable items, destructible scene objects, etc.; the status display area shows information such as player character attributes and health values.
[0091] Differentiated sampling frequencies are set for each game area according to its characteristics and importance. For example, in the skill release area, since the operations of skill release are relatively critical and change relatively frequently, a higher sampling frequency is set, such as 30 frames per second. This area contains skill icons, skill cooldown displays, etc., and it is necessary to capture the state and timing of skill release in real time. In the environmental interaction area, a medium sampling frequency is set, such as 15 frames per second. This area involves interactions with the game environment, such as resource collection, obstacle avoidance, etc., and a certain degree of real-time performance is required to ensure the timeliness of interactions. In the status display area: a lower sampling frequency is set, such as 5 frames per second. This area mainly displays status information such as the player's health value and experience value, which changes relatively slowly and does not require a too high sampling frequency.
[0092] Construct a state-event mapping table to classify and encode the player character attribute vectors (such as health value, attack power, etc.), BOSS state vectors (such as blood volume, skill release state, etc.), and environmental event vectors (such as item appearance, terrain change, etc.) output by the game engine. For example, encode the player character attribute vector as P = {p1, p2,..., p n}, where p1, p2,..., p n represent different attribute values; encode the BOSS state vector as B = {b1, b2,..., b m}; encode the environmental event vector as E = {e1, e2,..., e k}. According to the state-event mapping table, convert the encoded vectors into continuous state vectors. For example, combine the player character attribute vector P, BOSS state vector B, and environmental event vector E into a comprehensive state vector S = [P, B, E] to form the game engine API data.
[0093] In this embodiment, the game interface is divided into different game areas through an adaptive region segmentation algorithm, and differentiated sampling frequencies are set for each area, avoiding the data redundancy problem caused by full-screen sampling and significantly improving the efficiency of data collection.
[0094] The setting of differentiated sampling frequencies enables each game area to collect data at an appropriate frequency, avoiding data loss or redundancy caused by improper sampling frequencies. For example, the high sampling frequency in the skill release area can accurately capture the instantaneous state of skill release, improving the accuracy of analyzing skill usage.
[0095] By constructing a state-event mapping table, different types of vectors output by the game engine are converted into continuous state vectors, making the collected game engine API data more standardized and easier to analyze. Developers can conduct more in-depth game behavior analysis, strategy formulation, and auxiliary function development based on these continuous state vectors.
[0096] In some embodiments, in the above step S101, the cross - verification of the screen pixel stream and the game engine API data using the inter - frame difference check algorithm specifically includes:
[0097] Generate an alignment sequence for the parsed result of the screen pixel stream and the state vector of the game engine API data through an interpolation algorithm;
[0098] Perform two - way difference verification on the screen pixel stream and the game engine API data in the alignment sequence, and calculate to obtain a modal consistency index;
[0099] When the modal consistency index exceeds the preset deviation threshold, trigger an abnormal data repair process based on motion compensation.
[0100] In this embodiment, the screen pixel stream can intuitively reflect the visual information of the game interface, while the game engine API data provides detailed information about the internal state of the game. However, due to factors such as acquisition device errors, network latency, and game engine bugs, these two types of data may be inconsistent, affecting the accuracy of game data analysis, strategy formulation, and auxiliary functions. Traditional data verification methods can often only verify a single data source and cannot effectively detect the inconsistency between the two types of data. Therefore, a method capable of cross - verifying the screen pixel stream and the game engine API data simultaneously is needed.
[0101] Specifically, perform image processing and analysis on the collected screen pixel stream, extract key information such as the position of the player character, the skill release state, the display of the BOSS's health, etc., and convert it into a structured data format as the parsed result of the screen pixel stream. Obtain the player character attribute vector (such as health value, attack power, etc.), the BOSS state vector (such as health, skill release state, etc.), and the environmental event vector (such as the appearance of items, terrain changes, etc.) from the game engine to form the state vector of the game engine API data. Add timestamp information to the parsed result of the screen pixel stream and the state vector of the game engine API data for subsequent alignment operations. Since the acquisition frequencies of the screen pixel stream and the game engine API data may be different, use an interpolation algorithm (such as linear interpolation, cubic spline interpolation, etc.) to perform time alignment on the two types of data to generate an alignment sequence. For example, assume that the screen pixel stream is collected 30 frames per second, and the game engine API data is collected 10 times per second. Through the interpolation algorithm, the game engine API data is interpolated to the same time interval as the screen pixel stream to form an alignment sequence.
[0102] Starting from the parsing result of the screen pixel stream, according to the game logic and rules, the expected game engine API data state vector is deduced. For example, according to the position and movement direction of the player character in the screen pixel stream, property vectors such as the speed and acceleration of the player character are deduced. The deduced expected state vector is compared with the actual game engine API data state vector, and the forward difference value is calculated.
[0103] Starting from the game engine API data state vector, according to the game rendering rules and display logic, the expected screen pixel stream parsing result is deduced. For example, according to the health status of the BOSS in the game engine, pixel information such as the display length and color of the BOSS health bar on the screen is deduced. The deduced expected parsing result is compared with the actual screen pixel stream parsing result, and the reverse difference value is calculated. Considering the forward difference value and the reverse difference value comprehensively, the modal consistency index is calculated. For example, the weighted average method can be used to assign different weights according to the importance of the forward difference and the reverse difference, and the modal consistency index C = α1·D1 + β1·D2 is calculated, where C represents the modal consistency index, D1 and D2 represent the forward difference value and the reverse difference value respectively, and α1 and β1 represent the weight coefficients of the forward difference value and the reverse difference value respectively.
[0104] According to the specific requirements and data characteristics of the game, a preset deviation threshold is set. For example, for an action game, the deviation threshold of the modal consistency index can be set to 0.2. When the modal consistency index exceeds 0.2, it is considered that the data is abnormal. The modal consistency index is monitored in real time. When it exceeds the preset deviation threshold, the abnormal data detection mechanism is triggered to determine the location and type of the abnormal data. According to the movement law and state change in the game, a motion compensation algorithm is used to repair the abnormal data. For example, if the position data of the player character in the screen pixel stream is abnormal, the position data can be corrected according to the speed and acceleration information of the player character in the game engine; if the health status of the BOSS in the game engine API data is abnormal, the health status can be calibrated according to the display of the BOSS health bar in the screen pixel stream.
[0105] In this embodiment, by using the inter-frame difference verification algorithm to cross-verify the screen pixel stream and the game engine API data, the inconsistency between the two types of data can be found in time, and a correction is made using the abnormal data repair process based on motion compensation, which significantly improves the accuracy of the game data. For example, when collecting and analyzing the data of a shooting game, after adopting this method, the data accuracy is increased by about 30%, effectively avoiding wrong decisions and abnormal behaviors caused by data inconsistency.
[0106] Timely detection and repair of abnormal data can avoid system crashes and abnormal behaviors caused by data errors, enhancing the stability of the game system. For example, in a multiplayer online game, accurate game data can ensure smooth interaction between players and reduce problems such as lag and disconnection caused by inconsistent data.
[0107] Accurate and consistent game data provides a reliable basis for game data analysis. Developers can conduct more in-depth game behavior analysis, strategy formulation, and auxiliary function development based on this data. For example, by analyzing the skill release data of players and the blood volume change data of BOSS, the game difficulty and skill balance can be optimized.
[0108] In some embodiments, in the above step S102, taking the verified combat data as input, identifying dynamic obstacles on the battlefield through an improved YOLOv5 model, and generating environmental perception data using a three-dimensional space interpolation algorithm specifically includes:
[0109] Add a spatio-temporal convolution module after the YOLOv5 backbone network to form an improved YOLOv5 model. Use the improved YOLOv5 model to perform dynamic obstacle detection on the verified combat data, and fuse temporal features through the spatio-temporal convolution module to form an obstacle detection result;
[0110] Based on a cross-modal attention mechanism, adaptively weight and fuse the game API state vector and pixel stream features in the verified combat data to generate a joint feature representation with semantic consistency;
[0111] Adopt a three-dimensional adaptive kernel density estimation algorithm to construct a dynamically updated environmental heat map according to the obstacle detection result and the preset terrain resistance parameters, and generate environmental perception data.
[0112] In this embodiment, accurately identifying dynamic obstacles (such as moving enemies, falling objects, etc.) on the battlefield and generating accurate environmental perception data can help the game AI make more reasonable decisions and enhance the realism and challenge of the game. However, traditional obstacle detection methods often rely on static image analysis, are difficult to process temporal information, and have poor fusion effects on different modal data (such as game API state vectors and pixel stream features). In addition, the generated environmental perception data usually lacks spatial continuity and dynamic update capabilities, unable to meet the requirements of complex battlefield environments.
[0113] Specifically, a spatio-temporal convolution module (STC module) is added after the YOLOv5 backbone network (such as CSPDarknet53). The STC module consists of alternating temporal convolution layers and spatial convolution layers. The temporal convolution layers are used to extract temporal features, and the spatial convolution layers are used to extract spatial features. By fusing temporal and spatial features, the model's ability to detect dynamic obstacles is enhanced. The verified combat data (such as consecutive game frame images) is used as input, and the improved YOLOv5 model is used for dynamic obstacle detection. The STC module extracts and fuses temporal features from the input consecutive frame images to form obstacle detection results, including information such as the position, category, and movement trajectory of the obstacles.
[0114] Extract the game API state vectors (such as player attributes, enemy states, etc.) and pixel stream features (such as information about the color, texture, edges, etc. of the image) from the verified combat data. Based on the cross-modal attention mechanism, design an attention network that adaptively assigns weights to features of different modalities according to the correlation between the game API state vectors and the pixel stream features. By means of weighted summation, the features of the two modalities are fused to generate a joint feature representation with semantic consistency.
[0115] According to the obstacle detection results and the preset terrain resistance parameters (such as the movement speed influence factors of different terrains), use the three-dimensional adaptive kernel density estimation algorithm to construct a dynamically updated environmental heat map. This algorithm takes into account the position, size, movement direction of the obstacles and the terrain resistance, performs density estimation on each spatial position, and generates a heat map reflecting the danger level and passage difficulty of the battlefield environment. The environmental heat map is used as environmental perception data, which contains environmental information of each position in the battlefield, such as obstacle distribution and terrain resistance, providing comprehensive environmental perception for the game AI and players.
[0116] Exemplarily, a large amount of image data of game combat scenarios is collected, and the dynamic obstacles in the images are labeled to generate a training data set. At the same time, the training data is verified to ensure the accuracy and consistency of the data. The improved YOLOv5 model is trained using the labeled training data set, and the hyperparameters of the model (such as learning rate, batch size, etc.) are adjusted to optimize the performance of the model. The verified combat image data is input into the trained improved YOLOv5 model, and the model outputs the detection results of the obstacles. The state vectors such as the player's health value and attack power are obtained from the game API, and the pixel stream features are extracted from the combat images. The attention network is trained using the labeled data set (including the game API state vectors, pixel stream features, and their corresponding joint feature representations) so that the network can learn the correlation between different modality features. The game API state vectors and pixel stream features in the verified combat data are input into the trained attention network, and the network outputs adaptive weights to perform weighted fusion on the features of the two modalities to generate a joint feature representation. For example, the dimension of the fused feature vector is 256, which contains the comprehensive features of the player state and image information. According to the terrain characteristics of the game scene, different terrain resistance parameters are set. For example, the resistance parameter for flat ground is 1.0, and the resistance parameter for mountainous terrain is 1.5. The obstacle detection results and terrain resistance parameters are input into the three-dimensional adaptive kernel density estimation algorithm to calculate the density value at each spatial position and generate an environmental heat map. For example, in the heat map, the location of the obstacle has a darker color, indicating a higher degree of danger at that location; the location of the flat ground has a lighter color, indicating easier passage. The environmental heat map is output as environmental perception data to provide a decision-making basis for the game AI. For example, the game AI can plan the player's movement route based on the environmental perception data to avoid dangerous areas.
[0117] In this embodiment, the improved YOLOv5 model can better fuse temporal features by adding a spatio-temporal convolution module, improving the detection accuracy of dynamic obstacles. The cross-modal attention mechanism can adaptively assign weights to the features of different modalities, making the fused joint feature representation have better semantic consistency. The three-dimensional adaptive kernel density estimation algorithm takes into account the comprehensive influence of obstacles and terrain, and the generated environmental heat map can more accurately reflect the degree of danger and ease of passage in the battlefield environment.
[0118] In some embodiments, in the above step S103, the constructing a dynamic attribute relationship graph according to the environmental perception data and the pre-set character state data specifically includes:
[0119] The skill attributes of the player character and the skill attributes of the BOSS are respectively encoded into multi-dimensional feature vectors, and skill embedding representations are generated through self-supervised learning;
[0120] Calculate the synergy effect coefficients of each skill combination in a specific battlefield environment based on environmental perception data;
[0121] Construct a dynamic attribute relationship graph according to the skill embedding representation and the synergy effect coefficients. The vertex set of the dynamic attribute relationship graph includes each character skill node, and the edge set of the dynamic attribute relationship graph includes each restraint relationship weight matrix based on design values.
[0122] In this embodiment, accurately evaluating the attribute relationship between the player character skills and the BOSS skills, as well as the synergy effects of different skill combinations in different battlefield environments, is crucial for formulating effective combat strategies. However, traditional methods usually perform simple calculations based on fixed skill attribute values, ignoring the impact of environmental factors on skill effects and being difficult to accurately depict the complex restraint relationships between skills. Therefore, a method that can comprehensively consider environmental perception data and character status data and dynamically construct an attribute relationship graph is needed.
[0123] Specifically, encode the player character skill attributes and the BOSS skill attributes into multi-dimensional feature vectors respectively. For example, for an attack skill of a player character, its attribute vector can include dimensions such as attack power, attack range, cooldown time, energy consumption, etc.; for a defense skill of a BOSS, its attribute vector can include dimensions such as defense power, damage reduction ratio, duration, etc.
[0124] Adopt a self-supervised learning algorithm (such as contrastive learning) to train the skill attribute vectors to generate skill embedding representations. The goal of self-supervised learning is to make similar skills closer in the embedding space and dissimilar skills farther apart by constructing positive and negative sample pairs. For example, skills with similar attack effects should have a high similarity in the vector space of their embedding representations.
[0125] Design a synergy effect calculation model based on environmental perception data, including terrain information (such as flat ground, mountainous areas, water areas, etc.), weather information (such as sunny days, rainy days, foggy days, etc.), and obstacle distributions, etc. This model considers the interaction of different skill combinations in different environments and calculates the synergy effect coefficients of each skill combination by simulating combat scenarios or analyzing historical combat data. For example, in a rainy environment, the combination of water-based attack skills and lightning-based attack skills may produce additional chain reactions, increasing the overall damage output, and its synergy effect coefficient will increase accordingly.
[0126] Use each character skill node as the vertex set of the dynamic attribute relationship graph. Each skill node corresponds to a skill embedding representation, which is used to uniquely identify the skill. According to the skill embedding representation and the synergy coefficient, construct various restraint relationship weight matrices based on design values as the edge set of the dynamic attribute relationship graph. The restraint relationship weight matrix reflects the restraint intensity between different skills. For example, a skill with a high armor penetration attribute has a high restraint intensity against a skill with lower defense, and the corresponding edge weight is larger.
[0127] Exemplarily, collect the attribute data of all player character skills and BOSS skills in the game to construct a skill attribute database. Encode the attributes of each skill and convert it into a multi-dimensional feature vector. For example, the attribute vector of the player character skill "Flame Impact" is [Attack Power = 100, Attack Range = 5 meters, Cooldown Time = 8 seconds, Energy Consumption = 50]; the attribute vector of the BOSS skill "Rock Shield" is [Defense = 200, Damage Reduction Ratio = 30%, Duration = 10 seconds].
[0128] Input the skill attribute vectors into a self-supervised learning model for training. During the training process, construct positive and negative sample pairs. For example, skills with similar attack mechanisms are used as positive sample pairs, and skills with completely different attack mechanisms are used as negative sample pairs. After multiple rounds of training, obtain the embedding representation of each skill. For example, the embedding representation of "Flame Impact" is a 128-dimensional vector [0.12, 0.34,..., 0.78], and the embedding representation of "Rock Shield" is [0.45, 0.67,..., 0.23].
[0129] Assume that the current battlefield environment is rainy, the terrain is mountainous, and there are some obstacles. Design a synergy calculation model that combines rules and machine learning. The rule part considers the impact of environmental factors on skill effects. For example, rain will increase the damage output of water-based skills; the machine learning part uses historical battle data for training to learn the synergy of different skill combinations in different environments. Input the embedding representations of player character skills and BOSS skills and environmental perception data into the synergy calculation model to calculate the synergy coefficients of various skill combinations. For example, the synergy coefficient of the player character skill "Flame Impact" and the BOSS skill "Rock Shield" in the rainy mountain environment is 0.8, indicating that this skill combination has a certain synergy effect in this environment.
[0130] Take all player character skills and BOSS skills as the vertices of the graph, and each vertex corresponds to an embedded representation of a skill. Based on the skill embedded representation and the synergy coefficient, construct a restraint relationship weight matrix. For example, design a 10×10 restraint relationship weight matrix, and each element in the matrix represents the restraint strength between two skills. Assume that the restraint strength between the player character skill "Flame Impact" and the BOSS skill "Frost Breath" is 0.6, then the value at the corresponding position in the matrix is 0.6.
[0131] In this embodiment, by encoding skill attributes as multi-dimensional feature vectors and generating embedded representations, the attribute characteristics of skills can be more accurately characterized, avoiding the limitations of the traditional method of evaluating based on a single value. Calculating the synergy coefficient based on environmental perception data can fully consider the influence of environmental factors on the effect of skill combinations, improving the accuracy of synergy prediction. The constructed dynamic attribute relationship graph can intuitively display the restraint relationships and synergy effects between different skills, providing more comprehensive skill combination information for players and game AIs.
[0132] In some embodiments, in the above step S103, the time-varying restraint relationship between character skills is modeled using a bi-directional graph attention network, and the edge weight matrix is updated through a time sliding window mechanism, specifically including:
[0133] Use a spatio-temporal graph attention network to extract the node features of the dynamic attribute relationship graph, aggregate neighborhood information through a bi-directional message passing mechanism, and perform dynamic reasoning on the skill restraint relationship in combination with the gating weight, where the gating weight is jointly determined by the current combat phase index and the environmental complexity;
[0134] Based on the collected real-time combat log data, use a time sliding window to count the skill interaction frequencies corresponding to different combat scenario types, and dynamically adjust the edge weight parameters of the dynamic attribute relationship graph in combination with the meta-gradient descent algorithm;
[0135] When a new skill combination is detected, automatically expand the nodes of the dynamic attribute relationship graph and initialize the connection relationships.
[0136] In this embodiment, accurately evaluating the restraint relationship between character skills and considering its changes over time is crucial for formulating effective combat strategies. Traditional methods usually perform simple calculations based on fixed skill attribute values, ignoring the dynamics and time-variability of skill restraint relationships. With the complexity of game scenarios and the diversification of player strategies, a method that can model and update skill restraint relationships in real time is needed.
[0137] Specifically, a spatio-temporal graph attention network is constructed, which can process the spatial structure information and time series information of the dynamic attribute relationship graph simultaneously. The network consists of multiple graph attention layers and recurrent neural network layers. The graph attention layers are used to extract node features and aggregate neighborhood information, and the recurrent neural network layers are used to process time series data.
[0138] A bidirectional message passing mechanism is adopted. In the graph attention layer, each node not only receives the forward messages from its neighboring nodes but also receives the reverse messages. In this way, the node can more comprehensively perceive the information of its neighboring nodes and improve the accuracy of information aggregation.
[0139] Dynamic reasoning is performed on the skill restraint relationship by combining gating weights. The gating weights are jointly determined by the current combat phase index and the environmental complexity. The current combat phase index can be calculated based on factors such as combat time and the health of both sides, and the environmental complexity can be evaluated based on factors such as terrain, weather, and obstacle distribution. The gating weights are generated by a learnable neural network model and are used to control the contribution degree of different information to the inference of the skill restraint relationship.
[0140] Combat log data in the game is collected in real-time, including information such as skill release time, skill type, and skill effect. Using the time sliding window mechanism, the combat log data is divided into multiple time windows. Within each time window, the skill interaction frequencies corresponding to different combat scenario types are counted. For example, in the melee scenario, the interaction frequency between melee skills is counted; in the ranged scenario, the interaction frequency between ranged skills is counted. Combining with the meta-gradient descent algorithm, the edge weight parameters of the dynamic attribute relationship graph are dynamically adjusted according to the skill interaction frequencies. The meta-gradient descent algorithm learns a more robust edge weight update strategy through joint optimization on multiple tasks (different time windows). At the end of each time window, the loss function is calculated based on the skill interaction frequencies within that window, and the edge weight parameters are updated using the meta-gradient descent algorithm.
[0141] The skill release situation in the game is monitored in real-time. When a new skill combination is detected, the automatic expansion mechanism is triggered. The nodes of the dynamic attribute relationship graph are automatically expanded, and the new skill is added as a new node to the graph. At the same time, the connection relationship between the new node and other nodes is initialized, which can be set as an initial weak connection or initialized according to certain rules.
[0142] Exemplarily, the dynamic attribute relationship graph is input into the spatio-temporal graph attention network for training. During the training process, the network extracts node features through the graph attention layer and aggregates neighborhood information using a bidirectional message passing mechanism. At the same time, the gating weight is calculated by combining the current combat stage index and the environmental complexity to perform dynamic reasoning on the skill restraint relationship. For example, in the initial stage of the battle, the current combat stage index is low and the environmental complexity is high, and the gating weight may be more inclined to consider the impact of environmental factors on the skill restraint relationship. After training, the spatio-temporal graph attention network can infer the strength of the restraint relationship between various skills based on the input dynamic attribute relationship graph and the current combat state. For example, it is inferred that the restraint relationship strength of the player character skill "Flame Strike" against the BOSS skill "Rock Shield" is 0.7.
[0143] During the game operation, combat log data is collected in real time and stored in the database. The size of the time sliding window is set to 10 seconds, and the combat log data is divided into multiple time windows. Within each time window, the skill interaction frequency in different combat scenario types is counted. For example, in the melee scenario, the interaction frequency of the "Flame Strike" and "Heavy Strike" skills is counted as 5 times / 10 seconds; in the ranged scenario, the interaction frequency of the "Ice Arrow" and "Fireball" skills is counted as 3 times / 10 seconds. The loss function is calculated based on the skill interaction frequency, and the edge weight parameters of the dynamic attribute relationship graph are updated using the meta-gradient descent algorithm. For example, if the interaction frequency of the "Flame Strike" and "Rock Shield" skills is high and a strong restraint relationship is shown in multiple time windows, then increase the value of the corresponding element in their edge weight matrix.
[0144] In the game, when the player releases a new skill combination, such as the combination of "Flame Strike" and "Thunderbolt", the system detects this new skill combination. The "Thunderbolt" skill is added as a new node to the dynamic attribute relationship graph, and its connection relationship with other nodes is initialized. For example, the initial connection weight between "Thunderbolt" and "Flame Strike" is set to 0.2, indicating a certain potential restraint relationship between them.
[0145] In this embodiment, through the spatio-temporal graph attention network and the bidirectional message passing mechanism, node features can be more comprehensively extracted and neighborhood information can be aggregated, and dynamic reasoning is performed by combining the gating weight, improving the accuracy of skill restraint relationship reasoning. Based on the time sliding window mechanism and the meta-gradient descent algorithm, the edge weight matrix can be updated in real time according to the combat log data, enabling the edge weight matrix to timely reflect the time-varying characteristics of the skill restraint relationship. When a new skill combination is detected, the nodes of the dynamic attribute relationship graph are automatically expanded and the connection relationship is initialized, enabling the system to quickly adapt to new game situations. After introducing a new skill, the system can learn the restraint relationship between the new skill and other skills in a short time, providing more effective combat suggestions for players.
[0146] In some embodiments, in the above step S104, when the dynamic attribute relationship graph and the player's historical operation data are jointly input into the reinforcement learning decision model to generate a battle strategy tree, the search depth of the battle strategy tree is dynamically adjusted through the real-time calculated battle complexity index to generate an optimal battle strategy, which specifically includes:
[0147] Fuse the dynamic attribute relationship graph and the player's historical operation data into a unified state representation;
[0148] Input the unified state representation into the reinforcement learning decision model. In each decision cycle of the reinforcement learning decision model, use the improved Monte Carlo tree search algorithm to construct an initial battle strategy tree;
[0149] Determine the expected rewards of each path of the initial battle strategy tree through the node value evaluation function to generate a battle strategy tree;
[0150] According to the dynamically calculated battle complexity index, use the maximum search depth adjustment formula to adjust the maximum search depth of the battle strategy tree;
[0151] Based on the battle strategy tree, fuse the immediate search value of the reinforcement learning decision model and the predicted value of the deep Q network to output an optimal battle strategy;
[0152] When it is detected that the reward value of any branch of the battle strategy tree is lower than the preset reward threshold, trigger the pruning operation on this branch.
[0153] In this embodiment, traditional decision-making methods often rely on fixed rules or simple decision trees and are difficult to adapt to complex and changing battle environments. Reinforcement learning technology provides a new idea for solving this problem. However, existing reinforcement learning methods have problems such as low decision-making efficiency and poor strategy adaptability when dealing with dynamically changing battle scenarios and complex skill restraint relationships. Therefore, a method that can comprehensively consider dynamic attribute relationships and the player's historical operation data and dynamically adjust the decision-making strategy according to the battle complexity is needed.
[0154] Specifically, collect the player's historical operation data, such as the skill release order, movement trajectory, etc. Perform preprocessing on the collected data, including data cleaning, normalization, and other operations. Use a feature fusion algorithm to fuse the dynamic attribute relationship graph and the player's historical operation data into a unified state representation. For example, a neural network can be used to map the two types of data to the same feature space and fuse them through weighted summation or concatenation.
[0155] Build a reinforcement learning decision-making model, which takes a unified state representation as input and outputs the probability distribution of each possible action. The reinforcement learning decision-making model can be implemented using a deep Q-network (DQN), a policy gradient algorithm, etc. In each decision cycle of the reinforcement learning decision-making model, an improved Monte Carlo tree search algorithm is used to construct an initial battle strategy tree. The improved Monte Carlo tree search algorithm combines the prediction value of the reinforcement learning decision-making model on the basis of the traditional Monte Carlo tree search algorithm to optimize the selection and expansion of nodes. For example, when selecting a node, not only the number of visits and average reward of the node are considered, but also the predicted value of the reinforcement learning decision-making model for this node is considered.
[0156] Design a node value evaluation function. Specifically, the node value evaluation function satisfies
[0157]
[0158] where V(v) represents the expected reward of each path of the original battle strategy tree, Q(v) represents the cumulative battle reward obtained by the v-th node in the historical search, N(v) represents the number of times the v-th node is searched in the Monte Carlo tree search, c represents an exploration coefficient used to control the exploration tendency, N(parent(v)) represents the total number of times the parent node of the v-th node is visited, μ represents a complexity adjustment factor, C t represents the real-time battle complexity index at time t, ‖B t ‖2 represents the L2 norm of the BOSS state vector at time t, for example, the comprehensive strength of the state vector such as BOSS health, attack power, rage state flag, etc. represents the real-time restraint strength between the i-th player character skill and the j-th BOSS skill at time t, represents the cumulative value of the restraint relationships between all player character skills and all BOSS skills at time t, and α and β respectively represent ‖B t ‖2 and 's weight coefficients.
[0159] Use the node value evaluation function to evaluate each node of the initial battle strategy tree and determine the expected reward of each path. Generate a battle strategy tree according to the expected reward.
[0160] Adopt a maximum search depth adjustment formula to adjust the maximum search depth of the battle strategy tree according to the battle complexity index. The maximum search depth adjustment formula satisfies
[0161]
[0162] where d max represents the maximum search depth, d0 represents the base search depth, that is, the default search depth when the battle complexity is at a medium level, represents the low complexity threshold, below which it is regarded as a simple battle, represents the high complexity threshold, above which it is regarded as an extremely complex battle, Δd represents the depth adjustment amplitude, d min represents the lower limit of the search depth, represents the upper limit of the search depth, clip(·) represents the truncation function.
[0163] For example, when C t increases (such as when the BOSS uses a powerful move), the formula automatically increases d max , allowing the policy tree to explore deeper branches to handle complex situations and comprehensively evaluate skill combinations and movement positions. When C t the complexity decreases, quickly generate a strategy and reduce the depth to save computing resources.
[0164] During the search process of the battle strategy tree, fuse the immediate search value of the reinforcement learning decision model and the predicted value of the deep Q-network to comprehensively evaluate each node. According to the evaluation results, select the optimal action path and output the optimal battle strategy. Real-time detect the reward values of each branch of the battle strategy tree. When it is detected that the reward value of any branch of the battle strategy tree is lower than the preset reward threshold, trigger the pruning operation on that branch to reduce unnecessary searches and improve the decision-making efficiency.
[0165] Exemplarily, the dynamic attribute relationship graph contains the restraint relationships and attribute bonus information among 10 character skills, and the player historical operation data includes the skill release sequence and movement trajectory of the player in the past 5 minutes. Use a simple fully connected neural network to fuse the dynamic attribute relationship graph and the player historical operation data into a unified state representation. The input layer of the neural network contains 50 neurons, the output layer contains 20 neurons, and the hidden layer contains 30 neurons. Use the fused state representation to train a deep Q-network as the reinforcement learning decision model. The input layer of the deep Q-network has 20 neurons, the output layer has 5 neurons (corresponding to 5 possible actions), and the hidden layer contains 40 neurons. In each decision-making cycle, use the trained deep Q-network and the improved Monte Carlo tree search algorithm to construct the initial battle strategy tree. The initial maximum search depth is set to 10. During the search process of the battle strategy tree, fuse the immediate search value of the reinforcement learning decision model and the predicted value of the deep Q-network, and select the optimal action path. After the search, output the optimal battle strategy, such as "release skill A first, then move to position B, and then release skill C". Real-time detect the reward values of each branch of the battle strategy tree. Suppose the reward value of a certain branch is lower than the preset reward threshold of -0.5, then trigger the pruning operation on that branch.
[0166] In this embodiment, by fusing the dynamic attribute relationship graph and the player's historical operation data into a unified state representation and inputting it into the reinforcement learning decision model, various factors in the game can be considered more comprehensively, improving the accuracy and rationality of decision-making. Dynamically adjusting the search depth of the battle strategy tree according to the dynamically calculated battle complexity index can improve the search efficiency while ensuring the quality of decision-making. Through branch pruning operations, unnecessary branches can be deleted in a timely manner, reducing the scale of the strategy tree and improving the decision-making efficiency.
[0167] Referring to Figure 2 , an embodiment of the present invention provides a real-time analysis system 2 for a player character to fight against a BOSS. The system 2 specifically includes:
[0168] The first real-time analysis module 201 is used to synchronously collect the screen pixel stream and game engine API data in real time, and perform cross-verification on the screen pixel stream and game engine API data using the inter-frame difference verification algorithm to obtain verified battle data;
[0169] The second real-time analysis module 202 is used to take the verified battle data as input, identify dynamic battlefield obstacles through an improved YOLOv5 model, and generate environmental perception data using a three-dimensional space interpolation algorithm;
[0170] The third real-time analysis module 203 is used to construct a dynamic attribute relationship graph based on the environmental perception data and pre-set character status data, and model the time-varying restraint relationship between character skills using a bidirectional graph attention network, and update the edge weight matrix through a time sliding window mechanism;
[0171] The fourth real-time analysis module 204 is used to jointly input the dynamic attribute relationship graph and the player's historical operation data into the reinforcement learning decision model to generate a battle strategy tree, dynamically adjust the search depth of the battle strategy tree according to the real-time calculated battle complexity index, and generate an optimal battle strategy;
[0172] The fifth real-time analysis module 205 is used to online update the decision model parameters according to the strategy execution effect data of the battle strategy tree through the meta-gradient descent algorithm.
[0173] It can be understood that the content in the embodiment of the real-time analysis method for a player character to fight against a BOSS as Figure 1 shown is applicable to the embodiment of this real-time analysis system for a player character to fight against a BOSS. The functions specifically implemented in the embodiment of this real-time analysis system for a player character to fight against a BOSS are the same as those in the embodiment of the real-time analysis method for a player character to fight against a BOSS as Figure 1 shown, and the beneficial effects achieved are also the same as those achieved in the embodiment of the real-time analysis method for a player character to fight against a BOSS as Figure 1 shown.
[0174] It should be noted that for the content such as information interaction and execution process among the above systems, since they are based on the same concept as the method embodiments of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0175] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and are not described herein again.
[0176] Referring to Figure 3 , an embodiment of the present invention further provides a computer device 3, including: a memory 302, a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the real-time analysis method for the battle between a player character and a BOSS as described in any one of the above methods.
[0177] The computer device 3 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art can understand that Figure 3 merely an example of the computer device 3, which does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0178] The so-called processor 301 may be a Central Processing Unit (CPU), and this processor 301 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0179] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In other embodiments, the memory 302 may also be an external storage device of the computer device 3, such as a plug-in hard disk equipped on the computer device 3, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 302 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 302 is used to store an operating system, application programs, a Boot Loader, data, and other programs, such as the program code of the computer program, etc. The memory 302 may also be used to temporarily store data that has been output or is to be output.
[0180] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the real-time analysis method for the confrontation between a player character and a BOSS as described in any one of the above methods.
[0181] In this embodiment, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code to the photographing device / terminal device. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0182] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0183] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0184] In the embodiments disclosed in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical, mechanical or other forms.
[0185] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
Claims
1. A real-time analysis method for a player character to fight against a BOSS, characterized in that The method specifically includes: Real-time synchronously collect the screen pixel stream and game engine API data, and use the inter-frame difference verification algorithm to cross-verify the screen pixel stream and game engine API data to obtain verified combat data; Use the verified combat data as input, identify dynamic battlefield obstacles through an improved YOLOv5 model, and generate environmental perception data using a three-dimensional spatial interpolation algorithm; Construct a dynamic attribute relationship graph based on the environmental perception data and pre-set character state data, and use a bidirectional graph attention network to model the time-varying restraint relationship between character skills, and update the edge weight matrix through a time sliding window mechanism; Jointly input the dynamic attribute relationship graph and the player's historical operation data into the reinforcement learning decision model to generate a battle strategy tree, and dynamically adjust the search depth of the battle strategy tree through the real-time calculated battle complexity index to generate the optimal battle strategy; According to the strategy execution effect data of the battle strategy tree, online update the decision model parameters through the meta-gradient descent algorithm.
2. The method according to claim 1, wherein The real-time synchronous collection of the screen pixel stream and game engine API data specifically includes: Use an adaptive region segmentation algorithm to divide the game interface into multiple game regions, and the game regions include a skill release area, an environmental interaction area, and a status display area; Set different sampling frequencies for each game region, and use the corresponding sampling frequency to collect different screen pixel streams in each game region in real time; Construct a state-event mapping table, and convert the player character attribute vector, BOSS state vector, and environmental event vector output by the game engine into a continuous state vector according to the state-event mapping table to form game engine API data.
3. The method according to claim 2, wherein The use of the inter-frame difference verification algorithm to cross-verify the screen pixel stream and game engine API data specifically includes: Generate an alignment sequence by interpolating the parsing result of the screen pixel stream and the state vector of the game engine API data; Perform two-way difference verification on the screen pixel stream and game engine API data in the alignment sequence to calculate the modal consistency index; When the modal consistency index exceeds the preset deviation threshold, trigger an abnormal data repair process based on motion compensation.
4. The method according to claim 1, wherein The use of the verified combat data as input, identifying dynamic battlefield obstacles through an improved YOLOv5 model, and generating environmental perception data using a three-dimensional spatial interpolation algorithm specifically includes: Add a spatio-temporal convolution module after the YOLOv5 backbone network to form an improved YOLOv5 model, use the improved YOLOv5 model to detect dynamic obstacles in the verified combat data, and fuse the temporal features through the spatio-temporal convolution module to form an obstacle detection result; Based on the cross-modal attention mechanism, adaptively weight and fuse the game API state vector and pixel stream features in the verified combat data to generate a joint feature representation with semantic consistency; Use a three-dimensional adaptive kernel density estimation algorithm to construct a dynamically updated environmental heat map according to the obstacle detection result and the pre-set terrain resistance parameters to generate environmental perception data.
5. The method according to claim 1, wherein The construction of the dynamic attribute relationship graph based on the environmental perception data and the pre-set character state data specifically includes: Encode the skill attributes of the player character and the BOSS skill attributes into multi-dimensional feature vectors respectively, and generate skill embedding representations through self-supervised learning; Calculate the synergy effect coefficients of each skill combination in a specific battlefield environment based on the environmental perception data; Construct a dynamic attribute relationship graph according to the skill embedding representation and the synergy effect coefficient. The vertex set of the dynamic attribute relationship graph includes each character skill node, and the edge set of the dynamic attribute relationship graph includes each restraint relationship weight matrix based on the design values.
6. The method according to claim 5, wherein The time-varying restraint relationship between character skills is modeled by a bidirectional graph attention network, and the edge weight matrix is updated through a time sliding window mechanism, specifically including: Use a spatio-temporal graph attention network to extract the node features of the dynamic attribute relationship graph, aggregate the neighborhood information through a bidirectional message passing mechanism, and perform dynamic reasoning on the skill restraint relationship in combination with the gating weight, where the gating weight is jointly determined by the current combat phase index and the environmental complexity; Based on the collected real-time combat log data, use a time sliding window to count the skill interaction frequencies corresponding to different combat scenario types, and dynamically adjust the edge weight parameters of the dynamic attribute relationship graph in combination with the meta-gradient descent algorithm; When a new skill combination is detected, automatically expand the nodes of the dynamic attribute relationship graph and initialize the connection relationship.
7. The method according to claim 1, wherein Input the dynamic attribute relationship graph and the player's historical operation data into the reinforcement learning decision model together to generate a battle strategy tree, and dynamically adjust the search depth of the battle strategy tree through the dynamically calculated combat complexity index to generate an optimal battle strategy, specifically including: Fuse the dynamic attribute relationship graph and the player's historical operation data into a unified state representation; Input the unified state representation into the reinforcement learning decision model. In each decision cycle of the reinforcement learning decision model, use an improved Monte Carlo tree search algorithm to construct an initial battle strategy tree; Determine the expected rewards of each path of the initial battle strategy tree through a node value evaluation function to generate a battle strategy tree; According to the dynamically calculated combat complexity index, adjust the maximum search depth of the battle strategy tree using the maximum search depth adjustment formula; Based on the battle strategy tree, fuse the immediate search value of the reinforcement learning decision model and the predicted value of the deep Q network to output the optimal battle strategy; When it is detected that the reward value of any branch of the battle strategy tree is lower than the preset reward threshold, trigger the pruning operation on this branch.
8. A real-time analysis system for a player character to battle against a BOSS, characterized in that, The system specifically includes: The first real-time analysis module is used to synchronously collect the screen pixel stream and game engine API data in real time, and perform cross-verification on the screen pixel stream and game engine API data using the inter-frame difference verification algorithm to obtain verified combat data; The second real-time analysis module is used to take the verified combat data as input, identify dynamic battlefield obstacles through an improved YOLOv5 model, and generate environmental perception data using a three-dimensional space interpolation algorithm; The third real-time analysis module is used to construct a dynamic attribute relationship graph according to the environmental perception data and the preset character state data, and model the time-varying restraint relationship between character skills by a bidirectional graph attention network, and update the edge weight matrix through a time sliding window mechanism; The fourth real-time analysis module is used to jointly input the dynamic attribute relationship graph and the player's historical operation data into the reinforcement learning decision model to generate a battle strategy tree, dynamically adjust the search depth of the battle strategy tree through the battle complexity index calculated in real time, and generate an optimal battle strategy; The fifth real-time analysis module is used to online update the decision model parameters through the meta-gradient descent algorithm according to the strategy execution effect data of the battle strategy tree.
9. A computer device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory, which, when executed on the processor, implement the real-time analysis method for the player character to fight against the BOSS as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is run by the processor, it implements the real-time analysis method for the player character to fight against the BOSS as described in any one of claims 1 to 7.
Citation Information
Cited By
Floating population dynamic collaborative governance method based on five-member role conversion
CN120634822A
Game interaction configuration method and device, equipment and medium
CN121490399A