An Adaptive Operation Method for Complex Scenes of Mobile Games Based on Multi-Task Learning
Through the combination of multi-task learning and reinforcement learning, a multi-modal feature extraction network and task map are constructed, task priority and resource allocation are optimized, and the problem of insufficient multi-task coordination and multi-modal information fusion in the existing technology is solved, and the intelligence and adaptability of mobile game operations are improved.
Patent Information
- Application Number
- CN202411952950.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing deep learning and reinforcement learning methods are difficult to effectively coordinate multiple tasks and multimodal information in complex scenarios of mobile games, resulting in insufficient adaptability and resource allocation of the system in complex environments and the inability to dynamically adjust task priorities and operation strategies.
The multi-task learning framework is adopted, combining multi-modal data fusion, multi-task learning and reinforcement learning technology, and by building a multi-modal feature extraction network, task graph and dynamic weight allocation mechanism, optimizing operation strategies, generating task priorities, and using the Actor-Critic architecture for strategy optimization.
It improves the intelligence and adaptability of game operations, can reasonably schedule tasks and resources in complex scenarios, reduce conflicts, provide personalized operation plans, and improve players' immersion and operation comfort.
Smart Images

Figure CN119868963B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of complex scene operations in mobile games, and particularly to an adaptive operation method for complex scenes in mobile games based on multi-task learning. Background Art
[0002] With the rapid development of the mobile Internet, mobile games have become an important part of modern people's daily entertainment life. Especially today, with the continuous improvement of the hardware performance of smart phones, the types and complexity of mobile games have become increasingly diverse. Many mobile games not only rely on players' operating skills, but also involve complex game scenes, task allocation, real-time feedback, and strategy adjustment and other elements. In order to provide a more intelligent and personalized gaming experience, in recent years, more and more research has begun to explore how to use technologies such as artificial intelligence and deep learning to enhance the adaptive operation ability in games.
[0003] Traditional mobile game control methods often rely on players' intuitive operations. Complex scenes and tasks in games are usually processed by fixed rules and manually set control methods. For example, the actions of characters or objects in games are often pre-determined or achieved through simple rules. Although this method is simple and effective, for complex and changing game environments, its adaptability and flexibility are poor. In addition, as the complexity of mobile game scenes continues to increase, when players perform operations, they not only need to consider real-time visual information, but also need to make decisions according to constantly changing task requirements and game rules, which makes the traditional operation method gradually seem inadequate.
[0004] To solve this problem, some researchers have begun to try to introduce advanced technologies such as deep learning and reinforcement learning, and use big data analysis and intelligent algorithms to automatically optimize operation behaviors. Through training on a large amount of data, deep learning can, to a certain extent, imitate the human decision-making process and make autonomous adjustments according to changes in the environment. Reinforcement learning is particularly suitable for dealing with decision-making problems. By exploring the state space and action space, the strategy can be continuously optimized. However, existing applications of deep learning and reinforcement learning often have some limitations. First, many applications focus on the optimization of single tasks and are insufficient in coordinating and optimizing multiple tasks and multiple goals in complex scenarios. Second, existing methods usually rely on single-modal data for learning and decision-making. This method often ignores the synergistic effect of multi-modal information (such as visual information, operation records, and game rules, etc.) in game operations and is difficult to comprehensively improve the intelligence and adaptability of game operations.
[0005] At present, for the adaptive operation methods of mobile games based on deep learning and reinforcement learning, most research focuses on a specific task or a specific type of operation scenario. For example, some research uses image processing technology to analyze game scenes and performs automatic aiming or obstacle avoidance operations by detecting the positions of enemies or target objects. Other research focuses on training agents through reinforcement learning to optimize strategies in a single task or a relatively simple environment. Although these methods can improve the intelligence level of game operations to a certain extent, they often face the following problems: First, in complex mobile game scenes, there are not only multiple target tasks, but also dependencies and mutual constraints between tasks. Therefore, how to perform effective task optimization in a complex task dependency graph is an urgent problem to be solved. Second, existing technologies usually learn and predict only through a single modality (such as visual information or operation records), lacking effective fusion and optimization of multi-modal information, resulting in poor adaptability of the system in complex environments. Third, although reinforcement learning can help agents adjust strategies according to feedback, the resource allocation and task priority adjustment in multi-task scenarios of existing methods are still too simple to dynamically respond to real-time changes in complex scenarios.
[0006] In addition, most traditional deep learning and reinforcement learning methods adopt fixed loss functions and optimization strategies, lacking a dynamic adjustment mechanism for weights in different tasks and different scenarios. In complex mobile game scenes, there are often dependencies, competitions, and even mutual exclusion relationships between tasks. How to reasonably adjust the loss weights between tasks, optimize resource allocation, and dynamically adjust task priorities is the key to improving game operation performance and intelligence. However, existing technologies have not fully solved this problem. Most existing reinforcement learning methods ignore this complex relationship between tasks, resulting in the system being unable to dynamically adjust resource allocation and operation strategies according to the urgency, importance, or executability of tasks, and it is difficult to perform flexible and effective adaptive operations in complex multi-task environments.
[0007] Therefore, aiming at the limitations of current technologies, a new method is urgently needed. It can comprehensively consider various information sources such as the relationships between tasks, visual information, operation records, and game rules in the framework of multi-task learning, and optimize operation strategies and task execution in games. Especially when there are complex dependencies, mutual exclusions, and parallel relationships between tasks, how to dynamically adjust task priorities, resource allocation, and operation strategies through intelligent algorithms is the key to improving the intelligence level of mobile game operations. Summary of the Invention
[0008] An object of the present invention is to propose a method for adaptively operating complex scenarios of mobile games based on multi-task learning. The present invention combines multi-modal data fusion, multi-task learning, and reinforcement learning technologies to optimize the operation strategy in complex mobile game scenarios. By accurately extracting operation records, visual information, and game rules, a task graph is generated and the task priority is dynamically adjusted, improving the efficiency of task coordination and resource allocation.
[0009] A method for adaptively operating complex scenarios of mobile games based on multi-task learning according to an embodiment of the present invention includes the following steps:
[0010] S1. Collect data in complex scenarios of mobile games, extract operation records, visual information, and game rules, and preprocess the data to generate an original multi-modal data set;
[0011] S2. Based on the original multi-modal data set, construct a multi-modal feature extraction network, extract visual features, operation behavior features, and rule features, and generate a task graph through the spatial relationship in visual features, task constraints in rule features, and operation logic in operation behavior features;
[0012] S3. Enhance visual features, operation behavior features, and rule features, optimize the expression of operation behavior features using an improved contrast learning method, and optimize visual features and rule features respectively by combining multi-view fusion technology and embedded coding method, and perform feature fusion on the optimized features to generate multi-modal features;
[0013] S4. Take the multi-modal features and the task graph as inputs, construct a multi-task learning model, optimize task features, and adjust the task loss weight by combining a dynamic weight allocation mechanism to generate task priorities;
[0014] S5. Based on the task priorities, use a recursive optimization method to uniformly optimize the tasks with dependencies in the task graph according to the optimized task features to generate a global task optimization result;
[0015] S6. Take the global task optimization result as an input, combine the lighting parameters, obstacle parameters, and delay parameters in the simulation environment, and optimize the global task strategy and local scene strategy through a reinforcement learning model based on the Actor-Critic architecture.
[0016] Optionally, the S2 specifically includes:
[0017] S21. Construct a multi-modal feature extraction network, and the multi-modal extraction network consists of a bidirectional long short-term memory network, a ResNet34 network, and a multi-layer perceptron;
[0018] S22. Use the bidirectional long short-term memory network to extract operation records to generate operation behavior features:
[0019] F op = f LSTM (X op );
[0020] Among them, F op represents the operation behavior feature, X op represents the operation record, and f LSTM represents the bidirectional long short-term memory network;
[0021] S23. For visual information, use the ResNet34 network to extract the spatial features in the image and generate visual features:
[0022] F vis = f CNN (X vis );
[0023] Among them, F vis represents the visual feature, X vis represents the visual information, and f CNN represents the ResNet34 network;
[0024] S24. For the game rules, use a multi-layer perceptron to perform an embedded expression of the rule logic and generate rule features:
[0025] F rule = f MLP (X rule );
[0026] Among them, F rule represents the rule feature, X rule represents the game rules, and f MLP represents the multi-layer perceptron;
[0027] S25. Analyze the spatial relationship in the visual feature, the task constraint in the rule feature, and the operation logic in the operation behavior feature to generate a task graph, and the task graph describes the dependency relationship, parallel relationship, and mutual exclusion relationship between tasks.
[0028] Optionally, the S3 specifically includes:
[0029] S31. Use an improved contrast learning method to optimize the expression of the operation behavior feature;
[0030] S32. Use the multi-view fusion technology to optimize the visual feature, construct the visual feature of each view, and perform weighted fusion on the features of each view to generate the optimized visual feature:
[0031]
[0032] Among them, Represents the optimized visual feature, α m Represents the weight of the m-th perspective, Represents the visual feature of the m-th perspective, M represents the number of perspectives, and W represents the trainable parameters, Represents the visual feature of the j-th perspective, and exp represents the natural exponential function;
[0033] S33. Process the rule features, optimize the rule features using the embedded encoding method, and generate the optimized rule features;
[0034] S34. Perform feature fusion on the optimized features, calculate the importance weights of the features using the self-attention mechanism, and calculate the weight matrix of each feature:
[0035]
[0036] Among them, W i Represents the weight matrix of feature i, Q i and K i Represent the query and key mappings of the feature vector respectively, d represents the feature dimension, T represents the transpose operation, and softmax represents the normalization function;
[0037] S36. Generate multi-modal features:
[0038]
[0039] Among them, Represents the multi-modal feature, V i Represents the value mapping of the feature vector, and n represents the number of modalities.
[0040] Optionally, the S31 specifically includes:
[0041] S311. Based on the operation behavior feature F op , construct contrastive learning sample pairs, and the contrastive learning sample pairs include positive sample pairs and negative sample pairs And screen the hardest negative samples that are most similar to the positive sample features from the negative sample set to form a negative sample index set H(i);
[0042] S312. Use the main encoder and the momentum encoder to perform feature projection on the operation behavior feature F op , and generate a stable feature representation Z op :
[0043] θ′ = mθ′+(1 - m)θ;
[0044] Among them, θ' represents the momentum encoder weight, θ represents the main encoder weight, and m represents the momentum coefficient;
[0045] S313. Calculate the symmetric contrast loss for positive and negative sample pairs, including positive similarity optimization and negative similarity optimization: op Calculate the symmetric contrast loss for positive and negative sample pairs, including positive similarity optimization and negative similarity optimization:
[0046]
[0047] where L sym represents the symmetric contrast loss for positive and negative sample pairs, L op represents the unidirectional contrast loss, represents the positive sample feature representation generated by the momentum encoder, represents the reverse positive sample feature representation generated by the momentum encoder, represents the negative sample feature representation generated by the momentum encoder, sim represents the cosine similarity function, τ represents the temperature coefficient, H(i) represents the set of negative sample indices, and represent the feature representations generated by the momentum encoder;
[0048] S314. Perform multi-scale feature extraction on the operation behavior feature F op to generate operation behavior feature representations at different scales Calculate the symmetric contrast loss for each scale and sum them up:
[0049]
[0050] where L multi is the total multi-scale contrast learning loss, represents the operation behavior feature representation at the s-th scale, and P represents the total number of feature scales;
[0051] S315. Generate the optimized operation behavior feature by minimizing the total multi-scale contrast learning loss
[0052] Optionally, S4 specifically includes:
[0053] S41. Construct a multi-task learning model, input the multi-modal features into the shared feature extraction network to extract shared features. The multi-task learning model includes a shared feature extraction network and task branch networks. The shared feature extraction network is composed of multiple layers of convolutional neural networks, and the task branch networks are composed of shallow convolutional neural networks;
[0054] S42. According to the task nodes in the task graph, construct an independent task branch network for each task, and extract task features through the task branch network;
[0055] S43. According to the dependency, competition or mutual exclusion relationships between tasks in the task graph, adjust the task features through the relevance of the shared features:
[0056]
[0057] R ij = 1 - Mutex(t i , t j );
[0058] Among them, represents the task feature after shared feature adjustment, R ij represents the inhibition factor, W p , W s and W k respectively represent the weighted parameters of the parent task, shared feature, and child task, used to control the contribution of the feature to the adjustment, C i represents the set of child tasks, P i represents the set of parent tasks, F shared represents the shared feature, and respectively represent the specific feature of the parent task and the specific feature of the child task, Mutex(t i , t j ) represents the mutual exclusion relationship between tasks t i and t j ;
[0059] S43. Calculate the immediate utility for each task according to the task graph and the adjusted task features;
[0060] S44. Dynamically adjust the task loss weight according to the immediate utility:
[0061]
[0062] Among them, w i represents the loss weight, ∈ represents the smoothing factor to prevent the denominator from being zero, u i and u j respectively represent the immediate utilities of tasks t i and t j ;
[0063] S45. Extract the inter-task dependencies according to the task graph, construct a dependency chain for each task, recursively optimize the dependency chain from the parent task to the child task, and update the features of the nested model:
[0064]
[0065] Among them, F nested (t i ) represents the nested optimized feature of task t i , f nested represents the nested optimization function, and Indicate task characteristics;
[0066] S46. Generate task priorities based on immediate utility and inter-task dependencies.
[0067] Optionally, the S5 specifically includes:
[0068] S51. According to the task graph, extract the set of tasks that have a mutually exclusive relationship with the task, introduce mutual exclusion constraints on the nested characteristics of the task, and adjust the task characteristics through the mutual exclusion optimization function:
[0069]
[0070] Among them, f mutex represents the mutual exclusion optimization function, F nested (t i ) represents the optimized characteristics of task t i , F nested (t k ) represents the optimized characteristics of mutually exclusive tasks, sim represents the cosine similarity function, λ represents the weight coefficient of the mutual exclusion constraint, used to control the intensity of the penalty term, and Mutex(t i ) represents the set of mutually exclusive tasks;
[0071] S52. According to the task graph, extract the set of tasks that can be executed in parallel, based on the nested characteristics of the parallel task set, construct a global consistency optimization objective, jointly optimize the characteristics of the parallel task set, and update the task characteristics:
[0072]
[0073] Among them, L parallel represents the global consistency optimization objective function, Parallel(t i ) represents the set of parallel tasks, F nested (t k ) represents the optimized characteristics of mutually exclusive tasks, sim represents the cosine similarity function, F shared represents the shared feature, indicating the information shared between tasks;
[0074] S53. Calculate the resource requirements of the task according to the task graph and task priorities, dynamically adjust the resource allocation weights according to the task priorities and resource requirements, and generate the global task optimization result.
[0075] Optionally, the S6 specifically includes:
[0076] S61. Construct a reinforcement learning model based on the Actor-Critic architecture, take the global task optimization result and environmental parameters as inputs, and generate unified state features:
[0077] H ti= ReLU(W t ·F nested (t i ) + b t );
[0078]
[0079] H env = ReLU(W e ·E env + b e );
[0080]
[0081] where H ti represents the task features after embedding, F nested (t i ) represents the nested optimization features of task t i , H global represents the global task features, k represents the number of tasks, E env represents the set of environmental parameters, H env represents the global environmental features, S represents the unified state features, W t , W e and W f represent the weights of task embedding, environmental embedding, and feature fusion respectively, b t , b e and b f represent the corresponding biases respectively; μ and σ represent the normalization parameters;
[0082] S62. Output the policy through the Actor network:
[0083] π(a|s) = softmax(W actor ·S + b actor );
[0084] where π(a|s) represents the output policy, W actor and b actor represent the weights and biases of the Actor network respectively;
[0085] S63. Generate the state value function through the Critic network:
[0086] V(s) = W critic ·S + b critic ;
[0087] where V(s) represents the state value function, W critic and b critic represent the weights and biases of the Critic network respectively;
[0088] S64. Optimize the global task strategy, calculate the advantage function through immediate rewards and cumulative returns, update the Actor network parameters, and update the Critic network parameters by minimizing the Critic network loss function:
[0089]
[0090] A(s,a) = R t -V(s);
[0091]
[0092] where R t represents the cumulative return, γ represents the discount factor, r t represents the immediate reward, A(s,a) represents the global advantage function, α represents the learning rate, θ actor represents the Actor network parameters, represents the gradient of the parameter θ actor L critic represents the Critic network loss function, θ critic represents the Critic network parameters, represents the gradient of the parameter θ critic ;
[0093] S65. Extract the local state features of the environmental parameters, and generate the local scene policy by optimizing the local policy loss function:
[0094]
[0095] π local (a|s) = softmax(W actor-local ·S local +b actor-local );
[0096] L local = -log2π local (a|s)·A local (s,a);
[0097] where H L represents the illumination feature, H O represents the overall obstacle feature, H D represents the latency feature, μ and σ represent the normalization parameters, S local represents the local state feature, W f and b f respectively represent the weights and biases of the fusion network, [H L ,H O ,H D represents concatenating H L ,H O and HD Perform splicing, L local Represents the local policy loss function, A local (s,a) represents the local advantage function, π local (a|s) represents the local scenario policy, W actor-local And b actor-local Respectively represent the weights and biases of the Actor network under the environmental parameters;
[0098] S66. Output the optimized global task policy and local scenario policy.
[0099] The beneficial effects of the present invention are:
[0100] First of all, the present invention deeply mines and optimizes data of different modalities by constructing a multi-modal feature extraction network and combining an improved contrast learning method and a multi-perspective fusion technology. This method can not only effectively improve the expression ability of operation behavior features, visual features and rule features, but also better capture the details and complexity in the game scene. By fusing multi-modal information, the system can more comprehensively understand the game environment, and then make more reasonable decisions. This intelligent fusion of multi-modal data greatly improves the accuracy of game operations, can handle different types and complexities of game tasks, enables the system to adaptively adjust operation strategies in various complex scenarios, and thus optimizes the player's game experience.
[0101] Secondly, the present invention innovatively introduces a multi-task learning framework based on a task graph, and optimizes the processing of task features through a dynamic weight allocation mechanism. The task graph can accurately describe the task relationships in the game, including dependencies, mutual exclusions and parallel relationships between tasks. Through this mechanism, the system can not only reasonably allocate resources, but also optimize according to the priorities, urgencies and resource requirements of tasks. This means that the system can reasonably schedule between different tasks to ensure that key tasks can be processed first and will not be interfered by the execution of low-priority tasks. Dynamically adjusting task priorities and resource allocations makes operations in multi-task scenarios more efficient, reduces task conflicts and resource waste, and thus improves the overall working efficiency of the system.
[0102] In addition, the present invention adopts a reinforcement learning model based on the Actor-Critic architecture, which can not only optimize the global task policy, but also optimize the policy for environmental factors such as lighting, obstacles, and latency in the local scenario. Through continuous learning and adjustment, the system can flexibly respond to changes in the current environment to ensure efficient and accurate operations in complex and changing game scenes. Different from traditional static rules or predetermined strategies, the reinforcement learning model can adjust the strategy in real time according to the player's behavior and environmental changes to achieve true adaptive operations.
[0103] Finally, the present invention can greatly improve the intelligence and adaptability of game operations in practical applications, not only optimizing the execution efficiency of game tasks, but also enhancing the interactive experience of players. For example, in the scenario of parallel execution of multiple tasks, the system can reasonably schedule operation tasks to avoid a decrease in operation efficiency caused by conflicts between tasks; in a dynamically changing environment, the system can adjust operation strategies in real time according to environmental changes to ensure the smooth execution of tasks. More importantly, the system can gradually optimize the task execution strategy based on the player's behavior and feedback from the game scenario, providing a personalized operation plan, thereby enhancing the player's sense of immersion and the comfort of operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0105] Figure 1 is a flowchart of a method for adaptively operating complex scenarios of mobile games based on multi-task learning proposed by the present invention;
[0106] Figure 2 is a structural diagram of a multi-modal feature extraction network of a method for adaptively operating complex scenarios of mobile games based on multi-task learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0107] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0108] Referring to Figure 1 and Figure 2 , a method for adaptively operating complex scenarios of mobile games based on multi-task learning includes the following steps:
[0109] S1. Collect data in the complex scenarios of mobile games, extract operation records, visual information, and game rules, and preprocess the data to generate an original multi-modal data set;
[0110] S2. Based on the original multi-modal data set, construct a multi-modal feature extraction network, extract visual features, operation behavior features, and rule features, and generate a task graph through the spatial relationship in the visual features, task constraints in the rule features, and operation logic in the operation behavior features;
[0111] S3. Enhance the visual features, operation behavior features, and rule features. Optimize the expression of operation behavior features using an improved contrastive learning method, and optimize the visual features and rule features using multi-view fusion technology and embedded coding method respectively. Then, fuse the optimized features to generate multi-modal features.
[0112] S4. Take the multi-modal features and task graph as inputs, construct a multi-task learning model to optimize the task features, and adjust the task loss weights by combining a dynamic weight allocation mechanism to generate task priorities.
[0113] S5. Based on the task priorities, use a recursive optimization method to uniformly optimize the tasks with dependencies in the task graph according to the optimized task features, and generate the global task optimization results.
[0114] S6. Take the global task optimization results as inputs, and optimize the global task policy and local scene policy through a reinforcement learning model based on the Actor-Critic architecture by combining the lighting parameters, obstacle parameters, and delay parameters in the simulation environment.
[0115] In this embodiment, the S2 specifically includes:
[0116] S21. Construct a multi-modal feature extraction network, which consists of a bidirectional long short-term memory network, a ResNet34 network, and a multi-layer perceptron.
[0117] S22. Use the bidirectional long short-term memory network to extract the operation records and generate operation behavior features:
[0118] F op = f LSTM (X op );
[0119] Where F op represents the operation behavior features, X op represents the operation records, and f LSTM represents the bidirectional long short-term memory network.
[0120] S23. For the visual information, use the ResNet34 network to extract the spatial features in the image and generate visual features:
[0121] F vis = f CNN (X vis );
[0122] Where F vis represents the visual features, X vis represents the visual information, and f CNN represents the ResNet34 network.
[0123] S24. For the game rules, the rule logic is embedded and expressed through a multi-layer perceptron to generate rule features:
[0124] F rule = f MLP (X rule );
[0125] Among them, F rule represents the rule feature, X rule represents the game rule, and f MLP represents the multi-layer perceptron;
[0126] S25. Analyze the spatial relationship in the visual features, the task constraints in the rule features, and the operation logic in the operation behavior features to generate a task graph, which describes the dependency relationship, parallel relationship, and mutual exclusion relationship between tasks.
[0127] In this embodiment, the specific content of S3 includes:
[0128] S31. Optimize the expression of operation behavior features by using an improved contrastive learning method;
[0129] S32. Optimize the visual features by using a multi-view fusion technology, construct the visual features of each view, and perform weighted fusion on the features of each view to generate optimized visual features:
[0130]
[0131] Among them, represents the optimized visual feature, α m represents the weight of the m-th view, represents the visual feature of the m-th view, M represents the number of views, W represents the trainable parameter, represents the visual feature of the j-th view, and exp represents the natural exponential function;
[0132] S33. Process the rule features, optimize the rule features by using an embedded encoding method, and generate optimized rule features;
[0133] S34. Perform feature fusion on the optimized features, calculate the importance weights of the features by using a self-attention mechanism, and calculate the weight matrix of each feature:
[0134]
[0135] Among them, W i represents the weight matrix of feature i, Q i and K i respectively represent the query and key mappings of the feature vector, d represents the feature dimension, T represents the transpose operation, and softmax represents the normalization function;
[0136] S36. Generate multimodal features:
[0137]
[0138] Among them, represents multimodal features, V i represents the value mapping of the feature vector, and n represents the number of modalities.
[0139] In this embodiment, S31 specifically includes:
[0140] S311. Based on the operation behavior feature F op , construct a contrastive learning sample pair, and the contrastive learning sample pair includes a positive sample pair and a negative sample pair and screen the hardest negative samples most similar to the positive sample features from the negative sample set to form a negative sample index set H(i);
[0141] S312. Use the main encoder and the momentum encoder to perform feature projection on the operation behavior feature F op to generate a stable feature representation Z op by optimizing the weights of the momentum encoder:
[0142] θ′ = mθ′+(1 - m)θ;
[0143] Among them, θ' represents the weights of the momentum encoder, θ represents the weights of the main encoder, and m represents the momentum coefficient;
[0144] S313. Calculate the symmetric contrastive loss of the positive and negative sample pairs according to the stable feature representation Z op , including forward similarity optimization and reverse similarity optimization:
[0145]
[0146] Among them, L sym represents the symmetric contrastive loss of the positive and negative sample pairs, L op represents the unidirectional contrastive loss, represents the positive sample feature representation generated by the momentum encoder, represents the reverse positive sample feature representation generated by the momentum encoder, represents the negative sample feature representation generated by the momentum encoder, sim represents the cosine similarity function, τ represents the temperature coefficient, H(i) represents the negative sample index set, and represent the feature representations generated by the momentum encoder;
[0147] S314. For the operation behavior feature F opPerform multi-scale feature extraction to generate operation behavior feature representations at different scales Calculate the symmetric contrast losses at each scale and sum them up:
[0148]
[0149] where, L multi is the total multi-scale contrast learning loss, represents the operation behavior feature representation at the s-th scale, and P represents the total number of feature scales;
[0150] S315. Generate optimized operation behavior features by minimizing the total multi-scale contrast learning loss
[0151] In this embodiment, the S4 specifically includes:
[0152] S41. Construct a multi-task learning model, input multi-modal features into the shared feature extraction network to extract shared features. The multi-task learning model includes a shared feature extraction network and task branch networks. The shared feature extraction network is composed of multiple layers of convolutional neural networks, and the task branch networks are composed of shallow convolutional neural networks;
[0153] S42. According to the task nodes in the task graph, construct independent task branch networks for each task, and extract task features through the task branch networks;
[0154] S43. According to the dependencies, competitions or mutual exclusions between tasks in the task graph, adjust the task features through the relevance of the shared features:
[0155]
[0156] R ij = 1 - Mutex(t i , t j );
[0157] where, represents the task features adjusted by the shared features, R ij represents the inhibition factor, W p , W s and W k respectively represent the weighted parameters of the parent task, shared features and sub-tasks, used to control the contribution of the features to the adjustment, C i represents the set of sub-tasks, P i represents the set of parent tasks, F shared represents the shared features, and respectively represent the specific features of the parent task and the specific features of the sub-task, Mutex(t i , tj ) represents task t i and t j 's mutual exclusion relationship;
[0158] S43. Calculate the immediate utility for each task according to the task graph and the adjusted task characteristics;
[0159] S44. Dynamically adjust the task loss weight according to the immediate utility:
[0160]
[0161] where w i represents the loss weight, ∈ represents a smoothing factor to prevent the denominator from being zero, u i and u j respectively represent the immediate utilities of task t i and t j ;
[0162] S45. Extract the inter-task dependency relationship from the task graph, construct a dependency chain for each task, recursively optimize the dependency chain from the parent task to the child task, and update the features of the nested model:
[0163]
[0164] where F nested (t i ) represents the nested optimization feature of task t i , f nested represents the nested optimization function, and F ti represent task features;
[0165] S46. Generate task priorities according to the immediate utility and the inter-task dependency relationship.
[0166] In this embodiment, the S5 specifically includes:
[0167] S51. According to the task graph, extract the task set that has a mutual exclusion relationship with the task, introduce a mutual exclusion constraint on the nested feature of the task, and adjust the task feature through the mutual exclusion optimization function:
[0168]
[0169] where f mutex represents the mutual exclusion optimization function, F nested (t i ) represents the optimized feature of task t i , F nested (t k ) represents the optimized feature of the mutually exclusive task, sim represents the cosine similarity function, λ represents the weight coefficient of the mutual exclusion constraint, used to control the intensity of the penalty term, Mutex(ti ) represents a set of mutually exclusive tasks;
[0170] S52. Extract the set of tasks that can be executed in parallel according to the task graph. Based on the nested characteristics of the parallel task set, construct a global consistency optimization objective, jointly optimize the characteristics of the parallel task set, and update the task characteristics:
[0171]
[0172] where L parallel represents the global consistency optimization objective function, Parallel(t i ) represents the parallel task set, F nested (t k ) represents the optimized characteristics of the mutually exclusive tasks, sim represents the cosine similarity function, F shared represents the shared characteristics, indicating the information shared between tasks;
[0173] S53. Calculate the resource requirements of the tasks according to the task graph and task priorities, and dynamically adjust the resource allocation weights according to the task priorities and resource requirements to generate the global task optimization result.
[0174] In this embodiment, the S6 specifically includes:
[0175] S61. Construct a reinforcement learning model based on the Actor-Critic architecture, and use the global task optimization result and environmental parameters as inputs to generate unified state characteristics:
[0176] H ti = ReLU(W t ·F nested (t i ) + b t );
[0177]
[0178] H env = ReLU(W e ·E env + b e );
[0179]
[0180] where H ti represents the embedded task characteristics, F nested (t i ) represents the nested optimization characteristics of task t i , H global represents the global task characteristics, k represents the number of tasks, E env represents the set of environmental parameters, Henv denotes the global environmental feature, S denotes the unified state feature, W t 、W e and W f respectively denote the weights of task embedding, environmental embedding, and feature fusion, b t 、b e and b f respectively denote the corresponding biases, μ and σ denote the normalization parameters;
[0181] S62. Output the policy through the Actor network:
[0182] π(a|s) = softmax(W actor ·S + b actor );
[0183] where π(a|s) denotes the output policy, W actor and b actor respectively denote the weights and biases of the Actor network;
[0184] S63. Generate the state value function through the Critic network:
[0185] V(s) = W critic ·S + b critic ;
[0186] where V(s) denotes the state value function, W critic and b critic respectively denote the weights and biases of the Critic network;
[0187] S64. Optimize the global task policy, calculate the advantage function through the immediate reward and cumulative return, update the Actor network parameters, and update the Critic network parameters by minimizing the Critic network loss function:
[0188]
[0189] A(s,a) = R t - V(s);
[0190]
[0191] where R t denotes the cumulative return, γ denotes the discount factor, r t denotes the immediate reward, A(s,a) denotes the global advantage function, α denotes the learning rate, θ actor denotes the Actor network parameters, denotes the gradient of the parameter θ actor ,L critic denotes the Critic network loss function, θ criticDenote the Critic network parameters, Denote the parameter θ critic gradient;
[0192] S65. Extract the local state features of the environmental parameters, and generate the local scene policy by optimizing the local policy loss function:
[0193]
[0194] π local (a|s) = softmax(W actor-local ·S local +b actor-local );
[0195] L local = -log2π local (a|s)·A local (s,a);
[0196] Among them, H L denotes the illumination feature, H O denotes the overall obstacle feature, H D denotes the latency feature, μ and σ denote the normalization parameters, S local denotes the local state feature, W f and b f respectively denote the weights and biases of the fusion network, [H L ,H O ,H D denotes concatenating H L , H O and H D , L local denotes the local policy loss function, A local (s,a) denotes the local advantage function, π local (a|s) denotes the local scene policy, W actor-local and b actor-local respectively denote the weights and biases of the Actor network under the environmental parameters;
[0197] S66. Output the optimized global task policy and local scene policy.
[0198] Example 1:
[0199] In modern mobile games, complex game scenarios and diverse operation requirements place high demands on players' operation skills, reaction speed, decision-making ability, and adaptability to scene changes. Especially in some action games or strategy games, players need to make real-time decisions in a constantly changing environment and quickly execute corresponding operations. This not only requires players to have excellent operation skills but also requires the game system to intelligently adjust its operation prompts, game rules, and even scene elements to provide a more personalized and intelligent gaming experience.
[0200] Against this backdrop, a method for adaptive operation in complex scenarios of mobile games based on multi-task learning has emerged. In this embodiment, by combining multi-task learning and reinforcement learning, it can adaptively adjust the operation logic and strategies of the game according to the player's performance in the game and the dynamic changes in the game environment, thereby providing a smoother and more efficient gaming experience for players. Taking a typical real-time strategy game as an example, in this game, players need to control multiple units and perform tasks in a complex scenario. The game scenario includes multiple dynamically changing elements, such as obstacles, enemies, resource points, etc. Players' operations not only need to adjust tactics according to these changes but also need to reasonably allocate resources among multiple tasks and execute different operations.
[0201] In such a complex scenario, the game system needs to possess task allocation and resource scheduling, scene adaptive operation, and multi-modal feature extraction and processing. Specifically, the system first extracts features of the player's operation behavior, visual information in the game scene, and game rules through a multi-modal feature extraction network. Then, the system constructs a task graph and optimizes the tasks through a multi-task learning model, and finally adjusts the priority of the tasks based on the optimized task results. Through the reinforcement learning model with the Actor-Critic architecture, combined with scene parameters such as lighting, obstacles, and latency, the system continuously optimizes the global task strategy and local scene strategy, and finally provides the best operation suggestions for players.
[0202] To verify the effectiveness of the present invention, multiple experiments were conducted to evaluate the impact of the multi-task learning model and the reinforcement learning strategy optimization method on game performance and player experience in complex scenarios.
[0203] Table 1 Experimental comparison data of the traditional method and the method of the present invention
[0204] Index Traditional method Method of the present invention Task completion time (seconds) 40.5 32.8 Operation fluency score (out of 10) 7.5 9.3 Operation intelligence score (out of 10) 6.3 9.0 Overall satisfaction score (out of 10) 7.1 9.2 Average task completion accuracy 72.4 92.1 Average resource utilization rate 68.9 84.3 Stability (number of times without lags) 3 0
[0205] The comparison in terms of task completion time shows that the method of the present invention significantly shortens the time required for players to complete tasks. Specifically, players using the traditional method on average need 40.5 seconds to complete the task, while players using the method of the present invention only need 32.8 seconds, saving nearly 19% of the time. This difference indicates that through the optimization of multi-task learning and reinforcement learning, the game system can more intelligently adjust the execution order of tasks and resource allocation, thus improving the efficiency of task completion.
[0206] In terms of the operation fluency score, players using the method of the present invention received a high score of 9.3. In contrast, the score of the traditional method was 7.5. This gap reflects the optimization of the present invention method for the fluency and smooth experience of players during the game operation. By real-time adjusting the operation logic and game environment, the present invention can provide players with a more smooth control experience, making the operation of players in complex scenarios more natural and comfortable.
[0207] The operation intelligence score further proves the advantage of the method of the present invention in intelligent operation. The operation intelligence score of the traditional method was 6.3, while the score of the method of the present invention was 9.0, and the gap was very obvious. This difference reflects that the present invention realizes in-depth analysis and understanding of players' behaviors, game scenarios and rules through the combination of multi-modal feature extraction and multi-task learning, enabling the game system to make more intelligent adjustments and optimizations according to environmental changes and players' needs.
[0208] In terms of the overall satisfaction score, the method of the present invention also significantly improves players' game satisfaction, with a score of 9.2, while the traditional method was 7.1. This shows that the present invention not only brings improvement in operation efficiency, but also has a positive impact on the overall perception of players' experience. Players are more willing to accept the intelligent adjustment and optimization operation of the system during the game process. This improvement is not only reflected in the game process, but more in players' evaluation of the overall game experience.
[0209] In terms of the average task completion accuracy, players using the method of the present invention can complete tasks more accurately, and the task completion accuracy reaches 92.1%, which is about 20% higher than 72.4% of the traditional method. This improvement indicates that through the optimization of task features and the intelligent adjustment of task priorities by the multi-task learning model, players can be more focused and reduce invalid operations when performing tasks, thus improving the accuracy and efficiency of task execution.
[0210] The improvement in the average resource utilization rate further demonstrates the advantages of the method of the present invention in resource allocation and optimization in games. The resource utilization rate of the traditional method is 68.9%, while the resource utilization rate of players using the method of the present invention has increased to 84.3%. This increase indicates that through the strategy of combining multi-task learning and reinforcement learning, the game can schedule and allocate resources more effectively, helping players complete more tasks efficiently under limited resource conditions, thereby improving the overall execution efficiency of the game.
[0211] Finally, the comparison in terms of stability also further shows the advantages of the method of the present invention. Under the traditional method, there were 3 freezes during the game, while the method of the present invention achieved 0 freezes. This means that the present invention can maintain the smoothness of game operations while avoiding the common performance bottleneck problems in the traditional method, ensuring the continuity and stability of the game experience.
[0212] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. An adaptive operation method for complex scenarios of mobile games based on multi-task learning, characterized in that, It includes the following steps: S1. Collect data in complex scenarios of mobile games, extract operation records, visual information, and game rules, and preprocess the data to generate an original multimodal dataset; S2. Based on the original multimodal dataset, construct a multimodal feature extraction network to extract visual features, operation behavior features, and rule features, and generate a task graph through the spatial relationship in visual features, task constraints in rule features, and operation logic in operation behavior features; S3. Enhance visual features, operation behavior features, and rule features, optimize the expression of operation behavior features using an improved contrastive learning method, and optimize visual features and rule features using multi-view fusion technology and embedded coding methods respectively. Feature fusion is performed on the optimized features to generate multimodal features; S4. Use the multimodal features and the task graph as inputs to construct a multi-task learning model, optimize task features, and adjust task loss weights in combination with a dynamic weight allocation mechanism to generate task priorities; S5. Based on task priorities, use a recursive optimization method to uniformly optimize tasks with dependencies in the task graph according to the optimized task features to generate a global task optimization result; S6. Use the global task optimization result as an input, combine lighting parameters, obstacle parameters, and delay parameters in the simulation environment, and optimize the global task policy and local scene policy through a reinforcement learning model based on the Actor-Critic architecture.
2. The adaptive operation method for complex scenarios of mobile games based on multi-task learning according to claim 1, wherein The specific content of S2 includes: S21. Construct a multimodal feature extraction network, which consists of a bidirectional long short-term memory network, a ResNet34 network, and a multi-layer perceptron; S22. Use the bidirectional long short-term memory network to extract operation records and generate operation behavior features; ; Among them, represents the operation behavior feature, represents the operation record, represents the bidirectional long short-term memory network; S23. For visual information, use the ResNet34 network to extract spatial features in the image and generate visual features; ; Among them, represents visual features, represents visual information, represents the ResNet34 network; S24. For game rules, use a multi-layer perceptron to perform embedded expression on rule logic and generate rule features; ; Among them, represents a rule feature, represents a game rule, represents a multi-layer perceptron; S25. Analyze the spatial relationship in visual features, task constraints in rule features, and operation logic in operation behavior features to generate a task graph, which describes the dependency relationship, parallel relationship, and mutual exclusion relationship between tasks.
3. A method for adaptively operating complex scenes of mobile games based on multi-task learning according to claim 1, characterized in that, The specific content of S3 includes: S31. Optimize the expression of operation behavior features using an improved contrastive learning method; S32. Optimize visual features using multi-view fusion technology, construct visual features for each view, and perform weighted fusion on each view feature to generate optimized visual features; ; ; Among them, represents the optimized visual feature, represents the weight of the th perspective, represents the th visual feature of the perspective, represents the number of perspectives, represents the trainable parameter, represents the th visual feature of the perspective, represents the natural exponential function; S33. Process rule features, use the embedded coding method to optimize rule features, and generate optimized rule features; S34. Perform feature fusion on the optimized features, use the self-attention mechanism to calculate the importance weights of features, and calculate the weight matrix of each feature; ; Among them, represents the weight matrix of feature i, and respectively represent the query and key mappings of the feature vector, represents the feature dimension, T represents the transpose operation, represents the normalization function; S36. Generate multimodal features; ; Among them, represents multi-modal features, represents the value mapping of the feature vector, represents the number of modalities.
4. The adaptive operation method for complex scenarios of mobile games based on multi-task learning according to claim 3, wherein, The specific content of S31 includes: S311. Based on the operation behavior characteristics , construct contrastive learning sample pairs, where the contrastive learning sample pairs include positive sample pairs and negative sample pairs , and screen the hardest negative samples with the most similar features to the positive samples from the negative sample set to form a negative sample index set ; S312. Use the main encoder and the momentum encoder to perform feature projection on the operation behavior features to generate a stable feature representation by optimizing the weights of the momentum encoder : ; Among them, represents the momentum encoder weight, represents the main encoder weight, represents the momentum coefficient; S313. Calculate the symmetric contrastive loss of positive and negative sample pairs according to the stable feature representation, including positive similarity optimization and negative similarity optimization: ; ; Among them, represents the symmetric contrast loss of positive and negative sample pairs, represents the one-way contrast loss, represents the positive sample feature representation generated by the momentum encoder, represents the reverse positive sample feature representation generated by the momentum encoder, represents the negative sample feature representation generated by the momentum encoder, represents the cosine similarity function, represents the temperature coefficient, represents the set of negative sample indices, , and represent the feature representations generated by the momentum encoder; S314. Perform multi-scale feature extraction on the operation behavior characteristics to generate operation behavior feature representations at different scales and calculate the symmetric contrast losses at each scale and sum them up: ; Among them, The total loss of multi-scale contrastive learning, represents the operation behavior feature representation of the th scale, indicating the total number of feature scales; S315. Generate optimized operation behavior features by minimizing the total loss of multi-scale contrastive learning .
5. A method for adaptively operating complex scenes of mobile games based on multi-task learning according to claim 1, characterized in that, The specific content of S4 includes: S41. Construct a multi-task learning model, input multi-modal features into a shared feature extraction network to extract shared features. The multi-task learning model includes a shared feature extraction network and task branch networks. The shared feature extraction network is composed of multiple layers of convolutional neural networks, and the task branch networks are composed of shallow convolutional neural networks; S42. According to the task nodes in the task graph, construct an independent task branch network for each task, and extract task features through the task branch network; S43. According to the dependencies, competitions or mutual exclusions between tasks in the task graph, adjust the task features through the relevance of the shared features: ; ; Among them, represents the task feature adjusted by the shared feature, represents the inhibition factor, , and respectively represent the weighted parameters of the parent task, shared feature, and child task, used to control the contribution of the feature to the adjustment, represents the set of child tasks, represents the set of parent tasks, represents the shared feature, and respectively represent the specific feature of the parent task and the specific feature of the child task, represents the task and 's mutually exclusive relationship; S43. According to the task graph and the adjusted task features, calculate the immediate utility for each task; S44. Dynamically adjust the task loss weights according to the immediate utility; ; Among them, represents the loss weight, represents the smoothing factor to prevent the denominator from being zero, and respectively represent the task and i.e., the immediate utility; S45. Extract the dependencies between tasks according to the task graph, construct a dependency chain for each task, recursively optimize the dependency chain from the parent task to the child task, and update the features of the nested model: ; Among them, represents the nested optimization feature of the task , represents the nested optimization function and represents the task feature; S46. Generate task priorities according to the immediate utility and the dependencies between tasks.
6. The adaptive operation method for complex scenarios of mobile games based on multi-task learning according to claim 1, wherein The specific steps of S5 are as follows: S51. According to the task graph, extract the set of tasks that are mutually exclusive with the task, introduce mutual exclusion constraints on the nested features of the task, and adjust the task features through the mutual exclusion optimization function: ; Among them, represents the mutually exclusive optimization function, represents the task optimization feature, represents the optimization feature of mutually exclusive tasks, represents the cosine similarity function, represents the weight coefficient of the mutually exclusive constraint, used to control the intensity of the penalty term, represents the set of mutually exclusive tasks; S52. According to the task graph, extract the set of tasks that can be executed in parallel, based on the nested features of the parallel task set, construct a global consistency optimization objective, jointly optimize the features of the parallel task set, and update the task features: ; Among them, represents the global consistency optimization objective function, represents the set of parallel tasks, represents the optimization features of mutually exclusive tasks, represents the cosine similarity function, represents the shared features, indicating the information shared between tasks; S53. Calculate the resource requirements of the tasks according to the task graph and the task priorities, dynamically adjust the resource allocation weights according to the task priorities and resource requirements, and generate the global task optimization result.
7. A method for adaptively operating complex scenarios of mobile games based on multi-task learning according to claim 1, characterized in that The specific steps of S6 are as follows: S61. Construct a reinforcement learning model based on the Actor-Critic architecture, take the global task optimization result and environmental parameters as inputs, and generate unified state features: ; ; ; ; Among them, represents the task feature after embedding, represents the task 's nested optimization feature, represents the global task feature, represents the number of tasks, represents the set of environmental parameters, represents the global environmental feature, represents the unified state feature, 、 and respectively represent the weights of task embedding, environmental embedding, and feature fusion, 、 and respectively represent the corresponding biases, and represent the normalization parameters; S62. Output the policy through the Actor network: ; Among them, represents the output policy, and represent the weights and biases of the Actor network respectively; S63. Generate the state value function through the Critic network: ; Among them, represents the state value function, and represent the weights and biases of the Critic network, respectively; S64. Optimize the global task policy, calculate the advantage function through the immediate reward and cumulative return, update the parameters of the Actor network, and update the parameters of the Critic network by minimizing the loss function of the Critic network: ; ; ; ; Among them, represents the cumulative return, represents the discount factor, represents the immediate reward, represents the global advantage function, represents the learning rate, represents the Actor network parameters, represents the parameter gradient of, represents the Critic network loss function, represents the Critic network parameters, represents the parameter gradient of; S65. Extract the local state features of the environmental parameters, and generate the local scenario policy by optimizing the local policy loss function: ; ; ; Among them, represents the illumination feature, represents the overall obstacle feature, represents the delay feature, and represents the normalization parameter, represents the local state feature, and respectively represent the weights and biases of the fusion network, represents concatenating , and together, represents the local policy loss function, represents the local advantage function, represents the local scene policy, and respectively represent the weights and biases of the Actor network under the environmental parameters; S66. Output the optimized global task policy and local scenario policy.
Citation Information
Patent Citations
Structure inspection agent navigation method based on damage driving and multi-mode multi-task learning
CN116824303A
Knowledge association learning method and system based on knowledge graph and virtual reality
CN119166830A