A Dynamic Control Optimization Method for Mobile Games Based on Adaptive Deep Learning
Through adaptive deep learning technology, the adaptive convolutional network and adaptive memory network of space-time features are used, combined with the multi-dimensional feedback scoring mechanism, the adaptive and insufficient feedback problems of manipulation optimization in the existing technology are solved, and efficient manipulation optimization in complex game scenarios is achieved.
Patent Information
- Application Number
- CN202411605661.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-11-12
AI Technical Summary
The existing mobile game control optimization technology lacks the ability to recognize and adapt to the player's individual operating habits, and cannot provide a personalized and smooth control experience in high-complexity and high-response game scenes. The feedback and scoring mechanism is single, making it difficult to adjust the control parameters in real time.
Adaptive deep learning method is adopted to dynamically adjust the control parameters through space-time features adaptive convolution network, adaptive memory network and multi-dimensional feedback scoring mechanism, and generate accurate and personalized operation feedback, and optimize the control effect in real time.
It improves the smoothness and response speed of control, achieves high accuracy, high adaptability and good user experience, and provides a natural and smooth interactive operation experience.
Smart Images

Figure CN119548821B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of adaptive deep learning, and particularly to a method for optimizing the dynamic control of mobile games based on adaptive deep learning. Background Art
[0002] With the rapid development and popularization of mobile games, players' requirements for the game operation experience are constantly increasing. Especially in game scenarios with high complexity and high response speed requirements, an accurate, smooth, and personalized control experience has become an important measure of game quality. However, the existing mobile game control optimization technologies face many bottlenecks and technical difficulties in practical applications and are difficult to meet the continuously increasing user needs. Traditional mobile game controls mainly rely on preset fixed operation parameters and simple reaction speed adjustment strategies, lacking the ability to identify and adapt to players' individual operation habits. This method not only lacks intelligence and dynamic adaptability but also often fails to effectively respond to the diverse needs of different players and different scenarios, making players feel rigid, delayed, or even mis-touched during operation, thus affecting the game experience.
[0003] Existing control optimization methods generally rely on simple optimization models based on rules or traditional machine learning methods. These methods have obvious deficiencies in terms of adaptability and are difficult to quickly respond to players' dynamic operation needs in a real-time game environment. Specifically, traditional rule-based control optimization methods attempt to improve control sensitivity or response speed through preset parameter adjustments. However, due to the lack of a deep understanding of players' operation habits, such methods usually cannot provide a personalized operation experience in diverse game scenarios. In addition, although traditional machine learning methods have improved the ability to identify operation features to a certain extent, most of them are trained based on static data and cannot update and optimize the model in real time during actual operation, making it difficult to adapt to different players' operation preferences in real time. Therefore, during actual game play, players' operation experiences are often unsatisfactory, especially in game scenarios with high-frequency operations, significant changes in touch force, and frequent scene switches, where the existing technologies are difficult to maintain a smooth control effect.
[0004] To address the above problems, in recent years, some studies have introduced deep learning techniques to extract operation features and identify behavior patterns through methods such as convolutional neural networks and recurrent neural networks. However, these techniques are usually used for offline analysis and training of complex operations, and still have problems such as large computational complexity and poor adaptability in real-time environments. Although convolutional neural networks (CNNs) have high expressiveness in feature extraction, existing methods usually cannot adaptively adjust the size and shape of convolutional kernels to adapt to the diversity of touch operations. In addition, although traditional recurrent neural networks (RNNs) have certain advantages in memory feature and temporal feature extraction, their response to frequent changes in short-term operations is weak, and their ability to filter and remember historical data is insufficient, making it impossible to effectively associate and update the current operation and historical operations and focus on optimization in real-time operations, resulting in poor performance in changing game scenarios.
[0005] In dealing with operation intention prediction and multi-modal feature fusion, existing technologies also face many challenges. Self-supervised learning has certain applications in processing unlabeled data, but traditional self-supervised learning methods usually rely on static features and are difficult to effectively combine real-time scene features, operation intentions, and spatio-temporal features, resulting in less than ideal performance in personalized control optimization. When dealing with multi-modal feature fusion, existing technologies often use simple linear combinations or static weight assignments, lacking flexible adaptability to different scenarios, resulting in the inability to accurately judge feature priorities in specific scenarios. In addition, the design of the feedback scoring mechanism in existing technologies is also relatively single, usually only relying on a single operation success rate or response speed, lacking a comprehensive scoring mechanism for control accuracy, operation fluency, and player satisfaction, making it difficult to provide comprehensive feedback for subsequent adjustments.
[0006] Based on the above background, the limitations of existing technologies are mainly reflected in the following aspects: First, there is a lack of the ability to adaptively adjust touch features, resulting in the inability to extract and respond to dynamic features in operation data in real-time; second, there is a lack of an intelligent memory update mechanism for operation historical data, and existing recurrent neural networks are difficult to dynamically associate the features of the current operation with historical operations, resulting in inaccurate operation intention recognition; third, existing self-supervised learning methods are difficult to effectively generate adaptive multi-modal feature fusion and cannot comprehensively consider the importance of different features in different scenarios, resulting in an inaccurate control optimization process; fourth, there is a lack of a comprehensive feedback scoring mechanism, and existing technologies fail to make full use of feedback scoring for dynamic adjustment of operation parameters and adaptive learning, making it difficult to achieve control optimization based on real-time feedback. These defects make the existing mobile game control optimization technologies have obvious deficiencies in coping with high-frequency, complex, and diverse game operation requirements.
[0007] Therefore, how to provide a dynamic control optimization method for mobile games based on adaptive deep learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0008] An object of the present invention is to propose a dynamic control optimization method for mobile games based on adaptive deep learning. By introducing the adaptive deep learning method, the present invention adopts a spatio-temporal feature adaptive convolutional network, an adaptive memory network and a multi-dimensional feedback scoring mechanism to perform real-time dynamic optimization on the control of mobile games. This method can adaptively adjust the control parameters according to the touch force, operation intention and scene requirements, and generate accurate and personalized operation feedback. In complex and highly dynamic game scenarios, the present invention significantly improves the fluency and response speed of the control, has the advantages of high precision, high adaptability and good user experience, and provides a natural and fluent interactive operation experience for players.
[0009] A dynamic control optimization method for mobile games based on adaptive deep learning according to an embodiment of the present invention includes the following steps:
[0010] S1. Collect player operation data and scene background parameters, and generate feature data of a unified scale through data denoising, normalization processing and accidental touch data filtering;
[0011] S2. Adjust the size and shape of the convolutional kernel through a spatio-temporal feature adaptive convolutional network, and dynamically adjust the convolutional kernel and the time sampling step according to the touch force and trajectory to generate an adaptive spatio-temporal feature sequence;
[0012] S3. Construct an adaptive memory network, dynamically update the memory pool according to the time correlation between the current operation and the historical operation, and identify recent high-frequency operation patterns through a short-term focusing mechanism to generate operation intention features;
[0013] S4. Generate a self-supervised learning task through a self-supervised learning multi-layer perceptron, fuse the feature data of a unified scale, the adaptive spatio-temporal feature sequence and the operation intention features into multi-mode features, and adaptively adjust the feature weights according to the scene requirements to generate control parameters;
[0014] S5. Generate a feedback score in real time through a feedback scoring mechanism, and dynamically adjust the learning rate based on the feedback score to achieve adaptive adjustment;
[0015] S6. Establish a satisfaction evaluation mechanism based on the feedback score and data analysis, and adjust the control parameters in real time through a re-learning mechanism when detecting dissatisfied operations;
[0016] S7. Detect the response speed and operation accuracy, record the operation feedback and scoring results in real time, and achieve adaptive optimization based on the feedback loop.
[0017] Optionally, the S2 specifically includes:
[0018] S21. Initialize the extraction module of the spatio-temporal feature adaptive convolution network, input the collected touch data into the input layer of the adaptive convolution network, and extract the touch position, touch force, and trajectory dynamic features;
[0019] S22. Based on the real-time touch force F strength (t) and the trajectory change rate R trajectory (t), dynamically adjust the size parameter of the convolution kernel:
[0020]
[0021] Among them, K adaptive_size ( t ) represents the dynamic adaptive parameter of the convolution kernel size, α conv_scale represents the overall adjustment coefficient of the convolution kernel, λ path_mod represents the path adaptive adjustment coefficient, σ strength_bias represents the intensity offset coefficient, η response_mod represents the response adjustment coefficient of the touch intensity, θ decay represents the exponential decay coefficient, β adapt_smooth represents the adaptive smoothing coefficient, δ strength_norm represents the touch intensity normalization coefficient;
[0022] S23. Dynamically adjust the shape parameter K of the convolution kernel according to the continuity and direction change of the operation trajectory adaptive_shape (t) to make the convolution kernel shape adapt to the change direction of the touch path;
[0023] S24. When the change of the touch acceleration A feedback (t) is detected, dynamically adjust the sampling step:
[0024]
[0025] Among them, T dynamic_interval ( t ) represents the dynamic adaptive parameter of the time sampling step, γ base_interval represents the basic sampling step coefficient, δ accel_mod represents the acceleration adjustment coefficient, ∈ log_smooth represents the logarithmic smoothing coefficient, ν strength_weight represents the touch intensity weight coefficient, κ decay_rate represents the exponential decay coefficient;
[0026] S25. Normalize the spatio-temporal features obtained at different time sampling steps to generate a normalized spatio-temporal feature sequence;
[0027] S26. Integrate the adjusted convolution kernel size, shape parameters, and the normalized spatio-temporal feature sequence to generate an adaptive spatio-temporal feature sequence S sta (t).
[0028] Optionally, the S3 specifically includes:
[0029] S31. Initialize the adaptive memory network, input the current operation data and the historical operation sequence into the memory network, establish the temporal association between the current operation and the historical operations, and form multi-dimensional temporal features;
[0030] S32. Construct a dynamic memory pool, store the historical operation features highly relevant to the current operation in the memory pool, and the state of the dynamic memory pool is denoted as M dynamic (t), and the state M dynamic (t) is based on the non-linear weighted combination of the current operation feature V current (t) and the historical operation feature H history (t) to achieve dynamic update:
[0031]
[0032] where, α cur represents the weight coefficient of the current operation, β hist represents the weight coefficient of the historical operation, tanh represents the hyperbolic tangent function, λ rate represents the rate of change of temporal association, σ decay represents the adaptive decay coefficient of the dynamic memory pool, γ memory represents the fusion coefficient of the dynamic memory pool, δ stabilizer represents the state stability coefficient, and exp represents the exponential function;
[0033] S33. Introduce a short-term focusing mechanism to aggregate the operation frequencies at the current moment and several recent moments, define an aggregation time window W focus , and statistically calculate the operation frequency feature F focus (t) within the aggregation time window W freq to capture the high-frequency operation features in the short term;
[0034] S34. Generate the high-frequency pattern feature V dynamic (t) based on the dynamic memory pool state M freq (t) and the operation frequency feature F pattern (t):
[0035]
[0036] where, γ freq represents the weight coefficient of the operation frequency feature, δ mem represents the weight coefficient of the dynamic memory pool state, ηinteraction Represents the interaction adjustment coefficient of the frequency feature and the memory state, θ context Represents the context correlation coefficient, κ stabilizer Represents the mode stability coefficient;
[0037] S35. For the high-frequency mode feature V pattern (t) and the current operation feature V current (t) perform feature fusion to generate the preliminary operation intention vector V intent_raw (t);
[0038] S36. Smooth the preliminary operation intention vector V intent_raw (t) to eliminate noise and generate the final operation intention feature V intent (t).
[0039] Optionally, the specific steps of S4 include:
[0040] S41. Create a self-supervised learning task in the self-supervised learning multi-layer perceptron, construct a pseudo-label set for unlabeled data, and input the feature data matrix X unified with a unified scale into the self-supervised learning multi-layer perceptron, and construct a pseudo-label set Y pseudo based on the game scene, and simulate the player's operation behavior to obtain the preliminary learning objective;
[0041] S42. Fuse the adaptive spatio-temporal feature sequence S sta (t) and the operation intention feature V intent (t) to generate the preliminary multi-mode feature vector F multi ( t);
[0042] S43. Further integrate the multi-mode feature vector F multi (t) with the feature data matrix X unified with a unified scale to generate the multi-mode comprehensive feature vector F combined (t), which comprehensively reflects the player's operation characteristics and scene requirements by integrating data from different sources;
[0043] S44. Dynamically adjust the feature weights based on the current game scene, construct the scene adaptive weight factor matrix W scene (t), and adjust the weights in the multi-mode comprehensive feature to ensure the best fusion ratio of each feature in different scenes;
[0044] S45. Perform non-linear activation processing on the multi-mode comprehensive feature vector F combined (t), and generate the manipulation feature vector F control (t) through the activation function;
[0045] S46. The manipulation feature vector Fcontrol (t) is input into the feature reconstruction layer, and a non-linear transformation is performed through the combination of the feature reconstruction weight matrix and the context adjustment matrix to generate operation parameters:
[0046]
[0047] Among them, P control (t0 represents the operation parameter at time t, W control represents the feature reconstruction weight matrix, C context represents the context adjustment matrix, α represents the coefficient of the non-linear activation term, β represents the smoothing factor, γ represents the exponential decay rate coefficient, δ represents the historical stability coefficient, P control (t - 1) represents the operation parameter at time t - 1, P control (t - 2) represents the operation parameter at time t - 2.
[0048] Optionally, the S5 specifically includes:
[0049] S51. Construct a multi-dimensional feedback scoring mechanism, and generate a feedback score value F score (t) in real time through weighted combination of indicators such as the accuracy, response speed, and success rate of the player's operations, for quantifying the current control effect
[0050] S52. Define the feedback score change rate ΔF score (t), for monitoring the change trend of the feedback score, capturing the rate of increase or decrease of the score, and the feedback score change rate ΔF score (t) identifies the change trend of the operation performance, and thus starts the learning rate adjustment according to the change trend;
[0051] S53. Design an adaptive learning rate adjustment function to non-linearly adjust the learning rate according to the feedback score change trend:
[0052]
[0053] Among them, η ( t ) represents the dynamic learning rate at time t, η base represents the base learning rate, α adjust represents the learning rate adjustment amplitude coefficient, tanh represents the hyperbolic tangent function, β rate represents the score change rate adjustment coefficient, γ smooth represents the smoothing coefficient, δ stabilize represents the stability factor, F target represents the preset target score;
[0054] S54. Introduce a fast adjustment mechanism, when the feedback score value F score (t) is lower than the set minimum score threshold Fmin When the [specific condition], activate the learning rate increase adjustment to adapt to the changing requirements of the current operation performance. After the score stabilizes, automatically return to the regular adaptive adjustment mode;
[0055] S55. Construct a score smoothing module to perform mean smoothing on the score within a certain time window to reduce the direct impact of score fluctuations on the learning rate;
[0056] S56. During the entire feedback score generation and learning rate adjustment process, based on the smoothed dynamic learning rate η smooth ( t) achieve continuous adaptive optimization.
[0057] Optionally, the specific steps of S6 are as follows:
[0058] S61. Establish a comprehensive satisfaction evaluation model, and use the feedback score value of the player, the degree of operation flow, and the key parameter of response delay to generate a comprehensive satisfaction score M satisfaction (t) to quantify the current operation experience. The comprehensive satisfaction score M satisfaction (t) is calculated based on a multi-dimensional scoring function, and the weight of each scoring dimension is set according to the operation requirements;
[0059] S62. Design a satisfaction change trend analysis mechanism to dynamically monitor the comprehensive satisfaction scores M satisfaction (t) at consecutive moments, capture the change trend, and by calculating the increment ΔM of the comprehensive satisfaction score at the current moment satisfaction (t), if the score continues to decline, it is marked as a downward trend, otherwise it is marked as an upward trend. When it is detected that the downward trend continuously exceeds the set number threshold, it is judged that the player's satisfaction is decreasing, and trigger the dissatisfied operation detection;
[0060] S63. Establish a dissatisfied operation detection mechanism, and set a minimum satisfaction threshold M min , which is used to identify whether the player's satisfaction with the current control reaches the expectation. If the comprehensive satisfaction score M satisfaction (t) is lower than the minimum satisfaction threshold M min , mark the current control as a dissatisfied operation and enter the low satisfaction state;
[0061] S64. When detecting the low satisfaction state of a dissatisfied operation, activate the re-learning mechanism and use an increased re-learning rate η relearn to adjust the control parameters. The re-learning rate η relearn is set according to the satisfaction gap and change trend:
[0062] η relearn (t) =
[0063]
[0064] where η relearn ( t ) represents the re - learning rate at time t, η base represents the base learning rate, α adapt represents the learning rate increment coefficient, β trend represents the trend response coefficient, λ stability represents the stability adjustment coefficient, γ smooth represents the smoothing factor, δ balance represents the balance coefficient, M target represents the target satisfaction;
[0065] S65. During the re - learning process, continuously optimize the control parameters according to real - time feedback, feedback scores, and satisfaction trends;
[0066] S66. When the satisfaction score returns to the set target satisfaction M target or when the stable state is satisfied for several consecutive operations, gradually reduce the re - learning rate and restore to the normal learning rate mode.
[0067] The beneficial effects of the present invention are as follows:
[0068] First, the present invention adopts a spatio - temporal feature adaptive convolutional network, which can dynamically adjust the size and shape of the convolutional kernel, and adaptively capture operation details in real - time according to the touch force and trajectory changes, effectively avoiding the problem that the existing fixed convolutional kernel cannot adapt to the diversity of touches. This innovation enables the system to more accurately extract spatio - temporal features in the player's operations, capture changes in different touch forces and positions, maintain flexible adaptability to touch features in complex game scenarios, thereby improving the fluency and accuracy of control and bringing more natural operation feedback to players.
[0069] Second, by constructing an adaptive memory network, the present invention realizes accurate prediction of operation intentions and efficient memory update. The adaptive memory network can dynamically update the memory pool at different times, intelligently associate the time correlation between the current operation and historical operations, and effectively identify recent high - frequency operation patterns through a short - term focusing mechanism to generate operation intention features. This dynamic update mechanism of the memory network makes up for the defects of inflexible memory association and inaccurate operation prediction in the prior art, enabling the system to always maintain accurate recognition of the player's intentions during complex and changing game operations, providing a more reliable basis for the adjustment of operation parameters.
[0070] In addition, the present invention also generates an adaptive multi-modal feature fusion mechanism through self-supervised learning of a multi-layer perceptron, effectively fusing feature data of a unified scale, an adaptive spatio-temporal feature sequence, and operation intention features. Based on the requirements of the game scenario, the system can dynamically adjust the weights of multi-modal features, enabling flexible control of the contribution degrees of different features in different scenarios. This design enables the system to provide more targeted control optimization in high-dynamic game situations, ensuring reasonable distribution of the weights of each feature in complex scenarios, effectively improving the generation accuracy of control parameters, and thereby enhancing the adaptability of the system in diverse scenarios.
[0071] In terms of the feedback mechanism, the present invention constructs a multi-dimensional feedback scoring mechanism to generate feedback scores in real time and dynamically adjust the learning rate through the feedback scores, ensuring that the system can adaptively adjust in a timely manner when the operation performance fluctuates. The real-time acquisition of multi-dimensional feedback scores such as operation accuracy, response speed, and player satisfaction enables the system to comprehensively evaluate the current control effect and adaptively adjust the learning rate according to the change trend of the scores, thereby avoiding the problem that a single score cannot reflect the overall control quality. Through this mechanism, the present invention can achieve highly sensitive adaptive optimization in the feedback link. The system can promptly increase the learning rate when the player's operation deviates and adjust smoothly in the stable state, ensuring continuous improvement in control optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0073] Figure 1 is a flowchart of a method for dynamically optimizing the control of a mobile game based on adaptive deep learning proposed by the present invention;
[0074] Figure 2 is a schematic structural diagram of adjusting the learning rate and control parameters of the feedback scoring mechanism of a method for dynamically optimizing the control of a mobile game based on adaptive deep learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0076] Refer to Figure 1 and Figure 2 , a method for dynamically optimizing the control of a mobile game based on adaptive deep learning includes the following steps:
[0077] S1. Collect player operation data and scene background parameters, and generate feature data of a unified scale through data denoising, normalization processing, and accidental touch data filtering;
[0078] S2. Adjust the size and shape of the convolution kernel through a spatio-temporal feature adaptive convolutional network, and dynamically adjust the convolution kernel and time sampling step according to the touch force and trajectory to generate an adaptive spatio-temporal feature sequence;
[0079] S3. Construct an adaptive memory network, dynamically update the memory pool according to the time correlation between the current operation and historical operations, and identify recent high-frequency operation patterns through a short-term focusing mechanism to generate operation intention features;
[0080] S4. Generate a self-supervised learning task through a self-supervised learning multi-layer perceptron, fuse the feature data of a unified scale, the adaptive spatio-temporal feature sequence, and the operation intention features into multi-modal features, and adaptively adjust the feature weights according to the scene requirements to generate control parameters;
[0081] S5. Generate a feedback score in real time through a feedback scoring mechanism, and dynamically adjust the learning rate based on the feedback score to achieve adaptive adjustment;
[0082] S6. Establish a satisfaction evaluation mechanism based on the feedback score and data analysis, and adjust the control parameters in real time through a re-learning mechanism when detecting dissatisfied operations;
[0083] S7. Detect the response speed and operation accuracy, record the operation feedback and scoring results in real time, and achieve adaptive optimization based on the feedback loop.
[0084] In this embodiment, S2 specifically includes:
[0085] S21. Initialize the extraction module of the spatio-temporal feature adaptive convolutional network, input the collected touch data into the input layer of the adaptive convolutional network, and extract the touch position, touch force, and trajectory dynamic features;
[0086] S22. Based on the real-time touch force F strength (t) and the trajectory change rate R trajectory (t), dynamically adjust the size parameter of the convolution kernel:
[0087]
[0088] Among them, K adaptive_size ( t ) represents the dynamic adaptive parameter of the convolution kernel size, α conv_scale represents the overall adjustment coefficient of the convolution kernel, λ path_mod represents the path adaptive adjustment coefficient, σ strength_bias represents the intensity offset coefficient, ηresponse_mod Response adjustment coefficient representing touch intensity, θ decay Exponential decay coefficient, β adapt_smooth Adaptive smoothing coefficient, δ strength_norm Touch intensity normalization coefficient;
[0089] S23. Dynamically adjust the convolutional kernel shape parameter K adaptive_shape (t) according to the continuity and direction change of the operation trajectory, so that the convolutional kernel shape adapts to the change direction of the touch path;
[0090] S24. When detecting the change of touch acceleration A feedback (t), dynamically adjust the sampling step size:
[0091]
[0092] where T dynamic_interval ( t ) Dynamic adaptive parameter representing the time sampling step size, γ base_interval Base sampling step size coefficient, δ accel_mod Acceleration adjustment coefficient, ∈ log_smooth Logarithmic smoothing coefficient, ν strength_weight Touch intensity weight coefficient, κ decay_rate Exponential decay coefficient;
[0093] S25. Normalize the spatio-temporal features obtained at different time sampling steps to generate a normalized spatio-temporal feature sequence;
[0094] S26. Integrate the adjusted convolutional kernel size, shape parameter, and the normalized spatio-temporal feature sequence to generate an adaptive spatio-temporal feature sequence S sta (t).
[0095] In this embodiment, the S3 specifically includes:
[0096] S31. Initialize the adaptive memory network, input the current operation data and the historical operation sequence into the memory network, establish the time correlation between the current operation and the historical operation, and form multi-dimensional time series features;
[0097] S32. Construct a dynamic memory pool, store the historical operation features highly relevant to the current operation in the memory pool, and the dynamic memory pool state is represented as M dynamic (t), and the dynamic memory pool state M dynamic (t) is based on the non-linear weighted combination of the current operation feature V current (t) and the historical operation feature H history (t) to achieve dynamic update:
[0098]
[0099] Among them, α cur represents the weight coefficient of the current operation, β hist represents the weight coefficient of the historical operation, tanh represents the hyperbolic tangent function, λ rate represents the time correlation change rate, σ decay represents the adaptive decay coefficient of the dynamic memory pool, γ memory represents the fusion coefficient of the dynamic memory pool, δ stabilizer represents the state stability coefficient, exp represents the exponential function;
[0100] S33. Introduce a short-term focusing mechanism to aggregate the operation frequencies at the current moment and several recent moments, and define the aggregation time window W focus , and within the aggregation time window W focus statistically analyze the operation frequency feature F freq (t) to capture the high-frequency operation features in the short term;
[0101] S34. Generate the high-frequency pattern feature V dynamic (t) based on the dynamic memory pool state M freq (t) and the operation frequency feature F pattern (t):
[0102]
[0103] Among them, γ freq represents the operation frequency feature weight coefficient, δ mem represents the dynamic memory pool state weight coefficient, η interaction represents the interaction adjustment coefficient between the frequency feature and the memory state, θ context represents the context correlation coefficient, κ stabilizer represents the pattern stability coefficient;
[0104] S35. Perform feature fusion on the high-frequency pattern feature V pattern (t) and the current operation feature V current (t) to generate the preliminary operation intention vector V intent_raw (t);
[0105] S36. Smooth the preliminary operation intention vector V intent_raw (t) to eliminate noise and generate the final operation intention feature V intent (t).
[0106] In this embodiment, the specific content of S4 includes:
[0107] S41. Create a self-supervised learning task in the self-supervised learning multi-layer perceptron, construct a pseudo-label set for unlabeled data, and input the feature data matrix X of unified scale unified into the self-supervised learning multi-layer perceptron, and construct a pseudo-label set Y based on the game scenario pseudo Simulate the player's operation behavior to obtain a preliminary learning objective;
[0108] S42. Integrate the adaptive spatio-temporal feature sequence S sta (t) and the operation intention feature V intent (t) to generate a preliminary multi-modal feature vector F multi ( t);
[0109] S43. Further integrate the multi-modal feature vector F multi (t) and the feature data matrix X of unified scale unified to generate a multi-modal comprehensive feature vector F combined (t), which comprehensively reflects the player's operation characteristics and scene requirements by integrating data from different sources;
[0110] S44. Dynamically adjust the feature weights based on the current game scenario, construct a scene adaptive weight factor matrix W scene (t) to adjust the weights in the multi-modal comprehensive features to ensure the best fusion ratio of each feature in different scenarios;
[0111] S45. Perform non-linear activation processing on the multi-modal comprehensive feature vector F combined (t) to generate a manipulation feature vector F control (t) through an activation function;
[0112] S46. Input the manipulation feature vector F control (t) into the feature reconstruction layer, and perform non-linear transformation through the combination of the feature reconstruction weight matrix and the context adjustment matrix to generate operation parameters:
[0113]
[0114] where, P control (t) represents the operation parameter at time t, W control represents the feature reconstruction weight matrix, C context represents the context adjustment matrix, α represents the coefficient of the non-linear activation term, β represents the smoothing factor, γ represents the exponential decay rate coefficient, δ represents the historical stability coefficient, P control (t - 1) represents the operation parameter at time t - 1, P control ( t - 2) represents the operation parameter at time t - 2.
[0115] In this embodiment, S5 specifically includes:
[0116] S51. Construct a multi-dimensional feedback scoring mechanism, and generate a feedback score value F(t) in real time by calculating the weighted combination of the accuracy, response speed, and success rate of the player's operations, which is used to quantify the current control effect. score (t) for quantifying the current manipulation effect
[0117] S52. Define the feedback score change rate ΔF(t), which is used to monitor the change trend of the feedback score, capture the rising or falling rate of the score, and the feedback score change rate ΔF(t) identifies the change trend of the operation performance, so as to start the learning rate adjustment according to the change trend. score (t) for monitoring the change trend of the feedback score, capturing the rising or falling rate of the score, and the feedback score change rate ΔF score (t) identifies the change trend of the operation performance, thereby starting the learning rate adjustment according to the change trend;
[0118] S53. Design an adaptive learning rate adjustment function to non-linearly adjust the learning rate according to the feedback score change trend:
[0119]
[0120] where η ( t ) represents the dynamic learning rate at time t, η base represents the base learning rate, α adjust represents the learning rate adjustment amplitude coefficient, tanh represents the hyperbolic tangent function, β rate represents the score change rate adjustment coefficient, γ smooth represents the smoothing coefficient, δ stabilize represents the stability factor, F target represents the preset target score;
[0121] S54. Introduce a fast adjustment mechanism. When the feedback score value F(t) is lower than the set minimum score threshold F score (t), activate the learning rate increase adjustment to adapt to the change requirements of the current operation performance. After the score returns to stability, automatically return to the normal adaptive adjustment mode; min When the feedback score value F(t) is lower than the set minimum score threshold F, activate the learning rate increase adjustment to adapt to the change requirements of the current operation performance. After the score returns to stability, automatically return to the normal adaptive adjustment mode;
[0122] S55. Construct a score smoothing module to perform mean smoothing processing on the score within a certain time window to reduce the direct impact of score fluctuations on the learning rate;
[0123] S56. During the entire feedback score generation and learning rate adjustment process, realize continuous adaptive optimization based on the smoothed dynamic learning rate η(t). smooth ( t) to achieve continuous adaptive optimization.
[0124] In this embodiment, S6 specifically includes:
[0125] S61. Establish a comprehensive satisfaction evaluation model, and generate a comprehensive satisfaction score M by using the feedback score value of players, the degree of operation flow, and key parameters such as response delay, so as to quantify the current operation experience. The comprehensive satisfaction score M(t) is calculated based on a multi-dimensional scoring function, and the weight of each scoring dimension is set according to operation requirements; satisfaction (t), quantifying the current operation experience, the comprehensive satisfaction score M satisfaction (t) is calculated based on a multi-dimensional scoring function, and the weight of each scoring dimension is set according to operation requirements;
[0126] S62. Design a satisfaction change trend analysis mechanism to dynamically monitor the comprehensive satisfaction score M(t) at consecutive moments, capture the change trend, and by calculating the increment ΔM(t) of the comprehensive satisfaction score at the current moment, if the score continues to decline, it is marked as a downward trend, otherwise it is marked as an upward trend. When it is detected that the downward trend continuously exceeds the set number threshold, it is determined that the player's satisfaction is decreasing, and the dissatisfaction operation detection is triggered; satisfaction (t) for dynamic monitoring, capturing the change trend, by calculating the increment ΔM satisfaction (t) of the comprehensive satisfaction score at the current moment, if the score continues to decline, it is marked as a downward trend, otherwise it is marked as an upward trend. When it is detected that the downward trend continuously exceeds the set number threshold, it is determined that the player's satisfaction is decreasing, and the dissatisfaction operation detection is triggered;
[0127] S63. Establish a dissatisfaction operation detection mechanism, set a minimum satisfaction threshold M, which is used to identify whether the player's satisfaction with the current operation reaches the expectation. If the comprehensive satisfaction score M(t) is lower than the minimum satisfaction threshold M, mark the current operation as a dissatisfaction operation and enter the low satisfaction state; min , used to identify whether the player's satisfaction with the current operation reaches the expectation. If the comprehensive satisfaction score M satisfaction (t) is lower than the minimum satisfaction threshold M min , mark the current operation as a dissatisfaction operation and enter the low satisfaction state;
[0128] S64. When the low satisfaction state of the dissatisfaction operation is detected, activate the re-learning mechanism, and use an increased re-learning rate η to adjust the operation parameters. The re-learning rate η is set according to the satisfaction gap and the change trend: relearn Adjust the operation parameters, and the re-learning rate η relearn is set according to the satisfaction gap and the change trend:
[0129]
[0130] where η relearn ( t ) represents the re-learning rate at time t, η base represents the basic learning rate, α adapt represents the learning rate increase coefficient, β trend represents the trend response coefficient, λ stability represents the stable adjustment coefficient, γ smooth represents the smoothing factor, δ balance represents the balance coefficient, M target represents the target satisfaction;
[0131] S65. During the re-learning process, continuously optimize the operation parameters according to real-time feedback, feedback scores, and satisfaction trends;
[0132] S66. When the satisfaction score returns to the set target satisfaction Mtarget When a stable state is satisfied after a series of consecutive operations, gradually decrease the re - learning rate and restore to the normal learning rate mode.
[0133] Example 1:
[0134] To verify the feasibility of the present invention in implementation, the present invention is applied to a certain real - time strategy mobile game. This game requires players to quickly perform multi - finger operations in complex combat scenarios, such as commanding character movement, triggering skills, avoiding enemy attacks, etc. The game operation environment is complex and changeable, with extremely high requirements for operation accuracy, response speed, and control fluency. Existing technologies are difficult to adapt to such diverse operation requirements in real time. Players often encounter problems such as slow touch response, misjudgment of operation intentions, and inaccurate gesture recognition during operation, resulting in a poor gaming experience. Especially in emergency combat scenarios with high - frequency operations, the satisfaction and operation fluency of players are relatively low.
[0135] Apply the adaptive deep - learning dynamic control optimization method of the present invention in this game. By collecting the player's touch operation data (including touch position, force, sliding trajectory, etc.) and the background parameters of the current game scene (such as game difficulty, combat status, enemy position, etc.), after denoising and normalizing the touch data, input it into the system's adaptive convolutional network and adaptive memory network, so as to generate control parameters in real time and dynamically adjust the character movement and skill release in the game. The adaptive convolutional network adjusts the size and shape of the convolutional kernel according to the changes in touch force and touch position, extracts touch spatio - temporal features, enabling the system to recognize the operation intentions of different touch positions and forces. The adaptive memory network dynamically updates the memory pool according to the player's current operations and historical operation data, extracts recent high - frequency operation patterns through a short - term focusing mechanism, and thus accurately recognizes the player's intentions. Based on multi - mode feature fusion, the self - supervised learning mechanism generates control parameters, and scores the real - time operations through a feedback scoring mechanism, dynamically adjusting the learning rate according to the scoring trend, so that the system can respond to dissatisfied operations in real time.
[0136] In actual tests, we selected two groups of players to conduct a comparative test in the same game scenario. The first group of players used the conventional control mode, and the second group of players used the dynamic control optimization method of the present invention. The test scenario of the game was set as a combat level with a high degree of difficulty, a large number of enemies, and frequent scene changes, and players needed to continuously perform touch operations. The entire test process was carried out in a game test center in Beijing from May 10, 2024 to May 17, 2024. The test involved a total of 50 players, and each person completed 20 operation tasks.
[0137] Table 1 Comparison table of touch response speed data of players in the conventional group and the group of the present invention
[0138]
[0139]
[0140] Table 2 Comparison Table of Touch Accuracy Data of Players in the Conventional Group and the Group of the Present Invention
[0141]
[0142]
[0143] Table 3 Comparison Table of Satisfaction Ratings of Players in the Conventional Group and the Group of the Present Invention
[0144] Test Date Player Group Satisfaction Score (out of 5) May 10, 2024 Regular Group 3.9 May 10, 2024 Group of the Present Invention 4.8 May 11, 2024 Regular Group 3.7 May 11, 2024 Group of the Present Invention 4.6 May 12, 2024 Regular Group 3.8 May 12, 2024 Group of the Present Invention 4.7 May 13, 2024 Regular Group 3.6 May 13, 2024 Group of the Present Invention 4.7 May 14, 2024 Regular Group 3.8 May 14, 2024 Group of the Present Invention 4.6 May 15, 2024 Regular Group 3.7 May 15, 2024 Group of the Present Invention 4.8 May 16, 2024 Regular Group 3.8 May 16, 2024 Group of the Present Invention 4.7 May 17, 2024 Regular Group 3.8 May 17, 2024 Group of the Present Invention 4.7 Average Regular Group 3.8 Average Group of the Present Invention 4.7
[0145] In the dynamic control optimization test of the present invention, by analyzing the data of different player groups in terms of touch response time, operation accuracy, and user satisfaction, the significant improvement effect of the present invention in the mobile game control experience is fully demonstrated. First, in terms of touch response time (see Table 1), the average response time of the player group using the present invention is 84 milliseconds, which is 24.7% shorter than the 112 milliseconds of the conventional group on average, and the standard deviation of the response time is also significantly reduced, indicating that the application of the present invention in dynamic game scenarios significantly improves the stability and speed of control response. The fast response not only optimizes the player's operation experience but also improves the operation accuracy of the player in high-frequency operation scenarios.
[0146] In terms of operation accuracy (see Table 2), the average touch accuracy of the player group using the present invention reaches 98.9%, which is significantly higher than 85.3% of the conventional group. This shows the significant improvement effect of the present invention in complex touch operations, especially in multi-finger, high-speed sliding, and frequent switching operations, where the mis-touch rate is only 1.1%, significantly lower than 14.7% of the conventional group. These data indicate that the adaptive feature extraction and memory mechanism of the present invention can accurately capture the operation intention, effectively reduce mis-touch situations, and improve the operation precision.
[0147] In terms of player satisfaction (see Table 3), the average score of the player group using the present invention is 4.7 points (out of 5), showing a significant improvement compared to 3.8 points of the conventional group. This difference reflects the high recognition of the players for the control experience provided by the present invention. The improvement in satisfaction stems from the overall improvement of the control optimization method in terms of response speed, touch accuracy, and operation fluency. Especially in frequently changing game scenarios, players obtain a more natural and smooth interaction experience.
[0148] Therefore, it can be seen that the dynamic control optimization method of the present invention has significant advantages in terms of touch response speed, operation accuracy, and player satisfaction through multi-level deep learning technology, verifying the effectiveness and superiority of the present invention in complex game environments.
[0149] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A dynamic control optimization method for mobile games based on adaptive deep learning, characterized in that, It includes the following steps: S1. Collect player operation data and scene background parameters, and generate feature data of a unified scale through data denoising, normalization processing, and accidental touch data filtering; S2. Adjust the size and shape of the convolution kernel through a spatio-temporal feature adaptive convolutional network, and dynamically adjust the convolution kernel and time sampling step according to the touch force and trajectory to generate an adaptive spatio-temporal feature sequence; S3. Construct an adaptive memory network, dynamically update the memory pool according to the time correlation between the current operation and historical operations, and identify recent high-frequency operation patterns through a short-term focusing mechanism to generate operation intention features; S4. Generate self-supervised learning tasks through a self-supervised learning multi-layer perceptron, fuse the feature data of a unified scale, the adaptive spatio-temporal feature sequence, and the operation intention features into multi-mode features, and adaptively adjust the feature weights according to the scene requirements to generate control parameters; S5. Generate a feedback score in real time through a feedback scoring mechanism, and dynamically adjust the learning rate based on the feedback score to achieve adaptive adjustment; S6. Establish a satisfaction evaluation mechanism based on the feedback score and data analysis, and adjust the control parameters in real time through a re-learning mechanism when detecting dissatisfied operations; S7. Detect the response speed and operation accuracy, record the operation feedback and scoring results in real time, and achieve adaptive optimization based on the feedback loop; The specific content of S3 includes: S31. Initialize the adaptive memory network, input the current operation data and the historical operation sequence into the memory network, establish the time correlation between the current operation and historical operations, and form multi-dimensional time series features; S32. Build a dynamic memory pool, store historical operation features highly relevant to the current operation in the memory pool, and represent the state of the dynamic memory pool as M dynamic (t), where the state M dynamic (t) is based on the current operation feature V current (t) and the historical operation feature H history (t) through a non-linear weighted combination to achieve dynamic update: Among them, α cur represents the weight coefficient of the current operation, β hist represents the weight coefficient of the historical operation, tanh represents the hyperbolic tangent function, λ rate represents the time correlation change rate, σ decay represents the adaptive decay coefficient of the dynamic memory pool, γ memory represents the fusion coefficient of the dynamic memory pool, δ stabilizer represents the state stability coefficient, exp represents the exponential function; S33. Introduce a short-term focus mechanism to aggregate the operation frequencies at the current moment and several recent moments, and define an aggregation time window W focus , and within the aggregation time window W focus statistically calculate the operation frequency feature F freq (t) to capture the high-frequency operation features in the short term; S34. Based on the dynamic memory pool state M dynamic (t) and the operation frequency feature F freq (t) to generate the high-frequency pattern feature V pattern (t): Among them, γ freq represents the operation frequency feature weight coefficient, δ mem represents the dynamic memory pool state weight coefficient, η interaction represents the interaction adjustment coefficient of the frequency feature and the memory state, θ context represents the context correlation coefficient, κ stabilizer represents the pattern stability coefficient; S35. Feature fusion is performed on the high-frequency mode feature V pattern (t) and the current operation feature V current (t) to generate a preliminary operation intention vector V intent_raw (t); S36. Smooth the preliminary operation intention vector V intent_raw (t) to eliminate noise and generate the final operation intention feature V intent (t); The specific content of S4 includes: S41. Create a self-supervised learning task in the self-supervised learning multi-layer perceptron, construct a pseudo-label set for unlabeled data, and input the feature data matrix X of the unified scale unified into the self-supervised learning multi-layer perceptron, and construct a pseudo-label set Y based on the game scenario pseudo Simulate the player's operation behavior to obtain a preliminary learning goal; S42. Fuse the spatio-temporal feature sequence S sta (t) and the operation intention feature V intent (t) to generate a preliminary multi-mode feature vector F multi (t); S43. Integrate the multi-mode feature vector F multi (t) with the feature data matrix X of a unified scale unified further to generate a multi-mode comprehensive feature vector F combined (t), which comprehensively reflects the player operation characteristics and scenario requirements by integrating data from different sources; S44. Dynamically adjust the feature weights based on the current game scene, construct a scene adaptive weight factor matrix W scene (t), and adjust the weights in the multi-mode comprehensive features to ensure that each feature has the best fusion ratio in different scenes; S45. Nonlinearly activate the multi-mode comprehensive feature vector F combined (t) to generate a manipulation feature vector F control (t) through an activation function; S46. Input the manipulation feature vector F control (t) into the feature reconstruction layer, perform a non-linear transformation through the combination of the feature reconstruction weight matrix and the context adjustment matrix, and generate an operation parameter: Among them, P control (t) represents the operation parameter at time t, W control represents the feature reconstruction weight matrix, C context represents the context adjustment matrix, α represents the coefficient of the non-linear activation term, β represents the smoothing factor, γ represents the exponential decay rate coefficient, δ represents the historical stability coefficient, P control (t - 1) represents the operation parameter at time t - 1, P control (t - 2) represents the operation parameter at time t - 2.
2. The dynamic control optimization method for mobile games based on adaptive deep learning according to claim 1, characterized in that The specific content of S2 includes: S21. Initialize the extraction module of the spatio-temporal feature adaptive convolutional network, input the collected touch data into the input layer of the adaptive convolutional network, and extract the touch position, touch force, and trajectory dynamic features; S22. Based on the real-time touch force F strength (t) and the trajectory change rate R trajectory (t), dynamically adjust the size parameter of the convolution kernel: Among them, K adaptive_size (t) represents the dynamically adaptive parameter of the convolutional kernel size, α conv_scale represents the overall adjustment coefficient of the convolutional kernel, λ path_mod represents the path adaptive adjustment coefficient, σ strength_bias represents the intensity offset coefficient, η response_mod represents the response adjustment coefficient of the touch intensity, θ decay represents the exponential decay coefficient, β adapt_smooth represents the adaptive smoothing coefficient, δ strength_norm represents the touch intensity normalization coefficient; S23. Dynamically adjust the convolutional kernel shape parameter K adaptive_shape (t) according to the continuity and direction change of the operation trajectory, so that the convolutional kernel shape adapts to the change direction of the touch path; S24. When a change in the touch acceleration A feedback (t) is detected, dynamically adjust the sampling step size: Among them, T dynamic_interval (t) represents the dynamic adaptive parameter of the time sampling step, γ base_interval represents the basic sampling step coefficient, δ accel_mod represents the acceleration adjustment coefficient, ∈ log_smooth represents the logarithmic smoothing coefficient, ν strength_weight represents the touch intensity weight coefficient, κ decay_rate represents the exponential decay coefficient; S25. Normalize the spatio-temporal features obtained at different time sampling steps to generate a normalized spatio-temporal feature sequence; S26. Integrate the adjusted convolutional kernel size, shape parameters, and the normalized spatio-temporal feature sequence to generate an adaptive spatio-temporal feature sequence S sta (t).
3. A method for optimizing dynamic control of mobile games based on adaptive deep learning according to claim 1, characterized in that, The specific content of S5 includes: S51. Build a multi-dimensional feedback scoring mechanism, which generates a feedback score value F(t) in real time by calculating the weighted combination of the accuracy, response speed, and success rate indicators of the player's operations, and is used to quantify the current control effect score (t) S52. Define the feedback score change rate ΔF score (t) for monitoring the change trend of the feedback score and capturing the rising or falling rate of the score. The feedback score change rate ΔF score (t) identifies the change trend of the operation performance, so as to start the learning rate adjustment according to the change trend; S53. Design an adaptive learning rate adjustment function to non-linearly adjust the learning rate according to the change trend of the feedback score: Among them, η(t) represents the dynamic learning rate at time t, and η base represents the base learning rate, α adjust represents the learning rate adjustment amplitude coefficient, tanh represents the hyperbolic tangent function, β rate represents the scoring change rate adjustment coefficient, γ smooth represents the smoothing coefficient, δ stabilize represents the stability factor, F target represents the preset target score; S54. Introduce a quick adjustment mechanism. When the feedback score value F score (t) is lower than the set minimum score threshold F min , activate the adjustment of the learning rate increase to adapt to the changing requirements of the current operation performance. After the score returns to stability, automatically resume the normal adaptive adjustment mode; S55. Construct a score smoothing module to perform mean smoothing on the score within a certain time window to reduce the direct impact of score fluctuations on the learning rate; S56. During the entire feedback score generation and learning rate adjustment process, continuous adaptive optimization is achieved based on the smoothed dynamic learning rate η smooth (t).
4. A method for optimizing dynamic control of mobile games based on adaptive deep learning according to claim 1, characterized in that, The specific content of S6 includes: S61. Establish a comprehensive satisfaction evaluation model, and generate a comprehensive satisfaction score M by using the feedback score value of players, the degree of operation flow, and the key parameter of response delay to quantify the current operation experience. The comprehensive satisfaction score M satisfaction (t) is calculated based on a multi-dimensional scoring function, and the weight of each scoring dimension is set according to operation requirements; satisfaction (t) is calculated based on a multi-dimensional scoring function, and the weight of each scoring dimension is set according to operation requirements; S62. Design satisfaction change trend analysis mechanism, for the comprehensive satisfaction score M satisfaction at consecutive moments for dynamic monitoring to capture the change trend. By calculating the increment ΔM satisfaction (t) of the comprehensive satisfaction score at the current moment, if the score continues to decline, it is marked as a downward trend, otherwise it is marked as an upward trend. When it is detected that the downward trend continuously exceeds the set number threshold, it is determined that the player's satisfaction is declining, and the dissatisfaction operation detection is triggered; S63. Establish a dissatisfaction operation detection mechanism and set a minimum satisfaction threshold M min , which is used to identify whether the player's satisfaction with the current operation reaches the expectation. If the comprehensive satisfaction score M satisfaction (t) is lower than the minimum satisfaction threshold M min , mark the current operation as a dissatisfied operation and enter the low satisfaction state; S64. When a low satisfaction state with dissatisfaction operations is detected, activate the relearning mechanism and use an increased relearning rate η relearn Adjust the control parameter, the relearning rate η relearn Set according to the satisfaction gap and the change trend: Among them, η relearn (t) represents the re-learning rate at time t, η base represents the base learning rate, α adapt represents the learning rate increase coefficient, β trend represents the trend response coefficient, λ stability represents the stability adjustment coefficient, γ smooth represents the smoothing factor, δ balance represents the balance coefficient, M target represents the target satisfaction level; S65. During the re-learning process, continuously optimize the control parameters according to the real-time feedback, feedback score, and satisfaction trend; S66. When the satisfaction score returns to the set target satisfaction level M target or when several consecutive operations meet the steady state, gradually decrease the re-learning rate and resume the normal learning rate mode.
Citation Information
Patent Citations
Auxiliary control method and device for cloud game, storage medium and electronic equipment
CN117085314A
AI-based digital human generation method and digital human live broadcast system
CN118608663A