Large model parameter optimization and adaptive adjustment method based on reinforcement learning

Through enhanced learning optimization of large model parameters, the training oscillation and insufficient adaptability caused by improper learning rate in traditional methods are solved, and the efficient training and stable application of the model in complex environments is realized, and the training efficiency and actual performance of the model are improved.

CN120494025AInactive Publication Date: 2025-08-15天津仁爱学院
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510479716.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing large model parameter optimization method has too high learning rate in the early stage of training, which leads to oscillation, and the low learning rate in the later stage leads to slow training, and lacks adaptability to data distribution changes and task requirements, resulting in a degradation in model performance in practical applications.

Method used

The large model parameter optimization and adaptive adjustment method based on enhanced learning is adopted. By integrating the environment perception module, the enhancement learning module and the reward feedback mechanism, the learning rate and parameter update strategy are adjusted in real time, combined with the policy network and the value network, and the model parameters are optimized based on environmental feedback.

Benefits of technology

It improves training efficiency, shortens training time, enhances the model's adaptability and generalization ability in complex environments, and improves the prediction accuracy and stability of the model in actual tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494025A_ABST
    Figure CN120494025A_ABST
Patent Text Reader

Abstract

The invention provides a large model parameter optimization and self-adaptive adjustment method based on reinforcement learning, and relates to the technical field of optimization and self-adaptive adjustment of large model parameters based on reinforcement learning in machine learning, and the method comprises an overall architecture fusing a large model main body, a reinforcement learning module, an environment perception module and a reward feedback mechanism. The method is used for realizing dynamic optimization and adaptive adjustment of large model parameters, and can rapidly increase the learning rate according to environment feedback in the initial stage of large model training, so that the model parameters rapidly approach to the direction of an optimal solution, and the early-stage exploration time of training is greatly shortened. In the later stage of training, the learning rate can be accurately reduced, model oscillation is avoided, and it is ensured that the model is stably converged to a globally optimal solution. According to the dynamic adjustment mechanism, the number of iterations of training is effectively reduced, the training efficiency is greatly improved, a large number of computing resources and time cost can be saved, and for example, in large-scale image classification model training, the training time can be shortened by more than 30%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of optimizing and adaptively adjusting large model parameters based on reinforcement learning in machine learning, and more specifically, relates to a method for optimizing and adaptively adjusting large model parameters based on reinforcement learning. Background Art

[0002] In the development of artificial intelligence, large models have become a key technology. Currently, parameter optimization for large models primarily relies on traditional gradient descent methods and their improved algorithms (such as SGD, Adagrad, Adadelta, and Adam). These traditional methods have achieved some success in simple tasks and datasets. For example, SGD enabled rapid model convergence in early image recognition. However, as large models scale and their application scenarios become more complex, challenges arise. First, using a fixed learning rate strategy. While a high learning rate accelerates convergence in the early stages of training, it can easily lead to oscillations later on. A low learning rate slows training and consumes a lot of resources. Second, they lack the ability to adapt to dynamic runtime changes. In real-world applications, when data distribution shifts or task requirements change, model performance degrades due to the inability to adjust parameters in real time. This is particularly true for machine translation tasks in natural language processing when faced with new language expressions and vocabulary. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides a large model parameter optimization and adaptive adjustment method based on reinforcement learning to solve the above problems.

[0004] A large-scale model parameter optimization and adaptive adjustment method based on reinforcement learning includes an overall architecture integrating a large-scale model main body, a reinforcement learning module, an environmental perception module, and a reward feedback mechanism, which is used to achieve dynamic optimization and adaptive adjustment of large-scale model parameters.

[0005] Preferably, the environmental perception module collects characteristic information of input data in real time during the large model training and operation stage, and conducts a comprehensive evaluation of the environmental state in combination with the current operating state of the model, and quantifies it into a state vector input to the reinforcement learning module. The reinforcement learning module includes an intelligent agent, and the intelligent agent has a policy network and a value network. The policy network outputs actions representing the large model parameter adjustment operations based on the state vector input by the environmental perception module, and the value network evaluates the cumulative reward estimate of the intelligent agent's actions. The policy network of the intelligent agent adopts a deep neural network structure, and optimizes its own parameters through feedback from the learning environment to generate better actions.

[0006] Preferably, the value network of the intelligent agent is based on a deep neural network, and its output provides a reference for the decision-making of the policy network. The reward function of the reward feedback mechanism is designed based on indicators such as the reduction in model loss value, accuracy improvement, and model convergence speed, and the reward is back-propagated through the reinforcement learning algorithm to update the policy network and value network parameters. During the training phase, the reinforcement learning module adjusts the model parameters in real time according to environmental perception and reward feedback, including dynamically adjusting the learning rate and parameter update priority.

[0007] Preferably, in the application stage, the environmental perception module monitors changes in data and task requirements, and the reinforcement learning module adaptively adjusts the model parameters accordingly. The unique reinforcement learning algorithm is applied to design a special intelligent agent, action space and reward feedback mechanism, so that the reinforcement learning algorithm can accurately interact with the large model to explore and learn the optimal parameter adjustment strategy, customize the large model parameter optimization strategy according to the characteristics of the training stage and real-time performance feedback, break the traditional fixed pattern, dynamically adjust the learning rate and parameter update method, and give priority to updating key parameters.

[0008] Compared with the prior art, the present invention has the following beneficial effects: Improve training efficiency: Compared with the traditional parameter optimization method with a fixed learning rate, the present invention is based on a dynamic parameter adjustment strategy based on enhanced learning. In the early stages of large-scale model training, it can quickly increase the learning rate based on environmental feedback, so that the model parameters quickly approach the optimal solution, greatly shortening the early exploration time of training. In the later stages of training, the learning rate can be accurately reduced to avoid model oscillations and ensure that the model converges stably to the global optimal solution. This dynamic adjustment mechanism effectively reduces the number of training iterations, greatly improves training efficiency, and can save a lot of computing resources and time costs. For example, in large-scale image classification model training, the training time can be shortened by more than 30%.

[0009] Enhanced model adaptability: Previous technologies struggled to cope with the changing data distribution and diverse task requirements in real-world application scenarios. This invention uses an environmental perception module to monitor data and task characteristics in real time, while an enhanced learning module adjusts large model parameters in response to environmental changes. Taking intelligent recommendation systems as an example, when user behavior patterns or product attributes change, the model can quickly adapt to maintain the accuracy and effectiveness of recommendations. This significantly improves the model's generalization and adaptability in complex and changing real-world environments, enabling large models to operate stably and efficiently in a wider range of fields and scenarios.

[0010] Optimizing Model Performance: This invention employs a multi-dimensional reward-feedback mechanism that rewards the agent based on key performance indicators, such as model loss and accuracy. This enables the agent to learn parameter adjustment strategies that effectively optimize model performance, effectively improving the model's predictive accuracy and stability. In natural language processing text classification tasks, models optimized using this method can achieve 10%-15% higher classification accuracy than traditional methods, significantly enhancing the practical application value of large models. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is the overall architecture diagram of the present invention; Figure 2 It is the internal structure diagram of the reinforcement learning module of the present invention. DETAILED DESCRIPTION

[0012] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0013] See also Figure 1-Figure 2 The present invention provides a large model parameter optimization and adaptive adjustment method based on reinforcement learning, including an overall architecture that integrates a large model body, a reinforcement learning module, an environmental perception module, and a reward feedback mechanism, for realizing dynamic optimization and adaptive adjustment of large model parameters.

[0014] The environmental perception module collects feature information of input data in real time during the large model training and operation stages, and conducts a comprehensive assessment of the environmental status in combination with the current operating status of the model, quantifying it into a state vector and inputting it into the reinforcement learning module. The reinforcement learning module contains an intelligent agent, which has a policy network and a value network. The policy network outputs actions representing the large model parameter adjustment operations based on the state vector input by the environmental perception module. The value network evaluates the cumulative reward estimate of the intelligent agent's actions. The policy network of the intelligent agent adopts a deep neural network structure, and optimizes its own parameters through learning environment feedback to generate better actions.

[0015] The value network of the intelligent agent is based on a deep neural network, and its output provides a reference for the decision-making of the policy network. The reward feedback mechanism designs a reward function based on indicators such as the reduction in model loss value, accuracy improvement, and model convergence speed. The reward is backpropagated through the reinforcement learning algorithm to update the policy network and value network parameters. During the training phase, the reinforcement learning module adjusts the model parameters in real time based on environmental perception and reward feedback, including dynamic adjustment of the learning rate and parameter update priority.

[0016] During the application phase, the environmental perception module monitors changes in data and task requirements, and the reinforcement learning module adaptively adjusts model parameters accordingly. The unique reinforcement learning algorithm uses a specially designed intelligent agent, action space, and reward feedback mechanism, allowing the reinforcement learning algorithm to accurately interact with the large model to explore and learn the optimal parameter adjustment strategy. The large model parameter optimization strategy is customized according to the characteristics of the training phase and real-time performance feedback, breaking the traditional fixed pattern, dynamically adjusting the learning rate and parameter update method, and giving priority to updating key parameters.

[0017] This invention aims to address key issues in existing large-model parameter optimization techniques. First, addressing the drawbacks of the fixed learning rate strategy of traditional optimization algorithms, the present invention dynamically and intelligently adjusts the learning rate according to the different stages of large-model training. This ensures rapid convergence in the early stages of training and accurate approximation to the global optimal solution in the later stages, effectively improving training efficiency and quality and avoiding model oscillation or slow convergence caused by inappropriate learning rates.

[0018] Secondly, to address the changing data distribution and diverse task requirements faced by models in real-world application scenarios, this paper aims to leverage reinforcement learning mechanisms to enable large models to perceive environmental changes in real time and adaptively adjust parameters based on feedback. This allows large models to maintain excellent performance under varying real-world conditions, significantly improving their generalization and adaptability. This overcomes the performance degradation of traditional methods when faced with dynamic changes, and promotes the effective application of large models in more complex scenarios.

[0019] 1. Overview of the overall framework: This paper proposes a method for large-scale model parameter optimization and adaptive adjustment based on reinforcement learning. Its main architecture integrates a large-scale model, a reinforcement learning module, an environmental perception module, and a reward-feedback mechanism. The overall process involves continuously optimizing the large-scale model's parameters during training and practical application, leveraging the reinforcement learning module to adapt to different stages and complex and changing environments.

[0020] 2. Specific technical solutions: (1) Environmental Perception Module Data feature extraction: During the large-scale model training and operation phases, the environmental perception module collects feature information from input data in real time. For example, in natural language processing tasks, it analyzes and extracts information about the vocabulary distribution, grammatical structure, and semantic themes of the input text; in image recognition tasks, it extracts features such as image color, texture, and shape. This feature extraction allows for an accurate understanding of the data's characteristics, providing a basis for subsequent parameter adjustments.

[0021] Environmental state assessment: This comprehensively evaluates the model's environmental state by combining extracted data features with the model's current operational status, such as the model's loss, accuracy, and parameter updates. The environmental state is quantified as a series of state vectors, which serve as input to the reinforcement learning module, allowing the RL algorithm to understand the model's current status.

[0022] (2) Reinforcement Learning Module Agent Definition: In this invention, a reinforcement learning agent is responsible for determining the parameter adjustment strategy for the large model. The agent continuously learns and optimizes these parameter adjustment decisions by interacting with the environment (i.e., the system consisting of the large model and input data). The agent internally comprises a policy network and a value network.

[0023] The policy network outputs an action based on the state vector input from the environment perception module. This action represents a specific adjustment to the parameters of the larger model, such as the learning rate adjustment and the direction of parameter updates. The policy network uses a deep neural network structure and continuously learns from environmental feedback to optimize its parameters and generate more optimal actions.

[0024] The value network estimates the cumulative reward that the agent will receive over a period of time after taking a specific action. The value network, also based on a deep neural network, provides a reference for the policy network's decision-making, helping it determine the pros and cons of the current action.

[0025] Action Space Design: The action space encompasses all possible adjustments to large model parameters. For example, the learning rate can be increased or decreased within a certain range, the step size for parameter updates can be varied, and the update priority for different parameter groups can be set. By carefully designing the action space, the agent can flexibly explore various parameter adjustment strategies.

[0026] (3) Reward and feedback mechanism Reward Function Design: The reward function is key to guiding the agent to learn the correct parameter adjustment strategy. The reward function is designed based on multiple metrics, including the magnitude of the decrease in model loss, the improvement in accuracy, and the speed of model convergence. For example, when the model loss decreases rapidly and the accuracy improves, the agent is given a large positive reward; if the model oscillates or performance degrades, a negative reward is given. The reward function is designed to help the agent understand which parameter adjustment strategies will effectively improve model performance.

[0027] Reward Propagation and Learning: After taking an action, the agent receives corresponding rewards based on environmental feedback. These rewards are not only used to evaluate the effectiveness of the current action but are also back-propagated through the reinforcement learning algorithm to update the parameters of the policy and value networks. Through continuous trial and error and learning, the agent gradually masters the parameter adjustment strategy that optimizes model performance.

[0028] (IV) Implementation of parameter adjustment Parameter Adjustment During Training: During large-scale model training, the reinforcement learning module adjusts model parameters in real time based on environmental perception and reward feedback. For example, if the model converges slowly in the early stages of training, the agent, through the policy network, increases the learning rate. Later in training, as convergence approaches, the learning rate is reduced to prevent oscillation. Furthermore, the priority of parameter updates is dynamically adjusted based on the update progress of parameters at different layers, ensuring more timely and effective updates of important parameters.

[0029] Adaptive adjustment during the application phase: After a large model is deployed in a real-world application scenario, the environmental awareness module continuously monitors changes in data and task requirements. Once environmental changes are detected, the reinforcement learning module quickly responds by adjusting model parameters to adapt to the new environment. For example, in intelligent customer service scenarios, when encountering new question types or changes in user language style, the model can automatically adjust parameters to improve the accuracy and relevance of responses.

[0030] Dynamic parameter adjustment based on reinforcement learning: Breaking the traditional fixed parameter adjustment mode, using the intelligent decision-making mechanism of reinforcement learning to achieve dynamic adjustment of large model parameters during training and application, which is the core innovation of this invention.

[0031] Multi-dimensional environmental perception and assessment: Build a comprehensive environmental perception module to evaluate the model's environment from multiple dimensions, including data characteristics and model operation status, providing an accurate basis for parameter adjustment.

[0032] Customized reward feedback mechanism: Design a reward function based on model performance indicators to effectively guide the reinforcement learning agent to learn the optimal parameter adjustment strategy.

[0033] (1) Data preparation Data Collection: We collect a wide range of text data, including news, literature, academic papers, and social media comments, to build a large and rich corpus. To ensure data diversity and representativeness, we source data from diverse fields, languages, and timeframes. For example, we not only collect news reports from mainstream media, but also include specialized literature in niche fields and everyday online conversations on social media platforms.

[0034] Data preprocessing: Word Segmentation: Use professional word segmentation tools, such as deep learning-based word segmentation models or classic word segmentation algorithms (such as Jieba Segmentation), to split continuous text into individual word units. For English text, simple word segmentation can be achieved by means of spaces and punctuation marks; for languages like Chinese without natural delimiters, more complex algorithms are required to accurately identify word boundaries.

[0035] Part-of-Speech Tagging: Use part-of-speech tagging tools to tag the part of speech of each segmented word, such as nouns, verbs, adjectives, etc. This helps to understand the grammatical functions of words in a sentence and provides a basis for subsequent semantic analysis.

[0036] Stop Word Removal: Establish a stop word list and remove those words that frequently appear in the text but contribute little to semantic expression, such as "de", "shi", "zai", etc. This can reduce the amount of data, improve the processing efficiency of the model, and avoid noise interference.

[0037] Vector Representation: Convert the processed text into a vector form that can be understood by a computer. Common methods include the Bag of Words model and word embedding techniques such as Word2Vec and GloVe. These vectors not only contain the semantic information of words but also enable the measurement of the similarity between words through vector operations.

[0038] (2) Initialization of the Environmental Perception Module Setting of Data Feature Extraction Parameters: Word Frequency Statistics: Set the window size and frequency threshold for word frequency statistics. By counting the number of times each word appears within a certain window, obtain the lexical distribution characteristics of the text. This helps to understand the theme and common words of the text.

[0039] Analysis of Word Vector Dimensions: Determine the method for analyzing word vector dimensions, such as principal component analysis (PCA) or singular value decomposition (SVD), to extract the main characteristic components of word vectors, reduce dimensional redundancy, and retain key semantic information at the same time.

[0040] Setting of Model State Evaluation Metrics: Perplexity Calculation: Define the calculation method and update frequency of perplexity. Perplexity is an important indicator for measuring the prediction ability of a language model, which reflects the fitting degree of the model to the text. A lower perplexity indicates that the model has a stronger ability to understand and predict the text.

[0041] Accuracy Evaluation: Determine the specific method for calculating accuracy on the validation set. For example, in a text classification task, accuracy refers to the proportion of the number of samples correctly classified by the model to the total number of samples. At the same time, set the improvement threshold for accuracy. When the model accuracy reaches or exceeds this threshold, it is considered that the model performance has been significantly improved.

[0042] (3) Construction of reinforcement learning module Intelligent agent structure construction: The policy network uses a multi-layer perceptron (MLP) architecture, consisting of multiple hidden layers, each composed of multiple neurons. The input layer receives the state vector output by the environment perception module. After nonlinear transformations in the hidden layers, the output layer outputs the parameter adjustment action. To improve the model's generalization and training efficiency, a dropout layer is added between hidden layers to randomly drop neurons and prevent overfitting.

[0043] The value network is also built on an MLP. Its input, like the policy network, is the environment state vector. The value network uses feedback from the environment to estimate the cumulative reward the agent will receive over a period of time after taking a specific action. To speed up training, pre-training can be used to initialize the value network's parameters, bringing it close to the optimal solution.

[0044] Parameter initialization: Use a random initialization method (such as Xavier initialization or Kaiming initialization) to initialize the parameters of the policy network and value network to ensure that the initial parameter distribution is within a reasonable range, which is conducive to model convergence. At the same time, set optimizer parameters such as learning rate and momentum, which will affect the training speed and convergence of the model.

[0045] (IV) Training process Environmental Perception and State Input: During training, the environmental perception module operates in real time. For each batch of input text data, it first extracts data features, such as word frequency and principal components of word vectors. This is combined with the current model's perplexity and accuracy on the training and validation sets to generate a complete state vector. This state vector serves as input to the reinforcement learning agent, allowing it to understand the current environment in which the model operates.

[0046] Action Decisions and Parameter Adjustment: The agent adjusts the parameters of the larger model based on the actions output by the policy network. For example, when the policy network outputs an action to increase the learning rate, the agent increases the current learning rate by a certain percentage (e.g., 10%). When the policy network outputs an action to adjust the parameter update step size, the agent adjusts the parameter update step size based on the step size specified in the action. Furthermore, the agent can adjust the update priority of different layer parameters based on the actions, increasing the frequency and magnitude of parameter updates for layers that are more relevant to the current task.

[0047] Reward Feedback and Learning: The agent is rewarded based on the model's performance during training. If the model's perplexity decreases and its accuracy improves on the validation set, the agent receives a positive reward, the value of which varies depending on the magnitude of the decrease or increase. If the model overfits or experiences performance degradation, a negative reward is given. Based on reward feedback, the agent uses a reinforcement learning algorithm (such as Proximal Policy Optimization (PPO) or Deep Q-Network (DQN)) to update the parameters of its policy and value networks, enabling the agent to gradually learn a more optimal parameter adjustment strategy.

[0048] (V) Application stage Real-time context monitoring: When the model is deployed in real-world applications, such as intelligent writing assistance tools, the context perception module continuously monitors changes in the characteristics of user input text. By analyzing vocabulary usage, sentence structure, and writing topics, the module determines whether the user's writing style and needs have changed. For example, if a user frequently uses professional terminology over a period of time, this indicates that the user may be writing in a specialized field. The model needs to adjust parameters accordingly to better understand and support this writing style.

[0049] Adaptive parameter adjustment: Once the context perception module detects changes in a user's writing style or needs, it immediately passes the relevant information to the reinforcement learning module. Based on this information, the reinforcement learning module uses the policy network to adjust the parameters of the large model. For example, it can adjust the word vector representation of the language model to better adapt to the new vocabulary distribution, or adjust the model's attention mechanism parameters to better capture the semantic relationships in the user's text. Through this real-time adaptive parameter adjustment, the model can provide users with more accurate and tailored writing suggestions, such as grammatical correction, vocabulary recommendations, and sentence polishing.

[0050] 2. Image Recognition Task Example (1) Dataset preparation Data Collection: Collect large-scale image datasets covering a wide range of categories and scenarios. For example, the popular image recognition datasets CIFAR-10 and ImageNet. CIFAR-10 contains 60,000 color images from 10 different categories, while ImageNet is a large-scale dataset with over 14 million images covering over 20,000 categories. In addition to these public datasets, image data from specific fields, such as medical imaging and industrial inspection images, can also be collected based on specific application scenarios.

[0051] Data preprocessing: Normalization: Normalize the pixel values of the image, mapping the pixel value range from [0, 255] to [0, 1] or [-1, 1], so that the data distribution of different images is consistent, which is conducive to model training and convergence.

[0052] Cropping and Scaling: Images are cropped and scaled according to the model input requirements. For example, images can be cropped to a fixed size (e.g., 224x224 pixels) to accommodate the input size of a convolutional neural network. Furthermore, to increase data diversity, data augmentation techniques such as random cropping, flipping, and rotation can be used to expand the dataset and improve the model's generalization capabilities.

[0053] Image enhancement: Image enhancement algorithms, such as histogram equalization, contrast enhancement, and Gaussian blur, are used to improve image quality and visual quality. These operations can highlight key features in the image, reduce noise interference, and improve the model's ability to recognize images.

[0054] (2) Environment Perception and Enhanced Learning Module Settings Environmental perception module: Image feature extraction parameter settings: Color Histogram: Set the number of bins and calculation method for the color histogram. This tool uses statistics to determine the distribution of different colors within an image and to determine its color characteristics. The color histogram reflects the overall hue and color distribution of an image, making it useful for distinguishing between different image categories.

[0055] HOG features: Determine the parameters of the Histogram of Oriented Gradients (HOG), such as cell size, block size, and number of gradient directions. HOG features describe the edge and shape information of an image by calculating the gradient direction and magnitude of a local area of the image. They are widely used in object detection and image classification tasks.

[0056] Model state evaluation indicator setting: Accuracy calculation: Define the method for calculating accuracy on the training and validation sets. For image recognition tasks, accuracy refers to the ratio of images correctly classified by the model to the total number of images. Also, set the frequency for monitoring accuracy and the target for improvement to ensure timely evaluation of model performance changes.

[0057] Recall evaluation: In addition to precision, recall is also a metric to consider, especially in applications where a high miss detection rate is critical. Recall refers to the ratio of correctly classified samples to the actual number of samples. By comprehensively evaluating precision and recall, we can gain a more comprehensive understanding of model performance.

[0058] Reinforcement Learning Module: Agent Structure: A convolutional neural network (CNN) serves as the foundation for the policy network and value network. CNNs automatically extract local features from images and efficiently process image data through a combination of convolutional, pooling, and fully connected layers. In the policy network, the input layer receives the state vector (containing image features and model state information) output by the environment perception module. After multiple convolutional and fully connected layers, it outputs parameter-adjusted actions. The value network has a similar structure, but outputs action value estimates.

[0059] Parameter initialization and optimization: Use appropriate initialization methods (such as Kaiming initialization) to initialize CNN parameters, ensuring a well-distributed initial parameter distribution. During training, use optimization algorithms (such as the Adam optimizer) to update the parameters of the policy and value networks. A learning rate decay strategy is also implemented, gradually reducing the learning rate as training progresses to ensure convergence stability of the model.

[0060] (3) Training phase Environmental Perception and State Feedback: When training an image recognition model, the Environmental Perception module extracts features from each batch of input images in real time. Combined with the model's current state information, such as accuracy and recall, this module generates a state vector that is fed into the reinforcement learning module. For example, if the model's accuracy for certain image categories is low during training, the Environmental Perception module transmits this information as part of the state vector to the agent, allowing it to understand any model issues.

[0061] Action Decisions and Parameter Adjustment: The agent adjusts the parameters of the image recognition model based on the actions output by the policy network. These actions can include adjusting the size, number, and step length of convolution kernels, changing the number and connection structure of neurons in the fully connected layer, and adjusting the learning rate and regularization parameters. For example, if the policy network determines that the model is under-retrieving certain detailed features, it may decide to increase the number of convolution kernels in the convolution layer to enhance the model's ability to perceive details. If the model shows signs of overfitting, the agent adjusts the regularization parameters to improve the model's generalization ability.

[0062] Reward feedback and learning optimization: The agent is rewarded based on the model's performance on the training and validation sets. If both the model's precision and recall improve, a positive reward is given; if performance degrades or overfits, a negative reward is given. The agent uses a reinforcement learning algorithm (such as A2C or DDPG) to update the parameters of the policy and value networks based on reward feedback, continuously optimizing the parameter adjustment strategy to achieve gradually better performance during training.

[0063] (IV) Practical Application Real-time environmental monitoring and parameter adjustment: In real-world image recognition applications, such as face recognition systems in security surveillance, the environmental perception module continuously monitors factors such as input image quality, lighting conditions, and posture changes. When a change in lighting conditions is detected, the environmental perception module transmits this information to the reinforcement learning module. Based on the policy network's decisions, the reinforcement learning module adjusts the parameters of the face recognition model. For example, this involves adjusting the lighting compensation parameters in image preprocessing or the parameters of the lighting-related feature extraction layer in the convolutional neural network to adapt to the new lighting conditions, ensuring accurate recognition of the target person in varying lighting environments.

[0064] Multi-scenario adaptability optimization: The model needs to have different adaptability for different application scenarios, such as access control systems, video surveillance, and mobile device recognition. The environmental perception module provides decision-making basis for the reinforcement learning module by monitoring scene-related information, such as the frequency of people entering and exiting the access control system and the scene complexity in video surveillance. The reinforcement learning module dynamically adjusts model parameters based on this information, ensuring that the model maintains good performance in different scenarios. For example, in an access control system with frequent entry and exit, to improve recognition speed, the model can appropriately simplify the complex feature extraction process while ensuring a certain recognition accuracy rate. In complex video surveillance scenarios, the model needs to strengthen its robustness to multiple interference factors and improve its face recognition capabilities in different postures and occlusion situations by adjusting parameters.

[0065] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.

Claims

1. A large model parameter optimization and adaptive adjustment method based on reinforcement learning, characterized by It includes an overall architecture that integrates the large model body, reinforcement learning module, environmental perception module and reward feedback mechanism, which is used to achieve dynamic optimization and adaptive adjustment of large model parameters.

2. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 1 is characterized in that: The environmental perception module collects the characteristic information of input data in real time during the large model training and operation stages, and comprehensively evaluates the environmental status based on the current operating status of the model, quantifying it into a state vector and inputting it into the reinforcement learning module.

3. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 1 is characterized in that: The reinforcement learning module includes an intelligent agent, which has a policy network and a value network. The policy network outputs actions representing the large model parameter adjustment operations based on the state vector input by the environment perception module, and the value network evaluates the cumulative reward estimate of the intelligent agent's actions.

4. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 3 is characterized in that: The agent's policy network adopts a deep neural network structure, which optimizes its own parameters through learning environment feedback to generate better actions.

5. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 3 is characterized in that: The agent's value network is based on a deep neural network, and its output provides a reference for the policy network's decision-making.

6. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 1 is characterized in that: The reward feedback mechanism designs a reward function based on indicators such as the reduction in model loss value, accuracy improvement, and model convergence speed, and backpropagates the reward through the reinforcement learning algorithm to update the policy network and value network parameters.

7. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 1 is characterized in that: During the training phase, the reinforcement learning module adjusts model parameters in real time based on environmental perception and reward feedback, including dynamic adjustment of learning rate and parameter update priority.

8. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 1 is characterized in that: During the application phase, the environmental perception module monitors changes in data and task requirements, and the reinforcement learning module adaptively adjusts the model parameters accordingly.

9. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 1 is characterized in that: The unique reinforcement learning algorithm uses a specially designed intelligent agent, action space, and reward feedback mechanism, which enables the reinforcement learning algorithm to accurately interact with large models to explore and learn the optimal parameter adjustment strategy.

10. The large model parameter optimization and adaptive adjustment method based on reinforcement learning according to claim 1 is characterized in that: The customized large model parameter optimization strategy breaks the traditional fixed pattern based on the characteristics of the training phase and real-time performance feedback, dynamically adjusts the learning rate and parameter update method, and prioritizes updating key parameters.

Citation Information

Cited By

  • Reinforced learning training parameter automatic tuning system and method based on TensorBoard log driving and large model service

    CN121303240A

  • A reinforcement learning training parameter automatic tuning system and method based on TensorBoard log driving and large model service

    CN121303240B