Large model reasoning optimization method based on GRPO algorithm
The GRPO algorithm optimizes the large-model inference process and reduces the consumption of computing resources through the relative reward mechanism, solving the problem of high computing resources in the inference process of large-scale deep learning models, achieving more efficient inference speed and stability, and supporting the compatibility and autonomous inference capabilities of multiple model architectures.
Patent Information
- Application Number
- CN202510250008.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-25
AI Technical Summary
The existing large-scale deep learning models consume high computing resources during the inference process, resulting in delayed system response and affecting user experience, especially in scenarios with high real-time requirements, which become a bottleneck in system deployment and operation.
The GRPO algorithm is used to sample multiple candidate actions in each state, calculate their relative performance differences and generate relative reward signals, update the policy network, abandon the dependence on the value network in traditional reinforcement learning, reduce computing resource consumption, and optimize the inference process.
Significantly reduce computing resource consumption, improve inference speed and efficiency, support a variety of model architectures, enhance training stability and flexibility, reduce hardware requirements, and achieve independent inference capabilities.
Abstract
Description
Technical Field
[0001] The present invention is applied to the field of artificial intelligence, and specifically relates to a method for optimizing large model inference based on the GRPO algorithm. Background Art
[0002] Deep learning models, especially those based on the Transformer architecture (such as BERT, GPT, etc.), have shown excellent capabilities in processing natural language, computer vision, and other tasks. These models are trained on large-scale datasets and can capture complex patterns in the data for inference. However, these models usually contain a large number of parameters and computational amounts, so they consume a large amount of computational resources and time during inference. Especially when dealing with large-scale data, the inference speed and computational overhead become more prominent.
[0003] On the one hand is the scale of parameters, and on the other hand is the scale of context. The scale of model parameters used in mainstream large model services is at least in the order of hundreds of billions, and the most advanced ones reach the scale of trillions of parameters, which is a very large scale of parameters. The scale of context supported by typical large models ranges from the 2K scale of GPT3 to the 1 million context parameter scale of gamaner 1.5 released at the beginning of 2024, and basically grows by 2 - 3 orders of magnitude in about one year. Therefore, the growth of the context scale is also very significant. The growth of the parameter scale and the context scale will very directly bring huge increases in the video memory overhead, cache overhead, and computational overhead during the inference of large model deployment. This is a very direct challenge to the performance in terms of inference.
[0004] The inference of large models with architectures such as BERT and GPT usually needs to be carried out on hardware such as GPUs or TPUs, which has huge demands for computational resources and may lead to system response delays, affecting the user experience. This has become a bottleneck in system deployment and operation in many scenarios with high real-time requirements. Therefore, inference optimization technologies have emerged, aiming to optimize the inference process of these models by reducing the computational amount and improving the computational efficiency. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method for optimizing large model inference based on the GRPO algorithm in view of the deficiencies of the prior art.
[0006] To solve the above technical problem, a method for optimizing large model inference based on the GRPO algorithm of the present invention includes the following steps:
[0007] Action Sampling: At each state, sample multiple candidate actions from the policy network, where the candidate actions include actions sampled from the action distribution output by the policy network or action variants generated based on the current policy through a perturbation mechanism;
[0008] Relative reward calculation: Calculate the performance difference of each candidate action relative to other candidate actions to generate a relative reward signal, where the relative reward signal is based on the relative effect between candidate actions rather than absolute value;
[0009] Policy update: Update the policy network according to the relative reward signal, where the policy update is based on the relative performance of candidate actions rather than individual action values;
[0010] Enhanced stability: Through the intra-group relative reward mechanism, avoid the instability caused by inaccurate estimation of the value network in traditional reinforcement learning methods, and reduce the problems of gradient vanishing or explosion;
[0011] Inference optimization: On the premise of keeping the model parameter structure unchanged, optimize the inference process of the large model through the GRPO algorithm, reduce the consumption of computing resources, and improve the inference speed and efficiency.
[0012] As a possible implementation, further, the GRPO algorithm reduces the computing resource requirements in the following ways:
[0013] Abandon the practice of relying on the critic model to estimate the value of each action in traditional reinforcement learning, and instead adopt the intra-group relative reward mechanism;
[0014] By sampling multiple candidate actions and calculating their relative performance differences, reduce the dependence on the value network and reduce the computational complexity.
[0015] As a possible implementation, further, the GRPO algorithm is applicable to various large model inference scenarios, including but not limited to:
[0016] Real-time data processing and decision-making in autonomous driving systems;
[0017] Fast diagnosis and image analysis in medical image processing;
[0018] Low-latency data processing in Internet of Things devices;
[0019] Large-scale concurrent request processing in cloud computing and edge computing.
[0020] As a possible implementation, further, the GRPO algorithm improves the inference efficiency in the following ways:
[0021] Implement the training and inference of large language models on consumer-grade GPUs;
[0022] Compared with Hugging Face+FA2, reduce the VRAM occupancy by 80%;
[0023] Support the inference optimization of models with a maximum number of parameters of 15 billion.
[0024] As a possible implementation, further, the GRPO algorithm enhances performance in the following ways:
[0025] Through reinforcement learning, the model extends the thinking time and re-evaluates the initial method without manual intervention, achieving an improvement in autonomous reasoning;
[0026] Reproduce the performance enhancement of DeepSeek R1-Zero on consumer-grade GPUs, reducing hardware requirements.
[0027] As a possible implementation, further, the GRPO algorithm improves training stability in the following ways:
[0028] Through the intra-group relative reward mechanism, avoid the training difficulties caused by sparse reward signals in traditional reinforcement learning methods;
[0029] In the case of complex tasks and long time delays, accelerate model convergence through the relative reward mechanism.
[0030] As a possible implementation, further, the GRPO algorithm improves the flexibility of the inference framework in the following ways:
[0031] Support loading the LoRA model in vLLM by parsing the state dictionary instead of loading from disk, increasing the GRPO training speed by 1.5 times;
[0032] Add a batch generation function to reduce the random VRAM peak during batch generation.
[0033] A large model inference optimization system based on the GRPO algorithm, including:
[0034] Action sampling module: used to sample multiple candidate actions from the policy network in each state;
[0035] Relative reward calculation module: used to calculate the performance difference of each candidate action relative to other candidate actions and generate a relative reward signal;
[0036] Policy update module: used to update the policy network according to the relative reward signal;
[0037] Inference optimization module: used to optimize the inference process of the large model through the GRPO algorithm while keeping the model parameter structure unchanged, reducing computational resource consumption, and improving inference speed and efficiency.
[0038] As a possible implementation, further, it also includes:
[0039] Model compatibility module: used to support the compatibility of multiple model architectures, including but not limited to Llama, Phi, Mistral, Qwen;
[0040] Training stability module: Used to avoid the training difficulties caused by sparse reward signals in traditional reinforcement learning methods through the intra-group relative reward mechanism.
[0041] As a possible implementation, further, it also includes:
[0042] Inference efficiency improvement module: Used to implement the training and inference of large language models on consumer-grade GPUs, reducing VRAM occupancy;
[0043] Throughput optimization module: Used to improve throughput and reduce VRAM consumption by integrating vLLM.
[0044] The present invention adopts the above technical solutions and has the following beneficial effects:
[0045] Significantly reduce computing resource consumption: Through the intra-group relative reward mechanism of the GRPO algorithm, the practice of relying on the critic model in traditional reinforcement learning is abandoned, reducing the occupancy of memory and computing resources. Compared with HuggingFace+FA2, the VRAM occupancy is reduced by 80%, enabling the training and inference of large language models on consumer-grade GPUs (such as 7GB VRAM).
[0046] Improve inference efficiency and speed: The GRPO algorithm significantly improves the inference speed and efficiency of the model by optimizing the inference process, especially performing outstandingly in high-concurrency and low-latency application scenarios. On 1 A100 40GB GPU, using the Llama 3.23B Instruct model with dynamic 4-bit quantization, a throughput of up to 4000 tokens / s is expected.
[0047] Enhance model training stability: Through the intra-group relative reward mechanism, the GRPO algorithm avoids the training difficulties caused by sparse reward signals in traditional reinforcement learning methods, reduces the problems of gradient vanishing or explosion, and improves the stability and convergence speed of the training process.
[0048] Support multiple model architectures and compatibility: The GRPO algorithm supports multiple large model architectures (such as Llama, Phi, Mistral, Qwen, etc.), and is compatible with full-scale fine-tuning, QLoRA, and LoRA technologies, further improving the inference ability and application flexibility of the model.
[0049] Achieving the "Eureka Moment": The GRPO algorithm enables the model to extend the thinking time and re-evaluate the initial approach without manual intervention through reinforcement learning, achieving an improvement in autonomous reasoning. The "Eureka moment" of DeepSeek R1-Zero can be reproduced on consumer-grade GPUs with only 7GB of VRAM, significantly reducing the hardware requirements. Detailed implementation manners
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.
[0051] Example 1
[0052] A large model inference optimization method based on the GRPO algorithm includes the following steps:
[0053] Action sampling: At each state, sample multiple candidate actions from the policy network, where the candidate actions include actions sampled from the action distribution output by the policy network or action variants generated by a perturbation mechanism based on the current policy.
[0054] Relative reward calculation: Calculate the performance difference of each candidate action relative to other candidate actions to generate a relative reward signal, where the relative reward signal is based on the relative effects between candidate actions rather than absolute values.
[0055] Policy update: Update the policy network according to the relative reward signal, where the policy update is based on the relative performance of candidate actions rather than individual action values.
[0056] Stability enhancement: Through the intra-group relative reward mechanism, avoid the instability caused by inaccurate value network estimation in traditional reinforcement learning methods, and reduce the problems of gradient disappearance or explosion.
[0057] Inference optimization: On the premise of keeping the model parameter structure unchanged, optimize the inference process of the large model through the GRPO algorithm, reduce the consumption of computing resources, and improve the inference speed and efficiency.
[0058] The GRPO algorithm reduces the computing resource requirements in the following ways:
[0059] Abandon the practice of relying on a critic model to estimate the value of each action in traditional reinforcement learning, and instead adopt the intra-group relative reward mechanism.
[0060] By sampling multiple candidate actions and calculating their relative performance differences, reduce the dependence on the value network and lower the computational complexity.
[0061] The GRPO algorithm is applicable to various large model inference scenarios, including but not limited to:
[0062] Real-time data processing and decision-making in autonomous driving systems;
[0063] Fast diagnosis and image analysis in medical image processing;
[0064] Low-latency data processing in Internet of Things devices;
[0065] Processing of large-scale concurrent requests in cloud computing and edge computing.
[0066] The GRPO algorithm improves the inference efficiency in the following ways:
[0067] Implement training and inference of large language models on consumer GPUs;
[0068] Compared with Hugging Face+FA2, it reduces VRAM occupancy by 80%;
[0069] Supports inference optimization for models with a maximum of 15 billion parameters.
[0070] The GRPO algorithm achieves performance enhancement in the following ways:
[0071] Through reinforcement learning, the model can extend the thinking time and re-evaluate the initial method without manual intervention to achieve an improvement in autonomous inference;
[0072] Replicate the performance enhancement of DeepSeek R1-Zero on consumer GPUs and reduce hardware requirements.
[0073] As a possible implementation, further, the GRPO algorithm improves the training stability in the following ways:
[0074] Through the intra-group relative reward mechanism, avoid the training difficulties caused by sparse reward signals in traditional reinforcement learning methods;
[0075] In the case of complex tasks and long time delays, accelerate model convergence through the relative reward mechanism.
[0076] The GRPO algorithm improves the flexibility of the inference framework in the following ways:
[0077] Supports loading the LoRA model in vLLM by parsing the state dictionary instead of loading from disk, increasing the GRPO training speed by 1.5 times;
[0078] Add a batch generation function to reduce the random video memory peak during batch generation.
[0079] A large model inference optimization system based on the GRPO algorithm, including:
[0080] Action sampling module: used to sample multiple candidate actions from the policy network in each state;
[0081] Relative reward calculation module: used to calculate the performance difference of each candidate action relative to other candidate actions and generate a relative reward signal;
[0082] Policy update module: used to update the policy network according to the relative reward signal;
[0083] Inference optimization module: used to optimize the inference process of the large model through the GRPO algorithm while keeping the model parameter structure unchanged, reducing computational resource consumption, and improving inference speed and efficiency.
[0084] Model compatibility module: used to support the compatibility of multiple model architectures, including but not limited to Llama, Phi, Mistral, Qwen;
[0085] Training stability module: used to avoid the training difficulties caused by sparse reward signals in traditional reinforcement learning methods through the intra-group relative reward mechanism.
[0086] Inference efficiency improvement module: used to implement the training and inference of large language models on consumer-grade GPUs, reducing VRAM occupancy;
[0087] Throughput optimization module: used to improve throughput and reduce VRAM consumption by integrating vLLM.
[0088] Embodiment 2
[0089] An inference framework for optimizing large models based on the GRPO algorithm, and the specific implementation process is as follows:
[0090] Action sampling: At each state, sample multiple candidate actions from the policy network. These actions can be samples from the action distribution output by the policy network or generated based on the current policy through certain perturbation mechanisms.
[0091] Relative reward calculation: For each sampled candidate action, calculate its performance difference relative to other actions. The reward signal here is not just a simple scalar value, but is generated by comparing the relative effects between actions. For example, use the standard reinforcement learning reward calculation method, but the reward is judged only based on the relative effect, rather than the absolute value of each action.
[0092] Policy update: Update the policy network according to the relative reward. Since it does not rely on the absolute value evaluation of the critic model, the direction of policy update mainly depends on the relative performance of each group of candidate actions, rather than the individual action value.
[0093] Enhanced Stability: Through the mechanism of relative rewards within the group, the instability caused by inaccurate estimation of the value network in traditional methods can be effectively avoided during the training process. Additionally, since the update process depends on relative values rather than absolute values, it enables the training process to adaptively adjust the strategy, reducing the problems of gradient vanishing or explosion that may occur when overly relying on the critic network.
[0094] Comparison between GRPO and Traditional Reinforcement Learning Methods
[0095] Different from traditional reinforcement learning methods (such as A3C, PPO, etc.), GRPO avoids relying on a separate value model for the valuation of each action. Instead, it optimizes based on the relative performance of actions. This relative reward mechanism can effectively overcome the following problems that occur in traditional methods:
[0096] Computational Burden of the Critic Model: Traditional reinforcement learning methods usually require training a critic network to estimate the value of each action, which can lead to a huge computational burden, especially when dealing with large-scale state spaces and action spaces. GRPO, through the method of relative rewards within the group, eliminates the need for a value network, thus greatly reducing the computational pressure.
[0097] Scarcity of Reward Signals: In many practical tasks, especially in the training of language models, reward signals are often very sparse, making it difficult for the model to learn effective strategies from them. GRPO, through the relative reward mechanism, enables the performance of each action to receive relative feedback, making the training process more efficient, especially in complex tasks and cases with long time delays.
[0098] Convergence Speed Problem: Traditional reinforcement learning methods usually require a long time for policy optimization, and the training process may have instability problems. GRPO, through relative reward optimization, makes the policy update more stable, helping to accelerate the convergence process.
[0099] The innovation of the GRPO (Group Relative Policy Optimization) method lies in abandoning the practice of relying on a critic model to estimate the value of each action in traditional reinforcement learning and instead adopting a mechanism of intra-group relative rewards. Specifically, GRPO samples multiple possible actions at each state and adjusts the policy by comparing the performance of these actions in terms of relative rewards. In this way, the optimization of the model no longer depends on a single value network to predict the absolute value of each action, but is based on the relative merits among a set of candidate actions for feedback. This method can not only reduce the computational complexity, but also improve the stability and convergence speed of the optimization process. The core implementation idea of GRPO is to obtain a relative reward signal by sampling a set of candidate actions and calculating the performance differences among them under the current policy. These candidate actions can be different actions sampled from the policy model or action variants generated by perturbing the existing policy. The relative reward of each candidate action reflects the difference in the effect of this action relative to other actions in a given state, rather than its absolute value. This mechanism makes the training process no longer rely on the traditional value network (critic network), so the computational resource requirements can be significantly reduced in practical applications.
[0100] Introduction of the GRPO algorithm: The present invention introduces the GRPO algorithm in a large model training and inference platform, allowing users to train their own R1 inference models. The GRPO algorithm enables the model to learn autonomously through reinforcement learning without manual intervention, and optimizes the policy through intra-group relative rewards, reducing the memory and computational burdens brought by the critic model in traditional reinforcement learning.
[0101] Improvement in VRAM efficiency: By optimizing the GRPO process, the present invention significantly reduces the VRAM occupancy during the inference process. Compared with Hugging Face + FA2, the optimization method of the present invention reduces the VRAM occupancy by 80%, and the "epiphany moment" of DeepSeek R1-Zero can be reproduced on a consumer-grade GPU that only requires 7GB of VRAM. This breakthrough enables large-scale language model training and inference even in low-resource environments.
[0102] Model compatibility: The technical support of the present invention can convert models with a maximum number of parameters of 15 billion (such as Llama 3.1 (8B), Phi-4 (14B), Mistral (7B), Qwen2.5 (7B)) into inference models, and supports the compatibility of multiple model architectures to meet the needs of different fields.
[0103] Chain of Thought "Insight": "Insight Moment" Principle: The GRPO algorithm enables the model to learn to extend the thinking time and re-evaluate the initial method without manual intervention through reinforcement learning, achieving an improvement in autonomous reasoning. By optimizing the reward mechanism, this method allows the model to independently make reasoning decisions when faced with complex tasks. Compared with the method used by the Tiny-Zero team with 2 A100 GPUs (160GB VRAM), the present invention can reproduce the same "insight moment" with only 7GB VRAM, thus significantly reducing the hardware requirements. In addition, GRPO supports full-scale fine-tuning and is compatible with both QLoRA and LoRA technologies, further enhancing the reasoning ability. The researchers at DeepSeek observed the "insight moment" when training R1-Zero using pure Reinforcement Learning (RL). The model learned to extend the thinking time by re-evaluating its initial plan without any human guidance or predefined instructions. In a test case, we only trained Phi-4 for 100 steps using GRPO. The model without GRPO training lacked "thinking tokens", while the model trained with GRPO had them and the answers were also correct.
[0104] Application Scenarios of GRPO: The GRPO algorithm can be widely applied to the creation of customized reward models. Especially in professional fields such as law and medicine, it can automatically generate the reasoning process from input-output data, avoiding the cumbersome process of preparing reasoning chain data in advance.
[0105] Practical Suggestions for GRPO: When conducting local GRPO training, it is recommended to install diffusers and train for at least 300 steps or more to observe the improvement in rewards, using the latest version of vLLM. To ensure good results, it is recommended to train for at least 12 hours, use a model with a parameter count of no less than 1.5 billion, and configure a suitable chat template. The inference mine provides a training loss tracking function to help users effectively monitor the training progress.
[0106] Support for Other Reinforcement Learning Methods: The present invention supports a variety of generation-based reinforcement learning methods, including OnlineDPO, PPO, and RLOO, expanding the flexibility of application scenarios and model selection.
[0107] Integration with vLLM: The present invention enhances throughput and reduces VRAM consumption by integrating vLLM. Specifically, the throughput is increased by 20 times, and the VRAM consumption is reduced by 50%. On a single A100 40GB GPU, using the Llama 3.23B Instruct model with dynamic 4-bit quantization, a throughput of up to 4000 tokens / s is expected. Additionally, Unsloth optimizes memory management, reducing the double memory occupancy when co-loading with vLLM. vLLM is capable of loading the Unsloth dynamic 4-bit quantization model and automatically optimizing parameters to improve RAM, VRAM efficiency, and throughput. The inference framework automatically selects multiple parameters to balance memory (RAM), video memory (VRAM) efficiency, and maximum throughput (e.g., the number of pre-filled tokens in chunks, the maximum sequence number, etc.). We default to enabling -O3 optimization and prefix caching in vLLM. We found that on older GPUs, Flashinfer actually reduces the speed by 10%. The FP8 KV cache reduces the speed by 10% but doubles the throughput potential. The present invention allows for loading LoRA models in vLLM by parsing the state dict rather than loading from disk - this can increase the GRPO training speed by 1.5 times. The present invention adds a batch generation function to reduce random video memory (VRAM) peaks, especially during batch generation.
[0108] The above are the embodiments of the present invention. For those of ordinary skill in the art, according to the teachings of the present invention, any equivalent changes, modifications, substitutions, and variations made without departing from the principles and spirit of the present invention within the scope of the patent application of the present invention shall fall within the scope covered by the present invention.
Claims
1. A large model inference optimization method based on the GRPO algorithm, characterized in that, It includes the following steps: Action Sampling: In each state, sample multiple candidate actions from the policy network, where the candidate actions include actions sampled from the action distribution output by the policy network or action variants generated based on the current policy through a perturbation mechanism; Relative Reward Calculation: Calculate the performance difference of each candidate action relative to other candidate actions to generate a relative reward signal, where the relative reward signal is based on the relative effect between candidate actions rather than absolute value; Policy Update: Update the policy network according to the relative reward signal, where the policy update is based on the relative performance of candidate actions rather than individual action values; Stability Enhancement: Through the intra-group relative reward mechanism, avoid the instability caused by inaccurate value network estimation in traditional reinforcement learning methods, and reduce the problems of gradient vanishing or explosion; Inference Optimization: On the premise of keeping the model parameter structure unchanged, optimize the inference process of the large model through the GRPO algorithm, reduce the consumption of computing resources, and improve the inference speed and efficiency.
2. The large model inference optimization method based on the GRPO algorithm according to claim 1, characterized in that The GRPO algorithm reduces the computing resource requirements in the following ways: Abandon the practice of relying on a critic model to estimate the value of each action in traditional reinforcement learning, and instead adopt an intra-group relative reward mechanism; By sampling multiple candidate actions and calculating their relative performance differences, reduce the dependence on the value network and lower the computational complexity.
3. The large model inference optimization method based on the GRPO algorithm according to claim 1, characterized in that The GRPO algorithm is applicable to a variety of large model inference scenarios, including but not limited to: Real-time data processing and decision-making in autonomous driving systems; Fast diagnosis and image analysis in medical image processing; Low-latency data processing in Internet of Things devices; Large-scale concurrent request processing in cloud computing and edge computing.
4. A large model inference optimization method based on the GRPO algorithm according to claim 1, characterized in that, The GRPO algorithm improves the inference efficiency in the following ways: Implement the training and inference of large language models on consumer-grade GPUs; Compared with Hugging Face+FA2, reduce the VRAM occupancy by 80%; Support the inference optimization of models with a maximum number of parameters of 15 billion.
5. A large model inference optimization method based on the GRPO algorithm according to claim 1, characterized in that The GRPO algorithm achieves performance enhancement in the following ways: Through reinforcement learning, enable the model to extend the thinking time and re-evaluate the initial method without manual intervention, achieving an improvement in autonomous inference; Reproduce the performance enhancement of DeepSeek R1-Zero on consumer-grade GPUs, reducing the hardware requirements.
6. A large model inference optimization method based on the GRPO algorithm according to claim 1, characterized in that, The GRPO algorithm improves the training stability in the following ways: Through the intra-group relative reward mechanism, avoid the training difficulties caused by sparse reward signals in traditional reinforcement learning methods; In the case of complex tasks and long-time delays, accelerate the model convergence through the relative reward mechanism.
7. A large model inference optimization method based on the GRPO algorithm according to claim 1, characterized in that The GRPO algorithm improves the flexibility of the inference framework in the following ways: Support loading the LoRA model in vLLM by parsing the state dictionary instead of loading from disk, improving the GRPO training speed by 1.5 times; Add a batch generation function to reduce the random VRAM peak during batch generation.
8. A large model inference optimization system based on the GRPO algorithm, characterized in that, It includes: Action Sampling Module: Used to sample multiple candidate actions from the policy network in each state; Relative Reward Calculation Module: Used to calculate the performance difference of each candidate action relative to other candidate actions and generate a relative reward signal; Policy update module: used to update the policy network according to the relative reward signal; Inference optimization module: used to optimize the inference process of the large model through the GRPO algorithm while keeping the model parameter structure unchanged, reducing computational resource consumption, and improving inference speed and efficiency.
9. The large model inference optimization system based on the GRPO algorithm according to claim 1, characterized in that The system further includes: Model compatibility module: used to support the compatibility of multiple model architectures, including but not limited to Llama, Phi, Mistral, Qwen; Training stability module: used to avoid the training difficulties caused by sparse reward signals in traditional reinforcement learning methods through the intra-group relative reward mechanism.
10. A large model inference optimization system based on the GRPO algorithm according to claim 1, characterized in that, The system further includes: Inference efficiency improvement module: used to implement the training and inference of large language models on consumer-grade GPUs, reducing VRAM occupancy; Throughput optimization module: used to improve throughput and reduce VRAM consumption by integrating vLLM.