Government affair large model performance optimization method, device and equipment and medium
Through real-time monitoring and machine learning, identify the performance bottlenecks of government big models and generate optimization strategies, the problems of inference delay, resource consumption and dynamic load adaptability of government big models are solved, and the optimization and stability of model performance are achieved.
Patent Information
- Application Number
- CN202510624640.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-29
AI Technical Summary
In actual operation, the government big model faces problems such as delay in reasoning, high resource consumption, poor dynamic load adaptability, and the inability of traditional optimization strategies to adapt to dynamic changes in real time.
By monitoring the performance data of government affairs large models in real time, using machine learning algorithms to identify performance bottlenecks, and generating targeted optimization strategies, including model parameter adjustment, hardware resource allocation optimization and model structure simplification, forming closed-loop control to ensure the sustainability and stability of optimization effects.
The performance optimization of the government affairs model is achieved, reducing inference delays, improving resource utilization, enhancing dynamic load adaptability, and ensuring the stable operation of the model when load fluctuates.
Smart Images

Figure CN120386699A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large models, and particularly to a method, device, equipment and medium for optimizing the performance of a government affairs large model. Background Art
[0002] With the rapid development of artificial intelligence technology, government affairs large models are increasingly widely used in the fields of natural language processing, image recognition, recommendation systems, etc. However, government affairs large models face many performance challenges in actual operation:
[0003] 1. Inference latency: The increase in model scale leads to a significant increase in inference time, affecting the user experience.
[0004] 2. Resource consumption: The high occupancy of computing resources (such as GPUs, memory) limits the scalability of the model.
[0005] 3. Poor adaptability to dynamic load: When the load fluctuates (such as a sudden increase or decrease in the number of requests), it is difficult for the model performance to remain stable.
[0006] 4. Limitations of existing optimization means: Traditional static optimization strategies cannot adapt to dynamic changes during runtime in real time.
[0007] In summary, how to monitor the model performance in real time and dynamically adjust the optimization strategy is an urgent problem to be solved at present. Summary of the Invention
[0008] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for optimizing the performance of a government affairs large model, which can monitor the model performance in real time and dynamically adjust the optimization strategy. The specific solutions are as follows:
[0009] In the first aspect, the present application provides a method for optimizing the performance of a government affairs large model, including:
[0010] Obtain the initial monitoring data corresponding to the government affairs large model to be optimized, perform a preset data processing operation on the initial monitoring data to obtain the processed monitoring data, and determine the first performance data of the government affairs large model to be optimized based on the processed monitoring data;
[0011] Based on the first performance data of the government affairs large model to be optimized, use a preset machine learning algorithm to obtain the performance bottleneck of the government affairs large model to be optimized, and determine the target position of the performance bottleneck according to the performance bottleneck of the government affairs large model to be optimized;
[0012] Generate a model optimization strategy corresponding to the to-be-optimized government affairs large model according to the target location of the performance bottleneck, perform a preset optimization operation on the to-be-optimized government affairs large model using the model optimization strategy to obtain an optimized government affairs large model, verify the second performance data of the optimized government affairs large model, and determine the target government affairs large model based on the obtained verification result.
[0013] Optionally, the obtaining the initial monitoring data corresponding to the to-be-optimized government affairs large model includes:
[0014] Capture the inference latency data of the to-be-optimized government affairs large model through a preset API probe;
[0015] Obtain the utilization rate data of the GPU and CPU and the memory occupancy rate data through a hardware performance counter;
[0016] Determine the initial monitoring data corresponding to the to-be-optimized government affairs large model based on the inference latency data, the utilization rate data of the GPU and CPU, and the memory occupancy rate data of the to-be-optimized government affairs large model.
[0017] Optionally, the performance bottleneck includes a computing resource bottleneck, a memory bottleneck, and an algorithm bottleneck;
[0018] Correspondingly, the obtaining the performance bottleneck of the to-be-optimized government affairs large model by using a preset machine learning algorithm based on the first performance data of the to-be-optimized government affairs large model includes:
[0019] Analyze the first performance data of the to-be-optimized government affairs large model to determine the inference latency and GPU utilization rate of the to-be-optimized government affairs large model. If the inference latency and GPU utilization rate of the to-be-optimized government affairs large model meet the preset latency condition, obtain the computing resource bottleneck corresponding to the to-be-optimized government affairs large model;
[0020] Analyze the first performance data of the to-be-optimized government affairs large model to determine the change trend of the memory occupancy rate of the to-be-optimized government affairs large model. If the change trend of the memory occupancy rate of the to-be-optimized government affairs large model meets the preset memory occupancy condition, obtain the memory bottleneck corresponding to the to-be-optimized government affairs large model;
[0021] Analyze the first performance data of the to-be-optimized government affairs large model to determine the computational graph and parameter update situation of the to-be-optimized government affairs large model. If the computational graph and parameter update situation of the to-be-optimized government affairs large model meet the preset algorithm execution condition, obtain the algorithm bottleneck corresponding to the to-be-optimized government affairs large model;
[0022] Determine the performance bottleneck of the to-be-optimized government affairs large model according to the obtained resource bottleneck, memory bottleneck, and algorithm bottleneck.
[0023] Optionally, generating a model optimization strategy corresponding to the government affairs large model to be optimized according to the target position of the performance bottleneck includes:
[0024] Generating a target optimization strategy corresponding to the target position respectively according to the target position of the performance bottleneck;
[0025] Integrating the generated target optimization strategies to obtain a model optimization strategy corresponding to the government affairs large model to be optimized.
[0026] Optionally, verifying the second performance data of the optimized government affairs large model and determining the target government affairs large model based on the obtained verification result includes:
[0027] Obtaining target monitoring data of the optimized government affairs large model;
[0028] Determining the second performance data of the optimized government affairs large model based on the target monitoring data;
[0029] Judging whether the optimized government affairs large model meets a preset model optimization condition and a preset model stability condition according to the second performance data of the optimized government affairs large model;
[0030] If the optimized government affairs large model meets the preset model optimization condition and the preset model stability condition, determining the optimized government affairs large model as the target government affairs large model;
[0031] If the optimized government affairs large model does not meet the preset model optimization condition and / or the preset model stability condition, re-executing the step of obtaining the performance bottleneck of the government affairs large model to be optimized by using a preset machine learning algorithm based on the first performance data of the government affairs large model to be optimized.
[0032] Optionally, judging whether the optimized government affairs large model meets a preset model optimization condition and a preset model stability condition according to the second performance data of the optimized government affairs large model includes:
[0033] Comparing the second performance data of the optimized government affairs large model with the first performance data of the government affairs large model to be optimized;
[0034] Judging whether the optimized government affairs large model meets the preset model optimization condition based on the obtained comparison result, so as to determine the target government affairs large model according to the obtained first determination result.
[0035] Optionally, judging whether the optimized government affairs large model meets a preset model optimization condition and a preset model stability condition according to the second performance data of the optimized government affairs large model includes:
[0036] Monitor the change trend of the second performance data;
[0037] Based on the change trend of the second performance data, determine whether the optimized government large model meets the preset model stability condition, so as to determine the target government large model according to the obtained second determination result.
[0038] In a second aspect, the present application provides a performance optimization device for a government large model, including:
[0039] A data determination module, configured to obtain initial monitoring data corresponding to the government large model to be optimized, perform a preset data processing operation on the initial monitoring data to obtain processed monitoring data, and determine the first performance data of the government large model to be optimized based on the processed monitoring data;
[0040] A position determination module, configured to obtain the performance bottleneck of the government large model to be optimized by using a preset machine learning algorithm based on the first performance data of the government large model to be optimized, and determine the target position of the performance bottleneck according to the performance bottleneck of the government large model to be optimized;
[0041] A target government large model determination module, configured to generate a model optimization strategy corresponding to the government large model to be optimized according to the target position of the performance bottleneck, perform a preset optimization operation on the government large model to be optimized by using the model optimization strategy to obtain an optimized government large model, verify the second performance data of the optimized government large model, and determine the target government large model based on the obtained verification result.
[0042] In a third aspect, the present application provides an electronic device, including:
[0043] A memory, configured to store a computer program;
[0044] A processor, configured to execute the computer program to implement the performance optimization method of the government large model as described above.
[0045] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the performance optimization method of the government large model as described above is implemented.
[0046] In summary, the present application first obtains the initial monitoring data corresponding to the government affairs large model to be optimized, performs a preset data processing operation on the initial monitoring data to obtain the processed monitoring data, and determines the first performance data of the government affairs large model to be optimized based on the processed monitoring data; obtains the performance bottleneck of the government affairs large model to be optimized by using a preset machine learning algorithm based on the first performance data of the government affairs large model to be optimized, and determines the target position of the performance bottleneck according to the performance bottleneck of the government affairs large model to be optimized; generates a model optimization strategy corresponding to the government affairs large model to be optimized according to the target position of the performance bottleneck, performs a preset optimization operation on the government affairs large model to be optimized by using the model optimization strategy to obtain the optimized government affairs large model, verifies the second performance data of the optimized government affairs large model, and determines the target government affairs large model based on the obtained verification result. As can be seen from the above, the present application first obtains the initial monitoring data of the government affairs large model to be optimized, obtains the processed monitoring data through a preset data processing operation, determines the first performance data of the government affairs large model to be optimized based on this, then uses a preset machine learning algorithm to find the performance bottleneck and determine its target position based on the first performance data, then generates a model optimization strategy according to the target position, uses this strategy to perform a preset optimization operation on the government affairs large model to be optimized to obtain the optimized government affairs large model, then verifies the second performance data of the optimized government affairs large model, and finally determines the target government affairs large model according to the verification result. In this way, performance data during model operation, including key indicators such as inference latency, throughput, and resource utilization, is collected through a real-time monitoring module, and machine learning algorithms are used based on the monitoring data to identify performance bottlenecks and generate targeted optimization strategies, such as model parameter adjustment, hardware resource allocation optimization, and model structure simplification. In addition, the model operation parameters or hardware resource configurations are adjusted in real time, and a closed-loop control is formed through a feedback mechanism to ensure the sustainability and stability of the optimization effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0048] Figure 1 It is a flowchart of a method for optimizing the performance of a government affairs large model disclosed in the present application;
[0049] Figure 2 It is a schematic diagram of the monitoring and analysis stage process of a method for optimizing the performance of a government affairs large model disclosed in the present application;
[0050] Figure 3Schematic diagram of the process of the performance optimization and verification stage of a government affairs large model disclosed in this application;
[0051] Figure 4 Schematic diagram of the structure of a performance optimization device for a government affairs large model disclosed in this application;
[0052] Figure 5 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Currently, with the rapid development of artificial intelligence technology, government affairs large models are increasingly widely used in fields such as natural language processing, image recognition, and recommendation systems. However, government affairs large models face many performance challenges in actual operation: 1. Inference latency: The increase in model scale leads to a significant increase in inference time, affecting the user experience. 2. Resource consumption: The high occupancy of computing resources, such as GPUs and memory, limits the scalability of the model. 3. Poor adaptability to dynamic loads: When the load fluctuates (such as a sudden increase or decrease in the number of requests), it is difficult for the model performance to remain stable. 4. Limitations of existing optimization means: Traditional static optimization strategies cannot adapt to dynamic changes during runtime in real time. To solve the above technical problems, this application discloses a performance optimization method, device, equipment, and medium for a government affairs large model, which can monitor the model performance in real time and dynamically adjust the optimization strategy.
[0055] See Figure 1 As shown, the embodiments of the present invention disclose a performance optimization method for a government affairs large model, including:
[0056] Step S11: Obtain the initial monitoring data corresponding to the government affairs large model to be optimized, perform a preset data processing operation on the initial monitoring data to obtain the processed monitoring data, and determine the first performance data of the government affairs large model to be optimized based on the processed monitoring data.
[0057] In this embodiment, first, inference latency data of the government affairs large model to be optimized is captured through a preset API (Application Programming Interface) probe; utilization rate data of the GPU (Graphics Processing Unit) and CPU (Central Processing Unit) and memory occupancy rate data are obtained through hardware performance counters; based on the inference latency data of the government affairs large model to be optimized, the utilization rate data of the GPU and CPU, and the memory occupancy rate data, initial monitoring data corresponding to the government affairs large model to be optimized is determined. Specifically, as Figure 2 shown, real-time monitoring is performed in the large model running environment in various ways to collect performance data in real time. For example, during the model inference process, the start time and end time of each inference are captured through a preset API probe to calculate the inference latency data; information such as the utilization rate data of the GPU and CPU and the memory occupancy rate data is obtained through a preset hardware performance counter. After integrating the obtained inference latency data, the utilization rate data of the GPU and CPU, and the memory occupancy rate data, initial monitoring data is obtained.
[0058] It should be noted that after obtaining the initial monitoring data, preset data processing operations, namely cleaning and normalization processing, are performed on the collected raw data to determine the processed monitoring data. For example, significantly abnormal latency data, such as occasional extremely high latency caused by network fluctuations, is removed, and resource utilization rate data in different formats is uniformly converted into a percentage form to ensure the consistency and accuracy of the data.
[0059] Meanwhile, based on the processed monitoring data, the first performance data of the government affairs large model to be optimized is determined. For example, the average inference latency is calculated by statistically averaging the latencies of all inference requests within a certain time window; the 99th percentile latency is calculated to find the value at the 99% position in the latency distribution, reflecting the latency performance in extreme cases; the change trend of resource utilization rate is analyzed through methods such as moving average to identify the upward or downward trend of resource utilization rate, providing a basis for identifying performance bottlenecks.
[0060] Step S12: Based on the first performance data of the government affairs large model to be optimized, use a preset machine learning algorithm to obtain the performance bottleneck of the government affairs large model to be optimized, and determine the target position of the performance bottleneck according to the performance bottleneck of the government affairs large model to be optimized.
[0061] In this embodiment, a machine learning algorithm is used to deeply analyze the first performance data to automatically identify the performance bottleneck of the government affairs large model to be optimized. Among them, the performance bottleneck may include a computing resource bottleneck, a memory bottleneck, and an algorithm bottleneck.
[0062] In a specific embodiment, analyze the first performance data of the government affairs large model to be optimized to determine the inference latency and GPU utilization rate of the government affairs large model to be optimized. If the inference latency and GPU utilization rate of the government affairs large model to be optimized meet the preset latency conditions, obtain the corresponding computing resource bottleneck of the government affairs large model to be optimized. Specifically, by analyzing the correlation between the inference latency and the GPU utilization rate, if it is found that the GPU utilization rate is high while the inference latency is long, and reducing the GPU load can significantly shorten the latency, it is determined that there is a computing resource bottleneck. Further, by analyzing the GPU memory allocation situation, identify whether there is computing latency caused by insufficient memory, and finally obtain the computing resource bottleneck corresponding to the government affairs large model to be optimized.
[0063] In another specific embodiment, analyze the first performance data of the government affairs large model to be optimized to determine the changing trend of the memory occupancy rate of the government affairs large model to be optimized. If the changing trend of the memory occupancy rate of the government affairs large model to be optimized meets the preset memory occupancy conditions, obtain the corresponding memory bottleneck of the government affairs large model to be optimized. Specifically, monitor the changing trend of the memory occupancy rate of the government affairs large model to be optimized. If the memory occupancy rate continues to rise and is positively correlated with the increase in inference requests, it is determined that there is a problem of memory leakage or insufficient memory, and it is necessary to further analyze the memory allocation and recycling mechanism to obtain the memory bottleneck of the government affairs large model to be optimized.
[0064] In a third specific embodiment, analyze the first performance data of the government affairs large model to be optimized to determine the computational graph and parameter update situation of the government affairs large model to be optimized. If the computational graph and parameter update situation of the government affairs large model to be optimized meet the preset algorithm execution conditions, obtain the corresponding algorithm bottleneck of the government affairs large model to be optimized. Specifically, by analyzing the computational graph and parameter update situation of the model, identify the efficiency problem of the algorithm itself, that is, the algorithm bottleneck corresponding to the government affairs large model to be optimized. For example, some complex network structures may lead to excessive computational operations, increasing the inference latency; some unreasonable hyperparameter settings may lead to slow model convergence speed, affecting the training and inference efficiency.
[0065] Furthermore, after identifying the performance bottleneck, further locate the specific location and cause of the bottleneck. For example, for the computing resource bottleneck, by analyzing the computational graph of the model, find the layers or operations with a large amount of computation; for the memory bottleneck, use a memory analysis tool to locate the code segment with memory leakage or the part with unreasonable memory allocation.
[0066] Step S13: Generate a model optimization strategy corresponding to the to-be-optimized government affairs large model according to the target position of the performance bottleneck, perform a preset optimization operation on the to-be-optimized government affairs large model by using the model optimization strategy to obtain an optimized government affairs large model, verify the second performance data of the optimized government affairs large model, and determine the target government affairs large model based on the obtained verification result.
[0067] In this embodiment, as Figure 3 shown, generate a targeted model optimization strategy according to the performance bottleneck of the to-be-optimized government affairs large model. If the performance bottleneck is caused by a large amount of computation due to a high floating-point precision of the model, a strategy to reduce the model precision can be generated, such as converting FP32 to FP16, reducing the amount of computation, and improving the inference speed. At the same time, it is necessary to evaluate the impact of the precision reduction on the model performance to ensure that it is within an acceptable range. If the performance bottleneck is due to insufficient GPU resources, a strategy to dynamically adjust the GPU memory allocation can be generated to preferentially ensure the GPU memory requirements of high-load tasks. In addition, CPU and memory resources can be reasonably allocated according to the priority and resource requirements of the tasks to improve resource utilization. If the performance bottleneck is caused by a large amount of computation due to a complex model structure, a strategy of model pruning or quantization can be generated. For example, redundant neurons or connections in the model can be removed through pruning to reduce computational operations; the storage space and computational complexity of the model parameters can be reduced through quantization, and the inference efficiency can be improved on the premise of ensuring the model precision.
[0068] Next, generate the target optimization strategy corresponding to the target position according to the target position of the performance bottleneck respectively; integrate the generated various target optimization strategies to obtain the model optimization strategy corresponding to the to-be-optimized government affairs large model. Specifically, on the basis of generating optimization strategies for a single dimension, multiple optimization strategies are combined. For example, model parameter adjustment and hardware resource allocation optimization are performed simultaneously to achieve a better optimization effect. By comprehensively considering the mutual influence of each optimization strategy, the optimal strategy combination is generated to ensure the comprehensive improvement of the model performance.
[0069] Furthermore, after receiving the model optimization strategy, the model optimization strategy is immediately executed. The runtime configuration of the model can be modified to adjust the floating-point precision, hyperparameters, etc. of the model. For example, during the inference process, the model accuracy can be reduced by loading the predefined FP16 model weights; the training and inference processes of the model can be optimized by adjusting hyperparameters such as the learning rate and batch size. Resource allocation strategies such as GPU memory allocation and CPU affinity can also be dynamically adjusted through interaction with the operating system and hardware management interfaces. For example, more GPU memory can be allocated to high-priority tasks, and the CPU usage of low-priority tasks can be restricted to ensure reasonable resource allocation. The original model can also be replaced by loading the pruned or quantized model structure. For example, during the inference process, the pruned model can be used for calculation to reduce the amount of calculation; the quantized model parameters can be used to reduce the storage space and computational complexity. In addition to optimizing the to-be-optimized government affairs large model, the application of the to-be-optimized government affairs large model can also be optimized. For example, in scenarios where multiple tasks and multiple models run concurrently, the scheduling order and resource allocation of tasks can be dynamically adjusted according to the load conditions and priorities of each task. For example, when the load of a certain task suddenly increases, the dynamic scheduling module can timely allocate more computing resources to this task and adjust the resource allocation of other tasks to ensure the performance balance of each task and avoid performance degradation caused by resource contention. After optimizing the to-be-optimized government affairs large model, an optimized government affairs large model is obtained.
[0070] In this embodiment, target monitoring data of the optimized government affairs large model is obtained; second performance data of the optimized government affairs large model is determined based on the target monitoring data; it is judged whether the optimized government affairs large model meets a preset model optimization condition and a preset model stability condition according to the second performance data of the optimized government affairs large model; if the optimized government affairs large model meets the preset model optimization condition and the preset model stability condition, the optimized government affairs large model is determined as the target government affairs large model; if the optimized government affairs large model does not meet the preset model optimization condition and / or the preset model stability condition, the step of using a preset machine learning algorithm to obtain the performance bottleneck of the to-be-optimized government affairs large model based on the first performance data of the to-be-optimized government affairs large model is re-executed. Specifically, after the optimization strategy is executed, the second performance data of the optimized government affairs large model is continuously collected to verify the optimization effect according to the second performance data of the optimized government affairs large model.
[0071] In a specific embodiment, the second performance data of the optimized government affairs large model is compared with the first performance data of the government affairs large model to be optimized; based on the obtained comparison result, it is determined whether the optimized government affairs large model meets the preset model optimization conditions, so as to determine the target government affairs large model according to the obtained first determination result. Specifically, by comparing the second performance data of the optimized government affairs large model with the first performance data of the government affairs large model to be optimized, comparison results such as inference latency, throughput, resource utilization rate, etc. are obtained to evaluate the effectiveness of the optimization strategy. For example, if the inference latency is significantly reduced and the throughput is significantly increased, it indicates that the optimization strategy has achieved good results.
[0072] In another specific embodiment, in addition to the verification of performance indicators, the stability of the model also needs to be verified. It is necessary to monitor the change trend of the second performance data; based on the change trend of the second performance data, it is determined whether the optimized government affairs large model meets the preset model stability conditions, so as to determine the target government affairs large model according to the obtained second determination result. Specifically, by running the model for a long time, the change trend of performance indicators is monitored to ensure that the performance of the optimized model is stable during long-term operation and there are no performance fluctuations or abnormal situations. For example, observe whether the inference latency remains stable and whether the resource utilization rate fluctuates within a reasonable range during the 24-hour operation after optimization.
[0073] Finally, according to the results of performance verification and stability verification, the target government affairs large model is determined. If it is found that the optimization effect does not meet the expectations or new performance bottlenecks occur, the analysis phase will be re-entered to generate a new model optimization strategy, and the optimization and verification process will be performed on the government affairs large model again to form a closed-loop control and continuously improve the performance of the government affairs large model.
[0074] As can be seen from the above, in the embodiment of the present application, the initial monitoring data of the government affairs large model to be optimized is first obtained, and the processed monitoring data is obtained through preset data processing operations. Based on this, the first performance data of the government affairs large model to be optimized is determined. Then, based on the first performance data, the performance bottleneck is found using a preset machine learning algorithm and its target location is determined. Next, a model optimization strategy is generated according to the target location, and the preset optimization operation is performed on the government affairs large model to be optimized using this strategy to obtain the optimized government affairs large model. Then, the second performance data of the optimized government affairs large model is verified. Finally, the target government affairs large model is determined according to the verification result. In this way, the performance data during the operation of the model, including key indicators such as inference latency, throughput, and resource utilization rate, is collected through the real-time monitoring module. Based on the monitoring data, machine learning algorithms are used to identify performance bottlenecks and generate targeted optimization strategies, such as model parameter adjustment, hardware resource allocation optimization, and model structure simplification. In addition, the model operation parameters or hardware resource configurations are adjusted in real time, and a closed-loop control is formed through the feedback mechanism to ensure the persistence and stability of the optimization effect.
[0075] As can be seen from the previous embodiment, the present application discloses a method for optimizing the performance of a government affairs large model, which can monitor the model performance in real time and dynamically adjust the optimization strategy. Next, a detailed description of the method for optimizing the performance of the government affairs large model will be given.
[0076] First, the present application monitors in real time in the operating environment of the government affairs large model in various ways to collect the initial monitoring data of the government affairs large model in real time. After obtaining the initial monitoring data, perform preset data processing operations on the collected raw data, that is, cleaning and normalization processing, to determine the processed monitoring data. Based on the processed monitoring data, determine the first performance data of the government affairs large model to be optimized.
[0077] Next, use machine learning algorithms to deeply analyze the first performance data to automatically identify the performance bottlenecks of the government affairs large model to be optimized. Among them, the performance bottlenecks can include computing resource bottlenecks, memory bottlenecks, and algorithm bottlenecks. After identifying the performance bottlenecks, further locate the specific location and cause of the bottlenecks.
[0078] Finally, generate targeted model optimization strategies according to the performance bottlenecks of the government affairs large model to be optimized. After receiving the model optimization strategies, immediately execute the model optimization strategies. After the optimization strategies are executed, continue to collect the second performance data of the optimized government affairs large model to perform performance verification and stability verification on the optimization effect according to the second performance data of the optimized government affairs large model. According to the results of the performance verification and stability verification, determine the target government affairs large model. If it is found that the optimization effect does not meet the expectations or new performance bottlenecks appear, re-enter the analysis stage, generate new model optimization strategies, and perform the optimization and verification process on the government affairs large model again to form a closed-loop control and continuously improve the performance of the government affairs large model.
[0079] See Figure 4 As shown, an embodiment of the present invention discloses a method for optimizing the performance of a government affairs large model, including:
[0080] A data determination module 11, configured to obtain initial monitoring data corresponding to a government affairs large model to be optimized, perform preset data processing operations on the initial monitoring data to obtain processed monitoring data, and determine the first performance data of the government affairs large model to be optimized based on the processed monitoring data;
[0081] A position determination module 12, configured to obtain the performance bottlenecks of the government affairs large model to be optimized by using a preset machine learning algorithm based on the first performance data of the government affairs large model to be optimized, and determine the target position of the performance bottlenecks according to the performance bottlenecks of the government affairs large model to be optimized;
[0082] The target government affairs large model determination module 13 is configured to generate a model optimization strategy corresponding to the to-be-optimized government affairs large model according to the target position of the performance bottleneck, perform a preset optimization operation on the to-be-optimized government affairs large model by using the model optimization strategy to obtain an optimized government affairs large model, verify the second performance data of the optimized government affairs large model, and determine the target government affairs large model based on the obtained verification result.
[0083] As can be seen from the above, in this application, the initial monitoring data of the to-be-optimized government affairs large model is first obtained, the processed monitoring data is obtained through a preset data processing operation, and the first performance data of the to-be-optimized government affairs large model is determined accordingly. Then, based on the first performance data, a preset machine learning algorithm is used to find the performance bottleneck and determine its target position. Next, a model optimization strategy is generated according to the target position, and the preset optimization operation is performed on the to-be-optimized government affairs large model by using this strategy to obtain an optimized government affairs large model. Then, the second performance data of the optimized government affairs large model is verified, and finally, the target government affairs large model is determined according to the verification result. In this way, by using the real-time monitoring module to collect the performance data during the model operation, including key indicators such as inference latency, throughput, and resource utilization rate, the machine learning algorithm is used based on the monitoring data to identify the performance bottleneck and generate targeted optimization strategies, such as model parameter adjustment, hardware resource allocation optimization, and model structure simplification. In addition, the model operation parameters or hardware resource configuration are adjusted in real time, and a closed-loop control is formed through the feedback mechanism to ensure the continuity and stability of the optimization effect.
[0084] In some specific embodiments, the data determination module 11 may specifically include:
[0085] The inference latency data capture unit is configured to capture the inference latency data of the to-be-optimized government affairs large model through a preset API probe;
[0086] The utilization rate data and memory occupancy rate data acquisition unit is configured to acquire the utilization rate data and memory occupancy rate data of the GPU and CPU through a hardware performance counter;
[0087] The initial monitoring data determination unit is configured to determine the initial monitoring data corresponding to the to-be-optimized government affairs large model based on the inference latency data, the utilization rate data of the GPU and CPU, and the memory occupancy rate data of the to-be-optimized government affairs large model.
[0088] In some specific embodiments, the performance bottleneck includes a computing resource bottleneck, a memory bottleneck, and an algorithm bottleneck;
[0089] Correspondingly, the position determination module 12 may specifically include:
[0090] A computing resource bottleneck acquisition unit, configured to analyze first performance data of the to-be-optimized government affairs large model to determine the inference latency and GPU utilization rate of the to-be-optimized government affairs large model. If the inference latency and GPU utilization rate of the to-be-optimized government affairs large model meet a preset latency condition, obtain the computing resource bottleneck corresponding to the to-be-optimized government affairs large model;
[0091] A memory bottleneck acquisition unit, configured to analyze the first performance data of the to-be-optimized government affairs large model to determine the change trend of the memory occupancy rate of the to-be-optimized government affairs large model. If the change trend of the memory occupancy rate of the to-be-optimized government affairs large model meets a preset memory occupancy condition, obtain the memory bottleneck corresponding to the to-be-optimized government affairs large model;
[0092] An algorithm bottleneck acquisition unit, configured to analyze the first performance data of the to-be-optimized government affairs large model to determine the computation graph and parameter update situation of the to-be-optimized government affairs large model. If the computation graph and parameter update situation of the to-be-optimized government affairs large model meet a preset algorithm execution condition, obtain the algorithm bottleneck corresponding to the to-be-optimized government affairs large model;
[0093] A performance bottleneck unit, configured to determine the performance bottleneck of the to-be-optimized government affairs large model according to the obtained resource bottleneck, memory bottleneck, and algorithm bottleneck.
[0094] In some specific embodiments, the target government affairs large model determination module 13 may specifically include:
[0095] A target optimization strategy generation unit, configured to generate target optimization strategies corresponding to the target positions according to the target positions of the performance bottleneck;
[0096] A model optimization strategy acquisition unit, configured to integrate the generated target optimization strategies to obtain a model optimization strategy corresponding to the to-be-optimized government affairs large model.
[0097] In some specific embodiments, the target government affairs large model determination module 13 may specifically include:
[0098] A target monitoring data acquisition unit, configured to acquire target monitoring data of the optimized government affairs large model;
[0099] A second performance data determination unit, configured to determine second performance data of the optimized government affairs large model based on the target monitoring data;
[0100] An optimized government affairs large model judgment unit, configured to judge whether the optimized government affairs large model meets a preset model optimization condition and a preset model stability condition according to the second performance data of the optimized government affairs large model;
[0101] The first optimized government affairs large model determination unit is used to determine the optimized government affairs large model as the target government affairs large model if the optimized government affairs large model meets the preset model optimization condition and the preset model stability condition;
[0102] The second optimized government affairs large model determination unit is used to re - execute the step of obtaining the performance bottleneck of the to - be - optimized government affairs large model by using the preset machine learning algorithm based on the first performance data of the to - be - optimized government affairs large model if the optimized government affairs large model does not meet the preset model optimization condition and / or the preset model stability condition.
[0103] In some specific embodiments, the optimized government affairs large model judgment unit may specifically include:
[0104] The performance data comparison subunit is used to compare the second performance data of the optimized government affairs large model with the first performance data of the to - be - optimized government affairs large model;
[0105] The target government affairs large model determination subunit is used to judge whether the optimized government affairs large model meets the preset model optimization condition based on the obtained comparison result, so as to determine the target government affairs large model according to the obtained first determination result.
[0106] In some specific embodiments, the optimized government affairs large model judgment unit may specifically include:
[0107] The change trend monitoring subunit of the second performance data is used to monitor the change trend of the second performance data;
[0108] The target government affairs large model determination subunit is used to judge whether the optimized government affairs large model meets the preset model stability condition based on the change trend of the second performance data, so as to determine the target government affairs large model according to the obtained second determination result.
[0109] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 5 It is the structure diagram of the electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the use scope of the present application.
[0110] Figure 5Schematic diagram of the structure of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the performance optimization method of the government affairs large model disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0111] In this embodiment, the power supply 23 is used to provide working voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitations are made here.
[0112] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be transient storage or permanent storage.
[0113] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, and it can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of implementing the performance optimization method of the government affairs large model executed by the electronic device 20 disclosed in any of the foregoing embodiments, may further include a computer program capable of performing other specific tasks.
[0114] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the performance optimization method of the government affairs large model disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0115] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0116] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.
[0117] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0118] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0119] The technical solutions provided in this application have been introduced in detail above. Specific examples have been used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for optimizing the performance of a government affairs large model, characterized in that, Including: Obtain the initial monitoring data corresponding to the government affairs large model to be optimized, perform a preset data processing operation on the initial monitoring data to obtain the processed monitoring data, and determine the first performance data of the government affairs large model to be optimized based on the processed monitoring data; Obtain the performance bottleneck of the government affairs large model to be optimized by using a preset machine learning algorithm based on the first performance data of the government affairs large model to be optimized, and determine the target position of the performance bottleneck according to the performance bottleneck of the government affairs large model to be optimized; Generate a model optimization strategy corresponding to the government affairs large model to be optimized according to the target position of the performance bottleneck, perform a preset optimization operation on the government affairs large model to be optimized by using the model optimization strategy to obtain the optimized government affairs large model, verify the second performance data of the optimized government affairs large model, and determine the target government affairs large model based on the obtained verification result.
2. The performance optimization method of the government affairs large model according to claim 1, wherein The obtaining of the initial monitoring data corresponding to the government affairs large model to be optimized includes: Capture the inference latency data of the government affairs large model to be optimized through a preset API probe; Obtain the utilization rate data of the GPU and CPU and the memory occupancy data through a hardware performance counter; Determine the initial monitoring data corresponding to the government affairs large model to be optimized based on the inference latency data, the utilization rate data of the GPU and CPU, and the memory occupancy data of the government affairs large model to be optimized.
3. The performance optimization method of the government affairs large model according to claim 1, characterized in that The performance bottleneck includes a computing resource bottleneck, a memory bottleneck, and an algorithm bottleneck; Correspondingly, the obtaining of the performance bottleneck of the government affairs large model to be optimized by using a preset machine learning algorithm based on the first performance data of the government affairs large model to be optimized includes: Analyze the first performance data of the government affairs large model to be optimized to determine the inference latency and GPU utilization rate of the government affairs large model to be optimized. If the inference latency and GPU utilization rate of the government affairs large model to be optimized meet the preset latency condition, obtain the computing resource bottleneck corresponding to the government affairs large model to be optimized; Analyze the first performance data of the government affairs large model to be optimized to determine the change trend of the memory occupancy rate of the government affairs large model to be optimized. If the change trend of the memory occupancy rate of the government affairs large model to be optimized meets the preset memory occupancy condition, obtain the memory bottleneck corresponding to the government affairs large model to be optimized; Analyze the first performance data of the government affairs large model to be optimized to determine the computation graph and parameter update situation of the government affairs large model to be optimized. If the computation graph and parameter update situation of the government affairs large model to be optimized meet the preset algorithm execution condition, obtain the algorithm bottleneck corresponding to the government affairs large model to be optimized; Determine the performance bottleneck of the government affairs large model to be optimized according to the obtained resource bottleneck, memory bottleneck, and algorithm bottleneck.
4. The performance optimization method of the government affairs large model according to claim 1, wherein The generating of the model optimization strategy corresponding to the government affairs large model to be optimized according to the target position of the performance bottleneck includes: Generate a target optimization strategy corresponding to the target position according to the target position of the performance bottleneck respectively; Integrate the generated target optimization strategies to obtain the model optimization strategy corresponding to the government affairs large model to be optimized.
5. The performance optimization method of the government affairs large model according to any one of claims 1 to 4, characterized in that The second performance data for verifying the optimized government affairs large model, and determining the target government affairs large model based on the obtained verification results, includes: Obtaining the target monitoring data of the optimized government affairs large model; Determining the second performance data of the optimized government affairs large model based on the target monitoring data; Judging whether the optimized government affairs large model meets the preset model optimization conditions and the preset model stability conditions according to the second performance data of the optimized government affairs large model; If the optimized government affairs large model meets the preset model optimization conditions and the preset model stability conditions, determining the optimized government affairs large model as the target government affairs large model; If the optimized government affairs large model does not meet the preset model optimization conditions and / or the preset model stability conditions, re-executing the step of obtaining the performance bottleneck of the to-be-optimized government affairs large model by using the preset machine learning algorithm based on the first performance data of the to-be-optimized government affairs large model.
6. The performance optimization method of the government affairs large model according to claim 5, wherein, The judging whether the optimized government affairs large model meets the preset model optimization conditions and the preset model stability conditions according to the second performance data of the optimized government affairs large model includes: Comparing the second performance data of the optimized government affairs large model with the first performance data of the to-be-optimized government affairs large model; Judging whether the optimized government affairs large model meets the preset model optimization conditions based on the obtained comparison result, so as to determine the target government affairs large model according to the obtained first determination result.
7. The performance optimization method of the government affairs large model according to claim 5, characterized in that, The judging whether the optimized government affairs large model meets the preset model optimization conditions and the preset model stability conditions according to the second performance data of the optimized government affairs large model includes: Monitoring the change trend of the second performance data; Judging whether the optimized government affairs large model meets the preset model stability conditions based on the change trend of the second performance data, so as to determine the target government affairs large model according to the obtained second determination result.
8. An apparatus for optimizing the performance of a government affairs large model, characterized in that, Includes: A data determination module, configured to obtain the initial monitoring data corresponding to the to-be-optimized government affairs large model, perform a preset data processing operation on the initial monitoring data to obtain the processed monitoring data, and determine the first performance data of the to-be-optimized government affairs large model based on the processed monitoring data; A position determination module, configured to obtain the performance bottleneck of the to-be-optimized government affairs large model by using a preset machine learning algorithm based on the first performance data of the to-be-optimized government affairs large model, and determine the target position of the performance bottleneck according to the performance bottleneck of the to-be-optimized government affairs large model; A target government affairs large model determination module, configured to generate a model optimization strategy corresponding to the to-be-optimized government affairs large model according to the target position of the performance bottleneck, perform a preset optimization operation on the to-be-optimized government affairs large model by using the model optimization strategy to obtain the optimized government affairs large model, verify the second performance data of the optimized government affairs large model, and determine the target government affairs large model based on the obtained verification result.
9. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the performance optimization method of the government affairs large model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer program is executed by a processor, it implements the performance optimization method of the government affairs big model as described in any one of claims 1 to 7.
Citation Information
Cited By
Model bottleneck determination method and device, electronic equipment, storage medium and program
CN120872776A