Language model reasoning method and device, computer equipment and storage medium

By monitoring and visualizing the inference performance information of the language model in real time and updating the inference strategy based on this information, the system performance degradation caused by the excessive number of inference requests is solved, the inference accuracy and stability are improved, and the resource utilization is optimized.

CN120069084AActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510224609.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

When the number of inference requests processed simultaneously is too high, the inference performance of the language model may be affected, manifested as prolonged inference time, excessive use of CPU and memory resources, and possible system bottlenecks, resulting in a degradation of the model's real-time response capabilities and the accuracy and stability of the inference results.

Method used

By monitoring the inference performance information in the inference process of language model inference in real time and visually display it, obtaining the inference performance information type, determining the inference evaluation system based on the inference performance information type and inference performance information, updating the inference strategy, including the number of multiple inference requests and the processing order, to optimize resource usage and improve service quality.

Benefits of technology

The inference strategy is updated in real time based on the inference performance information, which effectively solves the problem of system performance degradation caused by the excessive number of inference requests, improves the accuracy and stability of inference, and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069084A_ABST
    Figure CN120069084A_ABST
Patent Text Reader

Abstract

The invention discloses a language model reasoning method and device, computer equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps: monitoring reasoning performance information in a language model reasoning process in real time, and carrying out the visual display; the reasoning performance information type is obtained, a reasoning evaluation system is determined based on the reasoning performance information type and reasoning performance information, the comprehensive score of each main evaluation index in the reasoning evaluation system is determined, a reasoning strategy is updated, and the reasoning strategy comprises the number of multiple reasoning requests and the processing sequence; the technical problem that the system performance is reduced due to the fact that the number of reasoning requests processed at the same time is too large is solved, and the technical effect of improving reasoning accuracy and stability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a language model inference method, apparatus, computer device, and storage medium. Background Art

[0002] As a cutting-edge technology in the field of artificial intelligence, language model inference has made remarkable progress in recent years, profoundly reshaping the boundaries of natural language processing. During the inference process, language models achieve in-depth understanding and analysis of input text through multi-level attention mechanisms and self-attention mechanisms. They can identify logical relationships between sentences and perform advanced thinking activities such as causal reasoning and conditional reasoning, thereby generating outputs that are logical and close to human thinking patterns.

[0003] In related technologies, multi-threading or multi-processing technologies are widely adopted to process multiple inference requests in parallel, which significantly improves the overall processing capacity and response speed of the system. Although multi-threading or multi-processing technologies can process multiple requests in parallel, they also bring potential challenges. When the number of requests processed simultaneously is too large, the inference performance of the language model may be affected, specifically manifested as an extension of the inference time, excessive occupation of CPU and memory resources, and possible system bottlenecks. These negative impacts not only reduce the real-time response ability of the model but may also affect the accuracy and stability of the inference results. Summary of the Invention

[0004] This application provides a language model inference method, apparatus, computer device, and storage medium to at least solve the problem of system performance degradation caused by too many inference requests processed simultaneously in related technologies.

[0005] This application provides a language model inference method, including:

[0006] Obtain multiple inference requests to be processed, preprocess the multiple inference requests to be processed to obtain multiple inference requests;

[0007] Load the language model and perform inference on the multiple inference requests based on the language model;

[0008] Real-time monitor the inference performance information during the language model inference process and visually display it;

[0009] Obtain the type of inference performance information, and determine an inference evaluation system based on the type of inference performance information and the inference performance information. Among them, the inference evaluation system includes multiple main evaluation indicators; the multiple main evaluation indicators include a latency evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator;

[0010] Obtain the actual values of each main evaluation indicator and the preset indicator range of each main evaluation indicator;

[0011] Determine the scores of each main evaluation index according to the actual values of each main evaluation index and the preset index ranges of each main evaluation index;

[0012] Obtain the weight values of each main scoring index, and calculate the comprehensive scores of each main scoring index based on the weight values of each main scoring index and the scores of each main scoring index;

[0013] Update the inference strategy based on the comprehensive scores of each main evaluation index. The inference strategy includes the number and processing order of multiple inference requests, and perform language model inference operations according to the updated inference strategy.

[0014] This application also provides a language model inference device, including:

[0015] An acquisition module, configured to acquire multiple inference requests to be processed, preprocess the multiple inference requests to be processed, and obtain multiple inference requests;

[0016] An inference module, configured to load a language model and perform inferences on multiple inference requests based on the language model;

[0017] A monitoring module, configured to monitor the inference performance information during the language model inference process in real time and display it visually;

[0018] A determination module, configured to obtain the type of inference performance information, determine an inference evaluation system based on the type of inference performance information and the inference performance information, where the inference evaluation system includes multiple main evaluation indexes; the multiple main evaluation indexes include a latency evaluation index, a throughput evaluation index, and a processor evaluation index; obtain the actual values of each main evaluation index and the preset index ranges of each main evaluation index; determine the scores of each main evaluation index according to the actual values of each main evaluation index and the preset index ranges of each main evaluation index; obtain the weight values of each main scoring index, and calculate the comprehensive scores of each main scoring index based on the weight values of each main scoring index and the scores of each main scoring index;

[0019] An update module, configured to update the inference strategy based on the comprehensive scores of each main evaluation index. The inference strategy includes the number and processing order of multiple inference requests, and perform language model inference operations according to the updated inference strategy.

[0020] This application also provides a computer device, including: a memory for storing a computer program; a processor for implementing the steps of the language model inference method in the following embodiments when executing the computer program.

[0021] Obtain multiple inference requests to be processed, preprocess the multiple inference requests to be processed, and obtain multiple inference requests;

[0022] Load a language model and perform inferences on multiple inference requests based on the language model;

[0023] Monitor the inference performance information during the language model inference process in real time and display it visually;

[0024] Obtain the types of inference performance information, and determine an inference evaluation system based on the types of inference performance information and the inference performance information. Among them, the inference evaluation system includes multiple main evaluation indicators; the multiple main evaluation indicators include a latency evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator;

[0025] Obtain the actual values of each main evaluation indicator and the preset indicator ranges of each main evaluation indicator;

[0026] Determine the scores of each main evaluation indicator according to the actual values of each main evaluation indicator and the preset indicator ranges of each main evaluation indicator;

[0027] Obtain the weight values of each main scoring indicator, and calculate the comprehensive score of each main scoring indicator based on the weight values of each main scoring indicator and the scores of each main scoring indicator;

[0028] Update the inference strategy based on the comprehensive scores of each main evaluation indicator. The inference strategy includes the number and processing order of multiple inference requests, and perform the language model inference operation according to the updated inference strategy. This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the language model inference method in the following implementation manners are realized.

[0029] Obtain multiple inference requests to be processed, preprocess the multiple inference requests to be processed, and obtain multiple inference requests;

[0030] Load a language model and perform inferences on multiple inference requests based on the language model;

[0031] Monitor the inference performance information during the language model inference process in real time and display it visually;

[0032] Obtain the types of inference performance information, and determine an inference evaluation system based on the types of inference performance information and the inference performance information. Among them, the inference evaluation system includes multiple main evaluation indicators; the multiple main evaluation indicators include a latency evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator;

[0033] Obtain the actual values of each main evaluation indicator and the preset indicator ranges of each main evaluation indicator;

[0034] Determine the scores of each main evaluation indicator according to the actual values of each main evaluation indicator and the preset indicator ranges of each main evaluation indicator;

[0035] Obtain the weight values of each main scoring index, and calculate the comprehensive score of each main scoring index based on the weight values of each main scoring index and the scores of each main scoring index;

[0036] Update the inference strategy based on the comprehensive scores of each main evaluation index. The inference strategy includes the number and processing order of multiple inference requests, and perform language model inference operations according to the updated inference strategy.

[0037] Through the language model inference method provided by this application, since the inference performance information during the language model inference process can be monitored in real time and visualized; obtain the type of inference performance information, determine the inference evaluation system based on the type of inference performance information and the inference performance information, determine the comprehensive scores of each main evaluation index in the inference evaluation system, and update the inference strategy. The inference strategy includes the number and processing order of multiple inference requests. Therefore, the inference strategy can be updated in real time according to the inference performance information to solve the technical problem of the decline in system performance caused by too many inference requests processed simultaneously, and achieve the technical effects of improving inference accuracy and stability. Brief Description of the Drawings

[0038] To more clearly illustrate the embodiments of this application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a flowchart of the language model inference method in the related art;

[0040] Figure 2 It is an application environment diagram of the language model inference method provided by the embodiments of this application;

[0041] Figure 3 It is a flowchart of the language model inference method provided by the embodiments of this application;

[0042] Figure 4 It is a structural block diagram of the language model inference device provided by the embodiments of this application;

[0043] Figure 5 It is an internal structure diagram of the computer device provided by the embodiments of this application. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the protection scope of the present application.

[0045] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0046] As a cutting-edge technology in the field of artificial intelligence, large language model reasoning has made remarkable progress in recent years, profoundly reshaping the boundaries of natural language processing. During the reasoning process, large language models achieve in-depth understanding and analysis of the input text through multi-level attention mechanisms and self-attention mechanisms. They can identify the logical relationships between sentences and perform advanced thinking activities such as causal reasoning and conditional reasoning, thus generating outputs that are logical and close to the human thinking mode. However, the existing technologies also face many challenges when dealing with large language model reasoning tasks. Due to the complex architecture and huge number of parameters of large language models, the demand for computing resources has increased sharply, requiring higher-performance hardware devices (such as GPUs and TPUs, etc.). Affected by factors such as model complexity, data processing, algorithm optimization, computing resource bottlenecks, and insufficient parallel processing, large language models will face problems of high latency and low throughput when performing reasoning tasks. For example, a large language model based on BERT is deployed on a server with limited computing resources. When the server receives a large number of reasoning requests, due to computing resource bottlenecks and insufficient parallel processing, the throughput of the model may drop sharply. This may cause some requests to not be processed in a timely manner and may even lead to service crashes.

[0047] In the field of large language model (LLM) inference, performance evaluation is crucial to ensure that the model achieves the expected results in practical applications. Among them, Latency and throughput are two core metrics for measuring inference performance, which directly reflect the model's response speed and processing power. Latency, that is, the delay, refers to the time required from when the input data is submitted to the model for processing until the model outputs the result. In real-time application scenarios, such as online chatbots or speech recognition systems, low latency is crucial because it determines the smoothness and response speed of the user experience. High latency will cause users to wait too long, affect the interaction experience, and even may lead to task failure. Throughput, on the other hand, refers to the amount of data that the model can process per unit time, which reflects the model's data processing ability and efficiency. High throughput means that the model can process more data in the same time, which is particularly important for application scenarios that need to process a large amount of data, such as large-scale text analysis, content recommendation systems, etc. The improvement of throughput can reduce the total time required to process large-scale data sets, improve the overall work efficiency and resource utilization rate.

[0048] In large language model (LLM) inference technology, the technology of inputting batch requests into the model for inference is called batch inference technology. This technology aims to improve the system's throughput and resource utilization rate by processing multiple inference requests simultaneously, thereby reducing the overall latency and computing cost. The batch inference technology aggregates the inference requests of multiple users. These requests may include various types of tasks such as text generation, question answering, sentiment analysis, etc. The system organizes these requests into one or more batches, and each batch contains a certain number of requests. Preprocess the requests in each batch. The preprocessing steps usually include text tokenization, vectorization, etc. Tokenization is to split the text into smaller units, such as words or sub-words, while vectorization is to convert these units into high-dimensional vectors so that the model can process them. This step ensures the consistency and processability of the input data. Input the preprocessed batch data into the trained large language model, and the model will perform inference tasks on this data and generate corresponding outputs. Since the batch inference technology can process multiple requests simultaneously, it can significantly improve the inference speed.

[0049] Suppose there is a large language model processing text classification requests submitted by users. Without using batch inference technology, the model needs to process each piece of text data submitted by users sequentially and output the corresponding classification results. This approach will lead to frequent model startups and a relatively small amount of data processed each time, thus reducing the inference efficiency. After adopting batch inference technology, the system can merge the text data submitted by multiple users into a single batch and input this batch of data into the model for processing all at once. For example, assume a batch contains 100 pieces of text data. Then the model will process these 100 pieces of text data simultaneously and output the corresponding classification results. In this way, the throughput and resource utilization rate of the model can be significantly improved, thereby reducing the overall latency and computing costs.

[0050] Please refer to Figure 1 , Figure 1 The flowchart for performing large language model inference tasks shown below first batch processes these requests, merging multiple requests into a single batch through an intelligent scheduling algorithm to fully utilize the parallel computing power of the model and improve the inference efficiency. Subsequently, the pre-trained large language model is loaded to perform efficient inference calculations on the merged batch of requests, generating the corresponding output results, and splitting the results into responses corresponding to individual requests to be returned to the client. At the same time, to monitor the inference performance of the large language model in real time, Prometheus is introduced as a tool for collecting and storing performance data. Prometheus captures key performance metrics of the server in real time through its powerful client or custom Exporter, such as CPU usage, memory occupancy, disk I / O, etc., as well as specific metrics such as the response time and processing success rate of inference requests. These performance data are stored in a local time series database for subsequent query and analysis. To more intuitively display the performance data, Grafana is used as a data visualization tool. Grafana is tightly integrated with Prometheus, allowing users to create custom dashboards to display the performance data in the form of charts, panels, etc. Through Grafana, users can intuitively see the inference performance trends and promptly discover potential performance bottlenecks. This process ensures the efficiency and accuracy when performing large language model inference tasks, and also provides strong support for performance monitoring and optimization.

[0051] In current inference technology practices, multi-threading or multi-processing technologies are widely adopted to process multiple inference requests in parallel. This strategy significantly improves the overall processing capacity and response speed of the system. Although multi-threading or multi-processing technologies can process multiple requests in parallel, they also bring potential challenges. When the number of requests processed simultaneously is too large, the inference performance of the model may be affected, specifically manifested as an extended inference time, excessive occupation of CPU and memory resources, and possible system bottlenecks. These negative impacts will not only reduce the real-time response ability of the model but may also affect the accuracy and stability of the inference results.

[0052] Based on historical load data such as GPU / CPU usage, memory occupancy, number of requests, response time, etc., and time series models to predict the load situation in the future for a period of time. According to the prediction results, make adjustment decisions in advance and dynamically adjust the number of requests to optimize resource usage and improve service quality. For example, during peak loads, appropriately reduce the number of requests to relieve system pressure, and during low loads, appropriately increase the number of requests to improve system utilization. Load prediction can identify future load trends in advance, allowing the system to make adjustments in advance and avoid dealing with problems only when the load peaks.

[0053] Load prediction, as a strategy for dynamically adjusting the number of requests, its core lies in predicting the future system load trend through historical load data and formulating adjustment strategies based on this to optimize resource usage and improve service quality. However, the accuracy of load prediction is the key to the success of this method, and it is affected by various complex factors. The integrity and accuracy of historical load data have a crucial impact on the prediction results. If there are missing, incorrect, or outlier values in the historical data, it will cause the prediction model to fail to accurately capture the true load characteristics of the system, leading to prediction biases. In addition, the sampling frequency and granularity of the data are also important factors affecting prediction accuracy. If the sampling frequency is too low or the granularity is too large, it will be difficult to capture the subtle changes in the load, thereby affecting the prediction accuracy.

[0054] In addition, the implementation and application of load prediction methods include challenges in aspects such as accurate interpretation of data, reasonable setting of model parameters, evaluation and optimization of model performance, etc. If not handled properly, it may lead to inaccurate prediction results and even mislead the subsequent decision-making process. Therefore, applying this method in real-time adjustment of data streams has certain risks and complexities.

[0055] As can be seen from the above description, traditional large language model inference methods often lack real-time performance monitoring and adaptive adjustment mechanisms. As a result, when facing continuously changing data input loads, they are unable to effectively balance resource utilization efficiency and service quality. For example, when sending a large number of requests to the model inference simultaneously, a sudden increase in requests may lead to overloading of CPU / GPU resources, extended response times, and even service interruptions. During off-peak hours, resources may be idle, resulting in resource waste. After introducing real-time monitoring and dynamic adjustment mechanisms, it is possible to reduce the number of requests during high loads, relieve resource pressure, and avoid overload. During low loads, the number of requests can be appropriately increased to improve resource utilization and reduce waste. By dynamically adjusting the request processing strategy, key indicators such as the response time and accuracy of the model inference can be ensured to remain within a reasonable range, enhancing the stability and reliability of the service.

[0056] To address the above technical problems, this application provides a language model inference method. By setting up an integrated performance monitoring module, it can capture the inference performance information of the language model in real time (such as response time, throughput, CPU / GPU usage rate, memory occupancy, etc.). Based on this real-time data, it is possible to intelligently and dynamically adjust the number of requests sent to the language model to achieve the purpose of optimizing resource utilization, improving service quality, and enhancing the user experience.

[0057] To enable those skilled in the art of this technical field to better understand the solution of this application, the following further details the application in combination with the accompanying drawings and specific implementation manners.

[0058] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the language model inference method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0059] A language model inference method proposed in this application can be applied, for example, Figure 2In the application environment shown. Among them, the terminal 10 communicates with the server 11 through the network. The above language model inference method can be applied to the server 11, and the server can be an independent server or a server cluster. Specifically, the server 11 can obtain multiple inference requests to be processed, preprocess the multiple inference requests to be processed to obtain multiple inference requests; load the language model, and perform inference on the multiple inference requests based on the language model; monitor the inference performance information during the language model inference process in real time and display it visually; obtain the type of inference performance information, and determine an inference evaluation system based on the type of inference performance information and the inference performance information, where the inference evaluation system includes multiple main evaluation indicators; the multiple main evaluation indicators include a latency evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator; obtain the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator; determine the score of each main evaluation indicator according to the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator; obtain the weight value of each main scoring indicator, and calculate the comprehensive score of each main scoring indicator based on the weight value of each main scoring indicator and the score of each main scoring indicator; update the inference strategy based on the comprehensive score of each main evaluation indicator, the inference strategy includes the number and processing order of multiple inference requests, and perform the language model inference operation according to the updated inference strategy. And display the inference performance information and the updated inference strategy on the terminal 10. The above language model inference method can update the inference strategy in real time according to the inference performance information to solve the technical problem of the decline in system performance caused by too many inference requests to be processed simultaneously, and achieves the technical effects of improving inference accuracy and stability.

[0060] Among them, the terminal 10 can be, but is not limited to, various electronic devices such as personal computers, laptop computers, smart phones, and tablet computers.

[0061] Such as Figure 3 As shown, an embodiment of the present application provides a language model inference method, and the method specifically includes the following steps:

[0062] Step 101: Obtain multiple inference requests to be processed, and preprocess the multiple inference requests to be processed to obtain multiple inference requests.

[0063] Specifically, obtain the to-be-processed text and the to-be-processed text type corresponding to multiple to-be-processed inference requests; divide the to-be-processed text according to the to-be-processed text type to obtain the first to-be-processed text and the second to-be-processed text; perform deletion processing on the first to-be-processed text in the to-be-processed text to obtain the to-be-processed text after deletion; perform replacement processing on the second to-be-processed text in the to-be-processed text after deletion to obtain the target to-be-processed text corresponding to multiple to-be-processed inference requests; determine whether the target to-be-processed text corresponding to multiple to-be-processed inference requests meets the input requirements; use the multiple to-be-processed inference requests corresponding to the target to-be-processed text that meets the input requirements as multiple inference requests.

[0064] The to-be-processed inference requests can be inference requests transmitted from a received external device through a network interface. The inference requests here include inference text information. Usually, the inference requests are sent to an inference model, and the inference model will process the inference text to obtain an inference conclusion. The to-be-processed text types here can include spaces, line breaks, special characters, etc. In this application, spaces and line breaks are set as the first to-be-processed text, and special characters (such as HTML tags, non-printable characters, etc.) are set as the second to-be-processed text. Here, the first to-be-processed text is deleted to ensure the compactness of the to-be-processed text. For the replacement processing of the second to-be-processed text, regular invalid characters can be used to replace the second to-be-processed text. The invalid characters here refer to characters that have no impact on the to-be-processed text to ensure that all to-be-processed requests follow the same format.

[0065] In an implementation manner, determining whether the target to-be-processed text corresponding to multiple to-be-processed inference requests meets the input requirements includes: obtaining the text length of the target to-be-processed text corresponding to multiple to-be-processed requests; comparing the text length of the target to-be-processed text corresponding to multiple to-be-processed requests with a preset length threshold; if the text length of the target to-be-processed text corresponding to multiple to-be-processed requests is greater than the preset length threshold, perform truncation processing on the target to-be-processed text corresponding to multiple to-be-processed requests; if the text length of the target to-be-processed text corresponding to multiple to-be-processed requests is less than the preset length threshold, perform padding processing on the target to-be-processed text corresponding to multiple to-be-processed requests; if the text length of the target to-be-processed text corresponding to multiple to-be-processed requests is equal to the preset length threshold, determine that the target to-be-processed text corresponding to multiple to-be-processed inference requests meets the input requirements.

[0066] The input requirements of the language model can be obtained. The input requirements here can be the data length, data volume, etc. that the language model can input and process. By comparing the text length of the target to-be-processed text with the preset length threshold, the comparison result is obtained, and according to different comparison results, the target to-be-processed text corresponding to multiple to-be-processed requests is truncated or padded to ensure that the length of the multiple to-be-processed requests input is within the range allowed by the language model.

[0067] In this application, an acquisition module can be created, and the acquisition module is responsible for receiving and processing the incoming set of inference requests to be processed. Specifically, the acquisition module first receives multiple inference requests to be processed from the user or the front-end system. These inference requests usually contain the text data to be processed or other relevant parameters. Subsequently, the acquisition module preprocesses these requests, including necessary steps such as data cleaning, format conversion, text alignment, etc., to ensure that the input data meets the input requirements of the language model. The multiple requests to be processed obtained after processing are sent to the language model in an orderly manner for subsequent inference calculations.

[0068] In a specific implementation, QtDesigner can be used to design the GUI (Graphical User Interface). This includes adding controls such as input boxes (QLineEdit), sliders (QSlider), buttons (QPushButton), etc., and we can design an intuitive and easy-to-use user interaction interface. Among them, the title bar: located at the top of the interface, displays title information such as "Language Model Inference Request Input". Inference request input area: Text input box: Provides multiple text boxes to allow users to input inference requests to be inferred. The serial number or inference request type can be marked next to each text box for users to distinguish. File upload button: Provides a file upload function, supporting users to upload files containing multiple inference requests (such as CSV, Excel, etc.), and the system will automatically parse and fill them into the text input boxes. Parameter setting area: Provides necessary parameter setting options, such as model selection, inference mode (real-time / batch), timeout time, etc. Request preview area: Displays the preview of multiple inference requests that the user has input or uploaded, including text content and parameter settings. The preprocessing of inference requests includes necessary steps such as data cleaning, format conversion, text alignment, etc., to ensure that the input data meets the input requirements of the language model. Submit button: Located at the bottom of the interface, after the user clicks, the acquisition module will preprocess the inference requests and send them to the language model for inference. Status display area: Displays the status information of the inference request processing, such as "Processing", "Processing completed", "Error occurred", etc., and provides necessary error prompts or processing suggestions.

[0069] In one implementation, a code example for creating the acquisition module interface and basic functions using python and PyQt5:

[0070]

[0071]

[0072]

[0073]

[0074]

[0075] The pseudo - code of the data pre - processing process is as follows:

[0076]

[0077]

[0078] Step 102: Load the language model and perform inferences on multiple inference requests based on the language model.

[0079] The language model in this application is specifically a large - language model. Large - language models (LLMs for short) are a type of deep - learning model trained with a large amount of data, aiming to understand and generate natural - language text. These models are usually built based on deep - neural - network architectures. For example, the Transformer architecture can handle long - range dependencies through self - attention mechanisms, thereby capturing the complexity and diversity of language. LLMs have a large number of parameters, which enables them to learn rich language features and patterns, and can be generalized to a variety of different tasks and adapt to different application scenarios, such as translation, summarization, question - answering, etc. It is precisely because of the above characteristics that large - language models have gradually become the object of extensive research in various fields.

[0080] In this application, an inference module is set up. After the inference module obtains multiple inference requests transmitted by the acquisition module, the inference module can start the inference service of loading the large - language model. Specifically, the inference module is first responsible for loading the pre - trained large - language model and configuring it to a state where it can receive input data for inference. Once the module is started, it will enter the standby state, ready to receive data from the acquisition module. After receiving multiple inference requests, the inference module will use the loaded large - language model to perform inferences. After the inference is completed, the inference module will return the generated results to the system or the user. These results may include generated text, classification labels, sentiment tendencies, etc., depending on the training objectives and task types of the model. It can be understood that for the same text with different inference tasks, the inference results obtained from the inference operations are different. For example, in a text - generation task, the inference module may return a generated text that is relevant and coherent to the input text; in a sentiment - analysis task, it may return a label representing the sentiment tendency of the text (such as positive, negative, or neutral).

[0081] The following is a code example for creating an inference module:

[0082]

[0083]

[0084]

[0085]

[0086]

[0087] Step 103: Monitor the inference performance information during the inference process of the language model in real time and display it visually.

[0088] This application can achieve real-time monitoring during the inference process of the language model by setting up a monitoring module, obtain the inference performance information during the inference process of the language model, and convert the inference performance information into a visual chart form for display, so that users can view the inference performance information simply and clearly.

[0089] The monitoring module aims to comprehensively and real-time monitor the inference performance of the language model, and visually display the inference performance information in the form of a chart. The inference performance information here can be used to set various key performance indicators, and then evaluate and update the inference strategy. Exemplarily, the inference performance information can include information such as latency, throughput, CPU usage, memory occupancy, GPU utilization, etc. Use professional performance monitoring tools (such as Prometheus, Grafana, etc.) or monitoring services provided by cloud service providers to obtain and visually display the inference performance information.

[0090] In one embodiment, the monitoring module in this application can adjust the sampling frequency of the language model based on the performance parameter indicators during the inference process of the language model.

[0091] Among them, adjusting the sampling frequency of the language model based on the performance parameter indicators during the inference process of the language model includes: obtaining the performance parameter indicators, where the performance parameter indicators include the sampling frequencies of multiple initial language models for multiple inference requests, multiple data collection durations, and multiple data collection amounts; calculating the standard deviation of the first performance parameter indicator based on the sampling frequencies of multiple initial language models for multiple inference requests; calculating the standard deviation of the second performance parameter indicator based on the multiple data collection durations of multiple inference requests; calculating the standard deviation of the third performance parameter indicator based on the multiple data collection amounts of multiple inference requests; determining the data fluctuation value based on the standard deviation of the first performance parameter indicator, the standard deviation of the second performance parameter indicator, and the standard deviation of the third performance parameter indicator; obtaining the data collection amount threshold and the data fluctuation threshold; and adjusting the sampling frequency of the language model based on the data fluctuation value, the data collection amount threshold, and the data fluctuation threshold.

[0092] The sampling frequency here refers to the number of samples extracted from a continuous signal and composed into a discrete signal by a large language model per unit time, which is expressed in Hertz (Hz). The reciprocal of the sampling frequency is the sampling period or sampling time, which is the time interval between samplings. Generally speaking, the sampling frequency refers to the number of signal samples that a computer can collect per unit time. The standard deviation of the first performance parameter index is the standard deviation calculated from the sampling frequencies of multiple initial large language models for multiple inference requests; the standard deviation of the second performance parameter index is the standard deviation calculated from the collection data durations of multiple inference requests; the standard deviation of the third performance parameter index is the standard deviation calculated from the data collection amounts of multiple inference requests. The standard deviation of each calculated performance parameter index can be set as the data fluctuation value of volatility. The data collection amount threshold and the data fluctuation threshold are preset, and the actual data collection amount is compared with the data collection amount threshold, and the above calculated data fluctuation value is compared with the data fluctuation threshold. When the actual data collection amount is greater than the data collection amount threshold or the processing efficiency is too low, the sampling frequency will be reduced. When the data fluctuation value is much smaller than the data fluctuation threshold, that is, when the data fluctuation is considered too smooth, the sampling frequency can be increased. In this way, when collecting performance data, a reasonable sampling frequency is set to ensure that both the details of performance changes can be captured and the inference of the large language model will not be burdened additionally due to overly frequent data collection.

[0093] Please refer to the following code to understand the process of adjusting the sampling frequency:

[0094]

[0095]

[0096]

[0097] After each adjustment, run the large language model again and collect performance data (inference performance information), and then analyze and evaluate the suitability of the sampling frequency again.

[0098] This module can be set in the top navigation bar, with "Inference Performance Monitoring" displayed as the module title. The refresh button allows users to manually refresh the data to obtain the latest inference evaluation metrics. The settings button provides access to configuration options such as changing the sampling frequency, selecting monitoring metrics, etc. The monitoring metric list lists all monitored inference performance metrics such as latency, throughput, CPU usage, memory occupancy, GPU utilization, etc. Users can view the detailed charts or data of a metric by clicking on the metric name. The main monitoring area uses chart forms such as line charts, bar charts, or pie charts to display the changing trends of metrics over time in real-time. Users can switch the displayed charts by clicking on the items in the monitoring metric list.

[0099] The following is a code example for obtaining real-time inference performance information:

[0100]

[0101]

[0102]

[0103]

[0104] Step 104: Obtain the type of inference performance information, and determine an inference evaluation system based on the type of inference performance information and the inference performance information. Among them, the inference evaluation system includes multiple main evaluation indicators; the multiple main evaluation indicators include a latency evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator.

[0105] When conducting the evaluation, multiple main evaluation indicators can be set. Specifically, in this application, the set main evaluation indicators include a latency evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator as the main indicators in the indicator system. The latency evaluation indicator can cover factors such as the data volume size and data processing speed that have an impact on the latency evaluation comparison; the processor evaluation indicator can cover factors such as CPU utilization, CPU temperature, GPU utilization, GPU temperature, memory usage, and storage capacity usage that have an impact on the processor evaluation indicator; the throughput evaluation indicator can cover factors such as response time, network status, and data processing speed that have an impact on the throughput evaluation indicator.

[0106] In this application, considering the impact of the latency evaluation indicator, the throughput evaluation indicator, and the processor evaluation indicator on the running state of the inference operation, an inference evaluation system is constructed based on the three dimensions of the latency evaluation indicator, the throughput evaluation indicator, and the processor evaluation indicator, which can achieve a more detailed evaluation.

[0107] Step 105: Obtain the actual values of each main evaluation indicator and the preset indicator range of each main evaluation indicator. Step 106: Determine the score of each main evaluation indicator according to the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator.

[0108] Specifically, determine the types of the preset index ranges corresponding to the actual values of each main evaluation index. The preset index range types of each main evaluation index include a first preset index range, a second preset index range, and a third preset index range. If the actual value of the main evaluation index corresponds to the first preset index range, the score of the main evaluation index is the preset value. If the actual value of the main evaluation index corresponds to the second preset index range, calculate the score of the main evaluation index based on the first calculation formula. If the actual value of the main evaluation index corresponds to the third preset index range, calculate the score of the main evaluation index based on the second calculation formula.

[0109] The actual value of each main evaluation index here is the actually obtained value of the current main evaluation index. In this application, three types of preset index ranges are set for each main evaluation index, including a first preset index range, a second preset index range, and a third preset index range. Among them, the first preset index range corresponds to the ideal range. When the actual value of the main evaluation index is within the first preset index range, it is considered that the main evaluation index is within the ideal range. At this time, the main evaluation index is set to the preset value. The preset value here can be 100. Of course, the preset value here can also be other values, as long as it represents an ideal state. The second preset index range corresponds to the warning range. At this time, the score of the main evaluation index needs to be calculated based on the first calculation formula. When the actual value of the main evaluation index is within the second preset index range, it is considered that the main evaluation index is in a position close to danger. The third preset index range corresponds to the danger range. At this time, the score of the main evaluation index needs to be calculated based on the second calculation formula. When the actual value of the main evaluation index is within the third preset index range, it is considered that the main evaluation index is in a dangerous position. It should be noted that the first preset index range, the second preset index range, and the third preset index range here are the settings of the actual ideal range, warning range, and danger range corresponding to the actual value of each main evaluation index.

[0110] Among them, if the actual value of the main evaluation index corresponds to the second preset index range, calculating the score of the main evaluation index based on the first calculation formula includes:

[0111] Obtain the first target upper limit value, the first target lower limit value, and the first calculation formula corresponding to the second preset index range; calculate the score of the main evaluation index according to the first target upper limit value, the first target lower limit value, and the first calculation formula;

[0112] Among them, the first calculation formula is as follows:

[0113]

[0114] Among them, D represents the score of the main evaluation index, X represents the actual value of the main evaluation index, G represents the lower limit value of the alarm score range, A represents the first target upper limit value, B represents the first target lower limit value, and Y represents the number of elements in the alarm score range.

[0115] Among them, if the actual value of the main evaluation index corresponds to the third preset index range, calculating the score of the main evaluation index based on the second calculation formula includes: obtaining the first target lower limit value corresponding to the second preset index range, the second target lower limit value corresponding to the third preset index range, and the second calculation formula; calculating the score of the main evaluation index according to the first target lower limit value, the second target lower limit value, and the second calculation formula;

[0116] Among them, the second calculation formula is as follows:

[0117]

[0118] Among them, D represents the score of the main evaluation index, X represents the actual value of the main evaluation index, B represents the first target lower limit value, C represents the second target lower limit value, and Z represents the number of elements in the danger score range.

[0119] In a specific implementation manner, taking the calculation of the score of the delay evaluation index as an example for description, assume that the preset index ranges of the defined delay evaluation index are as follows:

[0120] The first preset index range: < 50ms;

[0121] The second preset index range: 50ms to 100ms;

[0122] The third preset index range: greater than 100ms;

[0123] The actual value of the obtained delay evaluation index is 70ms, which is within the preset index range of the second type. Then, the score of the delay evaluation index is calculated according to the first calculation formula as follows:

[0124]

[0125] At this time, the calculated score of the delay evaluation index is 70. The preset index range of the delay evaluation index in this application is only for example, and specific numerical changes can be made according to the actual situation.

[0126] It should be noted that the above-mentioned first preset index range, second threshold index range, and preset index range refer to the actual values of each main evaluation index within the corresponding ideal range, warning range, and danger range. In order to unify the comprehensive scores of multiple main evaluation indexes under a unified measurement unit, in this application, the comprehensive scores of multiple calculated main evaluation indexes will be mapped to the same percentage system or decimal system, etc., under the scoring rules of the scoring system, that is, calculate the mapping percentage system or decimal system scoring rules, including the scoring of the safety range, the scoring of the warning range, and the scoring of the danger range. In this way, it is convenient to generate an update strategy according to the comprehensive scores of each main evaluation index later.

[0127] Step 107: Obtain the weight values of each main scoring index, and calculate the comprehensive score of each main scoring index based on the weight values of each main scoring index and the scores of each main scoring index.

[0128] In this application, a weight is assigned to each inference performance data including the data corresponding to the main evaluation index to represent its importance in the decision-making process. For example, the latency evaluation index may have a higher weight because it directly affects the user experience. Calculate the score of each performance data to obtain a comprehensive score. Set a comprehensive threshold range for the comprehensive score corresponding to each main evaluation index. Determine whether to increase, decrease, or keep the current request quantity unchanged by judging whether the comprehensive score corresponding to each main evaluation index is within the comprehensive threshold range. The comprehensive threshold range here is determined according to the floating range of the numerical values of the comprehensive thresholds of each main evaluation index. The numerical values of the comprehensive thresholds of the main evaluation indexes here can be set according to the actual situation and are not limited here. According to the deviation degree of the final score, dynamically adjust the step size of the request quantity. For example, when the system is approaching a dangerous state, the step size of reducing the request quantity should be larger. Continuously optimize the adjustment strategy according to the current performance trend.

[0129] In practical applications, after comparing the comprehensive scores of multiple main evaluation indicators with their corresponding comprehensive threshold ranges, there may be some main evaluation indicators that do not meet the performance requirements. The difference ratio between the comprehensive threshold of each main evaluation indicator that does not meet the performance requirements and the comprehensive score of the current main evaluation indicator can be calculated. Here, the difference ratio can be obtained by dividing the value obtained by subtracting the corresponding comprehensive threshold from the comprehensive score of the current main evaluation indicator by the comprehensive threshold. After obtaining the difference ratio of each main evaluation indicator, the sum of the absolute value of the maximum and the absolute value of the minimum of the difference ratios of the main evaluation indicators that do not meet the performance requirements can be calculated. When this sum exceeds a certain value, it is considered that the numerical span of the inference request quantity needs to be adjusted among different main evaluation indicators. The difference ratio corresponding to the main evaluation indicator with the largest absolute value of the difference ratio is obtained as the target difference ratio, and the inference request is updated according to the absolute value of the target difference ratio. Exemplarily, the number of updated inference requests can be: Updated inference request quantity = Current inference request quantity × (1 ± Target difference ratio). When this sum does not exceed a certain ratio, it is considered that the numerical span of the inference request quantity needs to be adjusted among different main evaluation indicators is small, and the number of inference requests can be updated according to the average value of the difference ratios corresponding to each main evaluation indicator. Exemplarily, the number of updated inference requests can be: Updated inference request quantity = Current inference request quantity × (1 ± Average value of the difference ratios). In this embodiment, when the number of inference requests needs to be increased, the plus sign is used in the formula, and when the number of inference requests needs to be decreased, the minus sign is used in the formula.

[0130] In one embodiment, after obtaining the scores of each main scoring indicator through the above steps, the comprehensive score of each main scoring indicator can be obtained by performing a weighted operation based on the score of each main scoring indicator and its corresponding weight value of the main scoring indicator. Finally, the weighted sum of each main scoring indicator is the final score, and the weight of the main scoring indicator here can be set according to the degree of influence of the main scoring indicator on the performance. For example, the greater the influence degree of the main scoring indicator on the performance, the greater the weight value set for it, and the lower the influence degree of the main scoring indicator on the performance, the smaller the weight value set for it.

[0131] In a feasible embodiment, the factors that affect each main evaluation indicator can be used as sub-evaluation indicators, and an inference evaluation tree system can be constructed based on the main evaluation indicators and sub-evaluation indicators.

[0132] Specifically, the inference performance information can be parsed based on the type of inference performance information to obtain the target inference performance information type corresponding to the inference performance information and the related inference performance information types related to the target inference performance information type; taking the target inference performance information type as the main node and the related inference performance information types related to the target inference performance information type as the sub-nodes, an inference evaluation system tree is constructed.

[0133] The inference information type here refers to the type of inference performance information that affects the inference performance of the system. Since different inference tasks have different attentions to different types of inference performance, the target inference performance information type that the current inference task pays more attention to needs to be selected here, which is equivalent to selecting the type of the main evaluation index. The related inference performance information type is the inference performance information type that affects the target inference performance information type, which is equivalent to the type of the sub-evaluation index that affects the main evaluation index. The inference evaluation system tree here can be constructed based on tree structures such as binary trees or multi-way trees, etc. Applying the inference evaluation system tree can organize the data into a hierarchical structure, making data management and retrieval more intuitive and convenient.

[0134] In a feasible implementation manner, the first judgment matrix of each main evaluation index and the second judgment matrix of each sub-evaluation index can be constructed respectively; based on the first judgment matrix and the second judgment matrix, calculate the comprehensive importance degree of each sub-evaluation index corresponding to each main evaluation index relative to other sub-evaluation indexes and the comprehensive importance degree of each main evaluation index relative to other main evaluation indexes; calculate the expectations of each sub-evaluation index and each main evaluation index according to the comprehensive importance degree and the optimism coefficient; determine the first weight of each sub-evaluation index in each main evaluation index and the second weight of each main evaluation index according to the proportion of the expectation of each sub-evaluation index in the sum of the expectations of the sub-evaluation indexes and the proportion of the expectation of each main evaluation index in the sum of the expectations of the main evaluation indexes.

[0135] Specifically, in traditional analytic hierarchy process, the judgment matrix is determined by pairwise comparison of the selected ones, and it is required that the elements in the judgment matrix are exact numbers. However, in practical applications, it is very difficult to represent the results of pairwise importance comparison of elements with exact numbers. Therefore, in this embodiment, the fuzzy theory is adopted when constructing the judgment matrix, that is, at this time, fuzzy concepts such as slightly, obviously, strongly, etc. can be used to represent the element comparison results, and then based on the judgment matrix and the weighting method, the first weight of the sub-indexes under each main main index is calculated, and the second weight of the main index is calculated. Among them, the first weight is determined by the judgment matrix constructed by pairwise comparison of the sub-indexes under each main index, and the second weight is determined by the judgment matrix constructed by pairwise comparison of the three main indexes.

[0136] The main evaluation index can be rated based on the sub - scores of each sub - evaluation index and the second weights of each sub - evaluation index to obtain a rating result; an evaluation matrix is constructed according to the rating result; a synthesis operation is performed based on each first weight and each evaluation matrix to obtain a synthesis operation result; and the comprehensive score of each main evaluation index is determined based on the synthesis operation result.

[0137] The sub - score of the sub - evaluation index here can be the actual data of the sub - evaluation index. The second weight is the influence of each sub - evaluation index on its corresponding main evaluation index. The sub - score corresponding to the sub - evaluation index and the second weight are multiplied to obtain the rating of each sub - evaluation index, resulting in a rating result. An evaluation matrix is constructed according to the rating result, and then the first weight is multiplied with each evaluation matrix to further obtain the comprehensive score of each main evaluation index. Subsequently, the inference strategy is adjusted according to the comprehensive score of each main evaluation index.

[0138] In this application, a determination module is set up to adjust the number of requests sent to the inference service according to the monitoring information. This module continuously analyzes the usage of key resources such as CPU, memory, GPU, etc., as well as performance indicators such as the response time and throughput of the inference service, providing a decision - making basis for the dynamic adjustment strategy module. Preset KPI thresholds are used to monitor and evaluate the performance indicators of the inference service in real - time. According to the comparison result between the preset threshold and the comprehensive score of the main evaluation index, the adjustment strategy is selectively triggered. The adjustment result is sent to the acquisition module to adjust information such as the number of input requests.

[0139] The following is a code example for creating the determination module:

[0140]

[0141]

[0142]

[0143]

[0144] The pseudo - code for making decisions based on performance data is as follows:

[0145]

[0146]

[0147]

[0148]

[0149] Step 108: Update the inference strategy based on the comprehensive scores of each main evaluation indicator. The inference strategy includes the number and processing order of multiple inference requests, and perform language model inference operations according to the updated inference strategy.

[0150] In this application, the indicator threshold of the final score can be preset in advance, and the value corresponding to the calculated final score is compared with the indicator threshold of the final score. Once it is detected that the performance indicator exceeds the threshold, the module will automatically formulate or adjust the request processing strategy. These strategies may include reducing the number of concurrent requests, preferentially processing high-priority requests, etc. Adjust the requests according to the strategy. At the same time, the module will display the effect after the implementation of the updated strategy in the monitoring module to ensure that the performance indicators of the inference service are restored within the preset threshold. Based on the monitoring results after the implementation of the strategy, the module will optimize and iterate the strategy to adapt to the changing workload and resource environment. Therefore, this scheduling strategy is simple and easy to implement, can quickly implement resource monitoring and adjustment of the scheduling strategy; has strong real-time performance, can be directly applied to real-time monitoring systems without waiting for historical data accumulation or model training to complete. This method can instantly judge whether the system state exceeds the preset range and immediately trigger the corresponding scheduling strategy. By calculating the comprehensive scores of each main scoring indicator, then adding up the comprehensive scores of each main scoring indicator to obtain a final score, and then comparing the final score with the indicator threshold, it is possible to consider the influence of multiple main scoring indicators while avoiding a large amount of calculations. In this way, an update plan can be obtained more quickly and the complexity of adjustment is reduced.

[0151] The present invention provides a language model inference method, which realizes comprehensive insight and dynamic adjustment of the model inference performance by real-time monitoring the usage of key hardware resources and the indicators corresponding to the core inference performance information during the inference process of the large language model. Specifically, by continuously tracking the usage status of hardware resources such as CPU / GPU usage rate and memory occupancy rate, and closely monitoring inference performance indicators such as response time, throughput, and latency, combined with the built-in updated strategy, it is possible to intelligently and dynamically adjust the number of requests sent to the model according to the real-time captured performance data, ensuring that each request can be processed in a timely manner and avoiding inference delays caused by resource competition or excessive load. This mechanism not only gives the system the insight to identify inference bottlenecks, but also enables it to flexibly adjust request allocation according to the current load situation and resource status to achieve multiple effects of optimizing resource utilization, reducing inference latency, and increasing throughput.

[0152] Moreover, through an integrated design, the complex process of inference optimization and adjusting the number of requests is encapsulated under a simple and user-friendly interface. Users do not need to delve into the underlying technical details. They only need to simply input the necessary information or content, and the system can automatically complete the subsequent inference process and immediately return accurate and high-quality answers. Such a system greatly simplifies the complexity and tediousness of the traditional large language model inference process, providing users with unprecedented efficiency and convenience.

[0153] In one embodiment, although the number of iterations (i.e., the output length) of an inference job cannot be known in advance, the execution time of each iteration can be predicted. The iteration time is mainly determined by some key parameter information, including: the hardware parameter information of the GPU cluster, the number of model parameters information, and the job input length. When the job analysis module receives a target inference request, it can obtain this information through the information acquisition sub-module, and then use the analysis sub-module to predict the time required for the first iteration based on the above information.

[0154] Due to the autoregressive mode of large language model inference, the input length for the first iteration determines the execution time for generating the first output character. For a job with a long input and a short output, the execution time for outputting the first character may dominate the entire job's execution time. Specifically, the time for the first iteration (i.e., the execution time for generating the first output character) is usually longer than that of subsequent iterations, and as the length of the input sequence increases, the time for the first iteration increases approximately linearly, while the increase in the time for subsequent iterations can be ignored. This is achieved through cache optimization in memory management. In the first iteration, the key-value tensors of all input characters need to be calculated and cached, which takes a longer time. In the subsequent iterations, only the key-value tensors of the newly generated characters in the input characters need to be calculated, and the other input characters are directly loaded from the key-value cache of the memory management module, so the number of characters to be calculated becomes smaller, and the time required for the iteration also becomes correspondingly smaller.

[0155] In this embodiment, the job information can be obtained by executing jobs with different input lengths and output lengths in advance multiple times and sampling. The job information includes the total time spent on executing the job with the corresponding output length. The data obtained from multiple executions is used as a training dataset to complete the training of the job analysis module. Thus, by leveraging the semi-information-agnostic feature of the large language model inference service, for the actual inference jobs submitted later, although the total number of iterations cannot be determined, the time for each iteration can be accurately predicted through the information of the job analysis module.

[0156] Finally, according to the predicted execution time, determine the target priority queue that the target inference job request needs to enter. In this way, according to different predicted execution times, multiple inference requests can be scheduled into queues with different priorities, and inference operations can be performed in batches according to queues with different priorities, which can improve the efficiency of large language model inference operations.

[0157] An embodiment of the present application provides a language model inference device. The language model inference device is specifically as Figure 4 shown. The language model inference device includes: an acquisition module 20, an inference module 21, a monitoring module 22, a determination module 23, an update module 24, and a large language inference model.

[0158] The acquisition module 20 is used to acquire multiple inference requests to be processed, preprocess the multiple inference requests to be processed, and obtain multiple inference requests;

[0159] The inference module 21 is used to load the language model and perform inferences on multiple inference requests based on the language model;

[0160] The monitoring module 22 is used to monitor the inference performance information during the language model inference process in real time and display it visually;

[0161] The determination module 23 is used to obtain the type of inference performance information, determine an inference evaluation system based on the type of inference performance information and the inference performance information. Among them, the inference evaluation system includes multiple main evaluation indicators; the multiple main evaluation indicators include a latency evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator; obtain the actual values of each main evaluation indicator and the preset indicator ranges of each main evaluation indicator; determine the scores of each main evaluation indicator according to the actual values of each main evaluation indicator and the preset indicator ranges of each main evaluation indicator; obtain the weight values of each main scoring indicator, and calculate the comprehensive scores of each main scoring indicator based on the weight values of each main scoring indicator and the scores of each main scoring indicator;

[0162] The update module 24 is used to update the inference strategy based on the comprehensive scores of each main evaluation indicator. The inference strategy includes the number and processing order of multiple inference requests, and perform language model inference operations according to the updated inference strategy.

[0163] For the description of the features in the corresponding embodiment of the language model inference device, reference can be made to the relevant description in the corresponding embodiment of the language model inference method, which will not be elaborated here one by one.

[0164] An embodiment of the present application also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the language model inference method.

[0165] Embodiments of the present application further provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above-described embodiments of the language model inference method when running.

[0166] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0167] Those skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0168] The above has introduced in detail a language model inference method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A language model inference method, characterized in that: include: Acquire multiple inference requests to be processed, and pre-process the multiple inference requests to be processed to obtain multiple inference requests; Load the language model and perform reasoning on multiple reasoning requests based on the language model; Real-time monitoring of the inference performance information during the language model inference process and visualization; Obtaining the type of inference performance information, and determining an inference evaluation system based on the type of inference performance information and the inference performance information, wherein the inference evaluation system includes a plurality of main evaluation indicators; the plurality of main evaluation indicators include a delay evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator; Obtaining the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator; Determine the score of each main evaluation indicator according to the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator; Obtaining weight values ​​of the respective main scoring indicators, and calculating comprehensive scores of the respective main scoring indicators based on the weight values ​​of the respective main scoring indicators and the scores of the respective main scoring indicators; An inference strategy is updated based on the comprehensive scores of the main evaluation indicators, the inference strategy including the number and processing order of multiple inference requests, and a language model inference operation is performed according to the updated inference strategy.

2. The language model inference method according to claim 1, characterized in that: The preprocessing of the to-be-processed reasoning request to obtain a plurality of reasoning requests comprises: Obtaining the pending texts and the types of the pending texts corresponding to the multiple pending inference requests; Dividing the texts to be processed according to the types of the texts to be processed to obtain a first text to be processed and a second text to be processed; deleting the first text to be processed in the texts to be processed to obtain a deleted text to be processed; Replacing the second to-be-processed text in the deleted to-be-processed text to obtain target to-be-processed texts corresponding to the multiple to-be-processed inference requests; Determine whether the target to-be-processed texts corresponding to the multiple to-be-processed inference requests meet the input requirements; The multiple inference requests to be processed corresponding to the target text to be processed that meets the input requirement are taken as multiple inference requests.

3. The language model inference method according to claim 2, characterized in that: The step of determining whether the target to-be-processed texts corresponding to the multiple to-be-processed inference requests meet the input requirements includes: Get the text length of the target to-be-processed text corresponding to the multiple to-be-processed requests; Comparing the text length of the target to-be-processed text corresponding to the plurality of to-be-processed requests with a preset length threshold; If the text length of the target to-be-processed text corresponding to the multiple to-be-processed requests is greater than a preset length threshold, the target to-be-processed text corresponding to the multiple to-be-processed requests is truncated; If the text length of the target to-be-processed text corresponding to the multiple to-be-processed requests is less than the preset length threshold, padding processing is performed on the target to-be-processed text corresponding to the multiple to-be-processed requests; If the text length of the target text to be processed corresponding to the multiple pending requests is equal to the preset length threshold, it is determined that the target text to be processed corresponding to the multiple pending inference requests meets the input requirement.

4. The language model inference method according to claim 1, characterized in that: Determining the score of each main evaluation indicator according to the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator includes: Determine the type of the preset indicator range of each main evaluation indicator corresponding to the actual value of each main evaluation indicator, wherein the preset indicator range type of each main evaluation indicator includes a first preset indicator range, a second preset indicator range, and a third preset indicator range; If the actual value of the main evaluation indicator corresponds to the first preset indicator range, the score of the main evaluation indicator is the preset value; If the actual value of the main evaluation indicator corresponds to the second preset indicator range, calculating the score of the main evaluation indicator based on the first calculation formula; If the actual value of the main evaluation indicator corresponds to the third preset indicator range, the score of the main evaluation indicator is calculated based on the second calculation formula.

5. The language model inference method according to claim 4, characterized in that: If the actual value of the main evaluation indicator corresponds to the second preset indicator range, calculating the score of the main evaluation indicator based on the first calculation formula includes: Obtaining a first target upper limit value, a first target lower limit value, and a first calculation formula corresponding to a second preset indicator range; Calculate the score of the main evaluation indicator according to the first target upper limit value, the first target lower limit value and the first calculation formula; The first calculation formula is as follows: Wherein, D represents the score of the main evaluation indicator, X represents the actual value of the main evaluation indicator, G represents the lower limit of the alarm score range, A represents the first target upper limit, B represents the first target lower limit, and Y represents the number of elements in the alarm score range.

6. The language model inference method according to claim 4, characterized in that: If the actual value of the main evaluation indicator corresponds to the third preset indicator range, calculating the score of the main evaluation indicator based on the second calculation formula includes: Obtain a first target upper limit value and a first target lower limit value corresponding to the second preset indicator range, a second target lower limit value and a second calculation formula corresponding to the third preset indicator range; Calculate the score of the main evaluation indicator according to the first target upper limit value, the first target lower limit value, the second target lower limit value and the second calculation formula; The second calculation formula is as follows: Where D represents the score of the main evaluation index, X represents the actual value of the main evaluation index, B represents the lower limit of the first target, C represents the lower limit of the second target, and Z represents the number of elements in the risk score range.

7. The language model inference method according to claim 1, characterized in that: The method further includes: adjusting the sampling frequency of the language model based on the performance parameter index in the language model reasoning process: The adjusting of the sampling frequency of the language model based on the performance parameter index in the language model reasoning process includes: Acquire performance parameter indicators, where the performance parameter indicators include sampling frequencies of multiple initial language models of multiple inference requests, multiple data collection durations, and multiple data collection amounts; Calculate a first performance parameter indicator standard deviation based on sampling frequencies of multiple initial language models of multiple inference requests; Calculate a second performance parameter indicator standard deviation based on multiple data collection durations of multiple reasoning requests; Calculating a third performance parameter indicator standard deviation based on the plurality of data collection amounts of the plurality of reasoning requests; Determine a data fluctuation value based on the first performance parameter indicator standard deviation, the second performance parameter indicator standard deviation, and the third performance parameter indicator standard deviation; Obtain data collection volume threshold and data fluctuation threshold; The sampling frequency of the language model is adjusted based on the data fluctuation value, the data collection amount threshold, and the data fluctuation threshold.

8. A language model inference device, characterized in that: include: An acquisition module, used for acquiring a plurality of pending inference requests, and preprocessing the plurality of pending inference requests to obtain a plurality of inference requests; The inference module is used to load the language model and perform inference on multiple inference requests based on the language model; The monitoring module is used to monitor the inference performance information of the language model inference process in real time and visualize it; A determination module is used to obtain the type of reasoning performance information, and determine the reasoning evaluation system based on the reasoning performance information type and the reasoning performance information, wherein the reasoning evaluation system includes multiple main evaluation indicators; the multiple main evaluation indicators include a delay evaluation indicator, a throughput evaluation indicator, and a processor evaluation indicator; obtain the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator; determine the score of each main evaluation indicator according to the actual value of each main evaluation indicator and the preset indicator range of each main evaluation indicator; obtain the weight value of each main scoring indicator, and calculate the comprehensive score of each main scoring indicator based on the weight value of each main scoring indicator and the score of each main scoring indicator; an update module is used to update the reasoning strategy based on the comprehensive score of each main evaluation indicator, the reasoning strategy includes the number of multiple reasoning requests and the processing order, and perform language model reasoning operations according to the updated reasoning strategy.

9. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the language model inference method as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the language model inference method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Commodity information reasoning method and device

    CN115374845A

  • Data model monitoring method and device, processor and electronic equipment

    CN115757035A

  • GPT question and answer model optimization method and device, electronic equipment and storage medium

    CN117035097A

  • Network security score determination method and device, electronic equipment and storage medium

    CN118018279A

  • Model visual reasoning method and device, electronic equipment and storage medium

    CN118898296A

Cited By

  • Service performance evaluation method, computer program product and electronic device

    CN120723342A

  • Service performance evaluation method, computer program product, and electronic device

    CN120723342B

  • Efficiency evaluation method and device of model reasoning system and electronic equipment

    CN121434705A

  • Efficiency evaluation methods, devices, and electronic equipment for model inference systems

    CN121434705B