Availability evaluation method and system for chat large model service
By constructing a sponge sample test dataset based on gradient and word block mapping, and optimizing the adversarial suffix using backpropagation and loss functions, the problem of text discreteness and autoregressive output uncertainty of the chat model is solved, and effective evaluation of its availability and dynamic monitoring of resource efficiency is achieved.
Patent Information
- Application Number
- CN202510780866.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The prior art cannot effectively construct sponge samples suitable for chat big models, making it difficult to evaluate their availability under high load pressures, and the text discreteness and autoregressive output uncertainty of chat big models make it impossible to directly apply traditional methods.
By constructing a sponge sample test dataset based on the relationship between gradient and word block mapping, the adversarial suffix is generated using backpropagation and discrete text optimization algorithms, combining EOS loss and uniform loss functions, the adversarial suffix is optimized to generate sponge samples suitable for chat big models, and its overhead is evaluated in online services.
Effectively evaluate the service stability and resource efficiency of the chat model, reduce the uncertainty of model output, increase the diversity and length of output content, improve the trigger accuracy of the adversarial samples, and ensure the availability of services in extreme cases.
Smart Images

Figure CN120297424A_ABST
Abstract
Description
Technical Field
[0001] A method and system for evaluating the usability of chat large model services, which are used to evaluate the usability of chat large model services, and belong to the technical fields of artificial intelligence security and large model security. Background Art
[0002] Chat large model: A chat large model is a highly optimized natural language processing system. Its core architecture and training process involve multiple key stages, including pre-training of the base model, fine-tuning of dialogue instructions, and alignment of human feedback reinforcement learning. The base model is usually based on a decoder-only architecture and is pre-trained on a large-scale text dataset. Although the base model has powerful language generation capabilities, it may lack task adaptability and controllability when directly used for conversations. Therefore, it needs to be adapted to the conversation scenario through supervised fine-tuning in a dialogue dataset. The model after instruction fine-tuning may still generate content that does not conform to human preferences, such as meaningless replies, harmful information, etc. Therefore, it needs to be further optimized through reinforcement learning. After training and optimization, the chat large model needs to be engineered and deployed to provide services in real scenarios.
[0003] The whole process of the large model inference stage: The large model inference process, that is, the process of the large model generating content, can be divided into two stages. The first is the pre-fill stage, where the large model processes all the received inputs in parallel through the self-attention mechanism, stores the cache of the corresponding KV vectors, and generates the first output word. In the second stage, the KV cache is used to generate output word blocks one by one in an autoregressive manner until the end symbol is generated. <eos>The generation stops when the chunk is complete or the model reaches the predefined maximum generation length. In the first stage, the input is processed in parallel at once, and in the second stage, the output is generated step by step in series. Therefore, it can be found that the inference time of the large model is mainly determined by the output length and has nothing to do with the input length.
[0004] The original sponge sample generation method mainly heuristically searches for samples that can improve the inference time of the target model through genetic algorithms, with low efficiency and poor results. The existing LLMEffiChecker method, through white-box attacks, finds suitable sponge samples based on gradients. The specific process is as follows: First, the input sequence is input into the target model, and the gradients of the backward propagation in the embedding layer are calculated. According to the gradients in the embedding space, the gradient magnitudes of each word in the vocabulary can be obtained. Therefore, the word with the largest absolute value of the negative gradient among all words in the input sequence is selected as the keyword chunk. Three levels of perturbations are performed on the identified keyword chunk to generate candidate samples, namely character-level perturbation (adding, deleting, or modifying a certain letter), word-level perturbation (replacing it with a chunk with a larger negative gradient), and sentence-level perturbation (replacing it with a synonym of the same part of speech according to the sentence grammar). The generated candidate samples are input into the target model again, and the one with the longest output length is selected as the optimal sample for the current round. Repeat the above steps until the output length reaches the maximum value or the maximum number of iterations, and the finally obtained sample is used as the sponge sample.
[0005] The prior art discloses a system load stress test method for large language models. This method mainly targets large models with a mixture of experts model architecture and uses the Gumble-Softmax technique to achieve white-box gradient optimization to activate more expert models, thereby increasing the inference overhead of the large model. This method only targets large models with a mixture of experts model architecture. Since there is a problem of pre-order dependence in the currently mainstream large models with only decoder architecture, this method cannot be directly applied to mainstream chat large models.
[0006] Currently, there is a lack of a dedicated evaluation method for the availability evaluation test of large model services. Existing methods mainly rely on natural samples in conventional dialogue datasets for resource consumption evaluation. Although such samples can reflect daily service scenarios, they cannot effectively simulate extreme situations under high load stress. When a large model faces carefully constructed adversarial samples, it may abnormally activate the inference mechanism and produce continuous outputs, resulting in computational resource overload and a sharp increase in response time. Such adversarial samples that consume hardware resources and increase time overhead are called sponge samples. Similar to the denial-of-service attack in traditional cyberspace security, when a large number of sponge samples are input into a large model, the service of the large model is always in a high-load state and consumes a large amount of hardware resources, resulting in the inability to respond to the input requests of normal users, and the availability of the large model service is significantly reduced. Therefore, it is necessary to construct appropriate sponge samples for testing large models to evaluate the availability of large models.
[0007] The original sponge samples mainly target traditional models, which are highly sensitive to different inputs. With only minor perturbations to the inputs, it is possible to cause a large number of outputs in the models. However, due to their huge number of parameters, large chat models are pre-trained with massive amounts of data and fine-tuned with a large number of dialogue instructions, and thus have a strong ability to understand user inputs. Minor perturbations at the character level or word level are easily automatically corrected by the powerful semantic understanding of the large model, resulting in the inability to directly apply the original sponge samples to large chat models.
[0008] However, there are two problems to be solved in constructing sponge samples for large chat models, which are specifically as follows: Gradient conduction obstacles caused by text discreteness: The input processing process of large models first maps discrete text to a continuous vector space through the embedding layer, but the process of mapping back from the embedding layer to the discrete text space is irreversible. Therefore, the gradients obtained through backpropagation cannot directly use the gradient descent method to optimize samples in the discrete space.
[0009] Uncertainty of autoregressive output content: Mainstream large chat models generally adopt the decoder-only architecture, and their text generation process has double uncertainties. On the one hand, it stems from the dependence on previous outputs of the autoregressive mechanism itself, and on the other hand, it comes from the randomness introduced by the model decoding strategy. Therefore, even the same input may lead to output content of different lengths, and it is very difficult to directly limit the large model to generate responses of a specific length. Summary of the Invention
[0010] Aiming at the problems in the above research, the purpose of the present invention is to provide a method and system for evaluating the usability of large chat model services, so as to solve the problems of gradient conduction obstacles caused by text discreteness and uncertainty of autoregressive output content in the prior art.
[0011] In order to achieve the above purpose, the present invention adopts the following technical solutions: A method for evaluating the usability of large chat model services, including the following steps: Step 1, construct a sponge sample test data set based on the mapping relationship between gradients and word chunks for each original input sample and the large chat model; Step 2, evaluate the usability of the large chat model based on the sponge sample test data set.
[0012] Furthermore, the specific steps of Step 1 are as follows: Step 1.1: First, concatenate each original input sample with a randomly initialized adversarial suffix to form the model input. After performing one-hot vector transformation and mapping based on the model input and the vocabulary of the chat large model, it is fed into the chat large model for forward inference to obtain the output sequence. Here, the original input sample is the user input in the publicly available dialogue dataset of the chat large model. The dialogue dataset includes the ShareGPT dataset and the Alpaca dataset. The adversarial suffix includes at least one token block; Step 1.2: Construct a total loss function that combines the EOS loss and the uniform loss based on the output sequence; Step 1.3: Then, based on the total loss function, obtain the gradient corresponding to the adversarial suffix through backpropagation, and use the obtained gradient to adopt the discrete text optimization algorithm of gradient search to obtain the top M largest negative gradients of the adversarial suffix as candidate adversarial suffixes; Step 1.4: Concatenate each candidate adversarial suffix with the original input sample and then input it into the chat large model again to obtain the corresponding output sequence, and evaluate the overhead of the chat large model according to the length of the output sequence; Step 1.5: If the overhead of the chat large model reaches the maximum or the number of iterations reaches the maximum, stop the optimization, and concatenate the candidate adversarial suffix with the largest output sequence length with the original input sample as the final sponge test sample. Otherwise, select the candidate adversarial suffix with the largest output sequence length to update the randomly initialized adversarial suffix, and then execute Step 1.1 again; Step 1.6: Construct a sponge sample test dataset based on all sponge test samples.
[0013] Further, the specific steps of Step 1.1 are as follows: Concatenate each original input sample with a randomly initialized adversarial suffix as the model input, and perform position indexing for each token block in the model input based on the vocabulary of the chat large model; Perform one-hot encoding on each position index based on the size of the vocabulary. After one-hot encoding, a one-hot vector is obtained. In the one-hot vector, the position index corresponding to the vocabulary is set to 1, and other positions are 0; Map the one-hot vector to a continuous real number space using an embedding layer, and then feed it into the chat large model for forward inference to obtain the output sequence.
[0014] Based on Vocabulary Size: Each large model has a unique vocabulary that maps text (e.g., "yes") to a numeric identifier (Token ID, such as 123) that the model can process. The vocabulary size refers to the total number of Token IDs in this mapping table. For example, the vocabulary of Llama 2 contains 32,000 different mappings. At this time, the one-hot encoding of the text "yes" (assuming its Token ID is 123) under the Llama 2 vocabulary is a vector of length 32,000, where only the position corresponding to ID 123 is 1, and all other positions are 0, represented as [0, 0, …, 0, 1, 0, 0, …, 0].
[0015] Furthermore, the total loss function in step 1.2 is as follows: Wherein, represents the total loss, represents the EOS loss, represents the uniform loss, represents the weight coefficient of the uniform loss, represents the total length of the output sequence of the chat large model, represents the th output position in the output sequence, and the chat large model for <eos>The probability of a chunk, represents the probability distribution of all chunks at the th output position in the output sequence, represents the uniform distribution of the th output position in the output sequence, represents the KL divergence, which is a metric for measuring and . The smaller the KL divergence value, the closer the two distributions are.
[0016] Furthermore, the specific steps of step 2 are as follows: Step 2.1: Deploy the chat large model based on the open-source library as an online service, or upload the large model to the large model cloud platform for one-click deployment as an online service. Among them, deploying the chat large model based on the open-source library as an online service means that the open-source libraries Ollama or vLLM deploy the chat large model online as the backend, and then access the backend large model through the chat large model visualization chat interaction interface open-source libraries OpenWebUI, LangChain, or LM Studio to provide a visual interaction interface for users to simulate the chat large model service in the real scenario. The large model deployment cloud platform includes the Inference Endpoint of HuggingFace; Step 2.2: Send the sponge sample test dataset into the online service one by one, and obtain the overhead evaluation metrics during the service operation, including the chat large model output sequence length, hardware resource consumption, service response time, and network throughput. The hardware resources include memory, video memory, and the number of floating-point calculations; Step 2.3: Compare the differences between each overhead metric and the overhead metrics of the normal samples in the normal sample dataset, and finally generate the service test result. Among them, the normal sample dataset includes the ShareGPT dataset and the Alpaca dataset.
[0017] A usability evaluation system for chat large model services, including: Dataset construction module: Construct a sponge sample test dataset based on the mapping relationship between gradients and chunks of each original input sample and the chat large model; Usability evaluation module: Conduct usability evaluation of the chat large model based on the sponge sample test dataset.
[0018] Furthermore, the specific implementation steps of the dataset construction module are as follows: Step 1.1: First, concatenate each original input sample with a randomly initialized adversarial suffix to form the model input. After performing one-hot vector conversion and mapping based on the model input and the vocabulary of the chat large model, input it into the chat large model for forward inference to obtain the output sequence. Here, the original input sample is the user input in the publicly available dialogue dataset of the chat large model. The dialogue dataset includes the ShareGPT dataset and the Alpaca dataset. The adversarial suffix includes at least one token block. Step 1.2: Construct a total loss function that combines the EOS loss and the uniform loss based on the output sequence. Step 1.3: Then, based on the total loss function, obtain the gradient corresponding to the adversarial suffix through backpropagation, and use the obtained gradient to adopt the discrete text optimization algorithm of gradient search to obtain the top M largest negative gradients of the adversarial suffix as candidate adversarial suffixes. Step 1.4: Concatenate each candidate adversarial suffix with the original input sample and then input it into the chat large model again to obtain the corresponding output sequence, and evaluate the overhead of the chat large model according to the length of the output sequence. Step 1.5: If the overhead of the chat large model reaches the maximum or the number of iterations reaches the maximum, stop the optimization, and concatenate the candidate adversarial suffix with the largest output sequence length with the original input sample as the final sponge test sample. Otherwise, select the candidate adversarial suffix with the largest output sequence length to update the randomly initialized adversarial suffix, and then execute Step 1.1 again. Step 1.6: Construct a sponge sample test dataset based on all sponge test samples.
[0019] Furthermore, the specific steps of Step 1.1 are as follows: Concatenate each original input sample with a randomly initialized adversarial suffix as the model input, and perform position indexing for each token block in the model input based on the vocabulary of the chat large model. Perform one-hot encoding on each position index based on the size of the vocabulary. After one-hot encoding, obtain one-hot vectors. In the one-hot vectors, the position index corresponding to the vocabulary is set to 1, and other positions are 0. Map the one-hot vectors to a continuous real number space using an embedding layer, and then input it into the chat large model for forward inference to obtain the output sequence.
[0020] Furthermore, the total loss function in Step 1.2 is: where represents the total loss, Represents the EOS loss, Represents the uniform loss, Represents the weight coefficient of the uniform loss, Represents the total length of the output sequence of the chat large model, Represents the th output position in the output sequence for the chat large model <eos>The probability of chunks, represents the probability distribution of all chunks at the th output position in the output sequence, represents the uniform distribution of the th output position in the output sequence, represents the KL divergence, which is a metric for measuring and . The smaller the KL divergence value, the closer the two distributions are.
[0021] Furthermore, the specific implementation steps of the usability evaluation module are as follows: Step 2.1: Deploy the chat large model based on the open-source library as an online service, or upload the large model to the large model cloud platform for one-click deployment as an online service. Among them, deploying the chat large model based on the open-source library as an online service means that the open-source libraries Ollama or vLLM deploy the chat large model online as the backend, and then connect to the backend large model through the chat large model visualization chat interaction interface open-source libraries OpenWebUI, LangChain, or LM Studio to provide a visual interaction interface for users to simulate the chat large model service in the real scenario. The large model deployment cloud platform includes the Inference Endpoint of HuggingFace; Step 2.2: Send the sponge sample test dataset to the online service one by one, and obtain the overhead evaluation metrics during the service operation, including the output sequence length of the chat large model, the hardware resource consumption situation, the service response time, and the network throughput. The hardware resources include memory, video memory, and the number of floating-point calculations; Step 2.3: Compare the difference between each overhead metric and the overhead metrics of the normal samples in the normal sample dataset, and finally generate the service test result. Among them, the normal sample dataset includes the ShareGPT dataset and the Alpaca dataset.
[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention constructs sponge samples adapted to the chat large model for evaluating the usability of the chat large model service, effectively filling the technical gaps in the quantification of service stability, the dynamic evaluation of resource efficiency, and the continuous monitoring of response timeliness in the prior art, specifically as follows: First, although the gradient obtained by backpropagation cannot be directly applied to the discrete text space, the mapping relationship between the gradient and each chunk in the adversarial suffix is established through the discrete text optimization algorithm based on gradient search, comparing the influence degree of the chunks at each position in the adversarial suffix on the loss function, and reducing the loss function by replacing the chunks among the chunks in the adversarial suffix; II. To overcome the uncertainty of the output content, the present invention effectively increases the output content of the large model by designing two loss functions. At the same time, it breaks the dependence between the output sequences of the chat large model and introduces more uncertainty into each generated input sequence, encouraging the chat large model to output longer and more diverse content. That is, for the general decoder architecture, the present invention delays by <eos>The emergence of chunks increases the length of the output of large models, and further breaks the previous dependence on large models with decoder architectures through the proposal of uniform loss, avoiding the double uncertainty in the generation process of chat large models, enabling the large models to output longer content and output content of consistent or specific lengths, thereby increasing the inference overhead of large models; III. The present invention constructs sponge test samples by adding optimized adversarial suffixes after the original input samples. By separating the optimized suffixes from the input content, it can more precisely trigger anomalies in the output stage of chat large models while maintaining the fluency of conversations, making it more suitable for chat large models. That is, the present invention separately optimizes the perturbations by adding adversarial perturbations after the original input samples, more accurately locates anomalies in the inference mechanism of chat large models, achieves the effect of increasing the length of output content and thus increasing the consumption of hardware resources, and solves the problem that in current research on constructing sponge samples, mainly small perturbations are added to the input, and the powerful language understanding ability of chat large models will automatically correct the small perturbations in the input part, resulting in poor resource consumption effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the sponge test sample generation process in the present invention; Figure 2 It is a schematic diagram of the structure for evaluating the service availability of chat large models in the present invention; Figure 3 It is a schematic diagram of a specific implementation case with L = 10 and K = 3. DETAILED DESCRIPTION OF THE INVENTION
[0024] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0025] The present invention mainly includes two parts. The first part is to produce a sponge sample test data set, and the second part is to evaluate the service availability of chat large models based on the produced test data set. Among them, when producing the sponge sample test data set, it can be based on the existing publicly available large model chat data set. Specifically, taking the ShareGPT data set as an example, it contains the conversation data set between users and the real - scenario chat large model ChatGPT. We select the user inputs as the original input data set of the present invention, and then construct the corresponding test data set through the sponge sample production method of the present invention.
[0026] A method for evaluating the service availability of chat large models includes the following steps: Step 1: Construct a sponge sample test data set based on the mapping relationship between gradients and chunks for each original input sample and chat large model; the specific steps are as follows: Step 1.1: First, concatenate each original input sample with a randomly initialized adversarial suffix to form the model input. After performing one-hot vector conversion and mapping based on the model input and the vocabulary of the chat large model, send it into the chat large model for forward inference to obtain the output sequence. Here, the original input sample is the user input in the publicly available dialogue dataset of the chat large model. The dialogue dataset includes the ShareGPT dataset and the Alpaca dataset. The adversarial suffix includes at least one token block. The specific steps are as follows: Concatenate each original input sample with a randomly initialized adversarial suffix as the model input, and perform position indexing for each token block in the model input based on the vocabulary of the chat large model. Perform one-hot encoding on each position index based on the size of the vocabulary. After one-hot encoding, obtain the one-hot vector. In the one-hot vector, the position index corresponding to the vocabulary is set to 1, and other positions are 0. Map the one-hot vector to a continuous real number space using the embedding layer, and then send it into the chat large model for forward inference to obtain the output sequence.
[0027] Step 1.2: Construct a total loss function that combines the EOS loss and the uniform loss based on the output sequence. The total loss function is: where, represents the total loss, represents the EOS loss, represents the uniform loss, represents the weight coefficient of the uniform loss, represents the total length of the output sequence of the chat large model, represents the th output position in the output sequence, and the chat large model for <eos>The probability of the chunk, represents the probability distribution of all chunks at the th output position in the output sequence. represents the uniform distribution at the th output position in the output sequence. represents the KL divergence, which is a metric for measuring and . The smaller the KL divergence value, the closer the two distributions are.
[0028] Step 1.3: Then, based on the total loss function, obtain the gradient corresponding to the adversarial suffix through backpropagation, and use the obtained gradient to adopt the discrete text optimization algorithm of gradient search to obtain the top M largest negative gradients of the adversarial suffix as candidate adversarial suffixes; Step 1.4: Concatenate each candidate adversarial suffix with the original input sample and then input it into the chat large model again to obtain the corresponding output sequence, and evaluate the overhead of the chat large model according to the output sequence length; Step 1.5: If the overhead of the chat large model reaches the maximum or the number of iterations reaches the maximum, stop the optimization, and concatenate the candidate adversarial suffix with the largest output sequence length with the original input sample as the final sponge test sample. Otherwise, select the candidate adversarial suffix with the largest output sequence length to update the randomly initialized adversarial suffix, and execute Step 1.1 again; Step 1.6: Construct a sponge sample test data set based on all sponge test samples.
[0029] Step 2: Conduct an availability assessment of the chat large model based on the sponge sample test data set.
[0030] Specifically, the length of the adversarial suffix is L, and the size of the vocabulary of the chat large model is V. For one of the original input samples, concatenate a randomly initialized adversarial suffix with a fixed length of L (L is defaulted to 20) as the model input. First, it is necessary to transform the text input into a one-hot digital encoding through a tokenizer. That is, as a preparatory process for forward inference, first transform the text model input (including user input and adversarial suffix) into the corresponding numbers because the model can only process numbers. Among them, the value of the number represents the position index of the chunk in the vocabulary. For example, "yes” corresponds to the 157th position in the vocabulary. Then transform the numbers into one-hot encoding, that is, only the corresponding position is 1, and all other positions are zero, with the length being the size of the vocabulary. Therefore, the one-hot encoding dimension of the adversarial suffix is , that is, with a length of L, and each position is a one-hot encoding with a length of V; thus, the input of the text model becomes one-hot digital type that can be processed by the model, and after passing through the embedding layer to become a continuous vector representation, it is sent to the chat large model for forward inference to obtain the output sequence of the word distribution. Calculate the EOS loss and the uniform loss respectively based on the output sequence, where the EOS loss refers to each output position <eos>The mean of the chunk probabilities. The uniform loss refers to the KL divergence between the word distribution at each output position and the uniform distribution. Then, the two loss functions are directly summed (with a weight coefficient of 1 for the uniform loss) as the final loss function. After calculating the loss function, backpropagation is performed to obtain the gradient corresponding to the adversarial suffix after one-hot encoding, which is still of the same dimension , where L is the length of the adversarial suffix and V is the size of the large model's word space. At this time, the meaning of the gradient can be understood as the contribution degree of each word at each of the L positions to the loss function. Therefore, in the present invention, the top K chunks with the largest negative gradient at each position are selected to replace the original chunks as candidate suffixes to achieve the purpose of reducing the loss function. To prevent the optimization from falling into a local optimal solution, the value of K is taken as 32, and then there appears candidate suffixes. To reduce the computational complexity, the present invention randomly selects M (M = 128) candidate suffixes from them as the final candidates. Finally, the selected candidate suffixes are concatenated with the original input and sent into the chat large model again to obtain the corresponding output sequence. The adversarial suffix corresponding to the longest output sequence is selected as the current optimal suffix to update the initial suffix, and the optimization process is continuously repeated until the model output reaches the maximum value or the number of iterations reaches the maximum, then the optimization stops. The final adversarial suffix is concatenated with the original output as the final sponge test sample. By the above method, each input in the original dataset is used to generate the corresponding sponge sample, and finally a sponge test sample dataset is produced. Another example, that is, the length of the adversarial suffix is L = 10 (all initialized as asterisks " ********** "), the vocabulary of the chat large model is V, and the gradient dimension corresponding to the adversarial suffix obtained by backpropagation is , that is, there are gradients corresponding to V chunks at each of these ten positions. Select the top K (such as K = 3 or K = 20) chunks to replace the original asterisks. For the first position, there will be " A********* ", " B********* ", etc., K of them; for the second position, there will be " *C******** ", " * D******** ", etc., K of them. Therefore, there will be K candidates at each of these ten positions. To prevent falling into a local optimal solution, K is generally large, so there are candidates, the number is too large, so randomly select M of them. For example, Figure 3 is a schematic diagram of the specific implementation case of L = 10 and K = 3.
[0031] The specific steps are as follows: Step 2.1: Deploy the chat large model based on an open-source library as an online service, or upload the large model to a large model cloud platform for one-click deployment as an online service. Among them, deploying the chat large model based on an open-source library as an online service means that the open-source libraries Ollama or vLLM deploy the chat large model online as the backend, and then provide a visual interaction interface to users through the chat large model visual chat interaction interface open-source libraries OpenWebUI, LangChain, or LM Studio to access the backend large model, so as to simulate the chat large model service in the real scenario. The large model deployment cloud platform includes the Inference Endpoint of HuggingFace; Step 2.2: Send the sponge sample test dataset into the online service one by one, and obtain the overhead evaluation metrics during the service operation, including the output sequence length of the chat large model, the hardware resource consumption situation, the service response time, and the network throughput. The hardware resources include memory, video memory, and the number of floating-point calculations; Step 2.3: Compare the difference between each overhead metric and the overhead metrics of the normal samples in the normal sample dataset, and finally generate the service test result. Among them, the normal sample dataset includes the ShareGPT dataset and the Alpaca dataset.
[0032] The second part evaluates the usability of the chat large model service based on the constructed sponge sample test dataset. First, we deploy the chat large model online as the large model backend engine through the Ollama open-source library, and then use the OpenWebUI open-source library of the large model visual chat interaction interface to access the backend large model and provide a visual front-end interaction interface to users to simulate the chat large model service in the real scenario. Then, the original dataset and the corresponding test dataset generated by the method of the present invention are respectively sent into the online chat large model service, and the response time from input to complete output, the network throughput, and the output sequence length of the chat large model are recorded at the front end, and the hardware resource consumption situation is recorded at the backend, including the usage ratio and consumption of video memory, the memory occupancy, and the number of floating-point operations. Compare the differences in the specific values of each metric between the normal dataset and the test dataset, and finally give a conclusion to evaluate whether the chat large model can resist extreme situations and whether its usability can be satisfied when used as a service.
[0033] The above are only representative embodiments among the numerous specific application scopes of the present invention, and do not constitute any limitation to the protection scope of the present invention. Any technical solutions formed by transformation or equivalent replacement fall within the scope of the protection of the present invention.< / eos> < / eos> < / eos> < / eos> < / eos> < / eos>
Claims
1. A method for evaluating the usability of chat large model services, characterized in that, It includes the following steps: Step 1: Construct a sponge sample test data set based on the mapping relationship between gradients and word chunks for each original input sample and the chat large model; Step 2: Conduct an availability assessment of the chat large model based on the sponge sample test data set.
2. The usability evaluation method for chat large model services according to claim 1, wherein The specific steps of Step 1 are as follows: Step 1.1: First, splice each original input sample with a randomly initialized adversarial suffix as the model input, and perform one-hot vector conversion and mapping based on the model input and the vocabulary of the chat large model, and then send it into the chat large model for forward inference to obtain an output sequence. Among them, the original input sample is the user input in the dialogue data set publicly disclosed by the chat large model, the dialogue data set includes the ShareGPT data set and the Alpaca data set, and the adversarial suffix includes at least one word chunk; Step 1.2: Construct a total loss function that combines the EOS loss and the uniform loss according to the output sequence; Step 1.3: Then, based on the total loss function, obtain the gradient corresponding to the adversarial suffix through backpropagation, and use the obtained gradient to adopt a discrete text optimization algorithm for gradient search to obtain the top M largest negative gradients of the adversarial suffix as candidate adversarial suffixes; Step 1.4: Splice each candidate adversarial suffix with the original input sample and then input it into the chat large model again to obtain the corresponding output sequence, and evaluate the overhead of the chat large model according to the output sequence length; Step 1.5: If the overhead of the chat large model reaches the maximum or the number of iterations reaches the maximum, stop the optimization, and splice the candidate adversarial suffix with the largest output sequence length with the original input sample as the final sponge test sample. Otherwise, select the candidate adversarial suffix with the largest output sequence length to update the randomly initialized adversarial suffix, and then execute Step 1.1 again; Step 1.6: Construct a sponge sample test data set based on all sponge test samples.
3. The usability evaluation method for chat large model services according to claim 2, wherein The specific steps of Step 1.1 are as follows: Splice each original input sample with a randomly initialized adversarial suffix as the model input, and perform position indexing on each word chunk in the model input based on the vocabulary of the chat large model; Perform one-hot encoding on each position index based on the size of the vocabulary. After one-hot encoding, obtain a one-hot vector, where the position index corresponding to the vocabulary in the one-hot vector is set to 1, and other positions are 0; Map the one-hot vector to a continuous real number space using an embedding layer, and then send it into the chat large model for forward inference to obtain an output sequence.
4. A method for evaluating the usability of a chat large model service according to claim 2, characterized in that, The total loss function in Step 1.2 is: Among them, represents the total loss, represents the EOS loss, represents the uniform loss, represents the weight coefficient of the uniform loss, represents the total length of the output sequence of the chat large model, represents the th output position in the output sequence, where the chat large model for <eos>The probability of a chunk, represents the probability distribution of all chunks at the th output position in the output sequence, represents the uniform distribution at the th output position in the output sequence, represents the KL divergence, which is an index for measuring and The smaller the KL divergence value, the closer the two distributions are.< / eos> 5. The usability evaluation method for chat large model services according to claim 4, characterized in that, The specific steps of Step 2 are as follows: Step 2.1: Deploy the chat large model as an online service based on an open-source library, or upload the large model to a large model cloud platform for one-click deployment as an online service. Here, deploying the chat large model as an online service based on an open-source library means that the open-source libraries Ollama or vLLM deploy the chat large model online as the backend, and then the chat large model visualization chat interface open-source libraries OpenWebUI, LangChain, or LM Studio are used to access the backend large model to provide a visual interaction interface for users, so as to simulate the chat large model service in the real scenario. The large model deployment cloud platform includes the Inference Endpoint of HuggingFace; Step 2.2: Send the sponge sample test dataset into the online service one by one, and obtain the overhead evaluation metrics during the service operation, including the output sequence length of the chat large model, the hardware resource consumption, the service response time, and the network throughput. The hardware resources include memory, video memory, and the number of floating-point calculations; Step 2.3: Compare the difference between each overhead metric and the overhead metrics of the normal samples in the normal sample dataset, and finally generate the service test result. Here, the normal sample dataset includes the ShareGPT dataset and the Alpaca dataset.
6. An availability evaluation system for chat large model services, characterized in that, Including: Dataset construction module: Construct the sponge sample test dataset based on the mapping relationship between gradients and word chunks of each original input sample and the chat large model; Usability evaluation module: Evaluate the usability of the chat large model based on the sponge sample test dataset.
7. The availability evaluation system for chat large model services according to claim 6, wherein The specific implementation steps of the dataset construction module are as follows: Step 1.1: First, splice each original input sample with a randomly initialized adversarial suffix to form a model input, and perform one-hot vector conversion and mapping based on the model input and the vocabulary of the chat large model, and then send it into the chat large model for forward inference to obtain an output sequence. Here, the original input sample is the user input in the dialogue dataset publicly available by the chat large model. The dialogue dataset includes the ShareGPT dataset and the Alpaca dataset, and the adversarial suffix includes at least one word chunk; Step 1.2: Construct a total loss function that combines the EOS loss and the uniform loss based on the output sequence; Step 1.3: Then, obtain the gradient corresponding to the adversarial suffix through backpropagation based on the total loss function, and use the obtained gradient to adopt the discrete text optimization algorithm of gradient search to obtain the top M largest negative gradients of the adversarial suffix as candidate adversarial suffixes; Step 1.4: Splice each candidate adversarial suffix with the original input sample and then input it into the chat large model again to obtain the corresponding output sequence, and evaluate the overhead of the chat large model according to the output sequence length; Step 1.5: If the overhead of the chat large model reaches the maximum or the number of iterations reaches the maximum, stop the optimization, and splice the candidate adversarial suffix with the largest output sequence length with the original input sample as the final sponge test sample. Otherwise, select the candidate adversarial suffix with the largest output sequence length to update the randomly initialized adversarial suffix, and then execute Step 1.1 again; Step 1.6: Construct a sponge sample test dataset based on all sponge test samples.
8. An availability evaluation system for chat large model services according to claim 7, characterized in that, The specific steps of step 1.1 are as follows: Concatenate each original input sample with a randomly initialized adversarial suffix as the model input, and perform position indexing for each token block in the model input in the vocabulary based on the chat large model; Perform one-hot encoding on each position index based on the size of the vocabulary. After one-hot encoding, a one-hot vector is obtained. In the one-hot vector, the position index corresponding to the vocabulary is set to 1, and other positions are 0; Map the one-hot vector to a continuous real number space using an embedding layer, and then send it into the chat large model for forward inference to obtain an output sequence.
9. An availability evaluation system for chat large model services according to claim 7, characterized in that, The total loss function in step 1.2 is: Among them, represents the total loss, represents the EOS loss, represents the uniform loss, represents the weight coefficient of the uniform loss, represents the total length of the output sequence of the chat large model, represents the th output position in the output sequence, where the chat large model for <eos>The probability of a chunk, represents the probability distribution of all chunks at the th output position in the output sequence, represents the uniform distribution at the th output position in the output sequence, represents the KL divergence, which is an index for measuring and The smaller the KL divergence value, the closer the two distributions are.< / eos> 10. The availability evaluation system for chat large model services according to claim 9, characterized in that, The specific implementation steps of the usability evaluation module are as follows: Step 2.1: Deploy the chat large model based on an open-source library as an online service, or upload the large model to a large model cloud platform for one-click deployment as an online service. Among them, deploying the chat large model based on an open-source library as an online service means that the open-source libraries Ollama or vLLM deploy the chat large model online as the backend, and then access the backend large model through the chat large model visualization chat interface open-source libraries OpenWebUI, LangChain, or LM Studio to provide a visual interaction interface for users to simulate the chat large model service in a real scenario. The large model deployment cloud platform includes the Inference Endpoint of HuggingFace; Step 2.2: Send the sponge sample test dataset into the online service one by one, and obtain overhead evaluation metrics during the service operation, including the length of the chat large model output sequence, hardware resource consumption, service response time, and network throughput. The hardware resources include memory, video memory, and the number of floating-point calculations; Step 2.3: Compare the difference between each overhead metric and the overhead metrics of the normal samples in the normal sample dataset, and finally generate a service test result. Among them, the normal sample dataset includes the ShareGPT dataset and the Alpaca dataset.
Citation Information
Patent Citations
System load pressure testing method for large language model
CN118035104A
Method for realizing multi-person identity interaction and memory based on large language model
CN118656459A
Image evaluation model training method and device, image processing method and device, computer equipment and readable storage medium
CN118747297A
Large language model optimization generation method based on optimal cue word selection
CN119476209A
Model searching method for sorting machine, sorting method and electronic equipment
CN119540619A
Cited By
Multi-modal large model availability evaluation method based on repeated output
CN121030417A
A multi-modal large model availability evaluation method based on repeated output
CN121030417B
Model security availability joint evaluation method based on dynamic multi-objective optimization
CN121658349A