Service performance evaluation method, computer program product and electronic device

By partitioning and caching various experimental test data in the data cache pool and performing various data processing, the problem of repeated calculation in multi-service evaluation is solved, the evaluation efficiency and resource utilization are improved, and more accurate evaluation results are provided.

CN120723342AActive Publication Date: 2025-09-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511215871.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-30
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

In the existing technology, there is repeated calculation in the data processing process of multi-service evaluation, which leads to reduced evaluation efficiency and waste of resources.

Method used

A data cache pool is used to partition and cache various experimental test data. Multiple data sets are processed using various data processing methods to execute performance evaluation experiments. The streaming response of the service to be evaluated is then collected for dynamic evaluation.

Benefits of technology

It improves the efficiency and resource utilization of service evaluation, provides more accurate and scientific evaluation results, and reduces duplicate calculations and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723342A_ABST
    Figure CN120723342A_ABST
Patent Text Reader

Abstract

The invention discloses a service performance evaluation method, a computer program product and electronic equipment, and relates to the technical field of performance evaluation.A data cache pool is used for caching various experimental test data in a partitioned mode, the various experimental test data are obtained by processing a plurality of data sets through various data processing modes, and based on the data cache pool, the data processing mode is optimized; and executing a performance evaluation experiment. Therefore, the technical problem that the service evaluation efficiency is reduced due to many repeated calculations in the data processing process of multi-service evaluation can be solved, and the technical effects of improving the evaluation efficiency and the resource utilization rate and providing powerful data support are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of performance evaluation, and in particular to a service performance evaluation method, a computer program product, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the large-scale implementation of artificial intelligence (AI) technology, pre-trained language models are penetrating numerous fields, including finance, healthcare, and technology. These scenarios present significant differences in demand, bringing new requirements for inference efficiency and resource consumption. Therefore, inference service evaluation technology plays a crucial role in optimizing inference on pre-trained language models. Currently, the data processing for multi-service evaluation involves numerous repeated calculations, which reduces service evaluation efficiency. Summary of the Invention

[0003] The present application provides a service performance evaluation method, a computer program product, an electronic device and a computer-readable storage medium to at least solve the problem in the related art that there are many repeated calculations in the data processing process of multi-service evaluation, which reduces the efficiency of service evaluation.

[0004] This application provides a service performance evaluation method, including: Create performance evaluation experiments for the services to be evaluated, including pre-trained language model services. Perform performance evaluation experiments based on a data cache pool, where the data cache pool is used to partition and cache various experimental test data, which are obtained by processing multiple data sets using various data processing methods. During the performance evaluation experiment, data is collected from the streaming responses of the service to be evaluated to perform dynamic evaluation of the service to be evaluated.

[0005] The present application also provides a computer program product, comprising: A creation module is used to create performance evaluation experiments corresponding to the services to be evaluated, including pre-trained language model services. An execution module is used to execute a performance evaluation experiment based on a data cache pool, wherein the data cache pool is used to partition and cache a variety of experimental test data, and the various experimental test data are obtained by processing a plurality of data sets through a plurality of different data processing methods; The evaluation module is used to collect data on the streaming response of the service to be evaluated during the performance evaluation experiment, so as to perform dynamic evaluation on the service to be evaluated.

[0006] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned service performance evaluation methods when executing the computer program.

[0007] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned service performance evaluation methods are implemented.

[0008] This application uses a partitioned data cache pool to cache multiple experimental test data, each of which is obtained by processing multiple data sets using a variety of different data processing methods. Performance evaluation experiments can then be performed based on the data cache pool. This solves the technical problem of multiple service evaluations, which often involves repeated calculations during data processing, leading to reduced service evaluation efficiency. This improves evaluation efficiency and resource utilization, while also providing robust data support. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0010] Figure 1 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 2 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 3 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 4 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 5 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 6 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 7 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 8 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 9 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 10 A flowchart of a service performance evaluation method provided in an embodiment of the present application; Figure 11 A schematic diagram of a computer program product provided in accordance with an embodiment of the present invention; Figure 12 A schematic diagram of a module of an electronic device provided in an embodiment of the present application; Figure 13 A schematic diagram of a computer-readable storage medium provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0011] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0012] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0013] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0014] See also Figure 1 and Figure 2 , an embodiment of the present application provides a service performance evaluation method, and the method is described in detail in conjunction with the execution process of the service performance evaluation method.

[0015] Service performance evaluation methods include: 010: Create a performance evaluation experiment corresponding to the service to be evaluated, including the pre-trained language model service; 020: Execute performance evaluation experiments based on the data cache pool. The data cache pool is used to partition and cache various experimental test data. The various experimental test data are obtained by processing multiple data sets using various data processing methods. 030: During the performance evaluation experiment, data is collected from the streaming responses of the service to be evaluated to perform dynamic evaluation of the service to be evaluated.

[0016] In the service performance evaluation method of the present embodiment, a data cache pool is partitioned to cache multiple experimental test data. These data are obtained by processing multiple data sets through various different data processing methods. Performance evaluation experiments are then executed based on the data cache pool. This solves the technical problem of repeated calculations during data processing for multi-service evaluations, which reduces service evaluation efficiency. This improves evaluation efficiency and resource utilization, while also providing robust data support.

[0017] Specifically, first determine the service to be evaluated. The service to be evaluated refers to the pre-trained language model service to be evaluated. The service to be evaluated can include multiple sub-services, which can be multiple different pre-trained language model services or multiple different sub-services within a single pre-trained language model service.

[0018] Create a performance evaluation experiment for the service to be evaluated. For different sub-services, you can create different sub-experiments based on actual application needs, which together form a performance evaluation experiment. Once created, you can begin executing the performance evaluation experiment. In a performance evaluation experiment, input data processing is a crucial step in ensuring the effectiveness of the evaluation.

[0019] Set up a data cache pool to perform performance evaluation experiments based on the experimental test data in the data cache pool. For example, based on the data set and data processing method required for the performance evaluation experiment, select the corresponding experimental test data from the data cache pool.

[0020] In the data cache pool, different experimental test data are cached in partitions. The data cache pool can include multiple cache partitions, and each type of experimental test data can be cached in a cache partition. The multiple experimental test data corresponds to multiple data sets and multiple data processing methods. Each experimental test data is obtained by processing a data set using a data processing method.

[0021] Various data sets such as the MNIST and CIFAR datasets. Various data processing methods, including basic and specialized methods.

[0022] Basic data processing methods include outlier processing, corruption detection, and sensitive information masking. Outlier processing, for example, filters out overly long text and pure noise audio. Corruption detection, for example, verifies file integrity (based on cyclic redundancy checks or message digest algorithms). Sensitive information masking, for example, automatically masks ID card and bank card numbers. Professional data processing includes word segmentation, stemming, subword segmentation, sliding window segmentation, and cross-modal alignment.

[0023] During the execution of a performance evaluation experiment, the service to be evaluated can be dynamically evaluated based on the execution status of the performance evaluation experiment. For example, during the execution of a performance evaluation experiment, the service to be evaluated generates a streaming response. Data collection can be performed on the streaming response of the evaluation model. Based on the data collection results, the service to be evaluated can be dynamically evaluated to determine whether the performance of the service to be evaluated meets the requirements.

[0024] In related technologies, when evaluating each service, a data processing method is selected to process the data set, and then test data is obtained for evaluation. When repeated experiments are performed multiple times, the processing needs to be repeated multiple times. In the multi-service evaluation process, the same data set and the same data processing method may be used for different services. Each evaluation requires a lot of time and resources to waste real-time calculations and complex data processing, which reduces the efficiency of service evaluation and wastes computing resources to a certain extent.

[0025] In the embodiment of the present application, a multi-zone cache is added to accelerate evaluation. Multiple data sets are processed using a variety of different data processing methods to obtain a variety of experimental test data. The various experimental test data are partitioned and cached in a data cache pool. When executing a performance evaluation experiment, the required data set and the experimental test data corresponding to the required data processing method can be directly read from the data cache pool. This can effectively improve evaluation efficiency and resource utilization, provide strong data support for multi-service evaluation scenarios, and improve the efficiency of multi-service evaluation.

[0026] The following describes the detailed process of creating a performance evaluation experiment.

[0027] See also Figure 2 and Figure 3 In some embodiments, creating a performance evaluation experiment corresponding to the service to be evaluated (ie, 010) includes: 011: Determine the experimental objectives corresponding to the service to be evaluated and define the corresponding experimental rules; 012: Verify the legitimacy of experimental objectives and experimental rules; 013: When the experimental objectives and experimental rules pass the legality verification, create a performance evaluation experiment based on the experimental objectives and experimental rules.

[0028] Specifically, the experimental objectives corresponding to the service to be evaluated are determined, and based on the experimental objectives, corresponding experimental rules can be defined through a configuration file.

[0029] In some embodiments, the experimental objectives include any one or more of the services to be evaluated, version relationships, core indicators, and experimental types, and the indicator data include any one or more of normal delay, success rate, and dynamic streaming delay; and / or the experimental rules include any one or more of version representation, traffic distribution strategy, indicator threshold, statistical significance requirements, experimental execution cycle, and indicator collection frequency.

[0030] Experimental objectives include one or more of the service being evaluated, version relationships, metric data, and experiment type. Version relationships refer to the relationship between different versions of the service being evaluated, which can be divided into baseline versions and candidate versions. The baseline version is the currently running version of the service being evaluated, while the candidate version is a new version of the service being evaluated. The evaluation of the service being evaluated involves evaluating the improvement of the candidate version over the baseline version.

[0031] The indicator data is used for subsequent dynamic evaluation and can be determined based on actual evaluation requirements. For example, the indicator data can include any one or more of normal delay, success rate, and dynamic streaming delay.

[0032] The experiment type refers to the type of performance evaluation experiment created. It can be determined based on actual evaluation requirements. For example, the experiment type can be A / B testing, canary release, etc.

[0033] Experiment rules include one or more of the following: version identifier, traffic allocation strategy, metric threshold, statistical significance requirement, experiment execution period, and metric collection frequency. The version identifier identifies the relationship between the versions mentioned above. The traffic allocation strategy refers to the strategy for allocating traffic to different versions. In an example, for an A / B test, 90% of the traffic is allocated to the baseline version and 10% to the candidate version.

[0034] The indicator threshold refers to the threshold corresponding to the indicator data, which can be set according to actual evaluation requirements. For example, you can set the delay increase to no more than 10% as the indicator threshold corresponding to dynamic streaming delay.

[0035] Statistical significance requirements are used to determine whether the indicator data is statistically significant, ensuring that the observed differences (such as the difference between the indicator data of the baseline version and the candidate version) are not caused by random fluctuations. In one example, for latency, a latency p99 value of < 0.05 is set as the statistical significance requirement.

[0036] The experiment execution cycle refers to the period of time between requests to the service being evaluated during the performance evaluation experiment. The metric collection frequency refers to the frequency at which data is collected to determine metric data during the performance evaluation experiment. In one example, the metric collection frequency is set to once every 30 seconds.

[0037] In related technologies, general indicators are often used for evaluation, and the evaluation results are not accurate enough.

[0038] In the embodiment of the present application, the evaluation dimensions are refined and the user attention indicators are further improved. The indicator data includes normal delay, success rate and dynamic streaming delay, and the evaluation is performed in combination with the statistical significance requirements. The accuracy of the evaluation results is relatively high.

[0039] After determining the experimental objectives and experimental rules, perform a validity check on them. If the experimental objectives and experimental rules fail the validity check, return to the steps of determining the experimental objectives for the service to be evaluated and defining the corresponding experimental rules. If the experimental objectives and experimental rules pass the validity check, you can create a performance evaluation experiment based on the experimental objectives and experimental rules.

[0040] It should be noted that the services being evaluated can be deployed in a container orchestration system, such as a Kubernetes (K8s) cluster. Kubernetes is an open-source platform used to manage containerized applications across multiple hosts in a cloud platform. Kubernetes aims to make containerized application deployment simple and efficient, providing a mechanism for application deployment, planning, updating, and maintenance.

[0041] When creating a performance evaluation experiment, the corresponding metadata is injected into the job or cronjob template in the Kubernetes cluster as environment variables or volume data. The metadata includes the experiment name, namespace, experiment target, and experiment rules. After that, resources are initialized and a temporary task is created to periodically collect data, completing the creation of the performance evaluation experiment.

[0042] A Job is a controller used in Kubernetes to manage one-time batch processing tasks. It ensures that one or more basic computing units (Pods) terminate after successfully completing their tasks. Its core goal is to ensure reliable task execution. Even if a node fails or a Pod terminates unexpectedly, the Job will automatically retry or reschedule the task.

[0043] A CronJob is a resource object used in Kubernetes to manage periodic tasks. It automatically triggers job execution based on a Cron expression. Essentially, it is a controller that creates jobs on a scheduled basis and is suitable for tasks that need to be executed repeatedly at fixed intervals.

[0044] like Figure 2As shown in the figure, when performing performance evaluation experiments, you can use traffic configuration tools (such as Istio) to configure traffic routing and proportionally distribute requests to different versions of the service to be evaluated. You can also associate the data source of the monitoring and alarm system (such as Prometheus) to collect, store, and query the corresponding indicator data of different versions of the service to be evaluated in real time.

[0045] The specific process of verifying the legitimacy of experimental objectives and experimental rules is described in detail below.

[0046] See also Figure 4 In some embodiments, before creating a performance evaluation experiment corresponding to the service to be evaluated (ie, 010), the service performance evaluation method further includes: 040: Add the access address of the service to be evaluated and configure the corresponding model name and access key; At this point, the validity of the experimental objectives and experimental rules is checked (i.e. 012), including: 0121: Initiate a request to the service to be evaluated based on the access address, model name, and access key; 0122: Determine whether the experimental objectives and experimental rules pass the validity check based on the request results.

[0047] Specifically, the service to be evaluated is deployed in a Kubernetes cluster and has an externally accessible interface, such as a Hypertext Transfer Protocol (HTTP) path or a Google Remote Procedure Call (gRPC) port. By adding the access address of the service to be evaluated and configuring the corresponding model name and access key, you can initiate a request to the service to be evaluated. The access address can be a Uniform Resource Locator (URL), for example, https: / / api.example.com / v1 / chat / completions.

[0048] When verifying the validity of the experimental objectives and experimental rules, a request is initiated to the service to be evaluated based on the access address, model name, and access key. The request result can be used to determine whether the experimental objectives and experimental rules pass the validity verification.

[0049] For example, based on the request results, it can be determined whether the service to be evaluated exists, whether the data corresponding to the indicator data can be collected, and whether the experimental rules are correct. If the service to be evaluated exists, the data corresponding to the indicator data can be collected, and the experimental rules are correct, then the experimental objectives and experimental rules have passed the validity verification. If the evaluation service does not exist, the data corresponding to the indicator data cannot be collected, or the experimental rules are incorrect, then the experimental objectives and experimental rules have failed the validity verification.

[0050] In this way, the reliability of the created performance evaluation experiment is guaranteed, and errors in the experimental objectives and experimental rules are avoided, which may affect the subsequent evaluation results and avoid invalid experiments.

[0051] See also Figure 2 and Figure 5 In some embodiments, before creating a performance evaluation experiment corresponding to the service to be evaluated (ie, 010), the service performance evaluation method further includes: 050: Add the access address of the service to be evaluated and configure the corresponding model name and access tag; At this point, based on the data cache pool, a performance evaluation experiment (i.e., 020) is performed, including: 021: Initiate a request to the service to be evaluated based on the access address, model name, and access tag, and read the target test data from the data cache pool and send it to the service to be evaluated; 022: Receive the streaming response of the service to be evaluated; At this point, data is collected from the streaming response of the service to be evaluated to perform a dynamic evaluation (i.e., 030) of the service to be evaluated, including: 031: Collect data from streaming responses according to the preset collection frequency to determine indicator data; 032: Dynamically evaluate the services to be evaluated based on indicator data.

[0052] Specifically, the service to be evaluated is deployed in a Kubernetes cluster and has an externally accessible interface. By adding the access address of the service to be evaluated and configuring the corresponding model name and access key, you can initiate a request to the service to be evaluated.

[0053] After the performance evaluation experiment starts, requests are made to the service being evaluated based on the access address, model name, and access key, according to the experiment execution cycle specified in the experiment rules. When making requests, the application programming interface (API) specifications defined by the service being evaluated must be followed to ensure service stability and data reliability.

[0054] When a request is initiated, the target test data is read from the data cache and sent to the service to be evaluated. The target test data refers to the experimental test data required to perform the performance evaluation experiment. The service to be evaluated receives the target test data, processes it, and generates a streaming response.

[0055] Receive streaming responses from the service to be evaluated and collect data based on the preset collection frequency set in the experiment rules. Determine metrics based on the collected data. Evaluate the service based on these metrics. Requests and data are collected for different versions of the service to be evaluated, and metrics are associated with the corresponding versions.

[0056] The specific process of reading target test data from the data cache pool is described in detail below.

[0057] See also Figure 6 In some embodiments, reading target test data from a data cache pool and sending it to a service to be evaluated (ie, 021) includes: 0211: Determine the target data set and target data processing method corresponding to the target test data; 0212: Determine the cache tag of the target test data based on the target data set and the target data processing method; 0213: Read the target test data from the data cache pool according to the cache tag and send it to the service to be evaluated.

[0058] Specifically, it can be understood that the target test data is obtained by processing the target dataset using the target data processing method. The target dataset and target data processing method corresponding to the target test data are determined. The cache tag of the target test data is determined based on the target dataset and the target data processing method. Based on the cache tag, a query can be performed in the data cache pool. Based on the query result, the target test data is read from the data cache pool and sent to the service to be evaluated.

[0059] See also Figure 7 In some embodiments, the cache tag includes a first tag and a second tag. Determining the cache tag of the target test data (i.e., 0212) based on the target data set and the target data processing method includes: Determine whether the target data set is included in multiple data sets corresponding to multiple experimental test data; Without including the target dataset, importing the target dataset to determine the first label; In the case of including the target dataset, the label of the target dataset is used as the first label; Determining whether a plurality of data processing methods corresponding to a plurality of experimental test data include a target data processing method; In the case where the target data processing method is not included, importing the target data processing method to determine the second label; In the case of including the target data processing method, the label of the target data processing method is used as the second label.

[0060] Specifically, for the multiple data sets and multiple data processing methods corresponding to the experimental test data included in the data cache pool, in order to facilitate user viewing, corresponding data set pools and data processing method pools are generated. The data set pools and data processing method pools are equivalent to index directories of data sets and data processing methods.

[0061] By querying the dataset pool, you can determine whether the target dataset is included in the multiple datasets corresponding to various experimental test data. If the target dataset is not included, import the target dataset into the dataset pool, and use the label of the imported target dataset as the first label. If the target dataset is included, the label of the target dataset in the data cache pool can be directly used as the first label.

[0062] Similarly, by querying the data processing method pool, it is possible to determine whether the target data processing method is included in the various data processing methods corresponding to the various experimental test data. If the target data processing method is not included, the target data processing method is imported into the data processing method pool, and the label of the imported target data processing method can be used as the second label. If the target data processing method is included, the label of the target data processing method in the data cache pool can be directly used as the second label.

[0063] The cache tag includes a first tag and a second tag. In one example, the cache tag is four-dimensional data, as shown below:

[0064] in, is the cache tag; is the label of the target dataset, that is, the first label; The label of the target data processing method, which is also the second label; The timestamp of the last update of the target dataset. The timestamp of the last update of the target data processing method.

[0065] See also Figure 7 In some embodiments, the target test data is read from the data cache pool according to the cache tag and sent to the service to be evaluated (i.e., 0213), including: Based on the cache tag, determine whether there is experimental test data corresponding to the target test data in the data cache pool; In the absence of corresponding experimental test data, the target data set is processed according to the target data processing method to obtain the target test data; Cache the target test data into the data cache pool, and return the step of determining whether there is experimental test data corresponding to the target test data in the data cache pool based on the cache tag; If corresponding experimental test data exists, the experimental test data is obtained as target test data and sent to the service to be evaluated.

[0066] Specifically, based on the determined cache tag, a query can be performed in the data cache pool to automatically detect whether there is experimental test data corresponding to the target test data in the data cache pool.

[0067] Although in the aforementioned implementation, the target dataset was imported when the dataset pool did not exist, and the target data processing method was imported when the data processing method pool did not exist, the target dataset may not be processed using the target data processing method, that is, the corresponding experimental test data may not exist in the data cache pool.

[0068] If the corresponding experimental test data does not exist in the data cache pool, the target dataset is first processed according to the target data processing method to obtain the target test data. The target test data is then cached in the data cache pool as experimental test data. The process returns to the step of determining whether experimental test data corresponding to the target test data exists in the data cache pool based on the cache tag. A further query is then performed to retrieve the corresponding experimental test data.

[0069] If the corresponding experimental test data exists in the data cache pool, the experimental test data is directly obtained as the target test data and sent to the service to be evaluated.

[0070] It is understood that the cache tag is a unique tag, and the corresponding experimental test data can be determined in the data cache pool based on the cache tag. In this way, the accuracy of the experimental test data read is guaranteed, and the performance evaluation experiment is guaranteed to be executed normally, ensuring the accuracy of the evaluation results.

[0071] See also Figure 8 In some embodiments, dynamically evaluating the service to be evaluated based on the indicator data (i.e., 032) includes: 0321: Compare the data differences between the baseline version of the service to be evaluated and the candidate version; 0322: Verify the statistical significance of data differences through statistical hypothesis testing algorithms; 0323: Verify whether the indicator data of the baseline version and the indicator data of the candidate version meet the indicator thresholds corresponding to the indicator data.

[0072] Specifically, during the performance evaluation experiment, metrics data for a baseline version of the service being evaluated and metrics data for a candidate version of the service being evaluated can be determined. The differences between the baseline and candidate versions are compared, and the statistical significance of the differences is verified using a statistical hypothesis testing algorithm, such as a t-test.

[0073] The core purpose of the t-test algorithm is to determine whether the difference between the means of two or more groups of data is statistically significant, that is, to determine whether the difference is real or merely caused by random fluctuations (such as the chance of sample selection).

[0074] For example, according to the statistical significance requirement set in the experimental rules: the delayed p99 value is < 0.05. If the delayed p99 value calculated based on the t-test algorithm is < 0.05, the statistical significance of the data difference between the baseline version's indicator data and the candidate version's indicator data is met. If the delayed p99 value calculated based on the t-test algorithm is > 0.05, the statistical significance of the data difference between the baseline version's indicator data and the candidate version's indicator data is not met.

[0075] In addition, you can also verify whether the indicator data of the baseline version and the indicator data of the candidate version meet the corresponding indicator thresholds. For example, taking dynamic streaming delay as an example, verify whether the increase in the dynamic streaming delay of the baseline version exceeds the corresponding increase threshold. If the increase in the dynamic streaming delay of the baseline version exceeds the corresponding increase threshold, it indicates that the dynamic streaming delay of the baseline version is abnormal; if the increase in the dynamic streaming delay of the baseline version does not exceed the corresponding increase threshold, it indicates that the dynamic streaming delay of the baseline version is normal.

[0076] In the embodiment of the present application, a dynamic evaluation is performed based on the indicator data combined with statistical significance to evaluate whether the data difference between the indicator data of the baseline version and the indicator data of the candidate version is real, which provides a scientific and objective basis for service evaluation, and the accuracy of the evaluation results is relatively high.

[0077] See also Figure 9 In some embodiments, a performance evaluation experiment (ie, 020) is performed based on the data cache pool, including: In the data cache pool, when the data set and / or data processing method corresponding to the experimental test data is updated, a cache partition is established in the data cache pool to cache the updated data set and / or updated test data corresponding to the data processing method as the experimental test data; Determine whether the performance evaluation experiment includes sub-experiments corresponding to the dataset and / or data processing method; In the case of including sub-experiments corresponding to the data set and / or data processing method, wait for the sub-experiments to be completed and clear the experimental test data corresponding to the data set and / or data processing method.

[0078] Without including the sub-experiments corresponding to the data set and / or data processing method, the experimental test data corresponding to the data set and / or data processing method are directly cleared.

[0079] Specifically, during the execution of the performance evaluation experiment, the data set and / or data processing method corresponding to the experimental test data in the data cache pool may be updated.

[0080] When the dataset corresponding to the experimental test data is updated, the data processing method associated with the dataset before the update is obtained and processed using the associated data processing method to obtain updated test data. A new cache partition is established in the data cache pool to cache the updated test data corresponding to the updated dataset as the experimental test data.

[0081] If the service to be evaluated includes multiple sub-services, the performance evaluation experiment will include multiple corresponding sub-experiments. Determine whether the performance evaluation experiment includes sub-experiments corresponding to the dataset before the update. If the performance evaluation experiment includes sub-experiments corresponding to the dataset before the update, create a dependency list for the corresponding sub-experiments and dynamically monitor whether all sub-experiments in the dependency list have completed. Once a sub-experiment has completed, remove the sub-experiment from the dependency list and clear the experimental test data corresponding to the dataset before the update.

[0082] In the case that the performance evaluation experiment does not include the sub-experiment corresponding to the data set, the experimental test data corresponding to the data set before the update can be directly cleared.

[0083] When the data processing method corresponding to experimental test data is updated, the datasets associated with the data processing method prior to the update are obtained and processed using the updated data processing method to obtain the updated test data. A new cache partition is established in the data cache pool to cache the updated test data corresponding to the updated data processing method as the experimental test data.

[0084] Determine whether the performance evaluation experiment includes sub-experiments corresponding to the data processing method before the update. If the performance evaluation experiment includes sub-experiments corresponding to the data processing method before the update, you can create a dependency list for the corresponding sub-experiments and dynamically monitor whether all sub-experiments in the dependency list have completed. Wait for the sub-experiments to complete, remove the sub-experiments from the dependency list, and clear the experimental test data corresponding to the data processing method before the update.

[0085] When the performance evaluation experiment does not include a sub-experiment corresponding to the data processing method, the experimental test data corresponding to the data processing method can be directly cleared.

[0086] In this way, the data cache pool can be smoothly upgraded when updated test data is introduced, providing strong data support for the evaluation of multiple inference services. In addition, it does not affect the normal execution of performance evaluation experiments, further improving the efficiency of multi-service evaluation.

[0087] In addition, when the data set corresponding to the experimental test data is updated, the timestamp of the data set will change, so that when the performance evaluation experiment is subsequently executed, the latest experimental test data can be queried when querying in the data cache pool according to the cache tag.

[0088] Similarly, when the data processing method corresponding to the experimental test data is updated, the timestamp of the data processing method will change, so that when the performance evaluation experiment is subsequently executed, the latest experimental test data can be queried when querying in the data cache pool according to the cache tag.

[0089] See also Figure 10 In some embodiments, the indicator data includes a dynamic streaming delay, and data is collected from the streaming response according to a preset collection frequency to determine the indicator data (i.e., 031), including: 0311: Get the generation timestamps of multiple tags in the streaming response in real time; 0312: Calculate the arrival delay between adjacent tags based on the generation timestamp; 0313: Determine the weights corresponding to the multiple tags according to the generation order of the multiple tags; 0314: Calculate dynamic streaming delay based on arrival delay and weight.

[0090] Specifically, the streaming response of the service to be evaluated includes multiple tags, which can be referred to as tokens. The following is an example of a tag being a token. When a token is generated, the generation timestamp of each token is captured and recorded in real time. Suppose the number of tokens is n and the token sequence is ,in Refers to the arrival time of the i-th token from the start of the request (that is, the generation timestamp, in seconds).

[0091] The arrival delay of the i-th token is defined as the difference between its arrival time and the arrival time of the previous token, that is:

[0092] in, Denote the latency of the \(i\)-th token, that is, the arrival latency between the \(i\)-th token and the \((i - 1)\)-th token. Denote the time when the request starts.

[0093] To control the steepness of the weight increase through parameters and take into account the certain influence of early latency (terms with smaller \(i\)), the weight formula for each token can be shown as follows:

[0094] where is the latency of the \(i\)-th token is the weight. \(0 < k < 1\), and the default value is \(0.1\), which is used to adjust the exponent size. The larger \(k\) (e.g., \(0.5\)), the faster the weight of the later stage (terms with larger \(i\)) grows, and the relatively greater the impact on the result. \(c(c>0)\) is a constant term. Adding the constant term \(c\) to the numerator of the weight function is used to increase the weight of early latency. By adding the constant term \(c\), the weight of the first latency is increased from \(0\) to a non-zero value, and the larger \(c\) is, the closer the weight of early latency is to that of the later stage, and the influence of each latency is more balanced.

[0095] Combined with the definition of latency, the overall latency, that is, the dynamic streaming latency, is:

[0096] The final formula for the dynamic streaming latency is:

[0097] In the related art, the total latency of the entire response completion is calculated, ignoring the different sensitivities of users to intermediate tokens and ending tokens during the streaming generation process. <000​​​​​​​​​​​​If the experiment termination condition is not met, return to the step of executing the performance evaluation experiment based on the data cache pool.

[0100] Specifically, during the performance evaluation experiment, after determining the indicator data, it can be used to determine whether the performance evaluation experiment meets the experiment termination conditions. The experiment termination conditions include a first termination condition and a second termination condition. The first termination condition corresponds to the termination of the performance evaluation experiment under normal execution, and the second termination condition corresponds to the termination of the performance evaluation experiment under abnormal conditions.

[0101] The first termination condition includes that the performance evaluation experiment is executed until the first preset time, the performance evaluation experiment is executed a preset number of times, the indicator data continuously meets the corresponding indicator threshold for the second preset time, the statistical significance of the candidate version meets the statistical significance requirements, etc.

[0102] The second termination condition includes the candidate version's indicator data not meeting the corresponding indicator threshold, the baseline version's indicator data being abnormal (such as a sudden increase in latency or a crash), and the user manually terminating the process.

[0103] Of course, the first termination condition and the second termination condition can be set and adjusted according to actual evaluation conditions.

[0104] If the performance evaluation experiment meets the experiment termination conditions, the performance evaluation experiment is terminated and the experimental results are output. If the performance evaluation experiment does not meet the experiment termination conditions, the steps of the performance evaluation experiment are executed based on the data cache pool and the performance evaluation experiment is continued until the experiment termination conditions are met.

[0105] In this embodiment, multiple experiment termination conditions are set. When the performance of the service being evaluated meets the standard, the experiment is automatically terminated. If the difference between the candidate version and the baseline version is not significant, the experiment can be prevented from running indefinitely. If the experiment is abnormal, the experiment is terminated promptly. This can avoid wasting resources and improve evaluation efficiency.

[0106] In some embodiments, outputting experimental results includes: When a persistent volume is configured, a report file corresponding to the experimental results is generated and saved to the persistent volume; When cloud storage is configured, generate a report file corresponding to the experimental results and upload it to the cloud storage; When email information is configured, a report file corresponding to the experimental results will be generated and sent to the corresponding email address.

[0107] Specifically, when a performance evaluation experiment terminates, its status is checked. This status can be either success or failure. For example, if the performance evaluation experiment triggers an early termination condition and terminates, the status is failure. Regardless of whether the performance evaluation experiment succeeds or fails, an experiment report is immediately generated, summarizing the indicator data comparison results and statistical significance conclusions.

[0108] The experimental results can be output in a variety of ways: when a persistent volume (PersistentVolumeClaim, PVC) is configured, a report file corresponding to the experimental results is generated and saved to the persistent volume for subsequent access.

[0109] If cloud storage is configured, generate a report file corresponding to the experimental results and upload it to the cloud storage platform, such as Amazon Simple Storage Service (AWS S3).

[0110] When email information is configured, a report file corresponding to the experimental results is generated and sent to the corresponding mailbox via a webhook.

[0111] Taking cloud storage as an example, the report file corresponding to the experimental results is as follows: report: format: "html" destination: type: "s3" bucket: "reports-bucket" path: "reports / {{ .Experiment.Name}} / {{ .Timestamp}}.html" credentials: secretName: "s3-creds" trigger: "on_termination" The parameter format is used to configure the report file format, which supports JavaScript Object Notation (JSON) or Hypertext Markup Language (HTML).

[0112] The parameter destination.type is used to configure the cloud storage type and matches the target cloud storage platform. For example, "s3" matches AWS S3.

[0113] The destination.bucket parameter is used to configure the destination bucket. The destination bucket must be created in advance and ensured to be accessible. Buckets are the fundamental containers used to organize and manage data in cloud storage services. They are similar to folders in a local file system, but their functionality and features are more suited to cloud scenarios. Buckets can hold massive amounts of files, support flexible permission management, and seamlessly integrate with various cloud services. They are the core unit for organizing data in cloud storage.

[0114] The destination.path parameter is used to configure the path of the report file in the storage bucket. It supports Go template syntax and dynamic variables, such as the experiment name and timestamp. {{ .Timestamp}} generates a timestamp to avoid file name conflicts.

[0115] The destination.credentials.secretName parameter is used to reference the K8sSecret name that stores authentication information.

[0116] The parameter trigger is used to configure the upload trigger timing. It can be configured as on_termination, which means automatic upload when the experiment terminates; it can also be configured as periodic, which means periodic upload, which needs to be combined with the interval field.

[0117] After the report file is sent, you can clean up the related resources of the performance evaluation experiment, such as temporary test Pods and traffic distribution strategies, to prevent residual resources from occupying the system.

[0118] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0119] See also Figure 11 , an embodiment of the present application also provides a computer program product 100. The computer program product 100 includes a creation module 10, an execution module 20 and an evaluation module 30. The creation module 10 is used to create a performance evaluation experiment corresponding to the service to be evaluated, and the service to be evaluated includes a pre-trained language model service. The execution module 20 is used to perform a performance evaluation experiment based on a data cache pool, wherein the data cache pool is used to partition and cache a variety of experimental test data, and the various experimental test data are obtained by processing a plurality of data sets using a plurality of different data processing methods. The evaluation module 30 is used to collect data on the streaming response of the service to be evaluated during the execution of the performance evaluation experiment, so as to dynamically evaluate the service to be evaluated.

[0120] In some embodiments, the creation module 10 is specifically used to determine the experimental objectives corresponding to the service to be evaluated and define the corresponding experimental rules; perform a validity check on the experimental objectives and experimental rules; and create a performance evaluation experiment based on the experimental objectives and experimental rules when the experimental objectives and experimental rules pass the validity check.

[0121] In some embodiments, the computer program product 100 further includes a configuration module. Before creating a performance evaluation experiment corresponding to the service to be evaluated, the configuration module is configured to add the access address of the service to be evaluated and configure the corresponding model name and access key. The creation module 10 is specifically configured to initiate a request to the service to be evaluated based on the access address, model name, and access key; and based on the request result, determine whether the experimental objectives and experimental rules pass the validity check.

[0122] In certain embodiments, before creating a performance evaluation experiment corresponding to the service to be evaluated, the configuration module is used to add the access address of the service to be evaluated and configure the corresponding model name and access tag. The execution module 20 is specifically used to initiate a request to the service to be evaluated based on the access address, model name, and access tag, read target test data from the data cache pool, and send it to the service to be evaluated; and receive a streaming response from the service to be evaluated. The evaluation module 30 is specifically used to collect data from the streaming response according to a preset collection frequency, determine indicator data, and dynamically evaluate the service to be evaluated based on the indicator data.

[0123] In some embodiments, the execution module 20 is specifically used to determine the target data set and target data processing method corresponding to the target test data; determine the cache tag of the target test data based on the target data set and target data processing method; read the target test data from the data cache pool according to the cache tag and send it to the service to be evaluated.

[0124] In some embodiments, the cache tag includes a first tag and a second tag. The execution module 20 is specifically configured to determine whether a target dataset is included in multiple datasets corresponding to the multiple experimental test data; if the target dataset is not included, import the target dataset to determine the first tag; if the target dataset is included, use the tag of the target dataset as the first tag; determine whether a target data processing method is included in multiple data processing methods corresponding to the multiple experimental test data; if the target data processing method is not included, import the target data processing method to determine the second tag; if the target data processing method is included, use the tag of the target data processing method as the second tag.

[0125] In some embodiments, the execution module 20 is specifically used to determine whether there is experimental test data corresponding to the target test data in the data cache pool based on the cache tag; if there is no corresponding experimental test data, the target data set is processed according to the target data processing method to obtain the target test data; the target test data is cached to the data cache pool, and the step of determining whether there is experimental test data corresponding to the target test data in the data cache pool based on the cache tag is returned; if there is corresponding experimental test data, the experimental test data is obtained as the target test data and sent to the service to be evaluated.

[0126] In some embodiments, the evaluation module 30 is specifically used to compare the data differences between the indicator data of the baseline version of the service to be evaluated and the indicator data of the candidate version; verify the statistical significance of the data differences through a statistical hypothesis testing algorithm; and respectively verify whether the indicator data of the baseline version and the indicator data of the candidate version meet the indicator thresholds corresponding to the indicator data.

[0127] In some embodiments, the execution module 20 is also used to establish a cache partition in the data cache pool to cache the updated test data corresponding to the updated data set and / or data processing method as experimental test data when the data set and / or data processing method corresponding to the experimental test data are updated in the data cache pool; determine whether the performance evaluation experiment includes a sub-experiment corresponding to the data set and / or data processing method; if the corresponding sub-experiment is included, wait for the sub-experiment to be completed and clear the experimental test data corresponding to the data set and / or data processing method; if the corresponding sub-experiment is not included, directly clear the experimental test data corresponding to the data set and / or data processing method.

[0128] In some embodiments, the indicator data includes dynamic streaming delay. The evaluation module 30 is specifically configured to obtain, in real time, the generation timestamps of multiple tags in the streaming response; calculate the arrival delays between adjacent tags based on the generation timestamps; determine the weights corresponding to the multiple tags based on the generation order of the multiple tags; and calculate the dynamic streaming delay based on the arrival delays and the weights.

[0129] In some embodiments, the computer program product 100 also includes a termination module. After executing the performance evaluation experiment based on the data cache pool, the termination module is specifically used to determine whether the performance evaluation experiment meets the experiment termination conditions based on the indicator data; if the experiment termination conditions are met, the performance evaluation experiment is terminated and the experimental results are output; if the experiment termination conditions are not met, the module returns to the step of executing the performance evaluation experiment based on the data cache pool.

[0130] In some embodiments, the termination module is specifically used to generate a report file corresponding to the experimental results and save it to the persistent volume when a persistent volume is configured; to generate a report file corresponding to the experimental results and upload it to the cloud storage when cloud storage is configured; and to generate a report file corresponding to the experimental results and send it to the corresponding mailbox when email information is configured.

[0131] For descriptions of features in the embodiment corresponding to the computer program product 100, reference may be made to the relevant descriptions of the embodiment corresponding to the service performance evaluation method, which will not be detailed here.

[0132] See also Figure 12 An embodiment of the present application further provides an electronic device 200, comprising a memory 210 and a processor 220, wherein the memory 210 stores a computer program, and the processor 220 is configured to run the computer program to execute the steps in any one of the above-mentioned service performance evaluation method embodiments.

[0133] See also Figure 13 An embodiment of the present application further provides a computer-readable storage medium 300, in which a computer program 310 is stored, wherein the computer program 310 is configured to execute the steps of any of the above-mentioned service performance evaluation method embodiments when running.

[0134] In an exemplary embodiment, the computer-readable storage medium 300 may include, but is not limited to, various media that can store the computer program 310, such as a USB flash drive, a read-only memory 210 (ROM), a random access memory 210 (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0135] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned service performance evaluation method embodiments are implemented.

[0136] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0137] The above is a detailed introduction to a service performance evaluation method, computer program product, electronic device, and computer-readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A service performance evaluation method, characterized in that: include: Creating a performance evaluation experiment corresponding to the service to be evaluated, wherein the service to be evaluated includes a pre-trained language model service; The performance evaluation experiment is performed based on a data cache pool, wherein the data cache pool is used to partition and cache a variety of experimental test data, and the various experimental test data are obtained by processing a plurality of data sets respectively through a plurality of data processing methods; During the execution of the performance evaluation experiment, data is collected from the streaming response of the service to be evaluated, so as to dynamically evaluate the service to be evaluated.

2. The service performance evaluation method according to claim 1, characterized in that: The step of creating a performance evaluation experiment corresponding to the service to be evaluated includes: Determine the experimental objectives corresponding to the service to be evaluated and define the corresponding experimental rules; Performing a validity check on the experimental objectives and the experimental rules; When the experimental objective and the experimental rules pass the legality check, the performance evaluation experiment is created based on the experimental objective and the experimental rules.

3. The service performance evaluation method according to claim 2, characterized in that: Before creating a performance evaluation experiment corresponding to the service to be evaluated, the service performance evaluation method further includes: Add the access address of the service to be evaluated and configure the corresponding model name and access key; The legitimacy verification of the experimental objectives and the experimental rules includes: Initiate a request for the service to be evaluated based on the access address, the model name, and the access key; Determine whether the experimental objective and the experimental rules pass the legality check according to the request result.

4. The service performance evaluation method according to claim 1, wherein: Before creating a performance evaluation experiment corresponding to the service to be evaluated, the service performance evaluation method further includes: Add the access address of the service to be evaluated and configure the corresponding model name and access tag; The performing of the performance evaluation experiment based on the data cache pool includes: Initiating a request to the service to be evaluated based on the access address, the model name, and the access tag, and reading target test data from the data cache pool and sending it to the service to be evaluated; receiving the streaming response of the service to be evaluated; The collecting data of the streaming response of the service to be evaluated to dynamically evaluate the service to be evaluated includes: Collect data from the streaming response according to a preset collection frequency to determine indicator data; Dynamically evaluate the service to be evaluated based on the indicator data.

5. The service performance evaluation method according to claim 4, characterized in that: The step of reading target test data from the data cache pool and sending the data to the service to be evaluated includes: Determining a target data set and a target data processing method corresponding to the target test data; Determining a cache tag of the target test data according to the target data set and the target data processing method; The target test data is read from the data cache pool according to the cache tag and sent to the service to be evaluated.

6. The service performance evaluation method according to claim 5, characterized in that: The cache tag includes a first tag and a second tag, and determining the cache tag of the target test data according to the target data set and the target data processing method includes: Determining whether the plurality of data sets corresponding to the plurality of experimental test data include the target data set; importing the target dataset without including the target dataset to determine the first label; In a case where the target data set is included, using a label of the target data set as the first label; Determining whether the plurality of data processing methods corresponding to the plurality of experimental test data include the target data processing method; In a case where the target data processing method is not included, importing the target data processing method to determine the second label; In the case where the target data processing method is included, the label of the target data processing method is used as the second label.

7. The service performance evaluation method according to claim 5, characterized in that: The step of reading the target test data from the data cache pool according to the cache tag and sending the data to the service to be evaluated includes: Based on the cache tag, determining whether the experimental test data corresponding to the target test data exists in the data cache pool; In the case where the corresponding experimental test data does not exist, performing data processing on the target data set according to the target data processing method to obtain the target test data; caching the target test data in the data cache pool, and returning to the step of determining whether the experimental test data corresponding to the target test data exists in the data cache pool based on the cache tag; In the case that the corresponding experimental test data exists, the experimental test data is obtained as the target test data and sent to the service to be evaluated.

8. The service performance evaluation method according to claim 4, characterized in that: The dynamically evaluating the service to be evaluated according to the indicator data includes: Comparing the data differences between the indicator data of the baseline version of the service to be evaluated and the indicator data of the candidate version; Verify the statistical significance of the data differences through statistical hypothesis testing algorithms; The indicator data of the baseline version and the indicator data of the candidate version are respectively checked to see whether they meet the indicator thresholds corresponding to the indicator data.

9. The service performance evaluation method according to claim 1, wherein: The performing of the performance evaluation experiment based on the data cache pool includes: In the data cache pool, when the data set and / or the data processing method corresponding to the experimental test data are updated, a cache partition is established in the data cache pool to cache the updated data set and / or updated test data corresponding to the data processing method as the experimental test data; Determining whether the performance evaluation experiment includes a sub-experiment corresponding to the data set and / or the data processing method; In the case of including a corresponding sub-experiment, waiting for the sub-experiment to be completed, and clearing the experimental test data corresponding to the data set and / or the data processing method; In the case where the corresponding sub-experiment is not included, the experimental test data corresponding to the data set and / or the data processing method is directly cleared.

10. The service performance evaluation method according to claim 4, characterized in that: The indicator data includes a dynamic streaming delay, and the data collection of the streaming response according to a preset collection frequency to determine the indicator data includes: Obtaining, in real time, generation timestamps of the plurality of tags in the streaming response; Calculating the arrival delay between adjacent markers according to the generation timestamps; Determining weights corresponding to the multiple tags according to the generation order of the multiple tags; Calculating the dynamic streaming delay based on the arrival delay and the weight; The calculation formula of the weight is: in, is the weight corresponding to the i-th tag among the multiple tags, 0 <k<1,c> 0.

11. The service performance evaluation method according to claim 4, characterized in that: After executing the performance evaluation experiment based on the data cache pool, the service performance evaluation method further includes: Determining whether the performance evaluation experiment meets the experiment termination condition according to the indicator data; When the experiment termination condition is met, terminating the performance evaluation experiment and outputting the experimental results; If the experiment termination condition is not met, return to the step of executing the performance evaluation experiment based on the data cache pool.

12. The service performance evaluation method according to claim 11, characterized in that: The output experimental results include: When a persistent volume is configured, a report file corresponding to the experimental results is generated and saved to the persistent volume; When cloud storage is configured, generate a report file corresponding to the experimental results and upload it to the cloud storage; When email information is configured, a report file corresponding to the experimental results is generated and sent to the corresponding mailbox.

13. A computer program product, characterized in that include: A creation module is used to create a performance evaluation experiment corresponding to the service to be evaluated, wherein the service to be evaluated includes a pre-trained language model service; an execution module, configured to execute the performance evaluation experiment based on a data cache pool, wherein the data cache pool is configured to cache a plurality of experimental test data in partitions, wherein the plurality of experimental test data are obtained by processing a plurality of data sets respectively through a plurality of different data processing methods; The evaluation module is used to collect data on the streaming response of the service to be evaluated during the execution of the performance evaluation experiment, so as to dynamically evaluate the service to be evaluated.

14. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the service performance evaluation method according to any one of claims 1 to 12 when executing the computer program.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the service performance evaluation method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Data processing method, device, equipment, medium and computer program product

    CN119090560A

  • Language model reasoning method and device, computer equipment and storage medium

    CN120069084A

  • Fusion platform construction method and system based on intelligent information processing and data analysis

    CN120086013A

  • Information acquisition and analysis method, device, equipment, medium and product

    CN120106232A

  • Large language model reasoning performance evaluation and optimization method, electronic equipment and storage medium

    CN120297409A