Method, system and device for evaluating algorithm quality and storage medium
Patent Information
- Application Number
- CN202611089528.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]然而,随着算法模型结构日益复杂、迭代频率不断加快,传统的算法质量的评测方法难以在有限的时间和资源下全面模拟并预判算法部署到生产环境后的真实表现
[0016] The aforementioned algorithm quality evaluation method, in the sample construction phase, constructs evaluation samples based on request logs from the production environment, ensuring that the evaluation samples cover diverse online interaction requests and complex business scenarios, thus improving the efficiency and comprehensiveness of sample construction. In the evaluation execution phase, by obtaining an evaluation sample set containing multiple evaluation samples, the algorithm under test and the benchmark algorithm are evaluated separately based on the evaluation sample set, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. This effectively assesses the performance change of the algorithm under test compared to the benchmark algorithm, providing a reliable data foundation for subsequent indicator calculation and comparative analysis. In the indicator calculation phase, through... The first indicator value for each evaluation indicator corresponding to the first evaluation result set and the second indicator value for each evaluation indicator corresponding to the second evaluation result set are calculated respectively. The evaluation indicators include technical indicators and algorithm indicators. The technical indicators are used to characterize the performance of the algorithm chain, and the algorithm indicators are used to characterize the effect of the algorithm operation. This allows for a quantitative evaluation of the algorithm's output results from two dimensions: algorithm chain performance and algorithm operation effect, avoiding the limitations of single-dimensional evaluation. In the result analysis stage, the evaluation results of the algorithm under test are determined based on the first and second indicator values of each evaluation indicator. This allows for the identification of potential quality problems in the algorithm under test based on indicator differences, and effectively predicts the performance of the algorithm under test after deployment to the production environment.
Smart Images

Figure CN122673067A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of evaluation technology, and in particular to a method, system, device and storage medium for evaluating algorithm quality. Background Technology
[0002] In the process of algorithm development and updates, in order to ensure the quality and stability of the algorithm, a set of evaluation processes are usually required, such as manually designing test cases and executing automated regression tests to verify the algorithm's functionality.
[0003] However, as algorithm model structures become increasingly complex and iteration frequencies accelerate, traditional algorithm quality evaluation methods struggle to fully simulate and predict the real-world performance of algorithms deployed to production environments within limited time and resources. For example, some tuned algorithms may perform normally when executing manually designed test cases, but when faced with diverse real-world requests and complex interactive environments online, their output data distribution trends may deviate from expectations. This forces developers to rely on real-world business feedback after algorithm deployment to identify potential problems, increasing debugging and correction costs. Therefore, a more efficient and comprehensive algorithm quality evaluation method is needed. Summary of the Invention
[0004] This application provides a method, system, apparatus, device, and storage medium for evaluating algorithm quality, which can improve evaluation efficiency. The above technical solution is as follows: In a first aspect, embodiments of this application provide a method for evaluating algorithm quality, applied to a server, including: Obtain the evaluation sample set; the evaluation sample set includes multiple evaluation samples, each of which is constructed based on the request logs in the production environment; Based on the evaluation sample set and preset control parameters, an evaluation request set is constructed and generated; the control parameters include at least one of environmental parameters, experimental parameters, and shunting parameters; the evaluation request set includes multiple evaluation requests; The evaluation request set is input into the algorithm under test and the benchmark algorithm respectively for evaluation, and the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm are obtained. Calculate the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set respectively; The evaluation results of the algorithm under test are determined based on the first and second indicator values of each evaluation indicator. The evaluation metrics include technical metrics and algorithm metrics. Technical metrics are used to characterize the performance of the algorithm chain, while algorithm metrics are used to characterize the effect of the algorithm operation. Technical metrics and algorithm metrics have a hierarchical dependency relationship when determining the evaluation results.
[0005] In one possible implementation, before obtaining the evaluation sample set, the method further includes: obtaining multiple request logs in the production environment; filtering the multiple request logs according to log filtering conditions to obtain multiple target request logs; the log filtering conditions include at least one of a first environment identifier, a first application identifier, and a first scenario identifier; modifying the request parameters in each target request log to generate each evaluation sample; storing each evaluation sample in an evaluation sample database; obtaining the evaluation sample set includes: filtering the evaluation samples in the evaluation sample database according to sample filtering conditions to obtain the evaluation sample set; the sample filtering conditions include at least one of a second environment identifier, a second application identifier, and a second scenario identifier.
[0006] In one possible implementation, the algorithm under test is not released to the production environment. The evaluation request set is input to both the algorithm under test and the benchmark algorithm for evaluation, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. This includes: if the pre-release environment and the production environment meet preset comparability conditions, the evaluation request set is input to both the algorithm under test in the pre-release environment and the benchmark algorithm in the production environment for evaluation, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm; if the pre-release environment and the production environment do not meet preset comparability conditions, the evaluation request set is input to both the algorithm under test and the benchmark algorithm under different pre-release branches in the pre-release environment for evaluation, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. The preset comparability conditions include that the difference in key environment configurations is within a preset difference range; the key environment configurations include at least one of the following: algorithm running configuration, algorithm experiment configuration, and algorithm model configuration.
[0007] In one possible implementation, the algorithm under test has been deployed to the production environment; the evaluation request set is input to the algorithm under test and the benchmark algorithm respectively for evaluation, to obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm, including: inputting the evaluation request set to the algorithm under test and the benchmark algorithm under different experimental buckets in the production environment for evaluation, to obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm.
[0008] In one possible implementation, the first indicator value includes a first technical indicator value and a first algorithm indicator value, and the second indicator value includes a second technical indicator value and a second algorithm indicator value. Based on the first and second indicator values of each evaluation indicator, the evaluation result of the algorithm under test is determined, including: determining whether the first and second technical indicator values corresponding to each technical indicator are both within a preset technical indicator range; if any first or second technical indicator value corresponding to any technical indicator is not within the preset technical indicator range, then the evaluation result of the algorithm under test is determined to be an abnormal link performance; if the first and second technical indicators corresponding to each technical indicator are both within the preset technical indicator range... Within the interval, the difference between the first and second technical indicator values corresponding to each technical indicator is calculated to obtain at least one technical indicator difference value. If any technical indicator difference value is greater than the corresponding technical indicator difference threshold, the evaluation result of the algorithm under test is determined to be an abnormal algorithm link performance. If the difference values of each technical indicator are less than or equal to the corresponding technical indicator difference threshold, indicator deviation analysis is performed on each algorithm indicator based on the first and second algorithm indicator values corresponding to each algorithm indicator. Indicator deviation analysis includes at least one of the following: outlier analysis, data distribution analysis, and difference analysis. The evaluation result of the algorithm under test is determined based on the results of the indicator deviation analysis.
[0009] In one possible implementation, outlier analysis includes: identifying whether there are outliers in the first and second algorithm indicator values corresponding to each algorithm indicator, and triggering anomaly attribution verification if outliers are found; data distribution analysis includes: determining whether there are abnormal distributions in the data distribution of the algorithm indicators in the first and second algorithm indicator values, and triggering anomaly attribution verification if abnormal distributions are found; difference analysis includes: calculating the difference between the first and second algorithm indicator values, and triggering anomaly attribution verification if the difference exceeds a preset difference threshold; determining the evaluation result of the algorithm under test based on the results of the indicator deviation analysis includes: determining the evaluation result of the algorithm under test based on the results of the anomaly attribution verification if the results of the indicator deviation analysis trigger anomaly attribution verification; and determining the evaluation result of the algorithm under test as passed if the results of the indicator deviation analysis do not trigger anomaly attribution verification; wherein, anomaly attribution verification includes at least one of environment configuration verification, request parameter verification, experiment hit verification, link configuration verification, and model parameter verification.
[0010] Secondly, embodiments of this application provide a method for evaluating algorithm quality, applied to a terminal, including: In response to a user's evaluation trigger action, an evaluation request is generated; A test request is sent to the server, causing the server to: respond to the test request and obtain a test sample set; the test sample set includes multiple test samples, each constructed based on request logs in the production environment; construct and generate a test request set based on the test sample set and preset control parameters; the control parameters include at least one of environmental parameters, experimental parameters, and traffic splitting parameters; the test request set includes multiple test requests; input the test request set to the algorithm under test and the benchmark algorithm respectively for testing, obtaining a first test result set corresponding to the algorithm under test and a second test result set corresponding to the benchmark algorithm; calculate the first indicator value of each test metric corresponding to the first test result set and the second indicator value of each test metric corresponding to the second test result set; determine the test result of the algorithm under test based on the first indicator value and the second indicator value of each test metric; and send the test result to the terminal. Receive evaluation results sent by the server; Display based on the evaluation results; The evaluation metrics include technical metrics and algorithm metrics. Technical metrics are used to characterize the performance of the algorithm chain, while algorithm metrics are used to characterize the effect of the algorithm operation. Technical metrics and algorithm metrics have a hierarchical dependency relationship when determining the evaluation results.
[0011] Thirdly, embodiments of this application provide an algorithm quality evaluation system, comprising: a terminal and a server; wherein, The terminal is used to respond to the user's evaluation trigger operation, generate an evaluation request, and send the evaluation request to the server; The server responds to evaluation requests and obtains an evaluation sample set. The evaluation sample set includes multiple evaluation samples, each constructed based on request logs from the production environment. Based on the evaluation sample set and preset control parameters, an evaluation request set is generated. The control parameters include at least one of environmental parameters, experimental parameters, and traffic splitting parameters. The evaluation request set includes multiple evaluation requests. The evaluation request set is input to the algorithm under test and the benchmark algorithm for evaluation, respectively, to obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm. The server calculates the first indicator value for each evaluation metric corresponding to the first evaluation result set and the second indicator value for each evaluation metric corresponding to the second evaluation result set. Based on the first and second indicator values of each evaluation metric, the server determines the evaluation result of the algorithm under test and sends the evaluation result to the terminal. The terminal is also used to receive evaluation results sent by the server and display them based on the evaluation results; The evaluation metrics include technical metrics and algorithm metrics. Technical metrics are used to characterize the performance of the algorithm chain, while algorithm metrics are used to characterize the effect of the algorithm operation. Technical metrics and algorithm metrics have a hierarchical dependency relationship when determining the evaluation results.
[0012] Fourthly, embodiments of this application provide an algorithm quality evaluation device, applied to a server, comprising: The acquisition module is used to acquire the evaluation sample set; the evaluation sample set includes multiple evaluation samples, each of which is constructed based on the request logs in the production environment; The execution module is used to construct and generate an evaluation request set based on the evaluation sample set and preset control parameters. The control parameters include at least one of environmental parameters, experimental parameters, and shunting parameters. The evaluation request set includes multiple evaluation requests. The evaluation request set is input to the algorithm under test and the benchmark algorithm respectively for evaluation, and a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm are obtained. The calculation module is used to calculate the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set. The determination module is used to determine the evaluation result of the algorithm under test based on the first and second indicator values of each evaluation indicator. The evaluation metrics include technical metrics and algorithm metrics. Technical metrics are used to characterize the performance of the algorithm chain, while algorithm metrics are used to characterize the effect of the algorithm operation. Technical metrics and algorithm metrics have a hierarchical dependency relationship when determining the evaluation results.
[0013] Fifthly, embodiments of this application provide an algorithm quality evaluation device, applied to a terminal, comprising: The response module is used to respond to user-triggered evaluation actions and generate evaluation requests; The sending module is used to send evaluation requests to the server, so that the server: responds to the evaluation request and obtains an evaluation sample set; the evaluation sample set includes multiple evaluation samples, each constructed based on request logs in the production environment; constructs and generates an evaluation request set based on the evaluation sample set and preset control parameters; the control parameters include at least one of environmental parameters, experimental parameters, and traffic splitting parameters; the evaluation request set includes multiple evaluation requests; inputs the evaluation request set to the algorithm under test and the benchmark algorithm respectively for evaluation, and obtains a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm; calculates the first indicator value of each evaluation metric corresponding to the first evaluation result set and the second indicator value of each evaluation metric corresponding to the second evaluation result set respectively; determines the evaluation result of the algorithm under test based on the first indicator value and the second indicator value of each evaluation metric; and sends the evaluation result to the terminal. The receiving module is used to receive the evaluation results sent by the server; The display module is used to display the evaluation results.
[0014] In a sixth aspect, embodiments of this application provide an electronic device, including: a processor and a memory; the memory stores a computer program, and the processor executes the computer program to implement the method steps provided in the first or second aspect of embodiments of this application.
[0015] In a seventh aspect, embodiments of this application provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps provided in the first or second aspect of embodiments of this application.
[0016] The aforementioned algorithm quality evaluation method, in the sample construction phase, constructs evaluation samples based on request logs from the production environment, ensuring that the evaluation samples cover diverse online interaction requests and complex business scenarios, thus improving the efficiency and comprehensiveness of sample construction. In the evaluation execution phase, by obtaining an evaluation sample set containing multiple evaluation samples, the algorithm under test and the benchmark algorithm are evaluated separately based on the evaluation sample set, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. This effectively assesses the performance change of the algorithm under test compared to the benchmark algorithm, providing a reliable data foundation for subsequent indicator calculation and comparative analysis. In the indicator calculation phase, through... The first indicator value for each evaluation indicator corresponding to the first evaluation result set and the second indicator value for each evaluation indicator corresponding to the second evaluation result set are calculated respectively. The evaluation indicators include technical indicators and algorithm indicators. The technical indicators are used to characterize the performance of the algorithm chain, and the algorithm indicators are used to characterize the effect of the algorithm operation. This allows for a quantitative evaluation of the algorithm's output results from two dimensions: algorithm chain performance and algorithm operation effect, avoiding the limitations of single-dimensional evaluation. In the result analysis stage, the evaluation results of the algorithm under test are determined based on the first and second indicator values of each evaluation indicator. This allows for the identification of potential quality problems in the algorithm under test based on indicator differences, and effectively predicts the performance of the algorithm under test after deployment to the production environment.
[0017] By adopting the above-mentioned algorithm quality evaluation method, an evaluation sample set from a real production environment is obtained. Based on the evaluation sample set, the algorithm under test and the benchmark algorithm are evaluated separately. The evaluation results are determined by comparing and analyzing the evaluation index values of the two algorithms. This method can achieve efficient and comprehensive evaluation of algorithm iteration. Potential problems can be discovered without waiting for the algorithm to go online and through real business feedback. This reduces the debugging and correction costs of algorithm iteration and improves the efficiency and comprehensiveness of algorithm evaluation. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A schematic diagram illustrating an application scenario of an algorithm quality evaluation method provided for an exemplary embodiment of this application; Figure 2 A schematic diagram illustrating the application environment of an algorithm quality evaluation method provided for an exemplary embodiment of this application; Figure 3 A flowchart illustrating an algorithm quality evaluation method provided for an exemplary embodiment of this application; Figure 4 A flowchart illustrating another algorithm quality evaluation method provided for an exemplary embodiment of this application; Figure 5 A schematic diagram of the structure of an algorithm quality evaluation device provided for an exemplary embodiment of this application; Figure 6 A schematic diagram of the structure of another algorithm quality evaluation device provided as an exemplary embodiment of this application; Figure 7 A schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application; Figure 8 This is a schematic diagram of the structure of another electronic device provided as an exemplary embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0021] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0022] First, combine Figure 1 This application describes the applicable scenarios for the embodiments of this application: An application plans to iterate and upgrade its existing algorithm B to a new algorithm A. To ensure that the new algorithm A meets expectations after its launch, it needs to be evaluated before the official release to predict its actual performance.
[0023] Please see Figure 1 The algorithm evaluation process provided in this application embodiment can be divided into: sample construction stage, evaluation execution stage, index calculation stage, and result analysis stage. The specific implementation process of each stage is as follows: During the sample construction phase, the server first obtains request logs from the production environment; then, it filters the request logs based on log filtering conditions such as the first environment identifier, the first application identifier, and the first scenario identifier to obtain target request logs for the specified environment, application, and scenario; next, it extracts core parameters such as application identifier and request attributes from each target request log, customizes these parameters according to the needs of the evaluation task, introduces control variables, and generates each evaluation sample; finally, it stores each generated evaluation sample in the evaluation sample database.
[0024] During the evaluation execution phase, the server first determines sample filtering conditions such as the second environment identifier, second application identifier, and second scenario identifier based on the environment, application, and scenario involved in the algorithm under test (A). Then, it filters the evaluation samples in the evaluation sample database according to these conditions to obtain an evaluation sample set. Next, it concatenates preset control parameters (such as environmental parameters, experimental parameters, and traffic splitting parameters) into each evaluation sample in the evaluation sample set, forming multiple evaluation requests to construct an evaluation request set. Finally, the evaluation request set is input to both the algorithm under test (A) and the benchmark algorithm (B) for evaluation, obtaining a first evaluation result set for algorithm A and a second evaluation result set for benchmark algorithm B. It is worth noting that the server can perform dual-run evaluations in the same or different environments. For example, when the differences in environment configuration between the pre-production and production environments are controllable, the server can input the evaluation request set into the test algorithm A in the pre-production environment and the benchmark algorithm B in the production environment for evaluation. When the differences in environment configuration between the pre-production and production environments are uncontrollable, the server can input the evaluation request set into the test algorithm A and the benchmark algorithm B under different pre-production branches in the pre-production environment for evaluation. Furthermore, if evaluation is involved in the production environment, a traffic control mechanism can be used to add evaluation labels to the evaluation traffic to prevent it from affecting online data statistics.
[0025] During the metric calculation phase, the server calculates the first metric value for each metric corresponding to the first evaluation result set and the second metric value for each metric corresponding to the second evaluation result set. The evaluation metrics include technical metrics and algorithmic metrics. The technical metrics characterize the performance of the algorithmic process, such as request success rate and request timeout rate. The algorithmic metrics characterize the effectiveness of the algorithm's operation, such as output value distribution, decision threshold distribution, recall rate, and precision rate.
[0026] During the results analysis phase, the server first checks whether the values of each technical indicator are normal. For example, it checks whether the values of each technical indicator are within the preset range and whether the difference between the first technical indicator value (calculated from the first evaluation result set) and the second technical indicator value (calculated from the second evaluation result set) is too large. If any technical indicator value is abnormal, the server determines that the evaluation result of the algorithm under test, A, is an algorithm link performance anomaly. If all technical indicator values are normal, the server checks whether the values of each algorithm indicator are normal. For example, it performs indicator deviation analysis on each algorithm indicator based on the first algorithm indicator value (calculated from the first evaluation result set) and the second algorithm indicator value (calculated from the second evaluation result set). If the result of the indicator deviation analysis triggers anomaly attribution verification, the server determines the evaluation result of the algorithm under test, A, based on the result of the anomaly attribution verification. If the result of the indicator deviation analysis does not trigger anomaly attribution verification, the server determines that the evaluation result of the algorithm under test, A, is passed.
[0027] The algorithm quality evaluation method provided in this application embodiment can be applied to, for example, Figure 2 In the application environment shown, terminal 10, equipped with the intelligent evaluation platform client, communicates with the intelligent evaluation platform server 20 via communication network 30. Data storage system 40 can store data that the server 20 needs to process, such as evaluation sample data. Data storage system 40 can be integrated onto server 20 or located on the cloud or other network servers. Terminal 10 can be, but is not limited to, various smartphones, tablets, personal computers, laptops, etc. Server 20 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0028] In some possible embodiments, terminal 10, in response to a user's evaluation trigger operation, generates an evaluation request and sends it to server 20 via communication network 30. Server 20, in response to the evaluation request, retrieves an evaluation sample set from data storage system 40; the evaluation sample set includes multiple evaluation samples, each constructed based on request logs in the production environment; it evaluates the algorithm under test and the benchmark algorithm based on the evaluation sample set, obtaining a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm; it calculates the first indicator value of each evaluation metric corresponding to the first evaluation result set and the second indicator value of each evaluation metric corresponding to the second evaluation result set; the evaluation indicators include technical indicators and algorithm indicators, where technical indicators characterize the performance of the algorithm chain and algorithm indicators characterize the effect of the algorithm operation; based on the first and second indicator values of each evaluation indicator, it determines the evaluation result of the algorithm under test; and sends the evaluation result to the terminal. Terminal 10 receives the evaluation result sent by the server and displays it based on the evaluation result.
[0029] It is worth noting that the aforementioned technical metrics include, but are not limited to, request success rate and request timeout rate. The aforementioned algorithmic metrics can be proxy metrics obtained by breaking down the goals expected to be achieved through algorithm iteration (such as the reasonableness of the output distribution, the coverage and accuracy of the output results), which can be directly calculated in the pre-release phase and are positively correlated with the goals. For example, algorithmic metrics include, but are not limited to, output value distribution, decision threshold distribution, recall rate, and precision rate.
[0030] In one embodiment, such as Figure 3 As shown, a method for evaluating algorithm quality is provided, and this method is applied to... Figure 2 The following steps are used as an example to illustrate the application environment shown: S301: Terminal 10 responds to the user's evaluation trigger operation and generates an evaluation request.
[0031] The evaluation trigger operation refers to the interactive action initiated by the user on terminal 10 to start the algorithm evaluation process, such as entering evaluation commands or clicking the evaluation button. The evaluation request contains the configuration information required for this evaluation task, such as the identifier of the algorithm under test, the identifier of the benchmark algorithm, and the evaluation range parameters.
[0032] Understandably, the algorithm under test refers to an algorithm that requires quality evaluation, such as a newly developed algorithm or an algorithm with adjusted parameters. The benchmark algorithm is a reference algorithm compared to the algorithm under test, such as an existing algorithm that is already running stably in a production environment. Evaluation scope parameters refer to configuration parameters used to limit the specific scope and conditions of this evaluation task, including but not limited to sample filtering conditions and control parameters for constructing evaluation requests.
[0033] S302: Terminal 10 sends an evaluation request to server 20.
[0034] Optionally, the terminal 10 sends the generated evaluation request to the server 20 through a network interface, so that the server 20, upon receiving the evaluation request, responds to the evaluation request and performs subsequent operations such as sample construction, evaluation execution, indicator calculation, and result analysis.
[0035] S303: Server 20 responds to the evaluation request and obtains the evaluation sample set.
[0036] The evaluation sample set includes multiple evaluation samples, each built based on request logs from the production environment. The production environment refers to the runtime environment in which the software system provides services and processes real data based on real requests after it has been officially launched. In this environment, the application processes real requests and generates corresponding request logs.
[0037] Optionally, server 20 filters and parameterizes request logs in the production environment at preset time intervals to obtain evaluation samples and stores them in the evaluation sample database. When performing sample evaluation, the evaluation samples in the evaluation sample database are filtered according to the environment, application, and scenario involved in the algorithm under test to obtain an evaluation sample set.
[0038] Specifically, during the sample construction phase, server 20 acquires multiple request logs from the production environment at preset time intervals. It then filters these logs based on log filtering conditions such as a first environment identifier, and / or a first application identifier, and / or a first scenario identifier, to obtain multiple target request logs. Next, it parameterizes the request parameters in each target request log to generate evaluation samples, which are then stored in the evaluation sample database. During the evaluation execution phase, server 20 filters the evaluation samples in the evaluation sample database based on sample filtering conditions such as a second environment identifier, and / or a second application identifier, and / or a second scenario identifier, to obtain an evaluation sample set.
[0039] Understandably, the log filtering conditions and sample filtering conditions described above can be the same or different. For example, during the sample construction phase, samples from multiple environments, applications, and scenarios may be collected, while during the evaluation execution phase, only samples from a specific environment, application, and scenario need to be used for evaluation. In other words, the sample range defined by the sample filtering conditions is a subset of the sample range defined by the log filtering conditions. Thus, by separating sample construction from sample selection, multiple evaluations with different purposes and scopes can be performed based on a single sample construction, avoiding redundant sample construction.
[0040] In this embodiment, server 20 pre-constructs evaluation samples, each corresponding to a real request log in the production environment. This ensures that the evaluation samples can cover diverse online interaction requests and complex interaction scenarios, improving the efficiency and comprehensiveness of sample construction. Furthermore, by separating sample construction from sample selection, multiple evaluations can be performed based on a single sample construction, further improving evaluation efficiency.
[0041] S304: Server 20 evaluates the algorithm under test and the benchmark algorithm based on the evaluation sample set, and obtains the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm.
[0042] To control variables during the evaluation process and ensure that the evaluation samples run only in the specified environment and experimental bucket, thus avoiding interference from other factors on the observability of the evaluation results, server 20 constructs an evaluation request set based on the evaluation sample set and preset control parameters. This evaluation request set is then input into the algorithm under test and the benchmark algorithm for evaluation, respectively, to obtain a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. The evaluation request set includes at least one of environmental parameters, experimental parameters, and load balancing parameters.
[0043] Optionally, server 20 can perform dual-run tests in the same or different environments: For algorithms yet to be released to the production environment, if the differences between the pre-production and production environments in key variable factors such as algorithm runtime configuration, algorithm experiment configuration, and algorithm model data are controllable (i.e., the two environments meet preset comparability conditions), then server 20 will input the evaluation request set into the algorithm under test in the pre-production environment and the benchmark algorithm in the production environment for evaluation, respectively, to obtain the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm. Through this dual-run evaluation in the pre-production and production environments, server 20 can directly compare the output differences between the algorithm under test and the benchmark algorithm running stably in the production environment using the same evaluation sample set before the algorithm under test goes live, and analyze the effect change of the algorithm under test relative to the benchmark algorithm based on the output differences, thereby predicting the actual performance of the algorithm under test after it is released to the production environment. If there are significant differences between the pre-production environment and the production environment in the aforementioned key variable factors (i.e., the two algorithms do not meet the preset comparability conditions), then server 20 will input the evaluation request set into the algorithm under test and the benchmark algorithm under different pre-production branches in the pre-production environment for evaluation, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. Through this dual-run evaluation executed in the pre-production environment, server 20 can use the same evaluation sample set to compare the differences in algorithm output under different experimental branches in the pre-production environment, thus effectively evaluating the performance change of the algorithm under test compared to the benchmark algorithm even when the pre-production environment and the production environment are not comparable.
[0044] For algorithms already deployed to the production environment, the server inputs the evaluation request set into both the algorithm under test and the benchmark algorithm in different experimental buckets within the production environment for evaluation, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. Through this dual-run evaluation in the production environment, the server can compare the differences in algorithm output across different experimental buckets using the same evaluation sample set. This allows for the assessment of the algorithm's performance relative to the benchmark algorithm after deployment to the production environment through traffic control. It's worth noting that for evaluations involving the production environment, the server can add evaluation tags to the evaluation traffic to prevent it from affecting online data statistics.
[0045] In this embodiment, the server 20 evaluates the algorithm under test and the benchmark algorithm based on the evaluation sample set, and obtains the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm. This can effectively evaluate the performance change of the algorithm under test compared to the benchmark algorithm, and provide a reliable data foundation for subsequent index calculation and comparative analysis.
[0046] S305: Server 20 calculates the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set.
[0047] The evaluation metrics include technical metrics and algorithmic metrics. Technical metrics characterize the performance of the algorithm chain, including but not limited to request success rate and request timeout rate. Algorithmic metrics characterize the effectiveness of the algorithm's operation, and can be surrogate metrics that can be directly calculated in the pre-release phase and are positively correlated with algorithm iteration, including but not limited to output value distribution, decision threshold distribution, recall rate, and precision rate. Technical metrics and algorithmic metrics have a hierarchical dependency in determining the evaluation results. Specifically, the metric values of technical metrics are used to determine whether the first and second evaluation result sets have a credible basis for comparison, and the metric values of algorithmic metrics are used, assuming the comparison basis is valid, to determine whether the operation effect of the algorithm under test meets expectations.
[0048] Optionally, the server 20 calculates the first evaluation result set and the second evaluation result set respectively using a pre-written indicator script to obtain the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set.
[0049] Specifically, for technical metrics, server 20 can, based on the first and second evaluation result sets respectively, calculate the proportion of evaluation requests that are successfully processed and return results to obtain the first request success rate and the second request success rate; and calculate the proportion of evaluation requests that do not receive a response within a preset time threshold to obtain the first request timeout rate and the second request timeout rate. For algorithm metrics, server 20 can, based on the first and second evaluation result sets respectively, divide the numerical parameters in the algorithm output results according to preset intervals, and calculate the proportion of evaluation requests in each interval to obtain the output value distribution; divide the threshold parameters on which the algorithm makes judgments according to preset intervals, and calculate the proportion of evaluation requests in each interval to obtain the judgment threshold distribution; calculate the proportion of different types of labels in the algorithm output results to obtain the output type distribution; and calculate the proportion of evaluation requests that fall back to the default strategy in the algorithm output results to obtain the default judgment proportion.
[0050] In this embodiment, server 20 calculates the first index value of each evaluation indicator corresponding to the first evaluation result set and the second index value of each evaluation indicator corresponding to the second evaluation result set. The evaluation indicators include technical indicators characterizing algorithm link performance and algorithmic indicators characterizing algorithm computational effect. This allows for a quantitative evaluation of the algorithm's output from two dimensions: algorithm link performance and algorithm computational effect, avoiding the limitations of single-dimensional evaluation. Furthermore, by establishing a hierarchical dependency between the two in the evaluation result determination, it is possible to first determine whether the evaluation process is affected by link anomalies based on the technical indicators, thereby determining whether the comparison basis is valid. Only when the technical indicators are normal (i.e., the comparison basis is valid) is the algorithmic indicator used to further analyze whether the algorithm computational effect meets expectations. This hierarchical determination logic avoids misjudgments of algorithmic indicators due to engineering link problems, improving the reliability of algorithm quality evaluation results. Simultaneously, the rapid determination of the preceding technical indicators can terminate invalid analysis in advance, improving the efficiency of algorithm quality evaluation.
[0051] S306: Server 20 determines the evaluation result of the algorithm under test based on the first and second indicator values of each evaluation indicator.
[0052] The first indicator value includes the first technical indicator value and the first algorithm indicator value calculated from the first evaluation result set. The second indicator value includes the second technical indicator value and the second algorithm indicator value calculated from the second evaluation result set.
[0053] Optionally, server 20 first determines whether the first and second technical indicator values corresponding to each technical indicator are normal. If the first and second technical indicator values are abnormal, the evaluation result of the algorithm under test is determined to be an algorithm link performance anomaly. If the first and second technical indicator values are normal, the server further determines whether the first and second algorithm indicator values corresponding to each algorithm indicator are normal. If the first and second algorithm indicator values are normal, the evaluation result of the algorithm under test is determined to be a pass. If the first and second algorithm indicator values are abnormal, an anomaly attribution verification is triggered, and the evaluation result is determined to be either an environment configuration anomaly or an algorithm model anomaly based on the result of the anomaly attribution verification.
[0054] Specifically, in the process of determining whether the first and second technical indicator values corresponding to each technical indicator are normal, the server 20 first determines whether the first and second technical indicator values corresponding to each technical indicator are both within the preset technical indicator range; if any first or second technical indicator value corresponding to any technical indicator is not within the preset technical indicator range, the evaluation result of the algorithm under test is determined to be abnormal link performance; if the first and second technical indicators corresponding to each technical indicator are both within the preset technical indicator range, the difference between the first and second technical indicator values corresponding to each technical indicator is calculated respectively to obtain at least one technical indicator difference value; if any technical indicator difference value is greater than the corresponding technical indicator difference threshold, the evaluation result of the algorithm under test is determined to be abnormal algorithm link performance; if the difference values of each technical indicator are less than or equal to the corresponding technical indicator difference threshold, the evaluation result of the algorithm under test is determined based on the first and second algorithm indicator values corresponding to each algorithm indicator.
[0055] In determining the evaluation result of the algorithm under test based on the first and second algorithm indicator values corresponding to each algorithm indicator, server 20 first performs indicator deviation analysis on each algorithm indicator based on the first and second algorithm indicator values corresponding to each algorithm indicator. For example, it identifies whether there are outliers in the first and second algorithm indicator values corresponding to each algorithm indicator for outlier analysis, determines whether there are abnormal distributions in the data distribution of the algorithm indicator in the first and second algorithm indicator values for data distribution analysis, and calculates the difference between the first and second algorithm indicator values for difference analysis. If the above indicator deviation analysis triggers anomaly attribution verification due to the detection of anomalies, the evaluation result is determined to be either an environment configuration anomaly or an algorithm model anomaly based on the result of the anomaly attribution verification. If no anomalies are detected in the above indicator deviation analysis, the evaluation result of the algorithm under test is determined to be a pass.
[0056] In this embodiment, the server 20 determines the evaluation result of the algorithm under test based on the first and second index values of each evaluation index, thereby being able to locate potential quality problems of the algorithm under test based on index differences and effectively predict the performance of the algorithm under test after it is deployed to the production environment.
[0057] S307: Server 20 sends the evaluation results to terminal 10.
[0058] The evaluation result is the final judgment conclusion determined by the server 20 based on the comparative analysis of the first indicator value and the second indicator value. It may include status information such as evaluation passed, abnormal link performance, and abnormal algorithm model, as well as auxiliary information such as detailed indicator comparison data and abnormal attribution verification data.
[0059] S308: Terminal 10 receives the evaluation results sent by server 20.
[0060] Optionally, terminal 10 receives the evaluation results returned by server 20 through a network interface.
[0061] S309: Terminal 10 displays results based on the evaluation results.
[0062] Optionally, the terminal 10 will display the evaluation results to the user in a visual form, such as displaying a comparison chart of the algorithm under test and the benchmark algorithm on various evaluation indicators, marking abnormal indicators, highlighting detected outliers or distribution differences, and presenting the final evaluation conclusion, such as evaluation passed, link performance abnormal, algorithm model abnormal, etc., so that the user can intuitively understand the quality status and potential problems of the algorithm under test.
[0063] In the aforementioned algorithm quality evaluation method, the server responds to the evaluation request and obtains an evaluation sample set from the real production environment. Based on the evaluation sample set, the algorithm under test and the benchmark algorithm are evaluated separately. The evaluation result is determined by comparing and analyzing the evaluation index values corresponding to the two algorithms. This method can achieve efficient and comprehensive evaluation of algorithm iteration. Potential problems can be discovered without waiting for the algorithm to go online and through real feedback. This reduces the debugging and correction costs of algorithm iteration and improves the efficiency and comprehensiveness of algorithm evaluation.
[0064] In one embodiment, such as Figure 4 As shown, another method for evaluating algorithm quality is provided, which is then applied to... Figure 2 The following steps are illustrated using server 20 as an example: S401: Retrieve multiple request logs from the production environment.
[0065] The production environment refers to the runtime environment in which the software system provides services and processes real data based on actual requests after it has been officially launched. In this environment, the application processes real requests and generates corresponding request logs.
[0066] Optionally, server 20 retrieves multiple request logs from the production environment at preset time intervals.
[0067] S402: Filter multiple request logs according to log filtering conditions to obtain multiple target request logs.
[0068] The log filtering conditions include at least one of the following: a first environment identifier for indicating the environment type of the log source, a first application identifier for indicating the application service of the log source, and a first scene identifier for indicating the interaction scenario of the log source.
[0069] S403: Modify the request parameters in the request logs of each target to generate each evaluation sample.
[0070] Optionally, server 20 obtains the original request parameters from any target request log, denoted as... ,in This represents various characteristic parameters in the request log of any given target, such as application identifier, scenario identifier, environment identifier, request attributes, request category, request location, and request timestamp. Then, server 20, based on the objectives of this evaluation task, customizes and modifies the original request parameters in the request log of any given target, generating a modified evaluation sample. The transformation process can be represented as follows: ,in To modify the function; Identify the modified feature parameters; if a certain feature parameter No modifications are required under the objectives of this evaluation. That is, the feature parameter remains unchanged; This refers to the control variables introduced as needed, which are used to ensure that the evaluation results are not affected by extraneous factors in order to achieve the control variable method. These include, but are not limited to, experimental parameters, diversion parameters, environmental labels, etc.
[0071] S404: Store each evaluation sample in the evaluation sample database.
[0072] Optionally, server 20 stores each evaluation sample in an evaluation sample database according to a preset data table structure. For example, each evaluation sample is treated as a record in the database, using the unique identifier (traceid) of the request log as the primary key, and associated with and storing core fields such as application identifier, scenario identifier, environment identifier, request attributes, request category, request location, request timestamp, experiment identifier, and extended parameters. Through this structured database storage method, the server can efficiently retrieve evaluation results based on conditions such as application identifier, scenario identifier, and environment identifier, providing support for subsequent evaluation sample set selection.
[0073] It is worth noting that, in order to maintain the novelty and real-time nature of the sample data, server 20 can be set to schedule tasks to update and iterate the evaluation sample database at preset time intervals. Specifically, after server 20 obtains multiple request logs from the production environment at preset time intervals, it generates new evaluation samples according to the sample construction process in S403, adds the new evaluation samples to the sample database, and deletes historical samples that have exceeded the preset retention period. Thus, through automated data update and iteration, the evaluation effect is ensured while reducing the cost of manually maintaining the sample database.
[0074] S405: Filter the evaluation samples in the evaluation sample database according to the sample filtering conditions to obtain the evaluation sample set.
[0075] The sample filtering conditions include at least one of the following: a second environment identifier for indicating the environment type of the sample source, a second application identifier for indicating the application service of the sample source, and a second scene identifier for indicating the interaction scene of the sample source.
[0076] Optionally, server 20 selects a subset of evaluation samples that meet the criteria from the evaluation sample database based on the environment, application, and scenario involved in the algorithm under test, and uses this subset as the input sample set for this evaluation. Understandably, the above sample filtering criteria can be a subset of the above log filtering criteria. For example, samples from multiple environments, applications, and scenarios may be collected during the sample construction phase, while only samples from a specific environment, application, and scenario need to be used for evaluation during the evaluation execution phase. By separating sample construction from sample selection, multiple evaluations for different purposes and scopes can be achieved based on a single sample construction, avoiding redundant sample construction.
[0077] In this embodiment, server 20 obtains multiple target request logs by acquiring real request logs from the production environment and filtering them based on log filtering conditions. The original request parameters in each target request log are then customized to generate evaluation samples. This not only ensures that the evaluation samples cover diverse online request scenarios, improving the efficiency and comprehensiveness of sample construction, but also eliminates interference from redundant factors in the evaluation results through parameter modification and the introduction of control variables, thus improving the reliability of sample construction. Furthermore, server 20 stores the samples in a structured database to achieve multiple evaluations with different scopes and purposes, and sets up scheduled tasks for automatic updates and iterations, effectively ensuring the novelty and real-time nature of the samples while reducing manual maintenance costs.
[0078] S406: Based on the evaluation sample set and preset control parameters, construct and generate an evaluation request set.
[0079] The control parameters include at least one of the following: environment parameters for specifying the target runtime environment of the evaluation request, experimental parameters for specifying the experimental configuration strategy, and distribution parameters for specifying the experimental bucket or branch to which the evaluation request belongs. The evaluation request set includes multiple evaluation requests, each corresponding to one evaluation sample and carrying corresponding control parameters.
[0080] Optionally, server 20 obtains the feature parameters corresponding to each evaluation sample in the evaluation sample set; then loads the corresponding parameter template according to the algorithm's request interface specification, and splices the preset control parameters for each evaluation sample according to the corresponding parameter template to ensure that the sample runs only in the specified environment and specified experimental bucket during the evaluation execution, avoiding other factors from interfering with the observability of the evaluation results; finally, each evaluation sample is encapsulated into an executable evaluation request, and all evaluation requests together constitute the evaluation request set.
[0081] S407: Input the evaluation request set into the algorithm under test and the benchmark algorithm respectively for evaluation, and obtain the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm.
[0082] Optionally, server 20 can perform dual-run tests in the same or different environments.
[0083] In one possible implementation, the algorithm under test is not deployed to the production environment. In this case, the evaluation execution environment depends on the comparability of the pre-production and production environments. Specifically, if the pre-production and production environments meet preset comparability conditions, the server 20 inputs the evaluation request set to the algorithm under test in the pre-production environment and the benchmark algorithm in the production environment for evaluation, respectively, to obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm. If the pre-production and production environments do not meet preset comparability conditions, the server 20 inputs the evaluation request set to the algorithm under test and the benchmark algorithm under different pre-production branches in the pre-production environment for evaluation, respectively, to obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm. The preset comparability conditions include that the difference in key environment configurations is within a preset difference range. Key environment configurations include at least one of the following: algorithm running configuration, algorithm experiment configuration, and algorithm model configuration. The preset difference range can be determined through pre-experiments or historical data.
[0084] Through this implementation, server 20 can flexibly choose evaluation strategies such as direct cross-environment comparison or comparison within the same environment branch, based on the actual configuration differences between the pre-production and production environments. When the environment configurations are comparable, the actual effect of the algorithm under test after deployment can be predicted most directly; when the environment configurations differ significantly, the effect of the algorithm under test relative to the benchmark algorithm can still be effectively evaluated through comparison within the same environment branch, avoiding the invalidation of evaluation results due to incomparable environments.
[0085] In another possible implementation, the algorithm under test has been deployed to the production environment. In this case, server 20 inputs the evaluation request set into the algorithm under test and the benchmark algorithm under different experimental buckets in the production environment for evaluation, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. Here, an experimental bucket refers to a traffic group that carries a specific experimental configuration in the production environment. Different experimental buckets share the same operating environment, and only the test variables (such as algorithm version and experimental parameters) differ.
[0086] Specifically, server 20 adds experiment identifiers and traffic distribution identifiers to the evaluation request set to ensure that evaluation requests fall into the designated experiment buckets. For example, the evaluation request set is input into the algorithm under test in experiment bucket A and the benchmark algorithm in experiment bucket B, and the output results of the two experiment buckets are obtained as the first evaluation result set and the second evaluation result set. At the same time, the server adds evaluation labels to the evaluation traffic through a traffic control mechanism to distinguish the evaluation traffic from the real online traffic and prevent the evaluation traffic from affecting online data statistics.
[0087] Through this implementation method, after the algorithm under test has been released to the production environment, the server 20 can directly observe the difference in the performance indicators of the algorithm under test and the benchmark algorithm in the actual production environment by comparing and evaluating different experimental buckets online, thus providing data support for the real-time adjustment of the online experimental configuration.
[0088] S408: Calculate the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set.
[0089] The evaluation metrics include technical metrics and algorithmic metrics. Technical metrics characterize the performance of the algorithm chain, while algorithmic metrics characterize the effectiveness of the algorithm computation. The technical metrics and algorithmic metrics have a hierarchical dependency in determining the evaluation results.
[0090] Specifically, the technical metrics include at least one of the following: request success rate and request timeout rate; the algorithm metrics include at least one of the following: output value distribution, decision threshold distribution, output type distribution, and default decision percentage. Among these, the request success rate represents the proportion of evaluation requests that are successfully processed and return results; the request timeout rate represents the proportion of evaluation requests that do not receive a response within a preset time threshold; the output value distribution represents the distribution of numerical parameters output by the algorithm in different intervals; the decision threshold distribution represents the distribution of threshold parameters used by the algorithm for decision-making in different intervals; the output type distribution represents the proportion of different types of labels output by the algorithm; and the default decision percentage represents the proportion of evaluation requests that fall back to the default strategy out of all evaluation requests.
[0091] Optionally, server 20 calculates the first and second evaluation result sets respectively using pre-written indicator scripts. For technical indicators, server 20 calculates the request success rate by statistically analyzing the proportion of evaluation requests that were successfully processed and returned results based on the two result sets, and calculates the request timeout rate by statistically analyzing the proportion of requests that did not receive a response within a preset time threshold. For algorithm indicators, server 20 calculates the output value distribution by statistically analyzing the proportion of evaluation requests in each interval based on the two result sets; it also calculates the judgment threshold distribution by statistically analyzing the proportion of different types of tags based on the preset intervals; and it calculates the default judgment proportion by statistically analyzing the proportion of evaluation requests that revert to the default strategy among all evaluation requests.
[0092] In this embodiment, the server can quantitatively evaluate the performance of the algorithm chain by calculating technical indicators such as request success rate and request timeout rate. This allows for accurate identification of any anomalies at the algorithm chain level, ensuring that the comparative analysis of algorithm indicators is based on reliable evaluation. Furthermore, by calculating algorithm indicators such as output value distribution, decision threshold distribution, output type distribution, and default decision ratio, the server can perform fine-grained quantification of the algorithm's operational effect from multiple dimensions, including the shape of the output value distribution, decision threshold offset, output type composition, and the degree of dependence on the default strategy. This provides comprehensive and quantifiable data support for determining the evaluation results, effectively ensuring the accuracy and comprehensiveness of the algorithm quality evaluation.
[0093] S409: Determine the evaluation result of the algorithm under test based on the first and second index values of each evaluation index.
[0094] Understandably, the first indicator value includes the first technical indicator value (such as the first request success rate, the first request timeout rate, etc.) and the first algorithm indicator value (such as the first output value distribution, the first decision threshold distribution, the first output type distribution, the first default decision percentage, etc.) calculated from the first evaluation result set. The second indicator value includes the second technical indicator value (such as the second request success rate, the second request timeout rate, etc.) and the second algorithm indicator value (such as the second output value distribution, the second decision threshold distribution, the second output type distribution, the second default decision percentage, etc.) calculated from the second evaluation result set.
[0095] Optionally, the server first checks whether the values of each technical indicator are normal. If the values of each technical indicator are normal, then it analyzes whether the values of each algorithm indicator are normal. The aforementioned technical indicators are numerical indicators, and the aforementioned algorithm indicators include numerical indicators and / or distributed indicators.
[0096] In one possible implementation, the server determines the evaluation result of the algorithm under test based on the first and second indicator values of each evaluation metric, including: determining whether the first and second technical indicator values corresponding to each technical indicator are both within a preset technical indicator range; if any first or second technical indicator value corresponding to any technical indicator is not within the preset technical indicator range, then the evaluation result of the algorithm under test is determined to be abnormal link performance; if the first and second technical indicators corresponding to each technical indicator are both within the preset technical indicator range, then the difference between the first and second technical indicator values corresponding to each technical indicator is calculated to obtain at least one technical indicator difference value; if any technical indicator difference value is greater than the corresponding technical indicator difference threshold, then the evaluation result of the algorithm under test is determined to be abnormal algorithm link performance; if all technical indicator difference values are less than or equal to the corresponding technical indicator difference thresholds, then the evaluation result of the algorithm under test is further determined based on the first and second algorithm indicator values corresponding to each algorithm indicator.
[0097] Understandably, technical metrics are numerical metrics. The range of technical metrics is the data range used to determine whether a single technical metric value is normal. For example, the normal range for request success rate can be set to [99%, 100%], and the normal range for request timeout rate can be set to [0%, 1%]. The technical metric difference threshold is the upper limit used to determine whether the difference between technical metrics is significant. For example, the difference threshold for request success rate can be set to 0.1%.
[0098] Through this implementation method, the server 20 first checks whether the technical indicator values corresponding to each technical indicator are normal. If the technical indicator values corresponding to each technical indicator are normal, it then analyzes whether the algorithm indicator values corresponding to each algorithm indicator are normal. This can improve the reliability of the evaluation results through hierarchical judgment and multi-dimensional quantitative evaluation, thereby improving the evaluation efficiency while ensuring the comprehensiveness of the evaluation.
[0099] In one possible implementation, the server determines the evaluation result of the algorithm under test based on the first algorithm indicator value and the second algorithm indicator value corresponding to each algorithm indicator, including: performing indicator deviation analysis on each algorithm indicator based on the first algorithm indicator value and the second algorithm indicator value corresponding to each algorithm indicator; the indicator deviation analysis includes at least one of the following: outlier analysis, data distribution analysis, and difference analysis; and determining the evaluation result of the algorithm under test based on the result of the indicator deviation analysis.
[0100] Outlier analysis includes identifying whether there are outliers in the first and second algorithm indicator values corresponding to each algorithm indicator, and triggering anomaly attribution verification if outliers are found. Outliers refer to indicator values that significantly deviate from the reasonable range or the algorithm's expected value. For example, if the normal range for a numerical algorithm indicator is [5, 10], and the numerical algorithm indicator value of a sample is 3, then that algorithm indicator value is an outlier.
[0101] Data distribution analysis includes: determining whether there are any abnormal distributions in the data distribution of the algorithm indicators between the first and second algorithm indicator values, and triggering anomaly attribution verification if abnormal distributions are found. These abnormal distributions include distributions outside a preset reasonable range within the interval distribution of algorithm indicator values (e.g., samples appearing in a high-value range that should not exist), or distribution differences between the two environments exceeding a preset distribution difference threshold.
[0102] The discrepancy analysis includes calculating the difference between the first algorithm's metric value and the second algorithm's metric value, and triggering anomaly attribution verification if the difference exceeds a preset difference threshold. Anomaly attribution verification includes at least one of the following: environment configuration verification, request parameter verification, experiment hit verification, link configuration verification, and model parameter verification. Specifically, environment configuration verification checks whether configuration differences between the pre-production and production environments lead to differences in evaluation results; for example, it checks whether environment variables, dependent service versions, etc., are consistent. Request parameter verification checks whether the request parameters of the evaluation samples that trigger anomaly attribution verification are correctly passed and whether they conform to the input parameter specifications. Experiment hit verification checks whether the evaluation samples correctly hit the specified experiment bucket or experiment strategy. Link configuration verification checks whether there are any anomalies in the configuration items (such as feature passing, model calling, configuration loading, etc.) in the algorithm engineering link. Model parameter verification checks whether the parameter configuration of the algorithm model (such as model version, model weights, hyperparameter settings, etc.) meets expectations.
[0103] Understandably, if the metric deviation analysis triggers anomaly attribution verification, server 20 determines the evaluation result based on the verification result: if the attribution verification indicates that the difference is caused by environmental or configuration factors, server 20 determines the evaluation result as an environmental configuration anomaly; if the attribution verification excludes environmental and configuration factors, server 20 determines the evaluation result as an algorithm model anomaly. If the metric deviation analysis does not trigger anomaly attribution verification, server 20 determines the evaluation result of the algorithm under test as passed.
[0104] For example, suppose the iterative goal of algorithm A under test is to optimize the quality of its output. Theoretically, its corresponding algorithm metric M (such as the distribution of output values) should exhibit the expected distribution trend. If server 20 finds that the actual performance of algorithm metric M does not exhibit the expected distribution trend during metric deviation analysis, but instead concentrates in an unexpected direction, it triggers anomaly attribution verification. For example, it sequentially performs environment configuration verification, request parameter verification, experiment hit verification, link configuration verification, and model parameter verification. If server 20 detects after anomaly attribution verification that the environment parameter configuration of the pre-production environment where algorithm A is located is inconsistent with the actual plan, it determines that the evaluation result is an environment configuration anomaly, and the user needs to correct the environment parameter configuration and re-execute the evaluation. If server 20 detects after model parameter verification that the model parameters of algorithm A under test in the pre-production environment are not adjusted as expected, causing the model to fail to load the optimized parameter combination as expected, it determines that the evaluation result is an algorithm model anomaly, and the user needs to correct the model parameter configuration and re-execute the evaluation. If no abnormalities are found in server 20 after all the above attribution checks, the evaluation result is determined to be an algorithm model abnormality, that is, the algorithm under test A itself failed to achieve the expected iteration goal. The user needs to further debug and optimize the algorithm under test A and then re-execute the evaluation.
[0105] Through this implementation method, server 20 performs indicator deviation analysis on each algorithm indicator from different dimensions such as outlier analysis, data distribution analysis, and difference analysis. When the indicator deviation analysis triggers anomaly attribution verification, it systematically investigates whether the anomaly of the algorithm indicator value is caused by environmental configuration problems, request parameter problems, experiment hit problems, link configuration problems, or problems of the model itself, which effectively improves the efficiency and accuracy of algorithm evaluation and anomaly location.
[0106] The aforementioned algorithm quality evaluation method, on the one hand, in the sample preparation stage, filters real request logs from the production environment based on log filtering conditions to obtain multiple target request logs. The original request parameters in each target request log are then customized to generate evaluation samples, improving the efficiency, comprehensiveness, and reliability of sample construction. On the other hand, in the evaluation execution stage, based on the evaluation sample set and preset control parameters, an evaluation request set is constructed. This set is then input into the algorithm under test and the benchmark algorithm for evaluation, obtaining a first evaluation result set for the algorithm under test and a second evaluation result set for the benchmark algorithm. This eliminates evaluation bias caused by different sample sets, ensuring the comparability of evaluation results and facilitating subsequent indicator calculation and comparison. The analysis provides a reliable data foundation. Furthermore, in the indicator calculation stage, the first indicator value for each evaluation indicator corresponding to the first evaluation result set and the second indicator value for each evaluation indicator corresponding to the second evaluation result set are calculated respectively. The evaluation indicators include technical indicators and algorithm indicators. The technical indicators characterize the performance of the algorithm chain, and the algorithm indicators characterize the effect of the algorithm operation. This allows for a quantitative evaluation of the algorithm's output results from two dimensions: algorithm chain performance and algorithm operation effect, avoiding the limitations of single-dimensional evaluation. Additionally, in the result analysis stage, the evaluation results of the algorithm under test are determined based on the first and second indicator values of each evaluation indicator. This allows for improved reliability of the result analysis through hierarchical judgment and multi-dimensional quantitative evaluation.
[0107] By adopting the above-mentioned algorithm quality evaluation method, constructing evaluation samples based on real request logs in the production environment, and using big data evaluation to predict trends and analyze quality of algorithm iterations, it is possible to achieve efficient and comprehensive evaluation of algorithm iterations. Potential problems can be discovered through real feedback without waiting for the algorithm to go live, reducing the debugging and correction costs of algorithm iterations, improving the efficiency and comprehensiveness of evaluation, and having a positive significance for quality assurance in the algorithmic era.
[0108] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0109] The inventive concept based on the above-mentioned algorithm quality evaluation method, such as... Figure 5 As shown in the illustration, this application also provides a first evaluation device 500 for implementing the above-described algorithm quality evaluation method applied to a server. The first evaluation device 500 includes: The acquisition module 501 is used to acquire the evaluation sample set; the evaluation sample set includes multiple evaluation samples, each of which is constructed based on the request logs in the production environment; The execution module 502 is used to construct and generate an evaluation request set based on the evaluation sample set and preset control parameters; the control parameters include at least one of environmental parameters, experimental parameters and shunting parameters; the evaluation request set includes multiple evaluation requests; the evaluation request set is input to the algorithm under test and the benchmark algorithm respectively for evaluation, and a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm are obtained; The calculation module 503 is used to calculate the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set. The determination module 504 is used to determine the evaluation result of the algorithm under test based on the first indicator value and the second indicator value of each evaluation indicator. The evaluation metrics include technical metrics and algorithm metrics. Technical metrics are used to characterize the performance of the algorithm chain, while algorithm metrics are used to characterize the effect of the algorithm operation. Technical metrics and algorithm metrics have a hierarchical dependency relationship when determining the evaluation results.
[0110] In one embodiment, the first evaluation device 500 further includes a sample construction module, configured to acquire multiple request logs in the production environment; filter the multiple request logs according to log filtering conditions to obtain multiple target request logs; the log filtering conditions include at least one of a first environment identifier, a first application identifier, and a first scenario identifier; modify the request parameters in each target request log to generate each evaluation sample; and store each evaluation sample in an evaluation sample database; the acquisition module 501 is specifically configured to: filter the evaluation samples in the evaluation sample database according to sample filtering conditions to obtain an evaluation sample set; the sample filtering conditions include at least one of a second environment identifier, a second application identifier, and a second scenario identifier.
[0111] In one embodiment, the algorithm under test is not released to the production environment; the execution module 502 is specifically used to: when the pre-release environment and the production environment meet preset comparability conditions, input the evaluation request set to the algorithm under test in the pre-release environment and the benchmark algorithm in the production environment for evaluation, and obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm; when the pre-release environment and the production environment do not meet preset comparability conditions, input the evaluation request set to the algorithm under test and the benchmark algorithm under different pre-release branches in the pre-release environment for evaluation, and obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm; wherein, the preset comparability conditions include the difference value of the key environment configuration being within a preset difference range; the key environment configuration includes at least one of the following: algorithm running configuration, algorithm experiment configuration, and algorithm model configuration.
[0112] In one embodiment, the algorithm under test has been deployed to the production environment; the execution module 502 is specifically used to: input the evaluation request set to the algorithm under test and the benchmark algorithm respectively for evaluation, and obtain the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm, including: inputting the evaluation request set to the algorithm under test and the benchmark algorithm under different experimental buckets in the production environment for evaluation, and obtaining the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm.
[0113] In one embodiment, the technical metrics include at least one of the following: request success rate and request timeout rate; the algorithm metrics include at least one of the following: output value distribution, judgment threshold distribution, output type distribution, and default judgment percentage; wherein, the request success rate is used to characterize the proportion of evaluation requests that are successfully processed and return results; the request timeout rate is used to characterize the proportion of evaluation requests that do not receive a response within a preset time threshold; the output value distribution is used to characterize the distribution of numerical parameters output by the algorithm in different intervals; the judgment threshold distribution is used to characterize the distribution of threshold parameters on which the algorithm makes judgments in different intervals; the output type distribution is used to characterize the proportion of different types of labels output by the algorithm; and the default judgment percentage is used to characterize the proportion of evaluation requests whose outputs fall back to the default strategy out of all evaluation requests.
[0114] In one embodiment, the first indicator value includes a first technical indicator value and a first algorithm indicator value, and the second indicator value includes a second technical indicator value and a second algorithm indicator value. The determining module 504 is specifically used to: determine whether the first technical indicator value and the second technical indicator value corresponding to each technical indicator are both within a preset technical indicator range; if the first technical indicator value or the second technical indicator value corresponding to any technical indicator is not within the preset technical indicator range, then the evaluation result of the algorithm under test is determined to be abnormal link performance; if the first technical indicator value and the second technical indicator value corresponding to each technical indicator are both within the preset technical indicator range, then the difference value between the first technical indicator value and the second technical indicator value corresponding to each technical indicator is calculated respectively to obtain at least one technical indicator difference value; if any technical indicator difference value is greater than the corresponding technical indicator difference threshold, then the evaluation result of the algorithm under test is determined to be abnormal algorithm link performance; if the difference values of each technical indicator are less than or equal to the corresponding technical indicator difference threshold, then the evaluation result of the algorithm under test is determined according to the first algorithm indicator value and the second algorithm indicator value corresponding to each algorithm indicator.
[0115] In one embodiment, the determining module 504 is specifically used to: perform indicator deviation analysis on each algorithm indicator based on the first algorithm indicator value and the second algorithm indicator value corresponding to each algorithm indicator; the indicator deviation analysis includes at least one of the following: outlier analysis, data distribution analysis, and difference analysis; and determine the evaluation result of the algorithm to be tested based on the result of the indicator deviation analysis.
[0116] In one embodiment, outlier analysis includes: identifying whether there are outliers in the first and second algorithm indicator values corresponding to each algorithm indicator, and triggering anomaly attribution verification if outliers are found; data distribution analysis includes: determining whether there are abnormal distributions in the data distribution of the algorithm indicators in the first and second algorithm indicator values, and triggering anomaly attribution verification if abnormal distributions are found; difference analysis includes: calculating the difference between the first and second algorithm indicator values, and triggering anomaly attribution verification if the difference exceeds a preset difference threshold; the determination module 504 is specifically used to: determine the evaluation result of the algorithm under test based on the result of the anomaly attribution verification if the result of the indicator deviation analysis triggers anomaly attribution verification; and determine that the evaluation result of the algorithm under test is passed if the result of the indicator deviation analysis does not trigger anomaly attribution verification.
[0117] In one embodiment, anomaly attribution verification includes at least one of environment configuration verification, request parameter verification, experiment hit verification, link configuration verification, and model parameter verification.
[0118] The inventive concept based on the above-mentioned algorithm quality evaluation method, such as... Figure 6As shown in the illustration, this application also provides a second evaluation device 600 for implementing the above-described algorithm quality evaluation method applied to a terminal. The second evaluation device 600 includes: The response module 601 is used to generate an evaluation request in response to the user's evaluation trigger operation; The sending module 602 is used to send an evaluation request to the server, so that the server: responds to the evaluation request and obtains an evaluation sample set; the evaluation sample set includes multiple evaluation samples, each of which is constructed based on request logs in the production environment; constructs and generates an evaluation request set based on the evaluation sample set and preset control parameters; the control parameters include at least one of environmental parameters, experimental parameters, and traffic splitting parameters; the evaluation request set includes multiple evaluation requests; inputs the evaluation request set to the algorithm under test and the benchmark algorithm respectively for evaluation, and obtains a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm; calculates the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set respectively; determines the evaluation result of the algorithm under test based on the first indicator value and the second indicator value of each evaluation indicator; and sends the evaluation result to the terminal. The receiving module 603 is used to receive the evaluation results sent by the server; Display module 604 is used to display the evaluation results.
[0119] Each module in the first evaluation device 500 and the second evaluation device 600 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0120] This application also provides a first electronic device, which may be a server, and its internal structure diagram is as follows: Figure 7 As shown. Please see below. Figure 7The first electronic device 70 includes a first processor 71, a first memory 72, a first input / output interface 73 (I / O), and a first communication interface 74. The first processor 71, first memory 72, and first I / O interface 73 are connected via a first system bus 75, and the first communication interface 74 is connected to the first system bus 75 via the first I / O interface 73. The first processor 71 of the first electronic device 70 provides computing and control capabilities. The first memory 72 of the first electronic device 70 includes a first non-volatile storage medium 76 and a first internal memory 77. The first non-volatile storage medium 76 stores a first operating system 761, a first computer program 762, and a database 763. The first internal memory 77 provides an environment for the operation of the first operating system 761 and the first computer program 762 stored in the first non-volatile storage medium 76. The database 763 of the first electronic device 70 is used to store evaluation sample data, etc. The first I / O interface 73 of the first electronic device 70 is used for exchanging information between the first processor 71 and external devices. The first communication interface 74 of the first electronic device 70 is used to communicate with an external terminal via a network connection. The first processor 71 of the first electronic device 70 executes a first computer program 762 to implement the aforementioned algorithm quality evaluation method applied to the server.
[0121] Furthermore, embodiments of this application also provide a second electronic device, which can be a terminal, and its internal structure diagram can be as follows. Figure 8 As shown. Please see below. Figure 8The second electronic device 80 includes a second processor 81, a second memory 82, a second input / output interface 83, a second communication interface 84, a display unit 85, and an input device 86. The second processor 81, the second memory 82, and the second input / output interface 83 are connected via a second system bus 87. The second communication interface 84, the display unit 85, and the input device 86 are connected to the second system bus 87 via the second input / output interface 83. The second processor 81 of the second electronic device 80 provides computing and control capabilities. The second memory 82 of the second electronic device 80 includes a second non-volatile storage medium 88 and a second internal memory 89. The second non-volatile storage medium 88 stores a second operating system 881 and a second computer program 882. The second internal memory 89 provides an environment for the operation of the second operating system 881 and the second computer program 882 stored in the second non-volatile storage medium 88. The second input / output interface 83 of the second electronic device 80 is used for exchanging information between the second processor 81 and external devices. The second communication interface 84 of the second electronic device 80 is used for wired or wireless communication with an external terminal. Wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. The second computer program 882 is executed by the second processor 81 to implement the aforementioned algorithm quality evaluation method applied to the terminal. The display unit 85 of the second electronic device 80 is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an e-ink display screen. The input device 86 of the second electronic device 80 can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad provided on the casing of the second electronic device 80, or an external keyboard, touchpad, or mouse, etc.
[0122] This application also provides a computer storage medium storing instructions that, when run on a computer or processor, cause the computer or processor to perform one or more steps in the above embodiments. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in the above-described computer-readable storage medium.
[0123] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted through the computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.
[0125] The embodiments described above are merely preferred embodiments of this application and are not intended to limit the scope of this application. Any modifications and improvements made by those skilled in the art to the technical solutions of this application without departing from the spirit of this application should fall within the protection scope defined by the claims.
[0126] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the features, data, and information involved in this specification were all obtained under full authorization.
[0127] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A method for evaluating algorithm quality, characterized in that, Applied to a server, the method includes: Obtain an evaluation sample set; the evaluation sample set includes multiple evaluation samples, each of which is constructed based on request logs in the production environment; Based on the evaluation sample set and preset control parameters, an evaluation request set is constructed and generated; the control parameters include at least one of environmental parameters, experimental parameters, and diversion parameters; the evaluation request set includes multiple evaluation requests; The evaluation request set is input into the algorithm under test and the benchmark algorithm respectively for evaluation, to obtain the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm; Calculate the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set respectively; The evaluation result of the algorithm under test is determined based on the first indicator value and the second indicator value of each evaluation indicator. The evaluation metrics include technical metrics and algorithm metrics. The technical metrics are used to characterize the performance of the algorithm chain, and the algorithm metrics are used to characterize the effect of the algorithm operation. The technical metrics and the algorithm metrics have a hierarchical dependency relationship when determining the evaluation result.
2. The method as described in claim 1, characterized in that, Before obtaining the evaluation sample set, the method further includes: Retrieve multiple request logs from the production environment; The multiple request logs are filtered according to the log filtering conditions to obtain multiple target request logs; the log filtering conditions include at least one of a first environment identifier, a first application identifier, and a first scenario identifier; Modify the request parameters in the request logs of each target to generate each evaluation sample; Each evaluation sample is stored in the evaluation sample database; The acquisition of the evaluation sample set includes: The evaluation samples in the evaluation sample database are filtered according to the sample filtering conditions to obtain the evaluation sample set; the sample filtering conditions include at least one of the second environment identifier, the second application identifier, and the second scenario identifier.
3. The method as described in claim 1, characterized in that, The algorithm under test was not deployed to the production environment. The step of inputting the evaluation request set into the algorithm under test and the benchmark algorithm respectively for evaluation, and obtaining the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm, includes: If the pre-release environment and the production environment meet the preset comparability conditions, the evaluation request set is input to the algorithm under test in the pre-release environment and the benchmark algorithm in the production environment for evaluation, respectively, to obtain the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm. If the pre-release environment and the production environment do not meet the preset comparability conditions, the evaluation request set is input into the algorithm under test and the benchmark algorithm under different pre-release branches in the pre-release environment for evaluation, and the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm are obtained. The preset comparability conditions include that the difference in key environment configurations is within a preset difference range; the key environment configurations include at least one of the following: algorithm running configuration, algorithm experiment configuration, and algorithm model configuration.
4. The method as described in claim 1, characterized in that, The algorithm under test has been deployed to the production environment. The step of inputting the evaluation request set into the algorithm under test and the benchmark algorithm respectively for evaluation, and obtaining the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm, includes: The evaluation request set is input into the algorithm under test and the benchmark algorithm in different experimental buckets in the production environment for evaluation, so as to obtain the first evaluation result set corresponding to the algorithm under test and the second evaluation result set corresponding to the benchmark algorithm.
5. The method as described in claim 1, characterized in that, The first indicator value includes a first technical indicator value and a first algorithm indicator value; the second indicator value includes a second technical indicator value and a second algorithm indicator value. The step of determining the evaluation result of the algorithm under test based on the first indicator value and the second indicator value of each evaluation index includes: Determine whether the first technical indicator value and the second technical indicator value corresponding to each of the technical indicators are both within the preset technical indicator range; If any of the technical indicators corresponds to a first technical indicator value or a second technical indicator value that is not within the preset technical indicator range, then the evaluation result of the algorithm under test is determined to be an abnormal link performance. If both the first technical indicator and the second technical indicator corresponding to each of the technical indicators are within the preset technical indicator range, then the difference between the first technical indicator value and the second technical indicator value corresponding to each of the technical indicators is calculated to obtain at least one technical indicator difference value. If any of the technical indicators has a difference value greater than the corresponding technical indicator difference threshold, then the evaluation result of the algorithm under test is determined to be an abnormal algorithm link performance. If the difference values of each technical indicator are all less than or equal to the corresponding technical indicator difference threshold, then indicator deviation analysis is performed on each algorithm indicator based on the first algorithm indicator value and the second algorithm indicator value corresponding to each algorithm indicator; the indicator deviation analysis includes at least one of the following: outlier analysis, data distribution analysis, and difference analysis; the evaluation result of the algorithm to be tested is determined based on the result of the indicator deviation analysis.
6. The method as described in claim 5, characterized in that, The outlier analysis includes: identifying whether there are outliers in the first algorithm indicator value and the second algorithm indicator value corresponding to each of the algorithm indicators, and triggering anomaly attribution verification if outliers are found; the data distribution analysis includes: determining whether there is an abnormal distribution in the data distribution of the algorithm indicator in the first algorithm indicator value and the second algorithm indicator value, and triggering anomaly attribution verification if an abnormal distribution is found; the difference analysis includes: calculating the difference between the first algorithm indicator value and the second algorithm indicator value, and triggering anomaly attribution verification if the difference exceeds a preset difference threshold; The step of determining the evaluation result of the algorithm under test based on the results of the index deviation analysis includes: If the result of the deviation analysis of the indicator triggers anomaly attribution verification, the evaluation result of the algorithm under test is determined based on the result of the anomaly attribution verification. If the deviation analysis of the indicators does not trigger anomaly attribution verification, the evaluation result of the algorithm under test is determined to be passed. The anomaly attribution verification includes at least one of the following: environment configuration verification, request parameter verification, experiment hit verification, link configuration verification, and model parameter verification.
7. A method for evaluating algorithm quality, characterized in that, Applied to a terminal, the method includes: In response to a user's evaluation trigger action, an evaluation request is generated; The evaluation request is sent to the server, causing the server to: respond to the evaluation request by obtaining an evaluation sample set; the evaluation sample set includes multiple evaluation samples, each of which is constructed based on request logs in the production environment; construct and generate an evaluation request set based on the evaluation sample set and preset control parameters; the control parameters include at least one of environmental parameters, experimental parameters, and traffic splitting parameters; the evaluation request set includes multiple evaluation requests; input the evaluation request set to the algorithm under test and the benchmark algorithm respectively for evaluation, obtaining a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm; calculate the first indicator value of each evaluation indicator corresponding to the first evaluation result set and the second indicator value of each evaluation indicator corresponding to the second evaluation result set; determine the evaluation result of the algorithm under test based on the first indicator value and the second indicator value of each evaluation indicator; and send the evaluation result to the terminal. Receive the evaluation results sent by the server; The results are displayed based on the aforementioned evaluation. The evaluation metrics include technical metrics and algorithm metrics. The technical metrics are used to characterize the performance of the algorithm chain, and the algorithm metrics are used to characterize the effect of the algorithm operation. The technical metrics and the algorithm metrics have a hierarchical dependency relationship when determining the evaluation result.
8. A system for evaluating algorithm quality, characterized in that, The algorithm quality evaluation system includes: a terminal and a server; wherein, The terminal is used to respond to the user's evaluation trigger operation, generate an evaluation request, and send the evaluation request to the server; The server is configured to respond to the evaluation request by obtaining an evaluation sample set; the evaluation sample set includes multiple evaluation samples, each of which is constructed based on request logs in the production environment; an evaluation request set is generated based on the evaluation sample set and preset control parameters; the control parameters include at least one of environmental parameters, experimental parameters, and traffic splitting parameters; the evaluation request set includes multiple evaluation requests; the evaluation request set is input to the algorithm under test and the benchmark algorithm respectively for evaluation, to obtain a first evaluation result set corresponding to the algorithm under test and a second evaluation result set corresponding to the benchmark algorithm; a first indicator value for each evaluation metric corresponding to the first evaluation result set and a second indicator value for each evaluation metric corresponding to the second evaluation result set are calculated respectively; the evaluation result of the algorithm under test is determined based on the first indicator value and the second indicator value of each evaluation metric; and the evaluation result is sent to the terminal. The terminal is also used to receive the evaluation results sent by the server and display them based on the evaluation results; The evaluation metrics include technical metrics and algorithm metrics. The technical metrics are used to characterize the performance of the algorithm chain, and the algorithm metrics are used to characterize the effect of the algorithm operation. The technical metrics and the algorithm metrics have a hierarchical dependency relationship when determining the evaluation result.
9. An electronic device, characterized in that, include: A processor and a memory; the memory stores a computer program, and the processor executes the computer program to implement the method steps of any one of claims 1-7.
10. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the method steps as claimed in any one of claims 1-7.