Model Ability Testing Method, Device, Electronic Device, Storage Medium and Product

By receiving update requests and updating the traffic configuration of the inference graph, using the service set to test the model's capability index, the cumbersome and time-consuming problem of model generalization ability testing in the prior art is solved, and a fast and objective model performance evaluation is achieved.

CN116302874BActive Publication Date: 2025-07-11INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310020722.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2025-07-11
Estimated Expiration
2043-01-06

AI Technical Summary

Technical Problem

In the prior art, the model generalization ability test operation is cumbersome, time-consuming and the test results are not objective, making it difficult to quickly and accurately evaluate the generalization ability of multiple versions of models.

Method used

By receiving update requests, reading parameter configuration information to determine the request type, and updating the traffic configuration of the inference graph according to the request type, using the updated inference graph to intercept and redistribute the user traffic of the to-test model, and using the service set to test the model's capability index, including recall and accuracy.

Benefits of technology

It realizes rapid and objective evaluation of the generalization capabilities and other performance of each model to be tested, simplifying the operation process, and improving testing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116302874B_ABST
    Figure CN116302874B_ABST
Patent Text Reader

Abstract

The model ability testing method, device, electronic device, storage medium and product provided by the present invention belong to the technical field of deep learning, and include: receiving an update request; in response to the update request, updating the traffic configuration of the inference graph; based on the updated inference graph, allocating the user traffic of each model to be tested to a service set, so that the service set can use each user traffic to determine the ability index of each model to be tested. The model ability testing method, device, electronic device, storage medium and product provided by the present invention maintain and update the traffic configuration of the inference graph in the shunt system update module, intercept and redistribute the user traffic of the model to be tested by using the updated inference graph, and then use the service set to test the ability index of each model to be tested. The operation is simple, and the performance such as the generalization ability of each model to be tested can be evaluated quickly and objectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular to a method, apparatus, electronic device, storage medium and product for testing model capabilities. Background Art

[0002] How to compare the generalization capabilities of neural network models with a large amount of data and then select a suitable model is very difficult. Even if a model is selected, after modifying and training multiple versions of the model, one still has to face the problem of how to compare the generalization capabilities of multiple versions of the model.

[0003] Currently, a dataset composed of a large amount of data is mainly used to test each model, and then the generalization ability ranking of each model is obtained.

[0004] However, the above method is cumbersome to operate, time-consuming, and the test results are not objective. Summary of the Invention

[0005] The model capability testing method, apparatus, electronic device, storage medium and product provided by the present invention are used to solve the defects of cumbersome operation, long time consumption and non-objective test results in the prior art, and realize simple operation, and can quickly and objectively evaluate the performance such as the generalization ability of each model to be tested.

[0006] The present invention provides a model capability testing method, including:

[0007] Receiving an update request;

[0008] In response to the update request, reading the parameter configuration information of the update request to determine the request type of the update request; the request type includes: traffic allocation, service full volume, and traffic rollback;

[0009] When it is determined that the request type is traffic allocation, determining a traffic allocation plan according to the parameter configuration information;

[0010] When it is determined that the traffic allocation plan is equal distribution, reading the total traffic of the main line bucket, the total traffic of the test bucket, and the number of test buckets from the parameter configuration information;

[0011] Determining the total traffic of the main line bucket as the target traffic of the main line bucket, and determining the target traffic of each test bucket according to the total traffic of the test bucket and the number of test buckets;

[0012] When it is determined that the traffic allocation plan is custom allocation, reading the target traffic of the main line bucket and the target traffic of the test bucket from the parameter configuration information;

[0013] Update the traffic configuration of the inference graph according to the target traffic of the main bucket and the target traffic of the test bucket;

[0014] Based on the updated inference graph, allocate the user traffic of each model to be tested to the service set, so that the service set can use each user traffic to determine the ability index of each model to be tested.

[0015] According to a model ability test method provided by the present invention, the step of updating the traffic configuration of the inference graph according to the target traffic of the main bucket and the target traffic of the test bucket includes:

[0016] When it is determined that rolling release is not set in the parameter configuration information, update the traffic configuration of the inference graph by using the target traffic of the main bucket and the target traffic of the test bucket;

[0017] When it is determined that rolling release is set in the parameter configuration information, read the time interval and step size from the parameter configuration information;

[0018] Calculate the rolling release strategy according to the target traffic of the main bucket, the target traffic of the test bucket, the time interval and the step size;

[0019] Write the rolling release strategy into the database;

[0020] Use the rolling release module to retrieve the rolling release strategy from the database;

[0021] When the update time is reached, read the traffic configuration information in the rolling release strategy;

[0022] Update the traffic configuration of the inference graph by using the traffic configuration information.

[0023] According to a model ability test method provided by the present invention, when it is determined that the request type is full service, after determining the request type of the update request, it further includes:

[0024] Read the first service information of the first target test bucket from the parameter configuration information, set the traffic of the first target test bucket to 100%, and clear the traffic of the test buckets and the main bucket other than the first target test bucket;

[0025] Based on the first service information, use the service set management module to update the main service of the service set;

[0026] Determine that the updated main service is the full service of the first target test bucket, so as to update the traffic configuration of the inference graph.

[0027] According to a model ability testing method provided by the present invention, after determining that the request type is traffic rollback and after determining the request type of the update request, it further includes:

[0028] Read the second service information of the second target test bucket from the parameter configuration information, and read the traffic information of the target test bucket;

[0029] Based on the second service information, clear the traffic of the target test bucket, and add the traffic information to the main line bucket to update the traffic configuration of the inference graph.

[0030] According to a model ability testing method provided by the present invention, the ability index includes recall rate and precision. Based on the updated inference graph, distributing the user traffic of each model to be tested to the service set for the service set to determine the ability index of each model to be tested using each user traffic includes:

[0031] Based on the updated inference graph, distribute the user traffic of each model to be tested to the service set;

[0032] Use the service set to test at least one user traffic, and obtain the recall rate and precision of the model to be tested corresponding to each user traffic.

[0033] According to a model ability testing method provided by the present invention, reading the parameter configuration information of the update request to determine the request type of the update request includes:

[0034] Parse the participating service information in the parameter configuration information;

[0035] Read the management record of the service set management module;

[0036] Based on the management record, when it is determined that the participating service information meets the preset conditions, determine the request type of the update request.

[0037] The present invention also provides a model ability testing device, including:

[0038] A receiving module, configured to receive an update request;

[0039] A response module, configured to, in response to the update request, read the parameter configuration information of the update request to determine the request type of the update request; the request type includes: traffic allocation, service full volume, and traffic rollback;

[0040] A first determination module, configured to, when it is determined that the request type is traffic allocation, determine a traffic allocation scheme according to the parameter configuration information;

[0041] The first reading module is configured to read the total flow rate of the main line bucket, the total flow rate of the test buckets, and the number of test buckets from the parameter configuration information when it is determined that the flow distribution scheme is equal distribution.

[0042] The second determination module is configured to determine the total flow rate of the main line bucket as the target flow rate of the main line bucket, and determine the target flow rate of each test bucket according to the total flow rate of the test buckets and the number of test buckets.

[0043] The second reading module is configured to read the target flow rate of the main line bucket and the target flow rate of the test buckets from the parameter configuration information when it is determined that the flow distribution scheme is custom distribution.

[0044] The update module is configured to update the flow configuration of the inference graph according to the target flow rate of the main line bucket and the target flow rate of the test buckets.

[0045] The allocation module is configured to, based on the updated inference graph, allocate the user traffic of each model to be tested to the service set, so that the service set can use each user traffic to determine the ability index of each model to be tested.

[0046] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for testing the model ability as described in any one of the above is implemented.

[0047] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for testing the model ability as described in any one of the above is implemented.

[0048] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for testing the model ability as described in any one of the above is implemented.

[0049] The method, device, electronic device, storage medium, and product for testing the model ability provided by the present invention maintain and update the flow configuration of the inference graph in the shunt system update module, intercept and redistribute the user traffic of the model to be tested by using the updated inference graph, and then use the service set to test the ability index of each model to be tested. The operation is simple, and the performance such as the generalization ability of each model to be tested can be evaluated quickly and objectively. Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 is one of the schematic flowcharts of the model ability test method provided by the present invention;

[0052] Figure 2 is the schematic flowchart of the inference graph creation / update method provided by the present invention;

[0053] Figure 3 is the second schematic flowchart of the model ability test method provided by the present invention;

[0054] Figure 4 is the schematic structural diagram of the model ability test device provided by the present invention;

[0055] Figure 5 is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the present invention clearer, the following clearly and completely describes the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0057] In recent years, deep learning has developed rapidly, and a variety of deep learning models have emerged in various fields, such as Faster RCNN, YOLO, ResNet, etc. in the field of computer vision, and Bert, XLNet, Transformer, etc. in the field of natural language processing.

[0058] The A / B test system uses a data-driven approach and utilizes the user data generated online to determine which service performs better. Usually, the process of A / B test is to take a small part of the online traffic, randomly distribute it to service A and service B, and then combine some statistical methods to obtain an accurate estimate of the relative effects of the two services. The A / B test system usually also supports the comparison of multiple services, that is, A / B / n test.

[0059] The traffic splitting process in the A / B test system aims to allocate online users to different buckets according to a fixed traffic ratio and maintain the allocation relationship of these buckets, so as to verify whether the relevant indicators have improved by comparison.

[0060] The following combines Figures 1 to 5 to describe the model ability test method, device, electronic device, storage medium and product provided by the embodiments of the present invention.

[0061] For the model ability test method provided by the embodiments of the present invention, the execution subject can be an electronic device or a software or functional module or functional entity in the electronic device that can implement the model ability test method. In the embodiments of the present invention, the electronic device includes but is not limited to the traffic splitting system update module. It should be noted that the above execution subject does not constitute a limitation to the present invention.

[0062] The inference graph (AIS-Inference Graph) has a traffic division function. The service provider can set the services to participate in and the traffic of each participating service, then create an inference graph, and update the inference graph when there is a need to update the inference graph later. Furthermore, based on the updated inference graph, the model ability is tested.

[0063] However, this testing method requires manually setting the traffic of each participating service, which is cumbersome and not objective.

[0064] Figure 1 is one of the schematic flowcharts of the model ability test method provided by the present invention. As Figure 1 shown, it includes but is not limited to the following steps:

[0065] First, in step S1, an update request is received.

[0066] The update request can be sent by the service provider to update the inference graph in the traffic splitting system update module. The service provider can create a service and update the traffic splitting system update module.

[0067] The update request carries parameter configuration information for updating the inference graph in the traffic splitting system update module.

[0068] Furthermore, in step S2, in response to the update request, the parameter configuration information of the update request is read to determine the request type of the update request; the request type includes: traffic allocation, service full volume, and traffic rollback.

[0069] The traffic splitting system update module responds to the received update request and updates the traffic configuration of the inference graph according to the configuration information carried by the update request.

[0070] The AIStattion Inference Platform is an inference service software based on Kubernetes. The Inference Graph (AIS-Inference Graph) is self-developed software of the AIStation Inference Platform and is used for the orchestration of inference services. It can connect services into a topology graph for execution. The AIS-Inference Graph provides a basic traffic division function. When the AIS-Inference Graph is used in the traffic splitting system for A / B testing in the inference platform, it can effectively solve the problem of quality comparison of multiple and multi-version models of inference services.

[0071] The parameter configuration information may include: participating service information and request type.

[0072] In response to the update request, the traffic splitting system update module reads the parameter configuration information carried in the update request and can obtain the request type. Different request types correspond to different parameter configuration information, and the update method is determined according to the settings in the parameter configuration information.

[0073] Using the parameter configuration information corresponding to the request type, update the traffic configuration of the inference graph. By responding to the update request from the service provider, further update the traffic configuration of the inference graph in the traffic splitting system update module, thereby providing a basis for the performance analysis of the model.

[0074] Optionally, reading the parameter configuration information of the update request to determine the request type of the update request includes:

[0075] Parse the participating service information in the parameter configuration information;

[0076] Read the management records of the service set management module;

[0077] When it is determined based on the management records that the participating service information meets the preset conditions, determine the request type of the update request.

[0078] The participating service information is set by the service provider and is related to the request type. If the request type is traffic allocation, the participating service information includes the traffic allocation scheme, as well as the service names, main service, and test service that need to participate in the traffic allocation; if the request type is full service, the participating service information includes the service names of the test buckets that need full service, the first service information; if the request type is traffic rollback, the participating service information includes the service names that need traffic rollback and the second service information. The second service information includes: service name and traffic information P.

[0079] Specifically, analyze the participating service information in the parameter configuration information, and read the management records of the service set management module to check the service relationships of the participating service information according to the management records. The check content may include:

[0080] (1) Each service in the participating service information belongs to the same service set;

[0081] (2) The participating service information includes the main service;

[0082] (3) In the case where the update request is traffic rollback or service full volume, verify that the service to be rolled back or the full volume service belongs to the test service.

[0083] Therefore, in the case where the update request is traffic allocation, the preset conditions may include: each service in the participating service information belongs to the same service set, and the participating service information includes the main service; in the case where the update request is traffic rollback or service full volume, the preset conditions may include: each service in the participating service information belongs to the same service set, the participating service information includes the main service, and the service to be rolled back belongs to the test service.

[0084] In the case where it is determined that the participating service information meets the preset conditions, determine that the update request is valid, and then perform the next operation; in the case where it is determined that the participating service information does not meet the preset conditions, determine that the update request is invalid, generate an invalid alarm to provide an invalid feedback to the service provider.

[0085] According to the model ability test method provided by the present invention, by analyzing and verifying the participating service information set by the service provider, it can effectively ensure the normal progress of the traffic configuration update of the inference graph and improve the security of the update process.

[0086] Further, in step S3, in the case where it is determined that the request type is traffic allocation, determine the traffic allocation plan according to the parameter configuration information;

[0087] If the request type is traffic allocation, the parameter configuration information further includes: the total traffic of the main bucket, the time interval, the step size, the traffic allocation plan; the traffic allocation plan includes: custom allocation and average allocation;

[0088] If the request type is service full volume, the parameter configuration information further includes: the first service of the test bucket that needs to be full volume;

[0089] If the request type is traffic rollback, the parameter configuration information further includes: the second service of the test bucket that needs traffic rollback.

[0090] In the case where it is determined that the request type is traffic allocation, determine the traffic allocation plan in the parameter configuration information;

[0091] Further, in step S4, when it is determined that the traffic allocation scheme is equal distribution, read the total traffic of the main bucket, the total traffic of the test buckets, and the number of test buckets from the parameter configuration information;

[0092] When the traffic allocation scheme is equal distribution, read the total traffic of the main bucket, the total traffic of the test buckets, and the number of test buckets from the parameter configuration information;

[0093] Further, in step S5, determine the total traffic of the main bucket as the target traffic of the main bucket, and determine the target traffic of each test bucket according to the total traffic of the test buckets and the number of test buckets;

[0094] Since each bucket contains only one service and there is only one main bucket, the target traffic of the main bucket is the total traffic of the main bucket in the parameter configuration information;

[0095]

[0096] Wherein, when the total traffic of the test buckets cannot be divided evenly by the number of test buckets, first add the quotient of dividing the total traffic of the test buckets by the number of test buckets to each test bucket respectively, and then randomly add the remainder to one of the test buckets.

[0097] Further, in step S6, when it is determined that the traffic allocation scheme is custom allocation, read the target traffic of the main bucket and the target traffic of the test buckets from the parameter configuration information;

[0098] When the traffic allocation scheme is custom allocation, read the traffic of the main bucket configured by the user and the traffic of each test bucket from the parameter configuration information as the target traffic that each test bucket needs to be configured.

[0099] Further, in step S7, update the traffic configuration of the inference graph according to the target traffic of the main bucket and the target traffic of the test buckets.

[0100] According to the target traffic of the main bucket and the target traffic of the test buckets, further maintain and update the traffic configuration of the inference graph in the shunt system update module, so as to provide a basis for the performance analysis of the model.

[0101] Optionally, the updating the traffic configuration of the inference graph according to the target traffic of the main bucket and the target traffic of the test buckets includes:

[0102] When it is determined that rolling release is not set in the parameter configuration information, update the traffic configuration of the inference graph by using the target traffic of the main bucket and the target traffic of the test buckets;

[0103] When it is determined that rolling release is set in the parameter configuration information, read the time interval and step size from the parameter configuration information;

[0104] Calculate the rolling release strategy according to the target flow of the main line bucket, the target flow of the test bucket, the time interval and the step size;

[0105] Write the rolling release strategy into the database;

[0106] Use the rolling release module to retrieve the rolling release strategy from the database;

[0107] When the update time is reached, read the traffic configuration information in the rolling release strategy;

[0108] Use the traffic configuration information to update the traffic configuration of the inference graph.

[0109] Rolling release can also be set in the parameter configuration information.

[0110] The calculation method of the rolling release strategy is: every preset time period, randomly roll the traffic of the step size to the test bucket until the traffic in the main line bucket and the test bucket both reach their respective target flows.

[0111] First, determine whether rolling release is set in the parameter configuration information. If it is determined that rolling release is not set in the parameter configuration information, directly read the target flow of the main line bucket and the target flow of the test bucket calculated in the previous process, and update the traffic configuration of the inference graph;

[0112] When it is determined that rolling release is set in the parameter configuration information, read the time interval and step size from the parameter configuration information, and read the target flow of the main line bucket and the target flow of the test bucket calculated in the previous process;

[0113] Calculate the rolling release strategy according to the target flow of the main line bucket, the target flow of the test bucket, the time interval and the step size;

[0114] Write the rolling release strategy into the database;

[0115] Use the rolling release module to retrieve the rolling release strategy from the database;

[0116] When the update time is reached, read the traffic configuration information in the rolling release strategy; the update time can be flexibly set according to the needs of the service provider. The time interval between adjacent update times is related to the delay set by the shunting system update module. The shorter the interval, the shorter the delay.

[0117] In the case where the update time has not been reached, retrieve the rolling release policy from the database again until the update time is reached, and read the traffic configuration information in the rolling release policy.

[0118] Use the traffic configuration information to update the traffic configuration of the inference graph. The rolling release module is a resident loop program. When it detects that the rolling release policy in the database reaches the update time, it updates the inference graph traffic configuration according to the traffic configuration information in the rolling release policy, and repeatedly reads the rolling release policy in the database to achieve real-time update of the inference graph traffic configuration.

[0119] According to the model ability test method provided by the present invention, by calculating the rolling release policy, and then using the rolling release module to perform real-time update of the inference graph traffic configuration, thereby changing the traffic division method of the inference graph, and further providing a basis for the performance analysis of the model.

[0120] Optionally, in the case where it is determined that the request type is full service, after determining the request type of the update request, it further includes:

[0121] Read the first service information of the first target test bucket from the parameter configuration information, set the traffic of the first target test bucket to 100%, and clear the traffic of the test buckets and the main line bucket other than the first target test bucket;

[0122] Based on the first service information, use the service set management module to update the main service of the service set;

[0123] Determine that the updated main service is the full service of the first target test bucket to update the traffic configuration of the inference graph.

[0124] The first target test bucket is the test bucket that needs to perform full service, and the first service information includes: service name and traffic information.

[0125] The service provider sets the participating service in the parameter configuration information and sends an update request of the traffic rollback type; the shunt system update module parses the parameter configuration information of the update request, reads the first service information of the first target test bucket that needs full service from the parameter configuration information, sets the traffic of this bucket to 100%, sets the traffic of other buckets to 0%, then updates the main service of the service set to the service that needs full service, determines that the updated main service is the full service of the first target test bucket, and finally updates the inference graph traffic configuration.

[0126] According to the model ability test method provided by the present invention, by setting the traffic of the test bucket that needs full service and other buckets, and then updating the traffic configuration of the inference graph, it provides a basis for the performance analysis of the model.

[0127] Optionally, when it is determined that the request type is traffic rollback, after determining the request type of the update request, the following is further included:

[0128] Read the second service information of the second target test bucket from the parameter configuration information, and read the traffic information of the target test bucket;

[0129] Based on the second service information, clear the traffic of the target test bucket, and add the traffic information to the main line bucket to update the traffic configuration of the inference graph.

[0130] The second target test bucket is the test bucket that needs to perform traffic rollback.

[0131] The service provider sets the participating services in the parameter configuration information and sends an update request of the full-service type; the shunt system update module parses the parameter configuration information of the update request, reads the service information of the test bucket that needs traffic rollback from the parameter configuration information, reads the traffic information P of this bucket, sets the traffic of this bucket to 0%, adds the traffic information P to the traffic of the main line bucket, and finally updates the traffic configuration of the inference graph.

[0132] According to the model ability test method provided by the present invention, by setting the traffic of the test bucket that needs to perform traffic rollback and other buckets, the traffic configuration of the inference graph is updated, providing a basis for the performance analysis of the model.

[0133] Further, in step S8, based on the updated inference graph, the user traffic of each model to be tested is allocated to the service set for the service set to determine the ability index of each model to be tested using each user traffic.

[0134] The model to be tested is a neural network model that needs to perform ability tests, and the user traffic is the traffic generated during the use of the model to be tested after deployment.

[0135] The updated inference graph can intercept the user traffic of each model to be tested and then allocate it to the service set. The service set is used to maintain and manage all inference services of the model layer.

[0136] The service set management module can use the service set to call the user traffic of the model to be tested and test the called user traffic, thereby obtaining the ability index of each model to be tested. The ability index of the model to be tested is used to characterize the performance of the model.

[0137] The ability index of the model to be tested may include the recall rate and precision of the model, and may also include the service level objective (SLO) value and service level agreement (SLA) value of the model.

[0138] The model ability testing method provided by the present invention maintains and updates the traffic configuration of the inference graph in the shunt system update module, intercepts and redistributes the user traffic of the model to be tested by using the updated inference graph, and then uses the service set to test the ability index of each model to be tested. The operation is simple, and the performance such as the generalization ability of each model to be tested can be evaluated quickly and objectively.

[0139] Optionally, the ability index includes recall rate and precision. Based on the updated inference graph, the user traffic of each model to be tested is allocated to the service set for the service set to determine the ability index of each model to be tested by using each user traffic, including:

[0140] Based on the updated inference graph, the user traffic of each model to be tested is allocated to the service set;

[0141] Using the service set to test at least one user traffic to obtain the recall rate and precision of the model to be tested corresponding to each user traffic.

[0142] The updated inference graph can intercept the user traffic of each model to be tested and then redistribute it to the service set.

[0143] The service set management module can use the service set to call the user traffic of each model to be tested, test each user traffic, and calculate the precision and recall rate of the historical user request traffic in real time according to the response data of the model to the user request and the manually marked response data to the request, so as to obtain the recall rate and precision of each model to be tested.

[0144] According to the model ability testing method provided by the present invention, the user traffic of the model to be tested is intercepted and redistributed, and then the generalization ability and service quality of multiple types and multiple versions of models during online services are compared.

[0145] Figure 2 is a schematic flowchart of the inference graph creation / update method provided by the present invention, as Figure 2 shown, including:

[0146] When the service provider needs to update the inference graph, the AIS - stattion inference platform reads the traffic participating in the service in the inference graph to create / update the inference graph traffic configuration.

[0147] However, Figure 2 the AIS - Inference Graph in [] is still far from the shunt system for A / B testing, and it cannot provide functions of the A / B testing system such as service set management, rolling release, traffic average distribution, service full volume, and traffic rollback.

[0148] Figure 3 This is the second process schematic diagram of the model ability test method provided by the present invention. As Figure 3 shown, Figure 3 Based on Figure 2 , a service set management module and a shunt system update module are added. The shunt system update module includes a rolling release module.

[0149] Among them, the service set management module is responsible for managing the services participating in the A / B test. Its functions include creating and deleting the main service and the test service, and maintaining the subordinate relationship between the main service and the test service. There can only be one main service in the service set, and there can be multiple test services.

[0150] The shunt system update module uses the services in the service set to create / update the shunt policy. When user traffic arrives, it decides which bucket (where "bucket" refers to service here) the traffic flows to according to the shunt policy.

[0151] The functions in the shunt system update module include:

[0152] The service provider sends an update request for the shunt system. After receiving the update request, the shunt system update module reads the parameter configuration information in the update request.

[0153] Parse the participating service information in the parameter configuration information, and then perform a service set relationship check on the participating service information. The participating service information includes: traffic information P and service name.

[0154] The content of the service set relationship check includes:

[0155] First, read the management records of the service set management module to verify whether each service in the participating service information belongs to the same service set; a service set contains multiple service collections;

[0156] Second, verify whether the participating service information contains the main service;

[0157] Next, when it is determined according to the parameter configuration information that the request type of the update request is traffic rollback, verify whether the service to be rolled back is a test service; there are three types of request types for the update request, including: traffic allocation, service full volume, and traffic rollback.

[0158] On the first hand, when the request type of the update request is traffic allocation, judge the traffic allocation scheme in the parameter configuration information of the update request. The traffic allocation scheme includes average allocation and custom allocation.

[0159] When the traffic allocation scheme is average allocation, read the total traffic of the main bucket and the total traffic of the test buckets, and the number of test buckets from the parameter configuration information;

[0160] Since each bucket contains only one service and there is only one main bucket, the target traffic of the main bucket is the total traffic of the main bucket in the parameter configuration information;

[0161]

[0162] Among them, when it cannot be divided evenly, first add the quotient to each test bucket respectively, and then randomly add the remainder to one of the test buckets.

[0163] When the traffic allocation scheme is custom allocation, read the traffic of the user-defined configuration main bucket and the traffic of each test bucket from the parameter configuration information as the target traffic to be configured for each bucket. After determining the target traffic of each bucket, enter the rolling release step.

[0164] First, determine whether rolling release is set in the parameter configuration information. If rolling release is not set, directly read the target traffic of the main bucket and the target traffic of the test bucket calculated in the previous process, and then update the inference graph in the shunt system update module.

[0165] If rolling release is set, update the inference graph traffic configuration according to the following steps:

[0166] Read the time interval and step size from the parameter configuration information, and read the target traffic of the main bucket and the target traffic of the test bucket calculated in the previous process;

[0167] Calculate the rolling release strategy according to the time interval, step size, and each target traffic, and then write the rolling release strategy into the database; among them, the calculation method of the rolling release strategy is: every preset time period, randomly roll the traffic of the step size to the test bucket until the traffic in the main bucket and the test bucket reaches their respective target traffic.

[0168] The rolling release module is a resident loop program. When it detects that the rolling release strategy in the database reaches the update time, update the inference graph traffic configuration according to the traffic configuration information in the rolling release strategy, and repeatedly read the rolling release strategy in the database to achieve real-time update of the inference graph traffic configuration.

[0169] If rolling release is not set in the parameter configuration information, directly update the inference graph traffic configuration according to the target traffic of each bucket.

[0170] In a second aspect, when the request type is full service volume, the first service information of the first target test bucket that needs to be in full volume is read from the parameter configuration information, the traffic of the first test bucket is set to 100%, the traffic of other test buckets and the main bucket is set to 0%, and then, using the first service information, the service set management module is controlled to update the main service of the service set to the full-service volume of the first target test bucket that needs to be in full volume. On this basis, the inference graph traffic configuration is updated.

[0171] In a third aspect, when the request type is traffic rollback, the second service information of the second target test bucket is read from the parameter configuration information, the traffic information P of the second target test bucket is read, the traffic of the second target test bucket is set to 0%, the traffic of the main bucket is added with the traffic information P, and on this basis, the inference graph traffic configuration is updated.

[0172] In addition, while updating the inference graph traffic configuration, the shunt system update module also needs to maintain an interaction relationship with the outside. In the shunt system update module, the services in the service set management are used to form the inference graph, and the service set relationship in the service set management is maintained according to the service configuration information of the inference graph in the shunt system update module.

[0173] According to the model ability test method provided by the present invention, by using the AIS-Inference Graph in the shunt system of the A / B test in the inference platform, it is convenient for service providers to use the A / B test system to compare the generalization ability and service quality when comparing multiple types and multiple versions of model online services.

[0174] When the inference graph needs to be updated, the service provider constructs the main service and then creates multiple test services, with each test service corresponding to a model to be tested.

[0175] The service provider sets the participating service set, the total traffic of the main bucket, the total traffic of the test buckets, the time interval, and the step length in the parameter configuration information, and sends a request of the traffic allocation type that needs to be evenly distributed and rolled out; the shunt system update module parses the parameter configuration information of the update request, enters the even distribution process to set the target traffic of each bucket, and then enters the rolling release process to set the rolling release strategy; the rolling release module reads the database in a loop, and when the strategy reaches the update time, the inference graph traffic configuration is updated.

[0176] The service provider sets the participating services in the parameter configuration information and sends an update request for the type of traffic rollback required; the shunting system update module parses the parameter configuration information of the update request, reads the service information of the test buckets that require the full volume of the service from the parameter configuration information, sets the traffic of this bucket to 100%, sets the traffic of other buckets to 0%, then updates the main service of the service set to the service that requires the full volume, determines that the updated main service is the full volume service of the first target test bucket, and finally updates the traffic configuration of the inference graph;

[0177] The service provider sets the participating services in the parameter configuration information and sends an update request for the type of full service volume required; the shunting system update module parses the parameter configuration information of the update request, reads the service information of the test buckets that require traffic rollback from the parameter configuration information, reads the traffic information P of this bucket, sets the traffic of this bucket to 0%, adds P to the traffic of the main bucket, and finally updates the traffic configuration of the inference graph.

[0178] The model ability test device provided by the present invention is described below, and the model ability test device described below can be mutually corresponding and referred to the model ability test method described above.

[0179] Figure 4 It is a structural schematic diagram of the model ability test device provided by the present invention, as Figure 4 shown, including:

[0180] A receiving module 401, configured to receive an update request;

[0181] A response module 402, configured to, in response to the update request, read the parameter configuration information of the update request to determine the request type of the update request; the request type includes: traffic allocation, full service volume, and traffic rollback;

[0182] A first determination module 403, configured to, when determining that the request type is traffic allocation, determine a traffic allocation scheme according to the parameter configuration information;

[0183] A first reading module 404, configured to, when determining that the traffic allocation scheme is equal distribution, read the total traffic of the main bucket, the total traffic of the test buckets, and the number of test buckets from the parameter configuration information;

[0184] A second determination module 405, configured to determine the total traffic of the main bucket as the target traffic of the main bucket, and determine the target traffic of each test bucket according to the total traffic of the test buckets and the number of test buckets;

[0185] A second reading module 406, configured to read the target flow rate of the main bucket and the target flow rate of the test bucket from the parameter configuration information when it is determined that the flow allocation scheme is custom allocation;

[0186] An update module 407, configured to update the flow configuration of the inference graph according to the target flow rate of the main bucket and the target flow rate of the test bucket;

[0187] An allocation module 408, configured to allocate the user traffic of each model to be tested to a service set based on the updated inference graph, so that the service set can determine the ability index of each model to be tested by using each user traffic.

[0188] During the operation of the device, a receiving module 401 receives an update request; an update module 402 responds to the update request and reads the parameter configuration information of the update request to determine the request type of the update request; the request type includes: flow allocation, service full volume, and flow rollback; a first determination module 403 determines a flow allocation scheme according to the parameter configuration information when it is determined that the request type is flow allocation; a first reading module 404 reads the total flow rate of the main bucket, the total flow rate of the test bucket, and the number of test buckets from the parameter configuration information when it is determined that the flow allocation scheme is average allocation; a second determination module 405 determines the total flow rate of the main bucket as the target flow rate of the main bucket, and determines the target flow rate of each test bucket according to the total flow rate of the test bucket and the number of test buckets; a second reading module 406 reads the target flow rate of the main bucket and the target flow rate of the test bucket from the parameter configuration information when it is determined that the flow allocation scheme is custom allocation; an update module 407 updates the flow configuration of the inference graph according to the target flow rate of the main bucket and the target flow rate of the test bucket; an allocation module 408 allocates the user traffic of each model to be tested to a service set based on the updated inference graph, so that the service set can determine the ability index of each model to be tested by using each user traffic.

[0189] The model ability test device provided by the present invention maintains and updates the flow configuration of the inference graph in the shunt system update module, intercepts and redistributes the user traffic of the model to be tested by using the updated inference graph, and then uses the service set to test the ability index of each model to be tested. The operation is simple, and the performance such as the generalization ability of each model to be tested can be evaluated quickly and objectively.

[0190] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute the model capability test method, which includes: receiving an update request; in response to the update request, reading the parameter configuration information of the update request to determine the request type of the update request; the request type includes: traffic allocation, full service, and traffic rollback; in the case where it is determined that the request type is traffic allocation, determining a traffic allocation plan according to the parameter configuration information; in the case where it is determined that the traffic allocation plan is equal distribution, reading the total traffic of the main line bucket and the total traffic of the test bucket, and the number of test buckets from the parameter configuration information; determining the total traffic of the main line bucket as the target traffic of the main line bucket, and determining the target traffic of each test bucket according to the total traffic of the test bucket and the number of test buckets; in the case where it is determined that the traffic allocation plan is custom allocation, reading the target traffic of the main line bucket and the target traffic of the test bucket from the parameter configuration information; updating the traffic configuration of the inference graph according to the target traffic of the main line bucket and the target traffic of the test bucket; based on the updated inference graph, allocating the user traffic of each model to be tested to the service set for the service set to use each user traffic to determine the capability index of each model to be tested.

[0191] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0192] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model ability test method provided by each of the above methods. The method includes: receiving an update request; in response to the update request, reading the parameter configuration information of the update request to determine the request type of the update request; the request type includes: traffic allocation, service full volume, and traffic rollback; when determining that the request type is traffic allocation, determining a traffic allocation scheme according to the parameter configuration information; when determining that the traffic allocation scheme is average allocation, reading the total traffic of the main line bucket and the total traffic of the test bucket, and the number of test buckets from the parameter configuration information; determining the total traffic of the main line bucket as the target traffic of the main line bucket, and determining the target traffic of each test bucket according to the total traffic of the test bucket and the number of test buckets; when determining that the traffic allocation scheme is custom allocation, reading the target traffic of the main line bucket and the target traffic of the test bucket from the parameter configuration information; updating the traffic configuration of the inference graph according to the target traffic of the main line bucket and the target traffic of the test bucket; based on the updated inference graph, allocating the user traffic of each model to be tested to a service set for the service set to use each user traffic to determine the ability index of each model to be tested.

[0193] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the model ability test method provided by each of the above methods. The method includes: receiving an update request; in response to the update request, reading the parameter configuration information of the update request to determine the request type of the update request; the request type includes: traffic allocation, service full volume, and traffic rollback; when determining that the request type is traffic allocation, determining a traffic allocation scheme according to the parameter configuration information; when determining that the traffic allocation scheme is average allocation, reading the total traffic of the main line bucket and the total traffic of the test bucket, and the number of test buckets from the parameter configuration information; determining the total traffic of the main line bucket as the target traffic of the main line bucket, and determining the target traffic of each test bucket according to the total traffic of the test bucket and the number of test buckets; when determining that the traffic allocation scheme is custom allocation, reading the target traffic of the main line bucket and the target traffic of the test bucket from the parameter configuration information; updating the traffic configuration of the inference graph according to the target traffic of the main line bucket and the target traffic of the test bucket; based on the updated inference graph, allocating the user traffic of each model to be tested to a service set for the service set to use each user traffic to determine the ability index of each model to be tested.

[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for testing model capabilities, characterized in that, Including: Receiving an update request; In response to the update request, reading parameter configuration information of the update request to determine a request type of the update request; The request type includes traffic allocation, service full volume, and traffic rollback; When determining that the request type is traffic allocation, determining a traffic allocation scheme according to the parameter configuration information; When determining that the traffic allocation scheme is equal distribution, reading a total traffic of a main line bucket, a total traffic of a test bucket, and a number of test buckets from the parameter configuration information; Determining the total traffic of the main line bucket as a target traffic of the main line bucket, and determining a target traffic of each test bucket according to the total traffic of the test bucket and the number of test buckets; When determining that the traffic allocation scheme is custom allocation, reading a target traffic of the main line bucket and a target traffic of the test bucket from the parameter configuration information; Updating a traffic configuration of an inference graph according to the target traffic of the main line bucket and the target traffic of the test bucket; Based on the updated inference graph, allocating user traffic of each model to be tested to a service set, so that the service set determines an ability index of each model to be tested by using each user traffic; The updating the traffic configuration of the inference graph according to the target traffic of the main line bucket and the target traffic of the test bucket includes: When determining that rolling release is not set in the parameter configuration information, updating the traffic configuration of the inference graph by using the target traffic of the main line bucket and the target traffic of the test bucket; When determining that rolling release is set in the parameter configuration information, reading a time interval and a step length from the parameter configuration information; Calculating a rolling release strategy according to the target traffic of the main line bucket, the target traffic of the test bucket, the time interval, and the step length; Writing the rolling release strategy into a database; Invoking the rolling release strategy from the database by using a rolling release module; When an update time is reached, reading traffic configuration information in the rolling release strategy; Updating the traffic configuration of the inference graph by using the traffic configuration information.

2. The model ability test method according to claim 1, characterized in that When determining that the request type is service full volume, after determining the request type of the update request, further including: Reading first service information of a first target test bucket from the parameter configuration information, setting the traffic of the first target test bucket to 100%, and clearing the traffic of test buckets other than the first target test bucket and the main line bucket; Based on the first service information, updating a main service of the service set by using a service set management module; Determining that the updated main service is a full volume service of the first target test bucket to update the traffic configuration of the inference graph.

3. The model ability test method according to claim 1, wherein When determining that the request type is traffic rollback, after determining the request type of the update request, further including: Reading second service information of a second target test bucket from the parameter configuration information, and reading traffic information of the target test bucket; Based on the second service information, clearing the traffic of the target test bucket, and adding the traffic information to the main line bucket to update the traffic configuration of the inference graph.

4. The model ability test method according to any one of claims 1-3, characterized in that The ability index includes recall rate and precision. Based on the updated inference graph, distributing the user traffic of each model to be tested to the service set for the service set to determine the ability index of each model to be tested using each user traffic, including: Based on the updated inference graph, distributing the user traffic of each model to be tested to the service set; Using the service set to test at least one user traffic and obtaining the recall rate and precision of the model to be tested corresponding to each user traffic.

5. The model ability test method according to any one of claims 1-3, characterized in that Reading the parameter configuration information of the update request to determine the request type of the update request, including: Parsing the participating service information in the parameter configuration information; Reading the management record of the service set management module; Determining the request type of the update request when it is determined that the participating service information meets the preset conditions based on the management record.

6. A model ability test device, characterized in that Including: A receiving module for receiving an update request; A response module for, in response to the update request, reading the parameter configuration information of the update request to determine the request type of the update request; The request type includes traffic allocation, service full volume, and traffic rollback; A first determination module for, when determining that the request type is traffic allocation, determining a traffic allocation plan according to the parameter configuration information; A first reading module for, when determining that the traffic allocation plan is average allocation, reading the total traffic of the main line bucket, the total traffic of the test bucket, and the number of test buckets from the parameter configuration information; A second determination module for determining the total traffic of the main line bucket as the target traffic of the main line bucket and determining the target traffic of each test bucket according to the total traffic of the test bucket and the number of test buckets; A second reading module for, when determining that the traffic allocation plan is custom allocation, reading the target traffic of the main line bucket and the target traffic of the test bucket from the parameter configuration information; An update module for updating the traffic configuration of the inference graph according to the target traffic of the main line bucket and the target traffic of the test bucket; A distribution module for, based on the updated inference graph, distributing the user traffic of each model to be tested to the service set for the service set to determine the ability index of each model to be tested using each user traffic; The update module is specifically used for: When it is determined that rolling release is not set in the parameter configuration information, updating the traffic configuration of the inference graph using the target traffic of the main line bucket and the target traffic of the test bucket; When it is determined that rolling release is set in the parameter configuration information, reading the time interval and step size from the parameter configuration information; Calculating a rolling release strategy according to the target traffic of the main line bucket, the target traffic of the test bucket, the time interval, and the step size; Writing the rolling release strategy to the database; Using the rolling release module to retrieve the rolling release strategy from the database; When the update time is reached, reading the traffic configuration information in the rolling release strategy; Updating the traffic configuration of the inference graph using the traffic configuration information.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the model ability test method according to any one of claims 1-5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the model ability test method according to any one of claims 1-5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the model ability test method according to any one of claims 1-5.