Model testing method and device, equipment and storage medium

Through benchmark testing and baseline testing, the benchmark test data set and baseline test data set are used to test the model to be tested, which solves the problems of difficult maintenance and insufficient accuracy of the diagnostic model, and achieves accuracy and reliability assurance before the model is released to meet current and historical development needs.

CN120705070AActive Publication Date: 2025-09-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511194728.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-26
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

In the existing technology, the maintenance of diagnostic models is difficult, the accuracy and reliability of diagnostic results are difficult to guarantee, and problem discovery is delayed, making accurate maintenance difficult to achieve.

Method used

By introducing benchmark tests and baseline tests, the model to be tested is tested using benchmark test data sets and baseline test data sets to ensure its diagnostic accuracy and reliability. This includes sampling server logs in the previous period and new logs added in the current period, dynamically testing the diagnostic accuracy and reliability of the model to be tested to ensure that the model meets current and historical development needs.

Benefits of technology

This ensures diagnostic accuracy before the model is released, reduces maintenance difficulty, improves the accuracy and reliability of the diagnostic model, avoids delayed problem discovery, and ensures that the model meets current and historical development needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705070A_ABST
    Figure CN120705070A_ABST
Patent Text Reader

Abstract

The invention provides a model testing method and device, equipment and a storage medium, and can be applied to the field of computers. The model test method comprises the following steps: inputting a benchmark test data set for a current time period into a test model and a to-be-tested model to obtain a plurality of first diagnosis results of the test model for any first test parameter of a plurality of first server logs and a plurality of second diagnosis results of the to-be-tested model for any first test parameter; obtaining a benchmark test result based on the plurality of first diagnosis results and the plurality of second diagnosis results; inputting the baseline test data set for the current time period into the to-be-tested model to obtain a plurality of third diagnosis results for any second test parameter of the plurality of second server logs; obtaining a baseline test result based on a plurality of baseline values for any second test parameter in the baseline test sample and a plurality of third diagnosis results; and under the condition that the benchmark test result and the baseline test result do not meet the preset condition, determining that the to-be-tested model is abnormal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a model testing method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of information technology, servers are increasingly being used in various fields. Their stability and reliability are crucial to the normal operation of businesses. Automated fault diagnosis systems have become a key technology for ensuring server stability and reliability. The diagnostic models used in these systems can quickly diagnose uploaded server logs and output diagnostic conclusions, assisting operations engineers in quickly diagnosing and repairing servers. Therefore, the accuracy and reliability of the diagnostic model's results directly impact server diagnostics.

[0003] In related technologies, in order to ensure the reliability of the diagnostic results of the diagnostic model, it is usually necessary to regularly check large batches of logs to ensure that the diagnostic accuracy of the diagnostic model meets the requirements. However, this means that the total number of test parameters that need to be maintained is more than tens of thousands, and the number of test parameters that need to be maintained changes dynamically with the total number of logs, which greatly increases the difficulty of maintaining the diagnostic model. Summary of the Invention

[0004] In view of the above problems, the present application provides a model testing method, apparatus, device, medium and program product.

[0005] According to a first aspect of the present application, a model testing method is provided, comprising: inputting a benchmark test data set for a current period into a test model and a model to be tested, respectively, to obtain a plurality of first diagnostic results of the test model for any first test parameter of a plurality of first server logs and a plurality of second diagnostic results of the model to be tested for any first test parameter of the plurality of first server logs, wherein the benchmark test data set includes the first server logs of the previous period and the first server logs obtained by sampling the newly added first server logs of the current period according to a preset ratio; obtaining a benchmark test result based on the number of identical diagnostic results for the same first server log in the plurality of first diagnostic results and the plurality of second diagnostic results; inputting a baseline test data set for the current period into the model to be tested, to obtain a plurality of third diagnostic results for any second test parameter of a plurality of second server logs, wherein the baseline test data set includes the second server logs of the previous period and the second server logs for development requirements of the current period; obtaining a baseline test result based on a plurality of baseline values ​​for any of the second test parameters in the baseline test sample and the number of identical values ​​for the same second server log in the plurality of third diagnostic results; and determining that an abnormality exists in the model to be tested if the benchmark test result and the baseline test result do not meet preset conditions.

[0006] The second aspect of the present application provides a model testing device, including: a first input module, a first acquisition module, a second input module, a second acquisition module and a first determination module. The first input module is used to input the benchmark test data set for the current time period into the test model and the model to be tested respectively, and obtain a plurality of first diagnostic results of the above-mentioned test model for any first test parameter of a plurality of first server logs and a plurality of second diagnostic results of the above-mentioned model to be tested for any first test parameter of the above-mentioned plurality of first server logs, wherein the above-mentioned benchmark test data set includes the first server logs of the previous time period and the first server logs obtained by sampling the newly added first server logs of the current time period according to a preset ratio; the first acquisition module is used to obtain the benchmark test results based on the number of identical diagnostic results for the same first server log in the above-mentioned plurality of first diagnostic results and the above-mentioned plurality of second diagnostic results. ; A second input module is used to input the baseline test data set for the current time period into the above-mentioned model to be tested, and obtain multiple third diagnostic results for any second test parameter of multiple second server logs, wherein the above-mentioned baseline test data set includes the second server logs in the previous time period and the second server logs for development needs of the current time period; a second acquisition module is used to obtain the baseline test result based on multiple baseline values ​​for any of the above-mentioned second test parameters in the baseline test sample and the number of identical values ​​for the same second server log in the above-mentioned multiple third diagnostic results; a first determination module is used to determine that the above-mentioned model to be tested has an abnormality when the above-mentioned benchmark test result and the above-mentioned baseline test result do not meet the preset conditions.

[0007] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0009] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0010] According to the model testing method provided by the present application, before the model to be tested is released, the diagnostic accuracy of the model to be tested before release is guaranteed. Based on this, a benchmark test is introduced. By testing the model, the diagnostic accuracy and reliability of the model to be tested can be flexibly and dynamically tested based on the benchmark test data set of the first server log included in the previous period and the first server log newly added in the current period sampled at a preset ratio; the baseline test is introduced. By using the baseline test sample, the model to be tested is subjected to a baseline test based on the baseline test data set of the second server log included in the previous period and the second server log used for the development needs of the current period, ensuring that the model to be tested before release meets the development needs of the current period, while ensuring that the model to be tested meets the historical development needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 An application scenario diagram of the model testing method according to an embodiment of the present application is shown;

[0012] Figure 2 A flow chart of a model testing method according to an embodiment of the present application is shown;

[0013] Figure 3 A flowchart of obtaining benchmark test results according to an embodiment of the present application is shown;

[0014] Figure 4 A flowchart of obtaining baseline test results according to an embodiment of the present application is shown;

[0015] Figure 5 A schematic diagram of a process of performing a benchmark test and a baseline test on a model to be tested according to an embodiment of the present application is shown;

[0016] Figure 6 A schematic diagram of a benchmark test according to an embodiment of the present application is shown;

[0017] Figure 7 A schematic diagram of a baseline test according to an embodiment of the present application is shown;

[0018] Figure 8 shows a structural block diagram of a model testing device according to an embodiment of the present application;

[0019] Figure 9 A block diagram of an electronic device suitable for implementing a model testing method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0020] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0021] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0023] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0024] To ensure the reliability of the diagnostic model's results, it's generally necessary to regularly verify large volumes of logs to ensure the model's accuracy meets requirements. For example, for a diagnostic system with 1 million logs, sampling at 1% of the total log volume requires over 10,000 sample logs each time to verify the reliability of the model's results. Furthermore, because test parameters include multiple parameters such as fault type, fault level, and model SN (Serial Number), the total number of test parameters that need to be maintained exceeds tens of thousands, and these parameters change dynamically with the total log volume. This significantly increases the difficulty of writing and maintaining test cases for the diagnostic model.

[0025] Furthermore, current verification of diagnostic model reliability primarily occurs after the model is released. Leveraging the diagnostic system's online and multi-platform capabilities, it connects to various systems, such as work order systems, to verify the accuracy of the model's output. This information is then fed back to model and algorithm developers, enabling them to optimize the model's diagnostic code or algorithms accordingly. Pre-release testing of diagnostic models also primarily involves manually maintaining a small number of test case samples and corresponding test code, conducting small batch tests.

[0026] Based on the online docking solution of the diagnostic system, the accuracy and reliability of the diagnostic model after release can be more effectively verified. However, there will be a lag in the discovery of problems with the diagnostic model, which may result in problems being discovered and optimized several months after going online, posing a hidden danger to the reliability of the diagnostic model.

[0027] In addition to delayed problem discovery, the docking system also has many pain points. For example, when the diagnostic system and the work order system are docked, there are complex correspondences between the two systems, such as one log corresponding to multiple work orders or multiple logs corresponding to one work order, making it difficult to distinguish. Furthermore, the maintenance source of the work order system and the diagnostic system log may not correspond. For example, during maintenance, an engineer may replace a component based on the red light on an on-site server component, but the server failure log uploaded to the diagnostic system may not contain a record of the corresponding component failure. Therefore, the above situations require manual analysis to resolve inaccuracies caused by the docking of the work order system.

[0028] Furthermore, the continued increase in diagnostic parameters has increased the difficulty of maintenance. As diagnostic models adapt to the needs of various diagnostic systems, the output parameters of the diagnostic models' diagnostic conclusions will continue to increase, ranging from a dozen to dozens. For example, the output parameters of the diagnostic model increased from 10 to 31, initially including only the basic model, serial number, fault level, and faulty component. Subsequently, parameters such as AI operation and maintenance accuracy, expert rule accuracy, and component SN were gradually added. However, many parameters do not exist in other systems, such as the work order system, making them impossible to verify through docking.

[0029] Based on the above, it is difficult to accurately maintain the diagnostic model in existing solutions. Therefore, the embodiment of the present application provides a model testing method to ensure the diagnostic accuracy of the diagnostic model before release.

[0030] Figure 1 An application scenario diagram of the model testing method according to an embodiment of the present application is shown.

[0031] like Figure 1As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0032] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0033] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0034] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0035] For example, the benchmark test data set for the current time period can be input into the test model and the model to be tested respectively through server 105 to obtain multiple first diagnostic results of the test model for any first test parameter of multiple first server logs and multiple second diagnostic results of the model to be tested for any first test parameter of multiple first server logs, so that the benchmark test result can be obtained based on the number of identical diagnostic results for the same first server log in the multiple first diagnostic results and the multiple second diagnostic results; the baseline test data set for the current time period is input into the model to be tested to obtain multiple third diagnostic results for any second test parameter of multiple second server logs, so that the baseline test result can be obtained based on the multiple baseline values ​​for any second test parameter in the baseline test sample and the number of identical values ​​for the same second server log in the multiple third diagnostic results, so that when the benchmark test result and the baseline test result do not meet the preset conditions, it is determined that the model to be tested has an abnormality.

[0036] It should be noted that the model testing method provided in the embodiment of the present application can generally be executed by the server 105. Accordingly, the model testing device provided in the embodiment of the present application can generally be set in the server 105. The model testing method provided in the embodiment of the present application can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the model testing device provided in the embodiment of the present application can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0038] The following will be based on Figure 1 The scene described by Figures 2 to 7 The model testing method of the embodiment of the present application is described in detail.

[0039] Figure 2 A flow chart of a model testing method according to an embodiment of the present application is shown.

[0040] like Figure 2 As shown, the method 200 includes operations S210 to S250.

[0041] In operation S210, the benchmark test data set used for the current time period is input into the test model and the model to be tested respectively to obtain multiple first diagnostic results of the test model for any first test parameter of multiple first server logs and multiple second diagnostic results of the model to be tested for any first test parameter of multiple first server logs.

[0042] The benchmark test data set includes the first server logs in the previous period and the first server logs obtained by sampling the newly added first server logs in the current period according to a preset ratio.

[0043] According to an embodiment of the present application, the model to be tested may be a diagnostic model used to diagnose server faults. Testing the model to be tested is to ensure the diagnostic accuracy of the model to be tested. The test model may represent a model used to detect the model to be tested.

[0044] According to an embodiment of the present application, since the version of the model to be tested is iterative, such as the version of the model to be tested is iteratively released in each cycle, in the process of testing the updated version of the model to be tested in the current cycle, the benchmark test data set used to test the previous released version of the model to be tested and the server logs newly added in the current cycle compared with the previous cycle can be used as the benchmark test data set of the current version of the model to be tested.

[0045] In one embodiment, since the log increment is large, the newly added server logs can be sampled according to a preset ratio, wherein the preset ratio is set according to needs.

[0046] Specifically, the current period can be represented as the current cycle. Because issues may arise after version updates of the model under test, the model under test needs to be tested to ensure its diagnostic accuracy. The model under test may be tested multiple times within the current period, but the benchmark dataset used for each test of the model under test remains unchanged.

[0047] According to an embodiment of the present application, the first server log in the previous time period represents all the first server logs in the benchmark test data set used to test the previous version of the model to be tested; the first server log newly added in the current time period represents the first server log newly added in the current time period relative to the time period for testing the previous version of the model to be tested.

[0048] For example, if the preset ratio is 1%, the number of first server logs in the benchmark test dataset of the previous cycle is 10,000, and the number of newly added logs between the current cycle and the previous cycle is 30,000. Therefore, the benchmark test dataset of the current cycle includes 1% of the newly added logs, that is, 300 of the newly added logs. Therefore, the number of first server logs in the benchmark test dataset of the current cycle is 10,300.

[0049] According to an embodiment of the present application, benchmarking the model to be tested is to test the entire model to be tested, that is, to ensure that the entire model to be tested meets a certain diagnostic quality. Therefore, the benchmark test data set used for benchmarking is sampled from the total amount of new logs.

[0050] Specifically, by inputting the benchmark test data set into the test model and the model to be tested respectively, multiple first diagnostic results of the test model for any first test parameter of multiple first server logs and multiple second diagnostic results of the model to be tested for any first test parameter can be obtained.

[0051] In operation S220 , a benchmark test result is obtained based on the number of identical diagnostic results for the same first server log among the plurality of first diagnostic results and the plurality of second diagnostic results.

[0052] According to an embodiment of the present application, in order to determine the diagnostic accuracy of the model to be tested, it is necessary to obtain a benchmark test result based on the number of identical diagnostic results for the same first server log in multiple first diagnostic results and multiple second diagnostic results, that is, the benchmark test result can characterize the diagnostic accuracy of the model to be tested.

[0053] In one embodiment, executing the above operations S210 and S220 is to perform a benchmark test on the model to be tested.

[0054] In operation S230 , the baseline test data set for the current period is input into the model to be tested, and a plurality of third diagnosis results for any second test parameter of the plurality of second server logs are obtained.

[0055] The baseline test data set includes the second server log in the previous period and the second server log for development requirements in the current period.

[0056] According to an embodiment of the present application, performing a baseline test on a model to be tested is to test whether the model to be tested meets the development requirements of the current period, that is, to ensure that the model to be tested meets the development requirements of the current period. Therefore, the baseline test dataset used for the baseline test includes a second server log based on the development requirements for the current period.

[0057] In one embodiment, the second server logs in the previous period may represent all second server logs in a baseline test dataset used to test the previous version of the model to be tested.

[0058] Specifically, by inputting the baseline test data set into the model to be tested, a plurality of third diagnostic results for any second test parameter of a plurality of second server logs can be obtained.

[0059] In operation S240 , a baseline test result is obtained based on a plurality of baseline values ​​for any second test parameter in the baseline test sample and the number of identical values ​​for the same second server log in the plurality of third diagnosis results.

[0060] According to an embodiment of the present application, in order to determine whether the model to be tested meets the development requirements, it is necessary to obtain a baseline test result based on multiple baseline values ​​for any second test parameter in the baseline test sample and the number of identical values ​​for the same second server log in multiple third diagnostic results, that is, the baseline test result can characterize whether the model to be tested meets the development requirements.

[0061] In one embodiment, executing the above operations S230 and S240 is to perform a baseline test on the model to be tested.

[0062] Based on the above, a baseline test is performed on the model under test using the baseline test dataset for the current period. This baseline test result can be used to determine whether the model under test meets the development requirements of the current period. Furthermore, because the baseline test dataset also includes the second server logs from the previous period, the baseline test result can also be used to determine whether the model under test also meets all previous development requirements, that is, whether it meets historical development requirements.

[0063] In operation S250 , if the benchmark test result and the baseline test result do not meet a preset condition, it is determined that an abnormality exists in the model to be tested.

[0064] According to an embodiment of the present application, the preset condition is set based on the diagnostic accuracy requirements of the model to be tested. If the benchmark test result and the baseline test result do not meet the preset condition, that is, if at least one of the baseline test result and the baseline test result does not meet the preset condition, it can be determined that the model to be tested has an abnormality.

[0065] According to an embodiment of the present application, when it is determined that an abnormality exists in the model to be tested, it may be indicated that the currently updated model to be tested does not meet the current requirements and the model to be tested needs to be updated again.

[0066] Specifically, the above operations S210 to S250 are performed before the model to be tested is released.

[0067] According to an embodiment of the present application, before the model to be tested is released, the diagnostic accuracy of the model to be tested before release is guaranteed. Based on this, a benchmark test is introduced. By testing the model, the diagnostic accuracy and reliability of the model to be tested can be flexibly and dynamically tested based on a benchmark test data set of the first server log included in the previous period and the first server log newly added in the current period sampled at a preset ratio; a baseline test is introduced. By using a baseline test sample, a baseline test data set of the second server log included in the previous period and the second server log used for the development requirements of the current period is used to perform a baseline test on the model to be tested, ensuring that the model to be tested before release meets the development requirements of the current period, and at the same time ensuring that the model to be tested meets the historical development requirements.

[0068] According to an embodiment of the present application, the above-mentioned model testing method further includes: when the benchmark test results and the baseline test results meet preset conditions, determining the model to be tested as the model to be released.

[0069] According to an embodiment of the present application, in the current period, when both the benchmark test results and the baseline test results meet the preset conditions, it can be indicated that the current model to be tested meets the current requirements, and thus the model to be tested can be determined as the model to be released.

[0070] In one embodiment, the current model to be tested meeting the current requirements may specifically be that the current model to be tested meets the diagnostic accuracy requirements and development requirements.

[0071] According to an embodiment of the present application, within the current time period, if it is determined that the benchmark test results and baseline test results of the model to be tested do not meet the preset conditions, the R&D personnel can update and optimize the model to be tested based on the benchmark test results, baseline test results, and development requirements. This allows the R&D personnel to re-perform the benchmark test and baseline test on the model to be tested after completing the update and optimization of the model to be tested. This allows the R&D personnel to re-perform the benchmark test and baseline test on the model to be tested until the benchmark test results and baseline test results of the model to be tested meet the preset conditions.

[0072] According to an embodiment of the present application, based on the set preset conditions, by performing benchmark tests and baseline tests on the model to be tested, it is ensured that the model to be tested is determined as the model to be released only when the benchmark test results and baseline test results of the model to be tested meet the preset conditions, so as to ensure the diagnostic accuracy of the model to be tested before release.

[0073] According to an embodiment of the present disclosure, the benchmark test data set for the current time period is respectively input into the test model and the model to be tested, and multiple first diagnostic results of the test model for any first test parameter of the multiple first server logs and multiple second diagnostic results of the model to be tested for any first test parameter of the multiple first server logs are obtained, including: inputting the multiple first server logs in the benchmark test data set into the test model, and obtaining the first diagnostic result of each of the multiple first test parameters for any first server log; inputting the multiple first server logs in the benchmark test data set into the model to be tested, and obtaining the second diagnostic result of each of the multiple first test parameters for any first server log.

[0074] According to an embodiment of the present application, a test model can be used to identify multiple first test parameters in a first server log. Thus, multiple first server logs in a benchmark test dataset are input into the test model, and the test model can output a first diagnostic result for each of the multiple first test parameters in any first server log.

[0075] In one embodiment, all first server logs in the benchmark test data set are input into the model to be tested respectively, and the first diagnostic results of each first server log for the plurality of first test parameters can be obtained.

[0076] According to an embodiment of the present application, the model to be tested can be used to identify multiple parameters in the first server log, the multiple first test parameters are part of the multiple parameters in the first server log that the model to be tested can identify, and the multiple first test parameters are core parameters that the model to be tested can identify.

[0077] According to an embodiment of the present application, multiple first server logs in a benchmark test dataset are input into a model under test, and diagnostic results for multiple parameters of any of the first server logs can be obtained. Because the test model identifies multiple first test parameters, and these multiple first test parameters are core parameters, second diagnostic results for each of the multiple first test parameters of any of the first server logs can be selected from the output of the model under test.

[0078] According to an embodiment of the present application, by inputting a benchmark test data set into a test model and then inputting the benchmark test data set into a model to be tested, a first diagnostic result for each of the multiple first test parameters of the test model for any first server log can be obtained, and a second diagnostic result for each of the multiple first test parameters of the model to be tested for any first server log can be obtained. Thus, based on the multiple first diagnostic results and the multiple second diagnostic results, a benchmark test of the model to be tested can be performed, thereby facilitating determination of the diagnostic accuracy of the model to be tested.

[0079] Figure 3 A flow chart for obtaining benchmark test results according to an embodiment of the present application is shown.

[0080] like Figure 3 As shown, the method 300 includes operations S310 to S330.

[0081] According to an embodiment of the present application, the benchmark test result may represent the diagnostic accuracy of the model to be tested for each of the plurality of first test parameters.

[0082] In operation S310 , for any first test parameter of any first server log, a first diagnosis result of any first test parameter is compared with a second diagnosis result to obtain a first comparison result of any first test parameter.

[0083] According to an embodiment of the present application, since a test model is used to test the diagnostic accuracy of the model to be tested, it is necessary to compare the first diagnostic result of the test model for any first test parameter of any first server log and the second diagnostic result of the model to be tested for the corresponding first test parameter of the corresponding first server log to obtain a first comparison result of the model to be tested and the test model for the first test parameter.

[0084] According to an embodiment of the present application, a benchmark test result of the model to be tested can be obtained based on the first comparison results of each of the multiple first test parameters of each first server log in the benchmark test data set.

[0085] In operation S320 , for any first test parameter, a first number of first server logs whose first comparison result indicates that the first diagnosis result is consistent with the second diagnosis result is determined.

[0086] According to an embodiment of the present application, when the first comparison result indicates that the first diagnosis result is consistent with the second diagnosis result, it can be determined that the diagnosis of the model to be tested for a first test parameter of a first server log is correct.

[0087] Specifically, for any first test parameter, the first number of first server logs whose first comparison result indicates that the first diagnostic result is consistent with the second diagnostic result can be determined, that is, the number of first server logs whose diagnosis of the model to be tested is correct for any first test parameter can be determined.

[0088] In operation S330 , a diagnostic accuracy rate of the model to be tested for any first test parameter is obtained according to the first number and the number of first server logs in the benchmark test data set.

[0089] According to an embodiment of the present application, based on the first number and the number of first server logs in the benchmark test data set, the diagnostic accuracy of the model to be tested for any first test parameter can be obtained to determine whether the diagnostic accuracy of the model to be tested meets the requirements.

[0090] In one embodiment, a ratio of the first number to the number of first server logs in the benchmark test data set may be calculated, and the ratio may represent the diagnostic accuracy of the test object satisfying any first test parameter.

[0091] According to an embodiment of the present application, multiple first test parameters may include log type, server failure type and server type; the test model is trained using the log type, server failure type and server type involved in the server log, and the test model is used to identify multiple first test parameters in the first server log.

[0092] In one embodiment, the plurality of first test parameters may further include a serial number and a diagnosis identifier.

[0093] Specifically, the log type can represent the type of the first server log; the server failure type can represent the failure type of the server corresponding to the first server log, such as memory failure, hard disk failure, etc.; the server type can represent the model of the server corresponding to the first server log; the serial number can represent the serial number of the server corresponding to the first server log.

[0094] In one embodiment, the model under test and the test model include expert rule diagnosis, algorithm rule diagnosis, and a decision module. Each log entry is sent to these two diagnostic components, each of which outputs a result. The decision model then determines whether the expert rule diagnosis or the algorithm rule diagnosis is correct and outputs the corresponding result. Thus, the diagnosis flag can indicate the degree of match between the expert rule diagnosis and the algorithm rule diagnosis in the model under test or the test model.

[0095] Since the multiple first test parameters are core parameters, the test model is trained only based on the multiple first test parameters involved in the server log to identify the multiple first test parameters in the first server log.

[0096] According to an embodiment of the present application, a first comparison result of a plurality of first test parameters of a plurality of first server logs between the model to be tested and the test model may be as shown in Table 1 below.

[0097] Table 1

[0098]

[0099] In Table 1, "pass" indicates that the diagnostic results of the model under test and the test model regarding the first test parameter are consistent, and "fail" indicates that the diagnostic results of the model under test and the test model regarding the first test parameter are inconsistent.

[0100] Among them, CPU stands for Central Processing Unit; BMC stands for Baseboard Management Controller.

[0101] According to an embodiment of the present application, the diagnostic accuracy of the model to be tested with respect to each of the multiple first test parameters may be as shown in Table 2 below.

[0102] Table 2

[0103]

[0104] In Table 2, test_log may represent the number of first server logs in the benchmark test dataset.

[0105] According to an embodiment of the present application, by comparing the diagnostic results of the model to be tested with the test model for any first test parameter of any first server log, the diagnostic accuracy of the model to be tested with respect to each of the multiple first test parameters can be determined, thereby indicating the diagnostic accuracy of the model to be tested with respect to the first test parameters based on the diagnostic accuracy, and used to determine whether the benchmark test results of the model to be tested meet the preset conditions, that is, to determine whether the diagnostic accuracy of the model to be tested meets the requirements. Furthermore, the diagnostic accuracy of the model to be tested can be determined based on the comparison of the diagnostic results of multiple first test parameters as core parameters.

[0106] Based on the above, benchmark testing is conducted using batch benchmark datasets. By comparing the diagnostic differences between the model under test and the test model, the diagnostic accuracy of the model under test for multiple first test parameters is tested. Benchmark testing is primarily designed to prevent significant declines in diagnostic accuracy due to serious logical defects introduced by changes to the model under test, such as software structure changes or diagnostic logic changes, thereby ensuring diagnostic reliability and accuracy.

[0107] According to an embodiment of the present application, the second server log for development requirements in the current period includes a server log for diagnostic requirements newly added in the current period and a server log for diagnostic errors newly added in the current period.

[0108] According to an embodiment of the present application, the second server log for development requirements in the current period is a second server log newly added in the current period relative to the previous period. Specifically, the server log for newly added diagnostic requirements in the current period may, for example, include the server log corresponding to a new model that the model under test needs to be adapted to; the server log for newly added diagnostic errors in the current period may, for example, include the log corresponding to rules that the model under test in the previous period did not cover, such as the log of diagnostic errors of the model under test published in the previous period.

[0109] For example, in the current period, the developer has sorted out 10 diagnostic error logs and 10 diagnostic requirement logs, so the second server log used for the development requirements of the current period can be the above 20 server logs.

[0110] According to an embodiment of the present application, since the second server log for development requirements in the current period includes the server log for new diagnostic requirements in the current period and the server log for new diagnostic errors in the current period, it is possible to detect whether the model to be tested meets the requirements for new diagnostic requirements and existing diagnostic errors.

[0111] According to an embodiment of the present disclosure, a baseline test data set for the current time period is input into the model to be tested, and multiple third diagnostic results for any second test parameters of multiple second server logs are obtained, including: multiple second server logs in the baseline test data set are input into the model to be tested, and the third diagnostic results for each of the multiple second test parameters of any second server log are obtained.

[0112] According to an embodiment of the present application, the model to be tested can be used to identify multiple parameters in the second server log, the multiple second test parameters are part of the multiple parameters in the second server log that the model to be tested can identify, and the multiple second test parameters are core parameters that the model to be tested can identify.

[0113] According to an embodiment of the present application, multiple second server logs in a baseline test dataset are input into the model under test to obtain diagnostic results for each of the multiple parameters in any second server log. Because the multiple second test parameters are core parameters, a third diagnostic result for each of the multiple second test parameters in any second server log can be obtained by selecting from the output of the model under test.

[0114] According to embodiments of the present application, by inputting a baseline test dataset into the model under test, a third diagnostic result can be obtained for each of the multiple second test parameters of the model under test for any second server log. Thus, based on the multiple third diagnostic results and the baseline test sample, a baseline test of the model under test can be performed, thereby facilitating determination of whether the model under test meets development requirements.

[0115] Figure 4 A flow chart for obtaining baseline test results according to an embodiment of the present application is shown.

[0116] According to an embodiment of the present application, the baseline test result may represent the diagnostic accuracy of the model to be tested for each of the plurality of second test parameters.

[0117] like Figure 4 As shown, the method 400 includes operations S410 to S430.

[0118] In operation S410 , for any second test parameter of any second server log, a baseline value for any second test parameter is compared with a third diagnosis result based on a baseline test sample to obtain a second comparison result for any second test parameter.

[0119] According to an embodiment of the present application, since the baseline test sample is used to test the diagnostic accuracy of the model to be tested with respect to development requirements, it is necessary to compare the baseline value of any second test parameter for any second server log with the third diagnostic result based on the baseline test sample to obtain a second comparison result of the baseline test sample and the model to be tested for the any second test parameter.

[0120] According to an embodiment of the present application, the plurality of second test parameters may include a server failure type and a baseline use case, and the baseline test sample includes a baseline value of the server failure type and a baseline value of the baseline use case for any second server log.

[0121] In one embodiment, the plurality of second test parameters may further include a serial number and a server type. Specifically, the baseline use case may represent a specific fault under a server fault type, such as a fault level (fault_level), fault analysis (reason), conclusion and suggestion (solution), and typical fault (comment).

[0122] According to an embodiment of the present application, the third diagnostic result of the model to be tested regarding the baseline use case of multiple second server logs and the baseline value of the baseline test sample regarding the baseline use case of multiple second server logs can be shown in Table 3 below.

[0123] Table 3

[0124]

[0125] In Table 3, for log 1, the diagnostic result of the model to be tested for the baseline use case is to give a conclusion suggestion. For example, if the diagnostic result is specifically to replace RAID card 1, the diagnostic result of the model to be tested for the baseline use case is inconsistent with the baseline value; for log 5, the diagnostic result of the model to be tested for the baseline use case is the level of hard disk failure. The baseline value of the baseline use case in the baseline test sample is a hidden danger. The diagnostic result of the model to be tested for the baseline use case is inconsistent with the baseline value. The model to be tested identifies the hard disk hidden danger in log 5 as a hard disk failure.

[0126] Among them, RAID card is Redundant Array of Independent Disks card; BBU is Baseband Unit.

[0127] According to an embodiment of the present application, taking the above Table 3 as an example, the second comparison results of multiple second test parameters of the model to be tested and the baseline test sample in Table 3 regarding multiple second server logs can be shown in the following Table 4.

[0128] Table 4

[0129]

[0130] According to an embodiment of the present application, a baseline test result of the model to be tested can be obtained based on the second comparison results of each of the multiple second test parameters of each second server log in the baseline test data set.

[0131] In operation S420 , for any second test parameter, a second number of second server logs whose second comparison results indicate that the baseline value is consistent with the third diagnosis result is determined.

[0132] According to an embodiment of the present application, the baseline test result may represent the diagnostic accuracy of the model to be tested for each of the plurality of second test parameters.

[0133] According to an embodiment of the present application, when the second comparison result indicates that the baseline value is consistent with the third diagnosis result, it can be determined that the diagnosis of a second test parameter of a second server log by the model to be tested is correct.

[0134] Specifically, for any second test parameter, the second number of second server logs whose second comparison results indicate that the baseline value is consistent with the third diagnostic result can be determined, that is, the number of first server logs whose diagnosis of the model to be tested is correct for any second test parameter can be determined.

[0135] In operation S430 , the diagnostic accuracy of the model to be tested for any second test parameter is obtained according to the second number and the number of second server logs in the baseline test data set.

[0136] According to an embodiment of the present application, based on the second number and the number of second server logs in the baseline test data set, the diagnostic accuracy of the model to be tested for any second test parameter can be obtained to determine whether the model to be tested meets the development requirements.

[0137] In one embodiment, a ratio of the second number to the number of second server logs in the baseline test data set may be calculated, and the ratio may represent the diagnostic accuracy of the test subject satisfying any second test parameter.

[0138] According to an embodiment of the present application, by comparing the third diagnostic result of the model to be tested for any first test parameter of any first server log with the baseline value for the first test parameter in the baseline test sample, the diagnostic accuracy of the model to be tested with respect to each of the multiple second test parameters can be determined, thereby indicating the diagnostic accuracy of the model to be tested with respect to the second test parameters based on the diagnostic accuracy, so as to determine whether the baseline test result of the model to be tested meets the preset conditions, that is, to determine whether the model to be tested meets the development requirements of the current period and the historical development requirements. Moreover, based on the comparison of the diagnostic results and baseline values ​​of the multiple second test parameters as core parameters, it can be determined whether the model to be tested meets the development requirements.

[0139] Figure 5 A schematic diagram of a process of performing a benchmark test and a baseline test on a model to be tested according to an embodiment of the present application is shown.

[0140] like Figure 5 As shown, based on the benchmark test data set, the model to be tested is benchmarked, specifically: the model to be tested outputs a first diagnostic result with respect to the benchmark test data set, the test model outputs a second diagnostic result with respect to the benchmark test data set, and by comparing the first diagnostic result and the second diagnostic result, a first comparison result is obtained, thereby obtaining a benchmark test result with respect to the model to be tested.

[0141] Based on the baseline test data set, a baseline test is performed on the model to be tested, specifically: the model to be tested outputs a third diagnostic result about the baseline test data set, and by comparing the third diagnostic result with the baseline test sample, a second comparison result is obtained, thereby obtaining a baseline test result about the model to be tested.

[0142] Figure 6 A schematic diagram of a benchmark test according to an embodiment of the present application is shown.

[0143] like Figure 6 As shown, the initial number of first server logs in the benchmark test data set can be 6000, that is, the first server logs in the benchmark test data set increase periodically on this basis. The periodic incremental logs specifically include: adding 1% of the newly added first server logs.

[0144] For example, the number of first-server logs in the benchmark test dataset increases dynamically with the total number of diagnostic system logs. The increase is a random sampling of 1% of the log increment. The number of first-server logs in the benchmark test dataset initially was 6,000, and has increased to 9,900 in the current period.

[0145] According to an embodiment of the present application, multiple first server logs in a benchmark test data set are input into a test model and also into a model to be tested. Based on the outputs of the test model and the model to be tested, a benchmark test comparison is performed for multiple first parameters to be tested, thereby obtaining the diagnostic accuracy of each of the multiple first parameters to be tested, specifically including: log type accuracy, fault type accuracy, machine model accuracy, serial number accuracy, and diagnostic identification accuracy. Based on this, a benchmark test result for the model to be tested can be obtained.

[0146] Figure 7 A schematic diagram of a baseline test according to an embodiment of the present application is shown.

[0147] like Figure 7 As shown, the initial number of second server logs in the baseline test data set can be 100, that is, the second server logs in the baseline test data set are increased period by period. The periodic incremental logs specifically include: adding the second server logs for the development needs of the current period.

[0148] For example, the number of second server logs in the baseline test dataset increases dynamically with the iteration of the model to be tested, with the increment in each cycle dynamically changing between 30 and 100. The number of second server logs in the baseline test dataset can be initially 100 and increases to 470 in the current period.

[0149] According to an embodiment of the present application, multiple second server logs from a baseline test dataset are input into the model under test. Based on the output of the model under test and the baseline test samples, a baseline test comparison is performed for multiple second test parameters. This allows the diagnostic accuracy of each of the multiple second test parameters to be obtained, specifically including: test case accuracy, fault type accuracy, machine model accuracy, and serial number accuracy. Based on this, a baseline test result for the model under test can be obtained.

[0150] According to an embodiment of the present application, the preset conditions include at least one of the following: among multiple first test parameters, the diagnostic accuracy of at least one first test parameter is less than a first preset threshold; among multiple second test parameters, the diagnostic accuracy of at least one second test parameter is less than a second preset threshold.

[0151] According to an embodiment of the present application, a first preset threshold is set for the diagnostic accuracy of the first test parameter, and a second preset threshold is set for the diagnostic accuracy of the second test parameter.

[0152] When the diagnostic accuracy of at least one first test parameter among multiple first test parameters is less than the first preset threshold or the diagnostic accuracy of at least one second test parameter among multiple second test parameters is less than the second preset threshold, it can be determined that there is a problem with the model to be tested.

[0153] Specifically, since the purposes of benchmark testing and baseline testing are different, the benchmark testing is to verify the overall quality of the entire model to be tested to ensure that the model to be tested does not have major bugs (errors), so it is sufficient to ensure that the diagnostic results of the first test parameter between the test model and the model to be tested are not too different. Therefore, for example, the first preset threshold can be set to 98%; the baseline test is to verify that the model to be tested meets the development requirements of the current period to ensure that the model to be tested can fully meet the development requirements of the current period. Therefore, the second preset threshold is set to 100%.

[0154] According to the embodiments of the present application, the setting of the preset conditions based on the first preset threshold and the second preset threshold can ensure that the diagnostic accuracy of the model to be tested meets the requirements and the model to be tested meets the development requirements.

[0155] According to an embodiment of the present application, the model to be tested in the current period is obtained by optimizing the model to be tested in the previous period using the development requirements for the current period.

[0156] According to an embodiment of the present application, the model to be tested in the previous period can represent the model to be tested in the previous version. For the current period, there are new development requirements, so it is necessary to optimize the model to be tested in the previous version based on the new development requirements. Thus, the optimized model to be tested is the model to be tested in the current period.

[0157] In one embodiment, while optimizing the model to be tested based on development requirements, the test model also needs to be optimized accordingly based on development requirements to ensure that the test model is compatible with the model to be tested.

[0158] Specifically, to implement benchmarking, a test model was developed alongside the model under test. This test model, similar to the principles of the diagnostic system, uses in-band and out-of-band text logs to perform data standardization, rule matching, and fault diagnosis, ultimately outputting key data such as fault type, machine model, and serial number. Unlike the diagnostic system's stringent requirements for development language, multi-platform functionality, code security, parameter output, and algorithmic models, the test model utilizes a scripting language, supports fixed output parameters, and regularly maintains diagnostic rules, enabling rapid development and iteration.

[0159] On this basis, the test model can be maintained, updated and iterated based on the development requirements of the model to be tested and the benchmark test results.

[0160] For example, during the benchmark test process, the reliability of the test model will also be reversely verified. For example, the diagnostic accuracy of the server failure type of a certain version of the model to be tested dropped to 97%. After testing, it was found that the model to be tested had newly added expansion card type adaptation. This adaptation was introduced to adapt to the new model, but the test model is not currently included. Therefore, the test model can be reversely promoted to adapt to new rules and new models.

[0161] At the same time, the test model has certain development advantages over the model to be tested: compared with the model to be tested, the test model has simpler functions, faster development and iteration, and lower fault tolerance.

[0162] Specifically, it includes several aspects: the test model does not need to include a complex algorithm model, but only needs to focus on developing and maintaining the corresponding diagnostic rules. The logic is relatively simple and the adaptation is fast. On the contrary, the model to be tested requires a lot of time to maintain and test the algorithm model; there is no need to pay attention to complex parameters. The test model only focuses on five key parameters: model, serial number, fault type, diagnostic identifier and log type. On the contrary, the model to be tested currently has dozens of parameters, and the identification and development of these parameters require a lot of time to develop and maintain; there is no need to pay attention to log decompression. After the business test model is executed, the decompressed log is directly passed to the test model. There is no need to pay attention to a variety of log types and log encryption methods. On the contrary, the model to be tested needs to identify and decompress various standard and non-standard formats; there is no need to pay attention to performance. The test diagnosis architecture and instructions are relatively free. There is no need to pay attention to CPU usage time and data reading and writing. On the contrary, the model to be tested needs to pay attention to performance parameters such as CPU, memory usage time, data reading and writing time.

[0163] According to the embodiments of the present application, since the model to be tested in the current time period is obtained by optimizing the model to be tested in the previous time period using the development requirements for the current time period, the model to be tested in the current time period is tested to test whether the diagnostic accuracy of the optimized model to be tested meets the requirements and whether it meets the development requirements, that is, to test whether the optimization of the model to be tested is successful.

[0164] According to an embodiment of the present application, the above-mentioned model testing method also includes: when the benchmark test results meet the preset conditions, determining that the diagnostic accuracy of the model to be tested is not affected by the model optimization of the current period; when the baseline test results meet the preset conditions, determining that the model to be tested meets the development requirements of the current period and the development requirements of the previous period.

[0165] According to an embodiment of the present application, when the benchmark test results meet the preset conditions, that is, the diagnostic accuracy of multiple first test parameters is greater than or equal to the first preset threshold, it can be determined that the diagnostic accuracy of the model to be tested is not affected by the model optimization in the current period.

[0166] For example, when testing a certain version of the model under test, the accuracy of server fault types in the benchmark results dropped to 91%. This result was then fed back to the developer of the model under test. Analysis revealed that this issue was caused by a newly added function for obtaining component SNs (serial numbers). The addition of this function to the model under test also resulted in significant adjustments to the logic for obtaining component types, which led to defects in the component type acquisition logic. This suggests that model optimization during this period may be affecting the diagnostic accuracy of the model under test.

[0167] Therefore, when the benchmark test results do not meet the preset conditions, the model to be tested should be adjusted in a timely manner to ensure that the diagnostic accuracy of the model to be tested is not affected by the model optimization in the current period, and to avoid diagnostic accuracy problems after the model to be tested is released.

[0168] According to an embodiment of the present application, when the baseline test results meet the preset conditions, that is, the diagnostic accuracy of multiple second test parameters is equal to the second preset threshold, it can be determined that the model to be tested meets the development requirements of the current period and the development requirements of the previous period.

[0169] According to the embodiments of the present application, the benchmark test and the baseline test are used to test the model to be tested from different aspects, so that based on the benchmark test results meeting the preset conditions, it can be determined that the diagnostic accuracy of the model to be tested is not affected by the model optimization of the current period; based on the baseline test results meeting the preset conditions, it can be determined that the model to be tested meets the development requirements of the current period and the development requirements of the previous period.

[0170] In one embodiment, when testing a certain version of the model to be tested, the accuracy of the server model in the benchmark test results dropped to 85% and was then fed back to the developer of the model to be tested. After analysis, it was found that this problem was caused by a new FRU (Field Replaceable Unit) log acquisition model introduced by the model to be tested. This FRU format only supports specific models and the output format is non-standard. The model to be tested does not strictly limit this method, which leads to the introduction of model acquisition adaptation problems.

[0171] Based on the model testing method of the present application, based on batch benchmark test data sets, through benchmark testing, before the model to be tested is released, problems introduced in the development of the model to be tested can be quickly discovered, thereby ensuring the diagnostic reliability and accuracy of the model to be tested; based on the incremental baseline test data set, through baseline testing, before the model to be tested is released, the version iteration requirements of the model to be tested can be quickly confirmed to be realized, and historical version iterations can be ensured to be unaffected, thereby ensuring the reliability of version iterations, that is, ensuring that the model to be tested meets the development requirements of the current period before release, and also ensuring that the model to be tested meets historical development requirements.

[0172] Based on the above model testing method, this application also provides a model testing device. Figure 8 The device is described in detail.

[0173] Figure 8 The figure shows a structural block diagram of a model testing device according to an embodiment of the present application.

[0174] like Figure 8 As shown, the model testing device 800 of this embodiment includes a first input module 810 , a first obtaining module 820 , a second input module 830 , a second obtaining module 840 and a first determining module 850 .

[0175] First input module 810 is used to input a benchmark test data set for the current time period into the test model and the model to be tested, respectively, to obtain multiple first diagnostic results of the test model for any first test parameter of multiple first server logs and multiple second diagnostic results of the model to be tested for any first test parameter of multiple first server logs, wherein the benchmark test data set includes the first server logs of the previous time period and the first server logs obtained by sampling the newly added first server logs of the current time period according to a preset ratio. In one embodiment, first input module 810 can be used to perform operation S210 described above, which will not be repeated here.

[0176] The first obtaining module 820 is used to obtain a benchmark test result based on the number of identical diagnostic results for the same first server log in the plurality of first diagnostic results and the plurality of second diagnostic results. In one embodiment, the first obtaining module 820 can be used to perform the operation S220 described above, which will not be repeated here.

[0177] Second input module 830 is configured to input a baseline test dataset for the current period into the model under test, thereby obtaining multiple third diagnostic results for any second test parameter of multiple second server logs, where the baseline test dataset includes the second server logs for the previous period and the second server logs for development requirements for the current period. In one embodiment, second input module 830 can be configured to perform operation S230 described above, which will not be further described here.

[0178] The second obtaining module 840 is used to obtain the baseline test result based on the multiple baseline values ​​for any second test parameter in the baseline test sample and the number of identical values ​​for the same second server log in the multiple third diagnostic results. In one embodiment, the second obtaining module 840 can be used to perform the operation S240 described above, which will not be repeated here.

[0179] The first determination module 850 is used to determine that the model under test has an abnormality when the benchmark test result and the baseline test result do not meet the preset conditions. In one embodiment, the first determination module 850 can be used to perform the operation S250 described above, which will not be repeated here.

[0180] According to an embodiment of the present application, the first input module 810 includes a first input unit and a second input unit.

[0181] The first input unit is used to input multiple first server logs in the benchmark test data set into the test model to obtain first diagnostic results for each of the multiple first test parameters of any first server log.

[0182] The second input unit is used to input multiple first server logs in the benchmark test data set into the model to be tested, and obtain second diagnostic results for each of the multiple first test parameters of any first server log.

[0183] According to an embodiment of the present application, the benchmark test result represents the diagnostic accuracy of the model to be tested for each of the multiple first test parameters; the first obtaining module 820 includes a first comparing unit, a first determining unit and a first obtaining unit.

[0184] The first comparison unit is configured to compare a first diagnostic result of any first test parameter with a second diagnostic result of any first test parameter in any first server log to obtain a first comparison result of the first test parameter.

[0185] The first determining unit is configured to determine, for any first test parameter, a first number of first server logs whose first comparison result indicates that the first diagnosis result is consistent with the second diagnosis result.

[0186] The first obtaining unit is configured to obtain a diagnostic accuracy rate of the model to be tested for any first test parameter according to the first number and the number of first server logs in the benchmark test data set.

[0187] According to an embodiment of the present application, the second input module 830 includes a third input unit.

[0188] The third input unit is used to input multiple second server logs in the baseline test data set into the model to be tested, and obtain third diagnostic results for each of the multiple second test parameters of any second server log.

[0189] According to an embodiment of the present application, the baseline test result represents the diagnostic accuracy of the model to be tested for each of the multiple second test parameters; the second obtaining module 840 includes a second comparison unit, a second determination unit and a second obtaining unit.

[0190] The second comparison unit is used to compare the baseline value of any second test parameter of any second server log with the third diagnostic result based on the baseline test sample to obtain a second comparison result of any second test parameter.

[0191] The second determining unit is configured to determine, for any second test parameter, a second number of second server logs whose second comparison results indicate that the baseline value is consistent with the third diagnosis result.

[0192] The second obtaining unit is configured to obtain the diagnostic accuracy of the model to be tested for any second test parameter according to the second number and the number of second server logs in the baseline test data set.

[0193] According to an embodiment of the present application, the model testing device 800 further includes a second determination module.

[0194] The second determining module is configured to determine the model to be tested as the model to be released if the benchmark test result and the baseline test result meet preset conditions.

[0195] According to an embodiment of the present application, the model testing device 800 further includes a third determination module and a fourth determination module.

[0196] The third determination module is used to determine that the diagnostic accuracy of the model to be tested is not affected by the model optimization in the current period when the benchmark test result meets the preset conditions.

[0197] The fourth determination module is used to determine whether the model to be tested meets the development requirements of the current period and the development requirements of the previous period when the baseline test results meet the preset conditions.

[0198] According to embodiments of the present application, any multiple modules among the first input module 810, the first obtaining module 820, the second input module 830, the second obtaining module 840, and the first determination module 850 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the first input module 810, the first obtaining module 820, the second input module 830, the second obtaining module 840, and the first determination module 850 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the first input module 810 , the first obtaining module 820 , the second input module 830 , the second obtaining module 840 and the first determining module 850 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0199] Figure 9 A block diagram of an electronic device suitable for implementing a model testing method according to an embodiment of the present application is shown.

[0200] like Figure 9 As shown, an electronic device 900 according to an embodiment of the present application includes a processor 901, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 902 or programs loaded from a storage unit 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.

[0201] Various programs and data required for the operation of the electronic device 900 are stored in the RAM 903. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in the one or more memories.

[0202] According to an embodiment of the present application, electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to bus 904. Electronic device 900 may also include one or more of the following components connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 908 including a hard disk; and a communication section 909 including a network interface card such as a LAN card or modem. Communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 910 as needed, so that computer programs read from the removable media can be installed into storage section 908 as needed.

[0203] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.

[0204] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.

[0205] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the model testing method provided in the embodiments of the present application.

[0206] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the processor 901 executes the computer program. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0207] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0208] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0209] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0210] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0211] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

[0212] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.

Claims

1. A model testing method, characterized in that: The method comprises: Inputting a benchmark test data set for the current period into a test model and a model to be tested, respectively, to obtain a plurality of first diagnostic results of the test model for any first test parameter of a plurality of first server logs and a plurality of second diagnostic results of the model to be tested for any first test parameter of the plurality of first server logs, wherein the benchmark test data set includes the first server logs of the previous period and the first server logs obtained by sampling the newly added first server logs of the current period according to a preset ratio; Obtaining a benchmark test result based on the number of identical diagnostic results for the same first server log in the plurality of first diagnostic results and the plurality of second diagnostic results; Inputting a baseline test data set for the current period into the model to be tested, and obtaining a plurality of third diagnostic results for any second test parameter of a plurality of second server logs, wherein the baseline test data set includes the second server logs for the previous period and the second server logs for development requirements of the current period; Obtaining a baseline test result based on a plurality of baseline values ​​for any one of the second test parameters in the baseline test sample and a number of identical values ​​for the same second server log in the plurality of third diagnostic results; When the benchmark test result and the baseline test result do not meet a preset condition, it is determined that an abnormality exists in the model to be tested.

2. The method according to claim 1, characterized in that Inputting the benchmark test data set for the current period into the test model and the model to be tested respectively, and obtaining a plurality of first diagnostic results of the test model for any first test parameter of a plurality of first server logs and a plurality of second diagnostic results of the model to be tested for any first test parameter of the plurality of first server logs, includes: Inputting a plurality of first server logs in the benchmark test data set into the test model to obtain a first diagnostic result for each of a plurality of first test parameters of any first server log; A plurality of first server logs in the benchmark test data set are input into the model to be tested, and a second diagnostic result of each of the plurality of first test parameters for any first server log is obtained.

3. The method according to claim 2, characterized in that The benchmark test result represents the diagnostic accuracy of the model to be tested for each of the plurality of first test parameters; and obtaining the benchmark test result based on the number of identical diagnostic results for the same first server log in the plurality of first diagnostic results and the plurality of second diagnostic results includes: For any first test parameter of any first server log, compare the first diagnostic result of the any first test parameter with the second diagnostic result to obtain a first comparison result of the any first test parameter; For any first test parameter, determining a first number of first server logs for which a first comparison result indicates that the first diagnostic result is consistent with the second diagnostic result; The diagnostic accuracy of the model to be tested for any one of the first test parameters is obtained according to the first number and the number of first server logs in the benchmark test data set.

4. The method according to claim 3, characterized in that The plurality of first test parameters include log type, server failure type and server type; The test model is trained using log types, server fault types, and server types involved in server logs, and the test model is used to identify multiple first test parameters in a first server log.

5. The method according to claim 1, wherein Inputting the baseline test data set for the current period into the model to be tested to obtain multiple third diagnostic results for any second test parameter of multiple second server logs includes: A plurality of second server logs in the baseline test data set are input into the model to be tested to obtain a third diagnostic result for each of the plurality of second test parameters of any second server log.

6. The method according to claim 5, characterized in that The baseline test result represents the diagnostic accuracy of the model to be tested for each of the plurality of second test parameters; the baseline test result is obtained based on the plurality of baseline values ​​for any of the second test parameters in the baseline test sample and the number of identical values ​​for the same second server log in the plurality of third diagnostic results, including: For any second test parameter of any second server log, comparing a baseline value for the any second test parameter with the third diagnostic result based on the baseline test sample to obtain a second comparison result for the any second test parameter; For any second test parameter, determining a second number of second server logs for which the second comparison result indicates that the baseline value is consistent with the third diagnosis result; The diagnostic accuracy of the model to be tested for any second test parameter is obtained according to the second number and the number of second server logs in the baseline test data set.

7. The method according to claim 5, characterized in that The plurality of second test parameters include a server failure type and a baseline use case, and the baseline test sample includes a baseline value of the server failure type and a baseline value of the baseline use case for any second server log.

8. The method according to claim 6, characterized in that The preset condition includes at least one of the following: Among the multiple first test parameters, the diagnostic accuracy of at least one first test parameter is less than a first preset threshold; Among the multiple second test parameters, the diagnostic accuracy of at least one second test parameter is less than a second preset threshold.

9. The method according to claim 1, characterized in that The method further comprises: In a case where the benchmark test result and the baseline test result meet the preset condition, the model to be tested is determined as a model to be released.

10. The method according to any one of claims 1 to 9, characterized in that The model to be tested in the current period is obtained by optimizing the model to be tested in the previous period using the development requirements for the current period.

11. The method according to claim 10, characterized in that The method further comprises: If the benchmark test result satisfies the preset condition, determining that the diagnostic accuracy of the model to be tested is not affected by the model optimization in the current period; In the case where the baseline test result meets the preset condition, it is determined that the model to be tested meets the development requirements of the current period and the development requirements of the previous period.

12. The method according to claim 1, characterized in that The second server log for development requirements in the current period includes a server log for newly added diagnostic requirements in the current period and a server log for newly added diagnostic errors in the current period.

13. A model testing device, characterized in that: The device comprises: A first input module is configured to input a benchmark test data set for the current time period into a test model and a model to be tested, respectively, to obtain a plurality of first diagnostic results of the test model for any first test parameter of a plurality of first server logs and a plurality of second diagnostic results of the model to be tested for any first test parameter of the plurality of first server logs, wherein the benchmark test data set includes the first server logs of the previous time period and the first server logs obtained by sampling the newly added first server logs of the current time period according to a preset ratio; A first obtaining module, configured to obtain a benchmark test result based on the number of identical diagnostic results for the same first server log in the plurality of first diagnostic results and the plurality of second diagnostic results; a second input module, configured to input a baseline test data set for a current period into the model to be tested, and obtain a plurality of third diagnostic results for any second test parameter of a plurality of second server logs, wherein the baseline test data set includes the second server logs for a previous period and the second server logs for development requirements of the current period; a second obtaining module, configured to obtain a baseline test result based on a plurality of baseline values ​​for any one of the second test parameters in the baseline test sample and a number of identical values ​​for the same second server log in the plurality of third diagnostic results; The first determining module is configured to determine that an abnormality exists in the model to be tested when the benchmark test result and the baseline test result do not meet a preset condition.

14. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.

15. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Test result judgment method and test result judgment device

    CN113778874A

  • Server fault diagnosis method and device, electronic equipment and storage medium

    CN114691403A

  • Training method, device and equipment of fault detection model under micro-service architecture

    CN116561635A

  • Server fault diagnosis method, product, computer equipment and storage medium

    CN118211170A

  • Immunochromatography rapid diagnostic kit and method of detection using the same

    KR1020180138451A