Diagnostic methods for root causes of test failure

By combining test environment configuration similarity and log semantic feature analysis, the system automates the diagnosis of the root causes of server test failures, solving the problems of low efficiency and high false positives in existing technologies, and achieving efficient and accurate root cause localization and solution matching.

CN120994451BActive Publication Date: 2026-01-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511508114.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-30
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

In existing technologies, the root cause diagnosis of server test failures is inefficient and has a high probability of misjudgment. Manual analysis relies on technical experience, and log analysis tools ignore the influence of the test environment, resulting in poor data comprehensiveness and difficulty in accurately identifying implicit relationships between log texts.

Method used

By capturing test failure signals from the server test machine, obtaining multi-source log information and test environment configuration data, using the weighted cosine similarity algorithm to calculate the test environment configuration similarity between test cases and historical cases, and combining it with the BERT model to perform log semantic feature analysis, the root cause of test failure is determined by comprehensive scoring, and corresponding solutions are matched.

Benefits of technology

It improves the efficiency and accuracy of root cause diagnosis of test failures, shortens the overall testing cycle, reduces reliance on human experience, and improves the accuracy and comprehensiveness of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994451B_ABST
    Figure CN120994451B_ABST
Patent Text Reader

Abstract

This invention discloses a method for diagnosing the root causes of test failures, relating to the field of electronic digital data processing technology. After capturing a test failure signal, the method combines multi-source log information from test cases with test environment configuration data for comprehensive analysis. This fully considers the impact of environment configuration on testing and fully mines the information in the logs to obtain more accurate diagnostic results. Then, based on the root cause of the test failure, a corresponding solution is matched to resolve the test failure problem. Therefore, this method can solve the technical problems in related technologies, such as low efficiency and high probability of misjudgment due to manual analysis, neglect of the impact of the test environment on test failures, and difficulty in identifying implicit connections between text in logs, making it difficult to use the analysis results for accurate root cause diagnosis. This method improves the efficiency of diagnosing the root causes of test failures, increases diagnostic accuracy, and shortens the overall test cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a method for diagnosing the root causes of test failures. Background Technology

[0002] In related technologies, manual or semi-manual analysis methods are inefficient when diagnosing root causes of server test failures. They rely heavily on the technical experience of technicians, are highly subjective, and have a high probability of misdiagnosis. In semi-manual analysis, the method of using log analysis tools for preliminary analysis before root cause diagnosis ignores the impact of the test environment on test failures, resulting in poor data comprehensiveness. Furthermore, log analysis tools struggle to identify implicit relationships between text in logs, making the analysis results unsuitable for accurate root cause diagnosis. Improvements are urgently needed. Summary of the Invention

[0003] This invention provides a method for diagnosing the root causes of test failures, which addresses the technical problems in related technologies, such as low efficiency and high probability of misjudgment in manual analysis, neglect of the impact of the test environment on test failures when using log analysis tools for preliminary analysis, poor data comprehensiveness, and difficulty in identifying implicit relationships between text in logs, making it difficult to use the analysis results for accurate root cause diagnosis.

[0004] This invention provides a method for diagnosing the root cause of test failures, comprising: capturing test failure signals from at least one server test machine, determining test cases based on the test failure signals, and obtaining multi-source log information and test environment configuration data of the test cases; calculating the test environment configuration similarity between the test cases and historical cases based on the test environment configuration data, and analyzing the multi-source log information to obtain corresponding log semantic features; calculating a comprehensive score for the test cases based on the test environment configuration similarity and log semantic features, determining the root cause of the test failure based on the comprehensive score, matching the corresponding solution based on the root cause of the test failure, and pushing the solution to at least one server test machine.

[0005] The present invention also provides a diagnostic device for the root cause of test failure, comprising: an acquisition module, configured to capture test failure signals from at least one server test machine and determine test cases based on the test failure signals, thereby acquiring multi-source log information and test environment configuration data of the test cases; an analysis module, configured to calculate the test environment configuration similarity between the test cases and historical cases based on the test environment configuration data, and analyze the multi-source log information to obtain corresponding log semantic features; and a diagnostic module, configured to calculate a comprehensive score for the test cases based on the test environment configuration similarity and log semantic features, thereby determining the root cause of the test failure based on the comprehensive score, matching corresponding solutions based on the root cause of the test failure, and pushing the solutions to at least one server test machine.

[0006] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described methods for diagnosing the root cause of test failure when executing the computer program.

[0007] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described methods for diagnosing the root cause of test failure.

[0008] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for diagnosing the root cause of test failure.

[0009] This invention, upon detecting a test failure signal, identifies the corresponding test case and performs a comprehensive analysis combining multi-source log information and test environment configuration data to diagnose the root cause of the failure. The test environment configuration data calculates the similarity between the test case and historical test cases, while the multi-source log information analyzes implicit textual relationships. By combining test environment configuration similarity and log semantic features, the root cause of the test case failure is determined. This fully considers the impact of environment configuration on testing and fully mines information from the logs to obtain more accurate diagnostic results. Based on the root cause, a corresponding solution is matched to resolve the test failure. Therefore, this invention addresses the technical problems of low efficiency and high error rate in manual analysis, the neglect of the test environment's impact on test failures when using log analysis tools for preliminary analysis, poor data comprehensiveness, and the difficulty of identifying implicit textual relationships in logs, making accurate root cause diagnosis difficult. This invention improves the efficiency and accuracy of root cause diagnosis, and shortens the overall testing cycle. Attached Figure Description

[0010] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating a diagnostic method for testing root causes of failure according to an embodiment of the present invention;

[0012] Figure 2 A flowchart illustrating a method for diagnosing the root cause of test failure according to an embodiment of the present invention;

[0013] Figure 3 This is a schematic diagram of a diagnostic device for testing the root causes of failures according to an embodiment of the present invention. Detailed Implementation

[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0015] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0016] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0017] Understandably, during server testing, accurate diagnosis of the root cause of a test case failure is crucial for ensuring testing efficiency and server quality when a test case results in a failure. In server testing, when a test case fails, it is necessary to comprehensively analyze and locate the root cause of the failure by integrating multi-dimensional log information, including test script logs, OS logs, and BMC logs.

[0018] Currently, the relevant technologies mainly use two analysis methods to diagnose the root causes of test failures: one is manual analysis, where testers need to check various logs line by line, and troubleshoot problems by manually comparing and judging based on experience. This method not only consumes a lot of time and manpower, but is also prone to deviations in analysis results due to differences in the experience of different personnel; the other is semi-automated analysis, which uses existing log analysis systems for preliminary screening and then manual determination of the final root cause.

[0019] However, relying on existing log analysis systems for initial screening has significant limitations. Log analysis tools focus solely on matching error text, ignoring the impact of test environment configurations such as hardware model and software version on errors. This results in solutions with extremely poor universality, making it difficult to handle diverse scenarios. Rule engines based on keyword matching cannot identify the contextual semantic relationships in logs, ignoring potential implicit connections between errors and preceding / following entries in server logs. This can lead to biased analysis and affect the accuracy of root cause localization. Furthermore, historical solutions under similar configurations are not effectively mined and reused. When the same or similar problems occur, repeated analysis is still required, severely reducing efficiency.

[0020] To overcome the limitations of related technologies, this invention proposes a diagnostic method for test failure root causes that integrates configuration similarity and semantic analysis, enabling rapid and accurate diagnosis of test failure root causes and providing efficient support for server testing in complex multi-configuration scenarios.

[0021] like Figure 1 As shown, an embodiment of the present invention provides a diagnostic method for the root cause of test failure, comprising the following steps:

[0022] In step S101, at least one test failure signal from a server test machine is captured, and test cases are determined based on the test failure signal to obtain multi-source log information and test environment configuration data for the test cases.

[0023] A test case is a specific set of test conditions, steps, and expected results used to verify whether a certain function of the software works normally as designed.

[0024] Multi-source log information refers to log data generated by different software components, services, or devices during testing and system operation, which may vary in format and content. Multi-source log information can consist of multiple parts, such as web servers, application servers, databases, caching services, message queues, etc. Each part generates its own log file.

[0025] The test environment configuration data includes all relevant settings and parameters of the server environment during the test. This may include the operating system version, software dependency library version numbers, database connection strings, memory and CPU allocation, network configuration, environment variables, etc.

[0026] In actual implementation, embodiments of the present invention can communicate with a server to obtain server log information and test environment configuration data in real time or periodically. To obtain more targeted data, embodiments of the present invention can also retrieve the corresponding test cases, along with multi-source log information and test environment configuration data from the server after capturing a test failure signal from the server test set.

[0027] For example, in embodiments of the present invention, the server test machine that has triggered the test failure signal can be obtained from the test management platform, server configuration data (various component models, quantities, OS, BMC version, etc.) can be pulled from the MySQL database of the test platform, and test script logs, OS logs, and BMC logs can be collected in real time and stored in the log pool.

[0028] In step S102, the similarity of the test environment configuration between the test cases and historical cases is calculated based on the test environment configuration data, and the multi-source log information is analyzed to obtain the corresponding log semantic features.

[0029] Understandably, many test failures are not due to code logic errors, but rather to incorrect or inconsistent environment configurations. For example, code may run correctly in a production environment but fail in a test environment, likely because the database version in the test environment is outdated or a configuration item is set incorrectly.

[0030] Logs from a single source may only reflect a localized aspect of the problem. This invention, by correlating log information from multiple sources to form a global view, facilitates a more accurate diagnosis of the root causes of test failures.

[0031] As one possible implementation method, embodiments of the present invention can utilize algorithms such as weighted cosine similarity to calculate the similarity between the current test environment configuration data and the test environment configuration data of historical cases. Embodiments of the present invention can also preprocess multi-source log information and perform semantic encoding on the preprocessed multi-source log information to analyze semantics and extract log semantic features, so as to comprehensively diagnose the root cause of test failure based on log semantic features and test environment configuration similarity in subsequent processes.

[0032] Optionally, in one embodiment of the present invention, analyzing multi-source log information to obtain corresponding log semantic features includes: cleaning multi-source log information and segmenting the cleaned multi-source log information using a pre-built log domain dictionary to obtain multiple log fragments; calculating the error type probability distribution based on the multiple log fragments; and extracting error events based on the error type probability distribution to obtain log semantic features.

[0033] Furthermore, embodiments of the present invention can preprocess multi-source log information:

[0034] ① Log cleaning: Filtering out noise such as timestamps and IP addresses from logs;

[0035] ②Word segmentation: Enhance segmentation using a log domain dictionary.

[0036] Semantic encoding is performed on the preprocessed multi-source log information: log fragments are input into the fine-tuned BERT model, and the error type probability distribution is output.

[0037] Event graph construction: Extract error events and generate directed edges by timestamp; merge duplicate events (such as multiple consecutive "memory allocation failures").

[0038] To achieve correlation analysis of multi-source log information, embodiments of this invention can statistically analyze the frequency of various error types over a period of time, forming a probability distribution, i.e., an error type probability distribution. Error types whose probability suddenly increases significantly are considered abnormal and important. Then, embodiments of this invention can extract each specific log record of these error types as an independent error event.

[0039] This allows the embodiments of the present invention to automatically focus on the most likely relevant, sudden error event from thousands of logs in the current test failure.

[0040] Based on the above steps, the embodiments of the present invention can extract features from log text that represent its core meaning and content, namely log semantic features.

[0041] Optionally, in one embodiment of the present invention, after extracting error events based on the error type probability distribution, the method further includes: generating corresponding directed edges based on the timestamps of the error events; constructing an event graph using the directed edges, and calculating semantic similarity based on the event graph.

[0042] In actual implementation, this embodiment of the invention can traverse all error events. For any two events (event A and event B), if the timestamp t_A of A is earlier than the timestamp t_B of B, and the time difference (t_B - t_A) is less than a preset threshold (e.g., 5 minutes), then a directed edge is created between them, pointing from A to B (i.e., A -> B). This connects discrete event points into a potential causal chain with a temporal order.

[0043] Furthermore, in this embodiment of the invention, all error events can be treated as nodes and all directed edges as edges, and they can be combined into a directed graph, namely an event graph, to show the potential dependencies and influence relationships between faults.

[0044] This invention can determine the similarity between the log semantic features corresponding to multi-source log information and the error events in the event graph within the context of the event graph's topological structure. This transforms error events in multi-source logs (test scripts, OS, BMC) into a weighted directed graph, where nodes represent error types and edges represent time or logical relationships (e.g., "PCIe error → driver loading failure").

[0045] In addition, embodiments of the present invention can also improve the accuracy of root cause localization by matching common subgraphs in the current error chain and historical cases using graph similarity algorithms.

[0046] The embodiments of the present invention can combine time sequence and semantic association to jointly infer the root cause, effectively avoiding false alarms caused by single-dimensional analysis (only looking at time or only looking at keywords), making the location results more reliable.

[0047] Optionally, in one embodiment of the present invention, before segmenting and cleaning the multi-source log information using a pre-built log domain dictionary, the method further includes: defining the hierarchical relationship of server components; and constructing a log domain dictionary based on the hierarchical relationship. The log domain dictionary can be a dictionary customized for a specific system or business domain, containing key terms and concepts, including vocabulary specific to server testing, to help the analysis algorithm more accurately understand log content and identify which logs are related to core business or critical services, thereby improving the accuracy of the analysis.

[0048] In practical implementation, this invention can describe the architecture of the entire server application system in a structured manner, clarifying the affiliation and dependencies of each component in the system, i.e., defining hierarchical relationships, and using key entities in the hierarchical relationships (such as service names, module names, API endpoints, and key error codes) as terms to construct a dictionary specific to this business system. This ensures that subsequent log analysis algorithms are no longer general text processing but rather domain-specific text processing. When segmenting and identifying logs, the algorithm can prioritize identifying these domain-specific terms, greatly improving the accuracy of log parsing and feature extraction.

[0049] Optionally, in one embodiment of the present invention, extracting error events based on error type probability distribution to obtain log semantic features includes: importing multiple historical test cases that meet preset case conditions; constructing a pre-trained model based on the multiple historical test cases to calculate the error type probability distribution using the pre-trained model.

[0050] In other embodiments, a large number of labeled historical test case logs can also be collected. These labels can be error types such as "database error," "network timeout," and "memory overflow." This data is used to fine-tune models such as BERT. BERT is a pre-trained language model capable of deeply understanding contextual semantics. Unlike traditional methods that only match keywords, it can understand high semantic similarity across different contexts.

[0051] Through fine-tuning, embodiments of the present invention can transform the BERT model into a model more suited to the needs of test failure root cause analysis. The fine-tuning process allows the model to learn which word combinations and grammatical structures correspond to which specific error types within a given server application context.

[0052] In this embodiment of the invention, the log domain dictionary enables a more accurate understanding of log content and reduces misinterpretations. Based on NLP models such as BERT, deep semantic analysis is performed on unstructured log text to identify error types (such as "memory leak" and "network timeout") and their causal chains (such as "service A crash → service B timeout"), overcoming the limitations of traditional keyword matching.

[0053] Optionally, in one embodiment of the present invention, calculating the test environment configuration similarity between test cases and historical cases based on test environment configuration data includes: extracting numerical features from the test environment configuration data and normalizing the numerical features to obtain numerical data; performing variable transformation on the categorical features in the test environment configuration data to obtain coded data; calculating the cosine similarity between the test cases and multiple historical test cases based on the numerical data and coded data, and obtaining the historical cases based on the cosine similarity.

[0054] When searching for similar cases in historical test cases based on the current test environment configuration data, this embodiment of the invention can first perform feature vectorization processing to normalize numerical features (such as memory capacity) to eliminate biases caused by units and numerical ranges. Next, this embodiment of the invention can perform One-Hot encoding on categorical features (such as component model and CPU model) to transform categorical relationships into spatial relationships. Finally, cosine similarity is calculated by concatenating the normalized numerical features and the One-Hot encoded categorical features into a long, uniform feature vector, and then the cosine similarity between the current environment vector and the environment vectors of all historical cases is calculated.

[0055] The embodiments of the present invention can employ a weighted cosine similarity algorithm, which combines structured data such as hardware model (e.g., CPU, memory) and software version (e.g., operating system, driver) to dynamically adjust the weights of each configuration feature (e.g., CPU model weight 0.6, memory weight 0.3) to ensure the accuracy of similarity calculation.

[0056] Optionally, in one embodiment of the present invention, obtaining historical cases based on cosine similarity includes: determining whether there is at least one case in the historical test cases with a cosine similarity greater than a preset similarity threshold; if there is at least one case, obtaining historical cases based on cosine similarity; if there is no at least one case, verifying the calculation validity of cosine similarity using historical test cases; if the calculation validity does not meet the preset validity condition, generating an error reminder; otherwise, generating a new case reminder.

[0057] In this embodiment of the invention, a similarity threshold standard can be set, such as 0.7. When the similarity is greater than the similarity threshold, the historical test case can be retained. If only one historical test case has a similarity greater than the similarity threshold, then that historical test case is the historical case similar to the test case corresponding to the current test environment. If multiple historical test cases have a similarity greater than the similarity threshold, then the multiple historical test cases can be sorted according to the similarity, and the historical test case with the highest similarity can be selected as the historical case.

[0058] If there are no historical test cases similar to the test cases corresponding to the current test environment, this embodiment of the invention can call multiple historical test cases with known similarity to each other to perform similarity calculation, in order to determine whether there are any errors in the similarity calculation process, and to verify the validity of the cosine similarity calculation.

[0059] If the cosine similarity calculation is valid, it means that the test case corresponding to the current test environment is a new test case that has failed. In this embodiment of the invention, a new case reminder can be generated, and after the user processes the test case, the test case can be stored as a historical test case.

[0060] If the calculation of cosine similarity is invalid, a corresponding error message can be generated to remind the user to check for errors in a timely manner.

[0061] By setting a similarity threshold, cases with a certain level of similarity can be selected from historical test cases. These cases are then sorted by similarity, and the historical test case with the highest similarity is selected as the first historical case. In the absence of highly similar cases, this embodiment of the invention can automatically verify the reasons for low similarity to facilitate the learning and optimization process by inputting new cases, or to handle faults and avoid affecting subsequent analysis.

[0062] In step S103, the test case is calculated by combining the test environment configuration similarity and log semantic features to determine the root cause of the test case failure based on the comprehensive score, and the corresponding solution is matched based on the root cause of the test failure and pushed to at least one server test machine.

[0063] The embodiments of the present invention can perform weighted fusion of test environment configuration similarity and log semantic features to obtain a comprehensive score for test cases, thereby determining the root cause of test failure, and using test environment configuration similarity and test failure root cause to match corresponding solutions, so that the solution can take into account the impact of environment configuration and increase the probability of error repair.

[0064] Optionally, in one embodiment of the present invention, the comprehensive score of the test case is calculated based on the test environment configuration similarity and log semantic features, and the root cause of the test failure of the test case is determined based on the comprehensive score. This includes: calculating the semantic similarity between the test case and historical cases based on log semantic features; assigning corresponding weights to the test environment configuration similarity, semantic similarity and historical validity scores corresponding to the historical cases; calculating the comprehensive score of the test case based on the weights, test environment configuration similarity, semantic similarity and historical validity scores, and determining the root cause of the test failure based on the comprehensive score.

[0065] In actual implementation, embodiments of the present invention can comprehensively determine the root cause of test failure based on log semantic features, test environment configuration similarity, and effectiveness scores of historical case processing.

[0066] For example, embodiments of the present invention can perform multi-dimensional fusion:

[0067] For each similar case, calculate the overall score:

[0068] Score=0.4×Config_Sim+0.4×Semantic_Match+0.2×Historical_Effectiveness.

[0069] Wherein, Config_Sim represents the similarity of the test environment configuration, Semantic_Match represents the semantic matching degree of log semantic features, and Historical_Effectiveness represents the effectiveness of historical case processing.

[0070] The semantic matching degree is calculated using the event graph subgraph isomorphism algorithm.

[0071] The embodiments of the present invention can design a weighted scoring model (α+β+γ=1) that integrates configuration similarity (α), semantic matching degree (β), and historical solution effectiveness (γ) to achieve collaborative decision-making based on multi-dimensional data.

[0072] Optionally, in one embodiment of the present invention, matching the corresponding solution based on the root cause of test failure includes: matching multiple historical solutions based on the root cause of test failure; sorting the multiple historical solutions based on the validity scores corresponding to the multiple historical solutions to obtain a solution recommendation table; determining the solution based on the solution recommendation table for the root cause of test failure, and storing the solution, the root cause of test failure, and the validity score input by the user based on the execution result of the solution, or determining a new solution based on the order of the solution recommendation table.

[0073] It is understandable that there may be multiple solutions for the same root cause of test failure. For example, there may be solutions determined after user access, solutions that match based on the similarity of historical test cases, and solutions that directly match the root cause of test failure.

[0074] After determining the root cause of the test failure, embodiments of the present invention can match multiple historical solutions and dynamically calculate the effectiveness score given by the user for each historical solution, or the effectiveness score converted from server test data after the server executes the historical solution, such as the success rate of problem resolution, resolution speed, or number of positive user feedback after the solution is executed, and sort the multiple historical solutions to obtain a solution recommendation table.

[0075] In this embodiment of the invention, the solution with the highest effectiveness score can be executed first, following the order of the solution recommendation table. If the solution solves the root cause of the test failure, the solution is confirmed to be effective. If the solution cannot solve the root cause of the test failure, the next solution is selected and executed based on the order of the solution recommendation table.

[0076] The effectiveness score of the embodiments of the present invention can be calculated with reference to a variety of data, so that the effectiveness score can better characterize the effect of the solution in solving the corresponding problem. This allows the embodiments of the present invention to more accurately and quickly select the most likely solution from a large number of historical solutions, thereby improving problem-solving efficiency and reducing reliance on human experience.

[0077] Optionally, in one embodiment of the present invention, after determining the solution for the root cause of test failure based on the solution recommendation table, the method further includes: executing the solution and receiving a validity score input by the user; determining whether the validity score is greater than or equal to a first preset score; if it is greater than or equal to the first preset score, optimizing the weight of the solution in the historical validity scores based on the validity score; determining whether the validity score is less than or equal to a second preset score; if it is less than or equal to the second preset score, correcting the root cause of test failure based on the actual root cause of test failure input by the user, and calculating the weight of the comprehensive score of the test cases based on the validity score, wherein the second preset score is less than the first preset score.

[0078] After implementing a solution, this embodiment of the invention can receive user feedback on the problem-solving process. If the user's feedback is high, exceeding a certain rating threshold, the solution is considered effective, and its weight in the historical effectiveness score can be increased. This ensures that when similar problems arise again, the solution will be ranked higher in the solution recommendation list, leading to faster problem resolution. If the user's feedback is low, below a certain rating threshold, the solution is considered ineffective, or it may not meet the user's expectations, or the identified root cause may not match the actual root cause. In this case, the solution's weight in the historical effectiveness score can be reduced, ensuring that when similar problems arise again, the solution will be ranked lower in the solution recommendation list, decreasing the probability of implementing the solution. Alternatively, the actual root cause input by the user can be received to adjust the weight in the overall score calculation.

[0079] Through iterative optimization, the optimal matching of solutions is achieved, and data can be updated in a timely manner in the event of identification errors, so as to continuously improve the identification efficiency, identification accuracy and effectiveness of solution matching in the embodiments of the present invention.

[0080] Optionally, in one embodiment of the present invention, calculating a comprehensive score for test cases based on test environment configuration similarity and log semantic features, and determining the root cause of test failure based on the comprehensive score, further includes: acquiring multi-source monitoring data from at least one server test machine; converting the multi-source monitoring data into corresponding feature vectors; acquiring the architectural topology of at least one server test machine to capture the dependencies between components of at least one server test machine based on the architectural topology; calculating a comprehensive score by combining the feature vectors, test environment configuration similarity, and log semantic features; and determining the root cause of test failure by combining the comprehensive score and dependencies.

[0081] The multi-source monitoring data consists of various runtime performance metrics collected from the server test machine, including but not limited to: CPU utilization, load, memory usage, swap partition status, disk I / O throughput, latency, network bandwidth, number of connections, packet loss rate, process resource usage, thread status, etc.

[0082] This invention can transform multidimensional monitoring data into a unified numerical vector representation through mathematical transformations. For example, the mean, variance, and trend within a sliding window are transformed into time-series features; quantiles, peak values, and outliers are transformed into statistical features; and periodic patterns extracted through Fourier transform are transformed into frequency domain features.

[0083] The embodiments of the present invention can construct a directed graph describing the connections and dependencies between components in a server system to capture the dependencies between components, that is, the quantified dependency strength of components in the topology, such as call frequency, data traffic, timeout dependency, transaction consistency requirements, resource contention, etc.

[0084] This invention can fuse features from multiple dimensions. Based on the comprehensive scoring of environmental configuration features and log semantic features, it can also fuse monitoring data features to calculate a comprehensive score that characterizes the root causes of test failures.

[0085] Building upon the existing comprehensive score and considering the dependencies between components allows for a more accurate pinpointing of the root cause. For example, even if a component's direct error log score is low, it may be the root cause if it is a common dependency of multiple faulty components.

[0086] The embodiments of the present invention can identify hidden faults such as performance bottlenecks and resource contention that cannot be detected by pure log analysis. Based on topological dependencies, it can reverse deduce the propagation path of faults so as to match solutions more effectively.

[0087] Optionally, in one embodiment of the present invention, determining the root cause of test failure by combining comprehensive scoring and dependency relationships includes: obtaining the causal dependency relationship between the components of the server test machine corresponding to the historical test cases and the corresponding root cause of test failure; determining the degree of impact of component failure on test cases based on the causal dependency relationship; performing counterfactual analysis on the root cause of failure based on the degree of impact to obtain the analysis results, and matching solutions based on the analysis results.

[0088] This invention can learn the causal relationship between component failures and test failures from historical cases. For example, a failure of component A can lead to a failure of component B, which in turn leads to test failure, thus determining the degree of impact of component failures on test cases. The degree of impact quantifies the magnitude of the effect of each component failure on test failures, for example, by statistically analyzing historical data to determine the probability or frequency of test failures caused by a particular component failure.

[0089] Furthermore, embodiments of the present invention can perform counterfactual analysis on the root causes of failure based on the degree of impact, obtain analysis results, and match solutions based on the analysis results. For example, embodiments of the present invention can assume that component A has not failed, infer whether the root cause of the test failure has occurred, and thus obtain the importance of component A for each candidate root cause of test failure, or determine whether component A is the main root cause of test failure.

[0090] Based on the above counterfactual analysis, embodiments of the present invention can determine the root cause of test failure and find corresponding solutions from historical cases. For example, embodiments of the present invention can rank the solutions based on the effectiveness score and, based on the ranking, upgrade the solutions involving component A.

[0091] By analyzing historical causal dependencies, the impact of component failures can be assessed, avoiding superficial analysis of current failure phenomena. This allows for precise identification of the root cause, matching solutions based on the analysis results, and quickly providing effective solutions to improve operational efficiency.

[0092] Combination Figure 2 As shown, the working principle of the test failure root cause diagnosis method of the present invention will be explained in detail with an example.

[0093] First, embodiments of the present invention can construct a configuration feature dictionary: defining the hierarchical relationship of server components (such as "CPU" including sub-features such as "model", "number of cores" and "clock frequency").

[0094] Initialize the knowledge base: Import 500+ historical cases (including configurations, logs, root causes, and solutions) to complete the first fine-tuning of the BERT model.

[0095] like Figure 2 As shown, embodiments of the present invention may include the following steps:

[0096] Step S201: Trigger a test failure signal.

[0097] Step S202: Acquire data. After triggering the test failure signal, collect structured test environment configuration data and multi-dimensional unstructured log data of the server under test in the current test environment in real time.

[0098] For example, embodiments of the present invention can obtain the current server test machine from the test management platform. Server configuration data (various component models, quantities, OS, BMC versions, etc.) is pulled from the test platform's MySQL database. Test script logs, OS logs, and BMC logs are collected in real time and stored in a log pool. Structured configuration data and multi-source log files are output.

[0099] Step S203: Calculate configuration similarity. Calculate the similarity between the current test environment configuration data and the test environment configuration data of historical cases in the knowledge base module. Convert the structured configuration data into feature vectors, and calculate the similarity between them and the configuration vectors in the historical case library using a weighted cosine similarity algorithm to obtain a set of similar configuration cases.

[0100] This invention can acquire server configuration data and perform feature vectorization, normalizing numerical features (such as memory capacity) and performing One-Hot encoding on categorical features (such as component model and CPU model). Then, similarity calculation is performed: historical configuration cases are loaded from the knowledge base, sorted according to weighted cosine similarity, and configuration cases with a similarity > 0.7 are retained. This yields a list of similar configurations and a similarity score.

[0101] Step S204, Log Semantic Analysis. Semantic analysis is performed on multi-dimensional log information to extract log semantic features. The unstructured logs are semantically encoded using a BERT model fine-tuned from a log-domain corpus, generating error type probability distributions and a directed event graph with timestamps.

[0102] This invention can preprocess and clean multi-source log text, such as OS logs, test script logs, and BMC logs, to filter out noise such as timestamps and IP addresses; it also enhances segmentation using a log domain dictionary. Semantic encoding is then performed, and the log fragments are input into a fine-tuned BERT model to output error type probability distributions. Next, an event graph is constructed, error events are extracted, and directed edges are generated based on timestamps; duplicate events (such as multiple consecutive "memory allocation failures") are merged. The result is error type labels and an event graph JSON.

[0103] Step S205: Configure joint analysis of similarity and log semantics. This embodiment of the invention can comprehensively determine the root causes of test failure based on log semantic features, test environment configuration similarity, and the effectiveness score of historical case processing. A weighted scoring model is used to fuse the similarity value, semantic matching degree, and historical case effectiveness, and a list of candidate root causes is output according to the comprehensive score.

[0104] For example, embodiments of the present invention can perform multi-dimensional fusion:

[0105] For each similar case, calculate the overall score:

[0106] Score=0.4×Config_Sim+0.4×Semantic_Match+0.2×Historical_Effectiveness.

[0107] Wherein, Config_Sim represents the similarity of the test environment configuration, Semantic_Match represents the semantic matching degree of log semantic features, and Historical_Effectiveness represents the effectiveness of historical case processing.

[0108] The semantic matching degree is calculated using the event graph subgraph isomorphism algorithm.

[0109] A candidate set of root cause diagnoses was obtained.

[0110] Step S206: Retrieve the knowledge base based on the analysis results. Retrieve the knowledge base used to store historical test cases and corresponding solutions to determine the root causes of test failures and recommend solutions.

[0111] This invention can process the candidate set of root cause diagnosis results as follows: 1. Retrieve historical solutions for the corresponding root cause from the knowledge base; 2. Rank the solutions according to their effectiveness scores and recommend the Top-1 solution; 3. Record the diagnosis result to the knowledge base (requires manual confirmation to take effect).

[0112] Step S207, Recommend a solution. Sort the solutions by effectiveness for recommendation.

[0113] Step S208, Feedback Optimization. The manually confirmed root causes and solutions are written back to the knowledge base, and model parameters and case weights are dynamically adjusted based on the feedback results. The calculation method for test failure root causes and the recommendation criteria for solutions are optimized based on the effectiveness evaluation of user feedback.

[0114] Humans rate the effectiveness of the recommended solutions (1-5 stars).

[0115] Processing: If the rating is ≥4 stars: the case will be automatically marked as a "high-quality case," and its Historical_Effectiveness weight will be increased. If the rating is ≤2 stars: manual intervention will be triggered, the root cause will be corrected, the case will be re-stored in the knowledge base, and the joint analysis model will be fine-tuned in reverse. The updated knowledge base and model parameters will be obtained.

[0116] Based on the above process, the embodiments of the present invention can be implemented as follows.

[0117] Scenario: A server test case "Memory Stress Test" failed. The test machine was configured as follows: "CPU: Intel Xeon E5-2680 v4, Memory: 32GB DDR4 2133MHz×4, OS: CentOS 7.9".

[0118] S1 Data Acquisition:

[0119] Configuration data: Structured feature vectors (CPU model: E5-2680 v4, memory capacity: 128GB, OS version: CentOS 7.9, etc.).

[0120] Log excerpt: OS log: "Out of memory: Kill process 1234 (memtest) score582 or sacrifice child"

[0121] Test script log: "[ERROR] Memory allocation failed, timed out after 5 consecutive attempts to allocate 4GB of space"

[0122] S2 configuration similarity calculation:

[0123] Three similar cases were matched from the knowledge base (similarity 0.82, 0.78, and 0.73), all of which had the same CPU model and memory capacity ≥ 64GB.

[0124] S3 Log Semantic Analysis:

[0125] Error label: "Memory overflow (92% probability)".

[0126] Event graph: "Memory allocation failed → process killed" (continuous timestamps, correlation 0.95).

[0127] S4 Joint Analysis:

[0128] Overall Score:

[0129] Case 1: 0.4×0.82+0.4×0.95+0.2×0.9 (historical solution valid)=0.87;

[0130] Case 2: 0.4×0.78+0.4×0.8+0.2×0.7=0.77;

[0131] Case 3: 0.4×0.73+0.4×0.75+0.2×0.6=0.70.

[0132] Output: Top-1 root cause "Inappropriate memory allocation strategy, large page memory not enabled".

[0133] S5 Solution Recommendation:

[0134] Recommended solution: "Execute echo 1> / sys / kernel / mm / hugepages / hugepages-2048kB / nr_hugepages to enable large page memory, then restart the test service."

[0135] Manual verification: After execution and testing, the scheme is marked as valid and stored in the knowledge base.

[0136] In summary, the embodiments of this invention can automatically complete root cause localization in a short time through automated configuration matching and semantic analysis, significantly improving diagnostic efficiency. It supports concurrent processing of multiple test failure cases, making it suitable for testing scenarios with large-scale server clusters and significantly shortening the overall testing cycle. By combining server configuration similarity with log context semantics (such as error correlation chains), it greatly reduces the analysis misjudgment rate and improves accuracy. The structured storage of historical case configurations, error patterns, and solutions forms a traceable knowledge base, solving the problems of strong reliance on and easy loss of traditional human experience. It reduces reliance on senior test engineers, allowing junior personnel to quickly locate problems with the help of the system, significantly reducing enterprise labor costs.

[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0138] like Figure 3 As shown, an embodiment of the present invention also provides a diagnostic device 10 for testing the root cause of failure, including: an acquisition module 100, an analysis module 200, and a diagnostic module 300.

[0139] Specifically, the acquisition module 100 is used to capture test failure signals from at least one server test machine and determine test cases based on the test failure signals in order to obtain multi-source log information and test environment configuration data of the test cases.

[0140] The analysis module 200 is used to calculate the similarity of test environment configuration between test cases and historical cases based on test environment configuration data, and to analyze multi-source log information to obtain corresponding log semantic features.

[0141] The diagnostic module 300 is used to calculate the comprehensive score of test cases based on the similarity of test environment configuration and log semantic features, to determine the root cause of test failure based on the comprehensive score, to match the corresponding solution based on the root cause of test failure, and to push the solution to at least one server test machine.

[0142] Optionally, in one embodiment of the present invention, the diagnostic module 300 includes: a matching unit, a sorting unit, and a determining unit.

[0143] The matching unit is used to match multiple historical solutions based on the root cause of test failure.

[0144] The sorting unit is used to sort multiple historical solutions based on their effectiveness scores to obtain a solution recommendation table.

[0145] The determination unit is used to determine solutions to the root causes of test failures based on the solution recommendation table, and to store the solution, the root causes of test failures, and the validity score of user input based on the execution results of the solution, or to determine new solutions based on the order of the solution recommendation table.

[0146] Optionally, in one embodiment of the present invention, the diagnostic device 10 for testing the root cause of failure further includes: a receiving module, a first judgment module, and a second judgment module.

[0147] The receiving module is used to execute the solution and receive the validity score of the user input.

[0148] The first judgment module is used to determine whether the validity score is greater than or equal to the first preset score. If it is greater than or equal to the first preset score, the weight of the solution in the historical validity score is optimized based on the validity score.

[0149] The second judgment module is used to determine whether the validity score is less than or equal to the second preset score. If it is less than or equal to the second preset score, the root cause of the test failure is corrected based on the actual test failure root cause input by the user, and the weight of the comprehensive score of the test case is optimized based on the validity score. The second preset score is less than the first preset score.

[0150] Optionally, in one embodiment of the present invention, the diagnostic module 300 includes: a first calculation unit, an allocation unit, and a second calculation unit.

[0151] The first calculation unit is used to calculate the semantic similarity between test cases and historical cases based on log semantic features.

[0152] The allocation unit is used to assign corresponding weights to the similarity of the test environment configuration, semantic similarity, and historical validity scores corresponding to historical cases.

[0153] The second calculation unit is used to calculate the comprehensive score of the test cases based on weights, test environment configuration similarity, semantic similarity, and historical validity scores, and to determine the root cause of test failure based on the comprehensive score.

[0154] Optionally, in one embodiment of the present invention, the analysis module 200 includes: a cleaning unit, a third calculation unit, and an extraction unit.

[0155] The cleaning unit is used to clean multi-source log information and use a pre-built log domain dictionary to segment the cleaned multi-source log information to obtain multiple log fragments.

[0156] The third calculation unit is used to calculate the error type probability distribution based on multiple log fragments.

[0157] The extraction unit is used to extract error events based on the error type probability distribution to obtain log semantic features.

[0158] Optionally, in one embodiment of the present invention, the analysis module 200 further includes a generation unit and a fourth calculation unit.

[0159] The generation unit is used to generate the corresponding directed edges based on the timestamp of the error event.

[0160] The fourth computational unit is used to construct an event graph using directed edges, and to calculate semantic similarity based on the event graph.

[0161] Optionally, in one embodiment of the present invention, the analysis module 200 further includes a definition unit, a construction unit, and a training unit.

[0162] The definition unit is used to define the hierarchical relationship of server components.

[0163] Building units are used to construct log domain dictionaries based on hierarchical relationships.

[0164] Optionally, in one embodiment of the present invention, the extraction unit includes an import unit and a training unit.

[0165] The import unit is used to import multiple historical test cases that meet the preset test case conditions.

[0166] The training unit is used to build a pre-trained model based on multiple historical test cases, and to use the pre-trained model to calculate the error type probability distribution.

[0167] Optionally, in one embodiment of the present invention, the analysis module 200 includes: a processing unit, a conversion unit, and a fifth calculation unit.

[0168] The processing unit is used to extract numerical features from the test environment configuration data and normalize the numerical features to obtain numerical data.

[0169] The transformation unit is used to transform categorical features in the test environment configuration data to obtain encoded data.

[0170] The fifth calculation unit is used to calculate the cosine similarity between test cases and multiple historical test cases based on numerical data and coded data, and to obtain historical cases based on the cosine similarity.

[0171] Optionally, in one embodiment of the present invention, the fifth calculation unit includes: a judgment subunit, a first acquisition subunit, a calculation subunit, and a reminder subunit.

[0172] The judgment subunit is used to determine whether there is at least one case in the historical test cases whose cosine similarity is greater than a preset similarity threshold.

[0173] The first acquisition subunit is used to obtain historical cases based on cosine similarity when at least one case exists.

[0174] The calculation subunit is used to verify the validity of the cosine similarity calculation using historical test cases when at least one case is not available.

[0175] The reminder subunit is used to generate an error reminder if the calculation validity does not meet the preset validity conditions; otherwise, it generates a new case reminder.

[0176] Optionally, in one embodiment of the present invention, the diagnostic module 300 further includes: a first acquisition unit, a conversion unit, a second acquisition unit, a sixth calculation unit, and a diagnostic unit.

[0177] The first acquisition unit is used to acquire multi-source monitoring data from at least one server test machine.

[0178] The transformation unit is used to transform multi-source monitoring data into corresponding feature vectors.

[0179] The second acquisition unit is used to acquire the architectural topology of at least one server test machine, so as to capture the dependencies between the components of at least one server test machine based on the architectural topology.

[0180] The sixth calculation unit is used to calculate a comprehensive score by combining feature vectors, test environment configuration similarity, and log semantic features.

[0181] The diagnostic unit is used to determine the root cause of test failure by combining comprehensive scores and dependencies.

[0182] Optionally, in one embodiment of the present invention, the diagnostic unit includes: a second acquisition subunit, a determination subunit, and an analysis subunit.

[0183] The second acquisition subunit is used to acquire the causal dependency relationship between the components of the server test machine corresponding to the historical test case and the root cause of the corresponding test failure.

[0184] Identify sub-units to determine the impact of component failures on test cases based on causal dependencies.

[0185] The analysis subunit is used to perform counterfactual analysis on the root causes of failure based on the degree of impact, obtain analysis results, and match solutions based on the analysis results.

[0186] For a description of the features in the embodiment of the diagnostic device for testing the root cause of failure, please refer to the relevant description of the embodiment of the diagnostic method for testing the root cause of failure, which will not be repeated here.

[0187] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the diagnostic method for testing root causes of failure.

[0188] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above embodiments of the method for diagnosing the root cause of test failure.

[0189] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0190] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the method for diagnosing the root cause of test failure.

[0191] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the method for diagnosing the root cause of test failure.

[0192] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0193] The above provides a detailed description of a method for diagnosing the root causes of test failures provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of these embodiments are only intended to aid in understanding the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the scope of protection of the claims of this invention.

Claims

1. A diagnostic method of testing a failure root cause, characterized by, The method comprises the following steps: capturing a test failure signal of at least one server test machine, and determining a test case based on the test failure signal to obtain multi-source log information and test environment configuration data of the test case; calculating a test environment configuration similarity of the test case and a historical case based on the test environment configuration data, and analyzing the multi-source log information to obtain corresponding log semantic features; calculating a comprehensive score of the test case based on the test environment configuration similarity and the log semantic features, determining a test failure root cause of the test case based on the comprehensive score, matching a corresponding solution based on the test failure root cause, and pushing the solution to at least one of the server test machines; wherein the calculation of the comprehensive score of the test case based on the test environment configuration similarity and the log semantic features, and the determination of the test failure root cause of the test case based on the comprehensive score further comprise: obtaining multi-source monitoring data of at least one of the server test machines; converting the multi-source monitoring data into corresponding feature vectors; obtaining an architecture topology relationship of at least one of the server test machines to capture a dependency relationship between components of at least one of the server test machines based on the architecture topology relationship; calculating the comprehensive score in combination with the feature vectors, the test environment configuration similarity and the log semantic features; and determining the test failure root cause in combination with the comprehensive score and the dependency relationship.

2. The diagnostic method for testing failure root cause according to claim 1, characterized in that, The matching of the corresponding solution based on the test failure root cause comprises: matching a plurality of historical solutions based on the test failure root cause; sorting a plurality of the historical solutions based on corresponding effectiveness scores of the plurality of the historical solutions to obtain a solution recommendation table, and determining a solution of the test failure root cause based on the solution recommendation table; executing the solution, and storing the solution, the test failure root cause and an effectiveness score input by a user based on an execution result of the solution, or determining a new solution based on an order of the solution recommendation table.

3. The diagnostic method for testing failure root cause according to claim 2, characterized in that, After determining the solution of the test failure root cause based on the solution recommendation table, the method further comprises: executing the solution, and receiving the effectiveness score input by the user; determining whether the effectiveness score is greater than or equal to a first preset score, and if the effectiveness score is greater than or equal to the first preset score, optimizing a weight of the solution in historical effectiveness scores based on the effectiveness score; determining whether the effectiveness score is less than or equal to a second preset score, and if the effectiveness score is less than or equal to the second preset score, correcting the test failure root cause based on an actual test failure root cause input by the user, and optimizing a comprehensive score calculation weight of the test case based on the effectiveness score, wherein the second preset score is less than the first preset score.

4. The diagnostic method for testing failure root cause according to claim 3, characterized in that, The calculation of the comprehensive score of the test case based on the test environment configuration similarity and the log semantic features, and the determination of the test failure root cause of the test case based on the comprehensive score comprise: calculate semantic similarity between the test case and the historical cases based on the log semantic features; respectively assign corresponding weights to the test environment configuration similarity, the semantic similarity and the historical effectiveness score corresponding to the historical cases; calculate a comprehensive score of the test case based on the weights, the test environment configuration similarity, the semantic similarity and the historical effectiveness score, and determine the test failure root cause based on the comprehensive score.

5. The diagnostic method for testing failure root cause according to claim 4, characterized in that, The analysis of the multi-source log information obtains corresponding log semantic features, which includes: cleaning the multi-source log information, and segmenting the cleaned multi-source log information using a pre-constructed log field dictionary to obtain a plurality of log segments; calculating an error type probability distribution based on a plurality of the log segments; extracting error events based on the error type probability distribution to obtain the log semantic features.

6. The diagnostic method for testing failure root cause according to claim 5, characterized in that, After extracting error events based on the error type probability distribution, it further includes: generating a corresponding directed edge based on the timestamp of the error event; constructing an event graph using the directed edge to calculate the semantic similarity based on the event graph.

7. The diagnostic method for testing failure root cause according to claim 5, wherein, Before segmenting the cleaned multi-source log information using a pre-constructed log field dictionary, it further includes: defining a hierarchical relationship of server components; constructing the log field dictionary based on the hierarchical relationship.

8. The diagnostic method for testing failure root cause according to claim 5, wherein, The extracting error events based on the error type probability distribution to obtain the log semantic features includes: importing a plurality of historical test cases that meet a pre-set case condition; constructing a pre-training model based on a plurality of the historical test cases to calculate the error type probability distribution using the pre-training model.

9. The diagnostic method for testing failure root cause according to claim 1, wherein, The calculating the test environment configuration similarity between the test case and the historical cases based on the test environment configuration data includes: extracting numerical features from the test environment configuration data and performing normalization processing on the numerical features to obtain numerical data; performing variable conversion on the category type features in the test environment configuration data to obtain encoding data; calculating the cosine similarity between the test case and a plurality of historical test cases based on the numerical data and the encoding data, and obtaining the historical cases based on the cosine similarity.

10. The diagnostic method of testing failure root cause according to claim 9, characterized in that, The obtaining the historical cases based on the cosine similarity includes: determining whether there is at least one case in the historical test cases whose cosine similarity is greater than a pre-set similarity threshold; if there is at least one of the cases, obtaining the historical cases based on the cosine similarity; if there is no at least one of the cases, verifying the calculation effectiveness of the cosine similarity using the historical test cases; generating an error reminder in the case where the calculation effectiveness does not meet a pre-set effective condition, otherwise, generating a new case reminder.

11. The diagnostic method for testing failure root cause according to claim 1, characterized in that, The determining the test failure root cause in combination with the comprehensive score and the dependency relationship includes: obtaining a causal dependency relationship between components of a server test machine corresponding to a historical test case and a corresponding test failure root cause; determining the impact degree of the failure of the component on the test case based on the causal dependency relationship; performing counterfactual analysis on the failure root cause based on the influence degree, to obtain an analysis result, and matching the solution based on the analysis result.

Citation Information

Patent Citations

  • Test case failure reason analysis method and device and electronic equipment

    CN110990575A

  • Test case analysis method and device, processor and electronic equipment

    CN116303029A