Containerized application fault detection method, equipment and medium
By building a guided analysis logic database and an operation and maintenance experience database, and combining a large language model for one-time analysis, the problem of low fault detection efficiency of cluster containerized applications is solved, and automated application abnormal data collection and fast and accurate fault handling are realized.
Patent Information
- Application Number
- CN202510602803.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-19
AI Technical Summary
In the deployment scenario of cluster containerized application, the fault detection efficiency caused by manual participation and multiple inspections in the existing technology is low, especially in a variety of resource environments to diagnose abnormalities and detect problems.
By building a guided analysis logic database and operation and maintenance experience database, using large language models and prompt words, one-time analysis is realized, automatic collection and inspection of application abnormal data, and fault detection is carried out by combining the matching of different data dimensions and pre-trained large language models.
It realizes automatic collection of abnormal data and abnormal cause inspection in cluster environments, improves operation and maintenance efficiency, reduces problem detection time, and improves the rapid accuracy and business continuity of fault handling.
Smart Images

Figure CN120508424A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of containerized deployment technology, and in particular to a method, device, and medium for detecting containerized application faults. Background Art
[0002] The rise of big language model technology has brought certain conveniences to the operation and maintenance of containerized applications. For example, big language models can identify common failure modes and potential problem trends by learning from historical operation and maintenance logs, error codes, system events, and other information. At the same time, by leveraging the powerful knowledge base and reasoning capabilities of big language models, effective solutions can be proposed for problems.
[0003] However, although the large language model can analyze data provided by users, users still need to identify and collect the data that may be needed, and need to sequentially troubleshoot several possible problems, allowing the large language model to perform separate analysis for each investigation. This is especially true in the scenario of cluster container application deployment, as it involves application anomaly diagnosis and problem troubleshooting in multiple resource environments. Therefore, manual participation and multiple troubleshooting methods lead to low efficiency in containerized application fault detection. Summary of the Invention
[0004] Embodiments of the present application provide a containerized application fault detection method, device, and medium to address the problem of low efficiency in containerized application fault detection.
[0005] The embodiments of this application adopt the following technical solutions:
[0006] On the one hand, an embodiment of the present application provides a containerized application fault detection method, which includes: when an application fails, collecting deployment data based on different data dimensions to obtain containerized deployment data of the application; extracting a data set to be analyzed from the containerized deployment data; matching the data dimensions in the data set to be analyzed in a guidance analysis logic database to obtain analysis logic prompt words of the data set to be analyzed; matching the data dimensions in the data set to be analyzed in an operation and maintenance experience database to obtain fault analysis prompt words of the data set to be analyzed; and performing fault detection on the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words based on a pre-trained large language model.
[0007] In one example, collecting deployment data based on different data dimensions to obtain containerized deployment data of the application specifically includes: determining a collection time period within a preset range from the error reporting time; and reading the deployment data within the collection time period based on different data dimensions to obtain the containerized deployment data of the application.
[0008] In one example, extracting the data set to be analyzed from the containerized deployment data specifically includes: performing data cleaning on the containerized deployment data to obtain cleaned containerized deployment data; the cleaned containerized deployment data has the same or different data dimensions as the containerized deployment data; and structuring the cleaned containerized deployment data to obtain the data set to be analyzed.
[0009] In one example, the data cleaning of the containerized deployment data to obtain cleaned containerized deployment data specifically includes: deleting data dimensions with empty dimension data in the containerized deployment data to obtain initial cleaned containerized deployment data; extracting preset key fields from the initial cleaned containerized deployment data to obtain cleaned containerized deployment data.
[0010] In one example, matching the data dimensions in the data set to be analyzed to obtain the analysis logic prompt words of the data set to be analyzed specifically includes: determining the data dimension combination in the data set to be analyzed; matching the data dimension combination in the dimension combination mapping relationship table to obtain the analysis logic prompt words of the data set to be analyzed; the dimension combination mapping relationship table is used to represent the analysis logic prompt words corresponding to different dimension combinations.
[0011] In one example, matching the data dimension combination to obtain the analysis logic prompt words of the data set to be analyzed specifically includes: matching the data dimension combination to obtain the analysis steps and analysis report format of the data set to be analyzed; and obtaining the analysis logic prompt words of the data set to be analyzed based on the analysis steps and the analysis report format.
[0012] In one example, matching the data dimensions in the data set to be analyzed to obtain the fault analysis prompt words of the data set to be analyzed specifically includes: matching each data dimension in the data set to be analyzed separately in a data dimension mapping relationship table, and determining the fault analysis prompt words of each data dimension as the fault analysis prompt words of the data set to be analyzed; the data dimension mapping relationship table is used to represent the fault analysis prompt words corresponding to different data dimensions.
[0013] In one example, the fault detection is performed on the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words based on the pre-trained large language model, specifically including: organizing the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words in a data format according to the large language model SDK; calling the large language model SDK to pass the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words to the pre-trained large language model to obtain a fault detection analysis report that conforms to the preset format.
[0014] On the other hand, an embodiment of the present application provides a containerized application fault detection device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the containerized application fault detection methods described above.
[0015] On the other hand, an embodiment of the present application provides a non-volatile computer storage medium for containerized application fault detection, which stores computer-executable instructions. The computer-executable instructions can execute any of the above-mentioned containerized application fault detection methods.
[0016] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:
[0017] Combining the rich operation and maintenance experience accumulated in the cluster container environment, we build a guidance analysis logic database and an operation and maintenance experience database, and set up a data set to be analyzed that meets the matching requirements. This enables a one-time analysis based on a large language model and prompt words to achieve automatic collection of application exception data and inspection of the cause of the exception. This is especially suitable for application exception diagnosis and problem troubleshooting in cluster applications involving multiple resource environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solution of the present application, some embodiments of the present application will be described in detail below with reference to the accompanying drawings, in which:
[0019] Figure 1 A flowchart of a containerized application fault detection method provided in an embodiment of the present application;
[0020] Figure 2 A schematic diagram of the structure of a containerized application fault detection device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] Some embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0023] Figure 1This is a flowchart of a containerized application fault detection method provided in an embodiment of the present application. This method can be applied to various business areas, such as internet finance, e-commerce, instant messaging, gaming, and government affairs. Certain input parameters or intermediate results in this process can be manually adjusted to help improve accuracy.
[0024] The analysis method involved in the embodiments of the present application can be implemented by a terminal device or a server, and the present application does not impose any special restrictions on this. For ease of understanding and description, the following embodiments are described in detail using a server as an example.
[0025] It should be noted that the server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make any specific restrictions on this.
[0026] Figure 1 The process in includes the following steps:
[0027] S101: When an application fails, deployment data is collected based on different data dimensions to obtain containerized deployment data of the application.
[0028] In some embodiments of the present application, data dimensions may include application status information, application-related resource information in the cluster, events generated by the application, application node logs, resource utilization, gateway logs, and other information.
[0029] For example, application status information may include whether the application in the cluster is ready and information about the application's CPU and memory resource utilization. Application-related resource information may include configurations such as the application's associated Service, Configmap, Secret, and access control ServiceAccount permissions.
[0030] The specific collection process includes:
[0031] First, determine the collection period within the preset range from the error reporting time. Then, read the deployment data within the collection period based on different data dimensions to obtain the containerized deployment data of the application.
[0032] S102: Extracting a data set to be analyzed from the containerized deployment data.
[0033] It should be noted that the containerized deployment data is extracted and analyzed based on a pre-set acquisition method.
[0034] In some embodiments of the present application, data cleaning is performed on the containerized deployment data to obtain cleaned containerized deployment data, and the cleaned containerized deployment data is structured to obtain a data set to be analyzed.
[0035] The data dimensions of the cleaned containerized deployment data and the containerized deployment data are the same or different.
[0036] Based on this, the data cleaning process is as follows:
[0037] In the containerized deployment data, data dimensions with empty dimension data are deleted to obtain initial cleansed containerized deployment data.
[0038] Preset key fields are extracted from the initial clean containerized deployment data to obtain the clean containerized deployment data.
[0039] Based on this, the process of extracting important fields and structuring them is then analyzed using a large language model.
[0040] For example, the application status field is extracted from the application information, and then the application data is structured and converted into the entity structure defined in the system as the return value of the method.
[0041] It should be noted that in some cases, the data dimension may be empty. For example, when the application fails to successfully generate an event within the collection time period, the event is empty.
[0042] S103: Matching the data dimensions in the data set to be analyzed in the guidance analysis logic database to obtain analysis logic prompt words of the data set to be analyzed.
[0043] In some embodiments of the present application, analysis logic words for different data dimensions are pre-set, and different data dimension combinations have different analysis steps. The analysis steps are set for the data dimension combination and can quickly diagnose the problem location of the data dimension combination.
[0044] Based on this, the matching process is as follows:
[0045] First, determine the data dimension combinations in the dataset to be analyzed. Then, match the data dimension combinations in the dimension combination mapping relationship table to obtain the analysis logic prompt words of the dataset to be analyzed.
[0046] The dimension combination mapping relationship table is used to represent the analysis logic prompt words corresponding to different dimension combinations.
[0047] It should be noted that, in the process of prompting the analysis steps, the output report format may also be prompted, for example, the report format includes the problem, the cause analysis of the problem and the solution.
[0048] For example, the analysis logic prompt words are as follows:
[0049] If you are an expert in troubleshooting containerized clusters, follow these steps to analyze:
[0050] 1. First check the application's own configuration (such as database connection, permissions, image version)
[0051] 2. Check dependent resources (whether the database and cache are normal and whether the network is connected)
[0052] 3. Finally, analyze the node and cluster environment (resource utilization, network fluctuations, scheduling strategy)
[0053] The output must include: problem location, direct cause, related factors, executable plan, and preventive measures
[0054] S104: Matching data dimensions in the data set to be analyzed in the operation and maintenance experience database to obtain fault analysis prompt words of the data set to be analyzed.
[0055] In some embodiments of the present application, in the data dimension mapping relationship table, each data dimension in the data set to be analyzed is matched separately, and the fault analysis prompt word of each data dimension is determined as the fault analysis prompt word of the data set to be analyzed.
[0056] The data dimension mapping relationship table is used to represent the fault analysis prompt words corresponding to different data dimensions.
[0057] It should be noted that the general process and overall approach to anomaly troubleshooting, summarized based on anomaly detection scenarios and requirements, combined with actual operation and maintenance experience, has led to the construction of analysis logic prompts and fault analysis prompts. This provides an overall approach for anomaly detection and troubleshooting using large language models, and guides the format of anomaly analysis reports generated by large language models.
[0058] S105: Performing fault detection on the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words according to the pre-trained large language model.
[0059] It should be noted that the analysis logic prompt words can also be equivalent to the system prompt words, and the fault analysis prompt words can be equivalent to the dialogue prompt words.
[0060] System prompts: These are primarily used to set the operating rules and patterns of large models or provide some general background information, guiding the model to process and generate answers in a specific manner. They play a macro-control role in the model's overall behavior and output style.
[0061] Conversation prompts: They focus more on guiding the specific conversation process and content direction, helping the model generate responses that are appropriate to the current conversation context, allowing the conversation to proceed naturally and smoothly.
[0062] The fault analysis prompt words are as follows:
[0063] When "Database connection failed" occurs:
[0064] 1. 80% of the cases are configuration errors (port / IP / key errors)
[0065] 2. 15% are due to abnormalities in the database service itself (such as instance restart and authentication service failure)
[0066] 3.5% of the time, network partitions caused connection interruptions
[0067] Prioritize checking the app configuration files and permission settings.
[0068] In some embodiments of the present application, the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words are organized in a data format according to the large language model SDK.
[0069] By calling the large language model SDK, the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words are passed to the pre-trained large language model to obtain a fault detection and analysis report that conforms to the preset format.
[0070] That is, a report can be obtained whose format and content meet the preset requirements, including specific problem analysis and feasible solutions for applying anomaly detection.
[0071] That is to say, when the user selects an application for anomaly analysis, the system automatically obtains the generated analysis logic prompt words, fault analysis prompt words and the data set to be analyzed obtained by the set method, organizes them according to the SDK format of the large language model and passes them to the large language model for intelligent analysis, and finally obtains the application anomaly detection report.
[0072] It should be noted that although the embodiments of this application are based on Figure 1 Steps S101 to S105 are described in sequence, but this does not mean that steps S101 to S105 must be performed in a strict order. Figure 1 The order shown in FIG1 is to introduce and explain step S101 to step S105 in order to facilitate those skilled in the art to understand the technical solution of the embodiment of the present application. In other words, in the embodiment of the present application, the order between step S101 to step S105 can be appropriately adjusted according to actual needs.
[0073] pass Figure 1This method is particularly suitable for intelligent troubleshooting of application anomalies in a cluster environment. It can ideally collect and extract data on abnormal applications in the cluster, check the causes of the anomalies, and provide problem repair suggestions.
[0074] Specifically: Combining the rich operation and maintenance experience accumulated in the cluster container environment, we build a guidance analysis logic database and an operation and maintenance experience database, and set up a data set to be analyzed that meets the matching requirements. This enables a one-time analysis based on a large language model and prompt words to achieve automatic collection of application abnormality data and abnormal cause inspection. This is especially suitable for application abnormality diagnosis and problem troubleshooting in cluster applications involving multiple resource environments.
[0075] Furthermore, this intelligent analysis method sets prompt words for data in different data dimensions, which can not only check the abnormal conditions of the application itself, but also make full use of resource information and log data related to the application to conduct comprehensive problem investigation, integrate various types of resource information and log data for in-depth analysis, and combine large language models to quickly locate problems and their causes, and propose effective problem-solving solutions, thereby significantly improving operation and maintenance efficiency, greatly reducing the time required for problem investigation, and making the operation and maintenance process easier and more efficient.
[0076] In summary, the powerful analytical capabilities of large language models can be used to quickly identify and locate sources of anomalies in complex cluster environments and accurately diagnose problems. At the same time, the intelligent analysis and fast and accurate fault handling capabilities provided by large models help reduce system downtime, improve service stability and availability, and ensure business continuity.
[0077] Based on the same idea, some embodiments of the present application also provide devices and non-volatile computer storage media corresponding to the above methods.
[0078] Figure 2 A schematic diagram of the structure of a containerized application fault detection device provided in an embodiment of the present application includes:
[0079] at least one processor; and,
[0080] a memory communicatively connected to the at least one processor; wherein,
[0081] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform any one of the above-mentioned containerized application fault detection methods.
[0082] Some embodiments of the present application provide a non-volatile computer storage medium for containerized application fault detection, which stores computer-executable instructions capable of executing any of the above-described containerized application fault detection methods.
[0083] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0084] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0085] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0086] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0087] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0089] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0090] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0091] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0092] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0093] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the technical principles of the present application should fall within the scope of protection of the present application.
Claims
1. A containerized application fault detection method, characterized in that: The method comprises: When an application fails, deployment data is collected based on different data dimensions to obtain containerized deployment data of the application. Extracting a data set to be analyzed from the containerized deployment data; In the guidance analysis logic database, data dimensions in the data set to be analyzed are matched to obtain analysis logic prompt words of the data set to be analyzed; Matching the data dimensions in the data set to be analyzed in the operation and maintenance experience database to obtain the fault analysis prompt words of the data set to be analyzed; According to the pre-trained large language model, fault detection is performed on the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words.
2. The method according to claim 1, characterized in that The collecting deployment data based on different data dimensions to obtain the containerized deployment data of the application specifically includes: Determine the collection time period within the preset range from the error reporting time; Based on different data dimensions, the deployment data within the collection time period is read respectively to obtain the containerized deployment data of the application.
3. The method according to claim 1, characterized in that Extracting the data set to be analyzed from the containerized deployment data specifically includes: performing data cleaning on the containerized deployment data to obtain cleaned containerized deployment data; wherein the cleaned containerized deployment data has the same or different data dimensions as the containerized deployment data; The cleaned containerized deployment data is structured to obtain a data set to be analyzed.
4. The method according to claim 3, characterized in that Cleaning the containerized deployment data to obtain cleaned containerized deployment data specifically includes: In the containerized deployment data, data dimensions whose dimension data are empty are deleted to obtain initial cleaned containerized deployment data; Preset key fields are extracted from the initial cleaning containerized deployment data to obtain cleaning containerized deployment data.
5. The method according to claim 1, wherein The matching of the data dimensions in the data set to be analyzed to obtain the analysis logic prompt words of the data set to be analyzed specifically includes: Determining a data dimension combination in the data set to be analyzed; In the dimension combination mapping relationship table, the data dimension combination is matched to obtain the analysis logic prompt words of the data set to be analyzed; the dimension combination mapping relationship table is used to represent the analysis logic prompt words corresponding to different dimension combinations.
6. The method according to claim 1, characterized in that The matching of the data dimension combination to obtain the analysis logic prompt words of the data set to be analyzed specifically includes: Matching the data dimension combinations to obtain analysis steps and analysis report formats for the data set to be analyzed; According to the analysis steps and the analysis report format, analysis logic prompt words of the data set to be analyzed are obtained.
7. The method according to claim 1, characterized in that The matching of the data dimensions in the data set to be analyzed to obtain the fault analysis prompt words of the data set to be analyzed specifically includes: In the data dimension mapping relationship table, each data dimension in the data set to be analyzed is matched separately, and the fault analysis prompt word of each data dimension is determined as the fault analysis prompt word of the data set to be analyzed; the data dimension mapping relationship table is used to represent the fault analysis prompt words corresponding to different data dimensions.
8. The method according to claim 1, characterized in that The performing of fault detection on the data set to be analyzed, the analysis logic prompt words, and the fault analysis prompt words based on the pre-trained large language model specifically includes: According to the large language model SDK method, the data set to be analyzed, the analysis logic prompt words and the fault analysis prompt words are organized into a data format; By calling the large language model SDK, the data set to be analyzed, the analysis logic prompt words and the fault analysis prompt words are passed to the pre-trained large language model to obtain a fault detection analysis report that conforms to a preset format.
9. A containerized application fault detection device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the containerized application fault detection method according to any one of claims 1 to 8.
10. A containerized application fault detection non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer-executable instructions can execute a containerized application fault detection method as described in any one of claims 1 to 8.