A container platform inspection method and device based on a large language model

CN122838152APending Publication Date: 2026-09-29JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611028681.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,现有的巡检方法中,通常直接采用规则引擎或阈值告警进行独立判断,并没有建立多源异构数据与大语言模型推理能力的深度协同机制

Benefits of technology

[0019]为达上述目的,本申请第四方面实施例提出了一种非临时性计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现如第一方面实施例所述的一种基于大语言模型的容器平台巡检方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838152A_ABST
    Figure CN122838152A_ABST
Patent Text Reader

Abstract

The application provides a container platform inspection method and device based on a large language model, and the application comprises the following steps: acquiring state data of a target resource of a container platform, and screening out an abnormal target object; querying log, event and performance index data associated with the abnormal target object, and assembling the data to form multi-dimensional running data; matching corresponding abnormal diagnosis knowledge from a preset expert knowledge base according to the type of the abnormal target object; filling the multi-dimensional running data and the abnormal diagnosis knowledge into a prompt word template to generate an inspection prompt word, and sending the inspection prompt word to a large language model for processing to obtain an inspection result. The application can quickly gather abnormal related multi-dimensional information, and realizes automatic diagnosis by combining expert knowledge and a large model, thereby improving the abnormal positioning efficiency of the container platform and the fault processing accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of container platform inspection technology, and in particular to a container platform inspection method and apparatus based on a large language model. Background Technology

[0002] Container platforms, as a core infrastructure of cloud-native architectures, are widely used in the deployment and management of various critical business systems. With the development of intelligent operations and maintenance, existing inspection systems have constructed a multi-dimensional health assessment mechanism through the collaborative operation of status monitoring, log analysis, and performance indicator collection. Specifically, this process covers the entire process from data collection and anomaly detection to fault location, including key steps such as calling interface services, retrieving time-series databases, and analyzing event streams, aiming to ensure the stability of cluster operation.

[0003] However, existing inspection methods typically rely on rule engines or threshold alarms for independent judgment, without establishing a deep collaborative mechanism between multi-source heterogeneous data and the reasoning capabilities of large language models. Due to the lack of semantic-level correlation analysis of status, logs, events, and performance data, and the absence of expert knowledge bases to assist decision-making, diagnostic accuracy is low when facing complex coupled faults, making it difficult to generate actionable repair suggestions, thus severely impacting operational efficiency and system reliability. Summary of the Invention

[0004] The present invention aims to at least partially solve one of the technical problems in the related art.

[0005] Therefore, the first objective of this invention is to propose a container platform inspection method based on a large language model.

[0006] Another objective of this invention is to propose a container platform inspection device based on a large language model.

[0007] The third objective of this invention is to provide a computer device.

[0008] The fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0009] To achieve the above objectives, a first aspect of the present invention proposes a container platform inspection method based on a large language model, comprising: S1, Obtain the status data of the target resource in the container platform, and filter out abnormal target objects based on the status data; S2, query the log data, event data, and performance indicator data associated with the target object of the anomaly, and assemble the status data, log data, event data, and performance indicator data to obtain multi-dimensional operating data; S3, Match the corresponding abnormal diagnosis knowledge from the preset expert knowledge base according to the type of the target object of the abnormality. The expert knowledge base stores common abnormal descriptions and handling solutions for different types of resources. S4. Fill the multidimensional operation data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and send the inspection prompt words to the large language model for processing to obtain the inspection results.

[0010] In one embodiment of the present invention, the step of obtaining status data of target resources in the container platform and filtering out abnormal target objects based on the status data includes: Call the container platform APIserver interface service to obtain a list of resources and their status; The list and its status are traversed and analyzed to identify entries with abnormal status, resulting in a list of abnormal hosts or abnormal services containing abnormal resource information and their status, which are used as the target objects of the anomalies.

[0011] In one embodiment of the present invention, the step of querying log data, event data, and performance indicator data associated with the abnormal target object, and assembling the status data, log data, event data, and performance indicator data to obtain multidimensional operational data includes: Based on the name of the target object of the anomaly, query the log data within the last 5 minutes in the ES log system, which stores host operating system logs and container platform service logs; Query the data of events with a severity level within the last 5 minutes in the Elasticsearch event system, which stores all event data of the container platform. The system queries the latest performance metrics data in the Prometheus time-series database that stores performance metrics data. The retrieved log data, event data, performance metrics data, and status data are then concatenated and assembled according to a fixed format to generate multi-dimensional operational data that includes a list of abnormal hosts, service names, log content, event causes, event content, and the latest performance metrics.

[0012] In one embodiment of the present invention, the step of matching corresponding anomaly diagnostic knowledge from a preset expert knowledge base according to the type of the anomaly target object, wherein the expert knowledge base stores common anomaly descriptions and handling schemes for different types of resources, including: Identify the resource category to which the target object of the anomaly belongs; Based on the resource category, common anomaly causes and solutions for the corresponding resource type are retrieved from the expert knowledge base, and the retrieved common anomaly causes and solutions are used as anomaly diagnostic knowledge that matches the type of the target object of the anomaly.

[0013] In one embodiment of the present invention, it further includes: Parse the metadata tags of the target object of the anomaly to determine the resource category corresponding to the target object of the anomaly; Based on the determined resource category, retrieve the anomaly causes and solutions pointed to by the anomaly characteristics of the corresponding type of resource in the expert knowledge base; Only the expert knowledge entries corresponding to the resource types involved in the current inspection item are output as the anomaly diagnosis knowledge.

[0014] In one embodiment of the present invention, the step of filling the multidimensional operational data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and sending the inspection prompt words to a large language model for processing to obtain inspection results, includes: Construct a prompt word template that includes role settings, analysis requirements, information placeholders, and knowledge base placeholders. The role is set as a professional K8S operations engineer, and the analysis requirements are to perform accurate and rigorous anomaly diagnosis based on resource status information, log data, event data, and performance indicators. The multidimensional operational data is filled into the information placeholders, and the anomaly diagnosis knowledge is filled into the knowledge base placeholders to generate complete inspection prompt words; The inspection prompts are sent to the large language model, and the diagnostic analysis and handling suggestions for the abnormal target object returned by the large language model are received as the inspection results. The inspection results are then sent to the operation and maintenance personnel.

[0015] In one embodiment of the present invention, it further includes: The configuration includes multiple inspection items such as host status inspection, Pod status inspection, workload status inspection, and Etcd status inspection to form an inspection toolset; Set up a periodic automatic execution strategy or a manual trigger command. When the preset execution cycle is detected or a manual trigger command is received, the inspection toolset is invoked to execute all inspection items in parallel or serially to complete a comprehensive inspection of the container platform's key resources and services.

[0016] To achieve the above objectives, a second aspect of the present invention provides a container platform inspection device based on a large language model, comprising: The anomaly filtering module is used to obtain the status data of the target resources in the container platform and filter out the abnormal target objects based on the status data. The data assembly module is used to query log data, event data, and performance indicator data associated with the target object of the anomaly, and to assemble the status data, log data, event data, and performance indicator data to obtain multidimensional operational data. The knowledge matching module is used to match the corresponding abnormal diagnosis knowledge from a preset expert knowledge base according to the type of the target object of the abnormality. The expert knowledge base stores common abnormal descriptions and handling solutions for different types of resources. The model diagnosis module is used to fill the multidimensional operational data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and send the inspection prompt words to the large language model for processing to obtain the inspection results.

[0017] This invention discloses a container platform inspection method and apparatus based on a large language model. By integrating multi-dimensional operational data and an expert knowledge base to generate prompt words, it achieves intelligent diagnosis and analysis of container platform anomalies, effectively improving the accuracy of inspections and operational efficiency.

[0018] To achieve the above objectives, a third aspect of this application provides a computer device, including a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading executable program code stored in the memory, for implementing a container platform inspection method based on a large language model as described in the first aspect embodiment.

[0019] To achieve the above objectives, a fourth aspect of this application proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a container platform inspection method based on a large language model as described in the first aspect embodiment.

[0020] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0021] Figure 1 This is a flowchart of a container platform inspection method based on a large language model according to an embodiment of the present invention; Figure 2 This is an architecture diagram of a container platform inspection method based on a large language model according to an embodiment of the present invention; Figure 3 This is a structural diagram of a container platform inspection device based on a large language model according to an embodiment of the present invention; Figure 4 It is a computer device according to an embodiment of the present invention. Detailed Implementation

[0022] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] The following description, with reference to the accompanying drawings, describes a container platform inspection method and apparatus based on a large language model according to an embodiment of the present invention.

[0025] Figure 1 This is a flowchart of a container platform inspection method based on a large language model according to an embodiment of the present invention, such as... Figure 1 As shown, it includes: S1, Obtain the status data of the target resource in the container platform, and filter out abnormal target objects based on the status data; Specifically, the process involves acquiring the status data of target resources within the container platform and filtering out abnormal target objects based on this data. This step aims to identify the scope of anomalies requiring in-depth diagnosis by perceiving and initially assessing the real-time status of various running entities within the container platform. The core technology involves establishing a communication connection with the container platform's interface service to obtain a list of target resources and their corresponding status information. Then, based on a preset status judgment logic, the acquired data is analyzed to identify target objects in abnormal operating states. The target resources encompass various entity types within the container platform, including hosts, workloads, storage components, and network components. The status data reflects the current operational health of these entities. Through this step, the system can quickly locate fault points from a massive amount of platform resources, providing accurate input objects for subsequent multi-dimensional data correlation analysis. As a specific implementation method, the container platform's API server interface can be called to obtain a list of resources and their statuses. A filtering mechanism can then be used to directly extract information about resources with abnormal statuses and their statuses, such as identifying abnormal host nodes node01 or node02, which can then be used as the benchmark for subsequent log, event, and performance indicator data queries.

[0026] This step, by proactively acquiring and filtering status data, enables rapid location and scope convergence of abnormal targets on the container platform, effectively avoiding the waste of resources in full data processing, improving the targeting and execution efficiency of the inspection process, and laying a solid data foundation for subsequent accurate diagnosis based on large language models.

[0027] S2, query the log data, event data, and performance indicator data associated with the target object of the anomaly, and assemble the status data, log data, event data, and performance indicator data to obtain multi-dimensional operating data; Specifically, the process involves querying log data, event data, and performance metrics associated with the abnormal target object, and then assembling these data to create multidimensional operational data. The aim is to construct an information set that comprehensively reflects the operational context of the target resource. This step establishes a mapping relationship between the abnormal target object and multidimensional operational data, retrieving relevant data records generated within a preset time window from the distributed storage system or monitoring system. Log data includes runtime text records of the host operating system and service components; event data contains system status changes and alarm information; and performance metrics represent resource utilization and time-series trends. Subsequently, a standardized data encapsulation mechanism is used to logically integrate the retrieved multi-source heterogeneous data with the initially acquired status data, forming structured multidimensional operational data. This ensures data integrity and relevance, providing a unified data input foundation for subsequent diagnostic analysis. As one implementation method, the container platform interface service can be invoked to determine the list of abnormal hosts, the log content within the last five minutes can be queried from the log system based on the service name, the event data of the level of severity within the same time range can be filtered from the event system, and the latest core performance indicators such as CPU utilization and memory utilization can be obtained from the time series database. Finally, the above-mentioned data are assembled into detailed data containing an abnormal list, detailed logs, event causes and performance values ​​in a fixed format.

[0028] This step effectively solves the technical problem that single-state data cannot fully present the system fault context through multi-dimensional data association query and standardized assembly, ensuring that the information input into the large language model has comprehensive, timely and structured characteristics, thereby significantly improving the accuracy of anomaly diagnosis and operation and maintenance efficiency.

[0029] S3, Match the corresponding abnormal diagnosis knowledge from the preset expert knowledge base according to the type of the target object of the abnormality. The expert knowledge base stores common abnormal descriptions and handling solutions for different types of resources. Specifically, in the step of matching corresponding anomaly diagnostic knowledge from a pre-defined expert knowledge base based on the type of the anomaly target object, the core lies in establishing a mapping and association mechanism between resource types and domain-specific expert knowledge to achieve targeted enhancement of the input information of the large language model. The expert knowledge base constructed in this step serves as a structured data storage unit, pre-including common anomaly patterns, potential cause analyses, and standardized processing solutions for different types of resources in the container platform. During the matching operation, the system searches and filters the knowledge base based on the specific type identifier of the anomaly target object determined in the previous steps, aiming to extract diagnostic rules or experience data highly relevant to that type, thereby forming anomaly diagnostic knowledge that can guide subsequent model reasoning. This matching process is not a simple keyword search, but rather a logical association based on the object classification system, ensuring that the acquired knowledge content accurately covers typical fault scenarios that the type of object may face. By introducing this mechanism, general operational experience can be transformed into contextual information that the model can understand, effectively reducing the reasoning search space of the large language model. As a specific implementation method, the expert knowledge base may contain descriptions and troubleshooting solutions for memory overflow anomalies of host-type resources, or causes and solutions for scheduling failure events of Pod-type resources. The system automatically retrieves the corresponding knowledge entries from the database as matching results based on whether the current inspection object is a host or a Pod. The specific construction form and storage medium of the knowledge base do not constitute a limitation on the scope of protection of this step.

[0030] This step significantly improves the accuracy and professionalism of anomaly diagnosis during the inspection process by introducing a pre-set expert knowledge base and performing type-based matching. It enables the large language model to reason based on high-quality domain knowledge that has been selected, effectively avoiding misjudgments caused by generalization of general knowledge, while also reducing the computational resource consumption caused by the model processing irrelevant information.

[0031] S4. Fill the multidimensional operation data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and send the inspection prompt words to the large language model for processing to obtain the inspection results.

[0032] Specifically, multidimensional operational data and anomaly diagnostic knowledge are populated into a prompt word template to generate inspection prompt words. These prompt words are then sent to a large language model for processing to obtain inspection results. This step aims to construct a standardized input context for the large language model, triggering its reasoning ability based on pre-trained knowledge and contextual information. The core of this process lies in semantically fusing heterogeneous operational status information with domain expert knowledge through a pre-defined logical structure, forming a prompt word template with a complete diagnostic background. The prompt word template defines role settings, task objectives, and analysis constraints. By using the assembled multidimensional operational data as the object to be analyzed and the matched anomaly diagnostic knowledge as a reference, it dynamically populates the corresponding positions in the template, thereby generating inspection prompt words containing specific fault scenarios and theoretical support. Subsequently, these inspection prompt words are sent as input parameters to the large language model. Utilizing the natural language understanding and logical deduction capabilities of the large language model, the input information is comprehensively evaluated, and inspection results containing anomaly cause analysis and handling suggestions are output. As one implementation method, the prompt word template can be set to require the model to play the role of a professional operations and maintenance engineer, and to fill in the template information field with detailed data such as the status, logs, events and performance indicators of the host or service, as well as common anomaly descriptions and solutions for this type of resource, and then request the large language model to return a rigorous diagnostic conclusion.

[0033] This step effectively combines the experience of operations and maintenance experts with the general reasoning capabilities of large language models through structured prompt word engineering. This not only ensures the completeness and relevance of the input information, but also significantly improves the accuracy and interpretability of automated inspection results, making the anomaly diagnosis process of the container platform more intelligent and in line with professional standards.

[0034] This invention proposes a container platform inspection method based on a large language model. By collecting and assembling resource status data, event data, log data, and performance data, and employing a common container platform anomaly expert knowledge base, prompt word generation, and request large language model processing, it achieves inspection operations on critical resources and services of the container platform. The entire solution is as follows: Figure 2 The system includes three core components: multidimensional data collection and assembly, construction of an expert knowledge base for common container platform anomalies, and large language model inspection.

[0035] Specifically, multi-dimensional data collection and assembly includes: Service or resource status data: Starting with status data, the container platform API server is called to obtain a list of resources and their statuses, filtering out abnormal resource information and their statuses. Log data: The log system has collected host operating system logs and container platform service logs, and stored them in Elasticsearch (ES). Log data for the last 5 minutes is queried based on the service or resource name. Event data: The event system has collected all event data from the container platform, and stored it in ES. Event data with a severity level for the last 5 minutes is queried based on the service or resource name. Performance metric data: The monitoring system has collected performance metric data, and stored it in the Prometheus time-series database. The latest performance metrics are queried based on the service or resource name. For example, for hosts, core performance metrics such as CPU utilization and memory utilization are queried. Data assembly: The above information is assembled into detailed data for this resource, using host status inspection as an example.

[0036] Specifically as follows: The list of abnormal hosts is as follows: node01, node02 The log information for host node1 in the last 5 minutes is as follows: Service name: kernel, Log content: outofmemory Service name: systemd, Log content: startedtimeservice The event information for host node1 in the last 5 minutes is as follows: Event cause: Rebooted event content: Node node1 has been rebooted The latest performance metrics for host node1 are as follows: CPU utilization: 85.01% Memory utilization: 88.1% Furthermore, an expert knowledge base for common container platform anomalies is built, collecting common anomalies to construct the knowledge base, with examples for host and Pod categories: Host-related: "OutofMemory" error in logs. Common causes and solutions: An abnormal program is consuming excessive memory, or the operating system kernel has crashed. "DiskPressure" error in event logs. Common causes and solutions: Check if a large number of exception logs have been generated. If a large number of files are not cleaned up on the disk, clean them up. Pod-related: "FailedScheduling" error in event logs. Common causes and solutions: Insufficient available CPU and memory resources on cluster nodes; too few nodes to meet the anti-affinity requirements of the Pod component. When selecting the knowledge base, due to output length limitations, select the appropriate expert knowledge based on the resource type involved in the inspection item.

[0037] Further, the large language model inspection involves: Prompt word assembly: The prompt word template reads: "You are a professional Kubernetes operations engineer with extensive experience. You can diagnose anomalies based on resource status information, log data, event data, and performance metrics. Please analyze the following information accurately and rigorously; otherwise, you will be penalized. The information is as follows: {Data assembled from multi-dimensional data collection and assembly}. The common anomaly knowledge base is as follows: {Expert knowledge base assembled from building a common anomaly expert knowledge base for container platforms}". The expert knowledge base from multi-dimensional data assembly and filtering is then populated into the prompt word template to generate the prompt word. Requesting the large model: The prompt word is sent to the large model, and the large model's returned results are sent to the operations personnel.

[0038] Furthermore, multiple inspection items form an inspection toolset, which describes the processing flow of an inspection item and realizes multiple inspection items, such as host status inspection, Pod status inspection, workload status inspection, Etcd status inspection, etc., forming an inspection toolset that supports periodic automatic execution or manual triggering of all inspection items.

[0039] This invention enables the inspection of critical resources and services on container platforms by collecting and assembling resource status data, event data, log data, and performance data, and processing it through a process including a common container platform anomaly expert knowledge base, prompt word generation, and requesting a large language model. This invention supports multi-dimensional data collection and assembly. It obtains the status of resources associated with the inspection items based on container platform interface services; it obtains collected event data based on the event system; it obtains collected log data based on the log system; and it obtains collected performance data based on the monitoring system. The collected data is then assembled according to a fixed format. This invention supports a common container platform anomaly expert knowledge base. It collects existing descriptions of common anomalies to establish a common container platform anomaly expert knowledge base. This invention supports prompt word generation and large language model inspection. The assembled multi-dimensional data and the common container platform anomaly expert knowledge base are combined to form prompt words, and the large language model is requested to obtain the inspection results and send them to the operations and maintenance personnel.

[0040] To achieve the above embodiments, such as Figure 3 As shown, this embodiment also provides a container platform inspection device 10 based on a large language model, including: An anomaly filtering module 100 is used to obtain the status data of target resources in the container platform and filter out abnormal target objects based on the status data. The data assembly module 200 is used to query log data, event data, and performance indicator data associated with the target object of the anomaly, and to assemble the status data, log data, event data, and performance indicator data to obtain multidimensional operational data. The knowledge matching module 300 is used to match corresponding abnormal diagnosis knowledge from a preset expert knowledge base according to the type of the target object of the abnormality. The expert knowledge base stores common abnormal descriptions and handling solutions for different types of resources. The model diagnosis module 400 is used to fill the multidimensional running data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and send the inspection prompt words to the large language model for processing to obtain the inspection results.

[0041] This invention discloses a container platform inspection device based on a large language model. By integrating multi-dimensional operational data and an expert knowledge base to generate prompt words, it achieves intelligent diagnosis and analysis of container platform anomalies, effectively improving the accuracy of inspections and operational efficiency.

[0042] To implement the methods of the above embodiments, the present invention also provides a computer device, such as... Figure 4 As shown, the computer device 600 includes a memory 601 and a processor 602; wherein, the processor 602 reads the executable program code stored in the memory 601 to run a program corresponding to the executable program code, so as to implement the various steps of the container platform inspection method based on a large language model described above.

[0043] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a container platform inspection method based on a large language model as described in the foregoing embodiments.

[0044] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0045] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A container platform inspection method based on a large language model, characterized in that, include: S1, Obtain the status data of the target resource in the container platform, and filter out abnormal target objects based on the status data; S2, query the log data, event data, and performance indicator data associated with the target object of the anomaly, and assemble the status data, log data, event data, and performance indicator data to obtain multi-dimensional operating data; S3, Match the corresponding abnormal diagnosis knowledge from the preset expert knowledge base according to the type of the target object of the abnormality. The expert knowledge base stores common abnormal descriptions and handling solutions for different types of resources. S4. Fill the multidimensional operation data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and send the inspection prompt words to the large language model for processing to obtain the inspection results.

2. The method as described in claim 1, characterized in that, The step of acquiring the status data of the target resource in the container platform and filtering out abnormal target objects based on the status data includes: Call the container platform APIserver interface service to obtain a list of resources and their status; The list and its status are traversed and analyzed to identify entries with abnormal status, resulting in a list of abnormal hosts or abnormal services containing abnormal resource information and their status, which are used as the target objects of the anomalies.

3. The method as described in claim 1, characterized in that, The query retrieves log data, event data, and performance metrics data associated with the abnormal target object, and assembles the status data, log data, event data, and performance metrics data to obtain multidimensional operational data, including: Based on the name of the target object of the anomaly, query the log data within the last 5 minutes in the ES log system, which stores host operating system logs and container platform service logs; Query the data of events with a severity level within the last 5 minutes in the Elasticsearch event system, which stores all event data of the container platform. The system queries the latest performance metrics data in the Prometheus time-series database that stores performance metrics data. The retrieved log data, event data, performance metrics data, and status data are then concatenated and assembled according to a fixed format to generate multi-dimensional operational data that includes a list of abnormal hosts, service names, log content, event causes, event content, and the latest performance metrics.

4. The method as described in claim 1, characterized in that, The step involves matching corresponding anomaly diagnostic knowledge from a preset expert knowledge base based on the type of the target object of the anomaly. This expert knowledge base stores common anomaly descriptions and handling solutions for different types of resources, including: Identify the resource category to which the target object of the anomaly belongs; Based on the resource category, common anomaly causes and solutions for the corresponding resource type are retrieved from the expert knowledge base, and the retrieved common anomaly causes and solutions are used as anomaly diagnostic knowledge that matches the type of the target object of the anomaly.

5. The method as described in claim 4, characterized in that, The method further includes: Parse the metadata tags of the target object of the anomaly to determine the resource category corresponding to the target object of the anomaly; Based on the determined resource category, retrieve the anomaly causes and solutions pointed to by the anomaly characteristics of the corresponding type of resource in the expert knowledge base; Only the expert knowledge entries corresponding to the resource types involved in the current inspection item are output as the anomaly diagnosis knowledge.

6. The method as described in claim 1, characterized in that, The process of filling the multidimensional operational data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and sending the inspection prompt words to the large language model for processing to obtain inspection results, includes: Construct a prompt word template that includes role settings, analysis requirements, information placeholders, and knowledge base placeholders. The role is set as a professional K8S operations engineer, and the analysis requirements are to perform accurate and rigorous anomaly diagnosis based on resource status information, log data, event data, and performance indicators. The multidimensional operational data is filled into the information placeholders, and the anomaly diagnosis knowledge is filled into the knowledge base placeholders to generate complete inspection prompt words; The inspection prompts are sent to the large language model, and the diagnostic analysis and handling suggestions for the abnormal target object returned by the large language model are received as the inspection results. The inspection results are then sent to the operation and maintenance personnel.

7. The method as described in claim 1, characterized in that, The method further includes: The configuration includes multiple inspection items such as host status inspection, Pod status inspection, workload status inspection, and Etcd status inspection to form an inspection toolset; Set up a periodic automatic execution strategy or a manual trigger command. When the preset execution cycle is detected or a manual trigger command is received, the inspection toolset is invoked to execute all inspection items in parallel or serially to complete a comprehensive inspection of the container platform's key resources and services.

8. A container platform inspection device based on a large language model, characterized in that, include: The anomaly filtering module is used to obtain the status data of the target resources in the container platform and filter out the abnormal target objects based on the status data. The data assembly module is used to query log data, event data, and performance indicator data associated with the target object of the anomaly, and to assemble the status data, log data, event data, and performance indicator data to obtain multidimensional operational data. The knowledge matching module is used to match the corresponding abnormal diagnosis knowledge from a preset expert knowledge base according to the type of the target object of the abnormality. The expert knowledge base stores common abnormal descriptions and handling solutions for different types of resources. The model diagnosis module is used to fill the multidimensional operational data and the anomaly diagnosis knowledge into the prompt word template to generate inspection prompt words, and send the inspection prompt words to the large language model for processing to obtain the inspection results.

9. A computer device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement a container platform inspection method based on a large language model as described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a container platform inspection method based on a large language model as described in any one of claims 1-7.