Fault processing method and device, storage medium, program product and electronic equipment

By obtaining and processing the fault information of the target platform, determining the reference text shard from multiple text shards, and using the fault processing model to generate a processing solution, the problem of low fault processing reliability in the existing technology is solved, and more accurate and efficient fault processing is achieved.

CN120196728AActive Publication Date: 2025-06-24ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510648788.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-24
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

When a platform fails, the reliability of fault processing in the prior art is low, and relying on manual processing is prone to misjudgment or omissions. Especially in complex fault diagnosis and emergency situations, it is difficult to ensure the accuracy of fault processing.

Method used

By obtaining the fault information of the system failure of the target platform, determining the reference text shard from multiple text shards based on the information, processing these text shards and fault information using the fault processing model, generating a target processing solution for the system failure of the target platform, and troubleshooting the target platform based on the scheme.

Benefits of technology

It improves the reliability of fault processing, and combines dynamic knowledge retrieval and generation model to generate more accurate and relevant fault processing solutions, reducing the possibility of manual misjudgment and omissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196728A_ABST
    Figure CN120196728A_ABST
Patent Text Reader

Abstract

The invention discloses a fault processing method and device, a storage medium, a program product and electronic equipment. The method comprises the steps of obtaining fault information of a system fault of a target platform, and obtaining target fault information; at least one reference text fragment is determined from the multiple text fragments based on the target fault information, and different text fragments comprise fault information of different system faults and fault processing schemes; the at least one reference text fragment and the target fault information are processed through a fault processing model, target information is obtained, and the target information at least comprises a target processing scheme for a system fault of the target platform; and performing fault processing on the target platform based on a target processing scheme in the target information. The technical problem that the fault processing reliability is low when the platform breaks down in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a method, apparatus, storage medium, program product, and electronic device for processing faults. Background Art

[0002] In the modern information technology system, a platform (such as an intelligent computing cloud platform), as the core infrastructure for data processing and job support, its stability and high availability are crucial. However, during the operation of the platform, various faults will inevitably occur. These faults not only affect the continuity of services but may also cause serious damage to user data and job operations. Currently, in the related art, when a platform fails, it usually relies on manual processing of system faults. This method depends on personal experience and level, and is prone to misjudgment or omission. Especially in complex fault diagnosis and emergency situations, manual processing may be difficult to ensure the accuracy of fault handling due to pressure and time constraints, thus there is a problem of low reliability in fault handling.

[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, storage medium, program product, and electronic device for processing faults, so as to at least solve the technical problem of low reliability in fault handling when a platform fails in the related art.

[0005] According to one aspect of the embodiments of this application, a method for processing faults is provided, including: obtaining fault information of a system fault of a target platform to obtain target fault information; determining at least one reference text slice from multiple text slices based on the target fault information, where different text slices contain fault information and fault handling solutions for different system faults; processing at least one reference text slice and the target fault information through a fault handling model to obtain target information, where the target information at least includes a target handling solution for the system fault of the target platform; performing fault handling on the target platform based on the target handling solution in the target information.

[0006] Further, before determining at least one reference text slice from multiple text slices based on the target fault information, the method further includes: obtaining an operation and maintenance knowledge document of the target platform; for the system faults recorded in the operation and maintenance knowledge document, extracting the fault analysis results and fault handling solutions of the system faults from the operation and maintenance knowledge document; determining fault information according to the system logs corresponding to the system faults; generating different text slices according to the fault analysis results, fault handling solutions, and fault information of different system faults to obtain multiple text slices.

[0007] Further, determine a query text vector according to the target fault information; obtain the text vectors corresponding to the text slices in multiple text slices; determine at least one reference text slice from the multiple text slices according to the similarity between the query text vector and the text vectors.

[0008] Further, generate a first prompt statement according to at least one reference text slice, the target fault information, and a preset first prompt template, where the first prompt statement is at least used to guide a fault handling model to generate a target fault analysis result and a target handling solution according to at least one reference text slice and the target fault information; input the first prompt statement into the fault handling model, and obtain the target fault analysis result and the target handling solution through the fault handling model; determine the target fault analysis result and the target handling solution as the target information.

[0009] Further, generate a second prompt statement according to the target information and a preset second prompt template, where the second prompt statement is used to guide the fault handling model to determine the source of the target information according to multiple text slices; input the second prompt statement into the fault handling model, and obtain the source information corresponding to the target information through the fault handling model; when the source information indicates that the target information belongs to multiple text slices, perform fault handling on the target platform based on the target handling solution.

[0010] Further, the text slice includes the function identifier of the executable function corresponding to the fault handling solution. When the source information indicates that the target information belongs to multiple text slices, performing fault handling on the target platform based on the target handling solution includes: determining the executable function corresponding to the target handling solution based on the function identifier corresponding to the target handling solution in the target information and the attribute information in the target database, where the target database is used to store the attribute information of the executable function; performing fault handling on the target platform based on the executable function corresponding to the target handling solution.

[0011] Further, the target handling solution at least includes a fault repair method. Performing fault handling on the target platform based on the target handling solution in the target information includes: performing fault repair processing on the target platform based on the first executable function in the executable function corresponding to the target handling solution, where the first executable function is used to execute the fault repair method.

[0012] Further, after performing a fault repair process on the target platform based on the first executable function in the executable function corresponding to the target processing solution, the method further includes: if the target processing solution further includes a repair verification method, performing a repair verification process on the target platform based on the second executable function in the executable function to obtain a repair verification result, where the second executable function is used to execute the repair verification method; in the case where the repair verification result indicates a repair failure, if the target processing solution further includes a rollback method, performing a rollback process on the target platform based on the third executable function in the executable function, where the third executable function is used to execute the rollback method.

[0013] Further, performing a fault process on the target platform based on the target processing solution in the target information includes: obtaining the fault level of the system fault of the target platform; determining the execution time point of the target processing solution based on the fault level; and performing a fault process on the target platform based on the target processing solution at the execution time point.

[0014] According to another aspect of the embodiments of the present application, there is also provided a method for processing a fault, including: obtaining fault information of a system fault of a target platform uploaded by a client to obtain target fault information; determining at least one reference text slice from multiple text slices based on the target fault information in a cloud server, where different text slices contain fault information of different system faults and fault processing solutions; processing at least one reference text slice and the target fault information through a fault processing model to obtain target information, where the target information at least includes a target processing solution for the system fault of the target platform; and feeding back the target information to the client, where the target information is used to instruct the client to perform a fault process on the target platform based on the target processing solution.

[0015] According to another aspect of the embodiments of the present application, there is also provided a device for processing a fault, including: a first obtaining unit, configured to obtain fault information of a system fault of a target platform to obtain target fault information; a determining unit, configured to determine at least one reference text slice from multiple text slices based on the target fault information, where different text slices contain fault information of different system faults and fault processing solutions; a first processing unit, configured to process at least one reference text slice and the target fault information through a fault processing model to obtain target information, where the target information at least includes a target processing solution for the system fault of the target platform; and a second processing unit, configured to perform a fault process on the target platform based on the target processing solution in the target information.

[0016] Further, the fault processing device further includes: a second acquisition unit, configured to acquire the operation and maintenance knowledge document of the target platform; a first extraction unit, configured to extract, from the operation and maintenance knowledge document, the fault analysis result and the fault handling solution of the system fault recorded in the operation and maintenance knowledge document; a second extraction unit, configured to determine fault information according to the system log corresponding to the system fault; a generation unit, configured to generate different text shards according to the fault analysis results, the fault handling solutions, and the fault information of different system faults, so as to obtain a plurality of text shards.

[0017] Further, the determination unit includes: a first determination subunit, configured to determine a query text vector according to the target fault information; a first acquisition subunit, configured to acquire the text vectors corresponding to the text shards in the plurality of text shards; a second determination subunit, configured to determine at least one reference text shard from the plurality of text shards according to the similarity between the query text vector and the text vectors.

[0018] Further, the first processing unit includes: a first generation subunit, configured to generate a first prompt statement according to at least one reference text shard, the target fault information, and a preset first prompt template, where the first prompt statement is at least used to guide the fault processing model to generate a target fault analysis result and a target handling solution according to at least one reference text shard and the target fault information; a first processing subunit, configured to input the first prompt statement into the fault processing model to obtain the target fault analysis result and the target handling solution through the fault processing model; a third determination subunit, configured to determine the target fault analysis result and the target handling solution as target information.

[0019] Further, the second processing unit includes: a second generation subunit, configured to generate a second prompt statement according to the target information and a preset second prompt template, where the second prompt statement is used to guide the fault processing model to determine the source of the target information according to the plurality of text shards; a second processing subunit, configured to input the second prompt statement into the fault processing model to obtain the source information corresponding to the target information through the fault processing model; a third processing subunit, configured to perform fault processing on the target platform based on the target handling solution when the source information indicates that the target information belongs to the plurality of text shards.

[0020] Further, the text shards include the function identifier of the executable function corresponding to the fault handling solution, and the third processing subunit includes: a determination module, configured to determine the executable function corresponding to the target handling solution based on the function identifier corresponding to the target handling solution in the target information and the attribute information in the target database, where the target database is used to store the attribute information of the executable function; a processing module, configured to perform fault processing on the target platform based on the executable function corresponding to the target handling solution.

[0021] Furthermore, the target processing solution at least includes a fault repair method. The second processing unit includes: a fourth processing subunit, configured to perform a fault repair process on the target platform based on a first executable function in an executable function corresponding to the target processing solution, where the first executable function is used to execute the fault repair method.

[0022] Furthermore, the fault processing device further includes: a third processing unit, configured to, if the target processing solution further includes a repair verification method, perform a repair verification process on the target platform based on a second executable function in the executable function to obtain a repair verification result, where the second executable function is used to execute the repair verification method; a fourth processing unit, configured to, when the repair verification result indicates a repair failure, if the target processing solution further includes a rollback method, perform a rollback process on the target platform based on a third executable function in the executable function, where the third executable function is used to execute the rollback method.

[0023] Furthermore, the second processing unit includes: a second acquisition subunit, configured to acquire the fault level of the system fault of the target platform; a fourth determination subunit, configured to determine the execution time point of the target processing solution based on the fault level; a fifth processing subunit, configured to perform a fault process on the target platform based on the target processing solution at the execution time point.

[0024] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a memory storing an executable program; a processor configured to run the program, where when the program runs, it executes the fault processing method of any one of the above.

[0025] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium storing a program, where when the program runs, it controls the device where the storage medium is located to execute the fault processing method of any one of the above.

[0026] According to another aspect of the embodiments of the present invention, there is also provided a computer program product including a computer program, where when the computer program is executed by a processor, it implements the fault processing method of any one of the above.

[0027] In the embodiments of the present application, by obtaining the fault information of the system fault of the target platform, the target fault information is obtained; based on the target fault information, at least one reference text fragment is determined from multiple text fragments, wherein different text fragments contain the fault information and fault handling solutions of different system faults; the fault handling model processes at least one reference text fragment and the target fault information to obtain the target information, wherein the target information at least includes the target handling solution for the system fault of the target platform; based on the target handling solution in the target information, the target platform is fault-handled. By combining the retrieval enhancement technology, when the fault information of the system of the target platform is obtained, relevant knowledge is first extracted from multiple text fragments through information retrieval, and then the obtained relevant knowledge is injected into the process of generating the fault handling model. This "retrieval-generation" collaborative mode not only retains the logical reasoning and natural language generation capabilities of the model but also breaks through the limitations of its static knowledge boundary, improving the accuracy of the generated target information, thereby improving the reliability of fault handling. Further, by setting different text fragments to contain the fault information and fault handling solutions of different system faults, the relevant information of the same system fault is summarized in the same text fragment, improving the relevance of the content in the same text fragment, thereby improving the retrieval accuracy during the retrieval process, achieving the purpose of combining the retrieval enhancement technology to generate a fault handling solution based on the text fragments of the distinguished fault knowledge for fault handling, realizing the technical effect of improving the reliability of fault handling, and further solving the technical problem of low reliability of fault handling in the related art when the platform fails. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0029] Figure 1 is a schematic diagram of a computer terminal provided in Embodiment 1 of the present application;

[0030] Figure 2 is the flowchart of the fault handling method provided in Embodiment 1 of the present application Figure 1 ;

[0031] Figure 3 is the flowchart of the fault handling method provided in Embodiment 1 of the present application Figure 2 ;

[0032] Figure 4 is the flowchart of the fault handling method provided in Embodiment 2 of the present application;

[0033] Figure 5It is a schematic diagram of a fault processing device provided in Embodiment 3 of the present application;

[0034] Figure 6 It is a structural block diagram of an electronic device provided in Embodiment 4 of the present application. Detailed implementation manners

[0035] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0037] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards in the relevant regions, and corresponding operation entrances are provided for the user to select authorization or rejection.

[0038] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0039] RAG (Retrieval-Augmented Generation): By combining dynamic knowledge retrieval with a generation model, the generation model can access external documents when answering questions, so as to provide accurate and rich responses based on relatively new and context-related knowledge.

[0040] Vector knowledge base: The vector knowledge base stores text data in the form of high-dimensional vectors, making semantic retrieval and similarity calculation more efficient and accurate, thereby enhancing the intelligent level of information retrieval.

[0041] Text chuncks: It refers to splitting long text into smaller semantic units or segments to more precisely match user queries and perform semantic analysis during information retrieval and processing, thereby enhancing the response accuracy and context relevance of the model.

[0042] Embodiment 1

[0043] According to the embodiments of the present application, a method for processing faults is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0044] The method embodiment provided by the first embodiment of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for processing faults is shown. As Figure 1 shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include, but is not limited to, processing devices such as a microcontroller unit (MCU) or a field programmable gate array (FPGA), and the processor set 102 may include a processor set, Figure 1 which is shown as 102a, 102b,..., 102n in it), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in it, or have a different configuration from Figure 1 shown.

[0045] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device).

[0046] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the fault processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned fault processing method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0047] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0048] The display can be a touch-screen liquid crystal display, which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0049] In the modern information technology system, a platform (such as an intelligent computing cloud platform) serves as the core infrastructure for data processing and job support, and its stability and high availability are of crucial importance. However, during the operation of the platform, various faults will inevitably occur. These faults not only affect the continuity of services but may also cause serious damage to user data and job operations. Currently, in related technologies, when a platform fails, it usually relies on manual handling of system faults. This method depends on personal experience and level and is prone to misjudgment or omission. Especially in complex fault diagnosis and emergency situations, manual handling may be difficult to ensure the accuracy of fault handling due to pressure and time constraints, thus resulting in the problem of low reliability of fault handling.

[0050] Against the above technical background, this application provides a method for handling faults as Figure 2 shown. Figure 2 is a flowchart of the method for handling faults provided in Embodiment 1 of this application Figure 1 . As Figure 2 shown, the method includes:

[0051] Step S201, obtain the fault information of the system fault of the target platform to obtain the target fault information.

[0052] Optionally, devices such as electronic devices, application systems, and servers can be used as the execution subject of this application. In this embodiment, the target processing system is used as the execution subject to execute the foregoing method for handling faults.

[0053] Optionally, the target platform refers to a platform with data processing capabilities. For example, the target platform can be an intelligent computing cloud platform, or a data center, or an Internet of Things platform, etc. The target platform varies according to different actual application scenarios, so no specific limitation is made here.

[0054] Optionally, the system faults of the target platform include but are not limited to faults at the hardware, software, or network levels detected in the target platform. The fault information of the system fault can be extracted from the system log. For example, when a fault occurs during the operation of the target platform, the target platform or a detection system used to detect faults in the target platform issues a warning message. After the target processing system receives this warning message, it collects the system warning log (also called the system log) of the target platform and extracts keywords from the system warning log, thereby determining the keywords as the fault information of the system fault of the target platform, that is, determining it as the target fault information.

[0055] Optionally, the extracted keywords are used to describe the fault characteristics of the system fault. The keywords may include at least one of the following: error code, error message, system component name, timestamp, resource status, operation or command, hostname or IP address, process ID, port number, etc. Among them, a system error or exception usually has a specific error code, which is issued by the system or application when encountering a problem and is an identifier used to indicate the error type or location. The error message is used to describe the specific situation of the fault occurrence. For example, memory overflow, connection timeout, file system corruption, certificate expiration, etc. The system component name refers to the name of the system component where the fault occurs. The timestamp refers to the time point when the fault occurs. The resource status is used to describe the current state of the system resources, such as disk full, CPU overload, etc. The operation or command refers to the operation or command that may cause the fault or the command being executed. The hostname or IP address refers to the name or IP address of the physical host (or virtual host) where the fault occurs. The process ID refers to the ID of the process involved in the fault. The port number refers to the port number involved in the fault.

[0056] In an optional embodiment, the fault information further includes at least one of the following: the fault type of the system fault, version information. Optionally, the target processing system can determine the fault type of the target platform according to the system warning log or the keywords extracted from the log. For example, obtain the preset fault type recognition rules, where the recognition rules include the determination conditions corresponding to each fault type. When the system warning log or the keywords meet the determination conditions of a certain fault type, determine that fault type as the fault type of the target platform. Another example is to input the system warning log or the keywords into a fault classification model to determine the fault type through the fault classification model, and the fault classification model is a neural network model.

[0057] Optionally, the version information includes at least one of the following: system version number, component version number, where the system version number refers to the version number of the system where the fault occurs, and the component version number refers to the version number of the system component where the fault occurs.

[0058] In an optional embodiment, when detecting that a system fault occurs in the target platform, the target processing system can also generate a fault handling work order to record the target fault information and the processing process of the system fault in the work order for subsequent traceability.

[0059] Step S202, determine at least one reference text fragment from multiple text fragments based on the target fault information, where different text fragments contain the fault information and fault handling solutions of different system faults.

[0060] Optionally, the target processing system can utilize the RAG technology to determine at least one reference text fragment from multiple text fragments. In the RAG framework, the vector knowledge base plays a key role. The vector knowledge base is a knowledge base stored by converting text data into the form of high-dimensional vectors. A piece of text, a document, or a knowledge unit is encoded into a numerical vector, and these vectors capture the semantic information of the text, making semantic retrieval efficient and accurate. When a user initiates a query, the query content is also converted into a vector and quickly matched and similarity calculated in the knowledge base to filter out relatively relevant knowledge entries. This vector-based semantic retrieval method is more intelligent than traditional keyword matching, can understand and extract deeper semantic relationships, thus providing accurate and context-related knowledge support for the large model and significantly improving the quality and professionalism of the generated results.

[0061] Therefore, in this embodiment, the target processing system can determine the text vectors corresponding to multiple text fragments, construct a vector knowledge base based on the text vectors corresponding to the multiple text fragments. Then, match based on the vector of the target fault information in the vector knowledge base to obtain at least one reference text fragment. Optionally, the reference text fragment refers to a text fragment whose vector similarity with the target fault information is greater than a preset similarity.

[0062] Optionally, during the process of constructing the vector knowledge base, the document will be pre-cleaned and sliced. In the RAG architecture, the slicing process of the document is an important step, which includes splitting the complete document into several smaller text units, called text chunks. The purpose of doing this is to improve the accuracy and efficiency of retrieval, because smaller text units are not only convenient for matching with the user's query, but also can capture semantic details more precisely. When a user initiates a query, the system will retrieve based on the query text in the sliced document set to find the chunks that are relatively relevant to the query text in terms of semantics and content. This fine-grained retrieval method can inject more accurate and context-related knowledge during the large model generation process to ensure the accuracy and reliability of the generated results. However, the size of the chunks and the relevance of the content in a single chunk have a greater impact on the accuracy of the retrieval results.

[0063] Therefore, in order to achieve more accurate and comprehensive knowledge retrieval and improve the appropriateness of the content in the text fragments, in an optional embodiment, the fault information and the fault handling solutions belonging to the same system fault are set in the same text fragment, that is, different text fragments contain the fault information and the fault handling solutions of different system faults.

[0064] In an optional embodiment, the system failures recorded in the text fragment may be system failures for the target platform. For example, the system failures recorded in the text fragment are system failures that have occurred on the target platform. Another example is that the system failures recorded in the text fragment are system failures that may occur on the target platform.

[0065] In an optional embodiment, the system failures recorded in the text fragment may be system failures for at least one reference platform. A reference platform refers to a platform whose system architecture similarity to the target platform is greater than a preset similarity. The at least one reference platform may include the target platform. For example, the system failures recorded in the text fragment are system failures that have occurred on the reference platform.

[0066] In an optional embodiment, the fault information in the text fragment at least includes keywords extracted from the system log corresponding to the system failure. And the content form of this keyword is the same as that of the keyword in the target fault information, so it will not be elaborated here.

[0067] In an optional embodiment, in addition to the foregoing keywords, the fault information in the text fragment may further include at least one of the following: the fault type of the system failure, and version information, so as to improve the accuracy during knowledge retrieval.

[0068] In an optional embodiment, the foregoing fault handling solution at least includes a fault repair method. Among them, the fault repair method refers to a method for repairing system failures. For example, the fault repair method may be to restart the service, apply patches, adjust configuration parameters, execute specific script commands, etc.

[0069] In an optional embodiment, in addition to the fault repair method, the foregoing fault handling solution may further include at least one of a repair verification method and a rollback method. The repair verification method refers to a method for verifying whether the fault has been resolved after executing the fault repair method. For example, the repair verification method may be to recheck the system log, run diagnostic tools, execute health check scripts, etc. The rollback method is used to roll back the system of the target platform to its state before the fault repair when the fault repair fails (such as, the fault is not resolved or side effects are generated).

[0070] In an optional embodiment, the text fragment may further include a fault analysis result of the system failure, and the fault analysis result is used to characterize the cause of the system failure. For example, an optional fault analysis result may be: "Due to insufficient server disk space, log recording fails, which in turn causes the service to restart automatically."

[0071] Step S203: Process at least one reference text shard and the target fault information through a fault handling model to obtain target information, where the target information at least includes a target handling solution for the system fault of the target platform.

[0072] Optionally, the fault handling model can be a large language model. The target handling system can combine these proprietary technical knowledge (i.e., the knowledge in the text shards) with a general large model (i.e., a pre-trained large language model) based on the RAG technology to achieve automatic fault analysis and solution generation. The semantic retrieval function of RAG can accurately obtain more relevant and timely repair suggestions from the continuously updated vector knowledge base, thereby improving the efficiency and accuracy of platform fault handling.

[0073] In an optional embodiment, the target handling system can generate a prompt statement based on at least one reference text shard and the target fault information, and then input the prompt statement into the fault handling model. Thus, the fault handling model can perform fault analysis and solution generation on the target fault information according to at least one reference text shard to obtain the target information.

[0074] In an optional embodiment, the target information includes a target handling solution for the system fault of the target platform. The target handling solution at least includes a fault repair method, and may also include at least one of the following: a repair verification method, a rollback method.

[0075] In an optional embodiment, when the text shards of multiple text shards include a fault analysis result, in addition to the above-mentioned target handling solution, the target information may also include a target fault analysis result for the system fault of the target platform.

[0076] Step S204: Perform fault handling on the target platform based on the target handling solution in the target information.

[0077] Optionally, after obtaining the target information, the target handling system can perform fault handling on the target platform based on the target handling solution in the target information. For example, obtain an executable function corresponding to the target handling solution in the target information, and then execute the executable function to perform fault handling on the target platform. Wherein, the number of executable functions is at least one, and the aforementioned fault handling at least includes performing fault repair handling on the target platform.

[0078] Optionally, the executable function corresponding to the target handling solution in the target information can be obtained based on a matching method. For example, there is a function identifier corresponding to the solution in the target handling solution. The target handling system can match and obtain the attribute information of the corresponding executable function from the target database based on the function identifier, and thus determine the executable function based on the attribute information. The target database stores the attribute information of multiple executable functions.

[0079] Optionally, the executable function corresponding to the target processing solution in the target information can be obtained based on a generation method. For example, by inputting the target processing solution into a function generation model, which can be a pre-trained and fine-tuned large language model, the function generation model can generate the executable function corresponding to the target processing solution.

[0080] In this solution, by combining retrieval enhancement technology, when the fault information of the system of the target platform is obtained, relevant knowledge is first extracted from multiple text fragments through information retrieval, and then the obtained relevant knowledge is injected into the fault handling model generation process. This "retrieval-generation" collaborative mode not only retains the logical reasoning and natural language generation capabilities of the language model but also breaks through the limitation of its static knowledge boundary, improving the accuracy of the generated target information, thereby improving the reliability of fault handling. Further, by setting the fault information and fault handling solutions of different system faults in different text fragments, the relevant information of the same system fault is summarized in the same text fragment, improving the relevance of the content in the same text fragment, thereby improving the retrieval accuracy in the retrieval process, achieving the purpose of generating a fault handling solution based on the text fragments of the distinguished fault knowledge by combining retrieval enhancement technology for fault handling, realizing the technical effect of improving the reliability of fault handling, and solving the technical problem of low reliability of fault handling in the related art when the platform fails.

[0081] How to improve the accuracy of the retrieval process in RAG is crucial. Therefore, in the fault handling method provided in the first embodiment of this application, before determining at least one reference text fragment from multiple text fragments based on the target fault information, the fault handling method further includes: obtaining the operation and maintenance knowledge document of the target platform; for the system faults recorded in the operation and maintenance knowledge document, extracting the fault analysis results and fault handling solutions of the system faults from the operation and maintenance knowledge document; determining the fault information according to the system log corresponding to the system fault; generating different text fragments according to the fault analysis results, fault handling solutions, and fault information of different system faults to obtain multiple text fragments.

[0082] In an optional embodiment, the operation and maintenance knowledge document of the target platform is used to record the relevant knowledge of the system faults that have occurred and been resolved on the target platform. For example, the operation and maintenance knowledge document of the target platform includes the fault analysis results and fault handling solutions in the system fault scenarios that have occurred on the target platform. That is, the knowledge in the operation and maintenance knowledge document is verified professional knowledge and has reliability.

[0083] Optionally, the operation and maintenance knowledge document of the target platform can be maintained by the user. For example, in the early stage of maintaining the operation and maintenance knowledge document (i.e., when there is relatively little knowledge in the document), when a system failure occurs on the target platform, the user records the system failure in the operation and maintenance knowledge document, and the user analyzes to obtain the failure analysis result and the failure handling solution, and records them together in the operation and maintenance knowledge document.

[0084] Optionally, the operation and maintenance knowledge document of the target platform can be maintained by the target processing system. For example, in the later stage of maintaining the operation and maintenance knowledge document (i.e., when there is relatively much knowledge in the document), when a system failure occurs on the target platform, the target processing system determines the failure analysis result and the failure handling solution corresponding to the occurred system failure based on the RAG technology, and then performs failure handling on the target platform according to the failure handling solution. If the failure is successfully repaired, the failure analysis result and the failure handling solution corresponding to the system failure are recorded in the operation and maintenance knowledge document. Otherwise, the operation and maintenance knowledge document is not updated.

[0085] In an optional embodiment, if the operation and maintenance knowledge document contains pictures, the target processing system can use OCR (Optical Character Recognition) to extract information from the pictures in the operation and maintenance knowledge document, and then supplement the extracted text information to the operation and maintenance knowledge document to obtain an updated operation and maintenance knowledge document. For example, the extracted text information is supplemented at a position adjacent to the picture (such as below the picture, below the picture), or for another example, the corresponding picture in the operation and maintenance knowledge document is replaced with the extracted text information.

[0086] Optionally, after obtaining the operation and maintenance knowledge document (or the updated operation and maintenance knowledge document), the target processing system can extract the failure analysis result and the failure handling solution of the system failure recorded in the operation and maintenance knowledge document. For example, based on the failure name or the failure identifier of the system failure, the failure analysis result and the failure handling solution corresponding to the same system failure are extracted from the operation and maintenance knowledge document.

[0087] Optionally, the target processing system can also determine the failure information according to the system log corresponding to the system failure. For example, keywords are extracted from the log content corresponding to the system failure in the system log, so as to determine the keywords as the failure information of the system failure. Optionally, the extracted keywords are used to describe the failure characteristics of the system failure. The keywords can include at least one of the following: error code, error message, system component name, timestamp, resource status, operation or command, hostname or IP address, process ID, port number, etc.

[0088] In an alternative embodiment, the fault information further includes at least one of the following: the fault type of the system fault, version information. Optionally, the version information includes at least one of the following: the system version number, the component version number, where the system version number refers to the version number of the system where the fault occurs, and the component version number refers to the version number of the system component where the fault occurs.

[0089] In an alternative embodiment, the fault handling solution at least includes a fault repair method.

[0090] In an alternative embodiment, in addition to the fault repair method, the aforementioned fault handling solution may further include at least one of a repair verification method and a rollback method.

[0091] Optionally, after determining the fault analysis result, fault handling solution, and fault information of a certain system fault, the target processing system may generate a text fragment corresponding to the system fault based on the fault analysis result, fault handling solution, and fault information of the system fault. For example, obtain a text fragment template, set the keywords in the fault information at the target positions in the text fragment template, where the target positions refer to the title and / or the first N lines in the text fragment template, N is a positive integer, and import the fault analysis result and fault handling solution into the first position in the text fragment template, so as to obtain the text fragment corresponding to the system fault. The first position refers to the position in the text fragment template other than the target positions. The first position and the target positions can be set according to actual needs, and an import rule may be preset in the target processing system. The import rule can be used to describe the method of setting the aforementioned keywords at the target positions and the method of importing the fault analysis result and fault handling solution into the first position. Among them, in the process of generating the text fragment, corresponding measures may also be taken to prevent security problems such as prompt injection. Prompt injection is an attack method against the model.

[0092] Optionally, after obtaining multiple text fragments, a vector knowledge base may be constructed based on the multiple text fragments for subsequent knowledge retrieval. For example, perform vector conversion processing on the text fragments through a text embedding model to obtain the text vectors corresponding to the text fragments, and construct a vector knowledge base based on the text vectors corresponding to the text fragments.

[0093] It should be noted that by generating different text fragments according to the fault information, fault analysis results, and fault handling solutions of different system faults, the rich fault instances and processing strategies of the target platform are effectively divided into different text fragments. Thus, on the one hand, the compatibility problems that may be brought by using general industry knowledge are avoided, and on the other hand, the relevance and richness of the content in the text fragments are improved, thereby effectively improving the knowledge retrieval efficiency and the quality of retrieval results.

[0094] To achieve accurate knowledge retrieval, in the fault processing method provided in the first embodiment of this application, determining at least one reference text fragment from multiple text fragments based on target fault information includes: determining a query text vector according to the target fault information; obtaining text vectors corresponding to the text fragments in the multiple text fragments; and determining at least one reference text fragment from the multiple text fragments according to the similarity between the query text vector and the text vectors.

[0095] In an optional embodiment, the target processing system can directly use a text embedding model to perform vector conversion on the target fault information to obtain a query text vector.

[0096] In an optional embodiment, after obtaining the target fault information, the target processing system can also generate a query statement based on the target fault information and a query statement template. For example, the query statement can be "Find the corresponding knowledge fragments according to the following content: keyword A, keyword B, keyword C", so as to use the text embedding model to perform vector conversion on the query statement to obtain a query text vector.

[0097] Optionally, after obtaining the query text vector, the target processing system can obtain the text vectors corresponding to the text fragments in the multiple text fragments and calculate the similarity between the query text vector and the text vectors, so as to determine at least one reference text fragment from the multiple text fragments according to the calculated similarity. For example, calculate the similarity between the query text vector and the text vectors according to the cosine similarity, and determine the text fragment corresponding to the text vector with a similarity greater than a preset threshold as the reference text fragment. Another example is to sort the text vectors of the multiple text fragments in descending order of similarity to obtain the sorted text vectors, and determine the first M text vectors in the sorted text vectors as the reference text fragments, where M is a positive integer greater than 1.

[0098] Optionally, the similarity between the text vector of a certain text fragment and the query text vector can be regarded as the similarity between the text fragment and the target fault information.

[0099] It should be noted that by quantifying the semantic similarity between texts during the retrieval process, the retrieval accuracy can be improved, and the historical cases and processing solutions highly relevant to the current system fault can be quickly located, thereby improving the reliability of fault processing.

[0100] In order to obtain accurate target information, in the fault handling method provided in the first embodiment of this application, at least one reference text fragment and target fault information are processed through a fault handling model to obtain target information, including: generating a first prompt statement according to at least one reference text fragment, target fault information, and a preset first prompt template, where the first prompt statement is at least used to guide the fault handling model to generate a target fault analysis result and a target handling solution according to at least one reference text fragment and target fault information; inputting the first prompt statement into the fault handling model, and obtaining a target fault analysis result and a target handling solution through the fault handling model; and determining the target fault analysis result and the target handling solution as target information.

[0101] Optionally, the first prompt template is used to provide clear guidance for the fault handling model, so that the generated fault analysis result and handling solution are both professional and closely fit the current fault scenario. The structure and content of the template can be adjusted according to the specific requirements of fault handling, so no specific limitation is made here, and only an exemplary description is given in this embodiment. For example, an optional first prompt template can be as follows:

[0102] "Please analyze the cause of the current system fault and provide detailed handling steps according to the following fault description and historical reference handling records:

[0103] Fault description: [target fault information]

[0104] Historical reference handling records: [at least one reference text fragment];

[0105] Please output the fault analysis result and handling solution in the following format:

[0106] Fault analysis: [fault analysis result]

[0107] Handling solution: [target handling solution]".

[0108] Optionally, the target handling system can fill in the target fault information at the "fault description" above, and fill in at least one reference text fragment at the "historical reference record handling" above, so as to obtain the first prompt statement.

[0109] After obtaining the first prompt statement, the target handling system can input the first prompt statement into the fault handling model, and obtain a target fault analysis result and a target handling solution through the fault handling model, where the target fault analysis result refers to the fault analysis result of the system fault of the target platform. Thus, the target fault analysis result and the target handling solution are determined as target information. Among them, the target fault analysis result can be presented to the user to facilitate the user to trace and supervise the fault handling process, and the target fault analysis result can be stored as a new knowledge fragment in the operation and maintenance knowledge document for updating the vector knowledge base.

[0110] It should be noted that by generating the first prompt statement and obtaining the target information based on the first prompt statement, it is possible to effectively guide the fault handling model to generate more accurate fault analysis results and handling solutions. By guiding the model to generate fault analysis results in addition to the target handling solution, it is convenient for users to understand the cause of the fault in a timely manner, which serves as a direct basis for subsequent task execution and manual supervision, and is also convenient for updating the operation and maintenance knowledge document as a new knowledge fragment, thereby improving the reliability of fault handling.

[0111] In order to further improve the reliability of fault handling, in the fault handling method provided in the first embodiment of the present application, the fault handling of the target platform based on the target handling solution in the target information includes: generating a second prompt statement according to the target information and a preset second prompt template, where the second prompt statement is used to guide the fault handling model to determine the source of the target information according to multiple text fragments; inputting the second prompt statement into the fault handling model, and obtaining the source information corresponding to the target information through the fault handling model; when the source information indicates that the target information belongs to multiple text fragments, performing fault handling on the target platform based on the target handling solution.

[0112] There will be hallucination problems in the process of large language model processing, that is, the large language model may contain some general hardware faults and solutions. When these unvalidated solutions are executed in a specific platform (such as the aforementioned target platform), it may cause further faults due to compatibility problems. Therefore, in order to address this difficulty, it is necessary to avoid using general knowledge for fault handling, that is, it is necessary to use the knowledge in multiple text fragments for fault handling.

[0113] Therefore, the target processing system can reorganize the target information returned by the large language model into a prompt statement, perform a secondary query in the large language model, and require the large language model to answer the source of the content in the query statement (specifically in which text fragment). When the query content is not in the text fragment, require the large language model to answer "don't know" to obtain the source information corresponding to the target information.

[0114] Optionally, the structure and content of the second prompt template can be adjusted according to the specific requirements of fault handling, so no specific limitation is made here, and only an exemplary illustration is given in this embodiment. For example, an optional second prompt template can be as follows:

[0115] "Please confirm whether the following provided fault analysis results and handling solutions are from the knowledge base document and clearly indicate their specific sources:

[0116] Fault analysis result: [Target fault analysis result]

[0117] Processing solution: [Target processing solution]

[0118] Please answer in the following format:

[0119] The analysis result comes from the shard [Shard title], and the specific shard is located at [Shard location]. The processing solution comes from the shard [Shard title], and the specific shard is located at [Shard location]. If the analysis result or the processing solution is not in the text shards, please answer 'Don't know'.

[0120] Optionally, the target processing system can fill the target fault analysis result in the target information into the above "Fault analysis result", and fill the target processing solution in the target information into the above "Processing solution" to obtain the second prompt statement.

[0121] After obtaining the second prompt statement, the target processing system can input the second prompt statement into the fault processing model, and obtain the source information corresponding to the target information through the fault processing model, so as to determine whether the target information belongs to multiple text shards according to the source information. For example, if the source information contains the target field (i.e., the above "Don't know"), it is determined that the target information does not belong to the above multiple text shards. If the source information does not contain the target field, it is determined that the target information belongs to the above multiple text shards.

[0122] Optionally, when the source information indicates that the target information belongs to multiple text shards, the target platform is fault processed based on the target processing solution, and the fault processing result can be recorded on the work order for subsequent traceability.

[0123] Optionally, when the source information indicates that the target information does not belong to multiple text shards, the target platform is not fault processed based on the target processing solution, and a prompt message can be generated and sent to the user. The prompt message can be used to indicate that the current system fault of the target platform cannot be automatically resolved and awaits manual processing.

[0124] It should be noted that through the above method, it is ensured that the fault processing solution and the fault analysis result on which the subsequent fault processing of the target platform is based are verified professional knowledge, rather than general industry knowledge, so as to avoid compatibility problems caused by using general industry knowledge and effectively improve the reliability of fault processing.

[0125] In order to achieve automatic execution of fault handling, in the fault handling method provided in Embodiment 1 of this application, the text fragment includes the function identifier of the executable function corresponding to the fault handling solution. In the case where the source information indicates that the target information belongs to multiple text fragments, performing fault handling on the target platform based on the target handling solution includes: determining the executable function corresponding to the target handling solution based on the function identifier corresponding to the target handling solution in the target information and the attribute information in the target database, where the target database is used to store the attribute information of the executable function; performing fault handling on the target platform based on the executable function corresponding to the target handling solution.

[0126] In an optional embodiment, an atomic operation library (i.e., the aforementioned target database) can be pre-configured. For example, the user can disassemble the execution processes of known fault repair methods, repair verification methods, and rollback methods into a series of independent and smaller operation and maintenance operation units, i.e., atomic operations. These operations are basic steps that cannot be further divided and can independently complete a certain function or task. That is, a fault repair method (or repair verification method, or rollback method) corresponds to at least one atomic operation. These atomic operations are represented by attributes such as Action Version, Action Code, Action Desc, Action Level, Action Target, Action Params, Action TimeOut, ActionMulti, Action Depend, Action Manual through functional configuration design, saved as data in a target format (e.g., JSON (JavaScript Object Notation) format), and then stored in the atomic operation library. Among them, the aforementioned attributes (i.e., Action Version, Action Code, Action Desc, Action Level, Action Target, Action Params, Action TimeOut, Action Multi, Action Depend, Action Manual) constitute the attribute information of the executable function. Optionally, the aforementioned Action Version represents the version information of the operation, used to track the update and change history of the operation; Action Code represents the unique code identifier of the operation (i.e., the aforementioned function identifier), used to match the identifier in the fault handling solution. Action Desc represents the detailed description of the operation, explaining the function and purpose of the operation, Action Level represents the urgency or permission level of the operation. For example, some operations may require manual approval to execute, while others can be executed automatically. Action Target represents the target resource or service of the operation, determining the execution scope of the operation, Action Params represents the parameters of the operation, including the configuration or input values required for the operation execution, such as IP addresses, network segments, etc.ActionTimeOut represents the execution timeout time of an operation, which is used to prevent the operation from taking too long to complete and affecting the system performance or stability. ActionMulti is used to mark whether the operation needs to be executed serially or in parallel with other operations to coordinate the multi-step repair process in the system. ActionDepend is used to represent the prerequisite dependencies of the operation, that is, the conditions that need to be met or other operations that the operation depends on before execution. ActionManual is used to specify whether the operation is limited to manual execution, such as physical operations like plugging and unplugging cables, which cannot be automated.

[0127] Optionally, the text fragment includes the function identifier of the executable function corresponding to the fault handling solution. The number of function identifiers in the text fragment may be one or more. In the case where the source information indicates that the target information belongs to multiple text fragments, the target information must also include the function identifier of the executable function. Therefore, the target processing system can match the attribute information of the executable function corresponding to the target processing solution from the target database based on the function identifier of the target processing solution in the target information, and then determine the executable function corresponding to the target processing solution according to the attribute information of the executable function corresponding to the target processing solution. For example, parsing the data in the target format in the atomic operation library into the target data structure, where the target data structure refers to the data structure that can be understood (or processed) by the program. For example, parsing a JSON string into a JavaScript object. After obtaining the target data structure, the atomic operations in the target data structure can be mapped to specific initial executable functions. For example, setting up a mapping table or function registry to associate the function identifier (ActionCode) with the actual initial executable function. Then, according to the mapping table, determine the initial executable function to which the function identifier corresponding to the atomic operation belongs, and pass the information defined in the attribute information (such as ActionParams) to the initial executable function to obtain the executable function. Since the number of function identifiers in the text fragment may be one or more, the number of finally obtained executable functions may also be one or more, that is, at least one executable function can be finally obtained.

[0128] In an optional embodiment, during the process of matching the attribute information of the executable function corresponding to the target processing solution from the target database, the target processing system can directly read the target database and match the attribute information corresponding to the function identifier of the target processing solution from the target database.

[0129] In an optional embodiment, during the process of matching the attribute information of the executable function corresponding to the target processing solution from the target database, the target processing system can also output the attribute information through the fault handling model. For example, in the case where it is determined that the provenance information indicates that the target information belongs to multiple text fragments, the user inputs a third prompt statement to the fault handling model. The third prompt statement is used to guide the large model to determine the matching attribute information according to the function identifier corresponding to the target processing solution. For example, the third prompt statement can be "Please output the attribute information of the executable function corresponding to the target processing solution according to the function identifier corresponding to the target information in the target information". After receiving the third prompt statement, the fault handling model can read the content in the target database through the target plugin, thereby obtaining the attribute information corresponding to the function identifier and outputting it to the user. Among them, the target plugin is used for interaction between the fault handling model and the target database, enabling the target processing model to read the attribute information stored in the target database. Optionally, read-only permissions can be set for the fault handling model to improve the data security in the target database and prevent injection attacks. Optionally, in order to avoid the query operation occupying the database resources for too long, a query timeout can also be set. If the query operation does not complete within the specified time, the system will automatically terminate the query so that the target database can still maintain a good response speed in a high-concurrency scenario.

[0130] Optionally, after obtaining the executable function, the target processing system can execute the executable function to perform fault handling on the target platform.

[0131] It should be noted that through the above method, the automatic determination of the executable function for fault handling is realized, avoiding the problems of low efficiency and low reliability existing when manually writing the executable function, thereby further improving the reliability of fault handling.

[0132] In order to effectively perform fault repair, in the fault handling method provided in Embodiment 1 of the present application, the target processing solution at least includes a fault repair method. Performing fault handling on the target platform based on the target processing solution in the target information includes: performing fault repair processing on the target platform based on the first executable function in the executable function corresponding to the target processing solution, where the first executable function is used to execute the fault repair method.

[0133] Optionally, the target processing solution at least includes a fault repair method. Among them, the fault repair method refers to a method for repairing system faults. For example, the fault repair method can be restarting the service, applying patches, adjusting configuration parameters, executing specific script commands, etc.

[0134] Optionally, the first executable function refers to an executable function whose function identifier is the same as the function identifier corresponding to the fault repair method in the target processing solution, and the first executable function is used to execute the fault repair method in the target processing solution.

[0135] The target processing system can determine the first executable function from at least one executable function corresponding to the target processing solution to perform fault repair processing on the target platform. Among them, when there are multiple first executable functions, the target processing system can execute the first executable functions according to the sorting of the function identifiers in the target processing solution.

[0136] It should be noted that through the above method, the automatic repair of the system fault of the target platform is realized, thereby improving the efficiency and reliability of fault repair.

[0137] In order to effectively perform repair verification and system rollback, in the fault processing method provided in the first embodiment of this application, after performing fault repair processing on the target platform based on the first executable function in the executable functions corresponding to the target processing solution, the method further includes: if the target processing solution further includes a repair verification method, performing repair verification processing on the target platform based on the second executable function in the executable functions to obtain a repair verification result, where the second executable function is used to execute the repair verification method; in the case where the repair verification result indicates that the repair fails, if the target processing solution further includes a rollback method, performing rollback processing on the target platform based on the third executable function in the executable functions, where the third executable function is used to execute the rollback method.

[0138] In an optional embodiment, the foregoing fault processing solution may further include at least one of a repair verification method and a rollback method in addition to the fault repair method. Among them, the repair verification method refers to a method used to verify whether the fault has been resolved after executing the fault repair method. For example, the repair verification method may be to recheck the system log, run a diagnostic tool, execute a health check script, etc. The rollback method is used to roll back the system of the target platform to its state before fault repair in the case of failed fault repair (such as, the fault is not resolved or side effects are generated).

[0139] In an optional embodiment, there are multiple executable functions.

[0140] Optionally, the second executable function refers to an executable function whose function identifier is the same as the function identifier corresponding to the repair verification method in the target processing solution, and the second executable function is used to execute the repair verification method in the target processing solution.

[0141] Optionally, if the target processing solution does not include a repair verification method, it is determined that the current fault processing is completed, and the subsequent verification process can be processed manually.

[0142] Optionally, if the target processing solution further includes a repair verification method, after performing the fault repair process, the target processing system may execute the second executable function among multiple executable functions to perform a repair verification process on the target platform and obtain a repair verification result. Among them, when there are multiple second executable functions, the target processing system may execute the second executable functions according to the sorting of the function identifiers in the target processing solution. The repair verification result is used to indicate whether the repair is successful.

[0143] Optionally, when the repair verification result indicates that the repair is successful, it is determined that the current fault handling is completed.

[0144] Optionally, when the repair verification result indicates that the repair fails, if the target processing solution does not include a rollback method, it is determined that the current fault handling is completed, and the subsequent verification process can be processed manually.

[0145] Optionally, when the repair verification result indicates that the repair fails, if the target processing solution further includes a rollback method, the target platform can be rolled back based on the third executable function among multiple executable functions. The third executable function refers to the executable function whose function identifier is the same as the function identifier corresponding to the rollback method in the target processing solution, and the third executable function is used to execute the rollback method in the target processing solution.

[0146] Optionally, to ensure the traceability of the fault handling process, from the entry of the target fault information, the generation process of the fault handling model, to the fault handling based on the target processing solution and the processing results, they can all be recorded in the work order to meet the requirements of manual supervision, query, termination, rollback, and traceability.

[0147] It should be noted that through the above method, after repairing the system fault of the target platform, the repair verification and the rollback operation in case of repair failure are automatically performed, thereby improving the security and integrity of the fault handling process, enabling the system to more intelligently cope with the uncertainties that may be brought about by the fault repair attempt, quickly taking remedial measures, avoiding secondary faults caused by improper fault repair, and at the same time reducing the workload of the operation and maintenance personnel, thereby improving the efficiency and reliability of fault repair.

[0148] To improve the flexibility of fault handling, in the fault handling method provided in the first embodiment of the present application, the fault handling of the target platform based on the target processing solution in the target information includes: obtaining the fault level of the system fault of the target platform; determining the execution time point of the target processing solution based on the fault level; and performing fault handling on the target platform based on the target processing solution at the execution time point.

[0149] In the operation and maintenance management of the platform, different types of faults have different priorities and handling urgencies due to their different impact scopes and severities. For example, some hardware faults may pose a serious threat to the stability of the current service and require immediate measures to be taken for repair. Such faults are defined as highly urgent faults and must be quickly resolved to ensure service continuity and user experience. On the other hand, although some hardware faults exist, they do not immediately have a devastating impact on the service. These faults can be effectively managed through the functions of the platform itself. For example, by implementing an elastic scaling strategy to temporarily increase resources to relieve pressure, or through an automatic isolation mechanism to remove the faulty device from the main service path to achieve dynamic balance and stability of the service. Such faults are classified as relatively non-urgent faults because the intelligent scheduling and resource management capabilities of the platform itself can alleviate the problem to a certain extent and enable the service to continue running smoothly. Therefore, in the face of different fault scenarios, different execution time points for fault handling can be determined. For example, for an urgent hardware fault that affects service stability, the system will quickly recommend or directly execute an immediately executable solution, while for a minor fault that can temporarily maintain service stability through elastic scaling or other means, the system will prompt for delayed handling. In addition, considering that some self-healing measures may interfere with the service, such as the need to restart the server or modify the network configuration, the target processing system can also intelligently evaluate its intrusion degree into the service and propose alternative suggestions.

[0150] Therefore, the target processing system can first obtain the fault level of the system fault of the target platform, and then determine the execution time point of the target processing plan based on the fault level.

[0151] Optionally, in the process of determining the fault level, the target processing system can adopt a preset level determination rule to determine the fault level of the system fault according to the log content of the system fault in the system log. For example, when the log content of the system fault hits the level determination rule of a certain fault level, it is determined that the hit fault level is the fault level of the system fault.

[0152] Optionally, in the process of determining the fault level, the target processing system can also input the log content of the system fault into a fault level classification model, perform fault identification through the fault level classification model, and obtain the fault level of the system fault. Among them, the fault level classification model can be a neural network model.

[0153] In the target processing system, there can be preset fault handling delay times corresponding to different fault levels (such as, immediate handling (i.e., no delay), delay for half an hour, delay until early morning, etc.). The target processing system can determine the execution time point according to the current time point and the fault handling delay time. For example, add the current time point and the fault handling delay time to obtain the execution time point.

[0154] After determining the execution time point, the target processing system may perform fault handling on the target platform based on the target processing solution at the execution time point.

[0155] It should be noted that by intelligently judging the fault level and reasonably planning the execution time point, high-priority faults can be immediately responded to, while low-priority faults are processed at a time when the impact on service availability is relatively small, thereby improving the flexibility of fault handling and reducing the potential risk to the service caused by improper fault handling, and enhancing the reliability of fault handling.

[0156] In an optional embodiment, the Figure 3 schematic diagram shown can be used to achieve automatic fault handling. In this embodiment, the target processing system may include an AI (Artificial Intelligence) agent to implement automated log analysis and fault handling through the AI agent. Figure 3 is the flowchart of the fault handling method provided in Embodiment 1 of the present application Figure 2 , as Figure 3 shown, the AI agent may include a corpus management module, a work order management module, an exception information processor, a problem analysis supervisor, and a self-healing solution processor. As Figure 3 shown, the operation and maintenance knowledge document of the target platform contains fault error messages and solutions for fault self-healing, which can be saved as corpora in the corpus. The corpus management module may slice the operation and maintenance knowledge document so that each single text slice has version information corresponding to the same system fault, problem error log keywords, fault repair methods, repair verification methods, and rollback methods, and then perform vector conversion processing on it to store the obtained text vector slices in the vector knowledge base for subsequent RAG processes. Optionally, the corpus management module may also refine the fault repair methods, repair verification methods, and rollback methods into atomic operations and store them in the atomic operation library (i.e., the above-mentioned target database) in a target format.

[0157] Optionally, as Figure 3As shown, when a system failure occurs in the target platform, the exception information processor can collect and analyze the system logs, extract the keywords in the system logs as the target fault information, and generate a query statement. Then, based on the query text vector of the generated query statement, knowledge retrieval is performed in the vector knowledge base to obtain at least one reference text fragment. After obtaining at least one reference text fragment, the fault handling model can process at least one reference text fragment and the target fault information to obtain the target information. Then, the problem analysis supervisor reorganizes the target information returned by the fault handling model into a query statement and performs a secondary query in the fault handling model, and requires the large model to answer the source of the content in the query statement (specifically in which text fragment). When the query content is not in the document slice, the large model is required to answer "don't know". If it is determined that the target information belongs to multiple text fragments, the fault handling model can determine the attribute information of the executable function corresponding to the target processing scheme from the atomic operation library. Furthermore, the self-healing scheme processor can determine the executable function according to the attribute information, and then perform specific self-healing operations on the servers or networks in the target platform according to the executable function. The overall self-healing process includes steps such as judging whether to execute the self-healing operation (i.e., determining the execution time point), executing the self-healing operation (i.e., performing fault repair processing), verifying the self-healing result, and rolling back after the repair fails. Each step will be pushed onto the stack and into the pipeline in sequence, so that the self-healing scheme processor can execute them in an orderly manner.

[0158] Optionally, to meet the traceability of the fault self-healing process, from the entry of the fault information, the analysis of the fault handling model, to the self-healing process and results, they will all be recorded in the work order. To meet the requirements of manual supervision, queryability, terminability, rollbackability, traceability, etc. As Figure 3 shown, when it is detected that a system failure occurs in the target platform, the work order management module creates a work order, then records the cleaning result of the system logs by the exception information processor in the work order. After that, the prompt statement generated according to at least one reference text fragment and the target fault information is recorded in the work order, and the work order management module can record the output results (i.e., the target information and the attribute information of the executable function) fed back by the fault handling model in the work order. During the fault self-healing process, if the fault repair is successful, the fault repair result is recorded in the work order and the work order is determined to be completed. If the fault repair fails, system rollback is performed, or as Figure 3As shown, manual intervention is selected. The manual intervention methods include, but are not limited to: adjusting the inquiry content and triggering the inquiry again; updating the atomic operation library and triggering the inquiry again; manual operation and maintenance work orders, etc. During the manual intervention process, the user can perform fault confirmation, manual repair, and manual verification, and the information generated during the manual intervention will also be recorded in the work order. In addition, when the operation level of the executable function is relatively high and manual blocking and waiting for approval to execute are required, etc., the self-healing solution processor can generate a reminder message to remind the user to approve the executable function before executing the executable function, so that when the user approves, the executable function starts to be executed to achieve fault repair.

[0159] In the embodiment of the present application, by combining the retrieval enhancement technology, when obtaining the fault information of the system of the target platform, relevant knowledge is first extracted from multiple text fragments through information retrieval, and then the obtained relevant knowledge is injected into the process of generating the fault processing model. This "retrieval-generation" collaborative mode not only retains the logical reasoning and natural language generation capabilities of the language model, but also breaks through the limitation of its static knowledge boundary, improves the accuracy of the generated target information, and thus can improve the reliability of fault handling. Further, by setting different text fragments to contain fault information and fault handling solutions for different system faults, the relevant information of the same system fault is summarized in the same text fragment, improving the relevance of the content in the same text fragment, and thus can improve the retrieval accuracy during the retrieval process, achieving the purpose of generating a fault handling solution based on the text fragments of the distinguished fault knowledge through the retrieval enhancement technology for fault handling, realizing the technical effect of improving the reliability of fault handling, and further solving the technical problem of low reliability of fault handling in the related art when a platform fails.

[0160] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0162] Embodiment 2

[0163] According to an embodiment of the present application, a method for processing a fault is further provided, as Figure 4 shown, the method includes:

[0164] Step S401, obtain the fault information of the system fault of the target platform uploaded by the client to obtain the target fault information.

[0165] Step S402, based on the target fault information in the cloud server, determine at least one reference text fragment from multiple text fragments, where different text fragments contain the fault information and fault handling solutions of different system faults; process at least one reference text fragment and the target fault information through a fault handling model to obtain target information, where the target information at least includes the target handling solution for the system fault of the target platform.

[0166] Step S403, feedback the target information to the client, where the target information is used to instruct the client to perform fault handling on the target platform based on the target handling solution.

[0167] Through the above solution, by combining retrieval enhancement technology, when obtaining the fault information of the system of the target platform, relevant knowledge is first extracted from multiple text fragments through information retrieval, and then the obtained relevant knowledge is injected into the process of generating the fault handling model. This "retrieval-generation" collaborative mode not only retains the logical reasoning and natural language generation capabilities of the language model but also breaks through the limitation of its static knowledge boundary, improving the accuracy of the generated target information, thereby improving the reliability of fault handling. Further, by setting the fault information and fault handling solutions of different system faults in different text fragments, the relevant information of the same system fault is summarized in the same text fragment, improving the relevance of the content in the same text fragment, thereby improving the retrieval accuracy in the retrieval process, achieving the purpose of combining retrieval enhancement technology to generate a fault handling solution based on the text fragments of the distinguished fault knowledge for fault handling, realizing the technical effect of improving the reliability of fault handling, and further solving the technical problem of low reliability of fault handling in the related art when a platform fails.

[0168] In the cloud server, the specific method for handling faults is the same as the method in Embodiment 1 and will not be elaborated here.

[0169] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0171] Embodiment 3

[0172] According to an embodiment of the present application, there is also provided a fault handling device for implementing the above-mentioned fault handling method, as Figure 5As shown in the figure, the device includes: a first acquisition unit 501, a determination unit 502, a first processing unit 503, and a second processing unit 504.

[0173] The first acquisition unit 501 is configured to acquire fault information of a system fault of a target platform to obtain target fault information;

[0174] The determination unit 502 is configured to determine at least one reference text fragment from multiple text fragments based on the target fault information, where different text fragments contain fault information and fault handling solutions of different system faults;

[0175] The first processing unit 503 is configured to process at least one reference text fragment and the target fault information through a fault handling model to obtain target information, where the target information at least includes a target handling solution for the system fault of the target platform;

[0176] The second processing unit 504 is configured to perform fault handling on the target platform based on the target handling solution in the target information.

[0177] In the fault handling device provided in Embodiment 3 of the present application, the first acquisition unit 501 acquires fault information of a system fault of a target platform to obtain target fault information; the determination unit 502 determines at least one reference text fragment from multiple text fragments based on the target fault information, where different text fragments contain fault information and fault handling solutions of different system faults; the first processing unit 503 processes at least one reference text fragment and the target fault information through a fault handling model to obtain target information, where the target information at least includes a target handling solution for the system fault of the target platform; the second processing unit 504 performs fault handling on the target platform based on the target handling solution in the target information. In this solution, by combining the retrieval enhancement technology, when the fault information of the system of the target platform is acquired, relevant knowledge is first extracted from multiple text fragments through information retrieval, and then the obtained relevant knowledge is injected into the generation process of the fault handling model. This "retrieval-generation" collaborative mode not only retains the logical reasoning and natural language generation capabilities of the language model but also breaks through the limitation of its static knowledge boundary, improving the accuracy of the generated target information, thereby improving the reliability of fault handling. Further, by setting different text fragments to contain fault information and fault handling solutions of different system faults, the relevant information of the same system fault is summarized in the same text fragment, improving the relevance of the content in the same text fragment, thereby improving the retrieval accuracy in the retrieval process, achieving the purpose of generating a fault handling solution based on the text fragments of the distinguished fault knowledge by combining the retrieval enhancement technology to perform fault handling, realizing the technical effect of improving the reliability of fault handling, and further solving the technical problem of low reliability of fault handling in the related art when a platform fails.

[0178] Optionally, in the fault processing device provided in Embodiment 3 of the present application, the fault processing device further includes: a second acquisition unit, configured to acquire an operation and maintenance knowledge document of a target platform; a first extraction unit, configured to extract, from the operation and maintenance knowledge document, a fault analysis result and a fault handling solution of the system fault recorded in the operation and maintenance knowledge document; a second extraction unit, configured to determine fault information according to a system log corresponding to the system fault; a generation unit, configured to generate different text slices according to the fault analysis results, fault handling solutions, and fault information of different system faults to obtain a plurality of text slices.

[0179] Optionally, in the fault processing device provided in Embodiment 3 of the present application, the determination unit 502 includes: a first determination subunit, configured to determine a query text vector according to target fault information; a first acquisition subunit, configured to acquire text vectors corresponding to text slices in a plurality of text slices; a second determination subunit, configured to determine at least one reference text slice from the plurality of text slices according to the similarity between the query text vector and the text vectors.

[0180] Optionally, in the fault processing device provided in Embodiment 3 of the present application, the first processing unit 503 includes: a first generation subunit, configured to generate a first prompt statement according to at least one reference text slice, target fault information, and a preset first prompt template, where the first prompt statement is at least used to guide a fault processing model to generate a target fault analysis result and a target handling solution according to at least one reference text slice and target fault information; a first processing subunit, configured to input the first prompt statement into the fault processing model to obtain a target fault analysis result and a target handling solution through the fault processing model; a third determination subunit, configured to determine the target fault analysis result and the target handling solution as target information.

[0181] Optionally, in the fault processing device provided in Embodiment 3 of the present application, the second processing unit 504 includes: a second generation subunit, configured to generate a second prompt statement according to the target information and a preset second prompt template, where the second prompt statement is used to guide the fault processing model to determine the source of the target information according to a plurality of text slices; a second processing subunit, configured to input the second prompt statement into the fault processing model to obtain source information corresponding to the target information through the fault processing model; a third processing subunit, configured to, when the source information indicates that the target information belongs to a plurality of text slices, perform fault processing on the target platform based on the target handling solution.

[0182] Optionally, in the fault processing device provided in the third embodiment of the present application, the text fragment includes a function identifier of an executable function corresponding to a fault processing solution. The third processing subunit includes: a determination module, configured to determine an executable function corresponding to the target processing solution based on the function identifier corresponding to the target processing solution in the target information and the attribute information in the target database, where the target database is used to store the attribute information of the executable function; a processing module, configured to perform fault processing on the target platform based on the executable function corresponding to the target processing solution.

[0183] Optionally, in the fault processing device provided in the third embodiment of the present application, the target processing solution at least includes a fault repair method. The second processing unit 504 includes: a fourth processing subunit, configured to perform fault repair processing on the target platform based on a first executable function in the executable function corresponding to the target processing solution, where the first executable function is used to execute the fault repair method.

[0184] Optionally, in the fault processing device provided in the third embodiment of the present application, the fault processing device further includes: a third processing unit, configured to, if the target processing solution further includes a repair verification method, perform repair verification processing on the target platform based on a second executable function in the executable function to obtain a repair verification result, where the second executable function is used to execute the repair verification method; a fourth processing unit, configured to, in the case that the repair verification result indicates repair failure, if the target processing solution further includes a rollback method, perform rollback processing on the target platform based on a third executable function in the executable function, where the third executable function is used to execute the rollback method.

[0185] Optionally, in the fault processing device provided in the third embodiment of the present application, the second processing unit 504 includes: a second acquisition subunit, configured to acquire the fault level of the system fault of the target platform; a fourth determination subunit, configured to determine the execution time point of the target processing solution based on the fault level; a fifth processing subunit, configured to perform fault processing on the target platform based on the target processing solution at the execution time point.

[0186] It should be noted here that the above-mentioned first acquisition unit 501, determination unit 502, first processing unit 503, and second processing unit 504 correspond to steps S201 to S204 in Embodiment 1. The above-mentioned units and the corresponding steps have the same implemented examples and application scenarios, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned modules, as a part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0187] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0188] Example 4

[0189] Embodiments of the present application may provide an electronic device, which may be any one of a group of electronic devices. Optionally, in this embodiment, the above-mentioned electronic device may also be replaced with a terminal device such as a mobile terminal.

[0190] Optionally, in this embodiment, the above-mentioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0191] In this embodiment, the above-mentioned electronic device may execute program codes corresponding to the steps in the fault processing method provided in any one of the above method embodiments.

[0192] Optionally, Figure 6 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 6 shown, the electronic device 60 may include: one or more ( Figure 6 only one is shown in the figure) processors 602, a memory 604. The electronic device 60 may further include a storage controller for controlling and managing the memory 604; the electronic device 60 may further include a peripheral interface for connecting a radio frequency module, an audio module, a display screen, etc.

[0193] Among them, the memory may be used to store software programs and modules, such as program instructions / modules corresponding to the fault processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned fault processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0194] The processor may call the information and application programs stored in the memory through a transmission device to execute the program codes corresponding to the steps in the fault processing method provided in any one of the above method embodiments.

[0195] Those of ordinary skill in the art can understand that Figure 6 the structure shown is only schematic, and the electronic device may also be a terminal device such as a smart phone, a tablet computer, a personal digital assistant, and a mobile Internet device (MID), a PAD, etc. Figure 6It does not limit the structure of the above electronic device. For example, the electronic device 60 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 6 and may have a configuration different from that shown in Figure 6 .

[0196] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0197] Embodiment 5

[0198] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the above storage medium may be used to store the program code executed by the fault processing method provided in the first embodiment above.

[0199] Optionally, in this embodiment, the above storage medium may be located in any one of the electronic devices in the electronic device group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0200] Embodiment 6

[0201] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and the computer program realizes the fault processing method provided in the first embodiment above when executed by a processor.

[0202] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0203] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0204] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0205] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0206] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0207] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical disks, etc., which can store program codes.

[0208] The above is only the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for handling a fault, characterized in that: include: Obtaining fault information of a system fault on a target platform, and obtaining target fault information; Determine at least one reference text segment from a plurality of text segments based on the target fault information, wherein different text segments contain fault information and fault handling solutions of different system faults; Processing the at least one reference text segment and the target fault information through a fault processing model to obtain target information, wherein the target information at least includes a target processing solution for the system fault of the target platform; Fault processing is performed on the target platform based on the target processing solution in the target information.

2. The method according to claim 1, characterized in that Before determining at least one reference text segment from a plurality of text segments based on the target fault information, the method further includes: Obtaining operation and maintenance knowledge documents of the target platform; For the system failure recorded in the operation and maintenance knowledge document, extract the failure analysis result and failure handling solution of the system failure from the operation and maintenance knowledge document; Determine the fault information according to the system log corresponding to the system fault; Different text segments are generated according to the fault analysis results, fault handling solutions and the fault information of different system faults to obtain the multiple text segments.

3. The method according to claim 1, characterized in that Determining at least one reference text segment from a plurality of text segments based on the target fault information includes: Determine a query text vector according to the target fault information; Obtaining text vectors corresponding to text segments among the multiple text segments; According to the similarity between the query text vector and the text vector, the at least one reference text segment is determined from the multiple text segments.

4. The method according to claim 2, characterized in that: The at least one reference text segment and the target fault information are processed by the fault processing model to obtain the target information, including: Generate a first prompt statement according to the at least one reference text segment, the target fault information and a preset first prompt template, wherein the first prompt statement is at least used to guide the fault processing model to generate a target fault analysis result and a target processing solution according to the at least one reference text segment and the target fault information; Inputting the first prompt statement into the fault processing model, and obtaining a target fault analysis result and the target processing solution through the fault processing model; The target fault analysis result and the target processing solution are determined as the target information.

5. The method according to any one of claims 1 to 4, characterized in that: Processing the fault of the target platform based on the target processing solution in the target information includes: generating a second prompt statement according to the target information and a preset second prompt template, wherein the second prompt statement is used to guide the fault processing model to determine the source of the target information according to the multiple text fragments; Inputting the second prompt statement into the fault processing model, and obtaining source information corresponding to the target information through the fault processing model; In a case where the source information indicates that the target information belongs to the multiple text fragments, fault processing is performed on the target platform based on the target processing solution.

6. The method according to claim 5, characterized in that The text segment includes a function identifier of an executable function corresponding to the fault handling solution, and when the source information indicates that the target information belongs to the multiple text segments, performing fault handling on the target platform based on the target handling solution includes: Determine an executable function corresponding to the target processing solution based on a function identifier corresponding to the target processing solution in the target information and attribute information in a target database, wherein the target database is used to store attribute information of executable functions; Fault processing is performed on the target platform based on an executable function corresponding to the target processing solution.

7. The method according to any one of claims 1 to 4, characterized in that The target processing scheme at least includes a fault repair method, and performing fault processing on the target platform based on the target processing scheme in the target information includes: Performing fault repair processing on the target platform based on a first executable function among the executable functions corresponding to the target processing solution, wherein the first executable function is used to execute the fault repair method.

8. The method according to claim 7, characterized in that After performing fault repair processing on the target platform based on the first executable function among the executable functions corresponding to the target processing solution, the method further includes: If the target processing scheme also includes a repair verification method, performing repair verification processing on the target platform based on a second executable function in the executable functions to obtain a repair verification result, wherein the second executable function is used to execute the repair verification method; In the case where the repair verification result indicates that the repair has failed, if the target processing solution also includes a rollback method, the target platform is rolled back based on a third executable function in the executable functions, wherein the third executable function is used to execute the rollback method.

9. The method according to any one of claims 1 to 4, characterized in that: Processing the fault of the target platform based on the target processing solution in the target information includes: Obtaining a fault level of a system fault of the target platform; Determining an execution time point of the target processing solution based on the fault level; Fault processing is performed on the target platform based on the target processing solution at the execution time point.

10. A method for handling a fault, characterized in that: include: Obtain the fault information of the system fault of the target platform uploaded by the client, and obtain the target fault information; In the cloud server, at least one reference text segment is determined from a plurality of text segments based on the target fault information, wherein different text segments contain fault information and fault handling solutions of different system faults; the at least one reference text segment and the target fault information are processed by a fault handling model to obtain target information, wherein the target information at least includes a target handling solution for the system fault of the target platform; The target information is fed back to the client, wherein the target information is used to instruct the client to perform fault processing on the target platform based on the target processing solution.

11. A fault processing device, characterized in that: include: A first acquisition unit is used to acquire fault information of a system fault of a target platform to obtain target fault information; A determination unit, configured to determine at least one reference text segment from a plurality of text segments based on the target fault information, wherein different text segments contain fault information and fault handling solutions of different system faults; A first processing unit, configured to process the at least one reference text segment and the target fault information through a fault processing model to obtain target information, wherein the target information at least includes a target processing solution for the system fault of the target platform; The second processing unit is used to perform fault processing on the target platform based on the target processing solution in the target information.

12. An electronic device, characterized in that: include: A memory storing an executable program; A processor is used to run the program, wherein the program executes the fault processing method of any one of claims 1 to 9 when running.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the storage medium is located is controlled to execute the fault processing method of any one of claims 1 to 9.

14. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the fault processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Fault diagnosis method and device, electronic equipment and storage medium

    CN117331730A

  • Fault repair method and device, storage medium and electronic equipment

    CN118445110A

  • Fault tree generation method and device, storage medium and electronic device

    CN119202188A

  • Fault root cause analysis method and device, electronic equipment and storage medium

    CN119597528A

  • System fault analysis and handling method, device, equipment and computer program product

    CN119938376A