A method and apparatus for alarm processing based on a large model

By constructing an agent operation and maintenance knowledge graph and the relationship between it and the large model, the target agent is activated to process alarm information, which solves the problem of low processing efficiency of the BERT model and achieves efficient alarm information processing.

CN119398158BActive Publication Date: 2025-10-21WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411479578.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-10-21
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The existing BERT model is inefficient in processing alarm information and is unable to cope with the processing requirements of high concurrency and large-scale data. It needs to be adjusted and trained for different data sets, resulting in low processing efficiency.

Method used

A large-scale model-based alarm processing system is introduced. By constructing an agent operation and maintenance knowledge graph, the relationship between alarm entities and agents is determined. The large-scale model is used to construct the association between entity description information and operation and maintenance objects, and the target agent is activated to process alarm information, thereby reducing computing power waste and improving processing efficiency.

Benefits of technology

It enables rapid location and processing of alarm information without wasting computing power, improving the efficiency and accuracy of alarm information processing and adapting to the processing needs of high concurrency and large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119398158B_ABST
    Figure CN119398158B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method and device for alarm processing based on a large model, applied to an alarm processing system having a plurality of agents, each agent being configured to process alarm information of a corresponding scene, the method comprising: obtaining at least one alarm information; inputting a first prompt word into a large model to obtain alarm entities corresponding to the at least one alarm information; the first prompt word comprising entity description information and the at least one alarm information; the entity description information being obtained by processing a second prompt word by the large model; the second prompt word comprising basic information of each operation and maintenance object and basic information of each agent; determining a target agent corresponding to the alarm entities based on an agent operation and maintenance knowledge graph; the agent operation and maintenance knowledge graph being obtained by processing the second prompt word by the large model; and the agent operation and maintenance knowledge graph representing an association relationship between operation and maintenance objects and agents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large models, and in particular to a method and device for alarm processing based on large models. Background Art

[0002] In today's information society, the rapid growth and complexity of data have brought new challenges to information processing, understanding and application.

[0003] Currently, the BERT model processes alert information based on its text understanding capabilities. However, the BERT model is not universally applicable and requires specific adjustments and training for different datasets before it can be used properly. Large datasets require a corresponding number of BERT models to be trained before they can be used properly. This results in inefficient alert processing and makes it difficult to handle high-concurrency and large-scale data processing.

[0004] In summary, how to improve the efficiency of processing alarm information is a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The embodiment of the present invention provides a method and device for alarm processing based on a large model, which is used to solve the problem of low efficiency in processing alarm information in existing memory.

[0006] In a first aspect, an embodiment of the present invention provides a method for alarm processing based on a large model, which is applied to an alarm processing system having multiple intelligent agents, each agent being used to process alarm information of a corresponding scenario, and the method includes: obtaining at least one alarm information; inputting a first prompt word into the large model to obtain an alarm entity corresponding to at least one alarm information; the first prompt word includes description information of each entity and at least one alarm information; the description information of each entity is obtained by processing the second prompt word by the large model; the second prompt word includes basic information of each operation and maintenance object and basic information of each agent; based on the agent operation and maintenance knowledge graph, determining the target agent corresponding to the alarm entity; the agent operation and maintenance knowledge graph is obtained by processing the second prompt word by the large model; the agent operation and maintenance knowledge graph represents the association relationship between the operation and maintenance object and the agent; activating the target agent to process the corresponding alarm information.

[0007] In the above technical solution, the concept of an alarm processing system including multiple agents is introduced. Each agent is responsible for processing alarm information in a specific scenario. In this way, the agent operation and maintenance knowledge graph is first constructed through a large model to extract the association relationship between the agent and the operation and maintenance object. Then, through the large model, the alarm information and the operation and maintenance object are associated to obtain the alarm entity. Then, combined with the agent operation and maintenance knowledge graph, the alarm is associated with the agent to determine the activated agent, thereby reducing a lot of computing power while finding the agent corresponding to the alarm information more quickly and conveniently, making it easier to process the alarm information according to the agent in the future.

[0008] Optionally, the second prompt word also includes a first description format of the entity description information; the large model processes the second prompt word, including: based on the basic information of each operation and maintenance object and the basic information of each agent in the second prompt word, obtaining the entity description information that conforms to the first description format through the large model; the first description format includes the entity name, entity type and entity description.

[0009] Optionally, the second prompt word also includes a second description format of the agent operation and maintenance knowledge graph and a first operation instruction for obtaining the agent operation and maintenance knowledge graph through each entity description information; after obtaining each entity description information that conforms to the first description format, it also includes: according to the first operation instruction, the large model obtains the agent operation and maintenance knowledge graph that conforms to the second description format through each entity description information; the second description format includes the agent as the source entity, the operation and maintenance object as the target entity, and a description of the operation and maintenance relationship between the source entity and the target entity.

[0010] Optionally, activating a target agent to process corresponding alarm information includes: building a chat room corresponding to the target agent based on a multicast mechanism; and any target agent collaboratively processes the alarm information through the chat room.

[0011] Optionally, based on the agent operation and maintenance knowledge graph, the target agent corresponding to the alarm entity is determined, including: for any alarm entity, if the alarm entity is an operation and maintenance object in the agent operation and maintenance knowledge graph, the activation coefficient of the alarm entity is set to a first value; if the alarm entity is an agent in the agent operation and maintenance knowledge graph, the activation coefficient of the alarm entity is set to a second value; the second value is higher than the first value; the activation value of the agent is determined according to the activation coefficients of each alarm entity pointing to the same agent; and the agent whose activation value meets the activation condition is determined as the target agent.

[0012] Optionally, the first prompt word also includes a third description format; the first prompt word is input into the large model to obtain an alarm entity corresponding to at least one alarm information, including: the first prompt word is input into the large model to obtain alarm description information that conforms to the third description format; the third description format includes the alarm entity as the alarm source, and the first correlation strength between the alarm entity and the alarm information; the activation value of the agent is determined according to the activation coefficient of each alarm entity pointing to the same agent, including: for any alarm entity, the target value of the alarm entity is determined according to the first correlation strength of the alarm entity and the activation coefficient of the alarm entity; the activation value of the agent is determined according to the target value of each alarm entity pointing to the same agent.

[0013] Optionally, it also includes: if the alarm entity is not an operation and maintenance object or agent in the agent operation and maintenance knowledge graph, determining the similarity between the alarm entity and the basic information of each agent; setting the similarity that meets the set threshold as the activation coefficient of the alarm entity; determining the target value of the alarm entity according to the first association strength of the alarm entity and the activation coefficient of the alarm entity, including: determining the target value of the alarm entity according to the first association strength of the alarm entity, the default association strength of the non-entity and the activation coefficient of the alarm entity.

[0014] Optionally, the second description format also includes a second association strength, which is used to characterize the strength of the operation and maintenance relationship between the source entity and the target entity; determining the target value of the alarm entity based on the first association strength of the alarm entity and the activation coefficient of the alarm entity, including: determining the target value of the alarm entity based on the first association strength of the alarm entity, the second association strength corresponding to the alarm entity in the agent operation and maintenance knowledge graph, and the activation coefficient of the alarm entity.

[0015] Optionally, the activation value of the agent is determined based on the target values ​​of each alarm entity pointing to the same agent, including: for the host agent, the convergence value of the host agent is determined based on the degree of aggregation of each alarm entity pointing to the host agent in the physical machine, rack and computer room; the activation value of the agent is determined based on the target value of each alarm entity pointing to the host agent and the convergence value of the host agent.

[0016] In a second aspect, an embodiment of the present invention provides an alarm processing device based on a large model, which is applied to an alarm processing system having multiple intelligent agents, each agent being used to process alarm information of a corresponding scenario, including: an acquisition unit for acquiring at least one alarm information; a processing unit for inputting a first prompt word into the large model to obtain an alarm entity corresponding to at least one alarm information; the first prompt word includes description information of each entity and at least one alarm information; the description information of each entity is obtained by processing the second prompt word by the large model; the second prompt word includes basic information of each operation and maintenance object and basic information of each agent; based on the agent operation and maintenance knowledge graph, the target agent corresponding to the alarm entity is determined; the agent operation and maintenance knowledge graph is obtained by processing the second prompt word by the large model; the agent operation and maintenance knowledge graph represents the association relationship between the operation and maintenance object and the agent; and the target agent is activated to process the corresponding alarm information.

[0017] Optionally, the second prompt word also includes a first description format of the entity description information; the processing unit is specifically used to: based on the basic information of each operation and maintenance object and the basic information of each agent in the second prompt word, obtain the entity description information that conforms to the first description format through the big model; the first description format includes the entity name, entity type and entity description.

[0018] Optionally, the second prompt word also includes a second description format of the agent operation and maintenance knowledge graph and a first operation instruction for obtaining the agent operation and maintenance knowledge graph through the description information of each entity; the processing unit is also used to: according to the first operation instruction, the large model obtains the agent operation and maintenance knowledge graph that conforms to the second description format through the description information of each entity; the second description format includes the agent as the source entity, the operation and maintenance object as the target entity, and a description of the operation and maintenance relationship between the source entity and the target entity.

[0019] Optionally, the processing unit is specifically configured to: construct a chat room corresponding to the target agent based on a multicast mechanism; and any target agent performs collaborative processing of alarm information through the chat room.

[0020] Optionally, the processing unit is specifically used to: for any alarm entity, if the alarm entity is an operation and maintenance object in the agent operation and maintenance knowledge graph, set the activation coefficient of the alarm entity to a first value; if the alarm entity is an agent in the agent operation and maintenance knowledge graph, set the activation coefficient of the alarm entity to a second value; the second value is higher than the first value; determine the activation value of the agent based on the activation coefficients of each alarm entity pointing to the same agent; and determine the agent whose activation value meets the activation condition as the target agent.

[0021] Optionally, the first prompt word also includes a third description format; the processing unit is specifically used to: input the first prompt word into the large model to obtain alarm description information that conforms to the third description format; the third description format includes the alarm entity as the alarm source, and the first correlation strength between the alarm entity and the alarm information; determine the activation value of the agent based on the activation coefficient of each alarm entity pointing to the same agent, including: for any alarm entity, determine the target value of the alarm entity based on the first correlation strength of the alarm entity and the activation coefficient of the alarm entity; determine the activation value of the agent based on the target value of each alarm entity pointing to the same agent.

[0022] Optionally, the processing unit is also used to: if the alarm entity is not an operation and maintenance object or agent in the agent operation and maintenance knowledge graph, determine the similarity between the alarm entity and the basic information of each agent; set the similarity that meets the set threshold as the activation coefficient of the alarm entity; the processing unit is specifically used to: determine the target value of the alarm entity based on the first association strength of the alarm entity, the default association strength of the non-entity and the activation coefficient of the alarm entity.

[0023] Optionally, the second description format also includes a second association strength, which is used to characterize the strength of the operation and maintenance relationship between the source entity and the target entity; the processing unit is specifically used to: determine the target value of the alarm entity based on the first association strength of the alarm entity, the second association strength corresponding to the alarm entity in the agent operation and maintenance knowledge graph, and the activation coefficient of the alarm entity.

[0024] Optionally, the processing unit is specifically used to: for the host agent, determine the convergence value of the host agent based on the degree of aggregation of each alarm entity pointing to the host agent in the physical machine, rack and computer room; determine the activation value of the agent based on the target value of each alarm entity pointing to the host agent and the convergence value of the host agent.

[0025] In a third aspect, an embodiment of the present application further provides a computing device, comprising: a memory for storing programs; a processor for calling the programs stored in the memory, and executing a large model-based alarm processing method as in the first aspect according to the obtained program.

[0026] In a fourth aspect, an embodiment of the present application further provides a computer-readable non-volatile storage medium, comprising a computer-readable program. When a computer reads and executes the computer-readable program, the computer executes a method of alarm processing based on a large model as in the first aspect.

[0027] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer device, the computer device executes the steps of the method for alarm processing based on a large model in the above-mentioned first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0029] Figure 1 A flowchart of a method for alarm processing based on a large model provided in an embodiment of the present invention;

[0030] Figure 2 A flow chart of a method for determining a target agent provided by an embodiment of the present invention;

[0031] Figure 3 A flow chart of a method for determining an agent activation value provided by an embodiment of the present invention;

[0032] Figure 4 A flow chart of a method for determining an activation value of a host agent provided by an embodiment of the present invention;

[0033] Figure 5 A flow chart of a method for constructing a chat room corresponding to an agent provided in an embodiment of the present invention;

[0034] Figure 6 A schematic diagram of the structure of an alarm processing device based on a large model provided by an embodiment of the present invention;

[0035] Figure 7 A schematic diagram of the structure of a computing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0036] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are merely some, rather than all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0037] In today's information-based society, the rapid growth and complexity of data pose new challenges to information processing. Traditional single servers or centralized systems often struggle to cope with the demands of high concurrency and large-scale data processing. Consequently, distributed systems and concurrent processing technologies have emerged. For example, the Go language has garnered widespread attention for its efficient concurrency model.

[0038] On the other hand, the development of artificial intelligence and machine learning has enabled systems to automatically learn and predict, raising the level of intelligent data processing. However, a single model is unable to cope with the high concurrency and large-scale data processing requirements. Therefore, embodiments of the present invention provide a method for alarm processing based on a large model to address the problem that a single model cannot cope with the high concurrency and large-scale data processing requirements.

[0039] like Figure 1 FIG. 1 is a flow chart of a method for alarm processing based on a large model provided by an embodiment of the present invention, the method comprising the following steps:

[0040] Step 101: Obtain at least one alarm information.

[0041] In an embodiment of the present invention, the alarm processing system includes multiple agents, wherein the agent is used to process the alarm information of the corresponding scene. In one possible case, when an alarm occurs, the alarm information is processed by activating all agents, but there is a problem. Although the alarm information is only associated with one of the agents, since the existing technology cannot associate the alarm information with the agent, it is still necessary to activate all agents to correspond to the alarm information, which will waste a lot of computing power. Therefore, the embodiment of the present invention needs to first determine which agents to activate, and then process the alarm information through the activated agents. Therefore, it is necessary to first obtain at least one alarm information to facilitate the subsequent association and determination based on the alarm information.

[0042] Step 102: Input the first prompt word into the large model to obtain at least one alarm entity corresponding to the alarm information.

[0043] In an embodiment of the present invention, a first prompt word is input into a large model, and an alarm entity corresponding to at least one alarm message is output, wherein the first prompt word includes description information of each entity and at least one alarm message, and the description information of each entity is obtained by processing the second prompt word by the large model, and the second prompt word includes basic information of each operation and maintenance object and basic information of each agent.

[0044] Step 103: Determine the target agent corresponding to the alarm entity based on the agent operation and maintenance knowledge graph.

[0045] In this embodiment of the present invention, before step 102, the second prompt word is input into the large model, which outputs description information for each entity and an agent operation and maintenance knowledge graph. The agent operation and maintenance knowledge graph represents the relationship between operation and maintenance objects and agents. Based on the agent operation and maintenance knowledge graph, the target agent corresponding to the alarm entity is determined. Since the target agent has an association with the alarm entity, the target agent is the agent that needs to be activated.

[0046] Step 104: Activate the target agent to process the corresponding alarm information.

[0047] In the embodiment of the present invention, the target agent is activated to process the corresponding alarm information, thereby specifically activating the agent to process the corresponding alarm information without wasting a lot of computing power.

[0048] From the above steps 101 to 104, it can be seen that the concept of an alarm processing system including multiple agents is introduced. Each agent is responsible for processing alarm information in a specific scenario. In this way, first, the agent operation and maintenance knowledge graph is constructed through the large model to extract the association relationship between the agent and the operation and maintenance object. Then, through the large model, the alarm information and the operation and maintenance object are associated to obtain the alarm entity. Then, combined with the agent operation and maintenance knowledge graph, the alarm is associated with the agent to determine the activated agent, thereby reducing a large amount of computing power while finding the agent corresponding to the alarm information more quickly and conveniently, so as to facilitate the subsequent processing of the alarm information according to the agent.

[0049] In this embodiment of the present invention, some alarm information can be directly associated with an agent, while some alarm information is at a lower level and can be associated with an operation and maintenance object, but cannot be directly associated with an agent. Therefore, the association between the operation and maintenance object and the agent is first determined, that is, the agent operation and maintenance knowledge graph. Then, the association between the alarm information and the entity is determined, thereby determining the alarm entity. The entity can be an operation and maintenance object or an agent. Then, based on the alarm entity, the corresponding target agent is determined, and the alarm information can be processed by activating the target agent.

[0050] The following describes how to build an agent operation and maintenance knowledge graph.

[0051] Optionally, the second prompt word is input into the big model, and the description information of each entity that conforms to the first description format is output, wherein the second prompt word includes the basic information of each operation and maintenance object, the basic information of each agent, the first description format of the entity description information, the second description format of the agent operation and maintenance knowledge graph, the first operation instruction of obtaining the agent operation and maintenance knowledge graph through each entity description information, and the second operation instruction of obtaining each entity description information. The first description format includes the entity name, entity type, and entity description. The second description format includes the agent as the source entity, the operation and maintenance object as the target entity, the operation and maintenance relationship description between the source entity and the target entity, and the second association strength. The second association strength is used to characterize the strength of the operation and maintenance relationship between the source entity and the target entity. For example, if the alarm processing system includes a host indicator agent, a micro-agent, and a DBagent, then the prompt is input into the big model to obtain the agent operation and maintenance knowledge graph. The prompt of the big model below is as follows:

[0052] Prompt starts:

[0053] #You are an AI assistant for the operation and maintenance knowledge graph. You need help abstracting the basic information of the following operation and maintenance objects and the basic information of each agent into the descriptive information of each entity and the agent operation and maintenance knowledge graph.

[0054] ##Target:

[0055] Given the basic information of each operation and maintenance object and each agent, identify all entities of these types and all relationships between the identified entities from the text.

[0056] ##Basic information of the operation and maintenance object:

[0057] First-level entity types: host metrics, business metrics, WeChat, database, network, middleware, and capacity anomalies.

[0058] Secondary entity type: Host metrics include host-related entities such as CPU, IO, and disk.

[0059] Business indicators include a decrease in success rate, an increase in time delay, a sudden drop in transaction volume, and a drop in transaction volume to zero.

[0060] WeChat includes WeChat jitter and Tencent Cloud anomalies.

[0061] Capacity anomalies include sudden increases in transaction volume and insufficient TDSQL space.

[0062] Among them, the first-level entity type is a summary of the second-level entity type.

[0063] ##step:

[0064] 1. You need to think step by step.

[0065] 2. You only need to identify the entities related to operation and maintenance. Entities include agents and operation and maintenance objects.

[0066] 3. Identify all entities. For each identified entity, extract the following information:

[0067] (1).entity_name:Entity name.

[0068] (2).entity_type: entity type, where you can select one first-level entity type and multiple second-level entity types, separated by commas.

[0069] (3).entity_description: Description of the entity.

[0070] Each entity output is formatted as JSON using the following format:

[0071] {{"entity_name":<entity name> ,"entity_type": <type>,"entity_description":<entity description>}}

[0072] 4. From the entities identified in step 3, identify all pairs (source_entity, target_entity) that are "significantly related" to each other. For each pair of related entities, extract the following information:

[0073] (1).source_entity: source entity, which is a specific agent.

[0074] (2).target_entity: target entity, which refers to the specific operation and maintenance objects.

[0075] (3).relationship_description: Explain why you think the source entity and target entity are related to each other.

[0076] (4).relationship_strength: integer, ranging from 1 to 10, named as the second association strength, used to indicate the strength of the relationship between the source entity and the target entity.

[0077] {{"source_entity":<source_entity> ,"target_entity":<target_entity> relationship_description:<relationship_description> relationship_strength<relationship_strength>}}

[0078] 5. Each agent needs to be a source_entity.

[0079] Here are the real data:

[0080] Basic information of each agent:

[0081] 1. tdsql agent, determines whether the tdsql database is healthy and whether the program error is caused by DB abnormality.

[0082] 2. WeChat agent is used to determine whether there are any abnormalities in the WeChat account of external partners.

[0083] 3. Host identification agent, used to determine whether the application host is healthy, including disk, CPU, IO, and memory.

[0084] The prompt ends.

[0085] It should be noted that the first description format of the entity description information is {{"entity_name":<entity name> ,"entity_type": <type>,"entity_description": <entitydescription>}}.

[0086] The second description format of the agent operation and maintenance knowledge graph is {{"source_entity":<source_entity> ,"target_entity":<target_entity> relationship_description:<relationship_description> relationship_strength<relationship_strength>}}

[0087] The first operation instructions for obtaining the agent operation and maintenance knowledge graph from the entity description information are steps 1, 2, and 3.

[0088] The second operation instructions for obtaining description information of each entity are steps 4 and 5.

[0089] The large model outputs the description information of each entity and the agent operation and maintenance knowledge graph based on the input information in the prompt. The description information of each entity is as follows:

[0090] {"entity_name":"tdsql agent","entity_type":"database","entity_description":"Judge whether the tdsql database is healthy and whether the program error is caused by a DB exception."},

[0091] {"entity_name":"DB exception","entity_type":"Database","entity_description":"The program error may be caused by a DB exception."},

[0092] {"entity_name":"WeChat agent","entity_type":"WeChat","entity_description":"Determine whether the external partner's WeChat account is abnormal."},

[0093] {"entity_name":"Host Identification Agent","entity_type":"Host Metrics","entity_description":"Determines whether the application host is healthy, including disk, CPU, IO, and memory."},

[0094] {"entity_name":"Disk","entity_type":"Host Metric","entity_description":"A host-type metric determined by the host identification agent."},

[0095] {"entity_name":"cpu","entity_type":"Host metric","entity_description":"A host-type metric determined by the host identification agent."},

[0096] {"entity_name":"io","entity_type":"Host indicator","entity_description":"A host-type indicator determined by the host identification agent."},

[0097] {"entity_name":"Memory","entity_type":"Host Metric","entity_description":"A host-type metric determined by the host identification agent."}

[0098] The agent operation and maintenance knowledge graph is as follows:

[0099] {"source_entity": "tdsql agent", "target_entity": "DB exception", "relationship_description": "The tdsql agent is used to determine whether the database is healthy and whether program errors are caused by DB exceptions.", "relationship_strength": 8}, {"source_entity": "Host identification agent", "target_entity": "Disk", "relationship_description": "When the host identification agent determines the health of the host, it includes the disk.", "relationship_strength": 7}, {"source_entity": "Host identification agent", "target_entity": "CPU", "relationship_description": "When the host identification agent determines the health of the host, it includes the CPU.", "relationship_strength": 7}, {"source_entity": "Host identification agent", "target_entity": "IO", "relationship_description": "When the host identification agent determines the health of the host, it includes the IO.", "relationship_strength": 7}, {"source_entity": "Host identification agent", "target_entity": "Memory", "relationship_description": "When the host identification agent determines the health of the host, it includes the memory.", "relationship_strength": 7},

[0100] {"source_entity": "WeChat agent", "target_entity": "WeChat", "relationship_description": "The WeChat agent is responsible for determining whether the external partner WeChat is abnormal, so it has a direct functional relationship with WeChat.", "relationship_strength": 8}.

[0101] The entity description information and the agent operation and maintenance knowledge graph are determined by the large model, so as to determine the association relationship between each entity and the agent. Then, the association relationship between the alarm information and the entity is determined according to the large model, so that the alarm information can be associated with the agent. The following introduces how to determine the association relationship between the alarm information and the entity according to the large model.

[0102] Optionally, the first prompt word is input into the large model to obtain alarm description information that conforms to a third description format, wherein the third description format includes the alarm entity serving as the alarm source and a first correlation strength between the alarm entity and the alarm information. The first prompt word includes description information of each entity, at least one alarm information, and a third operation instruction for obtaining the alarm description information that conforms to the third description format. The third description format includes the alarm entity serving as the alarm source and a first correlation strength between the alarm entity and the alarm information.

[0103] For example, suppose there are three alarms in the current system: CPU alarm, I / O alarm, and WeChat jitter. Then put the alarm information and entity description information into the prompt and start making a request to the big model. The prompt of the big model is as follows:

[0104] Prompt Start

[0105] #You are an AI assistant in the operation and maintenance knowledge graph, and need help abstracting the alarm information and entity description information to extract the relationship between the alarm information and the entity.

[0106] ##Target:

[0107] Given warning information and entity descriptions that may be relevant to the task, identify all entities of these types from the text, as well as the first relationship strength between the warning information and the entity, from 1-10.

[0108] ##Description of each entity:

[0109] {"entity_name":"tdsql agent","entity_type":"database","entity_description":"Judge whether the tdsql database is healthy and whether the program error is caused by a DB exception."},

[0110] {"entity_name":"DB exception","entity_type":"Database","entity_description":"The program error may be caused by a DB exception."},

[0111] {"entity_name":"WeChat agent","entity_type":"WeChat","entity_description":"Determine whether the external partner's WeChat account is abnormal."},

[0112] {"entity_name":"Host Identification Agent","entity_type":"Host Metrics","entity_description":"Determines whether the application host is healthy, including disk, CPU, IO, and memory."},

[0113] {"entity_name":"Disk","entity_type":"Host Metric","entity_description":"A host-type metric determined by the host identification agent."},

[0114] {"entity_name":"cpu","entity_type":"Host metric","entity_description":"A host-type metric determined by the host identification agent."},

[0115] {"entity_name":"io","entity_type":"Host indicator","entity_description":"A host-type indicator determined by the host identification agent."},

[0116] {"entity_name":"Memory","entity_type":"Host Metric","entity_description":"A host-type metric determined by the host identification agent."}

[0117] ##step

[0118] 1. You need to think step by step.

[0119] 2. You only need to identify entities related to operations and maintenance.

[0120] 3. For the alarm information and each entity, extract the following information:

[0121] (1).source_entity: source entity, which is each entity in the entity description information

[0122] (2).target_entity: target entity, which is the alarm information

[0123] (3) relationship_strength: integer, ranging from 1 to 10, indicating the strength of the second relationship between the alarm information and each entity.

[0124] {"source_entity":<source_entity> relationship_strength<relationship_strengt h>}

[0125] The warning information is as follows:

[0126] 1. [major [Panic] host indicator alarm]

[0127] Alarm content: CPU usage is greater than 90%.

[0128] 2. [major[Panic] keyword error warning]

[0129] Alarm content: Request to WeChat / api / wx / xxx interface is abnormal

[0130] 3. [Minor alarm io high]

[0131] Alarm content: io util is greater than 30%.

[0132] prompt ends

[0133] It should be noted that the large model outputs alarm description information that conforms to the third description format, where the alarm description information is as follows:

[0134] {"source_entity":"cpu","relationship_strength":10},

[0135] {"source_entity":"WeChat agent","relationship_strength":8},

[0136] {"source_entity":"io","relationship_strength":6}

[0137] The third operation instruction includes steps 1-3.

[0138] From the above example, we can see that the alarm entities output by the large model are cpu, WeChat agent and io, and then the target agent can be determined from each agent based on the alarm entity.

[0139] like Figure 2 FIG. 1 is a flow chart of a method for determining a target agent provided by an embodiment of the present invention, the method comprising the following steps:

[0140] Step 201: Determine the activation coefficient of each alarm entity.

[0141] In an embodiment of the present invention, if the alarm entity is an operation and maintenance object in the agent operation and maintenance knowledge graph, the activation coefficient of the alarm entity is set to a first value. If the alarm entity is an agent in the agent operation and maintenance knowledge graph, the activation coefficient of the alarm entity is set to a second value; the second value is higher than the first value. For example, if the alarm entities are cpu, WeChat agent, and io. Among them, cpu and io are operation and maintenance objects in the agent operation and maintenance knowledge graph, the activation coefficient of the alarm entity is set to the first value, wherein the first value is 0.8. WeChat agent is an agent in the agent operation and maintenance knowledge graph, then the activation coefficient of the alarm entity is set to the second value, wherein the second value is 1.

[0142] Step 202: Determine the activation value of the agent according to the activation coefficients of the alarm entities pointing to the same agent.

[0143] In the embodiment of the present invention, Figure 3 FIG. 1 is a flow chart of a method for determining an agent activation value according to an embodiment of the present invention. The method includes the following steps:

[0144] Step 301: For any alarm entity, determine a target value of the alarm entity according to the first association strength of the alarm entity and the activation coefficient of the alarm entity.

[0145] In an embodiment of the present invention, in one possible case, the alarm entities are all operation and maintenance objects or agents in the agent operation and maintenance knowledge graph, and the target value of the alarm entity is determined based on the first association strength of the alarm entity, the second association strength of the alarm entity and the activation coefficient of the alarm entity.

[0146] For example, if the alarm entity is cpu, where cpu is the operation and maintenance object in the agent operation and maintenance knowledge graph, the target value of the alarm entity is determined according to formula 1, as shown in formula 1:

[0147] Target value of alarm entity = first association strength of alarm entity * second association strength of alarm entity * activation coefficient of alarm entity Formula 1

[0148] Then if the first association strength of cpu is 10, the second association strength of cpu is 7, and the activation coefficient of cpu is 0.8, then the target value of cpu = 10*7*0.8 = 56.

[0149] For another example, if the alarm entity is the WeChat agent, where the WeChat agent is the agent of the agent operation and maintenance knowledge graph, the first association strength of the WeChat agent is 8, the second association strength of the WeChat agent is 8, and the activation coefficient of the WeChat agent is 1, then the target value of the WeChat agent = 8*8*1=64.

[0150] In another possible scenario, if the alarm entity is not an operation and maintenance object or agent in the agent operation and maintenance knowledge graph, the similarity between the alarm entity and the basic information of each agent is determined; the similarity that meets the set threshold is set as the activation coefficient of the alarm entity. The target value of the alarm entity is determined based on the first association strength of the alarm entity, the default association strength of the non-entity, and the activation coefficient of the alarm entity.

[0151] For example, if the alarm entity is a WeChat reminder, which is not an operation and maintenance object or agent in the agent operation and maintenance knowledge graph, then based on the WeChat reminder and the description of each agent, the similarity between the alarm entity and the basic information of each agent is determined, and then the similarity that meets the set threshold is set as the activation coefficient of the alarm entity. The similarity between the WeChat reminder and the basic information of the WeChat agent is 0.54, the similarity between the WeChat reminder and the basic information of the DBagent is 0.33, and the similarity between the WeChat reminder and the basic information of the host indicator agent is 0.32. The similarity between the WeChat agent and the WeChat reminder is the highest, meeting the set threshold. The target value of the WeChat reminder is then determined according to Formula 2, as shown in Formula 2:

[0152] Target value of alarm entity = first association strength of alarm entity * default association strength of non-entity * activation coefficient of alarm entity Formula 2

[0153] The first association strength of WeChat reminder is 7, the default association strength of non-entity is 4, and the activation coefficient of the alarm entity is 0.54. Then the target value of WeChat reminder is 7*4*0.54=15.12.

[0154] Step 302: Determine the activation value of the agent according to the target values ​​of the alarm entities pointing to the same agent.

[0155] In this embodiment of the present invention, for example, if the target value for CPU is 56, the target value for the host metric agent is 160, the target value for IO is 33.6, the target value for WeChat agent is 64, and the target value for WeChat reminder is 15.12, then the target values ​​of each alarm entity pointing to the same agent are accumulated to obtain the agent activation value. The target values ​​for CPU, host metric agent, and IO are accumulated to obtain the activation value of the host metric agent, which is 249.6. The target values ​​for WeChat agent and WeChat reminder are accumulated to obtain the activation value of WeChat agent, which is 79.12.

[0156] From the above steps 301 to 302, it can be seen that by considering different methods for determining the target value of the alarm entity in different situations, the target value of each alarm entity can be determined more accurately, and then the target value of the agent corresponding to the alarm entity can be determined more accurately.

[0157] Step 203: An agent whose activation value satisfies the activation condition is determined as a target agent.

[0158] In the embodiment of the present invention, for example, if the activation condition is to activate one agent, then based on the activation values ​​of the agents, the agent with the highest activation value is taken as the target agent.

[0159] For another example, if the activation condition is to activate three agents, then based on the activation values ​​of each agent, the agent with the highest three-dimensional activation values ​​is taken as the target agent.

[0160] As can be seen from steps 201 and 202 above, given that the number of agents associated with an alarm may be large, activating many agents simultaneously will generate a significant amount of computing power. Therefore, the agents that need to be activated can be determined by calculating the activation value of each agent. This allows for faster problem location while reducing the number of simultaneously activated agents.

[0161] like Figure 4 FIG. 1 is a flow chart of a method for determining an activation value of a host agent provided by an embodiment of the present invention. The method includes the following steps:

[0162] Step 401 : for a host agent, the aggregation degree of each alarm entity pointing to the host agent in the physical machine, rack and computer room is determined to determine the convergence value of the host agent.

[0163] In the embodiment of the present invention, host indicators are divided into the following levels:

[0164] 1. Instance level (such as virtual machines, containers); 2. Physical machine level; 3. Rack level; 4. Computer room level.

[0165] The purpose is to confirm from a large number of alarms whether the alarm is a single instance failure or concentrated in certain physical units.

[0166] Optionally, when one-third of the instances on a physical machine have problems, the physical machine is marked as failed.

[0167] Optionally, when one-third of the physical machines in a rack have problems, the rack is marked as failed.

[0168] Optionally, when one-third of the racks in a computer room have problems, the computer room is marked as faulty.

[0169] Optionally, an embodiment of the present invention proposes a multi-dimensional dictionary structure for storing and processing host hierarchical information. The dictionary structure can effectively map instance information to corresponding hierarchical nodes.

[0170] The following are the specific steps to build the dictionary structure:

[0171] The first step is information collection: First, collect instance information, including instance name, physical machine number, rack, and computer room.

[0172] The second step is dictionary structure: establish a three-level dictionary, organized from largest to smallest physical units:

[0173] {Computer room->Rack->Physical machine->Instance name->true / false}

[0174] The third step is the dictionary construction process:

[0175] a. Create a new blank dictionary `A` with the value {}.

[0176] b. For each instance of information, loop through it as follows:

[0177] The following is the cycle process

[0178] (1) Check if the dictionary `A` has the corresponding computer room key. If not, create a new computer room entry in the format: computer room: {rack: {physical machine: {instance: false}}}

[0179] Then exit this loop. If it exists, set dictionary `B` to the value of the corresponding computer room in dictionary `A`.

[0180] (2) Check if the corresponding rack key exists in dictionary `B`. If not, create a new rack entry in the format: rack:{physical machine:{instance:false}}

[0181] Then exit this loop. If it exists, set dictionary `B` to the value of the corresponding rack in dictionary `B`.

[0182] (3) Check if the corresponding physical machine key exists in dictionary `B`. If not, create a new physical machine entry in the format: physical machine:{instance:false}, and then exit this loop. If it exists, set the value of the corresponding key to `false`.

[0183] (4) If the physical machine key exists, then the dictionary A[computer room][rack][physical machine][instance name]=false is assigned.

[0184] This completes one instance traversal. Repeat this process until all instances are traversed. The final dictionary structure is as follows: Dictionary A [Data Center] [Rack] [Physical Machine] [Instance Name]. This dictionary structure is used to record and track the fault status of the host architecture, which helps to achieve efficient fault identification and location.

[0185] Optionally, the embodiment of the present invention can identify faults in instances, physical machines, racks, and computer rooms through multi-level traversal and alarm concentration analysis. The following describes the specific steps for locating alarm information to instances, physical machines, racks, and computer rooms:

[0186] The first step is to calculate the warning score for each instance. If the warning score is greater than 2, the instance is marked as an abnormal instance.

[0187] In the second step, the alarm of dictionary A is activated, and the hierarchical structure of dictionary A (i.e., each computer room, each rack, each physical machine, and each instance) is traversed.

[0188] In the third step, if the alarm score of an instance is greater than 2, set `Dictionary A[computer room][rack][physical machine][instance name] = true.

[0189] The fourth step is to traverse the dictionary A, analyze the alarm concentration, and build the physical machine list L1 to be activated and the faulty rack list L0.

[0190] For example, suppose there are alarms for the following instances:

[0191] -Guangzhou-Rack2-PhysicalMachine3-gz-1

[0192] -Guangzhou-Rack2-PhysicalMachine3-gz-2

[0193] -Guangzhou-Rack 2-Physical Machine 3-gz-3

[0194] -Guangzhou-Rack2-PhysicalMachine5-gz-4

[0195] -Guangzhou-Rack 3-Physical Machine 6-gz-5

[0196] -Shanghai-Rack2-PhysicalMachine7-sh-1

[0197] Step 5: Physical machine fault detection:

[0198] a. Create a list L1 of physical machines to be activated.

[0199] b. Loop through each rack level (rack name - physical machine dictionary).

[0200] c. At the physical machine level, if the alarm instance ratio is greater than 1 / 3, the physical machine name is added to the L1 list.

[0201] Here is an example:

[0202] a. Physical machine 3 in rack 2 in Guangzhou has three instances (gz-1, gz-2, and gz-3). The alarm scores of gz-1 and gz-2 are greater than 2. Therefore, physical machine 3 is marked as a failed physical machine and added to the L1 list.

[0203] b. Physical machine 5 in rack 2 in Guangzhou has one instance (gz-4) whose alarm score is greater than 2. Therefore, physical machine 5 is marked as a faulty physical machine and added to the L1 list.

[0204] The instances in Guangzhou-Rack 3-Physical Machine 6 and Shanghai-Rack 2-Physical Machine 7 do not reach 1 / 3, so they are not faulty physical machines. The L1 list result is: [Physical Machine 3, Physical Machine 5]

[0205] Step 6: Rack fault detection:

[0206] Create a faulty rack list L0.

[0207] If more than one-third of the physical machines in a rack are faulty, the rack name is added to the L0 list.

[0208] Here is an example:

[0209] Guangzhou rack 2 has three physical machines, all of which are faulty. Therefore, rack 2 is marked as a faulty rack and added to the L0 list.

[0210] Guangzhou rack 3 has no faults and is not added to the L0 list.

[0211] The L0 list result is: [Rack 2]

[0212] Step 7: Computer room fault detection:

[0213] If more than one-third of the racks in a computer room are faulty, the room is marked as a faulty room.

[0214] Here is an example:

[0215] The Guangzhou data center has a total of 10 racks. The L0 list contains one faulty rack (rack 2). The faulty rack ratio is less than 1 / 3, so the Guangzhou data center is marked as normal.

[0216] Through the above algorithm, the fault aggregation algorithm only aggregates faults at the Guangzhou-Rack 2 level. This enables multi-level alarm score calculation and fault identification from the instance level to the computer room level, and efficiently manages and analyzes alarm information through a multidimensional dictionary structure.

[0217] When the alarm convergence can converge to the corresponding level, it indicates that the anomaly may be caused by a basic component failure, which will trigger the host entity and then assign the corresponding weight according to the convergence level.

[0218] For example, the weight value at the physical machine level is 8, the weight value at the rack level is 16, and the weight value at the computer room level is 32.

[0219] Step 402: Determine the activation value of the host agent according to the target value of each alarm entity directed to the host agent and the convergence value of the host agent.

[0220] In the embodiment of the present invention, the activation value of the host agent is determined by adding the target value of each alarm entity directed to the host agent to the convergence value of the host agent.

[0221] It can be seen from the above steps 401 to 402 that, considering that the alarm information of the host indicator may be related to the hardware, on the basis of determining the target value of each alarm entity pointing to the host agent, the convergence value of the hardware corresponding to the host indicator is added, so that the activation value of the host agent can be determined more comprehensively and accurately.

[0222] like Figure 5 FIG. 1 is a flow chart of a method for constructing a chat room corresponding to an agent provided by an embodiment of the present invention. The method includes the following steps:

[0223] Step 501: Building a chat room corresponding to the target agent based on the multicast mechanism.

[0224] In the embodiment of the present invention, the agent is encapsulated into an agent worker. The worker receives two channels: a task channel, which is responsible for receiving messages, and a results channel, which is responsible for processing the messages and passing them back.

[0225] In order to achieve communication between agents, it is necessary to define a structure containing the following fields:

[0226] send: The name of the sender Agent.

[0227] receiver: The name of the receiver Agent. It can be empty, indicating that the message is sent to all Agents.

[0228] send_time: Sending time.

[0229] id: A unique ID generated.

[0230] Step 502: Any target agent performs collaborative processing of alarm information through the chat room.

[0231] In the embodiment of the present invention, the following are specific steps for implementing a Channel-based chat room:

[0232] Create a new task channel list and result channel list.

[0233] Create a multicast process.

[0234] Loop through the task channel and result channel lists, pair by pair, and then start the work agent's Ctrip.

[0235] a. Publishing topics: The publisher generates a topic (for example, which alarms are currently available), and uses a for loop to send the alarm information to all work agents through the task channel.

[0236] b. After each worker agent receives the task channel, it checks whether the Receiver field is itself (the default empty field indicates that the message is sent to everyone, that is, the agent chat room). If so, the agent's logical processing begins.

[0237] c. After receiving the topic, each worker agent starts processing the topic and returns the processing results to the superior through the result channel.

[0238] d. Completion notification: After the worker agent completes the topic processing, it sends the completion result to the result channel.

[0239] e. Result integration: The publisher listens to the results of the result channel. After receiving each returned result, it embeds the result into the prompt and checks whether it meets the requirements. If not, it continues to ask questions to the subscriber.

[0240] At this point, the multi-agent processing is completed.

[0241] From step 501 to step 502, it can be seen that the high concurrency characteristic of go is utilized to realize multicast, and the agent list is calculated through the agent activation algorithm to realize dynamic scheduling of agents and automatic processing of alarms.

[0242] Based on the same technical concept, the embodiment of the present application provides a device for alarm processing based on a large model, such as Figure 6 As shown, the device 600 is applied to an alarm processing system with multiple intelligent agents, each agent is used to process the alarm information of the corresponding scenario, including: an acquisition unit 601, used to obtain at least one alarm information; a processing unit 602, used to input a first prompt word into a large model to obtain an alarm entity corresponding to at least one alarm information; the first prompt word includes description information of each entity and at least one alarm information; the description information of each entity is obtained by processing the second prompt word by the large model; the second prompt word includes basic information of each operation and maintenance object and basic information of each agent; based on the agent operation and maintenance knowledge graph, the target agent corresponding to the alarm entity is determined; the agent operation and maintenance knowledge graph is obtained by processing the second prompt word by the large model; the agent operation and maintenance knowledge graph represents the association relationship between the operation and maintenance object and the agent; and the target agent is activated to process the corresponding alarm information.

[0243] Optionally, the second prompt word also includes a first description format of the entity description information; the processing unit 602 is specifically used to: based on the basic information of each operation and maintenance object and the basic information of each agent in the second prompt word, obtain the entity description information that conforms to the first description format through the large model; the first description format includes the entity name, entity type and entity description.

[0244] Optionally, the second prompt word also includes a second description format of the agent operation and maintenance knowledge graph and a first operation instruction for obtaining the agent operation and maintenance knowledge graph through the description information of each entity; the processing unit 602 is also used to: according to the first operation instruction, the large model obtains the agent operation and maintenance knowledge graph that conforms to the second description format through the description information of each entity; the second description format includes the agent as the source entity, the operation and maintenance object as the target entity, and a description of the operation and maintenance relationship between the source entity and the target entity.

[0245] Optionally, the processing unit 602 is specifically configured to: construct a chat room corresponding to the target agent based on a multicast mechanism; and any target agent performs collaborative processing of alarm information through the chat room.

[0246] Optionally, the processing unit 602 is specifically used to: for any alarm entity, if the alarm entity is an operation and maintenance object in the agent operation and maintenance knowledge graph, set the activation coefficient of the alarm entity to a first value; if the alarm entity is an agent in the agent operation and maintenance knowledge graph, set the activation coefficient of the alarm entity to a second value; the second value is higher than the first value; determine the activation value of the agent based on the activation coefficients of each alarm entity pointing to the same agent; and determine the agent whose activation value meets the activation condition as the target agent.

[0247] Optionally, the first prompt word also includes a third description format; the processing unit 602 is specifically used to: input the first prompt word into the large model to obtain alarm description information that conforms to the third description format; the third description format includes the alarm entity as the alarm source, and the first correlation strength between the alarm entity and the alarm information; determine the activation value of the agent based on the activation coefficient of each alarm entity pointing to the same agent, including: for any alarm entity, determine the target value of the alarm entity based on the first correlation strength of the alarm entity and the activation coefficient of the alarm entity; determine the activation value of the agent based on the target value of each alarm entity pointing to the same agent.

[0248] Optionally, the processing unit 602 is also used to: if the alarm entity is not an operation and maintenance object or agent in the agent operation and maintenance knowledge graph, determine the similarity between the alarm entity and the basic information of each agent; set the similarity that meets the set threshold as the activation coefficient of the alarm entity; the processing unit 602 is specifically used to: determine the target value of the alarm entity according to the first association strength of the alarm entity, the default association strength of the non-entity and the activation coefficient of the alarm entity.

[0249] Optionally, the second description format also includes a second association strength, which is used to characterize the strength of the operation and maintenance relationship between the source entity and the target entity; the processing unit 602 is specifically used to: determine the target value of the alarm entity based on the first association strength of the alarm entity, the second association strength corresponding to the alarm entity in the agent operation and maintenance knowledge graph, and the activation coefficient of the alarm entity.

[0250] Optionally, the processing unit 602 is specifically used to: for the host agent, determine the convergence value of the host agent based on the degree of aggregation of each alarm entity pointing to the host agent in the physical machine, rack and computer room; determine the activation value of the agent based on the target value of each alarm entity pointing to the host agent and the convergence value of the host agent.

[0251] Based on the same technical concept, the embodiment of the present application provides a computing device 700, such as Figure 7 As shown, it includes at least one processor 701 and a memory 702 connected to the at least one processor. The specific connection medium between the processor 701 and the memory 702 is not limited in the embodiment of the present application. Figure 7 For example, the processor 701 and the memory 702 are connected via a bus. The bus can be divided into an address bus, a data bus, a control bus, and the like.

[0252] In an embodiment of the present application, the memory 702 stores instructions that can be executed by at least one processor 701. By executing the instructions stored in the memory 702, the at least one processor 701 can perform the steps of the above-mentioned large model-based alarm processing method.

[0253] The processor 701 is the control center of the computing device and can connect the various parts of the computing device using various interfaces and lines. By running or executing instructions stored in the memory 702 and calling data stored in the memory 702, the method of alarm processing based on the large model is implemented. Optionally, the processor 701 may include one or more processing units. The processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understood that the above-mentioned modem processor can also be integrated into the processor 701. In some embodiments, the processor 701 and the memory 702 can be implemented on the same chip. In some embodiments, they can also be implemented separately on separate chips.

[0254] The processor 701 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method for alarm processing based on a large model disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0255] The memory 702 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 702 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 702 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer device, but is not limited thereto. The memory 702 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0256] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program that can be executed by a computer device. When the program runs on the computer device, the computer device executes the steps of the above-mentioned large model-based alarm processing method.

[0257] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0258] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0259] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0260] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0261] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.< / entitydescription> < / type> < / type>

Claims

1. A method for alarm processing based on a large model, characterized in that: Applied to an alarm processing system having multiple intelligent agents, each agent is used to process alarm information of a corresponding scenario, the method includes: Obtain at least one alarm information; Inputting a first prompt word into a large model to obtain an alarm entity corresponding to the at least one alarm message; the first prompt word includes description information of each entity and the at least one alarm message; the description information of each entity is obtained by processing the second prompt word by the large model; the second prompt word includes basic information of each operation and maintenance object and basic information of each agent; Determining a target agent corresponding to the alarm entity based on an agent operation and maintenance knowledge graph; the agent operation and maintenance knowledge graph is obtained by processing the second prompt word by the large model; the agent operation and maintenance knowledge graph represents the association relationship between the operation and maintenance object and the agent; the agent operation and maintenance knowledge graph entity includes the operation and maintenance object and the agent; activating the target agent to process the corresponding alarm information; Based on the agent operation and maintenance knowledge graph, the target agent corresponding to the alarm entity is determined, including: For any alarm entity, if the alarm entity is an operation and maintenance object in the agent operation and maintenance knowledge graph, the activation coefficient of the alarm entity is set to a first value; if the alarm entity is an agent in the agent operation and maintenance knowledge graph, the activation coefficient of the alarm entity is set to a second value; the second value is higher than the first value; the activation value of the agent is determined based on the activation coefficients of each alarm entity pointing to the same agent; the agent whose activation value meets the activation condition is determined as the target agent; The first prompt word further includes a third description format; the first prompt word is input into the large model to obtain an alarm entity corresponding to the at least one alarm information, including: Inputting the first prompt word into the large model to obtain alarm description information that conforms to the third description format; the third description format includes an alarm entity as an alarm source and a first correlation strength between the alarm entity and the alarm information; The activation value of the agent is determined based on the activation coefficients of each alarm entity pointing to the same agent, including: For any alarm entity, determining a target value of the alarm entity according to the first association strength of the alarm entity and the activation coefficient of the alarm entity; The activation value of the agent is determined based on the target values ​​of each alarm entity pointing to the same agent. For the host agent, the convergence value of the host agent is determined by the degree of aggregation of each alarm entity pointing to the host agent in the physical machine, rack, and computer room. The activation value of the agent is determined based on the target values ​​of each alarm entity pointing to the host agent and the convergence value of the host agent.

2. The method according to claim 1, wherein The second prompt word also includes a first description format of entity description information; The processing of the second prompt word by the large model includes: Based on the basic information of each operation and maintenance object and the basic information of each agent in the second prompt word, the entity description information that conforms to the first description format is obtained through the large model; the first description format includes entity name, entity type and entity description.

3. The method according to claim 2, wherein The second prompt word also includes a second description format of the agent operation and maintenance knowledge graph and a first operation instruction for obtaining the agent operation and maintenance knowledge graph through the description information of each entity; After obtaining the description information of each entity in accordance with the first description format, the method further includes: According to the first operation instruction, the large model obtains the agent operation and maintenance knowledge graph conforming to the second description format through the description information of each entity; The second description format includes an agent as a source entity, an operation and maintenance object as a target entity, and a description of the operation and maintenance relationship between the source entity and the target entity.

4. The method according to claim 1, wherein Activating the target agent to process the corresponding alarm information includes: Building a chat room corresponding to the target agent based on the multicast mechanism; Any target agent performs collaborative processing of alarm information through the chat room.

5. The method according to claim 1, wherein Also includes: If the alarm entity is not an operation and maintenance object or agent in the agent operation and maintenance knowledge graph, determine the similarity between the alarm entity and the basic information of each agent; Setting the similarity that meets the set threshold as the activation coefficient of the alarm entity; Determining a target value of the alarm entity according to the first association strength of the alarm entity and the activation coefficient of the alarm entity includes: A target value of the alarm entity is determined according to the first association strength of the alarm entity, the default association strength of the non-entity, and the activation coefficient of the alarm entity.

6. The method according to claim 3, wherein The second description format further includes a second association strength, where the second association strength is used to characterize the strength of the operation and maintenance relationship between the source entity and the target entity; Determining a target value of the alarm entity according to the first association strength of the alarm entity and the activation coefficient of the alarm entity includes: The target value of the alarm entity is determined according to the first association strength of the alarm entity, the second association strength corresponding to the alarm entity in the agent operation and maintenance knowledge graph, and the activation coefficient of the alarm entity.

7. A device for alarm processing based on a large model, characterized in that: Applied to an alarm processing system with multiple intelligent agents, each of which is responsible for processing alarm information for a specific scenario, including: An acquiring unit, configured to acquire at least one piece of alarm information; A processing unit is configured to input a first prompt word into a large model to obtain an alarm entity corresponding to the at least one alarm message; the first prompt word includes description information of each entity and the at least one alarm message; the description information of each entity is obtained by processing a second prompt word by the large model; the second prompt word includes basic information of each operation and maintenance object and basic information of each agent; based on an agent operation and maintenance knowledge graph, determine a target agent corresponding to the alarm entity; the agent operation and maintenance knowledge graph is obtained by processing the second prompt word by the large model; the agent operation and maintenance knowledge graph represents an association relationship between operation and maintenance objects and agents; entities in the agent operation and maintenance knowledge graph include operation and maintenance objects and agents; and activate the target agent to process the corresponding alarm message; The processing unit is further configured to, for any alarm entity, set an activation coefficient of the alarm entity to a first value if the alarm entity is an operation and maintenance object in the agent operation and maintenance knowledge graph; set an activation coefficient of the alarm entity to a second value if the alarm entity is an agent in the agent operation and maintenance knowledge graph; the second value is higher than the first value; determine an activation value of the agent based on the activation coefficients of each alarm entity pointing to the same agent; and determine an agent whose activation value meets the activation condition as a target agent; The processing unit is further configured to input the first prompt word into the large model to obtain alarm description information conforming to a third description format; the third description format includes an alarm entity as an alarm source and a first correlation strength between the alarm entity and the alarm information; The processing unit is further configured to determine, for any alarm entity, a target value of the alarm entity based on the first association strength of the alarm entity and the activation coefficient of the alarm entity; and determine an activation value of the agent based on the target values ​​of each alarm entity directed to the same agent; The processing unit is further configured to determine the convergence value of the host agent based on the degree of aggregation of the alarm entities pointing to the host agent in the physical machine, rack, and computer room; and determine the activation value of the agent based on the target value of each alarm entity pointing to the host agent and the convergence value of the host agent.

Citation Information

Patent Citations

  • Micro-service fault diagnosis method and device based on large language model and electronic equipment

    CN117891640A

  • Human-computer interaction method and device

    CN118551000A