Abnormality positioning method, system and device
Patent Information
- Application Number
- CN202510336592.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2026-09-22
Smart Images

Figure CN122802345A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and in particular to an anomaly location method, system, and apparatus. Background Technology
[0002] A 10,000-card cluster refers to a high-performance computing system composed of 10,000 or more accelerators, used for training and inference of fundamental large-scale artificial intelligence (AI) models. Accelerator cards can be graphics processing units (GPUs), neural network processing units (NPUs), or other possible processors. Currently, with the rapid development of AI technology, the construction of intelligent computing centers has entered a period of rapid growth. As a major evolutionary trend for future intelligent computing centers, how to reduce the impact of abnormal operation of intelligent computing clusters on model training and ensure the stability of intelligent computing clusters is a current research hotspot. Summary of the Invention
[0003] This application provides an anomaly location method, system, and apparatus to improve the operational stability of intelligent computing clusters.
[0004] To achieve the above objectives, this application adopts the following technical solution:
[0005] In a first aspect, an anomaly localization method is provided, applied to a first entity, which may be a terminal device, a chip in the terminal device, or a device for implementing the functions of the terminal device. The first entity may also be a network device, a chip in the network device, or a device for implementing the functions of the network device. The method includes: sending a first message indicating the correspondence between multiple data and multiple nodes, wherein the multiple data are generated by the operation of the multiple nodes, at least one of the multiple data includes at least one modality of data, and at least two of the multiple nodes are used to perform services with different functions; and receiving a second message indicating a first node, wherein the first node is an abnormal node determined among the multiple nodes according to the first message.
[0006] Therefore, this method involves a first entity sending a first message to a second entity to indicate the correspondence between multiple data points and multiple nodes. These multiple nodes are used to perform different functions of business, such as multiple nodes coming from at least two different functional domains. The multiple data points are data of at least one modality generated by the operation of multiple nodes. Thus, the second entity can determine the first node, i.e. the abnormal node, from among the multiple nodes based on the first message. This enables anomaly analysis of multi-domain and multi-modal operational data and accurate location of abnormal nodes, thereby improving the stability of complex intelligent computing clusters.
[0007] In one possible design, at least two of the multiple nodes belong to at least two different types of domains. The type of the domain is related to the business performed by the domain. In other words, the type of the node can be divided at the domain level to improve data processing efficiency.
[0008] Optionally, before sending the first information, the method of the first aspect further includes: determining multiple nodes in the topology information of at least two domains, and determining multiple data in the operational data of at least two domains; determining the first information based on the multiple data and the multiple nodes. The topology information is used to indicate all nodes contained in the domain and the relationships between the nodes, and the operational data is used to indicate all data generated by all nodes in the domain during operation. That is, the multiple data can be all or part of the operational data of the domain, and the multiple nodes can be all or part of the nodes in the domain, allowing for flexible configuration of the first information to adapt to more scenarios.
[0009] In one possible design, the first aspect of the method further includes: receiving third information, which is used to indicate a business anomaly; and sending first information, including: sending the first information based on the third information.
[0010] Optionally, the third information includes the identifier of the abnormal service, and the method of the first aspect further includes: determining multiple nodes in the topology information of at least two domains and determining multiple data in the running data of at least two domains based on the identifier of the abnormal service; and determining the first information based on the multiple data and the multiple nodes.
[0011] Therefore, the first entity can identify some nodes as multiple nodes from all nodes in the domain and some data as multiple data from all data in the domain by using abnormal situations (i.e., the identifiers of abnormal services). In this way, it can identify and send the first information, reduce the amount of data processing for the first entity, and improve the efficiency of subsequent anomaly location.
[0012] In one possible design, the first aspect of the method further includes sending a fourth message, which instructs the second entity to identify the malfunctioning node among multiple nodes based on the first message. In other words, the step of the second entity identifying the malfunctioning node can be triggered by the first entity, meaning the first entity can flexibly determine whether malfunction detection is necessary based on the collected data, thus avoiding redundant system operation.
[0013] In one possible design, any one of the multiple data sets is a vector structure, enabling the second entity to recognize and process the data from the multiple datasets.
[0014] Optionally, the vector structure includes at least one element, where one element corresponds to one modality of data in at least one modality.
[0015] In one possible design, at least one modality of data includes at least one of the following: log data, metric data, alarm data, stack data, resource networking data, or communication operator status data.
[0016] In one possible design scheme, the types of domains include computing domains, network domains, and storage domains.
[0017] Secondly, an anomaly localization method is provided, applied to a second entity. The second entity can be a terminal device, a chip in the terminal device, or a device for implementing the functions of the terminal device. The second entity can also be a network device, a chip in the network device, or a device for implementing the functions of the network device. The method includes: receiving first information, the first information indicating the correspondence between multiple data and multiple nodes, the multiple data being generated by the operation of the multiple nodes, at least one of the multiple data including at least one modality of data, and at least two of the multiple nodes being used to perform services with different functions.
[0018] Send a second message, which instructs a first node to be identified as an abnormal node among the plurality of nodes based on the first message.
[0019] In one possible design, at least two of the multiple nodes each belong to at least two different types of domains, and the type of the domain is related to the business that the domain performs.
[0020] In one possible design scheme, the second aspect of the method further includes: obtaining the prediction data of the second node based on the first information, wherein the second node is any one of multiple nodes; if the difference between the prediction data and the data corresponding to the second node is greater than a threshold, the third node is determined as the second node.
[0021] Optionally, obtaining the prediction data of the second node based on the first information includes: obtaining the data corresponding to the other nodes among the multiple nodes excluding the second node based on the first information; and obtaining the prediction data of the second node based on the data corresponding to the other nodes.
[0022] In one possible design, the second information also indicates the cause of the anomaly in the first node and / or suggestions for its repair.
[0023] In one possible design, any one of the multiple data points is a vector structure.
[0024] Optionally, the vector structure includes at least one element, where one element corresponds to one modality of data in at least one modality.
[0025] In one possible design, at least one modality of data includes at least one of the following: log data, metric data, alarm data, stack data, resource networking data, or communication operator status data.
[0026] In one possible design scheme, the types of domains include computing domains, network domains, and storage domains.
[0027] It is understandable that the technical effects of the method described in the second aspect can also refer to the relevant introduction of the method described in the first aspect above, and will not be repeated here.
[0028] Thirdly, an anomaly location system is provided. The system includes a first entity and a second entity. The first entity can be a terminal device, a chip in the terminal device, or a device for implementing the functions of the terminal device. The first entity can also be a network device, a chip in the network device, or a device for implementing the functions of the network device. The second entity can also be a terminal device, a chip in the terminal device, or a device for implementing the functions of the terminal device. The second entity can also be a network device, a chip in the network device, or a device for implementing the functions of the network device. Specifically, the first entity sends a first message to the second entity, indicating the correspondence between multiple data points and multiple nodes. The multiple data points are generated by the operation of the multiple nodes, at least one of the multiple data points includes data of multiple modalities, and at least two of the multiple nodes are used to perform services with different functions. The second entity receives the first message and sends a second message to the first entity, indicating a first node, which is a node with an abnormal operation determined among the multiple nodes based on the first message.
[0029] In one possible design, at least two of the multiple nodes each belong to at least two different types of domains, and the type of the domain is related to the business that the domain performs.
[0030] Optionally, the first entity is further configured to identify multiple nodes in the topology information of at least two domains, identify multiple data in the operational data of at least two domains, and determine first information based on the multiple data and multiple nodes.
[0031] In one possible design, the first entity is also used to receive third information, which indicates a business anomaly; and to send the first information, including: sending the first information based on the third information.
[0032] Optionally, the third information includes the identifier of the abnormal service, and the first entity is further used to determine multiple nodes in the topology information of at least two domains and multiple data in the running data of at least two domains based on the identifier of the abnormal service; and to determine the first information based on the multiple data and the multiple nodes.
[0033] In one possible design, the first entity is also used to send a fourth message, which instructs the first entity to identify the malfunctioning node among multiple nodes based on the first message.
[0034] In one possible design, the second entity is also used to obtain the predicted data of the second node based on the first information, where the second node is any one of multiple nodes; if the difference between the predicted data and the data corresponding to the second node is greater than a threshold, the third node is determined as the second node.
[0035] Optionally, the second entity is further configured to obtain data corresponding to other nodes among the multiple nodes, excluding the second node, based on the first information; and to obtain the predicted data of the second node based on the data corresponding to the other nodes.
[0036] In one possible design, the second information also indicates the cause of the anomaly in the first node and / or suggestions for its repair.
[0037] In one possible design, any one of the multiple data points is a vector structure.
[0038] Optionally, the vector structure includes at least one element, where one element corresponds to one modality of data in at least one modality.
[0039] In one possible design, at least one modality of data includes at least one of the following: log data, metric data, alarm data, stack data, resource networking data, or communication operator status data.
[0040] In one possible design scheme, the types of domains include computing domains, network domains, and storage domains.
[0041] It is understood that the technical effects of the system described in the third aspect can also refer to the relevant descriptions of the methods described in any of the first or second aspects above, and will not be repeated here.
[0042] Fourthly, a communication device is provided, the communication device including a module for performing the method described in any one of the first to second aspects.
[0043] In one possible design, the communication device described in the fourth aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the communication device described in the fourth aspect and other communication devices.
[0044] In one possible design, the communication device described in the fourth aspect may further include a memory. This memory may be integrated with the processor or disposed separately. The memory may be used to store instructions relating to the methods described in any of the first to second aspects.
[0045] In the embodiments of this application, the communication device described in the fourth aspect may be a network device, or a chip (system) or other component or assembly disposed in the network device, or a device containing the network device.
[0046] It is understood that the technical effects of the device described in the fourth aspect can also be referred to the relevant descriptions of the methods in any of the first to second aspects above, and will not be repeated here.
[0047] Fifthly, a communication device is provided. The communication device includes a processor coupled to a memory, the processor being configured to execute instructions stored in the memory such that the communication device performs the method described in any one of the first to second aspects.
[0048] In one possible design, the communication device described in the fifth aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the communication device described in the fifth aspect and other communication devices.
[0049] In the embodiments of this application, the communication device described in the fifth aspect may be a network device described in any one of the first to second aspects, or a chip (system) or other component or assembly disposed in the network device, or a device containing the network device.
[0050] Furthermore, the technical effects of the communication device described in the fifth aspect can be referred to the technical effects of the method described in any one of the first or second aspects, and will not be repeated here.
[0051] A sixth aspect provides a communication device, comprising: a processor and a memory; the memory being used to store instructions that, when executed by the processor, cause the communication device to perform the method as described in any one of the first to second aspects.
[0052] In one possible design, the communication device described in the sixth aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used by the communication device described in the fourth aspect to communicate with other communication devices.
[0053] In the embodiments of this application, the communication device described in the sixth aspect may be a network device described in any one of the first to second aspects, or a chip (system) or other component or assembly disposed in the network device, or a device containing the network device.
[0054] Furthermore, the technical effects of the communication device described in the sixth aspect can be referred to the technical effects of the method described in any one of the first or second aspects, and will not be repeated here.
[0055] A seventh aspect provides a chip comprising: a controller and an interface circuit, wherein the controller is configured to interact with other devices via the interface circuit to perform the method as described in any one of the first to second aspects.
[0056] Eighthly, a communication system is provided. The communication system includes a first manager for performing the method described in the first aspect, and a second manager for performing the method described in the second aspect.
[0057] A ninth aspect provides a computer-readable storage medium including a computer program or instructions that, when executed, cause the method described in any one of the first to second aspects to be performed.
[0058] A tenth aspect provides a computer program product comprising a computer program or instructions that, when run, cause the method described in any one of the first to second aspects to be performed. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of the architecture of the communication system provided in the embodiments of this application;
[0060] Figure 2 Flowchart of the anomaly localization method provided in the embodiments of this application Figure 1 ;
[0061] Figure 3 Flowchart of the anomaly localization method provided in the embodiments of this application Figure 2 ;
[0062] Figure 4 This is a schematic diagram of the architecture of the anomaly location system provided in the embodiments of this application;
[0063] Figure 5 Schematic diagram of the communication device provided in the embodiments of this application Figure 1 ;
[0064] Figure 6 Schematic diagram of the communication device provided in the embodiments of this application Figure 2 . Detailed Implementation
[0065] The technical solutions of this application embodiment can be applied to various communication systems, such as Wi-Fi systems, vehicle-to-everything (V2X) communication systems, device-to-device (D2D) communication systems, vehicle-to-everything (V2X) communication systems, fourth-generation (4G) mobile communication systems, such as long-term evolution (LTE) systems, worldwide interoperability for microwave access (WiMAX) communication systems, fifth-generation (5G) mobile communication systems, such as new radio (NR) systems, and future communication systems.
[0066] The technical terms used in this application will be explained below.
[0067] 1. Computing, Networking, and Storage: This refers to the integration of three major domains in a multi-card cluster: the computing domain, the network domain, and the storage domain. Each domain consists of multiple accelerator cards. For ease of understanding, the accelerator cards mentioned below will be described using NPUs as an example. The computing domain is responsible for data processing and task execution, such as iterating the training AI model matrix. The network domain is responsible for communication and data transmission, such as communication between NPU machines within each domain and communication between NPU machines in different domains. The storage domain is responsible for storing and accessing data, such as storing the model matrix file generated during the training phase.
[0068] Currently, the intelligent computing center is facing network and storage failure issues, which leads to low availability of AI training clusters. Frequent failures can cause interruptions and freezes in AI model training and inference, greatly increasing the cost of model training and inference. How to ensure stable operation is a key pain point that needs to be addressed.
[0069] In scenarios where computing, network, and storage converge, faults originate from various types of devices and data modalities, making it difficult to pinpoint the root cause and resulting in slow repair responses. For example, existing automated tools based on anomaly rule matching, time-series models, and sequence prediction technologies can only analyze data from a single domain (from a single type of device) and cannot perform root cause analysis of multi-source, multi-modal anomaly data. Multi-source, multi-modal data analysis primarily relies on manual intervention, which is inefficient, inaccurate, and lacks generalization ability.
[0070] To address the aforementioned technical problems, the inventors have proposed the technical solution described in this application. The technical solution will now be described in conjunction with the accompanying drawings.
[0071] This application will present various aspects, embodiments, or features relating to systems that may include multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.
[0072] Furthermore, in the embodiments of this application, words such as "exemplarily" and "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as an "example" in this application should not be construed as being better or more advantageous than other embodiments or designs. Rather, the use of the word "example" is intended to present the concept in a specific manner.
[0073] First, in this application, "for indicating" can include both direct and indirect indication. When describing "information" for indicating A, it can include whether the information directly indicates A or indirectly indicates A, but does not necessarily mean that the information carries A.
[0074] The information indicated by a given piece of information is called the information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as, but not limited to, directly indicating the information to be indicated, such as the information to be indicated itself or its index. It can also be indirectly indicated by indicating other information, where there is a relationship between the other information and the information to be indicated. It can also indicate only a part of the information to be indicated, while the other parts are known or pre-agreed upon. For example, the indication of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing the indication overhead to some extent. At the same time, common parts of various pieces of information can be identified and indicated uniformly to reduce the indication overhead caused by individually indicating the same information.
[0075] Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be repeated here. As described above, for example, when multiple pieces of information of the same type need to be indicated, the indication methods for different pieces of information may differ. In the specific implementation process, the required indication method can be selected according to specific needs. This application embodiment does not limit the selected indication method; therefore, the indication methods involved in this application embodiment should be understood to cover various methods that enable the party to be indicated to obtain the information to be indicated.
[0076] The information to be instructed can be sent as a whole or divided into multiple sub-information messages, and the sending period and / or timing of these sub-information messages can be the same or different. This application does not limit the specific sending method. The sending period and / or timing of these sub-information messages can be predefined, for example, according to a protocol, or configured by the transmitting device by sending configuration information to the receiving device. This configuration information can include, for example, but not limited to, one or a combination of at least two of radio resource control (RRC) signaling, medium access control (MAC) layer signaling, and physical layer signaling. MAC layer signaling includes, for example, a MAC control element (CE); physical (PHY) layer signaling includes, for example, downlink control information (DCI).
[0077] "Sending information" can be understood as one device sending information to another device, or it can also be understood as one logical module within a device sending information to another logical module. For example, "a network device sending information" can be understood as a network device sending information to another device (such as a terminal or other network device), or it can be understood as logical module 1 in the network device sending information to logical module 2 in the network device.
[0078] "Receiving information" can be understood as one device receiving information from another device, or it can be understood as a logical module within a device receiving information from another logical module. For example, "network device receiving information" can be understood as a network device receiving information from another device (such as a terminal or other network device), or it can be understood as logical module 1 in the network device receiving information from logical module 2 in the network device.
[0079] The phrase "sending information to... (e.g., a node)" or the related illustrations in the accompanying drawings can be understood as the destination of the information being a node. This can include sending information directly or indirectly to a node. Similarly, the phrase "receiving information from... (e.g., a node)," "receiving information from... (e.g., a node)," or "receiving information sent by (e.g., a node)," or the related illustrations in the accompanying drawings, can be understood as the source of the information being a node. This can include receiving information directly or indirectly from a node. Information may undergo necessary processing between the source and destination, such as format changes, but the destination can understand the valid information from the source. Similar expressions in this application can be interpreted similarly, and will not be elaborated further here.
[0080] Second, in the embodiments shown below, the first, second, and various numerical designations are merely distinctions for descriptive convenience and are not intended to limit the scope of the embodiments of this application. For example, to distinguish different indication information.
[0081] Third, "pre-defined," "pre-configured," or "pre-specified" can be achieved by pre-saving corresponding codes, tables, or other means of indicating relevant information in a device or entity (e.g., a network entity). This application does not limit the specific implementation method. "Saving" can refer to saving in one or more memories. These memories can be separate installations or integrated into an encoder or decoder, processor, or communication device. Alternatively, some memories can be separately installed, while others are integrated into the decoder, processor, or communication device. The type of memory can be any form of storage medium, and this application does not limit this.
[0082] For ease of understanding, the communication system is described below with reference to the accompanying drawings. Figure 1 This is a schematic diagram of the architecture of a communication system, which includes network devices, terminal devices, and a core network (CN).
[0083] The network devices may include network devices 101a to 101b, and the terminal devices may include terminal devices 102a to 102b. The terminal devices can be connected to the network devices wirelessly, and the network can be connected to the core network 103 via wired or wireless means.
[0084] In this embodiment, the terminal device can be connected to the network device wirelessly, and the network device can be connected to the core network via wired or wireless means. In one possible implementation, the first entity, the second entity, the node, and other devices involved in this application can be terminal devices or network elements in the core network. These devices can communicate with each other, and there are no specific limitations.
[0085] Terminal equipment can be a terminal with transceiver capabilities, or it can be a chip or chip system installed in the terminal equipment. This terminal equipment can also be referred to as User Equipment (UE), Access Terminal, Subscriber Unit, User Station, Mobile Station (MS), Mobile Station, Remote Station, Remote Terminal, Mobile Equipment, User Terminal, Terminal, Wireless Communication Equipment, User Agent, or User Device. The terminal devices in the embodiments of this application may be mobile phones, cellular phones, smartphones, tablets, wireless data cards, personal digital assistants (PDAs), wireless modems, handsets, laptop computers, machine-type communication (MTC) terminals, computers with wireless transceiver capabilities, virtual reality terminals, augmented reality terminals, smart home devices (e.g., refrigerators, televisions, air conditioners, electricity meters, etc.), intelligent robots, robotic arms, workshop equipment, wireless terminals in autonomous driving, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in telemedicine, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, vehicle-mounted terminals, and roadside units with terminal functions. The terminal device in this application can also be an onboard module, onboard unit, onboard component, onboard chip, or onboard unit built into a vehicle as one or more components or units. The terminal device can also be other devices with terminal functions; for example, it can be a device that performs terminal functions in D2D communication. The embodiments of this application do not limit the device form of the terminal device. The device used to implement the terminal function can be a terminal device; it can also be a device that supports the terminal in implementing the function, such as a chip system. This device can be installed in the terminal or used in conjunction with the terminal. In the embodiments of this application, the chip system can be composed of chips or can include chips and other discrete devices.
[0086] Network devices can be devices with wireless transceiver capabilities, or they can be chips or chip systems located in the access network (AN) of a communication system to provide access services to terminals. For example, network devices can be called radio access network (RAN) devices, and can be RAN devices of 5G or future mobile communication systems. In future mobile communication systems, network devices may also have other naming conventions, all of which are covered within the protection scope of the embodiments of this application, and this application does not impose any limitations on them. Alternatively, network equipment can also include 5G, such as a 5G base station (next-generation node B, gNB) in a new radio (NR) system, or one or a group of antenna panels (including multiple antenna panels) of a 5G base station. It can also be network nodes constituting a gNB, transmission and reception point (TRP) or transmission point (TP), or transmission measurement function (TMF), such as a central unit (CU), distributed unit (DU), CU-control plane (CP), CU-user plane (UP), or radio unit (RU), RSU with base station functionality, or wired access gateway, or core network elements of 5G, etc. Alternatively, network equipment can also include: access points (APs) in WiFi systems, wireless relay nodes, wireless backhaul nodes, various forms of macro base stations, micro base stations (also called small cells), relay stations, access points, wearable devices, vehicle-mounted equipment, etc.
[0087] CU and DU can be separate entities or included in the same network element, such as a baseband unit (BBU). RU can be included in radio frequency equipment or radio frequency units, such as remote radio units (RRUs), active antenna units (AAUs), or remote radio heads (RRHs). It is understood that network equipment can be CU nodes, DU nodes, or a combination of CU and DU nodes. Furthermore, CUs can be classified as network equipment in the access network (RAN) or the core network (CN), without limitation. In different systems, CUs (or CU-CPs and CU-UPs), DUs, or RUs may have different names, but their meanings will be understood by those skilled in the art. For example, in an ORAN system, a CU can also be called an O-CU (open CU), a DU can also be called an O-DU, a CU-CP can also be called an O-CU-CP, a CU-UP can also be called an O-CU-UP, and a RU can also be called an O-RU. For ease of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples. Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software and hardware modules. In the embodiments of this application, the form of the network device is not limited; the device used to implement the function of the network device can be the network device itself; it can also be a device capable of supporting the network device in implementing that function, such as a chip system. This device can be installed in the network device or used in conjunction with the network device.
[0088] The network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0089] The anomaly localization method of this application embodiment will be further described below with reference to the accompanying drawings. It is understood that this application uses the first entity and the second entity as examples of the execution subjects in the interaction illustration, but this application does not limit the execution subjects in the interaction illustration. In addition, the processing performed by a single execution subject can also be divided into multiple execution subjects, which can be logically and / or physically separated.
[0090] The interaction process between devices in the above-described communication system will be specifically described below through method embodiments. The anomaly localization method provided in this application embodiment can be applied to the above-described communication system and specifically applied to various scenarios involved in the above-described communication system, which will be described in detail below.
[0091] Figure 2 Flowchart of the anomaly localization method provided in the embodiments of this application Figure 1 This anomaly localization method is applicable to the aforementioned communication system, and is applied to the interaction between the first entity and the second entity, mainly involving the interaction between the first entity and the second entity. The first entity can be a terminal device, a chip within a terminal device, or a device used to implement the functions of the terminal device; the first entity can also be a network device, a chip within a network device, or a device used to implement the functions of the network device. Similarly, the second entity can be a terminal device, a chip within a terminal device, or a device used to implement the functions of the terminal device; the second entity can also be a network device, a chip within a network device, or a device used to implement the functions of the network device, and there are no specific limitations.
[0092] like Figure 2 As shown, the specific flow of this method is as follows:
[0093] S201, the first entity sends a first message to the second entity, and the second entity receives the first message from the first entity. The first message indicates the correspondence between multiple data and multiple nodes.
[0094] The first piece of information can be signaling or information elements within signaling. The signaling can reuse existing signaling to reduce implementation difficulty, or it can be newly defined signaling, decoupled from existing signaling to make the transmission of the first piece of information more flexible. The specific implementation is not limited. It should be noted that the correspondence between multiple data points and multiple nodes can be stored in a table or configuration file within a single piece of first information, and carried and sent to the second entity by only one piece of first information.
[0095] At least two distinct nodes among multiple nodes are used to perform different functionalities, which may include AI model training, AI model inference, computation, (network) communication, storage, etc. For example, a node can be a device within a domain (such as an NPU). The type of domain is related to the functionalities performed by that domain. For instance, the type of domain may include a computation domain and / or a network domain and / or a storage domain in a computation-network-storage system. That is, at least two nodes among multiple nodes may each come from at least two of the aforementioned domains. For ease of explanation, the computation domain, network domain, and storage domain are collectively referred to as the three domains below, and multiple nodes are described as nodes within these three domains, but other possible scenarios, such as multiple nodes coming from only two of the aforementioned domains, are not limited. It is understood that the nodes among multiple nodes can be classified and named according to their domain affiliation, such as computation nodes, network nodes, and storage nodes. Accordingly, a computation node refers to a node in the computation domain, a network node to a node in the network domain, and a storage node to a node in the storage domain. Furthermore, a computation domain may contain multiple computation nodes, a network domain may contain multiple network nodes, and a storage domain may contain multiple storage nodes.
[0096] Multiple data refers to the operational data generated by multiple nodes during operation. For example, suppose there are multiple nodes including node A, node B, and node C. Each node runs and generates corresponding operational data A, operational data B, and operational data C during operation. Then, operational data A, operational data B, and operational data C are also data in multiple data.
[0097] For example, at least one of the multiple data includes data of multiple modalities. Taking the runtime data A included in the multiple data as an example, runtime data A may include data of at least one of the following modalities: log data, indicator data, alarm data, stack data, resource networking data, or communication operator status data. Log data is used to describe information such as events, errors, or warnings during node operation; indicator data is used to describe node status; alarm data is used to trigger notifications of node operation abnormalities; stack data is used to describe the call stack information when the node executes the program; resource networking data is used to describe node resources and their connection relationships; and communication operator status data is used to describe the status information of various components (such as network interfaces and protocol stacks) in the domain or system where the node is located. Moreover, the data can be represented in different data modalities such as text data, structured data, or binary data. That is, runtime data A may include data of multiple modalities.
[0098] In one possible implementation, any one of the multiple data sets is a data structure that the second entity can recognize. For example, it can be a vector structure, which includes at least one element, and one element of the at least one element corresponds to one modality of data among multiple modalities. Taking the example of multiple data sets including runtime data A, and runtime data A including log data, metric data, and alarm data of several modalities, runtime data A = {a1, a2, a3}, where a1 is log data, a2 is metric data, a3 is alarm data, and so on.
[0099] For example, the first information indicates the correspondence between multiple data and multiple nodes, that is, the relationship between running data A and node A, running data B and node B, running data C and node C, and so on, which will not be elaborated further.
[0100] The above details the meaning of multiple data points, multiple nodes, and the correspondence between the multiple data points and multiple nodes. The following further explains how the first entity determines this information, that is, how the first information is determined.
[0101] In one possible implementation, multiple nodes can be all nodes in the three domains, or some nodes in the three domains, such as multiple nodes being some nodes in the three domains related to abnormal business; multiple data can also be all operational data related to all nodes in the three domains, or some operational data related to some nodes in the three domains, such as multiple data being data generated by the operation of some nodes related to abnormal business. The following describes the method by which the first entity determines the first information.
[0102] Case 1: Multiple nodes are all nodes in the three domains, and multiple data are all operational data related to the three domains.
[0103] For example, a first entity can periodically receive first messages from the computing domain, network domain, and storage domain. These first messages include topology information corresponding to the three domains and corresponding operational data. The period for receiving the first messages can be flexibly selected; for example, a shorter receiving period can be set when business stability requirements are high, while a longer receiving period can be set when the entity's (or system's) computing power decreases. There are no specific limitations. Alternatively, the first entity can periodically send second messages to the computing domain, network domain, and storage domain. These second messages request the topology information and operational data corresponding to the three domains. The first entity then receives the aforementioned first messages again. There are no specific limitations. The topology information corresponding to the three domains indicates all nodes within the three domains and their relationships, such as the layout and connections between nodes. The operational data corresponding to the three domains indicates all data generated by all nodes within the three domains during operation. Thus, the first entity can obtain all nodes in the three domains as multiple nodes and all operational data related to the three domains as multiple data sets, to further determine the correspondence between the multiple nodes and the multiple data sets.
[0104] Case 2: Multiple nodes are some of the nodes in the three domains, and multiple data are some of the operational data related to some of the nodes in the three domains.
[0105] For example, the first entity can determine multiple nodes related to the abnormal service from the topology information corresponding to the three domains, and multiple data related to the abnormal service from the operational data corresponding to the three domains, based on the information indicating the abnormal service. For instance, the first entity can receive third information indicating an abnormal service operation, determine the first information based on the third information, and send the first information. The third information can be sent by at least one of the computing domain, network domain, or storage domain, or by system maintenance personnel; there are no specific limitations.
[0106] An abnormal service refers to a service that freezes, is interrupted, or degrades during operation. In one possible implementation, the third information may include an identifier for the abnormal service, such as its ID or index. Therefore, the first entity can determine a subset of nodes from all nodes within the three domains that execute the abnormal service, treating them as multiple nodes, and obtain the operational data generated by these nodes during the execution of the abnormal service, treating them as multiple data sets. This further determines the correspondence between the multiple nodes and the multiple data sets, reducing the data processing load on the first entity and improving the efficiency of subsequent anomaly localization. The method by which the first entity obtains the topology information and operational data corresponding to the three domains can be referred to the explanation in Case 1 above, and will not be repeated here.
[0107] In one possible implementation, after the first entity acquires multiple data points, it can preprocess the data points based on a small model and built-in algorithms, and then fuse and embed them into a data structure that the second entity can recognize, such as the vector structure described above. For specific implementation details, please refer to existing technologies, which will not be elaborated further.
[0108] In one possible implementation, the first entity can further filter out key abnormal data from multiple data sets based on information such as log level, keywords, or pre-defined whitelists, update the key abnormal data sets with new data sets, determine the correspondence between the updated data sets and nodes, and send this information as the first piece of information to the second entity to reduce data transmission pressure.
[0109] In some possible implementations, the first information may also include information from multiple nodes and / or multiple data points, so that the second entity can obtain the required information from the first entity, without limitation.
[0110] S202, the second entity determines the first node based on the first information.
[0111] The first node is the node that is experiencing a malfunction among multiple nodes.
[0112] In one possible implementation, the second entity can make predictions for multiple nodes one by one based on the first information, obtaining the predicted data corresponding to each node. The predicted data for a node can refer to the data generated by the node assuming it is operating normally. Therefore, the second entity can compare the predicted data of a node with the data corresponding to the actual operation of that node from among the multiple data sets to determine whether the node is operating normally. It can be understood that if the predicted data of a node is the same as the actual operating data of that node, or the difference is small (e.g., less than a certain set threshold), it can be said that the node is operating normally or the deviation is within an acceptable error range. If the difference between the predicted data and the actual operating data of a node is large (e.g., greater than a certain set threshold), then the node is considered an abnormal operating node, and that node is identified as the first node.
[0113] The following example, using one of multiple nodes as the second node, illustrates how to determine the first node. Specifically, the second entity obtains the predicted data corresponding to the second node based on the first information in the following ways: The second entity determines the data corresponding to other nodes besides the second node based on the first information. For example, if the first information includes the correspondence between multiple data points and multiple nodes, then, in the case of determining the second node, other nodes can be determined among the multiple nodes, and the data corresponding to other nodes can be determined among the multiple data points based on the correspondence. Alternatively, if the first information only includes the correspondence between multiple data points and multiple nodes, the second entity can also obtain the information of multiple data points and / or multiple nodes from the three domains. The specific method can refer to the method in S201 where the first entity obtains multiple data points and / or multiple nodes, and then determine other nodes and the data corresponding to other nodes. Then, the second entity infers the predicted data of the second node based on the data corresponding to other nodes. This process can refer to existing technologies and will not be elaborated further.
[0114] In one possible implementation, the second entity may also determine the cause of the anomaly and / or repair suggestions corresponding to the first node based on the first node, and instruct the first entity accordingly.
[0115] In one possible implementation, the second entity can be triggered to perform the action of "determining the first node". For example, if the correspondence between the multiple data and multiple nodes indicated by the first information refers to some nodes in the three domains and the multiple data refers to some running data related to some nodes in the three domains (i.e., case 1 described in S201), then the second entity can be pre-programmed to perform "determining the first node" upon receiving the first information; or, if the correspondence between the multiple data and multiple nodes indicated by the first information refers to all nodes in the three domains and the multiple data refers to all running data related to some nodes in the three domains (i.e., case 2 described in S201), then the second entity can be instructed to determine the first node through the indication information.
[0116] For example, the method may also include:
[0117] Step A: The first entity sends the fourth message, and the second entity receives the fourth message from the first entity. The fourth message instructs the second entity to determine the abnormal node among multiple nodes based on the first message.
[0118] It is understandable that the fourth information can be a separately sent message or it can be carried together with the first information in a message. For example, the fourth information can be a binary number in a message carrying the first information. When the binary number is 0, it indicates that the second entity is unsure of the abnormal node. When the binary number is 1, it indicates that the second entity is sure of the abnormal node. There are no specific restrictions.
[0119] S203, the second entity sends second information to the first entity, the first entity receives the second information from the second entity, and the second information instructs the first node.
[0120] The second information may also include the cause of the anomaly and / or repair suggestions corresponding to the first node, so that the first entity can return the abnormal node and its corresponding cause of the anomaly and / or repair suggestions to the system operation and maintenance personnel to perform subsequent repair operations. The second information can be signaling or information elements in signaling. The signaling can reuse existing signaling to reduce the implementation difficulty, or it can be newly defined signaling, decoupled from the existing signaling, to make the transmission of the second information more flexible. There are no restrictions on the specific implementation.
[0121] In one possible implementation, the second entity can also obtain past node anomaly descriptions and anomaly repair experience, such as input by operations and maintenance personnel into the second entity or loaded from the cloud. During the process of the second entity determining the first node (i.e., performing anomaly localization), iteratively optimizes the anomaly causes and repair suggestions corresponding to the anomaly node by combining past node anomaly descriptions and anomaly repair experience, thereby improving the repair efficiency of abnormal operation.
[0122] In summary, the first entity can send a first message to the second entity, indicating the correspondence between multiple data points and multiple nodes. Then, the second entity receives the first message and sends a second message to the first entity, indicating the abnormal node identified among the multiple nodes based on the first message. These multiple nodes can perform at least two functional operations, such as nodes from the computing domain, network domain, and storage domain. Furthermore, the multiple data generated by these nodes include data of at least one modality. In other words, the anomaly localization method provided by this solution can perform unified mapping of abnormal data across multiple domains based on the node network topology, constructing an anomaly localization model that includes both the first and second entities. This enables anomaly analysis of multi-domain, multi-modal operational data and accurate localization of abnormal nodes, improving the stability of complex intelligent computing clusters.
[0123] The above combination Figure 2 The overall process of the anomaly localization method provided in the embodiments of this application is introduced below. Figure 3 This paper describes the process of the anomaly localization method provided in the embodiments of this application in a specific scenario.
[0124] Figure 3 Flowchart of the anomaly localization method provided in the embodiments of this application Figure 2This anomaly localization method is applicable to the aforementioned communication system and mainly involves the interaction between a first entity (such as an anomaly analysis module) and a second entity (such as an AI model module). For example, the first entity can send a first message to the second entity, indicating the correspondence between multiple data points and multiple nodes. Then, the second entity receives the first message and sends a second message to the first entity, indicating the abnormal node identified among the multiple nodes based on the first message. The multiple nodes can be nodes from the computing domain, network domain, and storage domain, and the multiple data generated by the operation of the multiple nodes include data of at least one modality. In other words, the first entity can perform unified mapping based on the node network topology for abnormal data in multiple domains to obtain the correspondence, and the second entity can combine the correspondence to perform anomaly analysis on the multi-domain, multi-modal operating data, accurately locating the abnormal node so that maintenance personnel can repair it, thereby improving the stability of complex intelligent computing clusters.
[0125] Specifically, such as Figure 3 As shown, the flow of this anomaly localization method is as follows:
[0126] S301, the anomaly analysis module sends data collection request messages to the computing domain, network domain, and storage domain.
[0127] The data collection request message can be the second message described in S201, which is only one possible name, or it can be any message that can achieve this function, without any specific limitation.
[0128] S302, the anomaly analysis module receives data sent from the computing domain, network domain, and storage domain.
[0129] The data sent by the computing domain, network domain, and storage domain may include the topology information and the operational data corresponding to the domain. For details, please refer to the relevant description in Case 1 of S201, which will not be repeated here.
[0130] S303, the anomaly analysis module receives alarm messages from the computing domain, network domain, and storage domain.
[0131] Alarm messages are used to indicate that at least one of the three domains has an operational anomaly. For example, it can be the second information described in S201, which is only one possible name. It can also be any message that can achieve this function, without any specific limitation.
[0132] It is understandable that steps S301-S303 are optional.
[0133] S304, the anomaly analysis module determines the first piece of information.
[0134] The anomaly analysis module can determine multiple data points and nodes, as well as their respective correspondences, from data sent in the computing, network, and storage domains based on the anomaly service ID, using this as primary information. Furthermore, the anomaly analysis module can filter key anomaly data from multiple data points based on log level, keywords, or pre-defined whitelists, updating the data and corresponding primary information accordingly. For details, please refer to the relevant description in S201; further elaboration is omitted here.
[0135] S305, the anomaly analysis module sends the first information to the AI model module, and the AI model module receives the first information from the anomaly analysis module.
[0136] For details, please refer to the relevant description in S201, which will not be repeated here.
[0137] S306, the AI model module determines the second information based on the first information, and the second information indicates the first node.
[0138] The first node is an abnormal node in the compute domain, network domain, and / or storage domain. For details, please refer to the relevant description in S202; further details will not be provided here.
[0139] S307, the AI model module sends the second information to the anomaly analysis module, and the anomaly analysis module receives the second information from the AI model module.
[0140] For details, please refer to the relevant description in S203, which will not be repeated here.
[0141] The above combination Figures 2-3 The anomaly localization method provided in the embodiments of this application is described in detail below. Figure 4 This document describes in detail an anomaly localization system for executing the anomaly localization method provided in the embodiments of this application. For example, as shown... Figure 4 As shown, the anomaly localization system 400 includes a first entity 401 and a second entity 402. The anomaly localization system 400 can be applied to the above-mentioned... Figures 2-3 The anomaly location method is used to achieve the corresponding functions. In addition, the technical effect of the anomaly location system 400 can be referred to the above-mentioned anomaly location method, which will not be repeated here.
[0142] The following will continue to combine Figures 5-6 This document describes in detail the communication apparatus used to execute the anomaly location method provided in the embodiments of this application.
[0143] Figure 5 This is a schematic diagram of the structure of the communication device provided in the embodiments of this application. Figure 1 For example, such as Figure 5 As shown, the communication device 500 includes a transceiver module 501 and a processing module 502. For ease of explanation, Figure 5Only the main components of the communication device are shown.
[0144] The communication device 500 can be applied to the above. Figures 2-3 The anomaly localization method is used to implement the corresponding functions. For example, the transceiver module 501 can be used to implement the above. Figures 2-3 The sending and receiving functions in the anomaly localization method can be implemented by the processing module 502. Figures 2-3 The exception location method includes functions other than sending and receiving.
[0145] Optionally, the transceiver module 501 may include a transmitting module ( Figure 5 (not shown in the image) and receiving module ( Figure 5 (Not shown in the diagram). The transmitting module implements the transmitting function of the communication device 500, and the receiving module implements the receiving function of the communication device 500.
[0146] Optionally, the communication device 500 may also include a storage module. Figure 5 (Not shown in the image), the storage module stores programs or instructions. When the processing module 502 executes the program or instructions, the communication device 500 can perform the aforementioned operations. Figures 2-3 The functions in the method shown.
[0147] It is understood that the communication device 500 may be a network device, or a chip (system) or other component or assembly that can be set in the network device, or a device that includes the network device. This application does not limit this.
[0148] Furthermore, the technical effects of the communication device 500 can be referenced from the technical effects of the above-mentioned anomaly location method, and will not be repeated here.
[0149] Figure 6 Schematic diagram of the communication device provided in the embodiments of this application Figure 2 For example, the communication device can be a terminal, or a chip (system) or other component or assembly that can be set in the terminal. Figure 6 As shown, the communication device 600 may include a processor 601. Optionally, the communication device 600 may also include a memory 602 and / or a transceiver 603. The processor 601 is coupled to the memory 602 and the transceiver 603, for example, they can be connected via a communication bus.
[0150] The following is combined with Figure 6 A detailed description of each component of the communication device 600 is provided below:
[0151] The processor 601 is the control center of the communication device 600. It can be a single processor or a collective term for multiple processing elements. For example, the processor 601 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0152] Optionally, the processor 601 can perform various functions of the communication device 600 by running or executing software programs stored in the memory 602 and calling data stored in the memory 602, such as performing the aforementioned functions. Figures 2-3 The anomaly localization method is shown.
[0153] In a specific implementation, as one example, processor 601 may include one or more CPUs, for example... Figure 6 CPU0 and CPU1 are shown in the diagram.
[0154] In a specific implementation, as one example, the communication device 600 may also include multiple processors, for example... Figure 6 The processors 601 and 604 are shown. Each of these processors can be a single-core processor or a multi-core processor. Here, "processor" can refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0155] The memory 602 is used to store the software program that executes the solution of this application, and is controlled by the processor 901 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0156] Optionally, the memory 602 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 602 may be integrated with the processor 601 or may exist independently and be connected via the interface circuit of the communication device 600. Figure 6 (Not shown in the image) is coupled to the processor 601, but this embodiment does not specifically limit this.
[0157] Transceiver 603 is used for communication with other communication devices. For example, if communication device 600 is a terminal, transceiver 603 can be used to communicate with a network device or with another terminal device. As another example, if communication device 600 is a network device, transceiver 603 can be used to communicate with a terminal or with another network device.
[0158] Alternatively, transceiver 603 may include a receiver and a transmitter. Figure 6 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0159] Optionally, the transceiver 603 can be integrated with the processor 601, or it can exist independently and be connected via the interface circuit of the communication device 600. Figure 6 (Not shown in the image) is coupled to the processor 601, but this embodiment does not specifically limit this.
[0160] Understandable, Figure 6 The structure of the communication device 600 shown does not constitute a limitation on the communication device. Actual communication devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0161] Furthermore, the technical effects of the communication device 600 can be referred to the technical effects of the methods described in the above method embodiments, and will not be repeated here.
[0162] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0163] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0164] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0165] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0166] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0167] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0168] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0169] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0170] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0171] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0172] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0173] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0174] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0175] In this application, descriptions such as "when," "under the circumstances," "if," and "if" all refer to the device taking corresponding actions under certain objective circumstances. They are not time limits, nor do they require the device to perform a judgment action during implementation, nor do they imply any other limitations.
[0176] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0177] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0178] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0179] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0180] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0181] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0182] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An anomaly localization method, characterized in that, The method includes: Send a first message, which indicates the correspondence between multiple data and multiple nodes. The multiple data are generated by the operation of the multiple nodes. At least one of the multiple data includes data of at least one modality. At least two of the multiple nodes are used to perform services with different functions. Receive second information, the second information indicating a first node, the first node being a node with abnormal operation determined among the plurality of nodes based on the first information.
2. The method according to claim 1, characterized in that, At least two of the plurality of nodes each belong to at least two different types of domains, the type of which is related to the business performed by the domain.
3. The method according to claim 2, characterized in that, Before sending the first message, the method further includes: The plurality of nodes are determined from the topology information of the at least two domains, and the plurality of data are determined from the operational data of the at least two domains; The first information is determined based on the multiple data and the multiple nodes.
4. The method according to claim 2 or 3, characterized in that, The method further includes: Receive third information, which is used to indicate a service anomaly; Sending a first message includes: Based on the third information, send one of the first messages.
5. The method according to claim 4, characterized in that, The third piece of information includes an identifier for abnormal service, and the method further includes: Based on the identifier of the abnormal service, determine the plurality of nodes in the topology information of the at least two domains, and determine the plurality of data in the running data of the at least two domains; The first information is determined based on the multiple data and the multiple nodes.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: A fourth message is sent, which instructs the second entity to determine the abnormal node among the plurality of nodes based on the first message.
7. An anomaly localization method, characterized in that, The method includes: Receive a first message, the first message indicating the correspondence between multiple data and multiple nodes, the multiple data being generated by the operation of the multiple nodes, at least one of the multiple data including at least one modality of data, and at least two of the multiple nodes being used to perform different functional services; Send a second message, which instructs a first node to be identified as an abnormal node among the plurality of nodes based on the first message.
8. The method according to claim 7, characterized in that, At least two of the plurality of nodes each belong to at least two different types of domains, the type of which is related to the business performed by the domain.
9. The method according to claim 7 or 8, characterized in that, The method further includes: Based on the first information, the predicted data of the second node is obtained, where the second node is any one of the plurality of nodes; If the difference between the predicted data and the data corresponding to the second node is greater than a threshold, the second node is determined as the first node.
10. The method according to claim 9, characterized in that, The step of obtaining the prediction data for the second node based on the first information includes: Based on the first information, obtain the data corresponding to the other nodes among the plurality of nodes, excluding the second node; Based on the data corresponding to the other nodes, obtain the prediction data for the second node.
11. The method according to any one of claims 7-10, characterized in that, The second information also indicates the cause of the anomaly in the first node and / or suggestions for repairing the first node.
12. The method according to any one of claims 1-6 or 7-11, characterized in that, Any one of the multiple data sets is a vector structure.
13. The method according to claim 12, characterized in that, The vector structure includes at least one element, and one element of the at least one element corresponds to data of one modality in the at least one modality of data.
14. The method according to any one of claims 1-6, 7-11, or 12-13, characterized in that, The data of at least one modality includes at least one of the following: log data, indicator data, alarm data, stack data, resource networking data, or communication operator status data.
15. The method according to any one of claims 2-6 or claim 8, characterized in that, The types of domains include computing domains, network domains, and storage domains.
16. An anomaly location system, characterized in that, The system includes a first entity and a second entity; The first entity is used to send a first message to the second entity. The first message indicates the correspondence between multiple data and multiple nodes. The multiple data are generated by the operation of the multiple nodes. At least one of the multiple data includes data of multiple modalities. At least two of the multiple nodes are used to perform services with different functions. The second entity is used to receive the first information and send second information to the first entity. The second information indicates the first node, which is a node with abnormal operation determined among the plurality of nodes based on the first information.
17. A communication device, characterized in that, The apparatus includes an entity for performing the method as described in any one of claims 1-16.
18. A communication device, characterized in that, The communication device includes a processor and a memory; the memory is used to store computer instructions, which, when executed by the processor, cause the communication device to perform the method as described in any one of claims 1-16.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instructions that, when executed, cause the method as described in any one of claims 1-16 to be performed.
20. A computer program product, characterized in that, Includes a computer program or instructions that, when executed, cause the method as described in any one of claims 1-16 to be performed.