Data circuit fault analysis method and device and medium
By generating multi-layer bearing details and using large language models to locate data circuit faults, the problems of slow fault positioning speed and low accuracy in the existing technology are solved, fast and accurate fault diagnosis is achieved, and network stability and user satisfaction are improved.
Patent Information
- Application Number
- CN202510344746.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the data circuit fault location speed is slow and the accuracy is low, which affects the efficiency of communication services and user satisfaction.
By obtaining the basic information and query requirements of the business, generating multi-layer bearer details, sending query instructions in layers, using large language models and SQL statements to locate fault levels and locations, and providing repair suggestions in combination with historical fault knowledge bases.
It realizes fast and accurate data circuit fault location, improves fault diagnosis efficiency and user satisfaction, reduces operating costs, and enhances network stability and reliability.
Smart Images

Figure CN120277121A_ABST
Abstract
Description
Technical Field
[0001] This application is at least related to the field of communication technologies, and particularly relates to a method, apparatus, and medium for analyzing data circuit failures. Background Art
[0002] With the rapid development of information technology, the communication industry has become an important pillar of modern society. Currently, the competition in the communication industry is becoming increasingly fierce, and users' requirements for communication services are also getting higher and higher. During the process of providing communication services, the occurrence of data circuit failures will bring great inconvenience to users' normal use of communication services.
[0003] Therefore, how to quickly and accurately locate data circuit failures and improve service quality and user satisfaction has become an urgent problem to be solved. Summary of the Invention
[0004] In view of the above deficiencies, this application provides a method, apparatus, and medium for analyzing data circuit failures to solve the following technical problems: how to quickly and accurately locate data circuit failures and improve service quality and user satisfaction.
[0005] In a first aspect, this application provides a method for analyzing data circuit failures, the method including:
[0006] Obtain the basic information and basic query requirements of the service to be queried;
[0007] Obtain the multi-layer bearing details of the service to be queried according to the basic information;
[0008] Generate multi-layer query instructions for the service to be queried according to the multi-layer bearing details and the basic query requirements;
[0009] Sequentially send the multi-layer query instructions layer by layer to obtain hierarchical query results, and locate the level and location where the service to be queried fails currently or in the future according to the hierarchical query results.
[0010] Further, obtaining the basic information and basic query requirements of the service to be queried specifically includes:
[0011] In response to receiving customer trouble reporting information, obtain the service involved in the customer trouble reporting information as the first service to be queried, and use a large language model to conduct a conversational Q&A with the customer to obtain the first basic information and the first basic query requirements of the first service to be queried. The first basic information includes customer identity information and the type of the customer trouble reporting service, and the first basic query requirements include the type of the customer trouble reporting service failure; or,
[0012] In response to generating a fault prediction work order based on the historical fault knowledge base, the business involved in the fault prediction work order is obtained as the second business to be queried. Fuzzy search is used and interacted with the user to obtain the second basic information and the second basic query requirements of the second business to be queried. The second basic information includes the fault prediction scope, and the second basic query requirements include the fault prediction indicators.
[0013] Furthermore, according to the basic information, the multi-layer bearing details of the business to be queried are obtained, specifically including:
[0014] Generate the first Structured Query Language (SQL) statement according to the first basic information, and use the first SQL statement to obtain the first multi-layer bearing details of the first business to be queried from the operator's business ledger system. The first multi-layer bearing details include the Service Router (SR) address or access switch address of the first business to be queried, as well as the core layer bearing device port, aggregation layer bearing device port, access layer bearing device port, and client bearing device address; or,
[0015] Generate the second Structured Query Language (SQL) statement according to the second basic information, and use the second SQL statement to obtain the second multi-layer bearing details of the second business to be queried from the operator's business ledger system. The second multi-layer bearing details include the core layer bearing device address, aggregation layer bearing device address, and access layer bearing device address of the second business to be queried.
[0016] Furthermore, according to the multi-layer bearing details and the basic query requirements, multi-layer query instructions for the business to be queried are generated, specifically including:
[0017] According to the first multi-layer bearing details and the first basic query requirements, match the preset first-layer query prompt words for each layer. The first-layer query prompt words for each layer include Virtual Local Area Network (VLAN) configuration check of the SR or access switch, Media Access Control (MAC) address table check, and Address Resolution Protocol (ARP) table entry check, as well as Ping test, port status query, traffic query, and Cyclic Redundancy Check (CRC) error query for the core layer, aggregation layer, and access layer, and Ping test for the client.
[0018] Obtain the first instruction template from the preset instruction library according to the first-layer query prompt words for each layer, and fill in the SR address or access switch address of the first business to be queried, as well as the core layer bearing device port, aggregation layer bearing device port, access layer bearing device port, and client bearing device address into the corresponding first instruction template to generate the first multi-layer query instruction for the first business to be queried; or,
[0019] According to the second multi-layer bearing details and the second basic query requirements, match the preset second-layer query prompt words for each layer. The second-layer query prompt words for each layer include fault prediction indicator query for the core layer, aggregation layer, and access layer.
[0020] Obtain the second instruction template from the preset instruction library according to the query prompt words of the second layer, and fill in the core layer bearing device address, aggregation layer bearing device address, and access layer bearing device address of the second service to be queried into the corresponding second instruction template to generate the second multi-layer query instruction for the second service to be queried.
[0021] Furthermore, send multi-layer query instructions layer by layer to obtain the hierarchical query results, and locate the level and location where the service to be queried currently has a fault according to the hierarchical query results. Specifically, it includes:
[0022] Send a VLAN configuration check instruction, a MAC address table check instruction, and an ARP table entry check instruction to the SR or access switch. If there is any error in the VLAN configuration, MAC address, or ARP table entry, locate the location where the first service to be queried currently has a fault as the SR or access switch configuration.
[0023] If the VLAN configuration, MAC address, and ARP table entries are all correct, send a core layer Ping test instruction to the core layer bearing device port, and judge whether the core layer bearing link is normal according to the core layer Ping test result. If not, send a core layer port status query instruction, a core layer traffic query instruction, and a core layer CRC error query instruction to judge the location where the first service to be queried currently has a fault in the core layer.
[0024] If the core layer bearing link is normal, send an aggregation layer Ping test instruction to the aggregation layer bearing device port, and judge whether the aggregation layer bearing link is normal according to the aggregation layer Ping test result. If not, send an aggregation layer port status query instruction, an aggregation layer traffic query instruction, and an aggregation layer CRC error query instruction to judge the location where the first service to be queried currently has a fault in the aggregation layer.
[0025] If the aggregation layer bearing link is normal, send an access layer Ping test instruction to the access layer bearing device port, and judge whether the access layer bearing link is normal according to the access layer Ping test result. If not, send an access layer port status query instruction, an access layer traffic query instruction, and an access layer CRC error query instruction to judge the location where the first service to be queried currently has a fault in the access layer.
[0026] If the access layer bearing link is normal, send an access layer Ping test instruction to the client bearing device, and judge whether the client bearing link is normal according to the client Ping test result. If not, locate the location where the first service to be queried currently has a fault as a client hardware fault. If so, locate the location where the first service to be queried currently has a fault as a client software fault.
[0027] Further, send multi-layer query instructions layer by layer to obtain multi-layer query results, and locate the level and location where the service to be queried will have a future failure according to the multi-layer query results, specifically including:
[0028] Send the core layer fault prediction index query instruction, the aggregation layer fault prediction index query instruction, and the access layer fault prediction index query instruction to the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device in sequence, and obtain the current status data of the fault prediction indexes of the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device;
[0029] According to the current status data of the fault prediction indexes of the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device, respectively predict the future status data of the fault prediction indexes of the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device within a preset future time period. The preset future time period is the time period during which a fault may occur in the future within the fault prediction range obtained from the historical fault knowledge base;
[0030] Obtain the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device corresponding to the future status data of the fault prediction index lower than the preset value, so as to locate the level and location where the second service to be queried will have a future failure.
[0031] Further, the method further includes:
[0032] According to the fault repair experience in the historical fault knowledge base, obtain the repair suggestions for the level and location where the fault occurs;
[0033] Send the level and location where the fault occurs and the repair suggestions to the data circuit fault repair personnel at the corresponding level.
[0034] Further, where:
[0035] The specific form of the multi-layer query instruction is the third Structured Query Language (SQL) statement. The third SQL statement is written by using a large language model according to the instruction template filled with query prompt words. The instruction template is a dialogue question used to instruct the large language model to generate an SQL statement, including the table name to be queried in the target device.
[0036] In a second aspect, the present application provides a data circuit fault analysis device, and the device includes:
[0037] An information acquisition unit, configured to acquire the basic information and basic query requirements of the service to be queried;
[0038] A bearing details unit, connected to the information acquisition unit, and configured to acquire the multi-layer bearing details of the service to be queried according to the basic information;
[0039] A query instruction unit, connected to the bearer details unit, is configured to generate multi-layer query instructions for the service to be queried according to the multi-layer bearer details and basic query requirements;
[0040] A fault location unit, connected to the query instruction unit, is configured to sequentially send multi-layer query instructions layer by layer to obtain multi-layer query results, and locate the layer and location where the fault of the service to be queried occurs currently or in the future according to the multi-layer query results.
[0041] Thirdly, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, the data circuit fault analysis method described above is implemented.
[0042] The present application provides a data circuit fault analysis method, device and medium. By querying the multi-layer bearer details of the service, generating query instructions respectively for each layer of the bearer service situation, and sending the query instructions layer by layer to locate the layer and location where the service fault occurs, it can help accurately locate the data circuit fault, improve the efficiency of fault diagnosis, and enable maintenance personnel to solve problems more pertinently, providing a strong guarantee for the safe and stable operation of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flowchart of a data circuit fault analysis method according to an embodiment of the present application;
[0044] Figure 2 is a schematic structural diagram of a data circuit fault analysis device according to an embodiment of the present application;
[0045] Figure 3 is a schematic diagram of a large language model interaction interface according to an embodiment of the present application;
[0046] Figure 4 is a schematic diagram of a diagnosis result display interface according to an embodiment of the present application;
[0047] Figure 5 is a schematic diagram of the internal structure of a large language model according to an embodiment of the present application;
[0048] Figure 6 is a schematic structural diagram of another data circuit fault analysis device according to an embodiment of the present application;
[0049] Figure 7 is a flowchart of another data circuit fault analysis method according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To enable those skilled in the art to better understand the technical solutions of the present application, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0051] It is understood that the specific embodiments and drawings described herein are only for explaining the present application and not for limiting the present application.
[0052] It is understood that, without conflict, the various embodiments in the present application and the various features in the embodiments may be combined with each other.
[0053] It is understood that for the convenience of description, only the parts related to the present application are shown in the drawings of the present application, and the parts unrelated to the present application are not shown in the drawings.
[0054] It is understood that each module and unit involved in the embodiments of the present application may correspond to only one entity structure, or may be composed of multiple entity structures. Alternatively, multiple modules and units may also be integrated into one entity structure.
[0055] It is understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present application may occur in a different order from that marked in the drawings.
[0056] It is understood that in the flowcharts and block diagrams of the present application, the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present application are shown. Among them, each block in the flowchart or block diagram may represent a module, unit, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart may be implemented by a hardware-based device for implementing the specified function, or may be implemented by a combination of hardware and computer instructions.
[0057] It is understood that the modules and units involved in the embodiments of the present application may be implemented in software or in hardware. For example, the modules and units may be located in the processor.
[0058] Embodiment 1:
[0059] As Figure 1 shown, the present application provides a method for analyzing data circuit faults, and the method includes:
[0060] S1. Obtain the basic information and basic query requirements of the service to be queried;
[0061] S2. Obtain the multi-layer bearing details of the service to be queried according to the basic information;
[0062] S3. Generate multi-layer query instructions for the service to be queried according to the multi-layer bearing details and the basic query requirements;
[0063] S4. Sequentially send the multi-layer query instructions layer by layer to obtain the hierarchical query results, and locate the layer and position where the service to be queried fails currently or in the future according to the hierarchical query results.
[0064] In this embodiment, the method generates query instructions separately for each layer of the service bearer by querying the multi-layer bearer details of the service, and sends the query instructions layer by layer to locate the layer and position where the service failure occurs. This can help accurately locate data circuit failures, improve the efficiency of fault diagnosis, enable maintenance personnel to solve problems more targeted, and provide a strong guarantee for the safe and stable operation of the network. As Figure 1 shown, the method is correspondingly applied to the device as Figure 2 shown.
[0065] Specifically, this embodiment provides a method for data circuit fault detection and processing analysis based on an artificial intelligence large model, which can provide an interface as Figure 3 shown. Users can interact with the system on this interface to query and perform performance prediction on specified service bearer devices, thereby realizing fault diagnosis. At the same time, the historical faults of the devices can be queried to assist in fault diagnosis. Finally, a detailed fault analysis result as Figure 4 shown can be provided. The large model includes each layer as Figure 5 shown. The functions of the large model include generating instructions, and the system uses the instructions to implement data collection as Figure 6 and 7 shown. At present, the positioning technology for data circuit faults is relatively backward, and problems such as slow positioning speed and low accuracy are often faced during the fault handling process. This not only affects service efficiency but also increases operating costs. At the same time, users' dependence on communication services is increasing continuously. Once a fault occurs, it will directly affect users' work and life. Realizing the rapid positioning of the fault point and cause analysis will help quickly restore communication services, improve user satisfaction, reduce operating costs, and enhance market competitiveness.
[0066] This embodiment provides an artificial intelligence-based data circuit fault analysis method, which is of great significance for ensuring network stability and reliability. By using artificial intelligence algorithms and large model technologies, a knowledge base is constructed from a large amount of historical data, and through training and learning, it is possible to achieve rapid and accurate diagnosis of circuit faults, avoiding the blindness and subjectivity in traditional methods, and significantly improving the accuracy and efficiency of fault location. In addition, this technology can also provide accurate fault information and processing suggestions for maintenance personnel, reduce the cost of fault handling, further enhance network stability and reliability. With the continuous development and application of large model technologies, it will promote the transformation and upgrading of the industry and provide a strong guarantee for the safe and stable operation of the network.
[0067] This embodiment provides a fault diagnosis system based on artificial intelligence and large model technologies. Through deep learning algorithms, it can automatically extract fault features from a large amount of historical data, thereby quickly and accurately identifying the fault type without human intervention and providing corresponding maintenance suggestions. The application of large model technologies can also enable the fault diagnosis system to have stronger generalization ability and adaptability. Even in the face of fault situations that have never been seen before, the system can put forward reasonable diagnostic opinions based on the learned knowledge base. This not only helps reduce maintenance costs, improve the safety and reliability of equipment, but also significantly shortens the fault handling time and improves the overall operation efficiency. Introducing artificial intelligence and large model technologies into the fault diagnosis system is a key step in improving the system's intelligence level and enhancing the fault handling ability, and is of great significance for promoting technological innovation and industrial upgrading in related fields.
[0068] More specifically, the system designs a set of intelligent fault analysis processes, including three stages: large network environment testing, aggregation layer operation status analysis, and access layer operation status analysis. Each stage locates faults through specific data metrics (such as port status, traffic, CRC, packet loss quantity, etc.), and finally forms a conclusion on the cause of the fault and treatment suggestions.
[0069] In the large network environment testing stage, it mainly evaluates the health status of the entire network infrastructure to ensure that there is no degradation in service quality due to core network problems. The key data metrics include: Link utilization rate: Check whether the links between core routers are overloaded; Delay and jitter: Measure the average delay time and its variation (i.e., jitter) between different core nodes; Packet loss rate: Evaluate whether there is data loss during data transmission from one core node to another; Routing table consistency: Confirm that the routing tables on all core routers are consistent and correct to avoid routing loops or incorrect routing.
[0070] During the analysis stage of the aggregation layer's operating status, it focuses on checking the connectivity and performance of the aggregation layer devices to determine whether there are configuration errors or hardware failures affecting the business. Important metrics include: Port status: Check whether the ports connected to the customer's business are in the UP (normal operation) state; Traffic pattern: Analyze the traffic pattern on a specific aggregation layer device to identify any anomalies or bottlenecks; CRC (Cyclic Redundancy Check) error count: Although more common in the access layer, it may also need to be monitored in the aggregation layer, especially when there are physical layer issues; MAC (Media Access Control) address learning: Check the MAC address table on the aggregation switch to ensure that the MAC addresses of the customer devices are correctly learned on the corresponding ports; ARP (Address Resolution Protocol) cache check: Ensure that the correct IP (Internet Protocol)-MAC mapping exists in the ARP cache.
[0071] During the analysis stage of the access layer's operating status, which is the network layer closest to the end users and directly related to the quality of the user experience, the focus is on detecting the specific physical and logical connection statuses. Relevant data metrics include: Physical port status: Confirm whether the port is in the UP or DOWN (stopped operation) state and whether there are frequent state switches; Real-time traffic: Monitor the traffic on the port to find any abnormal peaks or persistent low traffic phenomena; CRC check result: A high number of CRC errors usually indicates physical layer problems, such as cable damage or interface card failure; Online users / IP addresses: Check whether the customer hardware addresses are online, reflecting the connectivity of the layer 2 network; Ping (Packet Internet Groper) success rate: Ping the online IP of the client and return the success rate to evaluate the reachability of the layer 3 network.
[0072] By deploying the network device login connection and data collection module, combining the customized large model architecture with natural language processing technology, a conversational interaction based on semantic understanding is achieved, dynamically generating an SQL (Structured Query Language) instruction set adapted to multiple scenarios to complete efficient and accurate data collection; the core lies in the deep integration of the large model and domain knowledge, which is optimized through self-supervised pre-training and transfer learning, enabling it to directly parse unstructured natural language instructions and map them to structured data retrieval logic.
[0073] Such as Figure 5As shown, the large model adopts a hierarchical Transformer architecture, including a semantic encoding layer, a domain knowledge adaptation layer, a logical reasoning layer, and an instruction generation layer. The semantic encoding layer is based on the BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model) pre-trained model, and extracts the contextual semantic features of the instruction through a bidirectional attention mechanism to generate a sentence vector representation. The domain knowledge adaptation layer loads the professional corpus in the field of communication device operation and maintenance (including device parameter tables, fault case libraries, instruction manuals, etc.), and fine-tunes the model parameters through contrastive learning to improve the recognition ability of industry terms. The logical reasoning layer introduces a graph neural network (GNN, Graph Neural Network) to construct a device topology relationship graph, and combines reinforcement learning to dynamically optimize the SQL generation strategy to ensure that the generated instructions match the configuration logic of the target device. The instruction generation layer maps the semantic vector to an SQL syntax tree, generates a compliant query statement through a syntax constraint decoder, and automatically verifies its executability. The named entity recognition (NER, Named Entity Recognition) technology is used to extract the key entities in the instruction (such as device identification ID, parameter type), and the context attention mechanism is used to solve the polysemy problem (such as "port status" may refer to a physical port or a logical port). Based on the retrieval-augmented generation (RAG, Retrieval-augmented Generation) technology, the SQL template of similar cases is retrieved from the knowledge base as the context reference for the large model to generate instructions. When the user inputs a natural language instruction, the semantic encoding layer converts the instruction into a sentence vector and identifies the key entities. The logical reasoning layer combines the device topology graph to determine the IP address and access permission of the target router, and generates an adapted login instruction. The instruction generation layer outputs a structured SQL statement (such as SELECT port_status,crc_error_count FROM core_router_metrics WHERE circuit_id='**leased line' AND timestamp>=NOW()-INTERVAL'1HOUR';), and the system automatically executes the SQL instruction to collect data from the target device, and ensures the integrity and consistency of the data through the verification module. The hierarchical design (semantic encoding layer + domain knowledge adaptation layer + logical reasoning layer + instruction generation layer) of the customized large model architecture and its training method; the dynamic SQL generation mechanism, the prompt word generation technology based on RAG, and the logical constraint optimization combined with the device topology graph; the end-to-end closed-loop verification, embedding the syntax verification and permission verification modules in the instruction generation stage to avoid the load on the device caused by invalid queries.
[0074] The system uses AI (Artificial Intelligence) algorithms and large model technologies, combines real-time collected data and historical fault case libraries, and automatically generates intelligent fault analysis results and handling suggestions; it can not only quickly locate the fault point, but also provide specific repair suggestions and preventive measures according to the fault type and severity, significantly improving the efficiency and accuracy of fault handling.
[0075] The system cleans and standardizes the collected business operation status data, removes noise data, and ensures the accuracy and consistency of the data. The semantic encoding layer is used to convert the instructions into sentence vectors, and key entities are identified to extract features from the data, and key metrics related to faults (such as port status, traffic, CRC, packet loss quantity, etc.) are identified. On this basis, a large model (llama3.1-8b) based on the Transformer architecture is used for training in combination with the historical fault case library, enabling it to identify various fault modes. The large model simultaneously processes fault classification, fault location, and fault prediction tasks through multi-task learning (Multi-task Learning) to ensure the comprehensiveness and accuracy of the analysis. In addition, the retrieval-augmented generation (RAG) technology is introduced to retrieve similar fault cases from the knowledge base and generate a fault analysis report in combination with real-time data.
[0076] In the fault classification stage, the large model classifies the faults into types such as hardware faults, configuration errors, network congestion, etc. according to the extracted features and gives a confidence score. In the fault location stage, in combination with the device topology map, the large model can accurately locate the layer (such as the access layer, aggregation layer, core layer) and specific device where the fault occurs. In the fault prediction stage, based on time series analysis (such as the LSTM model), it predicts possible future faults and provides preventive maintenance suggestions. In the processing suggestion generation stage, the large model generates specific processing suggestions according to the fault type and severity (for example, suggesting replacing the device for hardware faults and providing detailed modification steps for configuration errors), and converts them into a user-friendly text format through natural language generation technology (NLG, Natural Language Generation), and also attaches visual charts (such as fault distribution maps, repair flowcharts) for technicians to quickly understand and execute.
[0077] Exemplarily, when the system detects that the status of a certain customer service port is DOWN and the traffic is zero, the large model first analyzes key indicators such as the port status, CRC error rate, and packet loss count, and combines with the historical fault case database to determine that the cause of the fault is "hardware failure of the access layer device". Subsequently, the large model generates a detailed fault report, including the fault type, location result (such as "access layer switch port failure"), handling suggestions (such as "replace the switch port module"), and preventive measures (such as "regularly check the hardware status of the device"). This report is displayed through the front-end interface, and technicians can quickly take actions according to the suggestions to restore the normal operation of the service.
[0078] Based on the Retrieval-Augmented Generation (RAG) technology and the capabilities of large models, an intelligent search driven by a dynamic knowledge base is constructed, supporting multi-dimensional fuzzy queries and semantic matching. It can quickly locate target customers and associated devices through natural language input and automatically trigger the data collection process, significantly improving search efficiency and accuracy.
[0079] In the RAG framework design, by integrating unstructured texts such as historical work order data, device configuration tables, fault case databases, and operation and maintenance manuals, a vectorized knowledge base is generated using large model embedding technology. The retrieval module adopts a dual-encoder model, semantically encodes the user input and the content of the knowledge base respectively, and filters the top-K relevant documents through cosine similarity calculation. The generation module is based on a large model with the GPT (Generative Pre-trained Transformer) architecture, combines the retrieval results with the user input, generates a structured search SQL query condition instruction, and dynamically optimizes the search strategy.
[0080] In terms of optimizing the search process, the RAG module supports semantic expansion functions. When the user enters a fuzzy condition (such as "** dedicated line failure"), the module automatically expands associated keywords (such as customer ID, device IP, associated circuit number) to improve the search coverage. At the same time, combined with the user's historical search records and the device topology relationship, the search weights are dynamically adjusted (for example, giving priority to displaying customers or high-priority devices that have been frequently accessed recently). In addition, a permission verification mechanism is embedded in the retrieval stage to ensure that users can only access data within their authorized scope and avoid information leakage.
[0081] The dynamic update mechanism of the knowledge base synchronizes data such as the processing results of operation and maintenance work orders and device configuration change records in real time, and updates the knowledge base vectors using incremental learning. At the same time, an active learning mechanism is introduced to label low-confidence retrieval results and feedback them to the large model to continuously optimize the retrieval accuracy.
[0082] Exemplarily, when a technician inputs "*Bank *Branch Private Line High Latency", the RAG module performs the following steps: Retrieve historical work orders related to "bank private line" and "high latency" (such as fault code E1023), device configurations (such as DNS configuration), and network topologies (such as the core router IP involved) from the knowledge base. The generation module combines the retrieval results, parses out potential association conditions: Customer Name LIKE "%*Bank *Branch%" AND Fault Type = "High Latency", and automatically completes the circuit number (such as 192HLW192xxx). The system accurately matches the target private line from the business ledger based on the parsing results and triggers the data collection module to collect real-time traffic, CRC, and latency data of the devices associated with the private line. The search result page also displays the handling suggestions for associated fault cases (such as "Adjust bandwidth allocation") for technicians to reference.
[0083] The system performs multi-level detections in sequence from the external network environment to the customer's internal network, such as for the external network environment, core router, service router, client, etc., to ensure comprehensive coverage of possible fault points; provide the specific implementation logic for multi-level detections, especially how to effectively integrate the detection results at each level to form a complete fault diagnosis report.
[0084] In one embodiment, S1: Obtain the basic information and basic query requirements of the business to be queried, specifically including:
[0085] In response to receiving customer trouble reporting information, obtain the business involved in the customer trouble reporting information as the first business to be queried, use a large language model to conduct a conversational Q&A with the customer to obtain the first basic information and first basic query requirements of the first business to be queried. The first basic information includes customer identity information and the type of the customer trouble reporting business, and the first basic query requirement includes the type of the customer trouble reporting business fault; or,
[0086] In response to generating a fault prediction work order based on the historical fault knowledge base, obtain the business involved in the fault prediction work order as the second business to be queried, use fuzzy search and interact with the user to obtain the second basic information and second basic query requirements of the second business to be queried. The second basic information includes the fault prediction scope, and the second basic query requirement includes the fault prediction index.
[0087] In this embodiment, as Figure 5 shown, a system corresponding to the method may include a business query module, a data collection module, an analysis module, and a content display module. The content display module may be as Figure 3 and 4 shown, or may be in other forms. As Figure 6As shown, the method starts with receiving customer trouble reports or fault prediction work orders, obtaining basic information such as billing codes, circuit codes, and IP addresses from the trouble reports or fault prediction work orders, and obtaining the bearing details of the faulty service or the service with a pending fault prediction by associating with the business ledger information. There is a large model historical fault knowledge base in the system, and RAG can assist the large model in generating fault prediction work orders, fault analysis and diagnosis results, and fault handling suggestions, etc.
[0088] Specifically, by deploying a network device login connection and data collection module, the system realizes the function of connecting to the devices carrying user services in real time and collecting service operation status information. The data collection module uses large models and natural language processing algorithms to generate SQL statements for retrieval. Only a conversational Q&A with the device is required. Through the natural language processing algorithm, sentence vectors are converted for semantic recognition and analysis, the required keywords are retrieved, the prompt words of the large model are converted, and then the corresponding instruction set is generated. Then, it is matched with the corresponding instruction library, and then converted into the corresponding SQL statement by the large model, and the data collection work is carried out in the form of instructions. When receiving a customer trouble report, the system locates the corresponding service bearing details in the business ledger system, and then initiates a network device connection process to log in to the corresponding service bearing device to collect key parameters such as the current service port status, service traffic, network quality parameters, and customer online status of the customer for further invocation.
[0089] Exemplarily, when receiving a customer trouble report, enter any one of the customer name, customer IP, product number, or circuit number in the search module, click search, and all customers that meet the query conditions are obtained according to fuzzy search. Click on the pop-up box to select one row of customers, and the information module disappears. At this time, the background network data collection module starts to collect information. The data collection module uses large models and natural language processing algorithms to generate SQL statements for retrieval. Only a conversational Q&A with the device is required. Through the natural language processing algorithm, sentence vectors are converted for semantic recognition and analysis, the required keywords are retrieved, the prompt words of the large model are converted, and then the corresponding instruction set is generated. Then, it is matched with the corresponding instruction library, and then converted into the corresponding SQL statement by the large model, and the data collection work is carried out in the form of instructions.
[0090] In one embodiment, S2. Obtain the multi-layer bearing details of the service to be queried according to the basic information, specifically including:
[0091] Generate a first Structured Query Language (SQL) statement based on the first basic information, and use the first SQL statement to obtain the first multi-layer bearing details of the first service to be queried from the operator service ledger system. The first multi-layer bearing details include the Service Router (SR) address or access switch address of the first service to be queried, as well as the core layer bearing device port, aggregation layer bearing device port, access layer bearing device port, and client bearing device address; or,
[0092] Generate a second SQL statement based on the second basic information, and use the second SQL statement to obtain the second multi-layer bearing details of the second service to be queried from the operator service ledger system. The second multi-layer bearing details include the core layer bearing device address, aggregation layer bearing device address, and access layer bearing device address of the second service to be queried.
[0093] In this embodiment, relevant content such as network device status, service operation status, and customer information is retrieved from the database through SQL retrieval. The service ledger system database contains the service details of all customers, such as customer name, IP address, product number, or circuit number, etc. When inputting the customer name or other identification information for search, the system will query this database to locate the corresponding service bearing details. The retrieval content of SQL can include: customer basic information, such as customer name, contact information, etc.; device status information: including but not limited to port status, traffic usage, CRC check result, packet loss rate, etc.; service configuration details: for example, the bandwidth size allocated to this customer, the specific service type used, etc.; historical faults and solutions: for similar problems that have occurred and their solutions, for reference in formulating new processing strategies.
[0094] Example SQL query: Suppose you want to find the relevant information of a specific customer, the SQL query:
[0095]
[0096] This example shows a combined query from the network_devices (network device) table and the customer_details (customer details) table, and filters out all relevant information belonging to "**", including its IP address, service type, bandwidth, port status, traffic data, and CRC value, etc. In this way, the system can dynamically generate corresponding SQL statements according to the keywords or conditions provided by the user, so as to effectively retrieve the required data from the database for further analysis or display.
[0097] In one embodiment, S3. Generate a multi-layer query instruction for the service to be queried according to the multi-layer bearing details and the basic query requirements, specifically including:
[0098] According to the first multi-layer bearing details and the first basic query requirements, match the preset first-layer query prompt words. The first-layer query prompt words include SR or virtual local area network (VLAN) configuration check of the access switch, media access control (MAC) address table check, and address resolution protocol (ARP) table entry check, as well as Ping test, port status query, traffic query, and cyclic redundancy check (CRC) error query for the core layer, aggregation layer, and access layer, and Ping test for the client.
[0099] Obtain the first instruction template from the preset instruction library according to the first-layer query prompt words, and fill in the SR address or access switch address of the first service to be queried, as well as the core-layer bearing device port, aggregation-layer bearing device port, access-layer bearing device port, and client bearing device address into the corresponding first instruction template to generate the first multi-layer query instruction for the first service to be queried; or,
[0100] According to the second multi-layer bearing details and the second basic query requirements, match the preset second-layer query prompt words. The second-layer query prompt words include query of fault prediction indicators for the core layer, aggregation layer, and access layer.
[0101] Obtain the second instruction template from the preset instruction library according to the second-layer query prompt words, and fill in the core-layer bearing device address, aggregation-layer bearing device address, and access-layer bearing device address of the second service to be queried into the corresponding second instruction template to generate the second multi-layer query instruction for the second service to be queried.
[0102] In this embodiment, different instructions will be generated according to different requirements of fault reporting query and fault prediction. In order to accurately generate various instructions, it is necessary to preset an instruction library in advance, and combine the instruction library with a large model to flexibly generate multi-layer query instructions for different requirements.
[0103] Specifically, the system analyzes the content of the conversational Q&A with the device and performs semantic recognition and analysis by converting sentence vectors through NLP (Natural Language Processing) technology. The required keywords are the words crucial for understanding the user's intention and constructing accurate data retrieval or operation commands, which may be directly extracted from the user's question or inferred through understanding the context. For example, when asking about the status of network devices, terms such as "port status", "traffic", "CRC", etc. are key metrics; while the customer name is the key information for locating specific customer records. Prompts play a key role in applications based on large models and are the inputs used to guide or specify the model to generate specific types of responses. There is a set of designed templates or guidelines to help generate these prompts, which can control the output of the large model according to fixed requirements such as fault query, traffic query, service report, etc., combined with specific keywords, to reduce the knowledge hallucinations brought by the large model. The instruction set is a series of operation commands designed to complete specific tasks (such as data collection, fault diagnosis, etc.). Keywords are extracted from the user's conversational Q&A through natural language processing algorithms, and these keywords are converted into prompts that the large model can understand. Then, specific instruction sets are generated from these prompts to ensure that the system can automatically execute complex tasks without manual intervention in each step. For example, when a user reports a problem, the system first determines the detailed information related to the problem through the search module and then generates a series of instructions to solve the problem based on this information. For instance, if the problem is about network latency, the system may generate a series of instructions to check the network status and traffic. The instruction library is a collection of a predefined series of standard operation commands or templates, storing all types of operation instructions that the system may need to execute. These instructions are carefully designed to meet the requirements of various business scenarios. Each instruction in the instruction library has a clear function, such as querying the status of a specific device, collecting a certain type of data, etc. The instruction library enables the system to quickly respond to various demands and ensures the consistency and accuracy of the execution process. The instruction set is a combination of several instructions selected from the instruction library according to specific task requirements. In other words, the instruction set is a series of instructions dynamically generated to complete a specific task. For example, when dealing with a network fault, it may be necessary to check multiple aspects of information (such as port status, traffic, CRC value, etc.), and at this time, an instruction set for this task will be formed.
[0104] To achieve the conversion from natural language queries to specific instruction sets, several key steps are usually involved: understanding the user's query intent, extracting necessary parameters and information, and mapping this information to predefined instruction templates. Suppose a user reports that their network connection speed has slowed down and provides their customer name "**" and IP address "192.168.1.1". The system needs to automatically generate a series of check commands based on this information to diagnose the problem. User input and semantic understanding: User's question: "We are **, and our network speed has slowed down. Please check our network status (IP: 192.168.1.1)." The system first uses natural language processing techniques to analyze this sentence, identify the user's intent as "check network status", and extract keywords such as "**" and "192.168.1.1". Conversion to prompt words: Based on the above analysis results, the system generates prompt words, such as: "Check the network status of customer '**' with IP address '192.168.1.1'." Matching the instruction library: The system looks up the matching instruction templates in the instruction library according to the information in the prompt words. The following templates may be found: Template 1: "Get the business port status of a specified customer", Template 2: "Query the traffic information of a specific IP address", Template 3: "Check the CRC value of a specific IP address". Generating the specific instruction set: Fill in the actual parameters (such as "**" and "192.168.1.1") using the templates to generate the specific instruction set: Instruction 1: "Get the business port status of customer '**'." Instruction 2: "Query the traffic information of IP address '192.168.1.1'." Instruction 3: "Check the CRC value of IP address '192.168.1.1'." Executing the instruction set: The system executes each operation in sequence according to the generated instruction set, generating the corresponding SQL, such as logging in to the corresponding network device and running commands to collect the required data.
[0105] The network bearing mode of the customer's business will also affect the generation of instructions, which in turn affects the instruction execution process. For example, determine whether the access mode of the customer's business is directly connected to the MAN aggregation router or to the dedicated line switch under the MAN aggregation router, automatically connect the device and start detection according to the following steps: whether the port status is UP or DOWN -> check whether there was traffic on this port 5 minutes ago -> check the CRC size during data transmission -> detect whether the customer has an IP / MAC address online -> ping the customer's online IP and return the success rate. All the above information is displayed in the network information display module. The system background checks in sequence from the external network environment, the core router CR (Core Router), the service router SR (Service Router), the client, etc., to detect whether the network data transmission from the highest-level network to the customer device is normal. After the above network information is collected, the data is imported into the intelligent fault diagnosis and analysis module for intelligent analysis: when the port is UP, there is traffic, and the CRC is zero, the coexistence of the three represents that the CR-SR side is normal in the progress bar; otherwise, if any of the three does not meet the above conditions, it means that there is a fault on the CR-SR side; when testing the large network environment, observe the status of other services under the same network. If it is normal, it means that the external network-CR side is normal: the analysis of the operation status of the aggregation layer detects whether the line from the transmission network to the customer is normal (usually the transmission device is offline or there are bad packets in the transmitted data), representing whether the cause of the fault is located on the SR-client side. Based on the information collected above, the intelligent analysis module performs fault location analysis on the above three network levels. If the whole process is normal, it means that the line is normal and there is a problem with the customer's internal network. Finally, the fault judgment conclusion is directly output to the content display module, greatly simplifying the past manual judgment process.
[0106] Whether the access mode of the customer's business is directly connected to the MAN aggregation router or to the dedicated line switch under the MAN aggregation router, there are some key differences in the detection methods, mainly because of their positions in the network architecture and the types of network devices involved.
[0107] The direct connection to the MAN aggregation router detection method includes: Port status check: First, confirm whether the port directly connected to the aggregation router is in the UP state, which usually involves executing commands such as show interfaces [interface name] to view the status of a specific port; Traffic monitoring: Check whether there is normal traffic in and out of this port. Commands such as show interface [interface name] traffic can be used to obtain real-time traffic data; CRC error check: Check whether there are cyclic redundancy check (CRC) errors, which may indicate problems at the physical layer. This can be achieved through commands such as show interfaces [interface name] counters errors; Ping test: Conduct a ping test on the customer IP address to verify the reachability from the core network to the customer and evaluate the latency; Routing table check: Ensure that the correct routing entries exist so that data packets can be correctly forwarded to the customer. The relevant routing information can be viewed through the command show ip route.
[0108] The detection method through the dedicated line switch includes: Port status check: Similar to the direct connection, it is also necessary to check the port status of the connection to the dedicated line switch. However, here it is also necessary to additionally check the status of the dedicated line switch itself and the link status between it and the aggregation router; VLAN (Virtual Local Area Network) configuration check: Since the dedicated line switch is usually used to support multiple customers, it is necessary to ensure the correct VLAN configuration. Commands such as show vlan brief can be used to confirm whether the relevant VLAN settings are correct; MAC address table check: View the MAC address table on the dedicated line switch to ensure that the MAC address of the customer device is correctly learned on the corresponding port. Commands such as show mac address-table can help complete this task; Traffic analysis: In addition to checking the traffic of the dedicated line switch port, it is also necessary to analyze the traffic pattern on the path from the customer to the aggregation router to identify any anomalies or bottlenecks; Layer 2 connectivity test: Since more Layer 2 network operations are involved, special attention should be paid to the Layer 2 connectivity test, including ARP cache check, STP (Spanning Tree Protocol) status, etc. For example, use show arp to view the ARP table entries and use show spanning-tree to check the status of the spanning tree protocol.
[0109] The direct access method focuses more on high-level routing and port status checks as it involves fewer intermediate devices. The method of accessing through a dedicated line switch requires more attention to Layer 2 network details, such as VLAN configuration, MAC address learning, and Layer 2 connectivity issues. Although both access methods need to perform basic port status and traffic monitoring, the latter involves more network layers (especially Layer 2), so more complex factors need to be considered during detection. This distinction helps to accurately locate the fault source and take corresponding solutions.
[0110] In one implementation, in S4, multi-layer query instructions are sequentially sent layer by layer to obtain a hierarchical query result, and the layer and location where the service to be queried currently has a fault are located according to the hierarchical query result. Specifically, it includes:
[0111] Send a VLAN configuration check instruction, a MAC address table check instruction, and an ARP table entry check instruction to the SR or access switch. If there is any error in the VLAN configuration, MAC address, or ARP table entry, locate the position where the first service to be queried currently has a fault as the SR or access switch configuration;
[0112] If the VLAN configuration, MAC address, and ARP table entries are all correct, send a core layer Ping test instruction to the core layer carrier device port, and judge whether the core layer carrier link is normal according to the core layer Ping test result. If not, send a core layer port status query instruction, a core layer traffic query instruction, and a core layer CRC error query instruction to judge the position where the first service to be queried currently has a fault in the core layer;
[0113] If the core layer carrier link is normal, send a aggregation layer Ping test instruction to the aggregation layer carrier device port, and judge whether the aggregation layer carrier link is normal according to the aggregation layer Ping test result. If not, send a aggregation layer port status query instruction, a aggregation layer traffic query instruction, and a aggregation layer CRC error query instruction to judge the position where the first service to be queried currently has a fault in the aggregation layer;
[0114] If the aggregation layer carrier link is normal, send an access layer Ping test instruction to the access layer carrier device port, and judge whether the access layer carrier link is normal according to the access layer Ping test result. If not, send an access layer port status query instruction, an access layer traffic query instruction, and an access layer CRC error query instruction to judge the position where the first service to be queried currently has a fault in the access layer;
[0115] If the access layer bearer link is normal, send an access layer Ping test command to the client bearer device, and judge whether the client bearer link is normal according to the client Ping test result. If not, locate the position where the first service to be queried has a current fault as a client hardware fault. If so, locate the position where the first service to be queried has a current fault as a client software fault.
[0116] In this embodiment, the system will perform layer-by-layer fault troubleshooting starting from the external network environment, passing through the core router (CR), service router (SR), etc. in sequence according to the network hierarchy, and finally reaching the client device. This method helps to narrow down the fault range and quickly locate the level where the problem lies. It is generally divided into several levels: External network environment: The main focus is on the connection quality provided by Internet services, including bandwidth usage, latency, etc.; Core router (CR): The key lies in evaluating the health of the core network, such as link utilization, whether there is packet loss, connectivity between core nodes, etc.; Service router (SR): It is more focused on the performance of specific service flows, such as the application effect of the quality of service (QoS) policy, performance indicators on specific service paths, etc.; Client: The last mile directly facing the user, with a focus on port status, the actual network speed experienced, CRC error rate, MAC address learning, whether the PING test is smooth, etc. Although each level will involve the above basic detection steps, the specific implementation details and key points of each step will be adjusted according to the network level where it is located. For example, at the client level, more attention may be paid to the status check of the physical layer and the layer 2 network, while at the CR level, more attention may be paid to routing and traffic management issues at layer 3 and above. Select appropriate detection methods based on the functional characteristics of each level and the types of problems that may occur, which can not only ensure comprehensive coverage but also improve efficiency and avoid unnecessary repetitive work.
[0117] Ping test: Perform a ping test on the management IP address of the transmission device. If it cannot be pinged, it may indicate that the device has gone offline. Measure the average latency time between different core nodes through the ping command, as well as the change in latency (i.e., jitter). High latency or unstable latency may indicate network congestion or other problems. Use ping or similar tools to test the connectivity between core layer devices and record the packet loss rate. A small amount of packet loss may be normal, but if the packet loss rate is high, it indicates that there may be serious network problems.
[0118] Use commands such as "show interfaces" to view the status of physical interfaces. Under normal circumstances, the ports connecting to the customer's services should be in the "UP / UP" state (i.e., both the physical layer and the protocol layer are normal). If it shows "DOWN", it may indicate a link problem or the device is not properly connected to the network. Check the CRC error count in the interface statistics. Commands such as "show interfaces [interface - name] counters errors" can display relevant data. A high CRC error count usually means that there are damaged data packets during data transmission. Look for input / output drops and input / output errors in the interface statistics. An increase in these values may indicate packet loss caused by network congestion or hardware failures. In some advanced network devices or transmission systems, the bit error rate can be monitored to evaluate the quality of data transmission. A high bit error rate is a sign of poor signal integrity and may lead to packet corruption. The information collected through the above methods can help determine whether the transmission device is disconnected from the network and whether there are bad packets in the transmitted data.
[0119] It is necessary to detect different parts of the network segment by segment (e.g., the external network - CR segment, the CR - SR segment, the SR - client segment, and the customer internal network segment), which can gradually narrow down the scope of the fault and finally locate the problem to a specific segment. This method helps to quickly identify and solve network faults. The following are the specific detection steps and key points for each segment:
[0120] External network - CR segment: Objective: Ensure the normal connection and service quality provided by the Internet service; Ping test: Ping the public IP address of the core router (CR) from an external public server or monitoring point to check the latency and packet loss rate; Traceroute: Use the traceroute command to trace each hop on the path from an external public node to the CR to identify potential bottlenecks or breakpoints; Bandwidth test: Use special tools to measure the bandwidth utilization and throughput to confirm whether it meets the standards stipulated in the contract; BGP (Border Gateway Protocol) status: In a multi - homed environment, check the status of the BGP session and the routing table to ensure there are no abnormal route announcements or session interruptions.
[0121] CR-SR Segment: Objective: Verify the stability and performance of the link between the Core Router (CR) and the Service Router (SR); Port Status Check: Execute the show interfaces [interface] command on the CR and SR respectively to confirm that the port status is UP / UP; Traffic Analysis: Use tools such as sflow or netflow to monitor the traffic pattern on this segment of the link and identify whether there is abnormal traffic or congestion; CRC Error Check: Check the CRC error count in the interface statistics of both ends of the device. A high value may indicate a physical layer problem; Ping Test: Conduct a ping test between the CR and SR and record the latency and packet loss situation.
[0122] SR-Client Segment: Objective: Evaluate the connection quality and performance between the Service Router (SR) and the client device; VLAN Configuration Check: If layer 2 switching is involved, check the relevant VLAN configuration to ensure the correct isolation of traffic for different customers; MAC Address Learning: View the MAC address table on the SR to confirm that the MAC address of the client device is correctly learned on the corresponding port; ARP Cache Check: Use the show arp command to ensure that there is a correct ARP entry on the SR pointing to the client device; Ping Test: Ping the IP address of the client device from the SR to evaluate the connectivity and latency.
[0123] By gradually troubleshooting the above three segments, the fault can be effectively located to a specific segment. For example, if the test results of the first two segments (External Network - CR and CR - SR) are both normal, while the third segment (SR - Client) shows obvious packet loss or increased latency, it can be initially determined that the problem lies in the segment from the SR to the client; further refined detection can help accurately find the specific reason, such as a physical fault or improper configuration of a certain access port causing the problem, or the port DOWN caused by a power outage at the user end (such faults account for 67% of the current total number of faults). This hierarchical method not only improves the efficiency of fault diagnosis but also enables maintenance personnel to solve problems more targeted.
[0124] In one embodiment, in S4, send multi - layer query instructions layer by layer to obtain the hierarchical query results, and locate the level and position where the service to be queried will have a future fault according to the hierarchical query results. Specifically, it includes:
[0125] Send the core layer fault prediction index query instruction, aggregation layer fault prediction index query instruction, and access layer fault prediction index query instruction to the core layer bearing device, aggregation layer bearing device, and access layer bearing device in sequence, and obtain the current status data of the fault prediction indexes of the core layer bearing device, aggregation layer bearing device, and access layer bearing device.
[0126] According to the current state data of the fault prediction indicators of the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device, respectively predict the future state data of the fault prediction indicators of the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device within a preset future time period. The preset future time period is the time period during which a fault may occur in the future within the fault prediction range obtained from the historical fault knowledge base;
[0127] Obtain the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device corresponding to the future state data of the fault prediction indicator lower than the preset value, so as to locate the level and location where the second service to be queried will have a fault in the future.
[0128] In this embodiment, in predicting possible future faults and providing preventive maintenance suggestions, the prediction of the service performance of the operator layer is mainly realized in segments. Based on the collected service operation state data, the system calls the automatic fault judgment and analysis module ( Figure 5 analysis module) to perform intelligent analysis and processing on the data and predict the future network state. The intelligent analysis mainly consists of three processes: Large network environment test: The operation state of the large network will also affect the customer perception. The system first tests the operation state above the core layer of the metropolitan area network to determine whether there is a large network fault that causes the decline of the customer service experience; Aggregation layer operation state analysis: Test the connectivity and operation state of the aggregation layer devices, analyze the customer IP online situation, and combine the access layer data to judge whether there are service configuration problems or device faults; Access layer operation state analysis: The access layer is the level that most intuitively reflects the customer service state. Analyzing the operation state parameters of the access port (such as port state, traffic) can directly judge the service state. The CRC and packet loss quantity reflect whether the customer's Internet access experience is good; The online situation of the customer's hardware address can directly reflect whether the layer 2 network connectivity between the access device and the customer device is normal; and so on. Combining the analysis and judgment conclusions of the above three steps, after intelligent analysis and summary, the final fault cause judgment conclusion is obtained, and a fault handling suggestion is given and returned to the foreground.
[0129] Dividing the fault diagnosis into three levels of judgment: large network environment test, aggregation layer operation state analysis, and access layer operation state analysis depends to a large extent on the key parameters collected from the service bearing devices in the first step. These key parameters provide the necessary data basis for the subsequent intelligent analysis. Specifically:
[0130] Large network environment test: This level is mainly to check whether there are problems in the entire network environment that may affect the business experience of all or some users. If problems are found in the core layer of the network (such as connection problems between core routers), then even if there is no problem with an individual user's device, their business experience will still be affected. Therefore, in the first-step data collection, information that can reflect the overall network health status is required, such as the operating status of core layer devices and link usage.
[0131] Analysis of the running status of the aggregation layer: The aggregation layer is located between the core layer and the access layer. It is responsible for aggregating traffic from multiple access points and directing it to the core layer. The analysis at this level mainly focuses on connectivity problems, device performance, and configuration errors, etc. In the first-step data collection, parameters related to aggregation layer devices need to be collected, such as the connectivity status of the devices, port status, traffic information, etc., in order to accurately evaluate whether there are problems in the aggregation layer that may cause service interruption or quality degradation.
[0132] Analysis of the running status of the access layer: The access layer is the network level closest to the end-users, directly related to the service quality and experience of individual users. The analysis focuses on detecting specific physical connections (such as port status), logical connections (such as the online status of IP addresses / MAC addresses), and network quality (such as CRC values and packet loss rates).
[0133] Therefore, in the data collection stage, it is necessary to obtain sufficiently detailed information about access layer devices, including but not limited to port status, real-time traffic, CRC check results, etc.
[0134] The fault judgment of these three levels is indeed based on the key parameters collected for the service-bearing devices in the first step. Only when the system can comprehensively and accurately collect relevant information covering all aspects from the core layer to the access layer can it effectively perform hierarchical fault location and analysis, so as to quickly and accurately find the root cause of the problem and give corresponding solutions. This hierarchical method not only helps to narrow down the scope of fault troubleshooting and improve the efficiency of problem-solving, but also helps maintenance personnel better understand the network structure and its mutual influence.
[0135] In one embodiment, the method further includes:
[0136] According to the fault repair experience in the historical fault knowledge base, obtain repair suggestions for the level and location where the fault occurs;
[0137] Send the level and location where the fault occurs and the repair suggestions to the data circuit fault repair personnel at the corresponding level.
[0138] In this embodiment, by comprehensively analyzing the conclusions of three levels (large network environment testing, aggregation layer operation status analysis, and access layer operation status analysis), the purpose is to accurately locate the cause of the fault and give effective handling suggestions. The following is a detailed example of a workflow and method to achieve this goal:
[0139] Data aggregation and preliminary analysis: It is necessary to aggregate the data collected from each level. This includes but is not limited to: the results of large network environment testing, such as the connection status between core routers, link usage, etc.; the results of aggregation layer operation status analysis, such as device connectivity, port status, traffic information, etc.; the results of access layer operation status analysis, such as physical connection status (port UP / DOWN), logical connection status (IP address / MAC address online status), network quality parameters (CRC value, packet loss rate), etc.
[0140] Hierarchical diagnosis: Conduct independent diagnosis for each level. Large network environment testing: If anomalies are found in the core layer (such as high latency, packet loss), the possible source of the problem is the entire network infrastructure, and it is necessary to further check the core layer devices or links; Aggregation layer operation status analysis: If no large network problems are found, check whether there are device failures or configuration errors in the aggregation layer. For example, overload or port failure of a certain aggregation switch may cause local service interruption; Access layer operation status analysis: Finally, if no problems are found in the previous two levels, focus on the access layer to check whether the port status of specific users is normal, whether there are a large number of CRC errors or packet loss phenomena, and whether user devices are correctly online, etc.
[0141] Comprehensive evaluation and decision-making: After completing the independent analysis of each level, conduct a comprehensive evaluation. If no obvious problems are found in all levels, but users still report a decline in service quality, then non-technical factors (such as application performance issues) may need to be considered; If a specific level shows clear anomalies, such as a certain port in the access layer frequently having CRC errors while other levels are normal, then it can be more certain that the problem lies in this specific port; When multiple levels show anomalies simultaneously, prioritize solving the problem with a larger impact range. For example, if there is an obvious bottleneck in the core layer, even if there are minor problems in the access layer, the core layer should be optimized first to improve the overall service.
[0142] Conclusions and Suggestions: Based on the above analysis, conclusions are drawn and corresponding suggestions are put forward. Conclusion: Clearly point out the level and specific location of the root cause leading to the problem (such as "Due to a failure of a certain port of a certain switch in the aggregation layer, some users are unable to access the Internet normally"); Suggestion: For hardware failures, it is recommended to replace or repair the faulty components (such as replacing the damaged switch port); For configuration errors, provide specific adjustment plans (such as modifying the VLAN configuration); If the performance degradation is caused by traffic overload, it is recommended to increase the bandwidth or optimize the traffic management strategy; If the problem involves the software or application level, relevant teams need to be contacted for further investigation and repair.
[0143] Through this method, not only can problems be systematically analyzed and solved, but also the proposed solutions can be ensured to be targeted and operable, which helps to quickly identify faults.
[0144] In one embodiment, where:
[0145] The specific form of the multi-layer query instruction is the third Structured Query Language (SQL) statement. The third SQL statement is written by using a large language model according to the instruction template filled with query prompt words. The instruction template is a dialogue question used to instruct the large language model to generate an SQL statement, which includes the table name of the data to be queried in the target device.
[0146] In this embodiment, the target device is the service-bearing device corresponding to the core layer, aggregation layer, access layer, etc. A network management system database can be established to store the status information of network devices, including but not limited to key parameters such as port status (UP / DOWN), traffic data, CRC value, online status, etc. When it is necessary to collect the service status of a specific device or a specific customer, the latest data will be obtained from here through SQL statements. A fault record database can be established. If there are historical fault records, relevant information can be retrieved from this database to assist in the diagnosis of the current problem, such as through effective measures taken in similar situations in the past. The instruction to collect data from the bearing device is connected to the actual production service device through a remote protocol or port connection. These devices are the hardware directly supporting and running customer services, which can include but are not limited to network infrastructures such as routers, switches, and servers. By interacting with these devices, the system can obtain key information about the service operation status in real time. By saving common instructions in the instruction library, the standardization of operations can be achieved, avoiding repetitive labor and improving work efficiency. Whenever there is a new similar task, the corresponding instruction can be directly called from the instruction library without having to rewrite it. Although the instruction set is customized for specific tasks, its basis is still the general instructions in the instruction library. Combined with the way of writing by the large language model, it not only ensures flexibility in the face of different tasks but also maintains a certain degree of unified specification, which helps to cope with complex and changeable task requirements. The instruction library is centrally managed and maintained. When a certain instruction needs to be updated or optimized, it only needs to be modified once in the instruction library, and all tasks depending on this instruction will automatically benefit. The instructions in the verified instruction library are usually more stable and reliable. Using these instructions to build the instruction set can reduce the probability of errors and improve the overall performance of the system. Through the large model, only natural language queries are needed, such as: Help me check the traffic situation. Through the semantic understanding of the large model, the corresponding SQL command is generated.
[0147] Embodiment 2:
[0148] As Figure 2 shown, this application provides a data circuit fault analysis device, and the device includes:
[0149] An information acquisition unit 1, configured to acquire the basic information and basic query requirements of the service to be queried;
[0150] A bearing details unit 2, connected to the information acquisition unit 1, and configured to acquire the multi-layer bearing details of the service to be queried according to the basic information;
[0151] A query instruction unit 3, connected to the bearing details unit 2, and configured to generate multi-layer query instructions for the service to be queried according to the multi-layer bearing details and the basic query requirements;
[0152] The fault location unit 4, connected to the query instruction unit 3, is used to sequentially send multi-layer query instructions in layers to obtain multi-layer query results, and locate the layer and location where the service to be queried fails currently or in the future according to the multi-layer query results.
[0153] In one embodiment, the information acquisition unit 1 specifically includes:
[0154] The customer fault reporting acquisition unit is used to, in response to receiving customer fault reporting information, obtain the service involved in the customer fault reporting information as the first service to be queried, use a large language model to conduct a conversational Q&A with the customer to obtain the first basic information and the first basic query requirements of the first service to be queried. The first basic information includes customer identity information and the type of the customer fault reporting service, and the first basic query requirements include the type of the customer fault reporting service failure; or,
[0155] The predicted work order acquisition unit is used to, in response to generating a fault prediction work order according to the historical fault knowledge base, obtain the service involved in the fault prediction work order as the second service to be queried, use fuzzy search and interact with the user to obtain the second basic information and the second basic query requirements of the second service to be queried. The second basic information includes the fault prediction scope, and the second basic query requirements include the fault prediction indicators.
[0156] In one embodiment, the bearer details unit 2 specifically includes:
[0157] The first bearer details unit, connected to the customer fault reporting acquisition unit, is used to generate a first structured query language (SQL) statement according to the first basic information, and use the first SQL statement to obtain the first multi-layer bearer details of the first service to be queried from the operator service ledger system. The first multi-layer bearer details include the service router (SR) address or access switch address of the first service to be queried, as well as the core layer bearer device port, aggregation layer bearer device port, access layer bearer device port, and client bearer device address; or,
[0158] The second bearer details unit, connected to the predicted work order acquisition unit, is used to generate a second structured query language (SQL) statement according to the second basic information, and use the second SQL statement to obtain the second multi-layer bearer details of the second service to be queried from the operator service ledger system. The second multi-layer bearer details include the core layer bearer device address, aggregation layer bearer device address, and access layer bearer device address of the second service to be queried.
[0159] In one embodiment, the query instruction unit 3 specifically includes:
[0160] The first query instruction unit includes:
[0161] The first prompt word subunit, connected to the first bearer detail unit, is used to match the preset first-layer query prompt words according to the first multi-layer bearer details and the first basic query requirements. The first-layer query prompt words include SR or virtual local area network (VLAN) configuration check of the access switch, media access control (MAC) address table check, and address resolution protocol (ARP) table entry check, as well as Ping test, port status query, traffic query, and cyclic redundancy check (CRC) error query for the core layer, aggregation layer, and access layer, and Ping test for the client.
[0162] The first instruction generation subunit, connected to the first prompt word unit, is used to obtain the first instruction template from the preset instruction library according to the first-layer query prompt words, and fill in the SR address or access switch address of the first service to be queried, as well as the core layer bearer device port, aggregation layer bearer device port, access layer bearer device port, and client bearer device address into the corresponding first instruction template to generate the first multi-layer query instruction for the first service to be queried; or,
[0163] The second query instruction unit includes:
[0164] The second prompt word subunit, connected to the second bearer detail unit, is used to match the preset second-layer query prompt words according to the second multi-layer bearer details and the second basic query requirements. The second-layer query prompt words include fault prediction index query for the core layer, aggregation layer, and access layer.
[0165] The second instruction generation subunit, connected to the second prompt word subunit, is used to obtain the second instruction template from the preset instruction library according to the second-layer query prompt words, and fill in the core layer bearer device address, aggregation layer bearer device address, and access layer bearer device address of the second service to be queried into the corresponding second instruction template to generate the second multi-layer query instruction for the second service to be queried.
[0166] In an embodiment, the fault location unit 4 includes a current fault location unit, connected to the first query instruction unit, for sequentially sending multi-layer query instructions to obtain hierarchical query results, and locating the layer and location where the service to be queried currently has a fault according to the hierarchical query results. Specifically, it includes:
[0167] The configuration fault location subunit, connected to the first instruction generation subunit, is used to send a VLAN configuration check instruction, MAC address table check instruction, and ARP table entry check instruction to the SR or access switch. If there is any error in the VLAN configuration, MAC address, or ARP table entry, locate the location where the first service to be queried currently has a fault as the SR or access switch configuration.
[0168] The core layer fault location subunit, connected to the configuration fault location subunit, is used to send a core layer Ping test instruction to the core layer carrier device port if the VLAN configuration, MAC address, and ARP table entry are all correct, and determine whether the core layer carrier link is normal according to the core layer Ping test result. If not, send a core layer port status query instruction, a core layer traffic query instruction, and a core layer CRC error query instruction to determine the location where the first service to be queried currently has a fault in the core layer;
[0169] The aggregation layer fault location subunit, connected to the core layer fault location subunit, is used to send an aggregation layer Ping test instruction to the aggregation layer carrier device port if the core layer carrier link is normal, and determine whether the aggregation layer carrier link is normal according to the aggregation layer Ping test result. If not, send an aggregation layer port status query instruction, an aggregation layer traffic query instruction, and an aggregation layer CRC error query instruction to determine the location where the first service to be queried currently has a fault in the aggregation layer;
[0170] The access layer fault location subunit, connected to the aggregation layer fault location subunit, is used to send an access layer Ping test instruction to the access layer carrier device port if the aggregation layer carrier link is normal, and determine whether the access layer carrier link is normal according to the access layer Ping test result. If not, send an access layer port status query instruction, an access layer traffic query instruction, and an access layer CRC error query instruction to determine the location where the first service to be queried currently has a fault in the access layer;
[0171] The client fault location subunit, connected to the access layer fault location subunit, is used to send an access layer Ping test instruction to the client carrier device if the access layer carrier link is normal, and determine whether the client carrier link is normal according to the client Ping test result. If not, locate the location where the first service to be queried currently has a fault as a client hardware fault. If so, locate the location where the first service to be queried currently has a fault as a client software fault.
[0172] In one embodiment, the fault location unit 4 includes a future fault location unit, connected to the second query instruction unit, which is used to sequentially send multi-layer query instructions to obtain multi-layer query results, and locate the layer and location where the service to be queried will have a fault in the future according to the multi-layer query results. Specifically, it includes:
[0173] The second instruction sending subunit, connected to the second instruction generating subunit, is used to sequentially send a core layer fault prediction index query instruction, an aggregation layer fault prediction index query instruction, and an access layer fault prediction index query instruction to the core layer carrier device, the aggregation layer carrier device, and the access layer carrier device, and obtain the current status data of the fault prediction indexes of the core layer carrier device, the aggregation layer carrier device, and the access layer carrier device;
[0174] A future indicator prediction subunit, connected to the second instruction sending subunit, is configured to respectively predict the future state data of the fault prediction indicators of the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device within a preset future time period according to the current state data of the fault prediction indicators of the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device. The preset future time period is the time period during which a fault may occur in the future within the fault prediction range obtained from the historical fault knowledge base;
[0175] A future fault location subunit, connected to the future indicator prediction subunit, is configured to obtain the core layer bearing device, the aggregation layer bearing device, and the access layer bearing device corresponding to the future state data of the fault prediction indicators lower than the preset value, so as to locate the level and location where the second service to be queried will have a fault in the future.
[0176] In one embodiment, the device further includes:
[0177] A repair suggestion unit, connected to the fault location unit 4, is configured to obtain repair suggestions for the level and location where a fault occurs according to the fault repair experience in the historical fault knowledge base;
[0178] A repair instruction unit, connected to the repair suggestion unit, is configured to send the level and location where a fault occurs and the repair suggestions to the data circuit fault repair personnel at the corresponding level.
[0179] In one embodiment, wherein:
[0180] The specific form of the multi-level query instruction is the third Structured Query Language (SQL) statement. The third SQL statement is written by using a large language model according to an instruction template filled with query prompt words. The instruction template is a dialogue question used to instruct the large language model to generate an SQL statement, which includes the table name of the data to be queried in the target device.
[0181] Example 3:
[0182] Embodiment 3 of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, it implements the data circuit fault analysis method as described in Embodiment 1, or implements the data circuit fault analysis device as described in Embodiment 2.
[0183] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program units, or other data. The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile discs (DVDs) or other optical disc storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0184] In addition, the present application may further provide a computer device including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the data circuit fault analysis method as described in Embodiment 1. The computer device may be the data circuit fault analysis device as described in Embodiment 2.
[0185] Among them, the memory is connected to the processor. The memory may adopt flash memory, read-only memory or other memories, and the processor may adopt a central processing unit or a single-chip microcomputer.
[0186] Embodiments 1-3 of the present application provide a data circuit fault analysis method, device and medium. By querying the multi-layer bearer details of services, query instructions are respectively generated for each layer of bearer services, and the query instructions are sent layer by layer to locate the layer and position where the service fault occurs, which can help accurately locate the data circuit fault, improve the efficiency of fault diagnosis, and enable maintenance personnel to solve problems more pertinently, providing a strong guarantee for the safe and stable operation of the network.
[0187] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principle of the present application. However, the present application is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present application, and these modifications and improvements are also regarded as the protection scope of the present application.
Claims
1. A method for analyzing data circuit faults, characterized in that The method includes: Obtaining the basic information and basic query requirements of the service to be queried; Obtaining the multi-layer bearing details of the service to be queried according to the basic information; Generating multi-layer query instructions for the service to be queried according to the multi-layer bearing details and basic query requirements; Sequentially sending the multi-layer query instructions layer by layer to obtain the hierarchical query results, and positioning the layer and location where the service to be queried fails currently or in the future according to the hierarchical query results.
2. The method according to claim 1, wherein Obtaining the basic information and basic query requirements of the service to be queried, specifically including: In response to receiving customer trouble report information, obtaining the service involved in the customer trouble report information as the first service to be queried, and using a large language model to conduct a conversational Q&A with the customer to obtain the first basic information and first basic query requirements of the first service to be queried. The first basic information includes customer identity information and the type of the customer trouble report service, and the first basic query requirement includes the type of the customer trouble report service failure; or, In response to generating a fault prediction work order according to the historical fault knowledge base, obtaining the service involved in the fault prediction work order as the second service to be queried, and using fuzzy search and interacting with the user to obtain the second basic information and second basic query requirements of the second service to be queried. The second basic information includes the fault prediction scope, and the second basic query requirement includes the fault prediction index.
3. The method according to claim 2, wherein Obtaining the multi-layer bearing details of the service to be queried according to the basic information, specifically including: Generating a first Structured Query Language (SQL) statement according to the first basic information, and using the first SQL statement to obtain the first multi-layer bearing details of the first service to be queried from the operator service ledger system. The first multi-layer bearing details include the Service Router (SR) address or access switch address of the first service to be queried, as well as the core layer bearing device port, aggregation layer bearing device port, access layer bearing device port, and client bearing device address; or, Generating a second Structured Query Language (SQL) statement according to the second basic information, and using the second SQL statement to obtain the second multi-layer bearing details of the second service to be queried from the operator service ledger system. The second multi-layer bearing details include the core layer bearing device address, aggregation layer bearing device address, and access layer bearing device address of the second service to be queried.
4. The method according to claim 3, wherein Generating multi-layer query instructions for the service to be queried according to the multi-layer bearing details and basic query requirements, specifically including: According to the first multi-layer bearing details and the first basic query requirements, matching the preset first layer-by-layer query prompt words. The first layer-by-layer query prompt words include the Virtual Local Area Network (VLAN) configuration check of the SR or access switch, Media Access Control (MAC) address table check, and Address Resolution Protocol (ARP) table entry check, as well as the Ping test, port status query, traffic query, and Cyclic Redundancy Check (CRC) error query of the core layer, aggregation layer, and access layer, and the Ping test of the client. Obtain a first instruction template from a preset instruction library according to the query prompt words of the first layer, and fill in the SR address or access switch address of the first service to be queried, as well as the core layer carrier device port, aggregation layer carrier device port, access layer carrier device port, and client carrier device address into the corresponding first instruction template to generate a first multi-layer query instruction for the first service to be queried; or, Match the preset query prompt words of the second layer according to the second multi-layer bearing details and the second basic query requirements. The query prompt words of the second layer include fault prediction index queries for the core layer, aggregation layer, and access layer. Obtain a second instruction template from a preset instruction library according to the query prompt words of the second layer, and fill in the core layer carrier device address, aggregation layer carrier device address, and access layer carrier device address of the second service to be queried into the corresponding second instruction template to generate a second multi-layer query instruction for the second service to be queried.
5. The method according to claim 4, characterized in that Send multi-layer query instructions layer by layer to obtain layer-by-layer query results, and locate the layer and location where the service to be queried currently has a fault according to the layer-by-layer query results. Specifically, it includes: Send a VLAN configuration check instruction, a MAC address table check instruction, and an ARP table entry check instruction to the SR or access switch. If there is an error in any of the VLAN configuration, MAC address, and ARP table entry, locate the location where the first service to be queried currently has a fault as the SR or access switch configuration. If the VLAN configuration, MAC address, and ARP table entries are all correct, send a core layer Ping test instruction to the core layer carrier device port, and judge whether the core layer carrier link is normal according to the core layer Ping test result. If not, send a core layer port status query instruction, a core layer traffic query instruction, and a core layer CRC error query instruction to judge the location where the first service to be queried currently has a fault in the core layer. If the core layer carrier link is normal, send an aggregation layer Ping test instruction to the aggregation layer carrier device port, and judge whether the aggregation layer carrier link is normal according to the aggregation layer Ping test result. If not, send an aggregation layer port status query instruction, an aggregation layer traffic query instruction, and an aggregation layer CRC error query instruction to judge the location where the first service to be queried currently has a fault in the aggregation layer. If the aggregation layer carrier link is normal, send an access layer Ping test instruction to the access layer carrier device port, and judge whether the access layer carrier link is normal according to the access layer Ping test result. If not, send an access layer port status query instruction, an access layer traffic query instruction, and an access layer CRC error query instruction to judge the location where the first service to be queried currently has a fault in the access layer. If the access layer carrier link is normal, send an access layer Ping test instruction to the client carrier device, and judge whether the client carrier link is normal according to the client Ping test result. If not, locate the location where the first service to be queried currently has a fault as a client hardware fault. If so, locate the location where the first service to be queried currently has a fault as a client software fault.
6. The method according to claim 4, wherein Send multi-layer query instructions layer by layer to obtain multi-layer query results, and locate the layer and location where the service to be queried will fail in the future, specifically including: Send the core layer fault prediction index query instruction, the aggregation layer fault prediction index query instruction, and the access layer fault prediction index query instruction to the core layer bearer device, the aggregation layer bearer device, and the access layer bearer device in sequence, and obtain the current status data of the fault prediction indexes of the core layer bearer device, the aggregation layer bearer device, and the access layer bearer device; According to the current status data of the fault prediction indexes of the core layer bearer device, the aggregation layer bearer device, and the access layer bearer device, respectively predict the future status data of the fault prediction indexes of the core layer bearer device, the aggregation layer bearer device, and the access layer bearer device within a preset future time period. The preset future time period is the time period during which a fault may occur in the future fault prediction range obtained from the historical fault knowledge base; Obtain the core layer bearer device, the aggregation layer bearer device, and the access layer bearer device corresponding to the future status data of the fault prediction index lower than the preset value, so as to locate the layer and location where the second service to be queried will fail in the future.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the repair suggestions for the layer and location where the fault occurs according to the fault repair experience in the historical fault knowledge base; Send the layer and location where the fault occurs and the repair suggestions to the data circuit fault repair personnel at the corresponding layer.
8. The method according to any one of claims 1 to 6, characterized in that Wherein: The specific form of the multi-layer query instruction is the third Structured Query Language (SQL) statement. The third SQL statement is written by using a large language model according to the instruction template filled with query prompt words. The instruction template is a dialogue question used to instruct the large language model to generate an SQL statement, which includes the table name of the data to be queried in the target device.
9. A data circuit fault analysis device, characterized in that, The device includes: An information acquisition unit, configured to acquire the basic information and basic query requirements of the service to be queried; A bearer details unit, connected to the information acquisition unit, configured to acquire the multi-layer bearer details of the service to be queried according to the basic information; A query instruction unit, connected to the bearer details unit, configured to generate multi-layer query instructions for the service to be queried according to the multi-layer bearer details and the basic query requirements; A fault location unit, connected to the query instruction unit, configured to send multi-layer query instructions layer by layer to obtain multi-layer query results, and locate the layer and location where the service to be queried fails currently or in the future according to the multi-layer query results.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it implements the data circuit fault analysis method according to any one of claims 1-8.