Traffic processing failure prediction method and device, storage medium and electronic equipment
By deploying a traffic processing agent in computing devices and analyzing operational log data using a predictive fault database, the problem of the inability to predict business traffic faults in existing technologies is solved. This enables early warning of faults and demonstration of solutions, thus avoiding losses caused by faults.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUSUR TECH CO LTD
- Filing Date
- 2024-11-07
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies cannot predict and warn of business traffic processing failures before they occur, making it impossible to avoid losses caused by failures.
By deploying a traffic processing agent in computing devices, operational log data is acquired and analyzed using a pre-defined predictive fault database to determine predicted fault details, including fault causes and solutions, and then displayed.
It enables us to understand potential fault information and solutions in advance before a fault occurs, thus avoiding fault occurrence and losses.
Smart Images

Figure CN119652776B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, storage medium and electronic device for predicting traffic processing faults. Background Technology
[0002] As software applications become increasingly complex, traditional software applications are gradually being replaced by microservice architectures. Microservice architecture refers to breaking down a large, complex software application into a series of small, independent microservices, each of which can be deployed and scaled independently. However, microservice architectures present challenges due to the complexity of communication between microservices. To address this, service mesh technology emerged. A service mesh manages communication between microservices by inserting a proxy layer, thereby reducing the burden of communication handling on the microservices themselves.
[0003] Currently, in a computer cluster, monitoring is typically done by monitoring the service traffic processing of the service mesh deployed on the computing devices. When a failure occurs in service traffic processing, Prometheus software can be used to collect, analyze, and display the log data of the service traffic processing to help users identify and resolve the problem. However, this method can only troubleshoot and resolve issues after a failure has occurred; it cannot prevent failures from happening in the first place, and therefore cannot prevent users from incurring losses due to service traffic processing failures. Summary of the Invention
[0004] In view of the above, embodiments of this application provide a method, apparatus, storage medium, and electronic device for predicting traffic processing faults, in order to at least partially solve the above-mentioned problems.
[0005] According to a first aspect of the embodiments of this application, a traffic processing fault prediction method is provided, applied to a computing device. The computing device deploys a traffic processing agent, which processes the service traffic sent and received by the computing device. The method of this embodiment includes: acquiring operational log data of the traffic processing agent; determining predicted fault log data and predicted fault details corresponding to the predicted fault data in the operational log data based on the operational log data and a preset predicted fault database. The predicted fault database includes multiple preset parameter ranges and predicted fault details corresponding to each preset parameter range. The preset parameter ranges refer to parameter ranges set with reference to target fault log data for early warning purposes. The predicted fault details include the fault causes and solutions corresponding to the target fault log data; and displaying the predicted fault log data and predicted fault details.
[0006] According to a second aspect of the embodiments of this application, a traffic processing fault prediction device is provided, applied to a computing device. The computing device deploys a traffic processing agent, which processes the service traffic sent and received by the computing device. The device of this application includes an acquisition module, a fault prediction module, and a display module. The acquisition module acquires the operation log data of the traffic processing agent. The fault prediction module determines, based on the operation log data and a preset fault prediction database, predicted fault log data and predicted fault details corresponding to the predicted fault data. The fault prediction database includes multiple preset parameter ranges and predicted fault details corresponding to each preset parameter range. The preset parameter ranges refer to parameter ranges set with reference to target fault log data for early warning. The predicted fault details include the fault causes and solutions corresponding to the target fault log data. The display module displays the predicted fault log data and predicted fault details.
[0007] According to a third aspect of the embodiments of this application, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0008] According to a fourth aspect of the present application, an electronic device is provided, including a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the method described in the first aspect.
[0009] This application provides a method, apparatus, storage medium, and electronic device for predicting traffic processing faults. The method is applied to a computing device, which deploys a traffic processing agent to process the service traffic sent and received by the computing device. The method includes: acquiring operational log data of the traffic processing agent; determining predicted fault log data and corresponding predicted fault details based on the operational log data and a preset fault prediction database. The fault prediction database includes multiple preset parameter ranges and corresponding predicted fault details. The preset parameter ranges are parameter ranges set with reference to target fault log data for early warning. The predicted fault details include the fault causes and solutions corresponding to the target fault log data; and displaying the predicted fault details. This application analyzes operational log data using preset parameter ranges in the fault prediction database to determine predicted fault log data that meets the preset parameter ranges and obtains predicted fault details. The predicted fault log data and predicted fault details are then displayed, allowing users to understand potential fault information and corresponding solutions in advance. This enables users to take measures to handle faults before they occur, effectively preventing fault occurrence and losses. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0011] Figures 1A to 1B This is a schematic diagram illustrating an application scenario of a traffic processing fault prediction method according to an exemplary embodiment of this application;
[0012] Figure 2 This is a flowchart illustrating the steps of a traffic processing fault prediction method according to an exemplary embodiment of this application;
[0013] Figure 3 This is a flowchart illustrating the steps of a traffic processing fault prediction method according to another exemplary embodiment of this application;
[0014] Figure 4 This is a structural block diagram of a traffic processing fault prediction device according to an exemplary embodiment of this application;
[0015] Figure 5 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of this application. Detailed Implementation
[0016] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.
[0017] Before describing the traffic processing fault prediction method of the embodiments of this application, the application scenarios of the traffic processing fault prediction method will be briefly described to facilitate understanding.
[0018] Figures 1A to 1B This is a schematic diagram illustrating an application scenario of a traffic processing fault prediction method according to an exemplary embodiment of this application. The traffic processing fault prediction method of this embodiment can be used in a computing device, see below. Figure 1A Computing devices can be compute nodes in a computer cluster, which can be a Kubernetes (K8s) cluster. A K8s cluster is a container orchestration and management cluster consisting of a management node and multiple compute nodes. Compute nodes, also known as workload nodes, can run containerized business applications, i.e., services. Management nodes are responsible for managing and controlling the containerized business applications running on the compute nodes, including deploying containerized business applications to the compute nodes. Both compute nodes and management nodes can be servers.
[0019] Reference Figure 1B Computing devices can also be dedicated data processors (DPUs) connected to computing nodes in a computer cluster. A DPU is a next-generation computing chip that is data-centric, I / O-intensive, and uses a software-defined technology approach to support infrastructure resource layer virtualization. It improves computing system efficiency, reduces the overall total cost of ownership, enhances data processing performance, and reduces performance overhead on other computing chips. Dedicated data processors can be plugged into computing nodes via PCIe slots to establish a connection between the DPU and the computing node. PCIe stands for PCI-Express, a bus and interface standard. Data transfer between the DPU and the computing node can occur through PCIe, allowing the DPU to process the business traffic sent and received by the computing node.
[0020] As business applications become increasingly complex, traditional business applications are gradually being replaced by microservice architectures. Microservice architecture refers to breaking down a large, complex business application into a series of small, independent microservices, each of which can be deployed and scaled independently. However, microservice architectures present the problem of complex communication between microservices. To address this, service mesh technology has emerged. A service mesh manages communication between microservices by inserting a proxy layer, thereby reducing the burden of communication handling on the microservices themselves.
[0021] A service mesh is a platform used to manage communication between microservices in a microservice architecture. It typically consists of two components: a control plane and a data plane. The control plane is responsible for tasks such as service discovery, load balancing, and traffic management, while the data plane is responsible for the actual request forwarding and processing.
[0022] In a computer cluster, service meshes are typically deployed on both management and compute nodes. The data plane is often deployed as Envoy, a traffic processing proxy. Envoy is an open-source, high-performance proxy commonly used as the data plane of the service mesh, responsible for the actual request forwarding and processing. Envoy's architecture can be divided into a data plane, a management plane, and a control plane. The control plane manages traffic routing and forwarding policies and configurations; the management plane obtains the latest traffic configuration information through standard APIs; and the data plane forwards and processes traffic based on the configuration rules issued by the control plane.
[0023] In one alternative implementation, to reduce the system resource consumption of hosts in the computer cluster, a Data Processing Unit (DPU) can be connected to the compute nodes, offloading the business traffic processing of microservices from the compute nodes to the DPU for execution. In this case, Envoy is deployed within the DPU. Here, the dedicated data processor (DPU) is not located within the computer cluster.
[0024] Traditional service meshes offer open-source log monitoring tools such as Prometheus, which can collect, analyze, and display the processing logs of traffic processing agents on each compute node. However, this approach can only troubleshoot and resolve issues after a failure has occurred, and cannot prevent failures from happening in the first place. Therefore, it cannot prevent users from suffering losses due to traffic processing failures. In view of this, embodiments of this application provide a traffic processing failure prediction method, apparatus, storage medium, and electronic device to solve the aforementioned technical problems.
[0025] Reference Figure 2 The figure shows a flowchart of a traffic processing fault prediction method according to an exemplary embodiment of the present application. As shown in the figure, this embodiment mainly includes the following steps:
[0026] S202. Obtain the operation log data of the traffic processing agent.
[0027] For example, the traffic processing agent can be Envoy, and the operation log data of Envoy can be obtained. The operation log data can include operation metrics, such as session latency of business traffic.
[0028] In one optional implementation, obtaining the operation log data of the traffic processing agent includes: obtaining the operation log data of the traffic processing agent at a preset acquisition frequency; and / or, obtaining the operation log data of the traffic processing agent in response to a received acquisition instruction.
[0029] For example, a timing module can be set in the computing device. This timing module can be used to set the frequency of acquiring the traffic processing agent's operation log data, i.e., a preset acquisition frequency. For example, the preset acquisition frequency could be 24 hours, 36 hours, or once a week, etc., and this embodiment does not impose any limitations on this. Alternatively, the system can detect user acquisition operations, receive corresponding acquisition instructions, and acquire the traffic processing agent's operation log data based on these instructions.
[0030] Furthermore, a data configuration module can be set up on the computing node. Users can configure the configuration parameter file for the preset log level through the data configuration module. When obtaining operation log data, the operation log data of the traffic processing agent can also be obtained according to the preset log level of the operation log data. For example, users can pre-classify the operation log data. The classification can be set according to the type or importance of the operation indicators contained in the operation log data. The importance refers to the degree of influence of the operation indicator on whether the business traffic processing can run normally. The greater the influence, the higher the importance. The target level of the obtained operation log data can be set. When obtaining the operation log data of the traffic processing agent, only the operation log data corresponding to the target level is obtained, so as to reduce the amount of data obtained and subsequently processed and improve data processing efficiency.
[0031] In this implementation, the operation log data of the traffic processing agent can be acquired at a preset acquisition frequency; and / or, in response to a received acquisition command, the operation log data of the traffic processing agent can be acquired. This allows users to control the frequency and timing of log analysis according to their needs, providing high flexibility.
[0032] S204. Based on the operation log data and the preset predicted fault database, determine the predicted fault log data in the operation log data, as well as the predicted fault details corresponding to the predicted fault data.
[0033] For example, the predicted fault database includes multiple preset parameter ranges and corresponding predicted fault details for each preset parameter range. The preset parameter range refers to a parameter range set with reference to target fault log data for early warning purposes. The predicted fault details include the fault causes and solutions corresponding to the target fault log data. The target fault log data can be historically processed fault log data, and the predicted fault details can include the fault causes and solutions corresponding to the historically processed fault log data. Alternatively, the target fault log data can be pre-inferred fault log data, along with corresponding fault causes and solutions, etc. This embodiment does not impose any limitations on this. The target fault log data can include fault indicators and corresponding fault thresholds. For example, a fault indicator could be that the session latency of business traffic exceeds a latency threshold, and the latency threshold can be a value flexibly set by those skilled in the art based on actual conditions.
[0034] Preset parameter ranges for early warning can be set based on fault indicators in the target fault log data. Specifically, the fault indicator can be used as the maximum or minimum threshold for the preset parameter range. For example, if the fault indicator is a session latency exceeding the latency threshold of 10ms, then the preset parameter range for session latency can be a session latency greater than 7ms and less than 10ms. The preset parameter range is a range close to the fault indicator. When the operation log data is found to fall within the preset parameter range, it indicates that the operation log data is about to reach the target fault log data, thus indicating that the data processing of the traffic processing agent is about to experience a fault corresponding to the target fault log data.
[0035] When identifying predicted fault log data and corresponding predicted fault details within the operational log data, operational metrics from the operational log data can be used as matching objects. Preset parameter ranges are then matched against the predicted fault database. If an operational log data point falls within at least one preset parameter range, it is determined to be predicted fault log data, and the predicted fault details corresponding to that preset parameter range are identified as the predicted fault details of the predicted fault log data. It should be noted that each preset parameter range may correspond to one or more predicted fault details. For example, if the preset parameter range is a session latency greater than 7ms and less than 10ms, the corresponding fault cause may be a communication network problem or a system memory problem with the computing device, such as excessive memory usage.
[0036] S206. Display the predicted fault log data and predicted fault details.
[0037] For example, if the computing device is a computing node, the predicted fault details can be displayed, and the predicted fault log data can also be displayed. This embodiment does not limit this. In an optional implementation, the computing device is a dedicated data processor (DPU). The DPU is coupled to the computing nodes in the target computer cluster. The DPU processes the service traffic sent and received by the computing nodes and displays the predicted fault log data and predicted fault details, including sending the predicted fault log data and predicted fault details to the computing nodes for display via the computing nodes. Therefore, the method of this embodiment can be applied to computing nodes or to dedicated data processors (DPUs) coupled to computing nodes, and has a wide range of applications. Furthermore, traditional Prometheus functionality requires a computer cluster network, meaning the control plane and data plane of the service mesh must be in the same cluster network. However, the DPU is not in a computer cluster network, so Prometheus software cannot collect and process log data in the DPU. In contrast, the method in this embodiment can be directly applied to the DPU, and the results are sent to the compute nodes after the entire analysis process is completed in the DPU, further reducing the CPU resource consumption of the host.
[0038] The display format is not limited to text or charts, allowing users to efficiently and intuitively understand the predicted fault causes and solutions. Furthermore, predicted fault log data can be sent to the computing node; this embodiment does not impose any restrictions on this. The computing node's data configuration module can be used to configure the configuration parameter file for preset alarm levels. Based on the preset alarm levels, the target alarm level for the predicted fault details can be determined. When displaying the predicted fault details, alarms can be triggered based on the target alarm level. The preset alarm levels can be set according to the importance of the fault indicators. Different alarm levels can correspond to different alarm methods, which may include different colors, different voice announcements, etc.; this embodiment does not impose any restrictions on this.
[0039] In one optional implementation, the computing device is a computing node in the target computer cluster, which also includes a management node. The method in this embodiment further includes: obtaining the address information of the computing device; generating a predicted fault analysis report for traffic processing of the computing device based on predicted fault log data, predicted fault details, and address information; and sending the predicted fault analysis report to the management node so that the predicted fault analysis report can be displayed by the management node.
[0040] For example, a computer cluster includes multiple computing nodes, each of which runs a traffic processing agent. Each computing node can perform traffic processing fault prediction and obtain predicted fault details. The computing nodes need to send the predicted fault details to the management node for centralized processing. Therefore, before sending the predicted fault details to the management node, the address information of the computing nodes, such as IP addresses, is obtained. Based on the predicted fault details and address information, or based on the predicted fault log data, predicted fault details, and address information, a predicted fault analysis report is generated. Then, the predicted fault analysis report is sent to the management node.
[0041] In this implementation, by adding the address information of the computing nodes to the predictive fault analysis report, the predictive fault status of traffic processing of each computing node can be quickly and accurately distinguished when the predictive fault analysis reports of the entire computer cluster are processed centrally.
[0042] In an optional implementation, the method of this embodiment further includes: obtaining update information for the predicted fault database, the update information including newly added preset parameter ranges and corresponding predicted fault details; and updating the predicted fault database based on the update information.
[0043] For example, the predicted fault database can be updated periodically. This can be achieved by providing users with an update information template through the data configuration module of the compute node. Users can then fill in the template with newly added preset parameter ranges and corresponding predicted fault details to generate a configuration parameter file. The predicted fault database is then updated based on the update information in the configuration parameter file, thus expanding the database. If the compute device is a DPU, the configuration parameter file is sent to the DPU via the compute node. The DPU updates the predicted fault database based on the update information in the configuration parameter file, expanding the database and improving its comprehensiveness, thereby making the prediction results more accurate.
[0044] In addition, if the computing device is a DPU, data can be copied using the high-speed network transmission channel between the DPU and the computing node. The high-speed network transmission channel is a high-speed network port used for internal communication between the DPU and the computing node, such as the PCIe interface, which is based on the hardware DMA implementation of the DPU, thereby improving the data interaction efficiency between the computing node and the DPU in the entire data processing process.
[0045] This embodiment sets up a predicted fault database in the computing device. The predicted fault database includes multiple preset parameter ranges and predicted fault details corresponding to each preset parameter range. The preset parameter range refers to the parameter range set with reference to the target fault log data and used for early warning. The operation log data is analyzed using the preset parameter ranges in the predicted fault database to determine the predicted fault log data that meets the preset parameter ranges and obtain the predicted fault details. The predicted fault details are then displayed so that users can understand the possible fault information and corresponding solutions in advance. Thus, measures can be taken to deal with the fault before it occurs based on the predicted fault details, which can effectively avoid the occurrence of faults and the losses caused by faults.
[0046] Reference Figure 3 The figure shows a flowchart of a traffic processing fault prediction method according to another exemplary embodiment of the present application. As shown in the figure, this embodiment mainly illustrates a specific implementation of step S204 of the above embodiment, which mainly includes the following steps:
[0047] S302. Obtain the operation log data of the traffic processing agent.
[0048] It should be noted that step S302 can be implemented with reference to the specific implementation method of step S202 above, and will not be described in detail here.
[0049] S304. Based on the preset fault prediction database, determine whether there is any log data in the operation log data that falls within the preset parameter range.
[0050] For example, the operation metrics in the operation log data are matched with the preset parameter ranges in the fault prediction database to determine whether there is log data in the operation log data that falls within the preset parameter range. For example, if the operation metrics in the operation log data include a session latency of 8ms, and the preset parameter range includes a session latency greater than 7ms and less than 10ms, then it means that there is log data in the operation log data that falls within the preset parameter range.
[0051] S306. If it exists, the log data falling within the preset parameter range will be identified as predicted fault log data, and the predicted fault details corresponding to the predicted fault log data will be determined.
[0052] For example, if the operation log data contains log data falling within one or more preset parameter ranges, such as a session delay of 8ms, it is determined to be predicted fault log data, and the predicted fault details corresponding to the preset parameter range are determined as the predicted fault details corresponding to the predicted fault log data. For example, if a session delay of 8ms is predicted fault log data, the corresponding predicted fault details may include the cause of the fault and the solution. The cause of the fault may include communication network problems or system memory problems of the computing device, such as excessive memory usage; the solution may include improving the network environment and releasing system memory. In an optional implementation, if the operation log data does not contain log data falling within the preset parameter range, it indicates that the traffic processing agent in the current computing device is operating normally, and no information is displayed; alternatively, the operation log data can also be displayed so that users can perform secondary troubleshooting and avoid omissions. This embodiment does not impose any limitations on this.
[0053] In one optional implementation, determining the predicted fault details corresponding to the predicted fault data includes: determining at least one fault cause corresponding to the predicted fault log data based on a preset predicted fault database; if there is a fault cause indicating a system fault in the computing device among the at least one fault cause, obtaining the system log data of the computing device; and determining the predicted fault details corresponding to the predicted fault log data from the predicted fault database based on the system log data and at least one fault cause.
[0054] For example, in the preset predicted fault database, each preset parameter range may correspond to one or more predicted fault details. For instance, when the preset parameter range is a session latency greater than 7ms and less than 10ms, the corresponding fault cause may be a communication network problem or a system memory problem of the computing device, such as excessive memory usage.
[0055] Determine whether there is a fault cause indicating a system failure of the computing device among at least one fault cause corresponding to the predicted fault log data, such as a system memory failure. If so, obtain the system log data of the computing device. For example, the system log data may include the operating status of the operating system of the computing device, the memory usage of the computing device, and the usage of computing cores on the computing device. This embodiment does not impose any restrictions on this.
[0056] Based on system log data, multiple causes of failure can be eliminated. For example, if the system log data determines that the computing device system has not failed, the cause of failure indicating a failure in the computing device system can be eliminated. Then, based on the remaining causes of failure after elimination, the predicted failure details corresponding to the predicted failure log data can be determined from the predicted failure database.
[0057] In this implementation, system log data from computing devices is collected, which can then be used as an aid to perform in-depth analysis of the predicted fault details in the predicted fault log data. This allows for a more accurate determination of the fault cause and solution, thus improving the accuracy of the analysis results.
[0058] In one optional implementation, determining the predicted fault details corresponding to the predicted fault log data from the predicted fault database based on system log data and at least one fault cause includes: determining a target fault cause from at least one fault cause based on system log data; determining a target solution corresponding to the target fault cause from the predicted fault database based on the target fault cause; and generating the predicted fault details corresponding to the predicted fault log data based on the target fault cause and the target solution.
[0059] For example, if the cause of the failure is a system memory failure, the memory usage of the current computing device can be determined based on the system log data. If the memory usage of the computing device is found to exceed a certain percentage threshold, which can be read from the predicted failure database, then the target failure cause can be determined to be a system memory failure. Based on the target failure cause, the corresponding target solution is matched from the predicted failure database, and then the predicted failure details corresponding to the predicted failure log data are generated based on the target failure cause and the target solution.
[0060] In this implementation, based on system log data, a target fault cause is determined from at least one fault cause, and then a target solution corresponding to the target fault cause is determined. Finally, based on the target fault cause and the target solution, predicted fault details corresponding to the predicted fault log data are generated. By utilizing system log data, the accuracy of log analysis can be improved, and the workload of users can be further reduced.
[0061] S08. Predicted fault log data and predicted fault details are displayed.
[0062] It should be noted that step S308 can be implemented with reference to the specific implementation method of step S206 above, and will not be described in detail here.
[0063] In this embodiment, the system uses a predicted fault database to determine whether any log data falls within a preset parameter range in the operation log data. If so, the log data falling within the preset parameter range is identified as predicted fault log data, and the corresponding predicted fault details are determined and displayed. Otherwise, it indicates that the traffic processing agent is operating normally. This embodiment only outputs the predicted fault log data and the corresponding predicted fault details, reducing the amount of data processing.
[0064] Reference Figure 4The diagram shows a structural block diagram of a traffic processing fault prediction apparatus according to an exemplary embodiment of the present application.
[0065] The traffic processing fault prediction device of this embodiment is applied to a computing device, which deploys a traffic processing agent to process the service traffic sent and received by the computing device. The device of this embodiment includes an acquisition module 402, a fault prediction module 406, and a display module 408.
[0066] The acquisition module 402 is used to acquire the operation log data of the traffic processing agent; the fault prediction module 406 is used to determine the predicted fault log data in the operation log data and the predicted fault details corresponding to the predicted fault data based on the operation log data and a preset predicted fault database. The predicted fault database includes multiple preset parameter ranges and the predicted fault details corresponding to each preset parameter range. The preset parameter range refers to the parameter range set with reference to the target fault log data and used for early warning. The predicted fault details include the fault causes and solutions corresponding to the target fault log data; the display module 408 is used to display the predicted fault log data and the predicted fault details.
[0067] In one optional implementation, the fault prediction module 406 is further configured to: determine whether there is log data in the operation log data that falls within a preset parameter range based on a preset fault prediction database; if so, determine the log data that falls within the preset parameter range as the fault prediction log data, and determine the fault prediction details corresponding to the fault prediction log data.
[0068] In one optional implementation, the fault prediction module 406 is further configured to: determine at least one fault cause corresponding to the predicted fault log data based on a preset predicted fault database; if there is a fault cause indicating a system fault in the computing device among the at least one fault cause, obtain the system log data of the computing device; and determine the predicted fault details corresponding to the predicted fault log data from the predicted fault database based on the system log data and at least one fault cause.
[0069] In one alternative implementation, the acquisition module 402 is further configured to: acquire the operation log data of the traffic processing agent according to a preset acquisition frequency; and / or, acquire the operation log data of the traffic processing agent in response to a received acquisition instruction.
[0070] In one optional implementation, the computing device is a computing node in the target computer cluster, which also includes a management node. The apparatus in this embodiment further includes a sending module, used to: obtain the address information of the computing device; generate a predicted fault analysis report for traffic processing of the computing device based on the predicted fault log data, predicted fault details and address information; and send the predicted fault analysis report to the management node so that the predicted fault analysis report can be displayed by the management node.
[0071] In one alternative implementation, the computing device is a dedicated data processor (DPU), which is coupled to the computing nodes in the target computer cluster. The DPU is used to process the service traffic sent and received by the computing nodes. The display module 408 is also used to send the predicted fault log data and predicted fault details to the computing nodes so that the predicted fault log data and predicted fault details can be displayed through the computing nodes.
[0072] In one optional implementation, the apparatus of this embodiment further includes an update module, configured to: obtain update information for the predicted fault database, the update information including newly added preset parameter ranges and corresponding predicted fault details; and update the predicted fault database based on the update information.
[0073] The traffic processing fault prediction device of this embodiment is used to implement the corresponding traffic processing fault prediction methods in the foregoing method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here. Furthermore, the functional implementation of each module in the traffic processing fault prediction device of this embodiment can be referred to the description of the corresponding part in the foregoing method embodiments, which will also not be repeated here.
[0074] Reference Figure 5 This diagram illustrates a structural schematic of an electronic device according to an exemplary embodiment of the present application. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.
[0075] like Figure 5 As shown, the electronic device may include: a processor 501, a memory 503, a communication bus 504, and a communication interface 505.
[0076] The processor 501, memory 503, and communication interface 505 communicate with each other via communication bus 504.
[0077] Communication interface 505 is used to communicate with other electronic devices or servers.
[0078] The processor 501 is used to execute the program 502, which can specifically execute the steps of any of the traffic processing fault prediction methods in the above embodiments.
[0079] Specifically, program 502 may include program code that includes computer operation instructions.
[0080] The processor 501 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The smart device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0081] Memory 503 is used to store program 502. Memory 503 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0082] Specifically, program 502 can be used to cause processor 501 to execute steps to implement any of the traffic processing fault prediction methods described in the embodiments. The specific implementation of each step in program 502 can be found in the corresponding descriptions of the steps and units executed in any of the traffic processing fault prediction methods described above, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments.
[0083] This application also provides a computer storage medium storing a computer program that, when executed by a processor, implements the traffic processing fault prediction method as described in any of the above method embodiments.
[0084] This application also provides a computer program product, including computer instructions that instruct a computing device to perform the operation corresponding to the traffic processing fault prediction method described in any of the above-described method embodiments.
[0085] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.
[0086] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0087] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0088] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.
Claims
1. A method for predicting traffic processing faults, characterized in that, Applied to a computing device, wherein a traffic processing agent is deployed in the computing device, the traffic processing agent is used to process the service traffic sent and received by the computing device, the method includes: Obtain the operation log data of the traffic processing agent; Based on the operation log data and a preset predicted fault database, predicted fault log data in the operation log data and predicted fault details corresponding to the predicted fault database are determined. The predicted fault database includes multiple preset parameter ranges and predicted fault details corresponding to each preset parameter range. The preset parameter ranges refer to parameter ranges set with reference to target fault log data for early warning. The predicted fault details include the fault cause and solution corresponding to the target fault log data. Determining the predicted fault details corresponding to the predicted fault database includes: based on the preset predicted fault database, determining at least one fault cause corresponding to the predicted fault log data; if the at least one fault cause indicates a fault in the computing device system, obtaining the computing device's system log data; determining a target fault cause from the at least one fault cause based on the system log data; determining a target solution corresponding to the target fault cause from the predicted fault database based on the target fault cause; and generating the predicted fault details corresponding to the predicted fault log data based on the target fault cause and the target solution. Displaying the predicted fault log data and the predicted fault details; wherein, displaying the predicted fault log data and the predicted fault details includes: determining the target alarm level of the predicted fault details according to a preset alarm level, and issuing an alarm for the predicted fault details according to the target alarm level when displaying the predicted fault details.
2. The traffic processing fault prediction method according to claim 1, characterized in that, The step of determining the predicted fault log data in the operation log data and the predicted fault details corresponding to the predicted fault database based on the operation log data and a preset predicted fault database includes: Based on a preset fault prediction database, determine whether there is any log data in the operation log data that falls within the preset parameter range; If it exists, the log data falling within the preset parameter range will be identified as predicted fault log data, and the predicted fault details corresponding to the predicted fault log data will be determined.
3. The traffic processing fault prediction method according to claim 1, characterized in that, The step of obtaining the operation log data of the traffic processing agent includes: Acquire the operation log data of the traffic processing agent according to a preset acquisition frequency; and / or, In response to the received acquisition command, the operation log data of the traffic processing agent is acquired.
4. The traffic processing fault prediction method according to claim 1, characterized in that, The computing device is a computing node in the target computer cluster, the target computer cluster further includes a management node, and the method further includes: Obtain the address information of the computing device; Based on the predicted fault log data, the predicted fault details, and the address information, a predicted fault analysis report for the traffic processing of the computing device is generated. The predicted fault analysis report is sent to the management node so that the predicted fault analysis report can be displayed through the management node.
5. The traffic processing fault prediction method according to claim 1, characterized in that, The computing device is a dedicated data processor (DPU), which is coupled to computing nodes in the target computer cluster. The DPU is used to process the service traffic sent and received by the computing nodes. The step of displaying the predicted fault log data and the predicted fault details includes: The predicted fault log data and the predicted fault details are sent to the computing node so that the predicted fault log data and the predicted fault details can be displayed through the computing node.
6. The traffic processing fault prediction method according to claim 1, characterized in that, The method further includes: Obtain updated information for the predicted fault database, including newly added preset parameter ranges and corresponding predicted fault details; The predicted fault database is updated based on the updated information.
7. A flow processing fault prediction device, characterized in that, An apparatus for use in computing devices, wherein a traffic processing agent is deployed in the computing device, and the traffic processing agent is used to process the service traffic sent and received by the computing device, the apparatus comprising: The acquisition module is used to acquire the operation log data of the traffic processing agent; A fault prediction module is used to determine predicted fault log data in the operation log data and predicted fault details corresponding to the predicted fault database based on the operation log data and a preset predicted fault database. The predicted fault database includes multiple preset parameter ranges and predicted fault details corresponding to each preset parameter range. The preset parameter ranges refer to parameter ranges set with reference to target fault log data for early warning. The predicted fault details include the fault cause and solution corresponding to the target fault log data. Determining the predicted fault details corresponding to the predicted fault database includes: determining at least one fault cause corresponding to the predicted fault log data based on the preset predicted fault database; if the at least one fault cause indicates a fault in the computing device system, obtaining the computing device's system log data; determining a target fault cause from the at least one fault cause based on the system log data; determining a target solution corresponding to the target fault cause from the predicted fault database based on the target fault cause; and generating predicted fault details corresponding to the predicted fault log data based on the target fault cause and the target solution. The display module is used to display the predicted fault log data and the predicted fault details; The data configuration module of the computing node is used to configure the configuration parameter file of the preset alarm level, so as to determine the target alarm level of the predicted fault details according to the preset alarm level; wherein, the preset alarm level is used to classify according to the importance of the fault indicators, and different alarm levels correspond to different alarm methods; correspondingly, the display module is also used to: determine the target alarm level of the predicted fault details according to the preset alarm level, and when displaying the predicted fault details, alarm the predicted fault details according to the target alarm level.
8. An electronic device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the method as described in any one of claims 1-6.
9. A computer storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.