Log processing method and system and storage medium
By using edge message clustering and cloud platform dynamic load balancing technology, combined with the AI analysis module to process system logs, the problem of insufficient carrying capacity of a single-node MQTT proxy server was solved, and stable and high-concurrency communication between industrial Internet of Things and Internet of Vehicles devices was achieved.
Patent Information
- Application Number
- CN202511025797.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-12
AI Technical Summary
A single-node message queue transport protocol proxy server cannot handle a large number of device connections, resulting in communication interruptions and delays.
Communication messages are routed to the cloud platform through the edge message cluster, the target MQTT cluster is selected using the cloud platform's dynamic load balancing strategy, and the system logs are processed through the AI analysis module to generate error data, enabling rapid fault location and stable communication.
It solves the bottleneck problem of single-node load, ensures stable communication under multi-device connection, reduces the pressure on processing nodes, and improves fault diagnosis efficiency and communication response speed.
Smart Images

Figure CN120639600A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a method, system, and storage medium for processing logs. Background Art
[0002] IoT devices refer to perception and control terminal devices connected to the Internet, which can be physical objects such as industrial sensors, vehicle networking units, and smart home controllers.
[0003] Taking Industrial Internet of Things (IIoT) equipment as an example, this equipment is installed at the production site and uses lightweight message transmission protocols to achieve machine-to-machine communication. It is required to support millions of concurrent connections and millisecond-level latency transmission capabilities. Vehicle-to-Everything (V2O) devices, as vehicle communication nodes, need to maintain stable, low-latency communication in high-speed mobile scenarios.
[0004] These devices rely on the Message Queuing Telemetry Transport Protocol (MQTP) for data exchange. This protocol, with its lightweight packet structure, low power consumption, and publish / subscribe model, is suitable for high-concurrency, low-latency communication scenarios in the Industrial Internet of Things and the Internet of Vehicles. However, a single-node MQTP proxy server cannot handle a large number of device connections. Summary of the Invention
[0005] The present application provides a method, system and storage medium for processing logs to solve the problem that a single-node message queue transmission protocol proxy server cannot handle a large number of device connections.
[0006] In a first aspect, the present application provides a method for processing logs, the method comprising:
[0007] receiving a communication message from a client device;
[0008] Routing the communication message to the cloud platform through the edge message cluster, and forwarding the communication message to the target MQTT cluster through the cloud platform, so that the target MQTT cluster and the terminal device process the communication message, the cloud platform is used to determine the target MQTT cluster;
[0009] Storing system logs generated by the target MQTT cluster;
[0010] extracting an error log from the system log;
[0011] The error log is analyzed by an AI analysis module to generate error data, where the error data includes the error type and the error cause.
[0012] In some feasible embodiments, routing the communication message to the cloud platform through the edge message cluster includes:
[0013] Configuring an edge message cluster on the client device to transmit the communication message via a first communication protocol;
[0014] Configuring a rule engine in the edge message cluster to forward the communication message to the cloud platform through a bridging function;
[0015] On the cloud platform, the target MQTT cluster is determined through a dynamic load balancing strategy based on the routing rule identified by the client device.
[0016] In some feasible embodiments, determining the target MQTT cluster by using a dynamic load balancing strategy includes:
[0017] Monitoring the connection load status of the MQTT cluster;
[0018] Calculating a routing weight according to the connection load state;
[0019] Based on the routing weight, a target MQTT cluster is determined, so as to forward the communication message to the target MQTT cluster through the cloud platform.
[0020] In some feasible embodiments, the system log includes debugging information, operation information, warning information and error data;
[0021] The extracting the error log from the system log includes:
[0022] Monitoring the storage path of the system log through a script module, wherein the script module is an extraction component for capturing logs;
[0023] Filtering the error data by using a preset error identifier;
[0024] Extract the log content containing the preset error identifier to determine the error log.
[0025] In some feasible embodiments, analyzing the error log by the AI analysis module to generate error data includes:
[0026] Inputting the error log into the AI analysis module so that the AI analysis module performs log type classification and error pattern recognition;
[0027] Obtaining associated information of the operation information, warning information, and error data;
[0028] An analysis is performed based on the associated information to generate the error type and error cause.
[0029] In some feasible embodiments, the method further includes:
[0030] Sending the error data to a notification module, so that the notification module pushes a formatted error report to a terminal device;
[0031] A feedback instruction is received based on the formatted error report.
[0032] In some feasible embodiments, the method further includes:
[0033] Obtain resource utilization of the MQTT cluster;
[0034] When the resource utilization rate is greater than the expansion threshold, creating a node and adding the node to the MQTT cluster;
[0035] When the resource utilization rate is less than or equal to the expansion threshold, the node corresponding to the resource utilization rate less than or equal to the expansion threshold is removed.
[0036] In a second aspect, the present application provides a system for processing logs, comprising:
[0037] A message routing module, configured to receive communication messages from client devices; route the communication messages to the cloud platform via an edge message cluster; and forward the communication messages to a target MQTT cluster via the cloud platform, so that the target MQTT cluster and the terminal device process the communication messages. The cloud platform is configured to determine the target MQTT cluster.
[0038] An agent module, configured to store system logs generated by the target MQTT cluster; and extract error logs from the system logs;
[0039] The intelligent agent module includes: an AI analysis module, which is used to analyze the error log to generate error data, and the error data includes the error type and the error cause.
[0040] In some feasible embodiments, the agent module further includes a script module, and the script module is used for the storage path of the system log;
[0041] Filtering the error data by using a preset error identifier;
[0042] Extract the log content containing the preset error identifier to determine the error log.
[0043] In a third aspect, the present application provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for processing logs.
[0044] It can be seen from the above technical solutions that the present application provides a method, system and storage medium for processing logs, the method comprising: receiving a communication message from a client device; routing the communication message to a cloud platform through an edge message cluster, and forwarding the communication message to a target MQTT cluster through the cloud platform, so that the target MQTT cluster and the terminal device process the communication message, the cloud platform is used to determine the target MQTT cluster; storing the system log generated by the target MQTT cluster; extracting the error log from the system log; analyzing the error log through an AI analysis module to generate error data, the error data including the error type and the error cause. The method disperses access pressure through an edge message cluster, and the cloud platform dynamically allocates communication messages to multiple target MQTT clusters according to real-time status, breaking through the single-node load bottleneck. After the target cluster generates a system log, the error log extraction and AI analysis module are used to quickly locate the fault, ensuring stable communication under multi-device connection. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1 A flowchart of a log processing method provided in an embodiment of the present application;
[0047] Figure 2 A schematic diagram of determining an error log provided in an embodiment of the present application;
[0048] Figure 3 Schematic diagram of the AI analysis process provided in this application embodiment;
[0049] Figure 4 A schematic diagram of the structure of a log processing system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.
[0051] The embodiments of this application are suitable for high-concurrency IoT communication scenarios, specifically, inter-cluster communication between Industrial IoT (IIoT) devices and high-speed data exchange between Vehicle-to-Everything (V2X) devices. IIoT devices are deployed in manufacturing sites and require millisecond-level command transmission between machine control systems. Vehicle-to-Everything (V2X) devices, as mobile communication nodes, require stable, low-latency communication even at high speeds.
[0052] In other words, it is necessary to support concurrent access of multiple devices, industrial control command transmission delay less than the preset value, and lightweight data exchange based on the Message Queue Telemetry Transport Protocol (MQTT). Traditional single-node MQTT proxy servers and single server physical connections cannot meet access requirements, static clusters cannot dynamically adjust node scale according to traffic fluctuations, and single-point overload will cause cascade communication interruption.
[0053] MQTT is a lightweight publish / subscribe messaging protocol designed for low-bandwidth, high-latency, or unreliable network environments.
[0054] To solve the above problems, Figure 1 As shown, some embodiments of the present application provide a method for processing logs, the method comprising:
[0055] S100: Receive a communication message from a client device.
[0056] The client device is a terminal device connected to the Internet of Things network, which is used to initiate data communication requests. The communication message carries device control instructions or status data. The format complies with the message queue telemetry transmission protocol specification. The message body includes a device identifier and a subject tag. In this embodiment, the communication information is a communication request initiated by the client device to the edge message cluster, and the request carries device control instruction data.
[0057] S200: Routing the communication message to the cloud platform through the edge message cluster, and forwarding the communication message to the target MQTT cluster through the cloud platform, so that the target MQTT cluster and the terminal device process the communication message, and the cloud platform is used to determine the target MQTT cluster.
[0058] The edge messaging cluster, located at the edge of the network and composed of multiple message broker nodes, is responsible for the initial reception and routing of communication messages. Upon receiving a request, the edge messaging cluster parses the device identifier field in the message and uses a rules engine to determine the forwarding path. The rules engine matches the pre-defined routing policy and activates the bridging component to transmit the communication message to the cloud platform.
[0059] The cloud platform is built on cloud infrastructure and has a built-in cluster decision unit that selects a target MQTT cluster based on real-time status. The target MQTT cluster consists of several MQTT server nodes, which ultimately process communication messages and maintain persistent connections with end devices.
[0060] The cloud platform monitors the status of each MQTT cluster node in real time, counting the number of active connections and resource load indicators. It then calculates routing weights based on a load balancing algorithm and determines the target MQTT cluster address. Communication messages are forwarded to the target MQTT cluster's master node, which then distributes the messages to service queues.
[0061] The target MQTT cluster establishes a persistent connection with the terminal device, performs message parsing and command response operations. The running status records generated by the target MQTT cluster are written to the log storage area, and the record items include timestamps and event codes.
[0062] S300: Storing system logs generated by the target MQTT cluster.
[0063] The system log records the operation records generated during the operation of the target MQTT cluster and stores them in the cloud platform log database. For example, the log storage area is divided into four independent storage partitions according to the event level, corresponding to debugging information, operation information, warning information and error data respectively.
[0064] S400: extracting an error log from the system log.
[0065] Error logs are records in the system log that include exception markers and are extracted through a preset filtering mechanism.
[0066] The error log extraction process monitors changes to files in the storage directory in real time, triggering a pre-defined error identifier matching process. Error identifier matching performs a keyword search against the log text, targeting entries that include error level tags. The extraction process encapsulates successfully matched error logs as data transfer objects and pushes them to the AI analysis module's input queue.
[0067] S500: Analyze the error log through the AI analysis module to generate error data, where the error data includes the error type and the error cause.
[0068] The AI analysis module, a machine learning engine installed on the cloud server, parses error logs and generates text descriptions, including the root cause of the error. System logs are stored in a categorized structure, with categories including operational status records, debugging process records, alarm notification records, and error event records. Error log extraction is performed by a standalone process that continuously scans the log storage directory. The AI analysis module executes a semantic understanding algorithm to analyze the grammatical structure and keyword sequences of the error log text. Error data is output as a structured data object, including fields for the error type code and the error cause description.
[0069] Specifically, the AI analysis module loads log text and performs word segmentation, identifying entities and actions within error descriptions. The semantic analysis engine builds a contextual association model to infer the causal chain of error events. The analysis results are converted into a standard error data format, including error type classification codes and a textual description of the root cause.
[0070] Taking the industrial IoT device communication scenario as an example, a sensor device sends a temperature control command message. The edge message cluster receives it and forwards it to the cloud platform. The cloud platform selects a second MQTT cluster based on the current cluster load status. This cluster then forwards the command to the workshop temperature control device to perform the adjustment. During execution, a parameter out-of-limit error is triggered, generating an error data record that is stored in the log partition. The error extraction process identifies the "ERR_TEMP_OVERFLOW" identifier and pushes the record to the AI analysis module. The analysis module parses the text to identify a temperature sensor failure, outputting the error type as a hardware anomaly and the cause as a sensor calibration failure.
[0071] This embodiment manages massive device connections, resolving the load bottleneck of single-node message queue transmission protocol proxy servers. Traffic diversion is achieved through edge message clustering, reducing pressure on processing nodes. The cloud platform optimizes cluster resource allocation to avoid communication delays caused by overloading a single cluster.
[0072] Proprietary processing of the target MQTT cluster enables more timely terminal device command responses. The system log's multi-level storage structure supports error event retrieval, and the extraction mechanism identifies key anomaly information. The AI analysis module's semantic parsing capabilities replace manual log review, improving error diagnosis efficiency. Maintaining communication connections during high-speed mobile Internet of Vehicles devices prevents connection interruptions caused by node overload in traditional solutions.
[0073] In some embodiments, routing the communication message to the cloud platform through the edge message cluster includes:
[0074] An edge message cluster is configured on the client device, and the communication message is transmitted through a first communication protocol, wherein the client device is an operating terminal of an Internet of Things application, for example, an on-board diagnostic device in an Internet of Vehicles or a sensor controller at an industrial site, and is responsible for generating communication messages.
[0075] When the client device initializes the communication connection, it calls the SDK to configure the edge messaging cluster access point address and establishes an encrypted transmission channel based on the WebSocket protocol and the first communication protocol. The device generates a communication message containing a temperature control instruction. The message header carries the device identifier IIoT_ZoneA_1001 and the subject tag / sensor / temp / control.
[0076] The first communication protocol uses the WebSocket long connection transmission method to establish a two-way data path between the device and the edge node, which is suitable for stable data transmission in a mobile network environment.
[0077] A rules engine is configured within the edge message cluster, which forwards communication messages to the cloud platform via a bridging function. The edge message cluster, located in a regional data center, comprises a cluster architecture consisting of multiple EMQX proxy nodes, each configured with the same virtual service address for device access. The rules engine, a message processing component within the cluster, pre-configures a routing policy table based on topic wildcards to determine the rules for forwarding messages to the cloud platform. The bridging function is implemented via the EMQX bridge plug-in, replicating and forwarding local cluster messages to the cloud access point, preserving the device identifier and payload content of the original message.
[0078] For example, the cloud platform receives a bridge message and extracts the zone code, ZoneA, from the device identifier as the basis for routing decisions. The routing decision unit queries a list of candidate clusters corresponding to the zone code, including clusters Alpha and Beta. The connection status monitor detects that cluster Alpha currently has 85,000 active connections and a CPU utilization rate of 72%, while cluster Beta has 62,000 connections and a CPU utilization rate of 58%. Based on the load balancing algorithm, the weight calculator outputs a higher priority score for cluster Beta, selecting it as the target MQTT cluster. The cloud platform repackages the message and forwards it to cluster Beta's message access port.
[0079] On the cloud platform, the target MQTT cluster is determined through a dynamic load balancing strategy based on the routing rule identified by the client device.
[0080] The cloud platform, built on a public cloud infrastructure, includes a cluster management unit and a routing decision unit. The latter stores real-time status data for each MQTT cluster. A dynamic load balancing strategy calculates cluster selection weights based on real-time collected connection counts and CPU load metrics. The target MQTT cluster is physically composed of service nodes deployed in independent regions, each with its own Internet Protocol address and message access port. The device identifier encoding rules adhere to the Industrial Internet of Things (IIoT) device encoding standard and consist of a device type code, region code, and unique serial number.
[0081] For example, the master node in cluster Beta receives a message, distributes it to the service node queue via a message queue, and then the terminal temperature control device performs the temperature adjustment action. In this process, the device identifier serves as a routing parameter to ensure closed-loop processing of the message within the regional cluster.
[0082] Taking the charging pile management scenario as an example, the on-board terminal of an electric vehicle sends a charging request message, and the device identifier includes the vehicle location code V2X_GPS_1134. After the edge cluster forwards the message to the cloud platform, the routing decision unit selects the target cluster deployed nearby based on the location code to reduce the command transmission delay.
[0083] In some embodiments, determining the target MQTT cluster by a dynamic load balancing strategy includes:
[0084] Monitoring the connection load status of the MQTT cluster;
[0085] Calculating a routing weight according to the connection load state;
[0086] Based on the routing weight, a target MQTT cluster is determined, so as to forward the communication message to the target MQTT cluster through the cloud platform.
[0087] Routing weight is a numerical evaluation result generated by the cloud platform's decision-making unit, used to quantitatively compare the message processing capabilities of different clusters. The weight calculation process uses a pre-defined algorithm based on the cluster's current idle resource ratio and maximum scalable capacity. The target MQTT cluster is selected as the service group with the highest routing weight and is designated by the cloud platform as the final message processor.
[0088] For example, the cloud platform decision-making unit periodically initiates status checks and sends status query requests to each regional cluster. The first cluster returns real-time data: 128,000 active connections, 65% average CPU utilization of service nodes, and 23% remaining available memory. The second cluster returns data: 94,000 connections, 41% CPU utilization, and 38% remaining memory. The decision-making unit loads a pre-set weight calculation formula and first calculates the processing headroom of the first cluster: the maximum designed connection capacity of 200,000 minus the current number of 128,000 connections, resulting in a headroom value of 72,000. It then calculates the CPU utilization conversion factor (100% - 65%) × 0.6 and the memory headroom factor (23% × 0.4), resulting in a resource score of 44.2. Similarly, the resource score of the second cluster is calculated as 54.6.
[0089] The decision-making unit compares the scores of the two clusters and determines that the second cluster is the optimal processor. The cloud platform redirects the newly arrived charging pile control message to the second cluster's access endpoint. After receiving the message, the second cluster's load balancer assigns it to node 03 for processing based on the real-time load status of its internal service nodes. The selected service node currently has 4512 connections (below the single node limit of 5000), and its CPU utilization is 39%, responding to the charging pile startup command.
[0090] Cluster resource changes are captured through a dynamic status check process, and the resource scoring model quantitatively reflects the actual processing capacity of service nodes.
[0091] like Figure 2 As shown, in some embodiments, the system log is a collection of log files generated during the operation of the message queue telemetry transmission cluster, including debugging information, operational information, warning information, and error data. Debug information records the execution of internal methods in the service node, operational information stores device connection and message forwarding operations, warning information identifies potential abnormal conditions, and error data records functional failure details. These four types of information are identified with different prefixes and stored in a unified file system.
[0092] The extracting the error log from the system log includes:
[0093] The storage path of the system log is monitored by a script module, which is an extraction component for capturing logs. The script module is set in a message cluster operating environment and periodically detects the log update status.
[0094] Filtering the error data by using a preset error identifier;
[0095] Extract the log content containing the preset error identifier to determine the error log.
[0096] During operation, the target message queue telemetry transmission cluster service node writes operation records to the corresponding log file in the storage path. The script module activates the file system monitoring program to monitor write operation changes in each log file in the storage path. When the monitoring program detects file content additions, it triggers the log analysis program to load the newly added text content.
[0097] The analysis program parses the log text and compares the text to be tested with a preset dictionary of error identifiers. This dictionary contains multiple levels of error keywords that are fully matched against the log level prefixes. During the identifier matching process, the analysis program skips debug information, runtime information, and warning information, and only performs subsequent operations on text lines that include the error level prefix. Successfully matched log text lines are fully extracted and encapsulated as structured data objects. The script module stores the data objects in a pending queue, completing the error log extraction operation.
[0098] like Figure 3 As shown, the target MQTT cluster, namely the EMQ service in the figure, continuously generates raw log streams during operation. Logs are categorized and stored in real time according to preset levels. These logs include debug information (debug), which records details of internal method calls and parameter passing; info (info), which stores routine operations such as device connection status and message forwarding volume; warning information (warning), which identifies non-blocking anomalies such as resource thresholds approaching; and error data (error), which records critical failures such as connection interruptions and command failures.
[0099] The Shell script daemon process deployed in the Linux operating system executes dual-threaded tasks. The first thread task is a monitoring thread that scans the file write events in the log storage directory in real time; the second thread task is an extraction thread that performs incremental reading for the error log file (error level).
[0100] The thread is then extracted and loaded with a preset error keyword dictionary (such as "ERROR", "FAIL", "CRASH", etc.), and feature matching is performed on the newly added log lines. The log text content is parsed line by line to identify record items that include error level identifiers, such as error1, error2,... in the figure. Debug, info, or warning logs that do not include identifiers are skipped. Successfully matched error log items are converted into a machine-processable format.
[0101] In some embodiments, analyzing the error log by the AI analysis module to generate error data includes:
[0102] Inputting the error log into the AI analysis module so that the AI analysis module performs log type classification and error pattern recognition;
[0103] Obtaining associated information of the operation information, warning information, and error data;
[0104] An analysis is performed based on the associated information to generate the error type and error cause.
[0105] The AI analysis module parses log semantics using a pre-trained model. It receives a structured error log dataset transmitted via an API interface. The dataset includes error identifier codes and original text content. The analysis engine first performs log classification, loading a pre-set classification model to perform feature encoding on the log text. The model outputs error type labels such as network anomalies, authentication failures, and memory overflows. After classification, the error pattern recognition program is launched. Error pattern recognition involves extracting specific error feature sequences from the log text, such as a keyword combination indicating an unresponsive connection interruption command. The semantic parsing model extracts key behavior descriptions and status parameters from the text to generate an error feature vector.
[0106] Correlation information, including operational status logs and warning log text that are causally related to the error event, serves as essential context for error analysis. The module synchronously requests a collection of correlation information from the storage module, obtaining operational status logs and warning logs for a specific time window before and after the error occurred. The analysis engine constructs a cross-log correlation matrix. The operational status log provides operational metrics such as device communication volume, while the warning log provides descriptions of early abnormality signs. The module performs multi-source information fusion calculations and establishes a causal relationship model based on operational metric trends and error feature vectors. Correlation verification identifies a set of root causes leading to the error and outputs a standardized code for the error type and a textual description of the error cause.
[0107] This embodiment addresses the lack of context in traditional log analysis and eliminates the risk of misjudgment in manual diagnosis. A pre-trained model enables intelligent parsing of log content, improving the accuracy of error classification. Multi-source log fusion overcomes the information limitations of a single error log. This shortens the time it takes to diagnose authentication failure errors in Industrial IoT scenarios where tens of thousands of devices are managed.
[0108] After the error data is generated, the error data is sent to the notification module, so that the notification module pushes a formatted error report to the terminal device; and then receives a feedback instruction based on the formatted error report.
[0109] After the AI analysis module completes the error log analysis, it sends the generated error data to the input interface of the notification module. The notification module is a message push service component integrated into the cloud platform and supports sending formatted data packets to the development terminal.
[0110] The notification module loads the report template engine, converts the error type code into a standard description text, performs semantic compression on the error cause text, adds a problem location suggestion entry, and the formatting engine combines text elements to generate an error report. The formatted error report is a graphic message generated by the error type, error cause, and location suggestion according to a predetermined template, and then pushed to the bound development terminal device through an encrypted channel.
[0111] Feedback instructions are responses sent by developers after receiving a report from a terminal device, such as confirmation processing or reanalysis instructions. For example, after receiving a report, the developer's terminal device displays error details and action options on the message interaction interface. The developer reviews the cause of the device connection interruption error and selects the confirmation processing option to generate a feedback instruction.
[0112] The terminal device encodes the operation instruction into a data message and transmits it back to the cloud platform notification module through the reverse channel. The cloud platform notification module parses the message content, updates the error handling status record, and completes the closed-loop process.
[0113] To achieve dynamic load balancing, in some embodiments, the method further includes:
[0114] Obtain the resource utilization of the MQTT cluster. Resource utilization is the proportion of computing resources consumed by the message queue transmission cluster during operation, including CPU usage and memory occupancy.
[0115] When resource utilization exceeds the expansion threshold, a node is created and added to the MQTT cluster. The expansion threshold is a pre-set resource pressure limit, triggering cluster expansion. A node is a single message broker server instance, serving as the smallest service unit that can be dynamically adjusted within the cluster.
[0116] The monitoring service periodically collects CPU usage and memory utilization data for each cluster node. If the average CPU usage exceeds the expansion threshold (e.g., 75%) for three consecutive collection cycles, or if the memory utilization remains above the set value, the automatic expansion process is triggered. The resource manager creates a new node instance through the cloud platform interface, injects cluster configuration parameters, and starts the service process. After the new node is initialized, it is added to the cluster's load balancing pool to offload tasks.
[0117] When the resource utilization rate is less than or equal to the expansion threshold, the node corresponding to the resource utilization rate less than or equal to the expansion threshold is removed.
[0118] When monitoring indicates a decrease in the total number of cluster connections and CPU utilization falls below a scaling threshold (e.g., 40%), the resource manager initiates a node reclaim process. It selects the least loaded instance from the currently online nodes and migrates its sessions to other nodes. Once the migration is complete, the node service is shut down, freeing up underlying computing resources. This entire process is executed autonomously according to pre-set rules, without requiring manual intervention.
[0119] This embodiment solves the problem of service unavailability caused by traffic bursts and eliminates the response delay of traditional manual capacity expansion. Real-time resource monitoring determines capacity expansion conditions and avoids communication interruptions caused by node overload.
[0120] In the peak charging scenario of the Internet of Vehicles, node expansion is completed within the preset time when traffic increases to ensure stable transmission of device commands.
[0121] Based on the above method of processing logs, such as Figure 4 As shown, some embodiments of the present application also provide a system for processing logs, including:
[0122] A message routing module, configured to receive communication messages from client devices; route the communication messages to the cloud platform via an edge message cluster; and forward the communication messages to a target MQTT cluster via the cloud platform, so that the target MQTT cluster and the terminal device process the communication messages. The cloud platform is configured to determine the target MQTT cluster.
[0123] An agent module, configured to store system logs generated by the target MQTT cluster; and extract error logs from the system logs;
[0124] The intelligent agent module includes: an AI analysis module, which is used to analyze the error log to generate error data, and the error data includes the error type and the error cause.
[0125] The message routing module is a service unit located at the edge node that receives and forwards communication messages. It includes a client protocol adapter and a cloud platform bridge. The message routing module receives communication messages initiated by client devices at the edge layer. The device identifier and message subject are parsed by the protocol and temporarily stored in a cache. The routing decision maker invokes a cloud platform interface to obtain the address of the target message queue telemetry transmission cluster. The message is then encapsulated and forwarded to the target cluster service node via a bridge. System logs are generated during message processing in the target cluster, including debug records, operating status, warnings, and error events. These log files are written to a cloud storage area.
[0126] The intelligent agent module runs on a comprehensive processing unit in the cloud, integrating log storage services, error extraction programs, and an AI analysis module. The AI analysis module uses pre-trained models to perform log semantic analysis and infer error causes.
[0127] The intelligent agent module activates the log monitoring service and sets up a script in the storage system to continuously scan for changes in the log directory. The error extraction unit identifies error level markers in the log file and isolates text lines containing pre-set error identifiers. The API communication unit transmits the extracted error logs to the AI analysis module's input queue. The analysis engine loads the error log text and performs multi-level processing: first, feature encoding and classification of the log content, then combining the log's operation log and warning log from the same time window to provide additional context, ultimately outputting structured error data including the error type code and root cause description.
[0128] In some embodiments, the agent module further includes a script module, and the script module is used for the storage path of the system log;
[0129] Filtering the error data by using a preset error identifier;
[0130] Extract the log content containing the preset error identifier to determine the error log.
[0131] The script module is an automated processing program deployed on the log storage node, serving as the error identification unit of the agent module. It runs as an independent service process on the log storage server. During initialization, the module binds the system log storage path and establishes a file system monitoring channel. The monitoring channel monitors file write events in the storage directory in real time, triggering the scanning process when new log content is detected.
[0132] The system log storage path is the directory location of the log file on the server disk, which is stored hierarchically by date and service node. The preset error identifier is a set of preset keywords used to match the error level tags in the log text.
[0133] The scanning program reads log text line by line, loading a preset list of error identifiers. This list includes standard error labels such as "ERROR" and "CRITICAL." During the text comparison process, the program skips debug, runtime, and warning lines, retaining only text that fully matches the error identifiers. Successfully matched log lines are extracted as independent data units, appended with timestamps and source node information, and transmitted to the AI analysis module queue.
[0134] It is understandable that the embodiments of this system can refer to the embodiments of the above-mentioned method embodiments, and will not be described in detail here.
[0135] The present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Based on the above-mentioned method for processing logs, some embodiments of the present application also provide a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for processing logs.
[0136] The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal, and a software distribution medium. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0137] It can be seen from the above technical solutions that the present application provides a method, system and storage medium for processing logs, the method comprising: receiving a communication message from a client device; routing the communication message to a cloud platform through an edge message cluster, and forwarding the communication message to a target MQTT cluster through the cloud platform, so that the target MQTT cluster and the terminal device process the communication message, the cloud platform is used to determine the target MQTT cluster; storing the system log generated by the target MQTT cluster; extracting the error log from the system log; analyzing the error log through an AI analysis module to generate error data, the error data including the error type and the error cause. The method disperses access pressure through an edge message cluster, and the cloud platform dynamically allocates communication messages to multiple target MQTT clusters according to real-time status, breaking through the single-node load bottleneck. After the target cluster generates a system log, the error log extraction and AI analysis module are used to quickly locate the fault, ensuring stable communication under multi-device connection.
[0138] Similar parts between the embodiments provided in this application can be referenced to each other. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods expanded based on the scheme of this application without expending creative work shall fall within the scope of protection of this application.
Claims
1. A method for processing logs, characterized in that: The method comprises: receiving a communication message from a client device; Routing the communication message to the cloud platform through the edge message cluster, and forwarding the communication message to the target MQTT cluster through the cloud platform, so that the target MQTT cluster and the terminal device process the communication message, the cloud platform is used to determine the target MQTT cluster; Storing system logs generated by the target MQTT cluster; extracting an error log from the system log; The error log is analyzed by an AI analysis module to generate error data, where the error data includes the error type and the error cause.
2. The method for processing logs according to claim 1, wherein: Routing the communication message to the cloud platform through the edge message cluster includes: Configuring an edge message cluster on the client device to transmit the communication message via a first communication protocol; Configuring a rule engine in the edge message cluster to forward the communication message to the cloud platform through a bridging function; On the cloud platform, the target MQTT cluster is determined through a dynamic load balancing strategy based on the routing rule identified by the client device.
3. The method for processing logs according to claim 2, wherein: Determining the target MQTT cluster by using a dynamic load balancing strategy includes: Monitoring the connection load status of the MQTT cluster; Calculating a routing weight according to the connection load state; Based on the routing weight, a target MQTT cluster is determined, so as to forward the communication message to the target MQTT cluster through the cloud platform.
4. The method for processing logs according to claim 1, wherein: The system log includes debugging information, operation information, warning information and error data; The extracting the error log from the system log includes: Monitoring the storage path of the system log through a script module, wherein the script module is an extraction component for capturing logs; Filtering the error data by using a preset error identifier; Extract the log content containing the preset error identifier to determine the error log.
5. The method for processing logs according to claim 4, characterized in that: The analyzing the error log by the AI analysis module to generate error data includes: Inputting the error log into the AI analysis module so that the AI analysis module performs log type classification and error pattern recognition; Obtaining associated information of the operation information, warning information, and error data; An analysis is performed based on the associated information to generate the error type and error cause.
6. The method for processing logs according to claim 1, wherein: The method further comprises: Sending the error data to a notification module, so that the notification module pushes a formatted error report to a terminal device; A feedback instruction is received based on the formatted error report.
7. The method for processing logs according to claim 1, wherein: The method further comprises: Obtain resource utilization of the MQTT cluster; When the resource utilization rate is greater than the expansion threshold, creating a node and adding the node to the MQTT cluster; When the resource utilization rate is less than or equal to the expansion threshold, the node corresponding to the resource utilization rate less than or equal to the expansion threshold is removed.
8. A system for processing logs, characterized in that: include: A message routing module, configured to receive communication messages from client devices; and routing the communication message to the cloud platform through the edge message cluster, and forwarding the communication message to the target MQTT cluster through the cloud platform, so that the target MQTT cluster and the terminal device process the communication message, the cloud platform is used to determine the target MQTT cluster; An agent module, configured to store system logs generated by the target MQTT cluster; extracting an error log from the system log; The intelligent agent module includes: an AI analysis module, which is used to analyze the error log to generate error data, and the error data includes the error type and the error cause.
9. The system for processing logs according to claim 8, wherein: The agent module further includes a script module, and the script module is used for the storage path of the system log; Filtering the error data by using a preset error identifier; Extract the log content containing the preset error identifier to determine the error log.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the log processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Storage device control method and electronic device
CN122337409A