Log management tracking method based on message middleware
By building a globally unique message identifier and log aggregation index system in a distributed system, the query difficulty problem caused by the dispersion of log information is solved, full-link tracking and efficient error handling are achieved, and the system reliability and operation and maintenance efficiency are improved.
Patent Information
- Application Number
- CN202511287340.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-10
AI Technical Summary
In distributed systems, existing technologies make it difficult to effectively integrate and analyze log information from different nodes, resulting in low efficiency in log query and error handling, and inability to provide a global link tracking view, affecting system reliability and the accuracy of business processes.
By building a logging mechanism in the message producer and consumer nodes, synchronously recording the message sending status, time, identifier and other information, establishing a globally unique message identifier, building a distributed log aggregation index system, and providing a visual log management interface, full-link tracking and exception handling across service nodes can be achieved.
Ensure that the entire process of each message is traceable, enhance system consistency and reliability, quickly locate problems, improve log analysis and error handling efficiency, and enhance system stability and operation and maintenance capabilities.
Smart Images

Figure CN120763007A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of message management, and in particular to a log management and tracking method based on message middleware. Background Art
[0002] With the widespread adoption of information technology and the expansion of enterprise scale, more and more systems and applications rely on distributed architectures for data processing and service provision. In this architecture, message-based middleware, as a key system component, undertakes asynchronous communication and data delivery, becoming an indispensable part of distributed systems. By decoupling application modules, message-based middleware improves system scalability, reliability, and performance.
[0003] In addition, similar patents such as CN118838724A disclose a log-based visual cross-system message link tracking method, system, medium and device, including: step S1: the service initiator generates a link unique identifier and passes it to the downstream service. The downstream service obtains the link representation and continues to pass it until the service call is completed; step S2: using the application service log to record the link unique identifier and call level information; step S3: using the message middleware service log to record the link unique identifier and the message unique identifier; step S4: using the log collection system to collect application service logs and message middleware service logs; step S5: using the message middleware management end to realize user query and positioning based on the service log collected by the log collection system. Although it can adopt traditional log recording and link storage methods and record message sending, receiving, processing and other operations through log files or databases, thereby providing certain log tracking capabilities, in a distributed environment, log management is often difficult because the distributed system involves multiple nodes and services, and log information is usually scattered in different systems, making log query and analysis difficult. It is impossible to effectively integrate log information from different nodes and it is difficult to provide a global link tracking view, thereby affecting the efficiency of log analysis and error handling. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems existing in the existing message log management process, the present invention provides a method for realizing full-process monitoring of the message life cycle by synchronously recording the sending status, time, identification and other information, and ensuring the accuracy and high availability of the business process through exception handling and status updates. At the same time, the constructed log aggregation index system can integrate the log data of each node, establish associations through global unique identification, and form a full-link tracking view across service nodes, which is convenient for rapid location and analysis of problems. The provided visual log management interface allows system administrators to efficiently retrieve and export log data and resend abnormal messages, thereby effectively improving the log analysis and error handling efficiency of distributed systems. The message middleware-based log management tracking method includes the following steps: In a distributed system where each service node client is connected to the same message middleware service and database service, a message sending log recording mechanism is built on the message producer node. This ensures that when a message producer sends a message to the message middleware, the message sending status, sending time, sender ID, complete sending content, globally unique message ID, business module ID, and business function name are synchronously written into the message generation log table according to a preset data structure. A message consumption logging mechanism is built on the message consumer node to implement duplication checking in the message consumption table based on the globally unique message identifier when the message is transmitted to the consumer node through the message middleware. If duplicate records exist, the consumption process is terminated. If not, the message reception status, consumer node identifier, and reception time are recorded, and the message is passed to the actual business processing logic. After the business processing logic is executed, the message status in the message consumption table is updated based on the execution result. If the business processing is successful, the status is marked as consumed. If an exception occurs, the error status code, detailed error information and stack trace data are recorded, and an error handling task record is generated at the same time. Build a distributed log aggregation indexing system to periodically collect messages from each service node to generate log tables and message consumption tables, and establish associated indexes through globally unique message identifiers to form a full-link message tracking view across nodes; A visual log management interface is constructed based on the message full-link tracking view to provide a log retrieval function based on a combination of multi-dimensional conditions, a log data export function, and an abnormal message retransmission control function.
[0005] The message sending and consumption log recording mechanism constructed in the present invention ensures that the entire process of each message from production to consumption can be accurately tracked. In particular, through the use of a globally unique message identifier, the problem of repeated message consumption is avoided, the message consistency and reliability of the system are enhanced, and the log recording after business processing is further strengthened. It not only ensures the accurate update of the message consumption status, but also provides detailed error information and stack traces when an exception occurs, which is convenient for quickly locating the problem and reducing the fault recovery time. Through the distributed log aggregation index system, the log information of each service node is regularly collected and integrated, and an associated index is established through a globally unique identifier, realizing the full-link tracking of messages across nodes, which not only helps to quickly locate the source of the problem, but also provides a comprehensive system status view, providing data support for system optimization and operation and maintenance. Finally, the visual log management interface provides multi-dimensional log retrieval, export and abnormal message retransmission control functions, so that operation and maintenance personnel can efficiently manage system logs, handle abnormal situations in real time, and further improve the efficiency of the system's log analysis and error handling. This series of steps provides comprehensive log management and error handling capabilities for the distributed system, ensuring the stable operation and high availability of the business.
[0006] Preferably, the message sending log recording mechanism includes: Implement a message pre-processing interceptor on the message producer node, which generates a composite globally unique message identifier containing a timestamp and a node identifier before sending the message; Construct a message content summary generation module, which extracts features from the original message content to generate a fixed-length content summary and stores the content summary and the full message content synchronously in the message generation log table; Implement a sending context association component that associates and stores business operation context metadata, including the identity of the user initiating the operation, the IP address of the operation source, and the operation session identifier, with the message sending record.
[0007] The present invention implements a message pre-processing interceptor and a message content summary generation module at the message producer node, which can significantly enhance the traceability and security of the system. By generating a composite globally unique message identifier containing a timestamp and a node identifier, it ensures that each message can be accurately identified and its source can be traced. In the content summary generation module, feature extraction of the original message content can not only improve the efficiency of message storage, but also effectively reduce the use of storage space while retaining the core information of the message. By storing the message summary and the complete content synchronously, the system can quickly verify the integrity and consistency of the message when performing log retrieval and data auditing. In addition, the sending context association component can associate information such as the identity of the user initiating the operation, the source IP address and the session identifier with the message record, which helps to trace the source and assign responsibility for the message when a problem occurs, thereby improving the audit and security of the system and ensuring that the sending behavior of each message is traceable and the information is not lost.
[0008] Preferably, the message consumption logging mechanism includes: Build a consumption state machine management module, which maintains the lifecycle status of message consumption, including the corresponding lifecycle states of pending consumption, consuming, consumed, consumption failure, and retrying, and automatically updates the status records in the message consumption table based on state transition rules; Implement a consumer idempotence guarantee component. This component uses database unique index constraints and optimistic locking mechanisms to ensure that the same message is consumed only once in a distributed environment. Develop a consumption context injection module, which injects message consumption context information, including consumption node environment variables, message routing path, and message retry count, into the business processing thread context before the message is passed to the business processing logic.
[0009] The consumption state machine management module constructed in the present invention plays a vital role in the message consumption process, ensuring that the life cycle of the message is strictly managed and controlled. By maintaining states such as "to be consumed", "consuming", "consumed", "consumption failed" and "retrying", the system can clearly understand the consumption progress and status of each message, avoiding the risk of message loss or repeated consumption. The consumption idempotence guarantee component further ensures that the same message is consumed only once in a distributed environment through the database unique index constraint and optimistic locking mechanism, avoiding repeated consumption due to network delays or system failures, and enhancing the stability and data consistency of the system. In addition, the consumption context injection module can inject key information such as consumption node environment variables, message routing path and number of retries into the business processing thread before the message is passed to the business processing logic, thereby providing richer context information for subsequent business processing. This mechanism greatly improves the processing efficiency of the system, reduces the need for manual intervention, and ensures the efficiency and accuracy of the message consumption process.
[0010] Preferably, the consumption state machine management module includes: Build a state transition rule engine that defines the legal state transition path for message consumption based on a finite state machine model and records the state change timestamp and trigger event at each state change. Implement a timeout status detection submodule, which periodically scans the message records in consumption, automatically marks the messages that exceed the preset processing time limit as processing timeout status, and triggers the timeout processing process; Develop a status rollback control submodule. When a retryable exception is detected in message processing, this submodule rolls back the message status to the pending consumption status and calculates the next consumption execution time according to the preset retry strategy.
[0011] The consumption state machine management module in the present invention can ensure that the state changes in the message consumption process comply with predetermined rules by constructing a state transition rule engine, and timestamps and records the triggering events for each state change, which provides detailed data support for subsequent message tracking and problem troubleshooting. Through this module, the system can accurately record the processing history of each message and capture key events at each state change, which is convenient for subsequent analysis and optimization. The timeout state detection submodule can timely mark timeout messages and trigger the timeout processing process when periodically scanning the message records in consumption, avoiding the situation where the message cannot be processed in time due to system problems or network problems during the consumption process. In addition, the state rollback control submodule can restore the message to the waiting state when it detects a retryable exception in the message processing, and automatically calculate the execution time of the next consumption according to the preset retry strategy, which greatly improves the reliability and fault tolerance of the system. By automatically processing timeout and abnormal messages, the entire message consumption process becomes more efficient and robust.
[0012] Preferably, the timeout state detection submodule includes: Build a dynamic timeout threshold configuration component. This component uses a machine learning algorithm to dynamically adjust the message processing timeout threshold for different business scenarios based on data corresponding to business type, message size, and historical processing time. It also determines the timeout message based on the message processing timeout threshold. Implement a hierarchical processing engine for timeout messages. This engine adopts different processing strategies for timeout messages based on the business importance of the message, including automatic retry, manual intervention reminder, and business process compensation. Develop a timeout processing effect evaluation module, which tracks the business impact indicators after timeout message processing, including business success rate, data consistency deviation rate, and user complaint rate, and optimizes the timeout processing strategy based on the evaluation results.
[0013] The construction of the timeout status detection submodule in the present invention enables the system to dynamically adjust the timeout threshold according to actual business needs and historical data of message processing when facing a large number of messages, which provides a more flexible strategy for message processing in different business scenarios. Through the machine learning algorithm, the timeout threshold can be automatically optimized according to data such as business type, message size and historical processing time, thereby effectively avoiding the problem of message backlog or loss caused by improper timeout settings during business peak periods or when processing complex tasks. In addition, the timeout message hierarchical processing engine can adopt different processing strategies according to the importance level of the message. For high-priority messages, the system can adopt an automatic retry strategy, and for low-priority messages, manual intervention reminders or business process compensation measures can be adopted to ensure the continuity of the business process and data consistency. Through the feedback of the timeout processing effect evaluation module, the business impact after the timeout message processing can be tracked in real time, the timeout processing strategy can be further optimized, the business losses caused by timeouts can be reduced, and user satisfaction and overall system performance can be improved.
[0014] Preferably, the distributed log aggregation indexing system includes: Build a log data collection channel. This channel uses message middleware to implement asynchronous data transmission between service nodes and log aggregation centers, and supports breakpoint-resume transmission based on watermarks. Develop a multi-dimensional index construction engine that builds a hybrid index structure of inverted index and B+ tree index based on globally unique message identifier, business module identifier, time range, and message status dimensions; Implement an incremental update mechanism for index data. This mechanism, based on database change data capture technology, senses changes in log data tables in real time and synchronizes incremental data to the index system.
[0015] In the construction of the distributed log aggregation index system, the present invention first realizes asynchronous data transmission between the service node and the log aggregation center through the message middleware, ensuring that the log data can be efficiently and stably transmitted from each service node to the aggregation center. The breakpoint resume function based on the watermark can effectively avoid data loss when the network is unstable or the service is interrupted, ensuring the integrity of the log and the high availability of the system. The development of the multi-dimensional index construction engine not only enables the log data to be flexibly queried according to multiple dimensions such as time and message identifier, but also improves the query efficiency and data retrieval accuracy through the hybrid structure of the inverted index and the B+ tree. By introducing the database change data capture technology, real-time monitoring and synchronization of incremental updates of log data are achieved, avoiding the performance bottleneck of full data updates and improving the efficiency of the system in processing data. Overall, the implementation of this step improves the response speed and data processing capabilities of the log management system, and provides strong data support for subsequent log analysis and visualization.
[0016] Preferably, the visual log management interface includes: Build a multi-dimensional conditional search engine that supports compound queries based on message content keywords, message status combinations, time intervals, and producer identification, and provides query condition saving and reuse functions; Develop a log data visualization component that generates visual analysis charts corresponding to message flow topology diagrams, processing time heat maps, and business throughput trend charts based on full-link tracking data; Implement an intelligent abnormal message positioning module, which automatically identifies suspicious message patterns based on preset abnormal pattern matching rules and provides abnormal root cause analysis suggestions.
[0017] The construction of the visual log management interface in the present invention provides users with intuitive and easy-to-use log retrieval and analysis functions. With the support of the multi-dimensional conditional retrieval engine, users can flexibly combine different query conditions, such as message content, time range, etc., to quickly locate relevant log data, significantly improving the accuracy and efficiency of data retrieval. The development of log data visualization components, especially the visualization of full-link tracking data, enables users to view the message flow path and its processing process in the entire system in real time, so as to intuitively discover performance bottlenecks or fault nodes. The display forms of various charts such as heat maps and trend charts make complex log data easier to understand and analyze, helping users to quickly identify abnormal states and performance bottlenecks of the system. The introduction of the abnormal message intelligent positioning module can automatically identify potential abnormal messages through preset pattern matching rules, and provide root cause analysis suggestions, which helps to quickly respond to and solve system problems and improve system stability and service quality.
[0018] Preferably, the log data visualization component includes: Build a dynamic topology generation engine that generates service call topology diagrams in real time based on message flow paths, and supports interactive functions such as node expansion / contraction, link highlighting, and performance indicator overlay display. Develop a time series data analysis module. This module, based on time series database technology, performs multi-dimensional analysis of message processing time, throughput, and error rate indicators, and provides automatic early warning functions for abnormal indicators. Implement a pivot table generator that supports user-defined data dimensions and aggregation methods, generates interactive pivot reports in real time, and supports data drill-down analysis.
[0019] The further refinement of the log data visualization component in the present invention, especially the construction of a dynamic topology generation engine, enhances the user's comprehensive understanding of the system service call link. The real-time generated service call topology diagram enables users to easily track the message flow path, and through interactive functions such as node expansion / contraction and link highlighting, analyze the interaction relationship between each service in the system in detail, and identify potential performance bottlenecks or abnormal nodes. The time series data analysis module uses time series database technology to accurately record key performance indicators in the message processing process, such as time consumption, throughput and error rate, and supports in-depth analysis in multiple dimensions. This function provides operation and maintenance personnel with a powerful real-time monitoring tool that can timely warn of abnormal indicators and prevent system failures. In addition, the introduction of the pivot table generator allows users to customize data dimensions and aggregation methods according to their own needs, generate interactive data reports in real time, provide data support for subsequent decision analysis, and greatly enhance the flexibility and operability of the log management system.
[0020] Preferably, the time series data analysis module includes: Build an indicator baseline learning engine. This engine automatically learns the normal fluctuation range of business indicators based on historical indicator data through a time series decomposition algorithm to generate a dynamic baseline model. Develop an anomaly pattern recognizer based on the Isolation Forest algorithm, Long Short-Term Memory Network, and other anomaly detection models to perform multi-dimensional anomaly detection on real-time indicator data and calculate anomaly confidence levels. Implement an anomaly propagation path analyzer. After detecting anomaly indicators based on a dynamic baseline model and anomaly confidence, the analyzer traces the anomaly propagation path in reverse based on the service call topology relationship and locates the root cause node of the anomaly.
[0021] The time series data analysis module in the present invention automatically learns and establishes a dynamic baseline model with the support of the indicator baseline learning engine, and can generate the normal fluctuation range of each indicator based on historical data. This mechanism helps the system identify which indicator data belongs to the normal range and which belongs to abnormal fluctuations, which helps to accurately judge the operating status of the system. The development of the abnormal pattern recognizer can perform multi-dimensional anomaly detection on real-time indicator data through a variety of anomaly detection models such as the isolation forest algorithm and the long short-term memory network. The efficient anomaly detection function can capture any abnormal events that may affect the stability of the system in a timely manner, and calculate the confidence level of the anomaly to ensure that the abnormal event can be quickly identified and processed. The introduction of the abnormal propagation path analyzer makes it possible to trace the propagation path of the anomaly based on the service call topology relationship after the anomaly is detected, and locate the root cause node of the anomaly. This function greatly improves the efficiency of system abnormality diagnosis, reduces troubleshooting time, and enhances the reliability and recovery capability of the system.
[0022] Preferably, the abnormal message retransmission control function includes: Build a message resending decision engine that automatically determines whether a message is suitable for resending based on multiple factors, including message error type, error occurrence count, and business impact, and generates resending recommendations. Developed a resend parameter configuration manager that allows users to customize key parameters for resending messages, including the maximum number of retries, retry interval strategy, and message priority adjustment rules; Implement a resend effect evaluation feedback system. After resending a message based on the resend suggestion, the system compares the processing results before and after the resend, evaluates the effectiveness of the resend operation, and feeds back the evaluation results to the visual log management interface.
[0023] The implementation of the exception message retransmission control function in this invention significantly enhances the system's automated response capabilities when faced with exceptions. Through the introduction of a message retransmission decision engine, it can automatically determine whether a message is suitable for retransmission based on multi-dimensional factors (such as error type and number of error occurrences) and generate retransmission recommendations. This automated decision-making mechanism not only reduces the need for manual intervention but also avoids the frequent retransmission of error messages, optimizing system processing efficiency. The development of a retransmission parameter configuration manager allows users to customize retransmission strategies based on specific business needs, including key parameters such as the number of retries and the retry interval, thereby achieving more refined message retransmission control. Finally, a retransmission effectiveness evaluation feedback system provides real-time assessment of the effectiveness of retransmission operations, providing intuitive feedback to users by comparing the processing results before and after retransmission. The evaluation results are displayed in a visual log management interface. This feedback mechanism ensures continuous optimization of the retransmission control strategy, improving the system's resilience and overall service quality in the event of an exception.
[0024] The present invention specifically has the following beneficial effects: (1) By recording the message sending status, sending time, sender ID, complete sending content, globally unique message ID, business module ID and business function name, it can fully reflect the key information in the message sending process. This mechanism provides a reliable data basis for subsequent message tracking and troubleshooting. In particular, the recording of the globally unique message ID and business function name can help track the flow path of each message and quickly locate problems when the system encounters problems. In addition, the recording of information such as the message sending status and time is helpful for analyzing system performance, such as detecting sending delays, frequent failure points, etc., thereby providing a basis for continuous optimization and improvement. When facing the complex environment of distributed systems, this message sending log recording mechanism can reduce the risk of information loss, ensure the integrity and reliability of information, and improve the maintainability and reliability of the system.
[0025] (2) Building a message consumption log recording mechanism at the message consumer node can effectively avoid duplicate consumption and ensure the accuracy of message processing. By performing duplicate verification on the globally unique message identifier at the consumer node, messages can be prevented from being consumed repeatedly, thereby reducing redundant business operations and avoiding data inconsistency problems caused by duplicate processing. This mechanism is particularly suitable for situations where messages may be delivered multiple times in a distributed system, ensuring that each message is processed only once. If a duplicate message is found, the consumption process is terminated immediately to avoid unnecessary waste of resources. At the same time, recording information such as the consumer node identifier and receipt time can provide accurate timestamps and source information for subsequent debugging and monitoring, helping developers quickly locate problems and optimize performance. In addition, the consumer node log can also help monitor the operating status of the system, for example, which consumer node has a slower processing speed and whether there is a queue backlog. This step enhances the transparency and traceability of the message consumption process, laying the foundation for the reliability and stability of the system.
[0026] (3) The mechanism for updating the message consumption table after the business processing logic is executed is mainly used to record the final processing status of each message. The key to this step is to accurately mark whether each message is successfully consumed and to promptly feedback information on handling exceptions. By marking successful messages as "consumed", the progress of message processing can be clearly reflected; and when business processing exceptions occur, recording detailed error information and stack trace data will help developers quickly find the root cause of the problem in subsequent maintenance and optimization. Stack trace data is particularly important. It provides an accurate execution path for debugging and can effectively shorten the time to locate the fault. In a distributed system, business processing involves multiple links and the probability of exceptions is high. Therefore, recording the status code and error details of the exception can provide a basis for the subsequent error handling mechanism, ensuring timely initiation of task retries, error notifications or other error correction measures. This step effectively enhances the system's fault tolerance and improves the accuracy and reliability of message processing.
[0027] (4) Construct a distributed log aggregation indexing system, so that the logs generated by different service nodes can be associated through globally unique message identifiers, forming a cross-node full-link tracking view. The advantage of this mechanism is that it can unify the log data generated in the entire distributed system to form a centralized log query and analysis platform. The cross-node full-link tracking view not only helps developers understand the complete life cycle of each message, from the producer node to the consumer node, and then to the business processing node, but also enables multi-dimensional log analysis. For example, users can track whether a message is delayed at a certain node or an error occurs at a certain node through this view. This mechanism provides strong data support for system troubleshooting and performance optimization, improving the visualization and monitoring capabilities of the system. Periodically collect message generation log tables and consumption log tables, associate them, and index them to quickly summarize and analyze the running status of each link of the system, reducing the workload of manually searching and analyzing logs and greatly improving work efficiency.
[0028] (5) The visual log management interface built based on the message full-link tracking view provides an intuitive and convenient log retrieval and management method. Users can use this interface to perform multi-dimensional condition combination log retrieval, such as accurate query by message identifier, time range, node identifier, etc., to quickly find the required log data. In addition, the log data export function can help users further analyze historical logs or generate reports, providing important data support for system operation and audit. The abnormal message retransmission control function provides an automated exception handling mechanism for the system. When a message consumption fails, it can automatically retransmit the message according to the preset rules, avoiding manual intervention and improving the robustness and efficiency of the system. This visual interface is particularly important for business operation teams, helping to intuitively understand the running status of the system, quickly locate and solve problems, and significantly improve the reliability, stability, and response speed of the service. Through the introduction of this interface, log management and exception handling of the system become more efficient and intelligent, helping to improve the overall operation and maintenance capability. BRIEF DESCRIPTION OF DRAWINGS
[0029] Other features, objects, and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings: Figure 1 The flowchart of the steps of the log management and tracking method based on the message middleware of the present application is shown in the figure. DETAILED DESCRIPTION
[0030] The present application will be further described below in conjunction with the drawings and examples, but it is not limited to the basis of the present application.
[0031] To achieve the above-mentioned purpose, please refer to Figure 1The embodiment of the present invention provides a message middleware-based log management and tracking method, comprising the following steps: S1: In a distributed system where each service node client is connected to the same message middleware service and database service, a message sending log recording mechanism is established at the message producer node. This ensures that when a message producer sends a message to the message middleware, the message sending status, sending time, sender ID, complete sent content, globally unique message ID, business module ID, and business function name are synchronously written into the message generation log table according to a preset data structure. In an embodiment of the present invention, in an e-commerce distributed system, each service node client is uniformly connected to the same message middleware service and database service. Taking the order creation process as an example, when a user submits an order, the order creation service starts working as a message producer node. Before sending the order creation message to the message middleware, the message sending log recording mechanism is started. First, the system generates a globally unique message identifier, such as "20241215142356_ORDER_CREATE_001", which is generated by combining the timestamp, the business function name and the auto-increment number. At the same time, the sender identifier is obtained, that is, the user ID "USER_1234" that triggers the order creation operation, as well as the business module identifier "ORDER_MODULE" and the business function name "Order Creation". Record the current sending time "20241215142356789" accurate to milliseconds, and confirm that the message sending status is "pending". Subsequently, the complete order creation message content, including data such as the order number, product information, and the user's delivery address, is synchronously written to the message generation log table along with the above information according to a preset data structure, such as a table containing fields such as message ID, delivery status, and time. For example, a new row is added to the log table, with each field filled with the corresponding information. This ensures that the relevant data related to the message sending is completely and accurately retained at the same time as the message is sent, providing basic information for subsequent tracking and management.
[0032] S2: Build a message consumption logging mechanism on the message consumer node to implement duplication verification in the message consumption table based on the globally unique message identifier when the message is transmitted to the consumer node through the message middleware. If duplicate records exist, the consumption process is terminated. If not, the message reception status, consumer node identifier, and reception time are recorded, and the message is passed to the actual business processing logic. In the embodiment of the present application, in the e-commerce system, the inventory update service acts as a message consumer node to receive order creation messages from the message middleware. The message consumption log recording mechanism is started immediately when the message arrives, and based on the global unique message identifier "20241215142356_ORDER_CREATE_001" carried by the message, a repetitive check is performed in the message consumption table. The system compares the existing records in the message consumption table one by one, and checks whether there is the same global unique message identifier. If it is found that there is already a record of the identifier, it means that the message has been consumed, and the current consumption process is terminated immediately to avoid repeated processing leading to data errors. If no record with the same identifier is found, the message receiving status is recorded as "received", the consumer node identifier "STOCK_UPDATE_NODE_001" is obtained, and the receiving time accurate to milliseconds "20241215142405321" is recorded. After completing the recording, the order creation message is passed to the actual business processing logic of inventory update, such as deducting the inventory according to the order commodity quantity. In the whole process, the message consumption log recording mechanism strictly controls the uniqueness and accuracy of message consumption, ensures that each message is processed only once, and maintains the consistency of system data.
[0033] S3: After the execution of the business processing logic is completed, the message state in the message consumption table is updated based on the execution result, and if the business processing is successful, the state is marked as consumed, and if an exception occurs, the error state code, detailed error information and stack trace data are recorded, and an error handling task record is generated; In the embodiment of the present application, after the execution of the inventory update business processing logic is completed, the system updates the message state in the message consumption table according to the execution result. If the inventory is successfully deducted according to the order commodity quantity, the message state of the corresponding record in the message consumption table is marked as "consumed", and the completion time "20241215142415678" is recorded. If an exception occurs, such as insufficient inventory to complete the deduction, the system records the error state code "STOCK_INSUFFICIENT_001", the detailed error information "inventory quantity is less than order commodity quantity", and at the same time, captures the complete stack trace data, including the function call level, code line number and other information of the error occurrence. In addition, an error handling task record is generated, including task identifier, error description, expected processing time and other contents, such as task identifier "TASK_20241215142416", and the expected processing time is set to 10 minutes later, so as to track and handle the abnormal situation later, and ensure that the message processing state is consistent with the actual business execution.
[0034] S4: A distributed log aggregation index system is constructed to periodically collect the message generation log table and the message consumption table of each service node, and an associated index is established through the global unique message identifier to form a cross-node message full-link tracking view; In an embodiment of the present invention, a distributed log aggregation indexing system collects information from the message generation log table and message consumption table of each service node in the e-commerce system on a 10-minute basis. Taking order creation messages as an example, the system collects records from the message generation log table of the order creation service node, which contains the globally unique message identifier "20241215142356_ORDER_CREATE_001," the sending status, and the sending time, as well as records from the message consumption table of the inventory update service node, which contains information such as the corresponding receipt status and processing result. Using the globally unique message identifier, the system establishes an associative index. In the index structure, records with the same identifier in the message generation table and the consumption table are associated to form a complete message chain. For example, the entire process of an order creation message, from sending, transmission, to consumption, is linked together, including the sender, receiving node, and processing result. Ultimately, a cross-node, full-chain message tracking view is constructed, allowing operations and maintenance personnel to clearly view the complete flow and processing of each message throughout the distributed system, facilitating system monitoring and troubleshooting.
[0035] S5: A visual log management interface is constructed based on the message full-link tracking view to provide a log retrieval function based on a combination of multi-dimensional conditions, a log data export function, and an abnormal message retransmission control function.
[0036] In an embodiment of the present invention, a visual log management interface is built based on the constructed message full-link tracking view. In the e-commerce system, operation and maintenance personnel can operate through this interface. In terms of log retrieval function, multi-dimensional condition combination retrieval is supported, such as entering the globally unique message identifier "20241215142356_ORDER_CREATE_001", or selecting the business module "order module" and the time interval "December 15, 2024 14:00-15:00", the system quickly filters out the log records that meet the conditions from the full-link tracking view, and displays them in a list form. Each record contains key information about message sending and consumption. The log data export function allows operation and maintenance personnel to export the filtered log records into CSV format files for offline analysis. For abnormal messages, such as messages marked as "processing failed", the interface provides an abnormal message retransmission control function. Operation and maintenance personnel can select an abnormal message and click the resend button. The system will automatically resend the message to the message middleware based on the message type and preset rules, and update the relevant records in the full-link tracking view after resending, and provide real-time feedback on the resending status, realizing visual management and efficient operation and maintenance of the entire life cycle of the message.
[0037] Furthermore, the message sending log recording mechanism includes: In the message producer node, a message preprocessing interceptor is implemented, which generates a composite global unique message identifier containing a timestamp and a node identifier before the message is sent; In the embodiment of the present application, in the message middleware system in the enterprise, the message producer node undertakes the task of sending various business messages. Taking the order creation business as an example, when a new order is generated, the message producer node needs to send the order related information in the form of a message. At this time, the message preprocessing interceptor starts to work. It is like a “security guard” before the message is sent, and intervenes in the processing at the moment when the message is about to be sent. The interceptor first acquires the current time stamp accurate to milliseconds of the system, for example, “20241205143215678”, and at the same time acquires the unique identifier of the message producer node itself, assuming that the node identifier is “NODE_003”. Then, the timestamp and the node identifier are combined according to a fixed format to generate a composite global unique message identifier “20241205143215678_NODE_003”. This identifier has uniqueness, which can ensure that each message in the entire message delivery system has a unique identity marker. Whether it is subsequent tracking and querying of the message, or troubleshooting of the message delivery process, the identifier can serve as a key clue to quickly locate the specific message, providing a basic guarantee for message management and tracking.
[0038] Further, a message content digest generation module is constructed, which extracts features from the original message content to generate a content digest of fixed length, and stores the content digest and the complete message content in the message generation log table synchronously; In this embodiment of the present invention, the message content summary generation module is responsible for processing the original message content. Taking the order creation message as an example, the original message content includes detailed data such as the order number "ORD20241205001," the product name "smart watch," the order amount "1999.00," and the buyer's information. The message content summary generation module uses a hash algorithm to extract features from this original message content. Specifically, the original message content is used as input to the hash algorithm, which generates a fixed-length content summary, such as "5f4dcc3b5aa765d61d8327deb882cf9." This summary is unique; even slight changes to the original message content will result in a different summary. After generating the content summary, the module stores it along with the full original message content in the message generation log table. Each record in the log table contains the generated composite globally unique message identifier, the content summary, and the full message content. For example, for a message identified as "20241205143215678_NODE_003," the log table will record the corresponding content summary "5f4dcc3b5aa765d61d8327deb882cf9" and the complete order creation message content. Storing the content summary allows for subsequent message comparison and integrity verification, eliminating the need to directly process the massive original message content. By simply comparing the summary, the user can quickly determine if the message has been tampered with. This also facilitates rapid retrieval of specific messages within a massive message volume, improving the efficiency of log management and message tracking.
[0039] Furthermore, a sending context association component is implemented, which associates and stores the business operation context metadata, including the identity of the user initiating the operation, the IP address of the operation source, and the operation session identifier, with the message sending record.
[0040] In an embodiment of the present invention, the context association component is used to associate business operation context metadata with message sending records during the message sending process. When a user initiates an order creation operation in the enterprise business system, the context association component starts to collect relevant metadata. The user identity of the operation initiator is "USER_0012", the operation source IP address is "192.168.1.100", and the operation session identifier is "SESSION_20241205001". The component associates and stores these metadata with the message sending records. In the message generation log table, in addition to recording the composite globally unique message identifier, content summary and complete message content, new fields are added to store the user identity of the operation initiator, the operation source IP address and the operation session identifier. For the order creation message described above, the log table records the message ID "20241205143215678_NODE_003," the summary "5f4dcc3b5aa765d61d8327deb882cf9," the complete order message content, the user initiating the operation "USER_0012," the source IP address "192.168.1.100," and the session ID "SESSION_20241205001." This associated storage approach allows for quick identification of the message source, operation initiator, and session context when message delivery anomalies or business operation tracing occur. This provides rich and accurate information for comprehensive troubleshooting and business process analysis, enabling effective tracking and management of messages throughout their lifecycle.
[0041] Furthermore, the message consumption logging mechanism includes: Build a consumption state machine management module, which maintains the lifecycle status of message consumption, including the corresponding lifecycle states of pending consumption, consuming, consumed, consumption failure, and retrying, and automatically updates the status records in the message consumption table based on state transition rules; In the embodiment of the present application, in the enterprise message middleware system, the consumption state machine management module is like a "command center" of message consumption, and the life cycle state of message consumption is comprehensively maintained. Taking the inventory update message as an example, after the message producer node sends the message, the initial state of the message in the message consumption table is marked as "to be consumed". Once the message is acquired by the consumption node and starts to be processed, the consumption state machine management module updates the state of the message in the consumption table to "being consumed" according to the preset state transition rule. If the message processing is successfully completed and the inventory update operation is successfully executed, the module updates the state to "consumed" again and records the specific time of the consumption completion. When the network interruption, business logic error and other conditions cause the message consumption failure, such as the inventory system cannot complete the update due to the data format mismatch, the module changes the state to "consumption failure" and records the failure reason. If the system sets an automatic retry mechanism, the message enters the retry process and the state is updated to "retrying". In the retry process, if the maximum retry number is reached and the message is still not successful, the state is finally kept as "consumption failure"; if the retry is successful, the state is updated to "consumed". In this way, the consumption state machine management module accurately controls the state change of the message in each stage, ensures that the message consumption process is clear and traceable, and facilitates the subsequent tracking and exception troubleshooting of the message processing.
[0042] Further, a consumption idempotency guarantee component is implemented, which is based on a database unique index constraint and an optimistic lock mechanism, and ensures that the same message is only consumed once in a distributed environment; In this embodiment of the present invention, the consumer idempotence guarantee component is key to ensuring reliable message consumption in a distributed environment. Taking inventory update messages as an example, in a distributed environment with multiple consuming nodes running simultaneously, to prevent duplicate consumption of the same message and resulting in inventory data errors, the component uses a database unique index constraint in conjunction with an optimistic locking mechanism. When designing the database table, a unique index is added to the message's unique identifier field (e.g., the composite globally unique message identifier "20241205143215678_NODE_003" generated in the previous step). When a consuming node attempts to consume a message, the database checks this index. If a record with the same identifier exists, the message has already been consumed, and the consumption operation is rejected. Furthermore, an optimistic locking mechanism is introduced, adding a version number field to the message table. When a consuming node retrieves a message for processing, it reads the current version number. For example, if the version number is 1, the update statement includes the current version number condition. This means that the update operation and the version number increment by 1 will only be executed if the message's version number in the database is still 1. If other consuming nodes have already completed consumption and updated their version numbers during this period, the update operation on the current node will fail, thus preventing duplicate consumption. This dual guarantee mechanism ensures that the same message can only be consumed once, regardless of the complexity of the distributed environment, thus ensuring the accuracy and consistency of business data.
[0043] Furthermore, a consumption context injection module is developed, which injects message consumption context information, including consumption node environment variables, message routing path, and message retry times, into the business processing thread context before the message is passed to the business processing logic.
[0044] In the embodiments of the present application, the consumption context injection module undertakes the key information transmission task before the message is transmitted to the business processing logic. Taking the order payment result message as an example, when the message reaches the consumption node, the consumption context injection module starts to work. The module first collects the message consumption context information. The consumption node environment variable contains the hardware configuration information of the current node, such as the memory size and the CPU model. The message routing path records the message from the producer node and the intermediate nodes it has passed through, such as “NODE_003→ROUTER_001→NODE_005”. The message retry number records the number of times the message has been tried to be consumed. If it has been tried to be consumed for 2 times due to network fluctuation, the retry number is 2. Then, the module injects these information into the business processing thread context for processing the message. When the business logic processes the order payment result message, the thread can directly obtain these information from the context. For example, according to the message routing path, it is judged whether the message has passed through the key node, so as to ensure the integrity of the message transmission. According to the retry number, the business processing strategy is adjusted. If the retry number is large, more strict verification logic can be added to avoid data abnormality caused by multiple retries. By accurately injecting the consumption context information into the business processing thread, the business logic can be processed based on more comprehensive information, the accuracy and adaptability of the message processing are improved, and the stable operation of the business process is ensured.
[0045] Further, the consumption state machine management module comprises: A state transition rule engine is constructed, which defines the legal state transition path of message consumption based on the finite state machine model, and records the state change timestamp and the triggering event at each state change; In the embodiment of the present application, in the enterprise message middleware system, the state transition rule engine regulates the state transition path of message consumption based on the finite state machine model, like a precise "traffic control map". Taking the processing of customer order message by the enterprise as an example, the initial state is "to be consumed", and according to the preset rule, only when the message is successfully acquired by the consumption node and starts to be processed, the state can be changed from "to be consumed" to "being consumed". If the message processing is completed and the business logic is executed correctly, such as the order is successfully generated and the inventory deduction is completed, the state is changed from "being consumed" to "consumed". Once the message processing error occurs, such as the order data format error cannot be parsed, the state is changed from "being consumed" to "consumption failure". Each time the state is changed, the engine accurately records the state change timestamp and the triggering event. For example, when the order message changes from "to be consumed" to "being consumed", the timestamp is recorded as "20241210102315", and the triggering event is "the consumption node NODE_007 acquires the message". If the message processing fails due to network failure, the state changes from "being consumed" to "consumption failure", the timestamp is updated to "20241210102508", and the triggering event is recorded as "network connection interruption, order data transmission failure". In this way, the trajectory of message state change is completely preserved, which provides detailed basis for subsequent analysis of message processing process and troubleshooting of abnormality.
[0046] Further, a timeout state detection submodule is implemented, which periodically scans the message records in the "being consumed" state, automatically marks the messages exceeding the preset processing time limit as "processing timeout" state, and triggers the timeout processing process; In the embodiment of the present application, the timeout state detection submodule plays the role of "time guardian" in the message middleware system. Taking the processing of complex supply chain coordination message by the enterprise as an example, the preset message processing time limit is 30 minutes, and the submodule scans the message records in the "being consumed" state once every 5 minutes. Assuming that a message about supplier delivery confirmation has been in the "being consumed" state for 35 minutes, exceeding the preset 30-minute time limit, the timeout state detection submodule immediately and automatically marks it as "processing timeout" state. After the marking is completed, the submodule quickly triggers the preset timeout processing process. First, an alarm notification is sent to the system administrator, detailing the unique identifier of the timeout message, the timeout time, and the current node information; then, the related information of the message is transferred to a special timeout processing queue, waiting for subsequent manual intervention or automatic retry processing. If the subsequent manual judgment determines that the message still needs to be processed, the state of the message is changed to "to be consumed" again, and it is included in the new round of message consumption process; if it is determined that it does not need to be processed, the state of the message is finally marked as "abandoned", ensuring that the message processing process will not be blocked due to the timeout message, and maintaining the efficient operation of the system.
[0047] Further, a state rollback control submodule is developed, which rolls back the message state to the to-be-consumed state when detecting that the message processing has a retryable exception, and calculates the next consumption execution time according to a preset retry strategy.
[0048] In the embodiment of the application, the state rollback control submodule is a "repair master" when the message processing is abnormal. Taking the enterprise financial settlement message processing as an example, when processing a customer order payment settlement message, if a retryable exception occurs due to temporary bank interface failure, the state rollback control submodule quickly intervenes. The submodule first rolls back the state of the message from "in consumption" to "to-be-consumed", and records the rollback timestamp and the exception reason, such as "20241210111020, bank interface response timeout". Then, the submodule calculates the next consumption execution time according to the preset retry strategy. If the preset strategy is fixed interval retry, and the interval time is 10 minutes, then the next consumption execution time of the message is calculated as "20241210112020"; if the exponential backoff retry strategy is used, and the first retry interval is 2 minutes, and the subsequent retry interval is doubled each time, then the corresponding retry time is calculated according to the current retry number. During the waiting period for retry, the submodule continuously monitors the message state, and automatically puts the message into the consumption process again at the scheduled time, so as to ensure that the message can continue to be correctly processed after the exception is recovered, and to guarantee the continuity of the business and the consistency of the data.
[0049] Further, the timeout state detection submodule comprises: A dynamic timeout threshold configuration component is constructed, which dynamically adjusts the message processing timeout threshold in different business scenarios based on the data corresponding to the business type, message size, and historical processing time length through a machine learning algorithm, and determines the timeout message according to the message processing timeout threshold; In an embodiment of the present invention, a dynamic timeout threshold configuration component within an e-commerce enterprise's message middleware system continuously collects and analyzes data to optimize message processing timeout thresholds. Taking three service types—"order creation messages," "logistics shipment confirmation messages," and "customer service inquiry response messages"—as examples, the component records the size of each message type. For example, the average size of order creation messages is 5KB, the average size of logistics shipment confirmation messages is 3KB, and the average size of customer service inquiry response messages is 2KB. It also calculates historical processing times, showing that over the past week, the average processing time for order creation messages was 20 seconds, the average processing time for logistics shipment confirmation messages was 35 seconds, and the average processing time for customer service inquiry response messages was 15 seconds. The component applies a machine learning algorithm to conduct in-depth analysis of this data. For the order creation service, given the high timeliness requirements and moderate message size, the algorithm predicts and dynamically adjusts its timeout threshold to 40 seconds based on historical data. For the logistics shipment confirmation service, due to its complex processing flow and multi-party system interactions, the timeout threshold is set to 60 seconds. For the customer service inquiry response service, which prioritizes timeliness, the timeout threshold is adjusted to 30 seconds. From then on, whenever a new message enters the consumption process, the component compares it with the adjusted timeout threshold based on the business type of the message. If an order creation message remains in the consumption state for more than 40 seconds, it is judged as a timed-out message, ensuring that the timeout judgment standard meets actual business needs.
[0050] Furthermore, a hierarchical processing engine for timeout messages is implemented. This engine adopts different processing strategies for timeout messages based on the business importance of the message, including automatic retry, manual intervention reminder, and business process compensation. In an embodiment of the present invention, a timeout message classification processing engine is used to implement differentiated processing of timeout messages based on the message service importance level. During the e-commerce promotion period, the "order payment success message" is set to the highest importance level, the "user comment review message" is set to the medium importance level, and the "promotional activity notification message" is set to the lower importance level. When a timeout order payment success message is detected, which is directly related to the completion of the transaction and the settlement of funds, the engine immediately starts the automatic retry strategy. According to the exponential backoff rule, the first retry interval is 10 seconds, and the subsequent retry interval is doubled each time, with a maximum of 3 retries. If the retry fails, a high-priority manual intervention reminder is sent to the payment system administrator, with detailed information such as the message unique identifier and the timeout reason, requiring them to investigate and handle it as soon as possible. For the timeout user comment review message, since the impact is relatively small, the engine will first perform an automatic retry with an interval of 30 seconds. If the retry is unsuccessful, a normal manual intervention reminder is sent to the content review team, waiting for manual processing. For timed promotion notification messages, since they are less timely and important, the engine directly executes the business process compensation strategy, skips the message and does not retry it, and instead resends the promotion information through other channels, ensuring that timed messages of different importance levels can be handled reasonably.
[0051] Furthermore, a timeout processing effect evaluation module is developed. This module tracks the business impact indicators after timeout message processing, including business success rate, data consistency deviation rate, and user complaint rate, and optimizes the timeout processing strategy based on the evaluation results.
[0052] In an embodiment of the present invention, a timeout handling effectiveness evaluation module comprehensively tracks and evaluates the business impact of timeout message processing. Taking the processing of timeout order payment success messages as an example, the module continuously monitors the business success rate, calculating the proportion of orders that successfully completed the payment process after processing the timeout message relative to the total number of orders. It also monitors the data consistency deviation rate, comparing the degree of discrepancy between data such as order amounts and transaction status in the payment system and the financial system. It also calculates the user complaint rate, recording the number of user complaints caused by payment issues. If the evaluation finds that after processing the timeout message, the business success rate is only 70%, the data consistency deviation rate reaches 15%, and the user complaint rate has increased significantly, the evaluation module will determine that the current handling strategy is ineffective. Based on this, the module recommends optimizing the timeout handling strategy, such as increasing the number of automatic retries for order payment success messages from 3 to 5, and adjusting the triggering conditions for manual intervention reminders to enable earlier intervention. For other types of timeout messages, the module also continuously optimizes the corresponding handling strategy based on changes in business impact indicators, ensuring that the message middleware system can continuously optimize its handling strategy when faced with timeout messages, minimizing the negative impact on the business.
[0053] Furthermore, the distributed log aggregation indexing system includes: Build a log data collection channel. This channel uses message middleware to implement asynchronous data transmission between service nodes and log aggregation centers, and supports breakpoint-resume transmission based on watermarks. In an embodiment of the present invention, a log data collection channel acts as a "data bridge" between service nodes and a log aggregation center within the vast system architecture of an e-commerce enterprise, enabling asynchronous data transmission based on message-based middleware. Numerous service nodes continuously generate massive amounts of log data, such as an order creation service node generating 2,000 log records per second and a payment service node generating 1,500 log records per second. The message-based middleware acts as a "dispatcher" for data transmission during this process. Service nodes encapsulate log data into message format and send it to the message-based middleware. The message-based middleware does not require service nodes to wait for a response from the log aggregation center; instead, it immediately returns an acknowledgment, allowing the service nodes to continue processing other tasks, thus achieving asynchronous transmission. For example, after an order creation service node sends a log message, it can immediately process the next order creation request without blocking or waiting, significantly improving overall system processing efficiency. Furthermore, the log data collection channel features a watermark-based resumable transmission function. When a brief network outage interrupts the transmission of some log data, the channel records the watermark at the time of the interruption. This mark contains information about the data transmission location, such as whether the 10,000th log record has been transmitted. After the network is restored, the channel restarts transmission of the remaining log data from the watermark, ensuring that no log data is lost. For example, if the payment service node's logs were being transmitted when the network was interrupted, transmission will resume from the interrupted point after recovery, ensuring that the log data reaches the log aggregation center intact, providing a reliable data foundation for subsequent log analysis and troubleshooting.
[0054] Furthermore, a multi-dimensional index construction engine was developed. This engine builds a hybrid index structure of inverted index and B+ tree index based on the globally unique message identifier, business module identifier, time range, and message status dimensions. In an embodiment of the present invention, a multi-dimensional index construction engine acts as an "intelligent search directory" in e-commerce enterprise log management, constructing a hybrid index structure based on multiple key dimensions. For example, each log message has a unique identifier, such as "20241111102345_ORDER_001," which persists throughout the entire message generation and processing process. The engine constructs an inverted index based on this identifier, associating all log records with the same identifier, facilitating quick querying of the complete processing flow of a specific message. Regarding business module identifiers, e-commerce systems include order modules, payment modules, and logistics modules. The engine constructs a B+ tree index for each business module identifier. Taking the order module as an example, the B+ tree index allows efficient querying of all log records generated by the order module within a specific time period. Regarding the time range dimension, the engine slices log data at a granularity of year, month, day, hour, minute, and second, constructing corresponding B+ tree indexes to facilitate rapid retrieval of logs within a specific time period. Similarly, an inverted index is constructed for message statuses, such as "pending," "processing," "completed," and "failed," allowing quick locating of all log records for a given status. By combining an inverted index with a B+ tree index, a hybrid index structure is formed. When searching for all log records of failed processing in the order module for the current day, the engine can quickly locate relevant log data from the hybrid index based on three dimensions: business module identifier (order module), time range (November 11, 2024), and message status (processing failure). This significantly improves log query efficiency and meets enterprises' needs for rapid log retrieval and analysis.
[0055] Furthermore, an incremental update mechanism for index data is implemented. This mechanism is based on database change data capture technology, which can perceive changes in log data tables in real time and synchronize incremental data to the index system.
[0056] In the embodiment of the present application, through the index data incremental update mechanism in the e-commerce enterprise log management system, like "real-time data synchronizer", based on database change data capture technology, real-time perception of log data table changes, when new log data is written into the log data table, for example, a new user returns a log record, the database change data capture technology will immediately detect this change and obtain the detailed information of the new log data, including global unique message identifier, business module identifier, timestamp, message content, etc. After the incremental update mechanism obtains the new data, it quickly synchronizes it to the index system. For the inverted index constructed based on the global unique message identifier, the information of the new log record is added to the index item corresponding to the identifier; for the B+ tree index constructed based on the business module identifier, the new log record is inserted into the appropriate tree node position according to the business module; in the index of time range and message state dimension, the index structure is also updated according to the rules. If the existing log record is modified, such as correcting the amount error in a certain order payment log, the database change data capture technology can also timely perceive the modification operation and synchronize the modified log data to the index system to update the related index item, ensuring that the index data and the log data table always remain consistent. Even in the peak data traffic period, when a large amount of log data is generated and changed every second, the index data incremental update mechanism can accurately synchronize the incremental data to the index system in real time, ensuring that the enterprise can quickly and accurately query the latest log data through the index at any time, providing strong support for business analysis and fault troubleshooting.
[0057] Further, the visual log management interface comprises: A multi-dimensional conditional retrieval engine is constructed, which supports composite queries based on message content keywords, message state combinations, time interval ranges, and producer identity identifiers, and provides query condition saving and reuse functions. In an embodiment of the present invention, a multi-dimensional conditional search engine acts as an "intelligent search manager" in the log management system of an e-commerce enterprise. For example, in a log query scenario after a major promotion, an operator wants to find all messages sent by certain designated producers within a specific time period, with a status of "processing failed" and the message content containing the keyword "payment exception." At this point, the multi-dimensional conditional search engine comes into play, allowing operators to enter multi-dimensional conditions on the interface: enter "payment exception" in the message content keyword field; check "processing failed" in the message status combo box; set the time interval range selection box to "November 11, 2024 00:00:00-November 11, 2024 23:59:59"; and fill in "payment service node A" and "payment service node B" in the producer identity column. After receiving these conditions, the engine quickly locates the log data that meets the conditions based on the previously constructed hybrid index structure. It first obtains all "processing failure" message records from the inverted index built based on the message status, and then combines the B+ tree index built based on the time range to filter out records within the specified time period, and then further filters according to the producer identity index, and finally matches the message containing the "payment exception" keyword in the message content index. Through this multi-dimensional compound query method, the engine can quickly lock the target message in the massive log data and present the query results to the operation personnel. In addition, the engine also provides query condition saving and reuse functions. The operation personnel can name and save the above complex query conditions as "Double Eleven Payment Anomaly Query". When a similar query is needed next time, the saved conditions can be directly called without re-entering, which greatly improves the query efficiency and facilitates enterprise personnel to quickly obtain the required log information for business analysis and problem troubleshooting.
[0058] Furthermore, we developed a log data visualization component. This component generates visual analysis charts corresponding to message flow topology diagrams, processing time heat maps, and business throughput trend charts based on full-link tracking data. In this embodiment of the present invention, the log data visualization component serves as a "visualization assistant" for log analysis for e-commerce companies. Based on end-to-end tracking data, the component begins generating various visualization charts after the promotion ends. The message flow topology diagram graphically illustrates the complete message flow from producer nodes (such as the order creation service node), through the message middleware, and to consumer nodes (such as the inventory deduction service node and the payment processing service node), based on the message transmission path and sequence between different service nodes. The diagram uses nodes of different colors and shapes to represent different services, and connecting lines to illustrate message transmission directions. This allows operators to intuitively visualize the flow of messages throughout the system and quickly identify key nodes and potential bottlenecks. The processing time heat map uses time as the horizontal axis and service nodes as the vertical axis, with color shading used to represent the message processing time of each service node. During the promotion, darker areas indicate nodes and time periods with longer processing times. For example, the color of the payment processing service node became significantly darker between 3:00 PM and 5:00 PM that day, indicating increased message processing time during this period and possible performance issues. By observing the heat map, operators can quickly locate nodes and time periods with extended processing times, providing a visual basis for optimizing system performance. The business throughput trend chart shows the changing trend in the number of business messages processed by the system over different time periods. On a given day, business throughput gradually increases from early morning, reaches a peak at noon, then decreases, and reaches a new peak in the evening. This chart allows operators to clearly understand the changing patterns of business traffic, rationally allocate system resources, and proactively respond to traffic peaks. These visualizations help enterprise personnel more intuitively and comprehensively understand the business operations underlying log data, assisting in decision-making.
[0059] Furthermore, an abnormal message intelligent positioning module is implemented, which automatically identifies suspicious message patterns based on preset abnormal pattern matching rules and provides abnormal root cause analysis suggestions.
[0060] In an embodiment of the present invention, the intelligent abnormal message location module acts as an "anomaly detector" in e-commerce enterprise log management. It pre-configures various anomaly pattern matching rules. For example, suspicious message patterns are defined as messages repeatedly containing the keyword "payment timeout" and consistently displaying the status "processing failure," or messages from the same producer processing a large number of messages exceeding a threshold in a short period of time. During a promotional event, the module continuously scans and analyzes log data. If it detects that a payment service node generates 50 messages within 10 minutes, all containing the keyword "payment interface connection interrupted" and displaying the status "processing failure," the module immediately identifies this message pattern as matching the pre-defined anomaly pattern and marks these messages as suspicious. The module then analyzes the root cause of the anomalies based on relevant information in the log data, such as the message flow path and the processing status of upstream and downstream service nodes. This analysis reveals that these abnormal messages are due to unstable connections caused by excessive traffic on the third-party payment interface during the promotional event. Based on this, the module generates root cause analysis recommendations, prompting operations and maintenance personnel to contact the third-party payment interface provider to optimize interface performance and implement an interface connection retry mechanism and timeout handling policy within the enterprise's internal system. In this way, the abnormal message intelligent positioning module can quickly detect abnormal messages and provide targeted analysis suggestions, helping enterprises to solve problems in a timely manner and ensure stable system operation.
[0061] Furthermore, the log data visualization component includes: Build a dynamic topology generation engine that generates service call topology diagrams in real time based on message flow paths, and supports interactive functions such as node expansion / contraction, link highlighting, and performance indicator overlay display. In an embodiment of the present invention, a dynamic topology generation engine acts as a "system cartographer" within the vast business systems of an e-commerce platform, generating a service call topology map in real time based on message flow paths. For example, the business process of a user placing an order for a product begins with the user submitting the order, and messages flow sequentially through multiple nodes, including the order creation service, inventory query service, payment processing service, and logistics dispatch service. The dynamic topology generation engine captures these message flow paths in real time, presenting each service node in the topology map with a specific shape and color. For example, the order creation service is represented by a blue rectangle, and the payment processing service is represented by a green circle. Arrows clearly indicate the message flow direction, thus constructing a complete service call topology map. The engine also provides rich interactive features. When operators want to drill down to the details of a service node, they can click on the node to expand it, which displays the node's specific configuration information, the number of messages currently being processed, and so on. Collapse makes the topology map more concise and easier to view the overall architecture. If performance issues are detected in the payment processing service, the link highlighting feature can highlight message links related to the payment processing service in red, allowing quick identification of affected upstream and downstream services. In addition, performance indicators can be superimposed and displayed on the topology map, such as marking the message processing time of each service node in digital form next to the node, or using arrows of different thicknesses to indicate the message throughput size, to help operators intuitively grasp the system operation status and quickly identify performance bottlenecks and potential risk points.
[0062] Furthermore, we developed a time series data analysis module. This module, based on time series database technology, performs multi-dimensional analysis on message processing time, throughput, and error rate indicators, and provides automatic warning functions for abnormal indicators. In the embodiment of the present application, the time series data analysis module analyzes the key indicators of the e-commerce platform message processing based on time series database technology. Taking the data during the big promotion event as an example, the module continuously collects message processing time consumption, throughput, error rate and other indicator data. For message processing time consumption, the time series database records the time from the entry of each message into the system to the completion of processing at a second level granularity; the throughput records the number of messages processed by the system per second; the error rate calculates the proportion of failed messages per second in the total number of messages. The module performs multi-dimensional analysis on these indicators. By comparing the message processing time consumption in different time periods, it is found that the processing time consumption increases significantly at 10 am and 8 pm every day; by analyzing the throughput data, it is found that the throughput reaches a peak in the first two hours after the event starts and then gradually decreases; by studying the error rate indicator, it is found that the error rate of payment-related messages increases during the peak of the event. At the same time, the module has an automatic abnormal indicator warning function, and the message processing time consumption exceeding 500 milliseconds, the throughput being lower than the normal level by 30%, and the error rate being higher than 5% are set as abnormal thresholds. When the message processing time consumption of the payment processing service exceeds 500 milliseconds for 10 consecutive times at a certain time, the module immediately issues a warning, notifies the operation and maintenance personnel in the form of a pop-up window and an email, and displays the relevant indicator data in red on the system interface, prompting the operation and maintenance personnel to promptly troubleshoot problems and ensure the stable operation of the system.
[0063] Further, a pivot table generator is implemented, which supports user-defined data dimensions and aggregation methods, generates interactive pivot reports in real time, and supports data drill-down analysis function.
[0064] In the embodiment of the present application, the pivot table generator provides e-commerce platform operators with a flexible data exploration tool. When analyzing order message data during the event, the operator can customize data dimensions and aggregation methods through the pivot table generator. For example, selecting “product category” and “order time (by hour)” as data dimensions and “order quantity” as the aggregation indicator, the generated interactive pivot report will clearly show the order quantity of different product categories in each hour. The operator can also further operate the report, such as changing the aggregation method to “total order amount”, and the report will be refreshed to show the total order amount of different product categories in each hour. If it is found that the order quantity or amount of a certain product category is abnormally high in a certain time period, the data drill-down analysis function can be used to click on the data cell, and the report will display the specific order message details of the product category in that time period, including order number, order user, payment amount and other detailed information. Through this flexible customization and drill-down analysis, the operator can deeply analyze the data from different angles and find business rules and potential problems, such as the hot-selling trend of certain products in a certain time period or the specific situation of abnormal orders, providing strong data support for formulating marketing strategies and optimizing business processes.
[0065] Furthermore, the time series data analysis module includes: Build an indicator baseline learning engine. This engine automatically learns the normal fluctuation range of business indicators based on historical indicator data through a time series decomposition algorithm to generate a dynamic baseline model. In an embodiment of the present invention, a metric baseline learning engine acts as a "data pattern explorer" within the e-commerce platform's message middleware system. Based on historical data from the three months preceding a major promotion, the engine conducts in-depth analysis of business metrics such as message processing time, throughput, and error rates. The engine applies a time series decomposition algorithm to decompose message processing time into components including long-term trends, seasonal variations, and random fluctuations. This analysis reveals that from 2:00 AM to 5:00 AM daily, due to low system traffic, message processing time remains generally low with minimal fluctuations. However, from 8:00 PM to 10:00 PM on Fridays, during peak shopping season, processing time increases significantly and exhibits regular fluctuations. Based on these analysis results, the engine generates a dynamic baseline model for the message processing time metric, defining the normal fluctuation range for this metric at different time points. For example, processing time should be between 100 and 200 milliseconds from 2:00 AM to 5:00 AM daily, while fluctuations from 300 to 500 milliseconds are permitted from 8:00 PM to 10:00 PM on Fridays. Similarly, for throughput metrics, the engine analyzes the different variations between weekdays and weekends, as well as the trends during the promotional pre-event period, promotional activity period, and promotional post-event period, to determine its dynamic baseline range. Similarly, for error rate metrics, the engine combines historical data with normal error patterns to generate a corresponding dynamic baseline model. These dynamic baseline models provide an accurate reference standard for subsequent anomaly detection.
[0066] Furthermore, an abnormal pattern recognizer was developed. This recognizer performs multi-dimensional anomaly detection on real-time indicator data and calculates the anomaly confidence level based on the isolation forest algorithm, long short-term memory network and other anomaly detection models. In an embodiment of the present invention, the anomaly pattern recognizer, acting as an "anomaly scout," utilizes multiple anomaly detection models, including the Isolation Forest algorithm and Long Short-Term Memory (LSTM) networks, to conduct multi-dimensional inspections of real-time metrics data from the e-commerce platform's message middleware. On the day of an event, as real-time message processing time, throughput, and error rate data continuously accumulates, the Isolation Forest algorithm, based on data distribution, quickly identifies data that is far from the majority of data points in the feature space and identifies these as potential anomalies. For example, if message processing time suddenly reaches 1500 milliseconds at a given moment, far exceeding the normal range set by the dynamic baseline model and isolated within the current data distribution, the Isolation Forest algorithm will initially flag this data as an anomaly. Simultaneously, the Long Short-Term Memory (LSTM) network model, leveraging its ability to memorize and predict time series data, predicts future metrics based on patterns learned from historical data. If the actual real-time data deviates significantly from the predicted value, such as predicting a throughput of 5000 messages / second at a given moment but only reaching 1000 messages / second, the LSTM model will also identify this as an anomaly. The anomaly pattern recognizer integrates the detection results of multiple models to calculate an anomaly confidence level. If both the Isolation Forest algorithm and the LSTM model detect a data anomaly, and other related indicators also show similar anomaly trends, the confidence level for the anomaly is determined to be high, such as 90%. If only a single model detects an anomaly, and other indicators show no obvious anomalies, the confidence level is relatively low. This multi-model fusion and confidence calculation approach ensures the accuracy and reliability of anomaly detection.
[0067] Furthermore, an anomaly propagation path analyzer is implemented. After detecting anomaly indicators based on a dynamic baseline model and anomaly confidence, the analyzer traces the anomaly propagation path in reverse based on the service call topology relationship and locates the root cause node of the anomaly.
[0068] In this embodiment of the present invention, when the anomaly pattern recognizer detects anomaly indicators based on a dynamic baseline model and anomaly confidence, the anomaly propagation path analyzer immediately activates, becoming an "anomaly tracking expert." Taking the example of an abnormally high message processing time for the payment service during an activity, the anomaly propagation path analyzer first traces the anomaly propagation path backwards from the payment service node where the anomaly occurred, based on the generated service call topology. An inspection of the topology reveals that the payment service relies on the user authentication service and the bank interface service, and interacts with the order creation service through messages. The analyzer further examines the metrics of the related services and finds that the response time of the user authentication service had increased slightly before the payment service anomaly, while the error rate of the bank interface service also increased. Continuing to trace back along the topology, the anomaly propagation path analyzer determined that the anomaly in the bank interface service was caused by excessive load on the third-party bank system during the activity, while the anomaly in the user authentication service was caused by a temporary performance bottleneck caused by a large number of users logging in simultaneously. After step-by-step analysis, the anomaly propagation path analyzer successfully located the root causes of the anomaly: the third-party bank system and the user authentication service. At the same time, the analyzer generates a detailed exception propagation path report, showing the entire propagation process from the root cause node to the payment service anomaly, providing clear problem-solving clues for operation and maintenance personnel, helping them to quickly solve problems and ensure the stable operation of the e-commerce platform during the promotion period.
[0069] Furthermore, the abnormal message retransmission control function includes: Build a message resending decision engine that automatically determines whether a message is suitable for resending based on multiple factors, including message error type, error occurrence count, and business impact, and generates resending recommendations. In the embodiment of the present application, in the process of running the message middleware system of the e-commerce platform, the message retransmission decision engine acts as a "smart referee" to accurately judge whether the message is suitable for retransmission based on multiple factors. For example, taking the order payment message during the big promotion period as an example, when a payment message fails to be processed, the engine first obtains the message error type. If it shows "temporary network interruption", it is a recoverable error. Then, the number of occurrences of this type of error in the past 10 minutes is counted. If there are only 2 times, it is at a low frequency. Then, the business impact degree is evaluated. Since this order involves a large transaction, the business impact degree is determined to be high. Based on the three factors, the engine automatically determines that the message is suitable for retransmission according to the preset decision rule, and generates a retransmission suggestion. On the contrary, if the message error type is "permanent invalidation of payment account information", even if the error occurs less frequently, it is a non-recoverable error and will have a fundamental impact on the business. Therefore, the engine determines that the message is not suitable for retransmission, and generates a suggestion not to retransmit, with detailed explanation. For other types of messages, such as logistics delivery notification messages, if the error type is "target address format error", the error occurs frequently, and the business impact is low. The engine also provides accurate retransmission decision suggestions for each failed message based on multiple factors, ensuring the rationality and efficiency of message processing.
[0070] Further, a retransmission parameter configuration manager is developed, which supports user customization of key parameters of retransmitted messages, including the upper limit of the number of retries, the retry interval strategy, and the message priority adjustment rule. In the embodiment of the present application, the retransmission parameter configuration manager provides e-commerce platform operation and maintenance personnel with a flexible message retransmission parameter setting tool. Before the big promotion, the operation and maintenance personnel customize the configuration through the manager according to the characteristics of different types of messages. For order creation messages, considering the high timeliness requirement, the upper limit of the number of retries is set to 3 to avoid occupying system resources due to excessive retries. The exponential backoff retry interval strategy is adopted, with a first retry interval of 5 seconds and a doubling interval for each subsequent retry, to balance the timeliness of retries and system load. The message priority adjustment rule is formulated to promote the priority of the order creation message to above the ordinary message when it is retransmitted, to ensure priority processing. For payment settlement messages, due to their high importance, the upper limit of the number of retries is set to 5 to give more retry opportunities. The fixed interval of 10 seconds is adopted for the retry interval strategy to ensure a stable retry rhythm. The message priority adjustment rule is set to directly promote to the highest priority when retransmitted to ensure that the messages related to fund transactions are quickly processed. For logistics status update messages, due to the relatively low timeliness, the upper limit of the number of retries is set to 2, and the retry interval strategy is random interval of 5-15 seconds to reduce the concentrated pressure on the system. The message priority remains unchanged. Through such customized configuration, the individual needs of message retransmission in different business scenarios are met.
[0071] Furthermore, a resending effect evaluation feedback system is implemented. After resending a message based on the resending suggestion, the system compares the processing results before and after the resending, evaluates the effectiveness of the resending operation, and feeds back the evaluation results to the visual log management interface.
[0072] In an embodiment of the present invention, after the message is resent on the e-commerce platform according to the recommendation of the message resending decision engine, the resending effect evaluation feedback system starts working. Taking an order payment message that failed to be processed and resent due to network fluctuations as an example, the system first records the detailed information of the message processing failure before resending, including the error type, failure time, amount involved, etc. After the message is successfully resent, the processing result after resending is obtained, such as the payment success time, transaction status, etc., and the processing results before and after resending are compared to calculate various evaluation indicators. If the payment is successful after resending and the transaction data is accurate, the resending operation is evaluated to be effective, and the resending success rate is recorded as 100%. At the same time, the processing time difference before and after resending is calculated to evaluate the impact of resending on processing efficiency. If it still fails after resending multiple times, the system will analyze the failure cause in detail, record the number of failures, the final error message, etc., evaluate the resending operation as invalid, and feedback the relevant data and analysis results to the visual log management interface. On the visual log management interface, operations personnel can directly see the evaluation results of each resent message, with green markings indicating successful resends and good results, and red markings indicating failed resends. Detailed evaluation data and analysis reports are also provided. This allows operations personnel to quickly understand the effectiveness of message resends, providing a strong basis for subsequent optimization of message resend strategies and system operations.
[0073] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.
[0074] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A log management and tracking method based on message middleware, characterized by: The following steps are involved: In a distributed system where each service node client is connected to the same message middleware service and database service, a message sending log recording mechanism is built on the message producer node. This ensures that when a message producer sends a message to the message middleware, the message sending status, sending time, sender ID, complete sending content, globally unique message ID, business module ID, and business function name are synchronously written into the message generation log table according to a preset data structure. A message consumption logging mechanism is built on the message consumer node to implement duplication checking in the message consumption table based on the globally unique message identifier when the message is transmitted to the consumer node through the message middleware. If duplicate records exist, the consumption process is terminated. If not, the message reception status, consumer node identifier, and reception time are recorded, and the message is passed to the actual business processing logic. After the business processing logic is executed, the message status in the message consumption table is updated based on the execution result. If the business processing is successful, the status is marked as consumed. If an exception occurs, the error status code, detailed error information and stack trace data are recorded, and an error handling task record is generated at the same time. Build a distributed log aggregation indexing system to periodically collect messages from each service node to generate log tables and message consumption tables, and establish associated indexes through globally unique message identifiers to form a full-link message tracking view across nodes; A visual log management interface is constructed based on the message full-link tracking view to provide a log retrieval function based on a combination of multi-dimensional conditions, a log data export function, and an abnormal message retransmission control function.
2. The message middleware-based log management and tracking method according to claim 1 is characterized in that: The message sending log recording mechanism includes: Implement a message pre-processing interceptor on the message producer node. This interceptor generates a composite globally unique message identifier containing a timestamp and a node identifier before the message is sent. Construct a message content summary generation module, which extracts features from the original message content to generate a fixed-length content summary and stores the content summary and the full message content in the message generation log table synchronously; Implement a sending context association component that associates and stores business operation context metadata, including the identity of the user initiating the operation, the IP address of the operation source, and the operation session identifier, with the message sending record.
3. The message middleware-based log management and tracking method according to claim 1, wherein: The message consumption logging mechanism includes: Build a consumption state machine management module, which maintains the lifecycle status of message consumption, including the corresponding lifecycle states of pending consumption, consuming, consumed, consumption failure, and retrying, and automatically updates the status records in the message consumption table based on state transition rules; Implement a consumer idempotence guarantee component. This component uses database unique index constraints and optimistic locking mechanisms to ensure that the same message is consumed only once in a distributed environment. Develop a consumption context injection module, which injects message consumption context information, including consumption node environment variables, message routing path, and message retry count, into the business processing thread context before the message is passed to the business processing logic.
4. The message middleware-based log management and tracking method according to claim 3 is characterized by: The consumption state machine management module includes: Build a state transition rule engine that defines the legal state transition path for message consumption based on a finite state machine model and records the state change timestamp and trigger event at each state change. Implement a timeout status detection submodule, which periodically scans the message records in consumption, automatically marks the messages that exceed the preset processing time limit as processing timeout status, and triggers the timeout processing process; Develop a status rollback control submodule. When a retryable exception is detected in message processing, this submodule rolls back the message status to the pending consumption status and calculates the next consumption execution time according to the preset retry strategy.
5. The message middleware-based log management and tracking method according to claim 4 is characterized in that: The timeout state detection submodule includes: Build a dynamic timeout threshold configuration component. This component uses a machine learning algorithm to dynamically adjust the message processing timeout threshold for different business scenarios based on data corresponding to business type, message size, and historical processing time. It also determines the timeout message based on the message processing timeout threshold. Implement a hierarchical processing engine for timeout messages. This engine adopts different processing strategies for timeout messages based on the business importance of the message, including automatic retry, manual intervention reminder, and business process compensation. Develop a timeout processing effect evaluation module, which tracks the business impact indicators after timeout message processing, including business success rate, data consistency deviation rate, and user complaint rate, and optimizes the timeout processing strategy based on the evaluation results.
6. The message middleware-based log management and tracking method according to claim 1, wherein: The distributed log aggregation indexing system includes: Build a log data collection channel. This channel uses message middleware to implement asynchronous data transmission between service nodes and log aggregation centers, and supports breakpoint-resume transmission based on watermarks. Develop a multi-dimensional index construction engine that builds a hybrid index structure of inverted index and B+ tree index based on globally unique message identifier, business module identifier, time range, and message status dimensions; Implement an incremental update mechanism for index data. This mechanism, based on database change data capture technology, senses changes in log data tables in real time and synchronizes incremental data to the index system.
7. The message middleware-based log management and tracking method according to claim 1, wherein: The visual log management interface includes: Build a multi-dimensional conditional search engine that supports compound queries based on message content keywords, message status combinations, time intervals, and producer identification, and provides query condition saving and reuse functions; Develop a log data visualization component that generates visual analysis charts corresponding to message flow topology diagrams, processing time heat maps, and business throughput trend charts based on full-link tracking data; Implement an intelligent abnormal message positioning module, which automatically identifies suspicious message patterns based on preset abnormal pattern matching rules and provides abnormal root cause analysis suggestions.
8. The message middleware-based log management and tracking method according to claim 7, wherein: The log data visualization component includes: Build a dynamic topology generation engine that generates service call topology diagrams in real time based on message flow paths, and supports interactive functions such as node expansion / contraction, link highlighting, and performance indicator overlay display. Develop a time series data analysis module. This module, based on time series database technology, performs multi-dimensional analysis of message processing time, throughput, and error rate indicators, and provides automatic early warning functions for abnormal indicators. Implement a pivot table generator that supports user-defined data dimensions and aggregation methods, generates interactive pivot reports in real time, and supports data drill-down analysis.
9. The message middleware-based log management and tracking method according to claim 8, wherein: The time series data analysis module includes: Build an indicator baseline learning engine. This engine automatically learns the normal fluctuation range of business indicators based on historical indicator data through a time series decomposition algorithm to generate a dynamic baseline model. Develop an anomaly pattern recognizer based on the Isolation Forest algorithm, Long Short-Term Memory Network, and other anomaly detection models to perform multi-dimensional anomaly detection on real-time indicator data and calculate anomaly confidence levels. Implement an anomaly propagation path analyzer. After detecting anomaly indicators based on a dynamic baseline model and anomaly confidence, the analyzer traces the anomaly propagation path in reverse based on the service call topology relationship and locates the root cause node of the anomaly.
10. The message middleware-based log management and tracking method according to claim 1, wherein: The abnormal message retransmission control function includes: Build a message resending decision engine that automatically determines whether a message is suitable for resending based on multiple factors, including message error type, error occurrence count, and business impact, and generates resending recommendations. Developed a resend parameter configuration manager that allows users to customize key parameters for resending messages, including the maximum number of retries, retry interval strategy, and message priority adjustment rules; Implement a resend effect evaluation feedback system. After resending a message based on the resend suggestion, the system compares the processing results before and after the resend, evaluates the effectiveness of the resend operation, and feeds back the evaluation results to the visual log management interface.
Citation Information
Patent Citations
Method and system e for improving message tracking capability of message-oriented middleware and monitoring modul
CN110019001A
Log analysis method and system based on OpenStack cloud computing
CN116755992A
Visual cross-system message link tracking method and system based on log, medium and equipment
CN118838724A
Transaction processing system and method based on message queue
CN119003102A
Software application performance optimization method based on big data analysis
CN119917390A
Cited By
Message sending method and device based on message middleware, equipment and storage medium
CN121567671A
Message-oriented middleware high-availability data transmission method and system
CN121603544A