Serverless architecture real-time log processing method and system

By introducing log routing tables and message queue partitioning routing technology into the serverless architecture, efficient log processing was achieved, solving the problems of real-time performance and low resource utilization, and improving the system's parallel processing capabilities and storage efficiency.

CN120804048AInactive Publication Date: 2025-10-17刘宇
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510874579.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a serverless architecture, log processing suffers from poor real-time performance and low resource utilization, making it difficult to meet the needs for real-time log observation and efficient resource utilization.

Method used

It employs a log routing table based on user session and function instance ID, combined with message queue partitioning routing technology, to achieve precise log filtering and push through session-level log flow control and WebSocket push, supporting efficient partition-level log subscription and parallel processing.

Benefits of technology

It significantly improves the real-time performance and resource utilization of log processing, achieves millisecond-level log visibility, optimizes storage costs, and enhances development and testing efficiency and problem-solving capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005470705220000071
    Figure BDA0005470705220000071
  • Figure BDA0005470705220000081
    Figure BDA0005470705220000081
  • Figure BDA0005470705220000091
    Figure BDA0005470705220000091
Patent Text Reader

Abstract

The invention provides a Serverless architecture real-time log processing method and a Serverless architecture real-time log processing system, which are used for solving the problems of poor log real-time performance and low resource utilization rate in a Serverless environment. The method is characterized in that a log routing table based on user sessions and function instance IDs is established, and accurate log screening and pushing are achieved through the session-level log flow control technology; meanwhile, a message queue partition routing technology based on a function ID is adopted, it is ensured that logs of the same function are routed to a fixed partition, and efficient partition-level log subscription and parallel processing are supported. The system comprises a log routing module, a log screening module, a message queue module, a log consumption module and a WebSocket pushing module. The method has the advantages that the real-time performance and parallelism of log processing are remarkably improved, the system resource utilization rate is optimized, and powerful support is provided for real-time log analysis and problem diagnosis under the Serverless architecture.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer system log processing, in particular to a real-time log processing method and system applied to a Serverless architecture environment. BACKGROUND

[0002] Log processing has important applications in the fields of computer system operation and maintenance and problem diagnosis. With the rapid development of cloud computing technology, Serverless architecture is widely adopted in various applications due to its high efficiency and flexibility. In the Serverless environment, effective log processing is crucial for system monitoring, performance optimization and fault diagnosis, directly affecting development efficiency and service quality.

[0003] Currently, log processing under Serverless architecture mainly adopts a centralized log storage and analysis scheme. This scheme usually collects logs generated by function instances into a centralized storage system, such as Elasticsearch or a log database, and then analyzes and retrieves them through batch processing or querying. Another common method is to use a distributed stream processing system, such as Apache Kafka or Amazon Kinesis, to collect and process log data in real time to support near-real-time log analysis.

[0004] However, these schemes have some significant problems in the Serverless environment. First, the real-time performance of log queries is poor, often with large delays, which is not conducive to real-time problem troubleshooting by developers. Second, it is difficult to achieve real-time observation of logs during development and testing, affecting development efficiency. In addition, the persistent storage of full logs results in high storage costs, especially in large-scale application scenarios, which is even more prominent.

[0005] To solve the above problems, some research has proposed log processing methods based on in-memory caching, trying to improve the speed of log retrieval. Although this method improves query performance to some extent, it still cannot meet the needs of real-time log pushing, and there are still deficiencies in resource utilization efficiency. Some other attempts have introduced lightweight log stream processing components, but they face challenges in accurately controlling log flow and ensuring processing sequence.

[0006] Therefore, there is an urgent need for a new method that can achieve efficient real-time log processing under Serverless architecture, which can meet the needs of real-time log observation, optimize resource utilization, and maintain the functional integrity of the original log system. Developing such a log processing scheme with real-time performance, efficiency and flexibility has become an important research direction in the field of Serverless technology. SUMMARY

[0007] The main purpose of the present application is to solve the problems of poor real-time performance and low resource utilization in log processing under the Serverless architecture. Specifically, the present application aims to provide a log processing method and system that can realize real-time log pushing, accurately control log flow, and improve system parallel processing capability. In addition, the present application also aims to optimize log storage cost, improve development and testing efficiency, and provide support for rapid diagnosis of online problems.

[0008] To achieve the above-mentioned purpose, the present application provides a Serverless architecture real-time log processing method and system, characterized by: establishing a log routing table based on user session and function instance ID, realizing accurate log filtering and pushing through session-level log flow control technology; at the same time, using function ID-based message queue partition routing technology to ensure that the logs of the same function are routed to a fixed partition, supporting efficient partition-level log subscription and parallel processing. In some embodiments, the method and system can be used in parallel with existing log persistence schemes to meet the needs of different scenarios.

[0009] Specifically, the method of the present application includes the steps of establishing a log routing table, log filtering, message queue partitioning, log consumption, and WebSocket pushing. The structure of the log routing table includes two parts: key and value, where the key is the combination of user session ID and function instance ID, and the value is the corresponding WebSocket connection object. This structure design ensures accurate directional transmission of logs.

[0010] Further, the system of the present application includes a log routing module, a log filtering module, a message queue module, a log consumption module, and a WebSocket pushing module. These modules work together to realize the whole process management from log generation to final pushing. Among them, the message queue module is responsible for partition storage based on function ID, ensuring that the logs of the same function are routed to a fixed partition.

[0011] Preferably, the present application also includes a connection monitoring module for monitoring the state of WebSocket connection and automatically cleaning related items in the log routing table when a disconnection is detected to optimize resource utilization.

[0012] Optionally, the log consumption module of the present application supports subscribing to the message queue partition corresponding to a specific function ID, realizing parallel log consumption at the partition level, and thus improving the processing efficiency of the system.

[0013] In one embodiment, during the log sending process, the present application first calculates the target partition according to the function ID, and then sends the log message to the specified partition. This method ensures that the logs of the same function are always routed to the same partition, guaranteeing the orderliness and integrity of the logs.

[0014] In some embodiments, the present application can be used in combination with traditional log persistence storage solutions, meeting the needs of real-time log processing and ensuring the possibility of long-term log analysis.

[0015] In addition, the present application can also dynamically adjust the log processing strategy according to different application scenarios, such as increasing log detail in development and testing stage, and optimizing resource occupation in production environment.

[0016] In other embodiments, the present application can be integrated into existing Serverless platforms to provide developers with one-stop log management and analysis tools.

[0017] In a preferred embodiment, the present application realizes real-time bidirectional communication between the client and the server through WebSocket long connection, ensuring that log data can be pushed to the client with minimal delay. At the same time, by using the persistence feature of the message queue, the log reliability is ensured when the network fluctuates or the client disconnects.

[0018] In another preferred embodiment, the present application adopts a distributed architecture design, supports horizontal expansion, and can dynamically adjust the number of processing nodes according to system load, ensuring high performance and high availability in large-scale Serverless application scenarios.

[0019] By adopting the above-mentioned scheme, the present application has the following beneficial effects: 1. Significantly improves the real-time performance of log processing, realizes millisecond-level log visibility, and effectively supports real-time problem diagnosis; 2. Through precise log flow control technology, the system resource utilization is greatly improved, and invalid log transmission is reduced; 3. Based on the message queue partition routing technology of function ID, the high parallelism and sequence of log processing are ensured; 4. Optimizes the log storage cost, reduces the long-term storage pressure through real-time processing and selective persistence; 5. Provides flexible log subscription and consumption mechanism, greatly improves the development and testing efficiency and problem troubleshooting ability.

[0020] In summary, the present application provides an innovative Serverless architecture real-time log processing scheme, which not only solves the key problems in the prior art, but also provides strong support for the development, testing and operation and maintenance of Serverless applications, and has important significance for improving the practicality and reliability of Serverless technology. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0022] Figure 1 is the overall architecture diagram of the real-time log processing system of the Serverless architecture.

[0023] Figure 2 is the system workflow diagram.

[0024] Figure 3 is the core component diagram of the session-level log flow control technology.

[0025] Figure 4 is the core component diagram of the message queue partition routing technology.

[0026] Figure 5 is the core component diagram of the log consumption and WebSocket push technology.

[0027] Figure 6 is the workflow diagram of log consumption and WebSocket push.

[0028] Figure 7 is the core component diagram of the connection monitoring and resource cleaning technology.

[0029] Figure 8 is the workflow diagram of connection monitoring and resource cleaning.

[0030] Figure 9 is the core component diagram of the partition-level parallel log consumption technology.

[0031] Figure 10 is the workflow diagram of partition-level parallel log consumption.

[0032] Figure 11 is the core component diagram of integration with existing log persistence systems.

[0033] Figure 12 is the workflow diagram of integration with existing log persistence systems. DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0035] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.

[0036] Embodiment 1: Overall architecture of a serverless architecture real-time log processing system.

[0037] The overall architecture of the serverless architecture real-time log processing system provided by the present application is shown in Figure 1 The system mainly includes a log routing module 101, a log filtering module 102, a message queue module 103, a log consumption module 104 and a WebSocket push module 105. These modules work together to realize the whole process management from log generation to final push, forming an efficient and real-time log processing system.

[0038] The log routing module 101 is the core component of the system, responsible for establishing and maintaining a log routing table based on user session and function instance ID. The module receives log information from the serverless function instance, and decides the flow direction of the log according to the routing table. The structure of the routing table adopts a key-value pair design, where the key is the combination of user session ID and function instance ID, and the value is the corresponding WebSocket connection object. This design ensures the accurate directional transmission of logs, greatly improving the efficiency of the system.

[0039] The log filtering module 102 closely cooperates with the log routing module 101 to filter the log according to the routing table. This process realizes session-level log flow control, avoids the transmission of irrelevant logs, and significantly reduces the occupation of system resources. The algorithm design of the filtering process focuses on efficiency to ensure high processing capacity.

[0040] The message queue module 103 adopts a partition routing technology based on function ID, ensuring that the logs of the same function are routed to a fixed partition. This design not only guarantees the order of the logs, but also supports efficient partition-level log subscription and parallel processing. In this embodiment, the message queue can be implemented by a distributed message queue system such as Apache Kafka, and the number of partitions can be dynamically adjusted according to the system size to balance the processing efficiency and resource occupation.

[0041] The log consumption module 104 is responsible for subscribing and consuming logs from specific partitions of the message queue. This module supports multi-instance deployment, each instance can subscribe to one or more partitions, achieving parallel log consumption at the partition level. The consumption process can adopt a batch processing strategy, and the batch size can be dynamically adjusted according to system load to balance real-time performance and processing efficiency.

[0042] The WebSocket push module 105 is responsible for pushing the consumed logs to the client through the WebSocket connection. This module maintains a connection pool, supporting a large number of concurrent connections. In order to ensure the reliability and real-time performance of the push, the system adopts a heartbeat mechanism and retransmission strategy. The heartbeat interval can be adjusted according to network conditions, and after multiple consecutive heartbeat failures, the system will determine that the connection is disconnected and trigger the reconnection mechanism.

[0043] The workflow of the system is as follows: 1. The serverless function instance generates logs. 2. The log routing module 101 receives the logs and queries the routing table to determine the target of the logs. 3. The log filtering module 102 filters the logs according to the routing table. 4. The filtered logs are sent to a specific partition of the message queue module 103. 5. The log consumption module 104 subscribes and consumes logs from the specified partition. 6. The WebSocket push module 105 pushes the consumed logs to the client in real time.

[0044] The innovation of this system lies in the introduction of session-level log flow control and message queue partition routing technology based on function ID. The combination of these two technologies not only realizes accurate delivery and efficient processing of logs, but also greatly improves the parallel processing capability of the system. Compared with traditional full log collection and batch processing methods, this system has significant improvements in log processing delay and resource utilization.

[0045] It should be noted that the specific parameter settings in this embodiment (such as the number of partitions, batch size, heartbeat interval, etc.) can be adjusted according to the actual application scenario and system size. For example, in a high-concurrency environment, the number of message queue partitions and consumer instances can be increased; in scenarios with higher real-time requirements, the size of log batch processing and the heartbeat interval can be reduced.

[0046] In addition, although this embodiment mentions using Apache Kafka as an example of message queue implementation, the technical solutions of the present application are not limited to specific message queue technologies. Other high-performance distributed message queue systems can also be used to implement similar functions.

[0047] In summary, the Serverless architecture real-time log processing system provided by the present application effectively solves the real-time and efficiency problems of log processing in the Serverless environment through an innovative technical solution. The system not only applies to the current Serverless architecture, but also has good scalability and can adapt to more complex distributed system environments in the future.

[0048] Embodiment 2: Log routing table establishment and maintenance embodiment.

[0049] The overall architecture of the Serverless architecture real-time log processing system provided by the present application is shown in Figure 1 The system mainly includes a log routing module 101, a log filtering module 102, a message queue module 103, a log consumption module 104 and a WebSocket push module 105. These modules work together to realize the whole process management from log generation to final push, forming an efficient and real-time log processing system.

[0050] The log routing module 101 is the core component of the system, responsible for establishing and maintaining a log routing table based on user session and function instance ID. The module receives log information from the Serverless function instance and decides the flow direction of the log according to the routing table. The structure of the routing table adopts a key-value pair design, where the key is the combination of the user session ID and the function instance ID, and the value is the corresponding WebSocket connection object. In specific implementation, a high-performance hash table data structure such as unordered_map in C++ or ConcurrentHashMap in Java can be used to ensure fast lookup and update operations. The update operation of the routing table should be atomic to ensure data consistency in a high-concurrency environment.

[0051] The log filtering module 102 closely cooperates with the log routing module 101 to filter the log according to the routing table. This process realizes session-level log flow control and avoids the transmission of irrelevant logs. The core of the filtering algorithm is fast lookup, which can be expressed as: IsRelevant(log)=RouteTable.contains(log.sessionId+log.functionId)(1) Where RouteTable is the routing table, and log.sessionId and log.functionId are the session ID and function instance ID of the log, respectively. The time complexity of this operation is O(1), which guarantees efficient processing speed.

[0052] The message queue module 103 employs a function ID-based partition routing technique, ensuring that logs of the same function are routed to a fixed partition. The partitioning strategy can use a simple hash function: PartitionId = Hash(functionId) % NumPartitions (2) where Hash is a hash function and NumPartitions is the total number of partitions. This design not only guarantees the order of logs but also supports efficient partition-level log subscription and parallel processing. In this embodiment, the message queue is implemented using Apache Kafka, but the system design should consider an adapter pattern to facilitate future replacement of the message queue.

[0053] The log consumption module 104 is responsible for subscribing to and consuming logs from specific partitions of the message queue. This module supports multi-instance deployment, and each instance can subscribe to one or more partitions, enabling parallel log consumption at the partition level. The consumption process uses a batch processing strategy, and the batch size can be dynamically adjusted through a configuration file to balance real-time performance and processing efficiency. In specific implementation, the producer-consumer pattern can be used in conjunction with a thread pool to manage multiple consumer instances.

[0054] The WebSocket push module 105 is responsible for pushing the consumed logs to the client through the WebSocket connection. This module maintains a connection pool, supporting a large number of concurrent connections. The implementation of the connection pool can consider using a non-blocking I / O model, such as Java NIO or Node.js asynchronous I / O, to improve concurrent processing capability. To ensure the reliability and real-time performance of the push, the system adopts a heartbeat mechanism and retransmission strategy. The format of the heartbeat message can be defined as a simple JSON structure:

[0055] The workflow of the system is shown in Figure 2 , and the specific steps are as follows: 1. The serverless function instance generates logs (step 1 in Figure 2 ). 2. The log routing module 101 receives the logs and queries the routing table to determine the target (step 2). 3. The log filtering module 102 filters the logs according to the routing table (step 3). 4. The filtered logs are sent to a specific partition of the message queue module 103 (step 4). 5. The log consumption module 104 subscribes to and consumes logs from the specified partition (step 5). 6. The WebSocket push module 105 pushes the consumed logs to the client in real time (step 6).

[0056] The core innovation of this system lies in the introduction of session-level log flow control and function ID-based message queue partition routing technology. Session-level log flow control significantly reduces the transmission of irrelevant logs by precisely controlling the flow of logs, thereby significantly reducing system resource usage. Function ID-based partition routing ensures the sequential nature of logs for the same function while improving the system's parallel processing capabilities. The combination of these two technologies not only solves the high log processing latency and low resource utilization issues of traditional methods, but also opens up new possibilities for real-time log analysis in serverless environments.

[0057] It's important to note that the specific parameter settings in this embodiment can be adjusted based on the actual application scenario and system scale. For example, in a high-concurrency environment, the number of message queue partitions and consumer instances can be increased; in scenarios with higher real-time requirements, the log batch size and heartbeat interval can be reduced. The system should provide a configuration interface to allow operations personnel to adjust these parameters based on actual needs.

[0058] Furthermore, while this embodiment uses Apache Kafka as a message queue implementation, the technical solution of the present invention is not limited to a specific message queue technology. Other high-performance distributed message queue systems, such as RabbitMQ or Apache Pulsar, can also be used to implement similar functions. System design should consider a modular and pluggable architecture to facilitate future technology upgrades and replacements.

[0059] In summary, the serverless real-time log processing system presented in this paper effectively addresses the real-time and efficiency challenges of log processing in serverless environments through innovative technical solutions. This system is not only applicable to current serverless architectures but also boasts excellent scalability, adapting to more complex distributed system environments in the future. Future development directions may include the introduction of machine learning algorithms for intelligent log analysis or integration with other monitoring systems to provide more comprehensive system observation capabilities.

[0060] Example 3: Session-level log flow control example.

[0061] In the third embodiment of the present invention, the implementation method of session-level log flow control technology will be described in detail. This technology is one of the core innovations of the present invention, which significantly improves system efficiency and resource utilization by accurately controlling log flow.

[0062] like Figure 3As shown, the session-level log steering technique mainly consists of two key components: log routing table 301 and log filter 302. Log routing table 301 is responsible for maintaining the correspondence between user sessions and function instances, while log filter 302 performs precise filtering of logs based on the information from the routing table.

[0063] The data structure of log routing table 301 adopts a key-value pair design, where the key is the combination of user session ID and function instance ID, and the value is the corresponding WebSocket connection object. This design allows the system to quickly locate the log target of a specific session and function instance. In actual implementation, a high-performance concurrent hash table can be used to store this information. For example, in a Java environment, the ConcurrentHashMap class can be used: ConcurrentHashMap<String,WebSocketSession>routeTable= new ConcurrentHashMap<>(); The format of the key can be defined as: String key=sessionId+":"+functionInstanceId;

[0064] This design ensures the uniqueness of the key while maintaining good readability.

[0065] The update operation of log routing table 301 is a key process. When a new user session is established or a function instance is started, the system needs to add the corresponding entry in the routing table. Conversely, when the session ends or the function instance terminates, the corresponding entry needs to be removed. These operations can be implemented through the following methods:

[0066] Log filter 302 is the core component of the session-level log steering implementation. It receives logs from Serverless function instances and decides whether to forward the log based on the information from log routing table 301. The filtering process can be represented as the following algorithm:

[0067] The time complexity of this filtering process is O(1), ensuring efficient processing speed, even in large-scale concurrent situations, maintaining good performance.

[0068] In practical applications, the log filter 302 can be implemented as a standalone microservice or embedded in the log processing pipeline. It can receive logs through a message queue, process them, and send the filtered results to the next processing stage. This design improves the modularity of the system, facilitating future expansion and maintenance.

[0069] An important feature of the session-level log flow control technique is its adaptability. The system can dynamically adjust the filtering strategy based on the current load. For example, in high-load situations, the strictness of the filtering can be increased, retaining only the most critical log information; while in low-load situations, the filtering criteria can be relaxed, allowing more logs to pass through. This adaptive strategy can be implemented through a simple load factor: FilterThreshold = BaseThreshold * (1 + LoadFactor) (3)

[0070] where BaseThreshold is the base filtering threshold, and LoadFactor is an indicator of the current system load. This formula ensures that the filtering becomes more stringent as the system load increases.

[0071] The session-level log flow control technique of this embodiment is closely integrated with the overall system architecture described in Embodiment 1. It plays a core role in the log routing module 101 and the log filtering module 102, ensuring the efficient operation of the entire system. At the same time, this technique also provides a foundation for the log routing table establishment and maintenance in Embodiment 2.

[0072] It should be noted that although this embodiment mainly discusses the WebSocket-based implementation, the core idea of this session-level log flow control technique can also be applied to other types of real-time communication mechanisms, such as Server-Sent Events (SSE) or long polling. System design should consider this scalability to reserve space for future technological evolution.

[0073] In addition, in actual deployment, the problem of routing table persistence may need to be considered to deal with system restart or fault recovery. One possible solution is to periodically save snapshots of the routing table to persistent storage, such as Redis or a distributed file system. This way, the routing state can be quickly recovered after system restart, minimizing service interruption time.

[0074] In summary, the session-level log flow control technology described in this embodiment is an innovative solution that significantly improves the efficiency and resource utilization of log processing in a Serverless environment by precisely controlling log flow. This technology not only applies to the specific implementation of the invention, but can also be extended to other scenarios that require precise control of data flow, showing broad application prospects.

[0075] Embodiment 4: Message queue partition routing embodiment.

[0076] The fourth embodiment of the present invention details the implementation method of the message queue partition routing technology. This technology is another core innovation of the present invention, which realizes efficient log processing and precise log positioning through a partitioning strategy based on function ID.

[0077] As shown in Figure 4 , the message queue partition routing technology mainly includes three key components: partition router 401, message queue 402, and partition consumer 403. The partition router 401 is responsible for determining the target partition of the log message according to the function ID, the message queue 402 is responsible for storing and managing partitioned log messages, and the partition consumer 403 is responsible for subscribing and consuming logs from a specific partition.

[0078] The core function of the partition router 401 is to route log messages to the correct partition. This process can be achieved through a hash function. Specifically, the following algorithm is used: PartitionId=Hash(FunctionId)%NumPartitions (4) Where Hash is a hash function, FunctionId is the unique identifier of the Serverless function, and NumPartitions is the total number of partitions of the message queue. This algorithm ensures that logs of the same function are always routed to the same partition, thereby guaranteeing the order and integrity of the logs. In actual implementation, an efficient hash function such as MurmurHash3 can be selected to calculate the partition ID.

[0079] The message queue 402 is the core storage component of this technology. In this embodiment, a distributed message queue system that supports partitioning, such as Apache Kafka, is selected. This type of system allows a topic to be divided into multiple partitions, and each partition can independently store and process messages, which exactly meets the partition routing requirements. System administrators can set an appropriate number of partitions according to actual needs to balance processing performance and resource consumption.

[0080] Partitioned consumers 403 are responsible for subscribing and consuming logs from specific partitions. This design allows multiple consumers to handle different partitions in parallel, significantly improving the system's processing capacity. Each partitioned consumer instance can be configured to focus on a specific partition, ensuring sequential processing of logs while maximizing the system's parallel processing capacity.

[0081] One significant advantage of this message queue partition routing technique is its ability to effectively handle sudden bursts of log streams. When a function suddenly generates a large number of logs, these logs will be routed to the same partition without affecting the processing of other partitions. This isolation ensures the stability and reliability of the system.

[0082] Another important feature is scalability. By increasing the number of partitions and consumer instances, the system can linearly expand its processing capacity. For example, if a partition is found to be under too much pressure, the following steps can be taken to increase the number of partitions: 1. Create a new partition 2. Update the number of partitions in the partition router 3. Rebalance existing log data 4. Start a new partitioned consumer instance to handle the added partition

[0083] This message queue partition routing technique is tightly integrated with the overall system architecture described in Embodiment 1, particularly in the message queue module 103. It also complements the session-level log flow control technique in Embodiment 3, forming an efficient log processing framework for the invention.

[0084] It is worth noting that although this embodiment mainly discusses the implementation based on a specific message queue system, the core idea of this partition routing can be applied to other distributed message queue systems. When choosing a specific message queue system, factors such as system size, performance requirements, and operational complexity should be considered.

[0085] In addition, to further improve the reliability of the system, a log backup mechanism can be introduced. For example, the replication factor of the message queue system can be configured to be greater than 1, ensuring that each partition has multiple copies. In this way, even if a storage node fails, the system can continue to operate normally.

[0086] In summary, the message queue partition routing technique described in this embodiment is an innovative solution that significantly improves the efficiency and scalability of log processing in a Serverless environment through intelligent partitioning strategies and parallel processing mechanisms. This technology not only applies to log processing, but also can be extended to other scenarios that require high throughput and low latency data processing, showing broad application prospects.

[0087] Example 5: Log consumption and WebSocket push implementation.

[0088] The fifth embodiment of the present application details the implementation method of log consumption and WebSocket push technology. This technology is the key link in the present application for realizing real-time log transmission. It ensures that log data can reach the client with minimal delay through efficient log consumption mechanism and real-time WebSocket push.

[0089] As shown in Figure 5 , the log consumption and WebSocket push technology mainly includes three core components: log consumer 501, WebSocket connection manager 502, and log pusher 503. The log consumer 501 is responsible for efficiently obtaining log data from the message queue, the WebSocket connection manager 502 is responsible for maintaining real-time connections with the client, and the log pusher 503 is responsible for pushing the consumed log data to the client in real time through WebSocket connection.

[0090] The core function of the log consumer 501 is to efficiently obtain log data from the message queue. Considering the burstiness and high concurrency characteristics of log generation in the Serverless environment, the log consumer adopts a batch consumption strategy. This strategy can be described by the following formula for the batch size: BatchSize=min(MaxBatchSize,max(MinBatchSize,QueueSize*ConsumptionRate)) (5) where MaxBatchSize and MinBatchSize are the upper and lower limits of the batch size, QueueSize is the number of messages in the current queue, and ConsumptionRate is a consumption rate dynamically adjusted according to system load. This dynamic batch strategy can adaptively adjust the consumption behavior when the system load changes, ensuring real-time performance and improving processing efficiency.

[0091] The WebSocket connection manager 502 is responsible for establishing and maintaining WebSocket connections with the client. In order to support large-scale concurrent connections, the connection manager adopts a non-blocking I / O model. This model allows a single thread to manage multiple connections, significantly improving the system's concurrent processing capability. The connection manager also implements a heartbeat mechanism to detect connection status and network quality. The sending interval of the heartbeat packet can be dynamically adjusted by the following formula: HeartbeatInterval=BaseInterval*(1+NetworkLatency / MaxLatency) (6) Where BaseInterval is the base heartbeat interval, NetworkLatency is the current network latency, and MaxLatency is the acceptable maximum latency. This dynamic adjustment strategy can adjust the heartbeat frequency in a timely manner when the network condition changes, ensuring the reliability of the connection and avoiding unnecessary network overhead.

[0092] The log pusher 503 is a key component that enables real-time log transmission. It receives log data from the log consumer and pushes data to the client through a WebSocket connection. To optimize transmission efficiency, the log pusher uses a compression transmission strategy. The choice of compression algorithm needs to balance between compression rate and computational overhead. In this embodiment, the GZIP algorithm is used, which can provide good compression effect in most scenarios while the computational overhead is relatively small.

[0093] The workflow of log consumption and WebSocket push is shown in Figure 6 The specific steps are as follows: 1. The log consumer 501 batch gets log data from the message queue (step 1). 2. The consumed log data is passed to the log pusher 503 after preliminary processing (step 2). 3. The log pusher 503 queries the WebSocket connection manager 502 to get the corresponding client connection (step 3). 4. The log pusher 503 compresses the log data and sends it to the client through the WebSocket connection (step 4). 5. The WebSocket connection manager 502 periodically sends heartbeat packets to maintain the connection state (step 5).

[0094] A significant advantage of this log consumption and WebSocket push technology is that it can achieve real-time log transmission. Under ideal network conditions, the end-to-end delay from log generation to client reception can be controlled within milliseconds. This is crucial for real-time monitoring and rapid problem diagnosis.

[0095] Another important feature is the reliability and fault tolerance of the system. The WebSocket connection manager implements an automatic reconnection mechanism that automatically attempts to reestablish the connection when a disconnection is detected. At the same time, the log consumer implements consumption site management, ensuring that log data is not lost or repeatedly consumed even in the event of system failure or network interruption.

[0096] This log consumption and WebSocket push technology is tightly integrated with the overall system architecture described in Embodiment 1, playing a core role in the log consumption module 104 and the WebSocket push module 105. It also complements the session-level log flow control technology in Embodiment 3 and the message queue partition routing technology in Embodiment 4, together forming the efficient real-time log processing framework of the present application.

[0097] It is worth noting that although this embodiment mainly discusses WebSocket-based implementation, the core idea of this real-time push can also be applied to other real-time communication technologies such as Server-Sent Events (SSE) or HTTP / 2-based server push. When choosing a specific push technology, factors such as client compatibility, network environment, performance requirements, etc. should be considered.

[0098] In addition, in order to further improve the usability of the system, a push failure retry mechanism can be considered. For example, when the WebSocket push fails, the system can temporarily store the failed logs in the local cache and automatically re-push after the connection is restored. This mechanism can effectively deal with temporary problems such as network fluctuations.

[0099] In summary, the log consumption and WebSocket push technology described in this embodiment is an innovative solution that realizes real-time transmission and display of logs in a Serverless environment through efficient consumption strategies and real-time push mechanisms. This technology is not only suitable for log processing, but can also be extended to other scenarios that require real-time data transmission, such as real-time monitoring, instant messaging, etc., showing broad application prospects.

[0100] Embodiment 6: Connection monitoring and resource cleanup implementation.

[0101] The sixth embodiment of the present application details the implementation method of the connection monitoring and resource cleanup technology. This technology is a key link in the present application to ensure system stability and efficient use of resources, which effectively prevents resource leakage and improves system reliability and performance through intelligent connection state monitoring and timely resource recovery.

[0102] As shown in Figure 7 , the connection monitoring and resource cleanup technology mainly includes three core components: connection monitor 701, resource cleaner 702, and state storage 703. The connection monitor 701 is responsible for real-time monitoring of the state of WebSocket connections, the resource cleaner 702 is responsible for cleaning related resources in a timely manner when the connection is disconnected, and the state storage 703 is responsible for maintaining the mapping relationship between connection states and related resources.

[0103] The core function of the connection monitor 701 is to monitor the status of WebSocket connections in real-time. Considering the transience and instability of connections in a serverless environment, the connection monitor adopts a multi-level monitoring strategy. First, it uses the ping-pong mechanism of the WebSocket protocol for basic connection detection. Second, it evaluates connection quality by analyzing data transfer frequency and network latency. Finally, it also considers heartbeat messages at the application layer. This multi-level monitoring strategy can be evaluated comprehensively by the following formula: where w1, w2, w3, w4 are the weights of each indicator, which can be adjusted according to the actual application scenario. PingPongStatus reflects the survival status of the basic connection, DataTransferRate represents the data transfer frequency, NetworkLatency represents the network delay, and HeartbeatStatus reflects the state of the application layer heartbeat. Through this formula, the system can comprehensively evaluate the health status of the connection and timely discover potential problems.

[0104] The resource cleaner 702 is responsible for cleaning related resources in a timely manner when the connection is disconnected. To avoid resource leakage and improve cleaning efficiency, the resource cleaner adopts a two-level cleaning strategy. The first level is immediate cleaning, which triggers cleaning operations immediately when the connection monitor detects that the connection is disconnected. The second level is periodic scanning and cleaning, which periodically scans all resources and cleans those that have been invalidated but have not been cleaned in time. This two-level cleaning strategy can be represented by the following pseudo code:

[0105] The state store 703 is a key component for efficient resource cleaning. It maintains the mapping relationship between connection status and related resources, and uses high-performance memory data structures such as concurrent hash tables to support high-concurrency read-write operations. The data structure of the state store can be represented as: StateStore:{ConnectionId→(ConnectionState,ResourceList)} (8) where ConnectionId is the unique identifier of the connection, ConnectionState is the current state of the connection, and ResourceList is the list of resources related to the connection.

[0106] The workflow of connection monitoring and resource cleaning is shown in Figure 8 , and the specific steps are as follows: 1. The connection monitor 701 continuously monitors the status of WebSocket connections (step 1). 2. When a connection state change is detected, the connection monitor 701 updates the state storage 703 (step 2). 3. The resource cleaner 702 periodically queries the state storage 703 for resource information that needs to be cleaned up (step 3). 4. The resource cleaner 702 performs the cleanup operation to release the resources that are no longer needed (step 4). 5. After the cleanup is completed, the resource cleaner 702 updates the state storage 703 (step 5).

[0107] One significant advantage of this connection monitoring and resource cleanup technique is that it effectively prevents resource leaks. In a serverless environment, function instances and connections often have short and unpredictable lifetimes, and traditional resource management methods can lead to resources not being released in a timely manner. This technique ensures that resources are immediately recycled after a connection is disconnected through real-time monitoring and proactive cleanup, greatly improving the resource utilization efficiency of the system.

[0108] Another important feature is the reliability and fault tolerance of the system. Through multi-level monitoring strategies and two-level cleanup mechanisms, the system can effectively deal with various abnormal situations, such as network interruptions and client crashes. Even in the case of partial component failure, periodic scanning and cleanup can serve as the last safeguard to ensure that the system eventually returns to a consistent state.

[0109] This connection monitoring and resource cleanup technique is tightly integrated with the overall system architecture described in Embodiment 1, particularly in the WebSocket push module 105. It also complements the session-level log flow control technique in Embodiment 3 and the WebSocket push technique in Embodiment 5, forming a highly efficient and reliable real-time log processing framework for the invention.

[0110] It is worth noting that although this embodiment mainly discusses WebSocket-based implementation, the core idea of this connection monitoring and resource cleanup can also be applied to other types of long connection or session management scenarios. For example, in HTTP long polling-based implementation, similar mechanisms can be used to manage the lifetimes of polling connections and related resources.

[0111] In addition, to further improve the reliability of the system, a distributed lock mechanism can be introduced to handle resource cleanup problems in multi-instance deployment scenarios. When the system is deployed in a cluster, multiple cleaner instances may attempt to clean up the same resource at the same time, and a distributed lock can avoid potential race conditions and data inconsistency problems.

[0112] In summary, the connection monitoring and resource cleaning technology described in this embodiment is an innovative solution that effectively solves the resource management problem in a Serverless environment through intelligent monitoring strategies and efficient cleaning mechanisms. This technology not only applies to real-time log processing systems, but can also be extended to other distributed systems that require fine-grained resource management, such as microservice architectures, IoT device management, etc., showing broad application prospects.

[0113] Embodiment 7: Partition-level parallel log consumption embodiment.

[0114] The seventh embodiment of the present application elaborates on the implementation method of the partition-level parallel log consumption technology. This technology is a key link in realizing high-throughput log processing in the present application, which significantly improves the processing capacity and scalability of the system through intelligent partition subscription and parallel processing mechanisms.

[0115] As shown in Figure 9 , the partition-level parallel log consumption technology mainly includes three core components: partition subscription manager 901, parallel consumption coordinator 902, and log processing executor 903. The partition subscription manager 901 is responsible for dynamically managing the subscription relationship between consumers and partitions, the parallel consumption coordinator 902 is responsible for coordinating the parallel processing of multiple consumer instances, and the log processing executor 903 is responsible for the actual log processing work.

[0116] The core function of the partition subscription manager 901 is to realize the dynamic mapping between consumers and partitions. Considering the dynamic change characteristics of the load in the Serverless environment, the manager adopts an adaptive partition allocation strategy. Specifically, the allocation strategy can be expressed as: Where P i is the number of partitions allocated to the i-th consumer, C i is the processing capacity of the i-th consumer (which can be estimated by historical processing speed), n is the total number of consumers, and N is the total number of partitions. This strategy can dynamically adjust the partition allocation according to the actual processing capacity of each consumer, ensuring optimal utilization of system resources.

[0117] The parallel consumption coordinator 902 is responsible for coordinating the parallel processing of multiple consumer instances. In order to maximize parallelism while ensuring consistency of processing, the coordinator adopts a concurrency control mechanism based on version vectors. The processing state of each partition can be represented as a version vector: V p = (v1, v2,..., v n ) (10) Where v iPi represents the processing progress of the i-th consumer on the partition. This mechanism allows multiple consumers to handle different log entries in parallel while ensuring the correctness of the processing order.

[0118] The log processing executor 903 is the component that actually performs the log processing. To improve processing efficiency, the executor adopts a pipeline processing model. The processing pipeline includes the following stages: 1. Log parsing: parsing the raw log into structured data. 2. Data enrichment: supplementing the missing context information in the log. 3. Rule matching: applying predefined processing rules. 4. Result aggregation: aggregating the processing results.

[0119] This pipeline model can be evaluated for its processing capacity by the following formula: Where T1, T2, T3, and T4 are the processing times of each stage respectively. By optimizing the bottleneck stage, the overall processing efficiency can be significantly improved.

[0120] The workflow of partition-level parallel log consumption is shown in Figure 10 The specific steps are as follows: 1. The partition subscription manager 901 allocates partitions based on the current system state (step 1). 2. The parallel consumption coordinator 902 assigns processing tasks to each consumer (step 2). 3. The log processing executor 903 obtains logs from the specified partition and processes them (step 3). 4. The processing results are returned to the coordinator for consistency checking (step 4). 5. The coordinator updates the processing state and decides the next operation (step 5).

[0121] One of the significant advantages of this partition-level parallel log consumption technology is its high scalability. By increasing the number of consumer instances, the system can linearly improve its processing capacity. For example, in a system containing 100 partitions, theoretically, 100 consumer instances can be deployed simultaneously, each handling one partition, achieving maximum parallel processing.

[0122] Another important feature is the system's resilience and fault tolerance. When a consumer instance fails, the partition subscription manager can quickly detect it and reassign the partition, ensuring continuous operation of the system. At the same time, the version vector-based concurrency control mechanism ensures that even during the reassignment process, there will be no duplication or omission of log processing.

[0123] This partition-level parallel log consumption technique is tightly integrated with the overall system architecture described in Embodiment 1, particularly playing a core role in the log consumption module 104. It also complements the message queue partition routing technique in Embodiment 4, together forming the high-throughput log processing framework of the present invention.

[0124] It is worth noting that although this embodiment mainly discusses partition-based parallel consumption, the core idea of this parallel processing can also be applied to other data-intensive processing scenarios. For example, in large-scale data analysis or machine learning tasks, similar partitioning and parallel processing strategies can be adopted to improve processing efficiency.

[0125] In addition, to further improve the adaptability of the system, an automated performance tuning mechanism can be introduced. For example, the system can automatically analyze processing bottlenecks through machine learning algorithms and dynamically adjust resource allocation at each stage of the pipeline to achieve optimal processing efficiency.

[0126] In summary, the partition-level parallel log consumption technique described in this embodiment is an innovative solution that significantly improves the efficiency and scalability of log processing in a Serverless environment through intelligent partition management and parallel processing mechanisms. This technology is not only suitable for log processing, but can also be extended to other scenarios that require high-throughput data processing, such as real-time data analysis, large-scale event processing, etc., showing broad application prospects.

[0127] Embodiment 8: Embodiment of integration with existing log persistence system.

[0128] The eighth embodiment of the present invention details the technical solution of integrating with existing log persistence systems. This technical solution ingeniously solves the contradiction between real-time log processing and long-term log storage, providing a comprehensive and flexible solution for log management in a Serverless environment.

[0129] As shown in Figure 11 , this embodiment mainly includes four core components: real-time log processing pipeline 1101, log persistence module 1102, log retrieval interface 1103, and log analysis engine 1104. The real-time log processing pipeline 1101 is responsible for processing real-time log streams, the log persistence module 1102 is responsible for storing log data to a persistent storage system, the log retrieval interface 1103 provides efficient log query functions, and the log analysis engine 1104 is responsible for complex log analysis tasks.

[0130] The real-time log processing pipeline 1101 is the integration point of the core technologies described in the previous embodiments of the present application. It contains innovative technologies such as session-level log flow control, message queue partition routing, and partition-level parallel consumption. The main goal of this pipeline is to achieve millisecond-level log processing delay to support real-time monitoring and rapid problem diagnosis. The efficiency of real-time processing can be expressed by the following formula:

[0131] where E realtime is the real-time processing efficiency, N processed is the number of processed logs, T total is the total processing time, and L avg is the average log processing delay.

[0132] The log persistence module 1102 is responsible for storing log data into the persistent storage system. Considering the huge amount of log data in a serverless environment, this module adopts a tiered storage strategy. Specifically, the storage strategy can be expressed as: where S(t) represents the storage strategy for logs before time t, S hot , S warm , and S cold represent hot storage, warm storage, and cold storage, respectively, and T hot and T warm are time thresholds. This tiered storage strategy can effectively balance storage costs and query performance.

[0133] The log retrieval interface 1103 provides efficient log query functions. To support complex query requirements, this interface implements a query engine based on inverted indexes. Query performance can be evaluated by the following formula: where P query is the query performance, T index is the index lookup time, and T fetch is the data acquisition time. By optimizing the index structure and cache strategy, query performance can be significantly improved.

[0134] The log analysis engine 1104 is responsible for performing complex log analysis tasks. This engine uses a scalable distributed computing framework to support custom analysis logic. The efficiency of the analysis task can be expressed by the following formula: where E analysis is the analysis efficiency, D processed is the amount of data processed, T analysis is the analysis time, and R resourcesis the amount of resources used.

[0135] The workflow of this embodiment is shown in Figure 12 and the specific steps are as follows: 1. The real-time log processing pipeline 1101 receives and processes real-time log streams (step 1). 2. The processed logs are sent to both the WebSocket push module and the log persistence module 1102 (step 2). 3. The log persistence module 1102 stores the logs into the appropriate storage layer (step 3). 4. The log retrieval interface 1103 retrieves logs from the storage system in response to user queries (step 4). 5. The log analysis engine 1104 performs complex analysis tasks, possibly involving large amounts of historical log data (step 5).

[0136] One significant advantage of this integration with existing log persistence systems is that it achieves a seamless combination of real-time processing and long-term storage. Through the real-time log processing pipeline, the system can support millisecond-level log visibility and real-time monitoring; at the same time, through the log persistence module, the system also retains the ability to perform in-depth analysis on historical logs. This dual-mode design provides users with great flexibility, meeting various needs from real-time fault diagnosis to long-term trend analysis.

[0137] Another important feature is the scalability and adaptability of the system. Through the hierarchical storage strategy and distributed computing framework, the system can easily cope with the growth of data volume and changes in analysis requirements. For example, the capacity of hot storage can be dynamically adjusted according to actual needs, or new analysis algorithms can be added to the analysis engine.

[0138] This integrated solution complements the previous embodiments well. It not only retains all the advantages of real-time log processing, but also expands the functional range of the system through persistent storage and advanced analysis capabilities. In particular, it perfectly fits into the overall architecture described in Embodiment 1 and can be seen as a further extension of the log consumption module 104 and the WebSocket push module 105.

[0139] It is worth noting that although this embodiment mainly discusses the integration with a specific type of persistent storage system, the core idea of this integration can be applied to various storage technologies. For example, different storage schemes such as relational databases, document databases, or object storage can be chosen according to specific needs.

[0140] In addition, to further improve the intelligent level of the system, machine learning techniques can be considered in the log analysis engine. For example, anomaly detection algorithms can be used to automatically identify abnormal patterns in logs, or predictive analysis techniques can be used to predict potential system problems.

[0141] In summary, the technical solution described in this embodiment, which integrates with existing log persistence systems, is an innovative solution that achieves the organic combination of real-time processing and long-term storage through clever design, providing a comprehensive and flexible solution for log management in a Serverless environment. This solution not only applies to Serverless architecture, but can also be extended to other scenarios that require both real-time and historical analysis, such as Internet of Things data processing, financial transaction analysis, etc., showing broad application prospects.

Claims

1. A real-time log processing method under a Serverless architecture, characterized in that: The following steps are involved: (a) Establish a log routing table based on user session and function instance ID; (b) screening the logs according to the log routing table; (c) Sending the filtered logs to the message queue; (d) partitioning the message queue based on function ID; (e) subscribing to and consuming logs from a specific partition of the message queue; (f) Push the consumed logs to the client through the WebSocket connection.

2. The method according to claim 1, characterized in that The structure of the log routing table includes: (a) Key: a combination of user session ID and function instance ID; (b) Value: the corresponding WebSocket connection object.

3. The method according to claim 1, characterized in that The following steps are also included: (a) Monitor WebSocket connection status; (b) When a WebSocket connection is detected to be disconnected, the relevant items in the log routing table are automatically cleared.

4. The method according to claim 1, wherein The step of partitioning the message queue based on the function ID includes: (a) Use the function ID as the partition key; (b) Ensure that logs for the same function are always routed to the same partition.

5. The method according to claim 1, characterized in that The following steps are also included: (a) Supports subscription to the message queue partition corresponding to a specific function ID; (b) Implement parallel log consumption at the partition level.

6. The method according to claim 1, characterized in that The following steps are also included: (a) When sending logs, the target partition is calculated based on the function ID; (b) Send log messages to the specified partition.

7. A real-time log processing system under Serverless architecture, characterized by: include: (a) a log routing module, used to establish and maintain a log routing table based on user session and function instance IDs; (b) a log screening module, configured to screen logs according to the log routing table; (c) a message queue module, which receives the filtered logs and stores them in partitions based on function IDs; (d) a log consumption module, configured to subscribe to and consume logs from a specific partition of the message queue; (e) WebSocket push module, used to push consumed logs to the client.

8. The system according to claim 7, characterized in that Also includes: (a) Connection monitoring module, used to monitor the WebSocket connection status and clean up related resources when the connection is disconnected.

9. The system according to claim 7, wherein: The message queue module also includes: (a) Partition routing unit, used to determine the target partition of the log based on the function ID.

10. The system according to claim 7, wherein: The log consumption module supports parallel log consumption at the partition level.