A design method for solving high concurrency problem of backend system using asynchronous flow

By designing asynchronous processes and employing a three-tier data recovery mechanism, the problem of data loss in traditional systems under high concurrency, high throughput, and low latency was solved, achieving efficient data processing and system optimization.

CN115438120BActive Publication Date: 2025-11-11HAINA ZHIYUAN DIGITAL TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210926701.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-03
Publication Date
2025-11-11
Estimated Expiration
2042-08-03

AI Technical Summary

Technical Problem

Traditional methods are becoming less effective in optimizing systems with high concurrency, high throughput, and low latency, making it difficult to meet the high concurrency requirements of Internet applications, and they also pose a high risk of data loss.

Method used

An asynchronous process design is adopted, using Kafka cluster, Redis cluster and Logstach as data transfer stations. Data integrity is guaranteed through a three-layer data recovery mechanism, including synchronous to asynchronous processing, multi-dimensional data monitoring and modular design.

Benefits of technology

Significantly reduce the time spent in the synchronization process, increase concurrency and throughput, ensure no data loss, and improve the efficiency of parallel system development and the quality of data recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438120B_ABST
    Figure CN115438120B_ABST
Patent Text Reader

Abstract

This invention discloses a design method for solving the high concurrency problem of backend systems using asynchronous processes. It involves building a highly available Kafka and Redis cluster and installing the Logstach plugin; inserting the dataset that needs to be saved in the synchronous business process into the Redis cluster; simultaneously writing this dataset to a log file in a specific JSON format string; sending the incremental dataset after processing normal business logic to a specific Topic in the Kafka cluster; simultaneously storing this incremental dataset in the Redis cluster; asynchronously executing the data persistence process through a Kafka consumer; verifying the successful persistence of data based on the dataset stored in the Redis cluster; and verifying the successful persistence of data based on the dataset collected from the Logstach logs. By using excellent and reliable middleware as a data transfer station, converting synchronous to asynchronous processing, and ensuring data integrity through multiple dimensions and gradients, the method achieves the goal of preventing data loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet applications, and more specifically to a design method for solving the high concurrency problem of backend systems using asynchronous processes. Background Technology

[0002] China is the country with the fastest-growing internet development. Due to its large population, valuable internet applications need to handle massive user requests and provide a user-friendly experience. This is especially true for industries such as e-commerce platforms, internet finance, insurance, and payment, where systems urgently require high concurrency, high throughput, and low latency. Traditional approaches to achieving high concurrency, high throughput, and low latency focus on optimizing business processes, databases, adding middleware, and increasing hardware resources. However, as user requests increase and business system functions become more complex, the effectiveness of these optimization methods decreases compared to the investment, eventually reaching a bottleneck. Summary of the Invention

[0003] To address the aforementioned issues, this invention provides a design method for solving the high concurrency problem of backend systems using asynchronous processes. It utilizes excellent and reliable middleware as a data transfer station, converts synchronous processes to asynchronous ones, and ensures data integrity through multiple dimensions and gradients.

[0004] Definitions:

[0005] 1. Redis: A high-performance, easily scalable key-value pair caching database.

[0006] 2. Kafka: A high-throughput distributed publish-subscribe messaging system.

[0007] 3. Topic: The messaging system categorizes message types by topic.

[0008] 4. Logstach: A real-time pipelined open-source log collection engine.

[0009] 5. key: A unique identifier for data storage.

[0010] 6. JSON: A lightweight data exchange format.

[0011] 7. set: Database data storage command.

[0012] To achieve the above technical objectives and effects, this invention provides the following technical solution: a design method for solving the high concurrency problem of backend systems using asynchronous processes, comprising the following steps:

[0013] Step 1: Build a highly available Kafka and Redis cluster as the data source for asynchronous processes and act as an asynchronous data processing relay station. Install Logstach to collect application logs.

[0014] Step 2: The data that needs to be stored in the database in the business process is called the first-stage dataset. The first-stage dataset is inserted into the Redis cluster, and the uniqueness of the data key is guaranteed based on the internal serial number.

[0015] Step 3: Write the first-stage dataset into a log file using a specific JSON format string, which is different from the ordinary log format.

[0016] Step 4: After processing other normal business logic, complete or update the first-stage dataset based on the processed results to obtain the final dataset, which is called the second-stage dataset.

[0017] Step 5: Send the second-stage dataset to a specific topic in the Kafka cluster using the producer; this is called the "database removal topic".

[0018] Step 6: Store the second-stage dataset into the Redis cluster. Its key value must be different from that of the first-stage dataset. The uniqueness of each data item is ensured based on the external serial number.

[0019] Step 7: The Kafka cluster consumer consumes the topic to be retrieved from the database and asynchronously executes the data write-to-database process;

[0020] Step 8: Periodically review whether the data has been successfully stored in the Redis cluster based on the first-phase dataset.

[0021] Step 9: Periodically review the first-stage dataset collected based on Logstach logs to check whether the data has been successfully stored in the database.

[0022] Preferably, the middleware cluster should be used independently and not shared with other application systems.

[0023] Preferably, the asynchronous process of the Kafka cluster is used to perform the main asynchronous data storage responsibility, the Redis cluster is used to perform the second layer of data protection, and the Logstach is used to perform the third layer of data protection.

[0024] Ideally, the three-tiered data recovery mechanism should work together to ensure data integrity and prevent loss.

[0025] Preferably, the three-layer data recovery mechanism is executed in staggered shifts as a fallback, but the data storage process must use the same set of code. This process is called the restore process, and the restore process must be compatible with three different data sources.

[0026] Ideally, producers and consumers in a Kafka cluster should run in separate processes, not in a single process with multiple threads.

[0027] Preferably, the first-stage dataset stored in Redis is used for the restore process, and the second-stage dataset is used for remedial measures in case of Kafka message loss, which is a fault tolerance mechanism for complex systems.

[0028] Preferably, the specific JSON format string mentioned in step three is: serializing the first-stage dataset into JSON format and adding a specific log identifier before the serialized string. This facilitates the Logstach log plugin in reading and recognizing these data logs, providing a data source for the third-layer log recovery mechanism.

[0029] Compared with the prior art, the beneficial effects of the present invention are:

[0030] Firstly, the synchronous-to-asynchronous data design can greatly reduce the time consumed by the synchronous process, which well meets the requirements of low-latency system indicators. At the same time, under the same conditions, the concurrency and throughput are also greatly improved.

[0031] Secondly, a three-layer data recovery mechanism is used to improve the robustness of asynchronous processing and ensure that data is not lost.

[0032] Thirdly, modular system design improves the efficiency of parallel development and troubleshooting of complex systems.

[0033] Fourthly, the unified design of the core restore process reduces the difficulty of collaboration between complex systems and improves the quality of data recovery. Attached Figure Description

[0034] Figure 1 This is an architecture diagram of a design method for solving the high concurrency problem of backend systems using asynchronous processes, according to the present invention.

[0035] Figure 2 This is a detailed flowchart of the present invention;

[0036] Figure 3 This is a detailed flowchart of the asynchronous processing flow in this invention. Detailed Implementation

[0037] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0038] A design approach that uses asynchronous processes to solve the high concurrency problem of backend systems, such as... Figure 1 , 2As shown, the specific steps include the following:

[0039] Step 1: Build a highly available Kafka and Redis cluster as the data source for asynchronous processes, acting as an intermediary for asynchronous data processing. Install Logstach to collect application logs. Logging is crucial and serves as the last layer of protection. Kafka and Redis clusters involve memory and network transmission when storing data; if underlying errors occur (operating system, network, etc.), data may be lost. Logs, however, are written directly to the local disk, making them relatively safer than the previous two methods. Therefore, this serves as a backup solution for intermediate storage before data is saved to the database.

[0040] Step 2: The data that needs to be stored in the database during the business process is called the first-stage dataset. This dataset is inserted into the Redis cluster, with the uniqueness of the data keys ensured by an internal serial number. The internal serial number is generated using a sequence number generator and must include a timestamp down to the millisecond level, allowing the extraction of the data's occurrence time. The first-stage dataset is stored in the Redis cluster using a set, allowing data to be stored in Redis cluster according to time intervals, with two-hour intervals per day. Normally, the first-layer Kafka data processing is completed within a few seconds. The second-layer Redis-based data review task can be executed periodically, checking whether the data in Redis over the past two hours has been processed correctly. If correct, the data from those two hours can be cleared because it has already been persisted in the database. This data review mechanism allows for control over Redis memory usage. If the data is incorrect, an asynchronous processing flow is executed using the data in Redis, and the second-layer data protection mechanism takes effect. After processing, the data is also cleared.

[0041] Step 3: Write the first-stage dataset into a log file using a specific JSON format string, which differs from the ordinary log format. Serialize the first-stage dataset into JSON format, and add a specific log identifier before the serialized string to facilitate the Logstach log plugin's reading and recognition of these data logs, providing a data source for the third-layer log recovery mechanism.

[0042] Step 4: After processing other normal business logic, complete or update the first-stage dataset based on the processed results to obtain the final dataset, which is called the second-stage dataset.

[0043] Step 5: Send the second-stage dataset to a specific topic in the Kafka cluster using the producer; this is called the "database delivery topic." If the same Kafka cluster as the business system is used, the topic name must be distinct from other business topics.

[0044] Step 6: Store the second-stage dataset in the Redis cluster. Its key must be different from the first-stage dataset, and the uniqueness of each data record is ensured by using an external serial number. This dataset primarily serves as a cache for the results of the first-stage data after synchronization. This is because the first-stage dataset stored in Redis is not the final data that needs to be saved.

[0045] Step 7: The Kafka cluster consumer consumes the database topic and asynchronously executes the data persistence process. Steps 1-6 are all synchronous processes; from this step onwards, all processes are asynchronous.

[0046] Step 8: Periodically review whether the data has been successfully stored in the Redis cluster based on the first-stage dataset.

[0047] Step 9: Periodically review the first-stage dataset collected based on Logstach logs to check whether the data has been successfully stored in the database.

[0048] The three-layer data recovery mechanism uses the same process and the same code, only running in three different processes with different asynchronous data entry points. Therefore, the asynchronous process needs to be compatible with all three scenarios. Figure 3 As shown, an asynchronous process needs to include the following steps:

[0049] 1. Determine whether the data to be processed is the first-stage dataset or the second-stage dataset. If it is the first-stage dataset, the second-stage dataset needs to be obtained through the cache in step 6 or normal business processes, such as external interfaces.

[0050] 2. The core process of converting synchronous to asynchronous processing should be encapsulated into an independent part call. This part is the data saving process, the most time-consuming part extracted from the synchronous process.

[0051] 3. If the previous step processed the dataset of the first stage, then the third step needs to try to obtain the dataset of the second stage again to complete the data update process. This process is the same as the previous step, which is also extracted from the synchronous process and encapsulated into an independent part call. If this step is not completed and the dataset of the second stage still cannot be obtained, then another process is needed to complete the compensation mechanism. This compensation mechanism is not related to whether the dataset is processed synchronously or asynchronously. It is an independent part of the complete business data closed loop.

[0052] 4. After the asynchronous process is completed, clean up the relevant data in the Redis cluster and record the completion of asynchronous data processing in the log.

[0053] In addition, monitoring is an essential and important part of actual production. In this invention, monitoring can have three dimensions: monitoring Redis memory size, monitoring Kafka cluster message consumption, and monitoring the status of database data persistence.

[0054] In this embodiment, the synchronous-to-asynchronous data conversion design significantly reduces the time consumed by the synchronous process, effectively meeting the low-latency system requirements. Simultaneously, under the same conditions, concurrency and throughput are also greatly improved. A three-layer data recovery mechanism enhances the robustness of asynchronous processing, ensuring no data loss. Modular system design improves the efficiency of parallel development and troubleshooting for complex systems. A unified design for the core restore process reduces the difficulty of collaboration between complex systems and improves data recovery quality.

[0055] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A design method for solving the high concurrency problem of backend systems using asynchronous processes, characterized in that, Includes the following steps: Step 1: Build a highly available Kafka and Redis cluster as the data source for asynchronous processes and act as an asynchronous data processing relay station. Install Logstach to collect application logs. Step 2: The data that needs to be stored in the database in the business process is called the first-stage dataset. The first-stage dataset is inserted into the Redis cluster, and the uniqueness of the data key is guaranteed based on the internal serial number. Step 3: Write the first-stage dataset into a log file using a specific JSON format string, which is different from the ordinary log format. Step 4: After processing other normal business logic, complete or update the first-stage dataset based on the processed results to obtain the final dataset, which is called the second-stage dataset. Step 5: Send the second-stage dataset to a specific topic in the Kafka cluster using the producer; this is called the "database removal topic". Step 6: Store the second-stage dataset into the Redis cluster. Its key value must be different from that of the first-stage dataset. The uniqueness of each data item is ensured based on the external serial number. Step 7: The Kafka cluster consumer consumes the topic to be retrieved from the database and asynchronously executes the data write-to-database process; Step 8: Periodically review whether the data has been successfully stored in the Redis cluster based on the first-phase dataset. Step 9: Periodically review the first-stage dataset collected based on Logstach logs to check whether the data has been successfully stored in the database.

2. The design method for solving the high concurrency problem of backend systems using asynchronous processes according to claim 1, characterized in that: The middleware cluster needs to be used independently and should not be shared with other application systems.

3. The design method for solving the high concurrency problem of backend systems using asynchronous processes according to claim 1, characterized in that: The asynchronous process of the Kafka cluster is used to perform the main asynchronous data storage task, the Redis cluster undertakes the second layer of data protection task, and Logstach undertakes the third layer of data protection task.

4. The design method for solving the high concurrency problem of backend systems using asynchronous processes according to claim 3, characterized in that: The three-tiered data recovery mechanism needs to work together to ensure data integrity and prevent loss.

5. The design method for solving the high concurrency problem of backend systems using asynchronous processes according to claim 4, characterized in that: The three-tiered data recovery mechanism is executed in staggered shifts as a fallback, but the data storage process must use the same set of code; this process is called the restore process.

6. The design method for solving the high concurrency problem of backend systems using asynchronous processes according to claim 1, characterized in that: Producers and consumers in a Kafka cluster need to run in separate processes; they cannot run in a single process using multiple threads.

7. The design method for solving the high concurrency problem of backend systems using asynchronous processes according to claim 5, characterized in that: The first-stage dataset stored in Redis is used for the restore process, while the second-stage dataset is used for remedial measures in case of Kafka message loss. It is a fault tolerance mechanism for complex systems.

8. The design method for solving the high concurrency problem of backend systems using asynchronous processes according to claim 1, characterized in that: The specific JSON format string mentioned in step three is as follows: serialize the first-stage dataset into JSON format and add a specific log identifier before the serialized string.

Citation Information

Patent Citations

  • System and method used for processing multi-phase distributed task scheduling

    CN103647834A

  • MongoDB log collection and analysis system based on Kafka message queue

    CN109241187A