A data processing method and architecture
By introducing a distributed publishing and subscription message system and a key-value storage system into the data processing system, the orderly processing is carried out according to the data type, the problem of unguaranteed data consumption order in the prior art is solved, and the accuracy and reliability of data processing are improved.
Patent Information
- Application Number
- CN202210088122.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-25
AI Technical Summary
In the prior art, consumers randomly read data from message middleware, resulting in unguaranteed data consumption order, which may lead to data loss and other problems, reducing the accuracy and reliability of data processing.
The data to be processed is stored through the distributed publishing and subscription message system, and is divided into incremental type, deletion type and modified type packet data according to the preset data type. The application server system reads these data and judges its type. If it is an increase type, it will be consumed directly; if it is a delete type or a modified type, it will be sorted and cached through the key-value storage system to ensure orderly consumption.
By judging data types and sorting processing, we ensure the orderliness of data consumption, avoid problems such as data loss, and improve the accuracy and reliability of data processing.
Smart Images

Figure CN114416717B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing technology, and in particular to a data processing method and architecture. Background Art
[0002] With the continuous development of data processing technology, for the processing of large amounts of data, the prior art usually adopts adding a message middleware between the producer and the consumer to temporarily cache the large amount of data produced by the data producer, in case the consumer cannot consume the large amount of data for a while. Among them, the producer and the consumer are interrelated, the producer is used to produce data, and the consumer is used to consume the data produced by the producer. However, since the consumer reads each piece of data randomly from the message middleware, and the consumer has different consumption execution time or execution efficiency for each piece of data, the consumption order of each piece of data cannot be guaranteed, which leads to data loss and the like during the data processing process, reducing the accuracy and reliability of data processing. Summary of the invention
[0003] The embodiments of the present invention provide a data processing method and architecture to improve the accuracy and reliability of data processing.
[0004] According to one aspect of the present invention, there is provided a data processing method, comprising:
[0005] The data to be processed is stored in a distributed publish-subscribe message system, wherein the data to be processed is divided into add-type message data, delete-type message data and change-type message data according to preset data types;
[0006] Reading the data to be processed in the distributed publish-subscribe message system through the application server system, and determining the data type of the data to be processed;
[0007] If the data type of the data to be processed is an incremental type, the data to be processed is the incremental type message data, and the data to be processed is consumed by at least one consumer in the application server system, and the obtained first consumption data is stored in the target database;
[0008] If the data type of the data to be processed is a deletion type or a modification type, then the data to be processed is the deletion type message data or the modification type message data, and the data to be processed is sorted and cached by a key-value storage system, and based on the results of the sorting and caching processing, at least one consumer in the application server system reads the data to be processed from the key-value storage system and performs consumption processing, and stores the obtained second consumption data in the target database.
[0009] According to another aspect of the present invention, there is provided a data processing architecture, comprising: a distributed publish-subscribe messaging system, a key-value storage system, and an application server system;
[0010] The data processing architecture is used to execute any data processing method described in the embodiments of the present invention.
[0011] The technical solution of the embodiment of the present invention is as follows: first, the data to be processed is stored in a distributed publish-subscribe message system, and the data to be processed is divided into increment type message data, delete type message data and change type message data according to preset data types; then, the data to be processed in the distributed publish-subscribe message system is read by an application server system, and the data type of the data to be processed is determined; finally, if the data type of the data to be processed is increment type, the data to be processed is increment type message data, and at least one consumer in the application server system performs consumption processing on the data to be processed, and the obtained first consumption data is stored in a target database; if the data type of the data to be processed is delete type or change type, the data to be processed is delete type message data or change type message data, and the data to be processed is sorted and cached by a key-value storage system, and at least one consumer in the application server system reads the data to be processed from the key-value storage system based on the results of the sorting and caching processing and performs consumption processing, and the obtained second consumption data is stored in the target database. This technical solution can process the corresponding data to be processed differently according to different data types by judging the data type to be processed by the application server system; it can also sort and cache the deletion type message data or the modification type message data through the key-value storage system, so that the application server system can consume the data to be processed in an orderly manner to avoid data loss during the data processing process, thereby improving the accuracy and reliability of data processing.
[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0014] Figure 1 is a schematic diagram of implementing a data processing method provided according to an embodiment of the present invention;
[0015] Figure 2 is a flow chart of a data processing method provided according to Embodiment 1 of the present invention;
[0016] Figure 3 is a flow chart of a data processing method provided according to Embodiment 2 of the present invention;
[0017] Figure 4 is a schematic diagram of implementing a data processing method provided according to Embodiment 2 of the present invention;
[0018] Figure 5 It is a structural diagram of a data processing architecture provided according to Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0021] In actual applications, in order to meet the flexible and changeable user data display needs of the application front end, considering that the application programming interface (API) real-time call to external system services and then secondary processing is a bit cumbersome, the real-time streaming technology can be used to directly obtain the original data, which can accumulate user data more directly and can more flexibly meet business needs.
[0022] With the continuous development of business, such as in the banking scenario, the number of users increases, and the number of transactions handled by users every day increases and changes, thus generating a large amount of variable data. However, the method of processing data through multi-partition message queues cannot currently guarantee that each variable data is delivered to consumers in order for consumption processing under the high number of transactions per second (TPS). In order to ensure the normal display of user-related business data and protect user rights, the timing and security requirements of data processing become more important.
[0023] The embodiment of the present invention proposes an incremental subscription and consumption component based on Binlog of a relational database (such as a Mysql database), wherein Binlog can be considered as a binary file for recording structured query statement information (i.e., SQL statement information) updated by the user on the database. After the producer's database starts the Binlog mechanism, the change data of the subscribed database will be written into the Binlog file. After the main library of the Mysql database (i.e., Mysql Master) interacts with the data synchronization framework (such as the Canal framework) through an interaction protocol, it starts to push the Binloig incremental log into the Canal framework. The Canal can parse the incremental log to obtain the structure and data changes of the main library of the Mysql database. The producer sends the Binlog log message subscribed to the Mysql database to a distributed publishing and subscription system (such as a Kafka cluster). A topic (i.e., Topic) in the Kafka cluster can include multiple partitions (i.e., Partitions), and a consumer can correspond to a Partition; the consumer can consume the message data in the Paitition in a multi-threaded manner. The threads start to execute or have different efficiencies, which may cause the order of consumer consumption data to be unguaranteed. In order to solve the problem of the timing of consumption data and improve consumer throughput, the consumer side can create multiple memory queues, and the data that needs to be sequentially routed can be routed to the same memory queue. Each consumer can include N memory queues. After consuming data from the partition of the Kafka cluster, the consumer saves the data with the same key (i.e. Key) to the same memory queue. Then, each consumer can open N threads, and each thread consumes only one memory queue to ensure the sequentiality of data consumption. It should be noted that the prerequisite for the above solution is that the producer data cached in the Kafka cluster is ordered.
[0024] Figure 1 FIG. 1 is a schematic diagram of implementing a data processing method according to an embodiment of the present invention. Figure 1As shown, the producer sends a set of ordered data (i.e., data 1, data 2, and data 3) produced to the cache in the same partition 1 in the Kafka cluster; consumer 1 corresponding to partition 1 reads the set of ordered data from partition 1 to the local memory queue 1, and the consumer thread corresponding to memory queue 1 (i.e., thread 1) is used to process the data and send the processed data to the corresponding database.
[0025] Embodiment 1
[0026] Figure 2 1 is a flowchart of a data processing method provided according to Embodiment 1 of the present invention. This embodiment is applicable to the case of orderly processing of a large amount of data. The method can be executed by a data processing architecture, which can be implemented in the form of hardware and / or software. Figure 2 As shown, the method includes:
[0027] S110. Storing the data to be processed through a distributed publish-subscribe message system, wherein the data to be processed is divided into add-type message data, delete-type message data and change-type message data according to preset data types.
[0028] In this embodiment, the distributed publish-subscribe message system may refer to a high-throughput message middleware for caching and transmitting data, which can be understood as a pipeline for transmitting real-time data and caching data, which can process data asynchronously and has high throughput data performance. Exemplarily, the distributed publish-subscribe message system may be a Kafka cluster. When the performance of a server is not enough to support the processing of certain tasks, one or more servers need to be added to jointly process the task, and the work content and process of these servers can be exactly the same; at this time, the collection of these servers can be considered to refer to a cluster; when a server in the cluster has a problem, the other servers can still work normally. The Kafka cluster can include multiple topics, which can be understood as classifying the received data, and each type of data can be called a topic; each topic can include multiple partitions, and each partition can include a certain storage space for storing corresponding data. The topic classification and the number of partitions in the Kafka cluster are not limited here, and can be flexibly set according to actual needs.
[0029] The data to be processed may refer to the data waiting to be processed obtained from the source database. The source database may refer to the database that provides the data to be processed, wherein the database may be understood as a warehouse established on a computer storage device that organizes, stores and manages data according to a data structure. The data may be divided into three corresponding data types (i.e., preset data types, such as add type, delete type and modify type) according to the operations of adding, deleting and modifying the data in the database, i.e., add type message data, delete type message data and modify type message data. Among them, the add operation may refer to adding data to the database, the delete operation may refer to deleting data in the database, and the modify operation may refer to modifying or updating data in the database.
[0030] The acquired data to be processed can be stored through the distributed publish-subscribe message system, wherein the data to be processed can be divided into three categories according to preset data types, namely, add-type message data, delete-type message data and change-type message data.
[0031] In one embodiment, before the data to be processed is stored in the distributed publish-subscribe messaging system, the data obtained from the source database may be converted into a data format. It is understandable that if the format of the data stored in the distributed publish-subscribe messaging system is different from the format of the data obtained from the source database, in order to ensure that the distributed publish-subscribe messaging system can store the data obtained from the source database, the data obtained from the source database may be converted into a format according to the data format corresponding to the distributed publish-subscribe messaging system to obtain the data to be processed that can be stored in the distributed publish-subscribe messaging system, that is, the data to be processed may be understood as data obtained by converting the data obtained from the source database into a format according to the data format of the distributed publish-subscribe messaging system. Among them, the data obtained from the source database may be a database operation log message, and the database operation log message may be understood as information data formed by recording historical operations on the data in the database, that is, the historical operation information on the data in the database within a certain period of time may be obtained through the database operation log message.
[0032] S120: Read the data to be processed in the distributed publish-subscribe message system through an application server system, and determine the data type of the data to be processed.
[0033] In this embodiment, the application server system may refer to a system for consuming and processing data to be processed; the application server system may be a system composed of multiple servers, each of which may be understood as a consumer for consuming and processing data. When the corresponding consumption process is triggered (the consumption process may be understood as a process for consuming and processing the data to be processed), the data to be processed cached in the distributed publish-subscribe message system may be read by the application server system. On this basis, after reading the data to be processed, the data type to which the data to be processed belongs may also be determined, so as to perform different data processing operations according to different data types, wherein the data type may include an add type, a delete type, and a modify type. The manner in which the application server system determines the data type to which the data to be processed belongs is not limited here.
[0034] S130. If the data type of the data to be processed is an incremental type, the data to be processed is the incremental type message data. The data to be processed is consumed by at least one consumer in the application server system, and the obtained first consumption data is stored in the target database, and the operation ends.
[0035] In this embodiment, if the data type of the data to be processed that is read is determined to be an increment type by the application server system, it can be determined that the data to be processed is an increment type message data. It can be understood that the increment type message data can indicate that an increment operation is performed on the data in the database, and since the increment operation on the data in the database is insensitive to the message time of the original data in the database, at this time, the data to be processed can be directly consumed by at least one consumer in the application server system, and the data after consumption processing (i.e., the first consumption data) is stored in the target database. Among them, the message time of the original data can refer to the data operation log message time, which can be understood as the time corresponding to the data operation when the original data was generated.
[0036] The target database and the source database are two corresponding concepts. The target database can be understood as a database that receives data read from the source database. On this basis, the data processing of this embodiment can be understood as data transfer processing between the source database and the target database. The target database can be considered as a database running on the consumer.
[0037] It is understandable that the data format of the target database may be different from the data format of the data to be processed. Therefore, before storing the first consumption data in the target database, the data format of the first consumption data may be converted to the data format of the target database so that the first consumption data after format conversion can be stored in the target database.
[0038] S140. If the data type of the data to be processed is a deletion type or a modification type, the data to be processed is the deletion type message data or the modification type message data, and the data to be processed is sorted and cached by a key-value storage system. Based on the results of the sorting and caching, at least one consumer in the application server system reads the data to be processed from the key-value storage system and performs consumption processing, and stores the obtained second consumption data in the target database.
[0039] In this embodiment, if the data type of the read data to be processed is determined by the application server system to be a delete type or a change type, it can be determined that the data to be processed is a delete type message data or a change type message data. It is understandable that the delete type message data or the change type message data can indicate operations such as deleting or modifying and updating the data in the database, and since the deletion or modification and update operations on the data in the database are sensitive to the message time of the original data in the database, for example, the update of the original data requires the message time corresponding to the original data to update the new data based on this message time, so at this time, the key value storage system can be used to sort and cache the data to be processed. The key value storage system can refer to a cache middleware that can support multiple key value type storage databases, and the key value storage system can include multiple key value storage nodes (nodes can be understood as servers that can support multiple key value type storage databases), and multiple nodes can share data and interact with each other. Exemplarily, the key value storage system can be a Redis cluster, and the Redis cluster can include multiple Redis nodes, which is not limited here.
[0040] It should be noted that each piece of data to be processed may correspond to a message time. On this basis, sorting can be understood as sorting all the data to be processed according to the order of the message time corresponding to each piece of data to be processed, using the ordered set type in the key-value storage system. The ordered set type can be represented as the Zest data type, and the function of the Zest data type can be understood as being able to arrange the data in a certain order.
[0041] Specifically, in this embodiment, the weight value of each piece of data to be processed is set according to the ordered set type and the order of the message time corresponding to each piece of data to be processed. The weight value can be used to characterize the message time of the piece of data to be processed, such as whether the message time is earlier or later. The weight value of the data to be processed with an earlier message time can be set to be smaller than the weight value of the data to be processed with a later message time. For example, assuming that there are 3 pieces of data to be processed, and the corresponding message times are 2 o'clock, 2:02 and 2:05 respectively, the weight value corresponding to the data to be processed with a message time of 2 o'clock can be set to be lower than the weight value of the data to be processed with a message time of 2:02, and the weight value corresponding to the data to be processed with a message time of 2:02 is lower than the weight value of the data to be processed with a message time of 2:05. On this basis, the corresponding data to be processed can be sorted in the order of the weight value from small to large. The specific numerical setting of the weight value is not limited here, and can be flexibly set according to actual needs. And after sorting the data to be processed, the data to be processed is cached to wait for the reading and consumption of the application server system. It should be noted that the data to be processed can be cached in at least one message queue pre-set in the key-value storage system according to the sorting result, and the number of consumers in the application server system can be set to be the same as the number of the message queues.
[0042] Through at least one consumer in the application server system, the data to be processed is read from the key-value storage system and consumed based on the result of sorting and caching the data to be processed, and the obtained second consumption data is stored in the target database. It should be noted that when reading the data to be processed from the key-value storage system, since it is read in real time, the latest data to be processed may not have completed sorting and caching. Therefore, in order to ensure that the data to be processed read is sorted data, a threshold interval can be set in advance, and only the data to be processed within the threshold interval is read from the key-value storage system; wherein the time of the right end point of the threshold interval is less than the current time. Exemplarily, assuming that the current time is 3:05, when reading the data to be processed from the key-value storage system in real time, in order to ensure that the data to be processed read is sorted data, the current time can be used as a reference, and only the data to be processed within the threshold interval between 2:50 and 3:00 can be read from the key-value storage system. By analogy, if the current time is 4 o'clock, the current time can be used as a reference, and only the data to be processed between 3:40 and 3:55 can be read from the key-value storage system. There is no limit on the size of the threshold interval here, and it can be flexibly set according to actual needs.
[0043] A data processing method provided in a first embodiment of the present invention comprises the following steps: firstly, storing data to be processed through a distributed publish-subscribe message system, wherein the data to be processed are divided into increment type message data, delete type message data and change type message data according to preset data types; then, reading the data to be processed in the distributed publish-subscribe message system through an application server system, and determining the data type of the data to be processed; finally, if the data type of the data to be processed is increment type, the data to be processed is increment type message data, and at least one consumer in the application server system performs consumption processing on the data to be processed, and stores the obtained first consumption data in a target database; if the data type of the data to be processed is delete type or change type, the data to be processed is delete type message data or change type message data, and sorting and caching the data to be processed through a key-value storage system, and reading and consuming the data to be processed from the key-value storage system based on the results of the sorting and caching processing through at least one consumer in the application server system, and storing the obtained second consumption data in the target database. The method determines the data type of the data to be processed by the application server system, and can perform different processing on the corresponding data to be processed according to different data types; it also sorts and caches the deletion type message data or the modification type message data through the key-value storage system, which can enable the application server system to consume the data to be processed in an orderly manner to avoid data loss during the data processing process, thereby improving the accuracy and reliability of data processing.
[0044] Embodiment 2
[0045] Figure 3 is a flow chart of a data processing method provided according to the second embodiment of the present invention, and the second embodiment is refined on the basis of the above embodiments. In this embodiment, the acquisition of the data to be processed, the sorting and caching of the data to be processed, and the processing after the first consumption data or the second consumption data is stored in the target database are described in detail. It should be noted that the technical details not described in detail in this embodiment can be referred to any of the above embodiments. Figure 3 As shown, the method includes:
[0046] S210: Convert the database operation log message obtained from the source database according to a preset format to obtain data to be processed.
[0047] In this embodiment, the preset format may refer to the data format corresponding to the cache data of the distributed publish-subscribe messaging system. The database operation log message obtained from the source database is converted into a data format according to the preset format, and the converted data is the data to be processed.
[0048] S220, storing the data to be processed in a distributed publish-subscribe messaging system.
[0049] In this embodiment, the data format of the data to be processed is the same as the data format corresponding to the distributed publish-subscribe messaging system, so the data to be processed can be stored in the distributed publish-subscribe messaging system. The data to be processed can be divided into add-type message data, delete-type message data and change-type message data according to the preset data type.
[0050] S230: Read the data to be processed in the distributed publish-subscribe message system through the application server system, and determine the data type of the data to be processed.
[0051] In this embodiment, the data to be processed in the distributed publish-subscribe message system is read by the application server system, and the data type of the data to be processed is determined according to the read data to be processed, wherein the data type may include an add type, a delete type and a modify type.
[0052] S240: If the data type of the data to be processed is an increment type, the data to be processed is increment type message data, and the data to be processed is consumed by at least one consumer in the application server system, and the obtained first consumption data is stored in the target database.
[0053] In this embodiment, if the data type of the data to be processed is an incremental type, it can be determined that the data to be processed is incremental type message data. At this time, the data to be processed can be consumed by at least one consumer in the application server system, and the obtained first consumption data can be stored in the target database.
[0054] S250. If the data type of the data to be processed is a delete type or a change type, the data to be processed is a delete type message data or a change type message data. The data to be processed is sorted based on the ordered set type in the key-value storage system and the message time corresponding to each piece of data to be processed, and the data to be processed is cached in a preset message queue according to the sorting result.
[0055] In this embodiment, if the data type of the data to be processed is a delete type or a change type, it can be determined that the data to be processed is delete type message data or change type message data; on this basis, the data to be processed can be sorted based on the ordered set type in the key-value storage system and the message time corresponding to each piece of data to be processed, and the data to be processed can be cached in a preset message queue in the key-value storage system according to the sorting result.
[0056] Optionally, the number of preset message queues is the same as the number of consumers in the application server system.
[0057] In order to ensure that the preset message queues in the key-value storage system correspond one-to-one to the consumers, the number of the preset message queues may be set to be the same as the number of consumers in the application server system.
[0058] Optionally, the data to be processed is sorted based on the ordered set type in the key-value storage system and the message time corresponding to each piece of data to be processed, including: for each piece of data to be processed, a corresponding weight value is set according to the order of the message time based on the ordered set type, wherein the weight value of the data to be processed with an earlier message time is less than the weight value of the data to be processed with a later message time; and the corresponding data to be processed is sorted from small to large according to the weight value.
[0059] Among them, for each piece of data to be processed, a corresponding weight value can be set according to the order of message time based on the ordered set type, wherein the weight value of the data to be processed with an earlier message time can be smaller than the weight value of the data to be processed with a later message time, such as the weight value of the data to be processed with a message time of 2 o'clock is smaller than the weight value of the data to be processed with a message time of 2:05; on this basis, the corresponding data to be processed can be sorted in order from small to large in terms of weight value.
[0060] S260. Using at least one consumer in the application server system, based on the results of the sorting and caching processing, read the data to be processed from the key-value storage system and perform consumption processing, and store the obtained second consumption data in the target database.
[0061] In this embodiment, at least one consumer in the application server system can sequentially read the data to be processed from the key-value storage system based on the results of sorting and caching processing and perform consumption processing, and store the obtained second consumption data in the target database.
[0062] Optionally, the data to be processed is read from the key-value storage system based on the results of sorting and caching processing, including: according to the results of sorting and caching processing, reading the data to be processed that meets a first preset condition from the key-value storage system, wherein the first preset condition is that the weight value is less than a first preset threshold and greater than a second preset threshold, and the first preset threshold is greater than the second preset threshold.
[0063] Among them, the settings of the first preset threshold and the second preset threshold can be flexibly set according to actual needs. The data to be processed that meets the first preset condition can be understood as the data to be processed whose weight value is between the second preset threshold and the first preset threshold. Among them, the first preset threshold is greater than the second preset threshold. In order to ensure that the data to be processed obtained from the key-value storage system in real time is already sorted, the current time can be used as a reference, and only the data to be processed whose weight value is less than the weight value corresponding to the current time can be read from the key-value storage system. Therefore, it can be understood that the first preset threshold and the second preset threshold in the first preset condition will be less than the weight value corresponding to the current time. On this basis, according to the results of sorting and caching processing, the data to be processed that meets the first preset condition can be read from the key-value storage system.
[0064] Optionally, when reading the data to be processed from the key-value storage system through the application server system, it also includes: verifying the weight value of the data to be processed through the application server system to determine the data to be processed whose weight value is less than the second preset threshold; triggering the key-value storage system to clean up the data to be processed in the local cache whose weight value is less than the second preset threshold.
[0065] Among them, in order to clean up the pending data that has not been consumed for a long time in the key-value storage system and alleviate the cache pressure of the key-value storage system, when reading the pending data from the key-value storage system through the application server system, the weight value of the pending data can be verified by the application server system to determine the pending data whose weight value is less than the second preset threshold; on this basis, the key-value storage system can be triggered to clean up the pending data in the local cache whose weight value is less than the second preset threshold.
[0066] S270: Add a data execution time field to the table of the target database, and determine the message time corresponding to each piece of data to be processed according to the data execution time field.
[0067] In this embodiment, in the database, a table can be understood as a collection of data sets, or as a two-dimensional data table that can store data. On this basis, a field can be understood as a column in the two-dimensional data table, that is, a data set with the same attributes, and each field can correspond to a unique name, which can be called a field name. The data execution time field can be understood as a field that represents the data execution time, where the data execution time can be understood as the message time of the data. Therefore, by adding a data execution time field to the table of the target database, the message time corresponding to each piece of data to be processed can be determined according to the data execution time field.
[0068] S280. For at least one piece of to-be-processed data belonging to the same data type, retain the to-be-processed data that meets the second preset condition according to the message time.
[0069] In this embodiment, when reading the data to be processed from the distributed publish-subscribe message system through the application server system, a reading failure may occur. In this case, during the next reading process, for at least one piece of data to be processed belonging to the same data type, the data to be processed that failed to be read last time and the updated data to be processed at the current time may be read at the same time. At this time, in order to retain the latest data to be processed (i.e., the data to be processed at the current time), the data to be processed that meets the second preset condition may be retained according to the message time of the data to be processed. The second preset condition may refer to the latest message time, i.e., the message time closest to the current time.
[0070] Exemplarily, for at least one piece of unprocessed data belonging to the same data type, assuming that the current time is 3:20, a piece of unprocessed data with a message time of 3:10 and a piece of unprocessed data with a message time of 3:15 are simultaneously read through the application server system. At this time, the unprocessed data with a message time of 3:15 can be retained (that is, the unprocessed data with the latest message time and closest to the current time of 3:20) is retained.
[0071] Optionally, for each piece of data to be processed, if the application server system detects that abnormal information is generated in the process of the data to be processed entering the key-value storage system, the abnormal information is transmitted to the distributed publish-subscribe message system, and the data to be processed is sent to the key-value storage system at preset intervals through the application server system until no abnormal information is detected.
[0072] The abnormal information may refer to information that the data to be processed cannot enter the key-value storage system normally. The abnormal information may be caused by the fact that the disk of the key-value storage system is full, the memory is full, or the network is abnormal, so that the data to be processed cannot enter the key-value storage system. The specific cause of the abnormal information is not limited here, as long as it is the reason that causes the data to be processed to be unable to enter the key-value storage system normally.
[0073] For each piece of data to be processed, if the application server system detects that abnormal information is generated during the process of the data to be processed entering the key-value storage system, the abnormal information can be transmitted to the distributed publish-subscribe message system for corresponding processing, and the application server system sends the data to be processed to the key-value storage system at preset intervals until no abnormal information is detected (i.e., the data to be processed is sent to the key-value storage system successfully). The preset time interval is to avoid excessive pressure on the key-value storage system to receive data due to frequent data transmission.
[0074] Embodiment 2 of the present invention provides a data processing method, which specifically describes the process of obtaining the data to be processed, sorting and caching the data to be processed, and processing after the first consumption data or the second consumption data is stored in the target database. This method can store the converted data to be processed in a distributed publish-subscribe message system by converting the data format; it also sorts the data to be processed according to the ordered set type in the key-value storage system and the message time of each data to be processed, so that the application server system can consume the data to be processed in order according to the sorting result, avoiding the loss of data during processing; in addition, the weight value of the data to be processed is verified by the application server system, so that the data to be processed that has not been consumed for a long time in the key-value storage system can be cleaned up to relieve the storage pressure of the key-value storage system.
[0075] In a specific embodiment, Figure 4 FIG. 1 is a schematic diagram of implementing a data processing method according to Embodiment 2 of the present invention. Figure 4 As shown, the specific implementation process of this method is as follows:
[0076] First, the data produced by the producer (i.e., the data to be processed) can be obtained by converting the database operation log messages obtained from the source database according to a preset format. The data to be processed can include add type message data, delete type message data, and change type message data.
[0077] Then, the data to be processed is stored in each partition (i.e., partition 1, partition 2, ..., partition n) in the distributed publish-subscribe messaging system (such as the Kafka cluster); consumers in the application server system can read the data to be processed in the distributed publish-subscribe messaging system in real time through the corresponding API interface and determine the data type of the data to be processed:
[0078] If the data type of the data to be processed is an incremental type, the data to be processed is incremental type message data, and the data to be processed is consumed by at least one consumer in the application server system, and the obtained first consumption data is stored in the target database.
[0079] If the data type of the data to be processed is a delete type or a change type, then the data to be processed is a delete type message data or a change type message data, and the corresponding weight value is set for the data to be processed based on the ordered set type (i.e., Zset data type) in the key-value storage system (such as a Redis cluster) and the operation log message time corresponding to each piece of data to be processed, and the data to be processed is sorted according to the weight value, and the data to be processed is cached in a preset message queue (i.e., preset message queue 1, preset message queue 2, ..., preset message queue n) according to the sorting result. On this basis, at least one consumer in the application server system reads the data to be processed from the key-value storage system based on the results of the sorting and caching processing and performs consumption processing, and the obtained second consumption data is stored in the target database.
[0080] Finally, after the first consumption data or the second consumption data is stored in the target database, the message time corresponding to each piece of data to be processed is determined according to the data execution time field added in the table of the target database, so as to retain the data to be processed that meets the second preset condition according to the message time for at least one piece of data to be processed belonging to the same data type.
[0081] It should be noted that when reading the data to be processed from the key-value storage system based on the results of sorting and caching, a delayed consumption time can be set, that is, based on the results of sorting and caching, the data to be processed that meets the first preset condition is read from the key-value storage system, where the first preset condition is that the weight value is less than the first preset threshold and greater than the second preset threshold.
[0082] The number of preset message queues and the number of consumers in the application server system can be set to be the same to improve the consumption efficiency of the consumers.
[0083] When reading the data to be processed from the key-value storage system through the application server system, in order to discard data that has not been consumed for a long time and prevent the Redis cluster from being blocked due to consumption anomalies, the weight value of the data to be processed can also be verified through the application server system to determine the data to be processed whose weight value is less than the second preset threshold, and trigger the key-value storage system to clean up the data to be processed in the local cache whose weight value is less than the second preset threshold.
[0084] For each piece of data to be processed, if the application server system detects that abnormal information is generated during the process of the data to be processed entering the key-value storage system, the abnormal information can be transmitted to the distributed publish-subscribe message system (such as the Kafka cluster) to stop the Kafka cluster from updating the Kafka Offset (where Kafka Offset can be understood as the amount of data consumed by the consumer during the consumption process, that is, the consumption Offset), and the application server system can be called by the Sleep method to send the data to be processed to the key-value storage system at preset intervals, and will not update the corresponding Kafka Offset in advance until no abnormal information is detected to avoid data loss. The preset time can be determined according to the configured Kafka consumption timeout, which is not limited here.
[0085] If the application server system captures an exception in the Redis cluster, it can repeatedly send the data to be processed to the Redis cluster to prevent the consumption thread in the Kafka cluster from being suspended and stopping consumption due to reasons such as a flash failure of the Redis cluster.
[0086] The upper limit of the data to be processed that the application server system obtains from the Redis cluster each time can be limited to avoid a large amount of changed data entering the Redis cluster cache and the corresponding server's memory overflow caused by consumers obtaining it at one time.
[0087] Embodiment 3
[0088] Figure 5 Schematic diagram of a data processing architecture provided according to Embodiment 3 of the present invention. Figure 5 As shown, the architecture includes: a distributed publish-subscribe messaging system 310, an application server system 320, and a key-value storage system 330;
[0089] The data processing architecture is used to execute any data processing method described in the embodiments of the present invention.
[0090] Optionally, in the architecture,
[0091] A distributed publish-subscribe messaging system 310, used to store data to be processed, wherein the data to be processed is divided into add-type message data, delete-type message data and change-type message data according to preset data types;
[0092] The application server system 320 is used to read the data to be processed in the distributed publish-subscribe message system 310 and determine the data type of the data to be processed;
[0093] If the data type of the data to be processed is an incremental type, the data to be processed is the incremental type message data, and at least one consumer in the application server system 320 consumes the data to be processed, and stores the obtained first consumption data in the target database;
[0094] If the data type of the data to be processed is a deletion type or a modification type, then the data to be processed is the deletion type message data or the modification type message data, and the data to be processed is sorted and cached by the key-value storage system 330, and the data to be processed is read from the key-value storage system 330 and consumed based on the results of the sorting and caching by at least one consumer in the application server system 320, and the obtained second consumption data is stored in the target database.
[0095] A data processing architecture is provided in the third embodiment of the present invention. First, the data to be processed is stored through a distributed publish-subscribe message system 310. The data to be processed is divided into increment type message data, delete type message data and change type message data according to preset data types; then, the data to be processed in the distributed publish-subscribe message system 310 is read through the application server system 320, and the data type of the data to be processed is determined; finally, if the data type of the data to be processed is increment type, the data to be processed is increment type message data, and at least one consumer in the application server system 320 consumes the data to be processed, and the obtained first consumption data is stored in a target database; if the data type of the data to be processed is delete type or change type, the data to be processed is delete type message data or change type message data, and the data to be processed is sorted and cached through a key-value storage system 330, and at least one consumer in the application server system 320 reads the data to be processed from the key-value storage system 330 based on the results of the sorting and caching processing and consumes the data, and the obtained second consumption data is stored in the target database. This architecture determines the data type of the data to be processed by the application server system 320, and can perform different processing on the corresponding data to be processed according to different data types; it also sorts and caches the deletion type message data or the modification type message data through the key-value storage system 330, which can enable the application server system 320 to consume the data to be processed in an orderly manner, thereby avoiding data loss during the data processing process, thereby improving the accuracy and reliability of data processing.
[0096] Optionally, the data to be processed is stored in a distributed publish-subscribe messaging system 310, including:
[0097] Convert the database operation log message obtained from the source database according to a preset format to obtain data to be processed;
[0098] The data to be processed is stored in the distributed publish-subscribe message system 310 .
[0099] Optionally, the key-value storage system 330 performs sorting and caching, including:
[0100] The data to be processed are sorted based on the ordered set type in the key-value storage system 330 and the message time corresponding to each piece of data to be processed, and the data to be processed are cached in a preset message queue according to the sorting result.
[0101] Optionally, sorting the data to be processed based on the ordered set type in the key-value storage system 330 and the message time corresponding to each piece of data to be processed includes:
[0102] For each piece of data to be processed, a corresponding weight value is set according to the order of the message time based on the ordered set type, wherein the weight value of the data to be processed with an earlier message time is less than the weight value of the data to be processed with a later message time;
[0103] The corresponding data to be processed are sorted in ascending order according to the weight values.
[0104] Optionally, reading the to-be-processed data from the key-value storage system 330 based on the results of the sorting and caching processing includes:
[0105] According to the results of sorting and caching processing, the data to be processed that meets the first preset condition is read from the key-value storage system 330, wherein the first preset condition is that the weight value is less than a first preset threshold and greater than a second preset threshold, and the first preset threshold is greater than the second preset threshold.
[0106] Optionally, the number of preset message queues is the same as the number of consumers in the application server system 320 .
[0107] Optionally, after storing the first consumption data or the second consumption data in the target database, the method further includes:
[0108] Adding a data execution time field to the table of the target database, and determining the message time corresponding to each piece of data to be processed according to the data execution time field;
[0109] For at least one piece of to-be-processed data belonging to the same data type, the to-be-processed data that meets the second preset condition is retained according to the message time.
[0110] Optionally, when the to-be-processed data is read from the key-value storage system 330 by the application server system 320, the method further includes:
[0111] Verifying the weight value of the data to be processed by the application server system 320, and determining the data to be processed whose weight value is less than the second preset threshold;
[0112] The key-value storage system 330 is triggered to clean up the data to be processed in the local cache whose weight value is less than the second preset threshold.
[0113] Optionally, for each piece of data to be processed, if it is detected through the application server system 320 that abnormal information is generated in the process of the data to be processed entering the key-value storage system 330, the abnormal information is transmitted to the distributed publish-subscribe message system 310, and the data to be processed is sent to the key-value storage system 330 at preset intervals through the application server system 320 until the abnormal information is no longer detected.
[0114] The data processing architecture provided in the embodiment of the present invention can execute the data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0115] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0116] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A data processing method, characterized in that: The method comprises: The data to be processed is stored in a distributed publish-subscribe message system, wherein the data to be processed is divided into add-type message data, delete-type message data and change-type message data according to preset data types; Reading the data to be processed in the distributed publish-subscribe message system through the application server system, and determining the data type of the data to be processed; If the data type of the data to be processed is an incremental type, the data to be processed is the incremental type message data, and the data to be processed is consumed by at least one consumer in the application server system, and the obtained first consumption data is stored in the target database; If the data type of the data to be processed is a deletion type or a modification type, the data to be processed is the deletion type message data or the modification type message data, the data to be processed is sorted and cached by a key-value storage system, and the data to be processed is read from the key-value storage system and consumed by at least one consumer in the application server system based on the results of the sorting and caching processing, and the obtained second consumption data is stored in the target database; The sorting and caching process is performed by the key-value storage system, including: Sorting the data to be processed based on the ordered set type in the key-value storage system and the message time corresponding to each piece of data to be processed, and caching the data to be processed into a preset message queue according to the sorting result; The step of sorting the data to be processed based on the ordered set type in the key-value storage system and the message time corresponding to each piece of data to be processed includes: For each piece of data to be processed, a corresponding weight value is set according to the order of the message time based on the ordered set type, wherein the weight value of the data to be processed with an earlier message time is less than the weight value of the data to be processed with a later message time; The corresponding data to be processed are sorted in ascending order according to the weight values.
2. The method according to claim 1, characterized in that The method of storing the data to be processed by a distributed publish-subscribe message system includes: Convert the database operation log message obtained from the source database according to a preset format to obtain data to be processed; The data to be processed is stored in the distributed publish-subscribe messaging system.
3. The method according to claim 1, characterized in that The step of reading the to-be-processed data from the key-value storage system based on the results of the sorting and caching processing includes: According to the results of sorting and caching processing, the data to be processed that meets the first preset condition is read from the key-value storage system, wherein the first preset condition is that the weight value is less than a first preset threshold and greater than a second preset threshold, and the first preset threshold is greater than the second preset threshold.
4. The method according to claim 1, characterized in that: The number of the preset message queues is the same as the number of consumers in the application server system.
5. The method according to claim 1, characterized in that After storing the first consumption data or the second consumption data in the target database, the method further includes: Adding a data execution time field to the table of the target database, and determining the message time corresponding to each piece of data to be processed according to the data execution time field; For at least one piece of to-be-processed data belonging to the same data type, the to-be-processed data that meets the second preset condition is retained according to the message time.
6. The method according to claim 3, characterized in that When the data to be processed is read from the key-value storage system by the application server system, the method further includes: Verifying the weight value of the data to be processed by the application server system, and determining the data to be processed whose weight value is less than the second preset threshold; The key-value storage system is triggered to clean up the to-be-processed data in the local cache whose weight value is less than the second preset threshold.
7. The method according to claim 1, characterized in that The method further comprises: For each piece of data to be processed, if the application server system detects that abnormal information is generated in the process of the data to be processed entering the key-value storage system, the abnormal information is transmitted to the distributed publish-subscribe message system, and the data to be processed is sent to the key-value storage system through the application server system at preset intervals until the abnormal information is not detected.
8. A data processing architecture, characterized in that: include: Distributed publish-subscribe messaging system, key-value storage system, and application server system; The data processing architecture is used to execute the data processing method described in any one of claims 1-7.
Citation Information
Patent Citations
Message processing system and processing method based on kafka
CN104754036A
Data synchronization method
CN108197263A