Data processing method and device

Decoupling data processing and storage in real-time systems through separate message queues improves stability and reduces development costs by allowing flexible task management and efficient handling of dimension changes.

CN120295768APending Publication Date: 2025-07-11BEIJING JINGDONG YUANSHENG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510330291.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In existing real-time data projects, dimensional data occupies a large amount of storage and is frequently adjusted, resulting in frequent modifications of real-time computing tasks, increasing developer workload and reducing task stability.

Method used

Decoupling data processing, dimension processing and data storage, processing through data message queues, including data processing and dimension processing, and using dimension supplement operators and data merging rules to achieve data decoupling processing.

Benefits of technology

Improve the performance and stability of computing tasks, reduce the workload of developers, reduce development costs, and improve data processing speed by changing space and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295768A_ABST
    Figure CN120295768A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the following steps: acquiring a data message, wherein the data message indicates source data; performing data processing on the source data to obtain first data corresponding to the source data, and sending the first data to a first message queue; obtaining the first data from the first message queue, and performing dimension processing on the first data to obtain second data corresponding to the first data; and sending the second data to a second message queue so as to process the second data through the second message queue. According to the implementation mode, data processing, dimension processing and data storage are decoupled, the performance and stability of a calculation task are improved, and the data processing speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular, to a data processing method and apparatus. Background Art

[0002] In existing real-time data projects, generally, a log extraction tool is used to extract log data of a production table to generate data messages of the production table, and then a real-time data calculation project processes the data messages. After calculating the fields required for a wide table, the data is written into a storage. In related technologies, the processing tasks are a combination of logical calculation and dimension calculation. However, dimension data occupies a large amount of storage and is frequently adjusted, which in turn causes frequent modification of real-time calculation tasks, thereby increasing the workload of developers and reducing the stability of real-time calculation tasks. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a data processing method and apparatus, which decouple data processing, dimension processing, and data storage, improve the performance and stability of calculation tasks, and improve the data processing speed.

[0004] To achieve the above object, according to one aspect of embodiments of the present invention, there is provided a data processing method, including:

[0005] Obtaining a data message, where the data message indicates source data;

[0006] Performing data processing on the source data to obtain first data corresponding to the source data, and sending the first data to a first message queue;

[0007] Obtaining the first data from the first message queue, and performing dimension processing on the first data to obtain second data corresponding to the first data;

[0008] Sending the second data to a second message queue to process the second data through the second message queue.

[0009] Optionally, performing dimension processing on the first data to obtain second data corresponding to the first data includes:

[0010] Performing dimension supplementation on the first data by using a dimension supplementation operator to obtain third data; the third data includes each data identifier in the first data and the data corresponding to each data identifier;

[0011] Performing data merging on each piece of data in the third data according to the data identifier to obtain the second data.

[0012] Optionally, performing data merging on each piece of data in the third data according to the data identifier to obtain the second data includes:

[0013] Partition the data using the data identifier to obtain multiple partitions, where each piece of data in each partition has the same data identifier;

[0014] Within a time window, merge the data in the same partition according to a preset merging rule to obtain the merged data corresponding to each partition;

[0015] Obtain the second data based on the merged data corresponding to each partition in each partition.

[0016] Optionally, merging the data in the same partition according to a preset merging rule to obtain the merged data corresponding to each partition includes:

[0017] Obtain the event time of each piece of data in the same partition, and use the data with the latest event time as the merged data corresponding to the partition; or

[0018] Obtain the arrival time of each piece of data in the same partition when it reaches the time window, and use the data with the latest arrival time as the merged data corresponding to the partition.

[0019] Optionally, the first data is obtained by performing data processing on the source data through a data processing task. The first message queue corresponds to the data processing task, and the data processing task is the main task or the backup task; before obtaining the first data from the first message queue, it further includes:

[0020] Determine whether the main task is available;

[0021] When the main task is available, obtain the first data from the first message queue corresponding to the main task using a dimension processing task;

[0022] When the main task is unavailable, obtain the first message queue corresponding to the backup task from the configuration file, and obtain the first data from the first message queue corresponding to the backup task using the dimension processing task.

[0023] Optionally, before obtaining the first message queue corresponding to the backup task from the configuration file, it further includes:

[0024] Store the backup task corresponding to the main task and the first message queue corresponding to the backup task in the configuration file.

[0025] Optionally, processing the second data through the second message queue includes:

[0026] Obtain the second data from the second message queue and store the second data in a database; or

[0027] Enable downstream consumers to consume the second data in the second message queue.

[0028] According to another aspect of the embodiments of the present invention, a data processing apparatus is provided, including:

[0029] An acquisition module, which acquires a data message, and the data message indicates source data;

[0030] A first processing module, which processes the source data to obtain first data corresponding to the source data, and sends the first data to a first message queue;

[0031] A second processing module, which acquires the first data from the first message queue and performs dimensionality processing on the first data to obtain second data corresponding to the first data;

[0032] A sending module, which sends the second data to a second message queue to process the second data through the second message queue.

[0033] According to another aspect of the embodiments of the present invention, an electronic device is provided, including:

[0034] One or more processors;

[0035] A storage device for storing one or more programs,

[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method provided by the present invention.

[0037] According to still another aspect of the embodiments of the present invention, a computer-readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, the data processing method provided by the present invention is implemented.

[0038] One embodiment of the above invention has the following advantages or beneficial effects: In the data processing method of the embodiment of the present invention, first, a data message is obtained, and the source data indicated by the data message is processed to obtain first data; the first data is obtained from the first message queue, and the first data is processed in terms of dimensions to obtain second data, and the second data is sent to the second message queue to process the second data through the second message queue to achieve the storage of the second data. This data processing method decouples data processing, dimensional processing, and data storage, improving the performance and stability of computing tasks; this method reduces the complexity of data processing, storage, and transmission consumption of dimensional data by sinking and processing dimensional data. When the dimensional data changes, at the dimensional processing layer, it is only necessary to update the unified version of the dimensional supplement operator, without paying attention to the changes in dimensions, reducing redevelopment; by adding a sinking layer for dimensional processing, downstream storage or consumers are unaware, improving the processing efficiency of the overall link; this method transfers the message of data storage to a new message queue in a space-for-time manner, improving the data processing speed and reducing the development cost at the same time.

[0039] The further effects of the above non-conventional optional methods will be described in conjunction with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0041] Figure 1 is a schematic diagram of the main process of a data processing method according to an embodiment of the present invention;

[0042] Figure 2 is a schematic diagram of the main process of another data processing method according to an embodiment of the present invention;

[0043] Figure 3 is a schematic diagram of the process of a data processing method according to an embodiment of the present invention;

[0044] Figure 4 is a schematic diagram of the main modules of a data processing device according to an embodiment of the present invention;

[0045] Figure 5 is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;

[0046] Figure 6 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The exemplary embodiments of the present invention will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present invention are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0048] It should be noted that in the technical solutions of the present disclosure, in terms of the collection, gathering, updating, analysis, processing, use, transmission, storage, etc. of the user's personal information, they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data, and to safeguard the security of the user's personal information, network security, and national security.

[0049] Currently, for real-time data projects, log extraction tools are used to extract the log data of production tables to generate data messages of production tables, and then real-time data computing projects, such as Storm, Flink, etc., process the production table messages. After operations such as logical processing, dimension data calculation, indicator calculation, data cleaning, data compression, etc., the fields required for the wide table are calculated and then written into storage, such as ES, Doris, MySQL, etc. In related technologies, the processing tasks are integrated with logical calculation and dimension calculation. If there are adjustments to dimension data, such as site information adjustment, etc., the processing tasks will be modified again. However, due to issues such as different developers and development quality, it may cause the state to fail to start, resulting in data loss; moreover, dimension data will occupy more storage, computing resources, and network bandwidth, thus causing the state to be too large, which will significantly reduce the stability of the task; data processing and data storage are integrated, and the coupling between data processing and data storage is too high. The storage performance cannot keep up with the computing performance, resulting in message delay and version conflicts in storage; the data of the primary and backup storage in the heavy guarantee task is inconsistent.

[0050] Based on the above technical problems, the embodiments of the present invention provide a data processing method, which can reduce the development workload of developers, improve development efficiency, reduce development costs, and improve the performance and stability of tasks.

[0051] Figure 1 is a schematic diagram of the main process of a data processing method according to an embodiment of the present invention, as Figure 1 shown. The data processing method includes the following steps:

[0052] Step S101: Obtain a data message, where the data message indicates source data;

[0053] Step S102: Process the source data to obtain the first data corresponding to the source data, and send the first data to the first message queue;

[0054] Step S103: Obtain the first data from the first message queue, and perform dimensional processing on the first data to obtain the second data corresponding to the first data;

[0055] Step S104: Send the second data to the second message queue to process the second data through the second message queue.

[0056] In the embodiment of the present invention, this data processing method can be applied to a real-time data processing project. First, obtain a data message. The data message can be sent by a producer or by other databases, file systems, or external interfaces. Among them, the producer can be a production table. It is possible to obtain the data message sent by the producer through a message queue, that is, the data message is sent by the producer to the message queue. The data message indicates the source data, that is, the data message can include the source data or the identifier of the source data. The source data includes multiple original fields, and the multiple original fields can be set according to business requirements. For example, the source data can be order data, and the order data includes multiple fields such as user id, order id, product category, and product quantity.

[0057] After obtaining the data message, use a data processing task to process the source data, that is, obtain the first data from the source data. Specifically, processing the source data includes at least one of data cleaning, data conversion, logical processing, rule engine, data statistics, and data compression. Then, send the first data to the first message queue. Among them, the data processing task can be a Flink (an open-source stream processing framework) data processing task, and the first message queue can be Kafka.

[0058] In the embodiment of the present invention, as Figure 2 shown, performing dimensional processing on the first data to obtain the second data corresponding to the first data includes:

[0059] Step S201: Use a dimension supplement operator to supplement the dimensions of the first data to obtain the third data; the third data includes each data identifier in the first data and the data corresponding to each data identifier;

[0060] Step S202: Merge the data in the third data according to the data identifier to obtain the second data.

[0061] In an embodiment of the present invention, the first data in the first message queue is consumed by a dimension processing task, that is, the first data is obtained from the first message queue, and then the first data is processed by a unified dimension supplement operator to obtain third data. That is, the third data is the first data after dimension supplement. The first data includes multiple pieces of data, each piece of data has a corresponding data identifier, and some of the multiple pieces of data may have the same data identifier. The third data includes each data identifier in the first data and the data corresponding to each data identifier. Each data identifier may correspond to one or more pieces of data, which is the same as the number of pieces of data corresponding to each data identifier in the first data, and each piece of data corresponding to each data identifier in the third data is obtained after the dimension data of the data corresponding to the data identifier in the first data is associated and supplemented. Among them, the dimension data is obtained according to the data identifier, and the dimension data is configured according to different services. For example, the dimension data may be organizational structure information, site information, user information, etc. For example, a waybill is sent from site A to site B, and the business data includes the site ids of sites A and B. When displaying to the user, the site information (i.e., dimension information) needs to be displayed. Therefore, when performing dimension processing, the dimension supplement operator is used to supplement dimension data such as the names, site types, and provinces of sites A and B.

[0062] After obtaining the third data, the data in the third data is merged according to the data identifier. For example, the data with the same data identifier can be merged into one piece of data, so as to obtain the second data. The second data may include each data identifier and one piece of data corresponding to each data identifier.

[0063] In an embodiment of the present invention, merging the data in the third data according to the data identifier to obtain the second data includes:

[0064] Using the data identifier to perform data partitioning on each piece of data to obtain multiple partitions, and each piece of data in each partition has the same data identifier;

[0065] Within a time window, the data in the same partition is merged according to a preset merging rule to obtain the merged data corresponding to each partition;

[0066] The second data is obtained according to the merged data corresponding to each partition in each partition.

[0067] In an embodiment of the present invention, when merging each piece of data in the third data according to the data identifier, the data is first partitioned according to the data identifier, that is, the data identifier is used as the partition key, and each piece of data with the same data identifier is grouped into the same partition, so that each piece of data in the same partition can be sent to a machine for processing. This process can be implemented based on a Flink task. By using the window function of Flink, data partitioning can be performed, and data merging can be carried out within a time window. Among them, Windows divides the stream into "buckets" of finite size, and calculations can be applied to the "buckets". Flink is a stream computing engine, and data is continuous. Batch processing is a special type of stream computing, where time windows are segmented from the stream, and each time window is equivalent to a space of finite size, aggregating the data to be processed. In an embodiment of the present invention, the data to be processed within a time window is each piece of data corresponding to each partition.

[0068] In an embodiment of the present invention, each piece of data in the same partition is merged according to a preset merging rule to obtain the merged data corresponding to each partition, including:

[0069] Obtain the event time of each piece of data in the same partition, and use the data with the latest event time as the merged data corresponding to the partition; or

[0070] Obtain the arrival time of each piece of data in the same partition when it reaches the time window, and use the data with the latest arrival time as the merged data corresponding to the partition.

[0071] In an embodiment of the present invention, after obtaining multiple partitions, within a time window, each piece of data in the same partition is merged according to a preset merging rule. It can be merged according to the event time (i.e., the data generation time) of each piece of data, that is, the data with the latest event time is used to overwrite the previous data to obtain the merged data corresponding to each partition; it can also be merged according to the arrival time of each piece of data when it reaches the time window, and the data with the latest arrival time is used to overwrite the previous data to obtain the merged data, so as to ensure that there is only one piece of data corresponding to the same data identifier, and there will be no conflict in data versions when writing to storage, and at the same time, the number of writes to storage can be reduced.

[0072] The second data is obtained according to the merged data corresponding to each partition in each partition. Each piece of data in the second data is the merged data, and each piece of data in the second data has a different data identifier.

[0073] In an embodiment of the present invention, the data processing task is the primary task or the backup task; before obtaining the first data from the first message queue by using the dimension processing task, it further includes: determining whether the primary task is available; in the case where the primary task is available, obtaining the first data from the first message queue corresponding to the primary task by using the dimension processing task; in the case where the primary task is unavailable, obtaining the first message queue corresponding to the backup task from the configuration file, and obtaining the first data from the first message queue corresponding to the backup task by using the dimension processing task.

[0074] In an embodiment of the present invention, the data processing task may be the primary task or the backup task, and the primary task and the backup task are two completely identical and isolated data processing tasks. That is, after obtaining the data message sent by the producer, the primary task and the backup task respectively perform data processing on the source data, and the first data obtained by each of them is respectively sent to their respective first message queues. Whether the first message queue from which the dimension processing task obtains the first data is the first message queue corresponding to the primary task or the first message queue corresponding to the backup task is determined according to the configuration file. The configuration file indicates that in the case where the primary task is available, the first data is obtained from the first message queue corresponding to the primary task, and in the case where the primary task is unavailable, the first data is obtained from the first message queue written by the backup task. That is, by pre-configuring the switching rule between the primary task and the backup task in the configuration file, it is possible to switch to the first message queue written by the backup task in a timely manner through the configuration file in the case where the primary task is unavailable, without the need for code modification and online process approval, and the downstream system is unaware, which can greatly improve the efficiency of fault resolution, better ensure data timeliness, and reduce the risk of data errors.

[0075] In an embodiment of the present invention, before obtaining the first message queue corresponding to the backup task from the configuration file, it further includes: storing the backup task corresponding to the primary task and the first message queue corresponding to the backup task in the configuration file. By pre-storing the backup task and the first message queue corresponding to the backup task in the configuration file, it is possible to achieve flexible switching between the primary task and the backup task through the configuration file, and the downstream storage or downstream consumer is unaware, thus improving the processing time of the overall link.

[0076] In an embodiment of the present invention, processing the second data through the second message queue includes: obtaining the second data from the second message queue and storing the second data in the database; or, enabling the downstream consumer to consume the second data in the second message queue.

[0077] In the embodiments of the present invention, after the second data is obtained, the second data is sent to the second message queue corresponding to the dimension processing task. The data synchronization task can be used to obtain the second data from the second message queue and store it in a database (such as ES, Elasticsearch, a search server based on Lucene), which improves the storage speed. Or it can be directly used by downstream consumers, that is, the downstream consumers consume the second data from the second message queue. This method uses the way of trading space for time to transfer the data storage messages to the second message queue to improve the data processing speed, because the storage speed of the database is much lower than the writing speed of the message queue. Through the data synchronization task for data storage, the writing performance of the database can be better and more fully utilized.

[0078] For real-time data processing projects, it can improve the data processing speed, reduce the version conflicts in data storage, decouple data storage, dimension supplementation, and data processing, realize the flexible assembly of tasks. Developers only need to focus on data processing, without additional concern about data dimension changes, out-of-order and performance issues during data storage, which is more friendly to downstream storage and downstream consumers, and will not cause switching costs due to upstream calculation exceptions.

[0079] Figure 3 It is a flowchart of a data processing method according to an embodiment of the present invention. The producer sends the data message generated based on the source data to the message queue; the main task or the standby task: uses the data processing task to obtain the data message from the message queue, so as to obtain the source data according to the data message and perform data processing on the source data. Obtaining the source data according to the data message includes obtaining it from the Redis (an open-source in-memory data structure storage system) cache and from the external storage through the interface service. The data processing process includes data cleaning, data transformation, logical processing, rule engine, data statistics, and data compression. And the first data obtained after data processing is sent to the corresponding message queue; the dimension processing task determines to obtain the first data from the message queue corresponding to the main task or the standby task according to the configuration file, and performs dimension processing on the first data. The dimension processing process includes dimension calculation, data transformation, and data merging, and the second data obtained after dimension processing is sent to the message queue, so that the data synchronization task obtains the second data from the message queue and stores it in the database to achieve data synchronization; the downstream consumers can also consume the second data from the message queue.

[0080] The data processing method according to the embodiment of the present invention first obtains a data message, processes the source data indicated by the data message to obtain first data; obtains the first data from the first message queue, and performs dimensional processing on the first data to obtain second data, and sends the second data to the second message queue to process the second data through the second message queue to realize the storage of the second data. This data processing method decouples data processing, dimensional processing, and data storage, improving the performance and stability of computing tasks; by sinking and processing dimensional data, this method reduces the complexity of dimensional data in data processing tasks and storage and transmission consumption. When dimensional data changes, in the dimensional processing layer, only the unified dimensional supplement operator version needs to be updated, without paying attention to dimensional changes, reducing redevelopment; by adding a sinking layer for dimensional processing, the primary and standby tasks can be flexibly switched, and downstream storage or consumers are unaware. When the primary task or standby task fails, the processing efficiency of the overall link is improved; this method uses the method of trading space for time to transfer the message of data storage to a new message queue, improving the data processing speed; through data synchronization tasks for data storage, the writing performance of the database can be better and more fully utilized, reducing development costs.

[0081] According to another aspect of the embodiment of the present invention, as Figure 4 shown, a data processing apparatus 400 is provided, including:

[0082] An acquisition module 401, which acquires a data message, and the data message indicates source data;

[0083] A first processing module 402, which processes the source data to obtain first data corresponding to the source data, and sends the first data to the first message queue;

[0084] A second processing module 403, which acquires the first data from the first message queue and performs dimensional processing on the first data to obtain second data corresponding to the first data;

[0085] A sending module 404, which sends the second data to the second message queue to process the second data through the second message queue.

[0086] In the embodiment of the present invention, the second processing module 403 is further configured to: use a dimensional supplement operator to supplement dimensions for the first data to obtain third data; the third data includes each data identifier in the first data and the data corresponding to each data identifier;

[0087] Merge the data in the third data according to the data identifier to obtain the second data.

[0088] In an embodiment of the present invention, the second processing module 403 is further configured to: partition each piece of data using a data identifier to obtain a plurality of partitions, where each piece of data in each partition has the same data identifier;

[0089] Within a time window, merge the data in the same partition according to a preset merging rule to obtain merged data corresponding to each partition;

[0090] Obtain second data based on the merged data corresponding to each partition in each partition.

[0091] In an embodiment of the present invention, the second processing module 403 is further configured to: obtain the event time of each piece of data in the same partition, and use the data with the latest event time as the merged data corresponding to the partition; or

[0092] Obtain the arrival time of each piece of data in the same partition when it reaches the time window, and use the data with the latest arrival time as the merged data corresponding to the partition.

[0093] In an embodiment of the present invention, the first data is obtained by performing data processing on source data through a data processing task. The first message queue corresponds to the data processing task, and the data processing task is the main task or the standby task; the second processing module 403 is further configured to: before obtaining the first data from the first message queue,

[0094] Determine whether the main task is available;

[0095] When the main task is available, obtain the first data from the first message queue corresponding to the main task using a dimension processing task;

[0096] When the main task is not available, obtain the first message queue corresponding to the standby task from the configuration file, and obtain the first data from the first message queue corresponding to the standby task using a dimension processing task.

[0097] In an embodiment of the present invention, the second processing module 403 is further configured to: before obtaining the first message queue corresponding to the standby task from the configuration file, store the standby task corresponding to the main task and the first message queue corresponding to the standby task in the configuration file.

[0098] In an embodiment of the present invention, the sending module 404 is further configured to: obtain the second data from the second message queue and store the second data in the database; or

[0099] Enable downstream consumers to consume the second data in the second message queue.

[0100] According to another aspect of an embodiment of the present invention, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the data processing method provided by the present invention.

[0101] According to still another aspect of an embodiment of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the data processing method provided by the present invention is implemented.

[0102] Figure 5 An exemplary system architecture 500 is shown to which the data processing method or data processing device according to an embodiment of the present invention can be applied.

[0103] As Figure 5 shown, the system architecture 500 may include terminal devices 501, 502, 503, a network 504, and a server 505. The network 504 is used to provide a medium for communication links between the terminal devices 501, 502, 503 and the server 505. The network 504 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0104] Users may use the terminal devices 501, 502, 503 to interact with the server 505 through the network 504 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 501, 502, 503, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0105] The terminal devices 501, 502, 503 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0106] The server 505 may be a server providing various services, such as a background management server (only as an example) that supports shopping websites browsed by users using the terminal devices 501, 502, 503. The background management server may analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - only as examples) to the terminal devices.

[0107] It should be noted that the data processing method provided by the embodiment of the present invention is generally executed by the server 505, and correspondingly, the data processing device is generally disposed in the server 505.

[0108] It should be understood, Figure 5The numbers of the terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0109] Reference is made below to Figure 6 , which shows a schematic structural diagram of a computer system 600 of a terminal device suitable for implementing the embodiments of the present invention. Figure 6 The terminal device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0110] As Figure 6 shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0111] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that the computer program read from it can be installed into the storage section 608 as required.

[0112] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 609 and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above functions defined in the system of the present invention are executed.

[0113] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0115] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes an acquisition module, a first processing module, a second processing module, and a sending module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the acquisition module can also be described as "the module for acquiring data messages".

[0116] As another aspect, the present invention further provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; it can also exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the device includes: acquiring a data message, where the data message indicates source data; performing data processing on the source data to obtain first data corresponding to the source data, and sending the first data to a first message queue; acquiring the first data from the first message queue, and performing dimensionality processing on the first data to obtain second data corresponding to the first data; sending the second data to a second message queue to process the second data through the second message queue.

[0117] According to the technical solution of the embodiments of the present invention, the data processing method of the embodiments of the present invention first acquires a data message, performs data processing on the source data indicated by the data message to obtain first data; acquires the first data from the first message queue, and performs dimensionality processing on the first data to obtain second data, and sends the second data to the second message queue to process the second data through the second message queue, so as to realize the storage of the second data. This data processing method decouples data processing, dimensionality processing, and data storage, improving the performance and stability of computing tasks; by sinking the dimensionality data for processing, this method reduces the complexity, storage, and transmission consumption of dimensionality data during data processing. When the dimensionality data changes, only the unified dimensionality supplement operator version needs to be updated at the dimensionality processing layer, without paying attention to the changes in dimensions, reducing redevelopment; by adding a sinking layer for dimensionality processing, the primary and standby tasks can be flexibly switched, and the downstream storage or consumers are unaware. When the primary task or the standby task fails, the processing efficiency of the overall link is improved; this method uses space to trade for time, transferring the data storage messages to a new message queue to improve the data processing speed; by performing data storage through data synchronization tasks, the writing performance of the database can be better and more fully utilized, reducing the development cost.

[0118] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data processing method, characterized in that, Including: Obtain a data message, where the data message indicates source data; Perform data processing on the source data to obtain first data corresponding to the source data, and send the first data to a first message queue; Obtain the first data from the first message queue, and perform dimensionality processing on the first data to obtain second data corresponding to the first data; Send the second data to a second message queue to process the second data through the second message queue.

2. The method according to claim 1, wherein Performing dimensionality processing on the first data to obtain second data corresponding to the first data includes: Use a dimensionality supplement operator to supplement the dimensions of the first data to obtain third data; the third data includes each data identifier in the first data and the data corresponding to each data identifier; Perform data merging on each piece of data in the third data according to the data identifier to obtain the second data.

3. The method according to claim 2, characterized in that, Performing data merging on each piece of data in the third data according to the data identifier to obtain the second data includes: Use the data identifier to perform data partitioning on each piece of data to obtain multiple partitions, and each piece of data in each partition has the same data identifier; Within a time window, perform data merging on each piece of data in the same partition according to a preset merging rule to obtain merged data corresponding to each partition; Obtain the second data according to the merged data corresponding to each partition in each partition.

4. The method according to claim 3, characterized in that, Performing data merging on each piece of data in the same partition according to a preset merging rule to obtain merged data corresponding to each partition includes: Obtain the event time of each piece of data in the same partition, and use the data with the latest event time as the merged data corresponding to the partition; or Obtain the arrival time of each piece of data in the same partition when it reaches the time window, and use the data with the latest arrival time as the merged data corresponding to the partition.

5. The method according to claim 1, wherein The first data is obtained by performing data processing on the source data through a data processing task, the first message queue corresponds to the data processing task, and the data processing task is the main task or the backup task; Before obtaining the first data from the first message queue, it further includes: Determine whether the main task is available; When the main task is available, use a dimensionality processing task to obtain the first data from the first message queue corresponding to the main task; When the main task is not available, obtain the first message queue corresponding to the backup task from the configuration file, and use the dimensionality processing task to obtain the first data from the first message queue corresponding to the backup task.

6. The method according to claim 5, wherein Before obtaining the first message queue corresponding to the backup task from the configuration file, it further includes: Store the backup task corresponding to the main task and the first message queue corresponding to the backup task in the configuration file.

7. The method according to claim 1, wherein Processing the second data through the second message queue includes: Obtain the second data from the second message queue and store the second data in a database; or Enable downstream consumers to consume the second data in the second message queue.

8. A data processing device, characterized in that, Including: An acquisition module that acquires a data message, where the data message indicates source data; The first processing module processes the source data to obtain first data corresponding to the source data, and sends the first data to a first message queue; The second processing module obtains the first data from the first message queue, and performs dimensionality processing on the first data to obtain second data corresponding to the first data; The sending module sends the second data to a second message queue to process the second data through the second message queue.

9. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, characterized in that, The program, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Cited By

  • A two-stage document recognition method and system based on an asynchronous decoupled architecture

    CN122526859A