Data processing method and device, equipment, medium and product
By identifying associated data field items in streaming data processing and performing key-value matching in a preset cache, combined with distributed caching and event-driven updates, the problem of high latency in real-time streaming data processing is solved, achieving high-concurrency processing and low-latency data splicing effects.
Patent Information
- Application Number
- CN202511023781.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies suffer from high latency and difficulty in supporting high-concurrency, fast queries in real-time streaming data processing. In particular, when dealing with large-scale real-time streaming data throughput, database storage and caching methods are insufficient to effectively improve concurrent access performance.
By identifying relevant data fields based on the business scenario of acquiring the stream data to be processed, and matching the relevant data values with the data values as keys in the preset relevant field data cache area, the data is concatenated. Combined with distributed caching and event-driven update mechanisms, the timeliness and accuracy of the data are ensured.
It improves the performance of concurrent processing of related data matching, effectively reduces data processing latency, is suitable for high-throughput streaming data processing, and ensures the real-time performance and accuracy of data processing.
Smart Images

Figure CN120873025A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of computer technology, and in particular to a data processing method, apparatus, device, medium and product. Background Technology
[0002] Current methods for associating real-time streaming data store the associated data in a database. When the real-time streaming data arrives, the database is queried and the data is joined. This approach is not suitable for large-scale real-time streaming data association. Real-time streaming data has a high throughput, and storing associated data in a database makes it difficult to support high-concurrency, fast queries, resulting in high latency. While loading associated data using caching can improve concurrent access performance and reduce latency, it is difficult to cache large amounts of associated data at low cost. Summary of the Invention
[0003] The present invention provides a data processing method, apparatus, device, medium and product that can improve the effect of concurrent processing of associated data, effectively reduce data processing latency, and is applicable to high-throughput streaming data processing.
[0004] In a first aspect, embodiments of the present invention provide a data processing method, the method comprising:
[0005] Obtain the stream data to be processed, and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed;
[0006] Using the data value of the first data field item in the stream data to be processed as the key, match at least one associated data value corresponding to a data field item in the preset associated field data cache.
[0007] The data in the stream to be processed is concatenated with the associated data values to obtain the data concatenation result.
[0008] In a second aspect, embodiments of the present invention provide a data processing apparatus, the apparatus comprising:
[0009] The data field item acquisition module is used to acquire the stream data to be processed and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed.
[0010] The data value matching module is used to match at least one associated data value corresponding to a data field item in the preset associated field data cache, using the data value of the first data field item in the stream data to be processed as the key value.
[0011] The data splicing result generation module is used to splice the data in the stream data to be processed with the associated data values to obtain the data splicing result.
[0012] Thirdly, embodiments of the present invention also provide a computer device, the computer device comprising:
[0013] One or more processors;
[0014] Memory, used to store one or more programs;
[0015] When the above one or more programs are executed by one or more processors, the one or more processors implement the data processing method provided in any embodiment of the present invention.
[0016] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data processing method provided in any embodiment of the present invention.
[0017] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the data processing method provided in any embodiment of the present invention.
[0018] The embodiments of the above invention have the following advantages or beneficial effects:
[0019] This invention involves acquiring stream data to be processed and determining at least one data field associated with the stream data based on the business data processing scenario corresponding to the stream data. Using the data value of the first data field in the stream data as a key, at least one associated data value corresponding to the data field is matched in a preset associated field data cache. The data in the stream data to be processed is then concatenated with the associated data value to obtain a concatenated data result. This invention solves the problem of high latency in current stream data processing, improves the concurrent processing effect of associated data matching, effectively reduces data processing latency, and is suitable for high-throughput stream data processing. Attached Figure Description
[0020] Figure 1 This is a flowchart of a data processing method provided in an embodiment of the present invention;
[0021] Figure 2 This is a flowchart of a data processing method provided in an embodiment of the present invention;
[0022] Figure 3 This is a flowchart of a data processing method provided in an embodiment of the present invention;
[0023] Figure 4 This is a flowchart of cache data maintenance provided by an embodiment of the present invention;
[0024] Figure 5This is a schematic diagram of the structure of a data processing device provided in an embodiment of the present invention;
[0025] Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of the present invention;
[0026] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0027] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0028] Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of the present invention. This embodiment is applicable to scenarios involving data processing and analysis. The method can be executed by a data processing device, which can be implemented in software and / or hardware and integrated into a computer device with application development capabilities.
[0029] like Figure 1 As shown, the data processing method in this embodiment includes the following steps:
[0030] S110. Obtain the stream data to be processed, and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed.
[0031] Streaming data is a continuously generated, real-time flowing data sequence with dynamically changing scale. In this embodiment, the streaming data to be processed can be streaming data generated in real time by the business system during the enterprise's business operations, such as streaming data generated by transaction business. The process of obtaining the streaming data to be processed can be through receiving it through a Kafka cluster, obtaining it by calling a preset API interface, or obtaining it by changing the data capture tool to synchronize it to the streaming processing engine in real time. This embodiment does not limit the method of obtaining the streaming data to be processed.
[0032] In scenarios such as financial business scenarios, the to-be-processed streaming data usually cannot be directly used as a data source for downstream real-time analysis applications. Instead, it is necessary to associate offline data to supplement information such as dimension fields. The downstream applications then process and analyze the streaming data after association and splicing. Therefore, after obtaining the to-be-processed streaming data, at least one data field item associated with the to-be-processed streaming data is determined according to the data splicing rules corresponding to the business data processing scenario corresponding to the to-be-processed streaming data. The data splicing rules include data association rules, which are used to associate fields, complete or convert the format of multi-source data according to the field requirements of downstream data analysis, and finally generate a data stream that meets the requirements for use by downstream applications. Business data processing scenarios such as financial business data processing scenarios can include processing scenarios in processes such as user interaction, transaction execution, risk control, and business analysis. For example, in the real-time transaction flow business scenario, if the to-be-processed streaming data is a user ID, then at least one associated data field item includes fields such as the customer address in the customer profile data.
[0033] S120. Use the data value of the first data field item in the to-be-processed streaming data as a key value to match the associated data values corresponding to at least one data field item in the preset associated field data buffer.
[0034] The first data field item can be a preset association identifier for matching the associated data value of the to-be-processed streaming data, such as user identity information or order number, etc. Use the data value of the first data field item in the to-be-processed streaming data as a key value to match, in the cached data of the preset associated field data buffer, the associated data values corresponding to at least one data field item with the same data value as the first data field item. The preset associated field data buffer can be a distributed cache, which can rely on a distributed KV storage system to achieve multi-node sharing of cached data.
[0035] Exemplarily, in the splicing scenario of real-time transaction flow and user information flow, the splicing condition is "user ID is the same". Then, extract <key: user ID, field 1: transaction amount, field 2: transaction time> from the to-be-processed transaction flow for association with fields such as "user name" and "risk level" with the same key in the user information flow cached in the preset associated field data buffer, and determine the data values corresponding to fields such as "user name" and "risk level".
[0036] S130. Splice the data in the to-be-processed streaming data with the associated data values to obtain a data splicing result.
[0037] Exemplarily, the splicing search fields extracted from the to-be-processed streaming data are <key: user ID, field 1: transaction amount, field 2: transaction time>; the fields matching this key in the user information flow cached in the preset associated field data buffer are user name and risk level.
[0038] The concatenated result data stream contains data values for the fields User ID (key), Transaction Amount, Transaction Time, User Name, and Risk Level.
[0039] The data concatenation result not only retains the real-time business data of the transaction flow, but also supplements user attribute information through key association, forming a complete data flow that meets the needs of downstream application analysis, such as statistical analysis of the transaction amount of high-risk users by region.
[0040] The technical solution of this embodiment involves acquiring the stream data to be processed and determining at least one data field item associated with the stream data based on the business data processing scenario corresponding to the stream data; using the data value of the first data field item in the stream data as a key, matching the associated data value corresponding to at least one data field item in a preset associated field data cache; and concatenating the data in the stream data with the associated data value to obtain the data concatenation result. This embodiment solves the problem of high latency in current stream data processing, improves the concurrent processing effect of associated data matching, effectively reduces data processing latency, and is suitable for high-throughput stream data processing.
[0041] Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of the present invention. This embodiment belongs to the same inventive concept as the data processing methods described in the previous embodiments, and further describes the process of loading data into a preset associated field data cache. This method can be executed by a data processing device, which can be implemented in software and / or hardware and integrated into a computer device with application development capabilities.
[0042] like Figure 2 As shown, the data processing method in this embodiment includes the following steps:
[0043] S210. Load data corresponding to at least one or all data field items from the preset associated data storage area into the preset associated field data cache area.
[0044] like Figure 3 As shown, Figure 3 The data processing flowchart shows that step P1 involves maintaining cached data. Related data is loaded into a distributed cache to support subsequent data association processing, enabling high-concurrency and low-latency queries when processing related data field items.
[0045] like Figure 4 As shown, Figure 4It is a flowchart for maintaining cached data and an extended step of step P1. In step P1.1, cached data is loaded. During initialization, part or all of the data is pre-loaded from a preset associated data storage area into the cache. The loaded data is the data corresponding to at least one data field item. The loading can be based on methods such as random sampling and loading late-updated records, and at least one subset of the full set of associated data in the preset associated data storage area is selected and loaded into the cache. The preset associated data storage area can be a KV database, a non-relational database with a key-value (Key-Value) as the core data model, such as DynamoDB, Berkeley DB, Cassandra, and HBase, etc.
[0046] S220. Detect data update events corresponding to at least one data field item.
[0047] In step P1.2, updated data messages are accessed. For detecting change events of associated data content, that is, data update events corresponding to at least one data field item, such as modification of customer address information, etc., data update events will be generated.
[0048] S230. Based on the data information in the data update event, update the data corresponding to the preset associated field data cache area and / or the preset associated data storage area.
[0049] In step P1.3, the cached data in the preset associated field data cache area is updated. For the updated data messages accessed in P1.2, combined with the scenario characteristics of real-time stream data splicing, it supports configuring two strategies for updating cached data, including but not limited to direct update and partial update:
[0050] Strategy 1. For the direct update strategy, directly write the updated data message into the preset associated field data cache area.
[0051] Strategy 2. For the partial update strategy, query whether the Key and the first data field item to which the update message belongs exist in the preset associated field data cache area and / or the preset associated data storage area. If it exists, update the data to <key, new Value>, otherwise write the key-value pair <key, Value> corresponding to the data into the cache.
[0052] In step P1.4, the stored data in the preset associated data storage area is updated. When updating the stored data, directly update the record to which the key belongs in the associated data storage unit through the updated data message accessed in P1.2.
[0053] This embodiment uses an event-driven approach to update associated data in a timely manner. At the same time, it uses an update caching mechanism to update the latest associated content in a timely manner, ensuring the accuracy of real-time streaming data association. By configuring direct update and partial update strategies, it not only ensures the timely maintenance of associated data, but also allows for the pre-loading of future hot access data into the cache, improving access efficiency and reducing cache penetration.
[0054] In one optional implementation, data in a preset associated field data cache area is cleaned up according to a preset data cache expiration time and a preset data cache size threshold.
[0055] In step P1.5, cache data cleanup is performed. Expired cached data is cleared to prevent invalid data from residing in the cache for extended periods, thus avoiding unlimited cache data expansion and ensuring cache access performance. This step uses a configuration-based approach to clear invalid data, and the configuration strategies include, but are not limited to, the following:
[0056] Strategy 1: Clean up cached data that has not been queried or accessed for a long time, such as cached data that has not been accessed for a preset data cache expiration time, in order to release storage resources.
[0057] Strategy 2: When the total data capacity of all data cached in the preset associated field data cache area exceeds the preset data cache size threshold, clean up the data added to the cache earlier.
[0058] Strategy 3: Configure a cleanup mechanism to be triggered during off-peak business periods to clean up some invalid data. This embodiment can promptly clean up invalid access data and ensure cache performance by configuring a cleanup strategy.
[0059] S240. Obtain the stream data to be processed, and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed.
[0060] like Figure 3 As shown, in step P2, real-time stream data is accessed, which involves acquiring the stream data to be processed and determining the associated field items. The real-time stream data to be concatenated is accessed from the upstream data source, and at least one data field item associated with the stream data to be processed is determined according to the concatenation rules corresponding to the business data processing scenario.
[0061] S250: Using the data value of the first data field item in the stream data to be processed as the key, match at least one associated data value corresponding to a data field item in the preset associated field data cache.
[0062] Perform real-time data association in step P3. According to the splicing condition, extract the splicing search field of the real-time stream data to be spliced, that is, the first data field item. For example, if the stream data to be processed includes fields <key, field 1, field 2,...>, where key is the first data field item.
[0063] Perform cache data search in step P4. According to the splicing search field extracted in step P3, search for the associated data value Value corresponding to at least one data field item of the Key from the cache data in the preset associated field data buffer, that is, the associated supplementary data element.
[0064] In an optional implementation manner, when at least one associated data value corresponding to the data field item is not matched in the preset associated field data buffer, perform matching in the preset associated data storage area to obtain at least one associated data value corresponding to the data field item; update the at least one associated data value corresponding to the data field item found to the preset associated field data buffer.
[0065] Perform hit judgment search in step P5. If the search in step P4 from the cache data is a hit, directly execute step P8, that is, step S260, otherwise execute step P6.
[0066] In step P6, the cache search is not a hit. Search for the associated data value Value corresponding to at least one data field item of the Key from the preset associated data storage area.
[0067] In step P7, update the associated data value <key, Value> found in P6 to the preset associated field data buffer for subsequent query access.
[0068] S260. Splice the data in the stream data to be processed with the associated data value to obtain a data splicing result.
[0069] In step P8, splice the real-time stream data field value, that is, the stream data to be processed, with the associated data value, that is, the found Value value, and output the real-time stream data splicing result, such as <key, field 1, field 2,..., Value>.
[0070] In an optional implementation manner, send the data splicing result to the target message queue associated with the stream data to be processed. For example, send the data splicing result to the message queue of the downstream application.
[0071] The technical solution of this embodiment involves: pre-loading data corresponding to at least one data field item from a preset associated data storage area into a preset associated field data cache area; detecting data update events corresponding to at least one data field item; updating the data corresponding to the preset associated field data cache area and / or the preset associated data storage area based on the data information in the data update events; acquiring the stream data to be processed and determining at least one data field item associated with the stream data to be processed according to the business data processing scenario corresponding to the stream data to be processed; using the data value of the first data field item in the stream data to be processed as a key, matching the associated data value corresponding to at least one data field item in the preset associated field data cache area; and concatenating the data in the stream data to be processed with the associated data value to obtain the data concatenation result. This technical solution solves the problem of high latency in current stream data processing, improves the concurrent processing effect of associated data matching, effectively reduces data processing latency, and is suitable for high-throughput stream data processing. Through an event-driven approach, the cache data can be updated in a timely manner, and the latest associated content is updated promptly through the cache update mechanism, ensuring the accuracy of stream data association during processing.
[0072] In a specific data processing instance, it can be done through, for example... Figure 5The data processing device shown performs data processing and includes a real-time streaming data access unit, a real-time data association unit, a cached data unit, an associated data storage unit, an update data access unit, a data update unit, a cache cleanup unit, and a real-time association output unit. Specifically, the real-time streaming data access unit is used to access high-throughput real-time streaming data and, through a distribution operation, distributes the real-time streaming data to be concatenated to downstream real-time data association units; the real-time data association unit is used to query supplementary data elements from the cached data unit based on the association concatenation conditions of the accessed real-time streaming data and supplement them to form new real-time streaming data; the cached data unit is used to cache associated data. Initially, it loads some data from the associated data storage unit for quick querying and access by the real-time data association unit, and simultaneously receives signals from the cache cleanup unit to promptly clear expired cached data and receives update event streams from the data update unit to update the cached data in a timely manner; the associated data storage unit stores large-scale associated data, for example, using a KV database (such as HBase) as data storage. This module is used to provide suitable data to the cached data unit to reduce cache penetration during real-time data association. It also receives update event streams from the data update unit, driving timely updates of related data; update data access unit. For changes in related data content, such as modifications to customer address information, an update event stream is generated. This unit receives update data event messages and distributes them to downstream data update units; data update unit. It receives related data update messages distributed by the update data access unit, driving updates to records in the related data storage unit, and simultaneously driving timely updates to records in the cache data unit. For updates to cached data, this unit can update cache records through configuration methods; cache cleanup unit. Based on set cache expiration time, cache size threshold, business peak strategy, and other information, it configures and cleans up cached data unit data to prevent unlimited expansion of cached data unit data; real-time correlation output unit. It outputs real-time streaming data correlation results to the target end's message queue, etc.
[0073] Figure 6 This is a schematic diagram of a data processing device provided in an embodiment of the present invention. This embodiment is applicable to data processing scenarios. The data processing device can be implemented in software and / or hardware and integrated into a computer terminal device with application development capabilities.
[0074] like Figure 6 As shown, the data processing device includes: a data field item acquisition module 310, a data value matching module 320, and a data splicing result generation module 330.
[0075] The data field item acquisition module 310 is used to acquire the stream data to be processed and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed; the data value matching module 320 is used to use the data value of the first data field item in the stream data to be processed as the key value to match the associated data value corresponding to at least one data field item in the preset associated field data cache; and the data splicing result generation module 330 is used to splice the data in the stream data to be processed with the associated data value to obtain the data splicing result.
[0076] The technical solution of this embodiment involves acquiring the stream data to be processed and determining at least one data field item associated with the stream data based on the business data processing scenario corresponding to the stream data; using the data value of the first data field item in the stream data as a key, matching the associated data value corresponding to at least one data field item in a preset associated field data cache; and concatenating the data in the stream data with the associated data value to obtain the data concatenation result. This embodiment solves the problem of high latency in current stream data processing, improves the concurrent processing effect of associated data matching, effectively reduces data processing latency, and is suitable for high-throughput stream data processing.
[0077] In one alternative embodiment, the apparatus further includes:
[0078] The data loading module is used to preload some or all of the data corresponding to at least one data field item from the preset associated data storage area into the preset associated field data cache area.
[0079] The data update event handling module is used to detect data update events corresponding to at least one data field item; based on the data information in the data update event, it updates the data corresponding to the preset associated field data cache and / or the preset associated data storage area.
[0080] The storage area data matching module is used to perform matching in the preset associated data storage area when no associated data value corresponding to at least one data field item is matched in the preset associated field data cache area, so as to obtain the associated data value corresponding to at least one data field item; and update the associated data value corresponding to the matched at least one data field item to the preset associated field data cache area.
[0081] The splicing result sending module is used to send the data splicing result to the target message queue associated with the stream data to be processed.
[0082] The data cleanup module is used to clean up data in the preset associated field data cache area according to the preset data cache expiration time and preset data cache size threshold.
[0083] The data processing apparatus provided in the embodiments of the present invention can execute the data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0084] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 7 A block diagram of an exemplary computer device 12 suitable for implementing embodiments of the present invention is shown. Figure 7 The computer device 12 shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as intelligent controllers and servers, mobile phones, and other terminal devices.
[0085] like Figure 7 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0086] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0087] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0088] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 7 Not shown; usually referred to as a "hard drive"). Although Figure 7Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0089] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0090] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with the computer device 12, and / or with any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although... Figure 7 As not shown, it may be used in conjunction with computer device 12 with other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, artificial intelligence systems, tape drives, and data backup storage systems.
[0091] Processing unit 16 executes various functional applications and data processing by running programs stored in system memory 28, such as implementing the data processing method provided in this embodiment, which includes:
[0092] Obtain the stream data to be processed, and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed;
[0093] Using the data value of the first data field item in the stream data to be processed as the key, match at least one associated data value corresponding to a data field item in the preset associated field data area;
[0094] The data in the stream to be processed is concatenated with the associated data values to obtain the data concatenation result.
[0095] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the data processing method provided in any embodiment of this invention, the method comprising:
[0096] Obtain the stream data to be processed, and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed;
[0097] Using the data value of the first data field item in the stream data to be processed as the key, match at least one associated data value corresponding to a data field item in the preset associated field data area;
[0098] The data in the stream to be processed is concatenated with the associated data values to obtain the data concatenation result.
[0099] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0100] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0101] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0102] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, Python, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0103] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the data processing method provided in any embodiment of this application.
[0104] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, Python, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0105] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0106] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A data processing method, characterized in that, include: Obtain the stream data to be processed, and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed; Using the data value of the first data field item in the stream data to be processed as the key, the associated data value corresponding to the at least one data field item is matched in the preset associated field data cache area; The data in the stream data to be processed is concatenated with the associated data value to obtain the data concatenation result.
2. The method according to claim 1, characterized in that, Before acquiring the stream data to be processed, the method further includes: Data corresponding to at least one data field item is preloaded from a preset associated data storage area into the preset associated field data cache area, either partially or entirely.
3. The method according to claim 2, characterized in that, The method further includes: Detect the data update event corresponding to at least one of the data field items; Based on the data information in the data update event, update the data corresponding to the preset associated field data cache and / or the preset associated data storage area.
4. The method according to claim 1, characterized in that, The method further includes: When no associated data value corresponding to the at least one data field item is matched in the preset associated field data cache, a matching is performed in the preset associated data storage area to obtain the associated data value corresponding to the at least one data field item. Update the associated data value corresponding to the at least one matched data field item to the preset associated field data cache area.
5. The method according to claim 1, characterized in that, The method further includes: The data concatenation result is sent to the target message queue associated with the stream data to be processed.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Clean up the data in the preset associated field data cache area according to the preset data cache expiration time and preset data cache size threshold.
7. A data processing apparatus, characterized in that, include: The data field item acquisition module is used to acquire the stream data to be processed and determine at least one data field item associated with the stream data to be processed based on the business data processing scenario corresponding to the stream data to be processed. The data value matching module is used to match the associated data value corresponding to the at least one data field item in the first data field item in the stream data to be processed as the key value in a preset associated field data cache area. The data splicing result generation module is used to splice the data in the stream data to be processed with the associated data value to obtain the data splicing result.
8. A computer device, characterized in that, The computer device includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the data processing method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the data processing method as described in any one of claims 1-6.