Dimension table association processing method and device, electronic equipment and storage medium

By performing keyby partitioning of the output of the Calc operator in the big data real-time computing engine, the problem of low cache hit rate in dimension table correlation scenarios is solved, and higher computing performance and lower storage query pressure are achieved.

CN119938714APending Publication Date: 2025-05-06中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510054947.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In real-time big data calculation, the cache hit rate is low in the dimension table correlation scenario, resulting in a degradation in real-time task computing performance and an increase in the pressure of external dimension table storage query.

Method used

By performing keyby partitioning operations on the output of the Calc operator in the data flow real-time calculation engine, the associated primary key associated with the dimension table is used as the partition condition, and the data is then partitioned and input to the LookupJoin dimension table association operator.

Benefits of technology

Improves cache hit rate, reduces the storage of useless keys, improves task processing performance, and reduces query pressure on external storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938714A_ABST
    Figure CN119938714A_ABST
Patent Text Reader

Abstract

The invention discloses a dimension table association processing method and device, electronic equipment and a storage medium, and the method comprises the steps: in a data stream real-time calculation engine, redistributing first consumption queue data to a Calc operator according to a Source data source operator; and according to the partitioning operation on the Calc operator, obtaining grouped data after the partitioning operation, inputting the grouped data into a LookupJoin dimension table association operator, and outputting the grouped data to second consumption queue data through a sink operator. By means of the method and device, on one hand, dimension table association is optimized, useless key storage is reduced, and on the other hand, the cache hit rate is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data real-time computing technology, and in particular to a dimension table association processing method, device, electronic device, and storage medium. Background Art

[0002] In the financial field, real-time risk control, real-time marketing and other scenarios have increasingly higher requirements for the timeliness of data. Currently, real-time computing engines based on stream computing are usually used to meet real-time requirements.

[0003] During real-time computing and processing, each piece of streaming data can be associated with an external dimension table data source to supplement the dimension of the streaming data, which is called dimension table association.

[0004] In the scenario of big data real-time computing dimension table association, the industry generally adopts the dimension data caching mechanism. When processing the actual scenario where the driving flow of real-time computing is relatively large and the association key distribution is relatively scattered, the hit rate of the dimension data in the cache will be very low. Summary of the invention

[0005] The embodiments of the present application provide a dimension table association processing method, device, electronic device, and storage medium to optimize dimension table association scenarios and improve cache hit rate in a cache.

[0006] The present application embodiment adopts the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a dimension table association processing method, wherein the processing method includes:

[0008] In the data stream real-time computing engine, the data of the first consumer queue is redistributed to the Calc operator according to the Source data source operator;

[0009] According to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator.

[0010] In some embodiments, according to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator, including:

[0011] According to the keyby operation after the Calc operator, the associated primary key associated with the dimension table is used as a condition of the keyby operation to partition the data, thereby obtaining grouped data after the keyby operation;

[0012] The grouped data after the partition operation is hashed and grouped according to the associated primary key associated with the dimension table, and the data of the same group falls into the same instance and then input into the LookupJoin dimension table association operator, and is output to the consumer queue topic through the sink operator.

[0013] In some embodiments, the method further comprises:

[0014] Put the data with the same dimension table association key into the same instance of the LookupJoin dimension table association operator for processing, so that the dimension table association key is a subset of the key distributed by the Calc operator;

[0015] The instance of the LookupJoin dimension table association operator caches a portion of the dimension data. When the dimension table is associated, the cache space accessed changes from the first key to the second key. The first key is all the keys, and the second key includes a portion of the all the keys.

[0016] In some embodiments, in the data stream real-time computing engine, reallocating the first consumer queue data to the Calc operator according to the Source data source operator includes:

[0017] According to the Source data source operator subscribing to and consuming the first data message of the source Kafka Topic, the first data message of the Kafka Topic is redistributed to the Calc operator;

[0018] According to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator, including:

[0019] When the instance of the LookupJoin dimension table association operator consumes a message, if the corresponding dimension data does not exist in the cache, it will go to the external storage for query, and the dimension data after the query will be put into the cache;

[0020] If the corresponding dimension data is found in the cache, the dimension data in the cache is used directly;

[0021] According to the dimension data in the cache, a second data message is output to the Kafka Topic through the sink operator.

[0022] In some embodiments, the partitioning operation after the Calc operator is performed, and the grouped data after the partitioning operation is obtained and then input into the LookupJoin dimension table association operator, further comprising:

[0023] In the LookupJoin dimension table association operator, the cache hit rate of the dimension table association is improved. The cache hit rate refers to the ratio of the number of times data is read from the cache to the total number of times data is read. The total number of times data is read is the sum of the number of times data is read from the cache and the number of times data is read from the dimension table storage.

[0024] In some embodiments, the real-time computing engine includes any one or more operators of the Source data source operator, the Calc operator, the LookupJoin dimension table association operator, and the sink operator.

[0025] In some embodiments, the real-time computing engine includes at least one of the following: Flink, Spark Streaming.

[0026] In a second aspect, an embodiment of the present application further provides a dimension table association processing device, wherein the processing device includes:

[0027] A processing module, used for reallocating the first consumer queue data to the Calc operator according to the Source data source operator in the data stream real-time calculation engine;

[0028] The operation module is used to obtain the grouped data after the partition operation according to the partition operation of the Calc operator, input the grouped data into the LookupJoin dimension table association operator, and output the grouped data to the second consumer queue data through the sink operator.

[0029] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to perform the above method.

[0030] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes the above method.

[0031] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: in the data stream real-time computing engine, the first consumer queue data is reallocated to the Calc operator according to the Source data source operator. Then, according to the partitioning operation of the Calc operator, the grouped data after the partitioning operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator. Add a partitioning operation after the Calc operator. The partitioning operation can partition the data after Calc processing according to different keys. Only part of the keys need to be processed in the instance of the LookupJoin operator, and each instance of the LookupJoin operator accesses an independent key space. Through the above-mentioned dimension table association, the storage of useless keys is reduced, and the cache hit rate can be further improved, thereby improving the processing performance of the task, and reducing the query pressure of the external memory. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0033] Figure 1 This is a flowchart of a method for processing a dimensional table association in an embodiment of the present application;

[0034] Figure 2 It is a topological diagram of the task of the scenario associated with the dimension table in the related technology;

[0035] Figure 3 It is a partition operation task topology diagram in the dimension table association processing method in the embodiment of the present application;

[0036] Figure 4 This is a schematic diagram of the structure of a dimensional table association processing device in an embodiment of the present application;

[0037] Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0039] For dimension table association, if each stream data directly queries the external dimension table, the query pressure on the external dimension table storage will be very high, so the production scenario generally adds a dimension table cache mechanism. There are two main parameters for setting the dimension table cache strategy for real-time computing tasks: one parameter is to set the cache expiration time, and the other parameter is the maximum number of data to be saved in the cache. However, this method of setting the dimension table cache cannot guarantee the cache hit rate.

[0040] For big data real-time computing dimension table association scenarios, the dimension data cache mechanism is usually used. For actual scenarios where the driving flow of real-time computing is relatively large and the distribution of association keys is relatively scattered, the dimension data hit rate in the cache will be very low. The dimension data in the cache will be constantly swapped in and out, which will lead to a decrease in real-time task computing performance and even back pressure. In addition, it will also cause the query pressure of external dimension table storage to increase. Therefore, how to reduce the frequency of swapping in and out of dimension data in the cache and improve the cache hit rate is a problem that needs to be optimized.

[0041] In view of the above shortcomings, a dimension table association processing method is provided in an embodiment of the present application, which can improve the cache hit rate by calculating the dimension table association scenario in real time. At present, when calculating the dimension table association in real time, all instances of the dimension table association operator access all cache spaces, which will cause a relatively large cache pressure, and the data in the cache is constantly swapped in and out, and the cache hit rate is usually relatively low. If the data can be partitioned in advance, it can be ensured that the cache primary key space accessed by each instance of the dimension table association operator is independent, so that the frequency of the dimension data cached in the independent space being swapped in and out can be reduced, thereby improving the cache hit rate.

[0042] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0043] The present application embodiment provides a dimension table association processing method, such as Figure 1 As shown, a schematic flow chart of a method for processing dimensional table association in an embodiment of the present application is provided, and the method at least includes the following steps S110 to S120:

[0044] Step S110: In the data stream real-time computing engine, the first consumer queue data is reallocated to the Calc operator according to the Source data source operator.

[0045] The business value of data will decrease rapidly over time, so data should be processed and calculated as soon as possible after it occurs. Data processing delays are very obvious in business scenarios with high timeliness requirements, such as real-time risk control and real-time marketing. Real-time computing can effectively shorten data processing delays and meet business needs. In addition, a unit used to process data in the real-time computing engine, such as reading data, calculating data, correlating data, and outputting data, are all data processing units.

[0046] Dimension table refers to the fact that real-time computing is an event-triggered computing mode, and the triggering data source is an unbounded data stream, which can also be called a driving stream. In actual production scenarios, it is usually necessary to supplement the information of the data stream. The supplementary information is usually attribute information, such as customer personal information, customer account card information, etc. This information is generally placed in a key-value database that stores data in a key-value data model as a dimension table. Dimension table association refers to the process of associating unbounded data streams with dimension table data by associating primary keys during real-time computing.

[0047] The Source data source operator consumes the message queue and redistributes the consumption queue data to the Calc operator.

[0048] It can be understood that the Source data source operator subscribes to and consumes data messages from the source Kafka Topic. The Calc operator is a field calculation and processing operator. Usually, the where condition in the SQL case will be converted into the Calc operator.

[0049] Step S120, according to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator.

[0050] According to the partition operation of the Calc operator, when performing the partition operation, the key is the associated primary key of the dimension table association, and the LookupJoin operation is performed after the partition operation grouping.

[0051] It can be understood that the LookupJoin dimension table association operator, specifically the external dimension table storage associated with the Kafka Topic data stream will be converted into a LookupJoin operator. The sink operator is used as an output to the external storage operator.

[0052] Specifically, for the partitioning operation of the Calc operator, i.e., the keyby operation, in the dimension table association scenario, the dimension table association has a natural keyby feature. The natural keyby feature means that when the tables are associated, the data in the two tables can only be associated if they have the same associated primary key, so the associated primary key of the dimension table association can be directly used as the keyby condition. In this way, after the Calc operator is processed, the keyby operation is performed, and the key of the keyby is the associated primary key of the dimension table association, and then LookupJoin and subsequent data processing are performed. Since the keyby operation uses the dimension table association primary key as the key, and because the dimension table association has a hash attribute, when performing the keyby operation, the associated key of the dimension table association is used for hashing, which can partition the data and ensure semantic correctness.

[0053] Through the above method, a keyby partitioning operation is added after the Calc operator, and the keyby partitioning operation can partition the data after Calc processing according to different keys. In the instance of the LookupJoin operator, only part of the keys need to be processed, and each instance of the LookupJoin operator accesses an independent key space, which reduces the storage of useless keys, improves the cache hit rate, and thus improves the processing performance of the task, and can also reduce the query pressure of the external memory.

[0054] Through the above method, the effect achieved by the keyby partitioning operation is that the Calc operator will perform hash grouping according to the associated primary key of the dimension table when distributing data. The data in the same group will fall into the same instance, that is, the data with the same dimension table associated key will be placed in an instance of the LookupJoin operator for processing.

[0055] Since data with the same dimension table association key will be placed in an instance of the LookupJoin operator, the dimension table association key processed by the instance of the LookupJoin operator is the partial key distributed by the Calc operator. Therefore, the instance of the LookupJoin operator only needs to cache the part of the dimension data it needs. The cache space accessed when the dimension table is associated is also optimized from all keys to partial keys. In this way, the cache hit rate is improved, which can improve the processing performance of real-time tasks and the timeliness of output result data, and at the same time reduce the access pressure of dimension table storage.

[0056] Different from the related technologies, in the actual scenario where the driving flow of real-time calculation is relatively large and the distribution of associated keys is relatively scattered, the hit rate of dimension data in the buffer will be very low, and the dimension data in the cache will be constantly swapped in and out, which will lead to the degradation of real-time task computing performance and even back pressure problems. Through the above method, the associated primary key associated with the dimension table is used as the condition of the keyby partition operation to partition the data, which can ensure that the cache primary key space accessed by each instance of the dimension table association operator is independent, so that the frequency of swapping in and out of the cached dimension data in the independent space can be reduced, thereby improving the cache hit rate.

[0057] The present application provides a method for improving cache hit rate in a real-time calculation dimension table association scenario, which is used to solve the problem of low cache hit rate of dimension table association.

[0058] like Figure 2 As shown in the figure, taking the typical real-time computing engine Flink as an example, a Flink SQL case is used to demonstrate the adopted technical solution. A Flink SQL case is given. The function implemented by this SQL is to first filter out the data with the amount field greater than 10,000 in the data stream, and then associate the dimension table to supplement the dimension. The resources given to the Flink SQL task corresponding to the case are as follows, with a total concurrency of 200, and the concurrency given to the reading data operator that consumes the Kafka Topic is 100. The SQL script content is as follows:

[0059] INSERT INTO kafka_topic

[0060] SELECT *

[0061] FROM source table

[0062] LEFT JION dim table

[0063] FOR SYSTEM_TIME AS OF source_table.proctime

[0064] ON source table.card=dim table.card

[0065] WHERE source table.amt>10000

[0066] When a Flink SQL task is running, the Flink computing engine will parse, verify, optimize, and execute the SQL statement content. During the execution phase, you can see the SQL topology diagram. Figure 2The figure below is a topology diagram based on the Flink SQL case. This topology diagram represents the actual execution process of SQL.

[0067] Please continue to refer to Figure 2 The operator concurrency involved is only for example and is not intended to limit the protection scope in the embodiments of the present application.

[0068] Parallelism: operator concurrency. A Flink program consists of multiple operators. An operator includes multiple instances that are executed in parallel, and each instance processes a subset of the operator's input data. The number of parallel instances of an operator is called the concurrency of the operator.

[0069] Source: data source operator, which subscribes to and consumes data messages from the source Kafka Topic. In the Flink computing engine, it corresponds to the Source operator. The concurrency of this operator is 100. Figure 2 shown.

[0070] Calc: Field calculation and processing operator. The where condition in the SQL case will be converted into the Calc operator. The concurrency of this operator is 200. Figure 2 shown.

[0071] LookupJoin: Dimension table association operator, which uses polling method. Kafka Topic data stream associated with external dimension table storage will be converted into LookupJoin operator. The concurrency of this operator is 200. Figure 2 shown.

[0072] Sink: Output to the external storage operator. The insert into statement in the SQL case will be converted into the Sink operator. The concurrency of this operator is 200. For example, Figure 2 shown.

[0073] Please continue to refer to Figure 2 First, the Source operator reads the data from the Kafka Topic. Because the Source operator and the downstream Calc operator have different concurrency, after the Source operator instance reads the data, it will redistribute the data to the Calc instance. After the Calc operator completes the processing, it will be handed over to the LookUpJoin operator, and finally output to the KafkaTopic through the Sink operator.

[0074] There is no keyby semantics in front of the LookUpJoin operator, so the data is not partitioned. It can also be found from the topology diagram that the Calc operator and the LookUpJoin operator are connected together. After the instance processing of the Calc operator is completed, it directly enters the corresponding instance processing of the LookUpJoin operator, and no data redistribution is performed between threads. In this way, if the source system data flow consumed by the task is relatively large during the actual production environment, the data partitioning will not be performed after the Calc operator is completed. In other words, each instance of the Calc operator may consume the data stream message corresponding to any associated primary key in the data source. The dimension table key space accessed by each instance of the LookUpJoin operator is actually the entire cache space, and all dimension data needs to be cached. However, the cache size space is limited and cannot store all dimension data of the dimension table. This will cause the cached dimension data to be constantly swapped in and out, which puts great pressure on the dimension table cache of the Flink task, and the cache hit rate is relatively low, and the job processing performance will be slow or even back pressure will occur. At the same time, the number of query requests received by the external cache storage will also increase, which may cause the query performance to slow down, and further cause the task to read the dimension table data slowly. In severe cases, the task may be constantly under back pressure, and the consumption speed will be extremely slow, which will eventually affect the timeliness of the task and have an impact on the business.

[0075] The keyby operation is mainly used for data partitioning in the implementation of the computing engine. It divides the data stream into different partitions, and the data in each partition will be sent to the same partition for processing. This partitioning method is implemented by calculating the hash value of the key and performing a modulo operation to ensure that data with the same key is assigned to the same partition. The keyby operation improves processing efficiency, especially when processing massive data, and can process data in each partition in parallel.

[0076] First, add a keyby operation after the Calc operator. When doing keyby, the key is the associated primary key of the dimension table. After performing keyby grouping, perform the LookupJoin operation. For the SQL case given above, according to the embodiment of this application, after adding the keyby operation, Figure 3 As shown in the topology diagram, we can see that the keyby operation is performed after the Calc operator. The keyby operation does not generate an additional actual operator. It is mainly used to specify that the data stream should be partitioned according to one or more fields.

[0077] Furthermore, combined with the Flink SQL case, the effect achieved by keyby is that when distributing data, the Calc operator will perform hash grouping according to the associated primary key card associated with the dimension table. The data in the same group will fall into the same instance, that is, the data with the same dimension table associated key will be placed in an instance of the LookupJoin operator for processing. In this way, the dimension table associated key processed by the instance of the LookupJoin operator is the partial key distributed by the Calc operator. Therefore, the instance of the LookupJoin operator only needs to cache the part of the dimension data it needs. The cache space accessed when the dimension table is associated is also optimized from all keys to partial keys. In this way, the cache hit rate is improved, which can improve the processing performance of real-time tasks and the timeliness of output result data, while also reducing the access pressure of dimension table storage.

[0078] In one embodiment of the present application, according to the partitioning operation of the Calc operator, the grouped data after the partitioning operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator, including: according to the keyby operation after the Calc operator, the associated primary key associated with the dimension table is used as the condition of the keyby operation to partition the data to obtain the grouped data after the keyby operation; the grouped data after the partitioning operation is hashed and grouped according to the associated primary key associated with the dimension table, and the data of the same group falls into the same instance and then input into the LookupJoin dimension table association operator, and output to the consumer queue topic through the sink operator.

[0079] In the embodiment of the present application, the effect achieved by keyby is that the Calc operator will perform hash grouping according to the associated primary key card associated with the dimension table when distributing data, and the data in the same group will fall into the same instance, that is, the data with the same dimension table associated key will be placed in an instance of the LookupJoin operator for processing. In this way, the dimension table associated key processed by the instance of the LookupJoin operator is the partial key distributed by the Calc operator, so the instance of the LookupJoin operator only needs to cache the part of the dimensional data it needs, and the cache space accessed when the dimension table is associated is also optimized from all keys to partial keys.

[0080] The Source operator reads the data of the Kafka Topic and finally outputs it to another Kafka Topic through the Sink operator.

[0081] In one embodiment of the present application, the method further includes: placing data with the same dimension table association key into the same instance of the LookupJoin dimension table association operator for processing, so that the dimension table association key is a subset of the key distributed by the Calc operator; the instance of the LookupJoin dimension table association operator caches a portion of the dimension data, and the cache space accessed when the dimension table is associated changes from the first key to the second key, the first key is all keys, and the second key includes a portion of all keys.

[0082] Based on the keyby partitioning operation, the Calc operator will perform hash grouping according to the associated primary key of the dimension table when distributing data. Since the dimension table associated key processed by the instance of the LookupJoin operator is part of the key distributed by the Calc operator, the instance of the LookupJoin operator only needs to cache the part of the dimension data it needs, and the cache space accessed when the dimension table is associated is also optimized from all keys to partial keys.

[0083] In one embodiment of the present application, in the data stream real-time computing engine, the first consumer queue data is reallocated to the Calc operator according to the Source data source operator, including: according to the Source data source operator, subscribing to and consuming the first data message of the source Kafka Topic, the first data message of the Kafka Topic is reallocated to the Calc operator; according to the partitioning operation of the Calc operator, the grouped data after the partitioning operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator, including: when the instance of the LookupJoin dimension table association operator consumes the message, if the corresponding dimension data does not exist in the cache, it will go to the external storage for query, and the dimension data after the query will be put into the cache; if the corresponding dimension data is found in the cache, the dimension data in the cache is directly used; according to the dimension data in the cache, the second data message output to the Kafka Topic through the sink operator.

[0084] In the process of caching dimension data of the LookUpJoin operator, when an instance of the LookUpJoin operator consumes a message, if the corresponding dimension data does not exist in the cache, it will go to the external storage for query, and put the dimension data after the query into the cache; if the corresponding dimension data can be found in the cache, the dimension data in the cache is directly used without querying the external storage.

[0085] In one embodiment of the present application, after the partitioning operation after the Calc operator is performed, the grouped data after the partitioning operation is obtained and then input into the LookupJoin dimension table association operator, and further includes: improving the cache hit rate of the dimension table association in the LookupJoin dimension table association operator, wherein the cache hit rate refers to the ratio of the number of times data is read from the cache to the total number of times data is read, and the total number of times data is read is the sum of the number of times data is read from the cache and the number of times data is read from the dimension table storage.

[0086] The cache hit rate refers to the ratio of the number of times data is read from the cache to the total number of times data is read. Here, the total number of times read = the number of times read from the cache + the number of times read from the dimension table storage.

[0087] Therefore, a keyby operation is added after the Calc operator. Keyby can partition the data after Calc processing according to different keys. In this way, the instance of the LookupJoin operator only needs to process some keys. Each instance of the LookupJoin operator accesses an independent key space, which reduces the storage of useless keys, improves the cache hit rate, and thus improves the processing performance of the task. It can also reduce the query pressure on the external memory.

[0088] In one embodiment of the present application, the real-time computing engine includes any one or more operators of the Source data source operator, the Calc operator, the LookupJoin dimension table association operator, and the sink operator.

[0089] Source: Data source operator, subscribes to and consumes data messages from the source Kafka Topic. Calc: Field calculation and processing operator. LookupJoin: Dimension table association operator, which uses polling to associate the Kafka Topic data stream with the external dimension table storage and is converted into a LookupJoin operator. Sink: Output to external storage operator.

[0090] In one embodiment of the present application, the real-time computing engine includes at least one of the following: Flink, Spark Streaming.

[0091] The value of data cannot be reflected in business scenarios with high timeliness requirements, such as real-time risk control and real-time marketing in financial scenarios. Real-time computing engines for stream computing include but are not limited to Flink and Spark Streaming.

[0092] The embodiment of the present application also provides a dimension table association processing device 400, such as Figure 4 As shown, a schematic diagram of the structure of a dimension table association processing device in an embodiment of the present application is provided. The dimension table association processing device 400 at least includes: a processing module 410 and an operation module 420, wherein:

[0093] In one embodiment of the present application, the processing module 410 is specifically used to: in the data stream real-time calculation engine, reallocate the first consumer queue data to the Calc operator according to the Source data source operator.

[0094] The business value of data will decrease rapidly over time, so data should be processed and calculated as soon as possible after it occurs. Data processing delays are very obvious in business scenarios with high timeliness requirements, such as real-time risk control and real-time marketing. Real-time computing can effectively shorten data processing delays and meet business needs. In addition, a unit used to process data in the real-time computing engine, such as reading data, calculating data, correlating data, and outputting data, are all data processing units.

[0095] Dimension table refers to the fact that real-time computing is an event-triggered computing mode, and the triggering data source is an unbounded data stream, which can also be called a driving stream. In actual production scenarios, it is usually necessary to supplement the information of the data stream. The supplementary information is usually attribute information, such as customer personal information, customer account card information, etc. This information is generally placed in a key-value database that stores data in a key-value data model as a dimension table. Dimension table association refers to the process of associating unbounded data streams with dimension table data by associating primary keys during real-time computing.

[0096] The Source data source operator consumes the message queue and redistributes the consumption queue data to the Calc operator.

[0097] It can be understood that the Source data source operator subscribes to and consumes data messages from the source Kafka Topic. The Calc operator is a field calculation and processing operator. Usually, the where condition in the SQL case will be converted into the Calc operator.

[0098] In one embodiment of the present application, the operation module 420 is specifically used to: according to the partition operation of the Calc operator, obtain the grouped data after the partition operation and then input it into the LookupJoin dimension table association operator, and output it to the second consumer queue data through the sink operator.

[0099] According to the partition operation of the Calc operator, when performing the partition operation, the key is the associated primary key of the dimension table association, and the LookupJoin operation is performed after the partition operation grouping.

[0100] It can be understood that the LookupJoin dimension table association operator, specifically the external dimension table storage associated with the Kafka Topic data stream will be converted into a LookupJoin operator. The sink operator is used as an output to the external storage operator.

[0101] Specifically, for the partitioning operation of the Calc operator, i.e., the keyby operation, in the dimension table association scenario, the dimension table association has a natural keyby feature. The natural keyby feature means that when the tables are associated, the data in the two tables can only be associated if they have the same associated primary key, so the associated primary key of the dimension table association can be directly used as the keyby condition. In this way, after the Calc operator is processed, the keyby operation is performed, and the key of the keyby is the associated primary key of the dimension table association, and then LookupJoin and subsequent data processing are performed. Since the keyby operation uses the dimension table association primary key as the key, and because the dimension table association has a hash attribute, when performing the keyby operation, the associated key of the dimension table association is used for hashing, which can partition the data and ensure semantic correctness.

[0102] In one embodiment of the present application, the operation module 420 is also used to

[0103] According to the keyby operation after the Calc operator, the associated primary key associated with the dimension table is used as a condition of the keyby operation to partition the data, thereby obtaining grouped data after the keyby operation;

[0104] The grouped data after the partition operation is hashed and grouped according to the associated primary key associated with the dimension table, and the data of the same group falls into the same instance and then input into the LookupJoin dimension table association operator, and is output to the consumer queue topic through the sink operator.

[0105] In one embodiment of the present application, the operation module 420 is also used to

[0106] Put the data with the same dimension table association key into the same instance of the LookupJoin dimension table association operator for processing, so that the dimension table association key is a subset of the key distributed by the Calc operator;

[0107] The instance of the LookupJoin dimension table association operator caches a portion of the dimension data. When the dimension table is associated, the cache space accessed changes from the first key to the second key. The first key is all the keys, and the second key includes a portion of the all the keys.

[0108] In one embodiment of the present application, the processing module 410 is also used to

[0109] According to the Source data source operator subscribing to and consuming the first data message of the source Kafka Topic, the first data message of the Kafka Topic is redistributed to the Calc operator;

[0110] The operation module 420 is also used to

[0111] When the instance of the LookupJoin dimension table association operator consumes a message, if the corresponding dimension data does not exist in the cache, it will go to the external storage for query, and the dimension data after the query will be put into the cache;

[0112] If the corresponding dimension data is found in the cache, the dimension data in the cache is used directly;

[0113] According to the dimension data in the cache, a second data message is output to the Kafka Topic through the sink operator.

[0114] In one embodiment of the present application, it also includes: a cache hit rate module for

[0115] In the LookupJoin dimension table association operator, the cache hit rate of the dimension table association is improved. The cache hit rate refers to the ratio of the number of times data is read from the cache to the total number of times data is read. The total number of times data is read is the sum of the number of times data is read from the cache and the number of times data is read from the dimension table storage.

[0116] It can be understood that the above-mentioned dimension table association processing device can implement each step of the dimension table association processing method provided in the above-mentioned embodiment. The relevant explanations about the dimension table association processing method are applicable to the dimension table association processing device and will not be repeated here.

[0117] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 5 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.

[0118] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0119] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0120] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a dimension table association processing device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0121] In the data stream real-time computing engine, the data of the first consumer queue is redistributed to the Calc operator according to the Source data source operator;

[0122] According to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator.

[0123] The above application Figure 1The method performed by the dimension table association processing device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0124] The electronic device may also perform Figure 1 The method executed by the dimensional table association processing device in the Figure 1 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0125] The present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, enable the electronic device to execute Figure 1 The method performed by the dimensional table association processing device in the embodiment shown is specifically used to perform:

[0126] In the data stream real-time computing engine, the data of the first consumer queue is redistributed to the Calc operator according to the Source data source operator;

[0127] According to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator.

[0128] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0130] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0132] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0133] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0134] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0135] It should also be noted that the terms "include", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, commodity or device comprising the element.

[0136] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0137] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A dimension table association processing method, wherein: The processing method comprises: In the data stream real-time computing engine, the data of the first consumer queue is redistributed to the Calc operator according to the Source data source operator; According to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator.

2. The method of claim 1, wherein: According to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator, including: According to the keyby operation after the Calc operator, the associated primary key associated with the dimension table is used as a condition of the keyby operation to partition the data, thereby obtaining grouped data after the keyby operation; The grouped data after the partition operation is hashed and grouped according to the associated primary key associated with the dimension table, and the data of the same group falls into the same instance and then input into the LookupJoin dimension table association operator, and is output to the consumer queue topic through the sink operator.

3. The method according to claim 2, further comprising: Put the data with the same dimension table association key into the same instance of the LookupJoin dimension table association operator for processing, so that the dimension table association key is a subset of the key distributed by the Calc operator; The instance of the LookupJoin dimension table association operator caches a portion of the dimension data. When the dimension table is associated, the cache space accessed changes from the first key to the second key. The first key is all the keys, and the second key includes a portion of the all the keys.

4. The method of claim 1, wherein: In the data stream real-time computing engine, reallocating the first consumer queue data to the Calc operator according to the Source data source operator includes: According to the Source data source operator subscribing to and consuming the first data message of the source Kafka Topic, the first data message of the KafkaTopic is redistributed to the Calc operator; According to the partition operation of the Calc operator, the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, and output to the second consumer queue data through the sink operator, including: When the instance of the LookupJoin dimension table association operator consumes a message, if the corresponding dimension data does not exist in the cache, it will go to the external storage for query, and the dimension data after the query will be put into the cache; If the corresponding dimension data is found in the cache, the dimension data in the cache is used directly; According to the dimension data in the cache, a second data message is output to the Kafka Topic through the sink operator.

5. The method of claim 1, wherein: The partition operation after the Calc operator is performed, and the grouped data after the partition operation is obtained and then input into the LookupJoin dimension table association operator, further comprising: In the LookupJoin dimension table association operator, the cache hit rate of the dimension table association is improved. The cache hit rate refers to the ratio of the number of times data is read from the cache to the total number of times data is read. The total number of times data is read is the sum of the number of times data is read from the cache and the number of times data is read from the dimension table storage.

6. The method of claim 1, wherein: The real-time computing engine includes any one or more operators of the Source data source operator, the Calc operator, the LookupJoin dimension table association operator, and the sink operator.

7. The method according to any one of claims 1 to 6, wherein: The real-time computing engine includes at least one of the following: Flink, Spark Streaming.

8. A dimension table association processing device, wherein: The processing device comprises: A processing module, used for reallocating the first consumer queue data to the Calc operator according to the Source data source operator in the data stream real-time calculation engine; The operation module is used to obtain the grouped data after the partition operation according to the partition operation of the Calc operator, input the grouped data into the LookupJoin dimension table association operator, and output the grouped data to the second consumer queue data through the sink operator.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.