Data processing method and device, equipment, medium and product
By generating a hash table and loading the dimension table data into the cache, combining real-time association and data change processing mechanisms, the problem of inaccurate association caused by out-of-order flow data is solved, and the accurate association between dimension table data and flow data is achieved.
Patent Information
- Application Number
- CN202510117452.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
When the prior art associates stream data with dimension table data, it fails to effectively process the out of order of stream data, resulting in inaccurate association.
By obtaining the dimension table, generating a hash table, loading the dimension table data into the cache, and associating it with the real-time stream data based on the hash table. If the dimension table data is changed, the hash table is updated, and after the dimension table data is updated, the relevant stream data is associated with the updated dimension table data.
Real-time perception of the changes in the dimension table data, accurately correlating the dimension table data with real-time stream data, solving the problem of inaccurate association caused by unordered stream data.
Smart Images

Figure CN120045562A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data processing method, apparatus, device, medium, and product. Background Art
[0002] In the field of big data processing, since streaming data lacks some dimensional information, it often needs to be associated with a dimension table. When associating the dimension table data with the streaming data, in order to cope with the changes in the dimension table data, it is necessary to refresh the dimension table data in a timely manner. However, the related technology does not consider the disorder of the streaming data and cannot accurately associate the streaming data with the dimension table data. Summary of the Invention
[0003] This application mainly provides a data processing method, apparatus, device, medium, and product.
[0004] The technical solution of this application is implemented as follows:
[0005] A data processing method, the method includes:
[0006] Obtain a dimension table; the dimension table is used to store dimension table data;
[0007] Generate a hash table based on the dimension table; the hash table is used to load the dimension table data into the cache;
[0008] Associate the dimension table data with real-time streaming data based on the hash table.
[0009] In the above solution, the generating a hash table based on the dimension table includes:
[0010] Set a first field and a second field; the first field is used to associate the real-time streaming data with the dimension table; the second field is used to identify the backfill field in the dimension table;
[0011] Write the dimension table data into the hash table based on the first field, the second field, and the dimension table.
[0012] In the above solution, the associating the dimension table data with real-time streaming data based on the hash table includes:
[0013] Obtain the real-time streaming data;
[0014] Generate a real-time wide table based on the hash table and the real-time streaming data;
[0015] Associate the dimension table data with the real-time streaming data based on the real-time wide table.
[0016] In the above solution, after generating the hash table based on the dimension table, it includes:
[0017] If the dimension table data changes, update the hash table to obtain an updated hash table and updated dimension table data;
[0018] Obtain the stream data and the first time during the update of the hash table; the first time indicates the time when the dimension table data changes; associate the stream data with the updated dimension table data based on the first time.
[0019] In the above solution, the associating the stream data with the updated dimension table data based on the first time includes:
[0020] Determine first data from the stream data based on the first time; the event time of the first data is greater than or equal to the first time;
[0021] Associate the first data with the updated dimension table data.
[0022] In the above solution, the method further includes:
[0023] Determine second data from the stream data based on the first time; the event time of the second data is less than the first time;
[0024] Associate the second data with the dimension table data.
[0025] A data processing device, the device includes:
[0026] An acquisition unit, configured to acquire a dimension table; the dimension table is used to store dimension table data;
[0027] A processing unit, configured to generate a hash table based on the dimension table; the hash table is used to load the dimension table data into a cache;
[0028] The processing unit is configured to associate the dimension table data with real-time stream data based on the hash table.
[0029] An electronic device, including: a processor and a memory for storing a computer program that can run on the processor,
[0030] Wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of the above.
[0031] A storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the steps of the method described in any one of the above.
[0032] A computer product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the steps of the method described in any one of the above.
[0033] The present invention provides a data processing method, apparatus, device, medium and product, which includes obtaining a dimension table; the dimension table is used to store dimension table data; generating a hash table based on the dimension table; the hash table is used to load the dimension table data into a cache; and associating the dimension table data with real-time stream data based on the hash table. That is to say, this application realizes real-time perception of changes in dimension table data and accurately associates dimension table data with real-time stream data by obtaining a dimension table, generating a hash table based on the dimension table, loading the dimension table data into a cache, and associating the dimension table data with real-time stream data based on the hash table; it solves the problem in the related art that the out-of-order of stream data is not considered and the stream data cannot be accurately associated with the dimension table data. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic flowchart of a data processing method provided by this application;
[0035] Figure 2 It is a schematic flowchart of dimension table data update provided by this application;
[0036] Figure 3 It is a schematic flowchart of another data processing method provided by this application;
[0037] Figure 4 It is a schematic structural diagram of a data processing apparatus provided by this application;
[0038] Figure 5 It is a schematic structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to make the objectives, technical solutions and advantages of this application clearer, the technical solutions of this application will be further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limitations to this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0040] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0041] The terms "first / second / third" involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.
[0043] In the related art, a fact table is a table in a data warehouse used to store fact data in business processes. It contains detailed information about specific business events or activities, such as sales transactions, order processing, etc. Each row usually represents a single fact record. Fact tables usually contain numerical data, such as sales amount, quantity, profit, etc., and will contain foreign keys associated with one or more dimension tables to obtain dimension-related detailed information. A dimension table is a table in a data warehouse used to store information describing the fact data environment. It contains dimension attributes related to business processes, such as time, location, product, etc. Dimension tables provide various dimension information for analyzing fact data and can be shared by multiple fact tables.
[0044] Fact tables store various fact data that occur in business processes, while dimension tables store dimension attributes used to describe fact data. They together constitute the core table structure in the data warehouse and support the needs of data analysis and reporting. In stream processing, a data stream can be regarded as an unbounded fact table. Due to the lack of some dimension information, it often needs to be associated with a dimension table. To cope with changes in dimension table data, there are already two solutions for refreshing dimension table data: querying the dimension table in real time or locally caching the dimension table and refreshing it regularly.
[0045] Real-time query generally targets one or a batch of stream data and queries the database where the dimension table is located in real time. This method will frequently query dimension table data, putting great pressure on the database and having low efficiency, and is prone to backpressure.
[0046] Locally caching the dimension table and refreshing it regularly means loading the dimension table data into the cache (usually based on the distributed processing engine Flink), then associating the stream data and the dimension table data in the cache, and refreshing the cache regularly. Although this method improves the query efficiency, it cannot capture changes in dimension table data in a timely manner and cannot accurately associate stream data with dimension table data.
[0047] Embodiments of this application provide a data processing method. Referring to Figure 1 as shown, the method includes the following steps:
[0048] Step S101: Obtain a dimension table.
[0049] Among them, the dimension table is used to store dimension table data.
[0050] It is understandable that the dimension table contains dimension attributes related to business processes, such as time, geographical location, user information, and product information. The dimension table provides various dimension information for analyzing fact data and can be shared by multiple fact tables. The dimension table is stored in the database. In the embodiment of this application, taking the dimension table data stored in Mysql as an example,
[0051] Step S102: Generate a hash table based on the dimension table.
[0052] Among them, the hash table is used to load the dimension table data into the cache.
[0053] In practical applications, a hash table can be generated through a program to load the dimension table data into the cache. Set the configuration parameters of the program for initialization. The configuration parameters can include: the database where the dimension table data is located, the Internet Protocol (IP) address and port of the database, the username and password required to connect to the database, the database name / table name / index name, the associated fields, the backfill fields, the number of threads, and the waiting data duration B. The associated fields can be one or more fields used to associate real-time stream data with the dimension table. The associated fields are used in the dimension table to establish an association with the fact table and are usually the primary key or surrogate key of the dimension table. The associated fields can ensure that the stream data and the dimension table can be accurately and quickly associated during stream data processing without being affected by duplicate dimension table data. The backfill fields can be one or more fields used to identify the backfill fields in the dimension table. The backfill fields can provide detailed information about dimension attributes, such as name, description, type, etc., and can help users analyze the dimension attributes of the data.
[0054] After the configuration parameters are set, the program can read the full amount of dimension table data from the database where the dimension table is located, use the value of the associated field as the key, and the value of the backfill field as the value, write the dimension table data into the hash table, place it in the local cache, and can also discard unnecessary fields to reduce cache occupancy.
[0055] Step S103: Associate the dimension table data with the real-time stream data based on the hash table.
[0056] In practical applications, when real-time stream data arrives, according to the associated fields in the configuration parameters, the stream data can be quickly matched to the corresponding key in the hash table with reference to the data time, and the value of the record in the cache is updated to the backfill field of the stream data to quickly associate the dimension table data with the real-time stream data.
[0057] As can be seen from the above, in this application, a dimension table is obtained, a hash table is generated based on the dimension table, the dimension table is used to store dimension table data, the dimension table data is loaded into the cache, and the dimension table data is associated with the real-time stream data based on the hash table to achieve real-time perception of changes in the dimension table data and accurately associate the dimension table data with the real-time stream data; this solves the problem in the related art that the out-of-order of the stream data is not considered and the stream data cannot be accurately associated with the dimension table data.
[0058] In some embodiments of this application, generating a hash table based on the dimension table includes:
[0059] Set a first field and a second field; the first field is used to associate the real-time stream data with the dimension table; the second field is used to identify the backfill field in the dimension table;
[0060] Write the dimension table data into the hash table based on the first field, the second field, and the dimension table.
[0061] In practical applications, in practical applications, the configuration parameters of the program can be set, and the hash table is generated through the program, and the dimension table data is loaded into the cache. The configuration parameters can include: the database where the dimension table data is located, the Internet IP address and port of the database, the username and password required to connect to the database, the database name / table name / index name, the association field, the backfill field, the number of threads, and the waiting data duration B.
[0062] The association field can be one or more fields, which are used to associate the real-time stream data with the dimension table. The association field is used to establish an association with the fact table in the dimension table and is usually the primary key or surrogate key of the dimension table. The association field can ensure that the stream data and the dimension table can be accurately and quickly associated during the stream data processing, without being affected by the duplication of the dimension table data. The first field can be understood as the association field.
[0063] The backfill field can be one or more fields, which are used to identify the backfill field in the dimension table. The backfill field can provide detailed information about the dimension attributes, such as name, description, type, etc., and can help users analyze the dimension attributes of the data. The second field can be understood as the backfill field.
[0064] After the configuration parameters are set, the program can read the full amount of dimension table data from the database where the dimension table is located, use the value of the association field as the key and the value of the backfill field as the value, write the dimension table data into the hash table, place it in the local cache, and can also discard the unnecessary fields to reduce the cache occupancy.
[0065] In some embodiments of this application, associating the dimension table data with the real-time stream data based on the hash table includes:
[0066] Obtain the real-time stream data;
[0067] Generate a real-time wide table based on a hash table and real-time stream data;
[0068] Associate the dimension table data with the real-time stream data based on the real-time wide table.
[0069] In practical applications, when real-time stream data arrives, a real-time wide table can be formed when the real-time stream data and dimension table data are associated in real time based on the hash table. The real-time wide table can associate the business entity data and various related dimension data into one table.
[0070] In some embodiments of the present application, after generating a hash table based on the dimension table, it includes:
[0071] If the dimension table data changes, update the hash table to obtain the updated hash table and the updated dimension table data;
[0072] Obtain the stream data and the first time during the hash table update; the first time indicates the time when the dimension table data changes;
[0073] Associate the stream data with the updated dimension table data based on the first time.
[0074] In practical applications, it is possible to monitor in real time whether the dimension table data has changed. Once it changes, the changed data is immediately synchronized to the cache. A data change awareness module can be set in Mysql. For example, the Flink Change Data Capture (CDC) tool can be used as the data change awareness module. The data change awareness module will act as a slave node of Mysql to synchronize the binary log (binlog) in real time. The binlog records all data change operations of Mysql. Monitor the data change operations and capture the changed data. A dimension table update blocking module can also be set to update the dimension table data in the cache.
[0075] Reference Figure 2 As shown, record the dimension table data change timestamp as timestamp A, and record the time when the data change awareness module senses the change as timestamp A'. In the embodiments of the present application, timestamp A' is approximately regarded as timestamp A. Record the time when the dimension table data in the cache is updated as T, that is, the time when the hash table is updated is recorded as T. The first time can be understood as the time corresponding to timestamp A'.
[0076] Due to the out-of-order problem of streaming data, after obtaining the timestamp A' sent by the data change awareness module, before the dimension table data in the cache is updated completely (before time T), the dimension table update blocking module will take the data in the incoming streaming data whose event time is greater than or equal to A' as blocked data and write it into the message queue of the Remote Dictionary Server (Redis); the data with an event time less than A' is regarded as normal data and associated with the dimension table data in the cache.
[0077] After the dimension table data in the cache is updated completely, the dimension table update blocking module stops writing the streaming data into Redis and reads the blocked data in the Redis queue and associates it with the updated dimension table data in the cache. At this time, the dimension table update blocking module no longer blocks data, and the processing of streaming data resumes normal.
[0078] Within a period of time after the dimension table data in the cache is updated completely, there may be data with an event time less than the timestamp A arriving. It can be regarded as late data. The late data should be associated with the dimension table data whose dimension table expiration time is A. According to the waiting data duration B in the configuration parameters, the late data can be associated with the dimension table data. After waiting for B time, it is considered that all the late data has arrived. To reduce the occupancy of the cache, the original records are deleted and only the new records are retained.
[0079] If the dimension table data changes again within the waiting time B, the same processing as above can be performed, and the above new records become the original records to solve the problem of inaccurate association caused by the out-of-order of streaming data.
[0080] In some embodiments of the present application, associating the streaming data with the updated dimension table data based on the first time includes:
[0081] Determining the first data from the streaming data based on the first time; the event time of the first data is greater than or equal to the first time;
[0082] Associating the first data with the updated dimension table data.
[0083] In practical applications, due to the out-of-order problem of streaming data, after obtaining the timestamp A' sent by the data change awareness module, before the dimension table data in the cache is updated completely (before time T), the dimension table update blocking module will take the data in the incoming streaming data whose event time is greater than or equal to A' as blocked data and write it into the message queue of the Remote Dictionary Server (Redis); the data with an event time less than A' is regarded as normal data and associated with the dimension table data in the cache. The first data can be understood as blocked data.
[0084] After the dimension table data update in the cache is completed, the dimension table update blocking module stops writing streaming data to Redis, reads the blocking data in the Redis queue, and associates it with the updated dimension table data in the cache. At this time, the dimension table update blocking module no longer blocks data, and the streaming data processing resumes normal.
[0085] In some embodiments of the present application, the method further includes:
[0086] Determine the second data from the streaming data based on the first time; the event time of the second data is less than the first time;
[0087] Associate the second data with the dimension table data.
[0088] In practical applications, due to the out-of-order problem of streaming data, after obtaining the timestamp A' sent by the data change perception module, before the dimension table data in the cache is updated (before time T), the dimension table update blocking module will use the data in the incoming streaming data with an event time greater than or equal to A' as blocking data and write it to the Redis message queue; the data with an event time less than A' is used as normal data and associated with the dimension table data in the cache. The second data can be understood as normal data.
[0089] From the above, it can be seen that the embodiments of the present application capture the changes in the dimension table data in real time through the data change perception module to automatically and timely load the latest data of the dimension table into the cache; divide the streaming data in the cache update window period into blocking data and normal data, associate the normal data with the data before the dimension table change, write the blocking data to Redis first, and then associate it with the updated dimension table data after the cache update is completed to solve the problem of inaccurate association caused by out-of-order streaming data.
[0090] In a realizable scenario, referring to Figure 3 as shown, the data processing method of the embodiments of the present application can be implemented in the following manner:
[0091] 1. Configure program parameters and perform initialization.
[0092] 2. Load the dimension table data into the cache.
[0093] 3. Connect to the real-time streaming data and perform real-time association.
[0094] 4. The motion data update perception module perceives and synchronizes the changes in the dimension table data.
[0095] 5. When the dimension table data cache is synchronized, the dimension table update blocking module processes the out-of-order data.
[0096] 6. After the dimension table data synchronization is completed, take out the blocking data and wait for the late data.
[0097] Based on the same inventive concept as described above, Figure 4 FIG. is a schematic structural diagram of a data processing apparatus provided by an embodiment of the present invention. The data processing apparatus 400 includes:
[0098] An obtaining unit 401, configured to obtain a dimension table; the dimension table is used to store dimension table data;
[0099] A processing unit 402, configured to generate a hash table based on the dimension table; the hash table is used to load the dimension table data into a cache;
[0100] The processing unit 402 is configured to associate the dimension table data with real-time stream data based on the hash table.
[0101] In some embodiments of the present application, the processing unit 402 is configured to set a first field and a second field; the first field is used to associate the real-time stream data with the dimension table; the second field is used to identify a backfill field in the dimension table;
[0102] Write the dimension table data into the hash table based on the first field, the second field, and the dimension table.
[0103] In some embodiments of the present application, the processing unit 402 is configured to obtain real-time stream data;
[0104] Generate a real-time wide table based on the hash table and the real-time stream data;
[0105] Associate the dimension table data with the real-time stream data based on the real-time wide table.
[0106] In some embodiments of the present application, the processing unit 402 is configured to, if the dimension table data changes, update the hash table to obtain an updated hash table and updated dimension table data;
[0107] Obtain the stream data and a first time during the hash table update; the first time indicates the time when the dimension table data changes; associate the stream data with the updated dimension table data based on the first time.
[0108] In some embodiments of the present application, the processing unit 402 is configured to determine first data from the stream data based on the first time; the event time of the first data is greater than or equal to the first time;
[0109] Associate the first data with the updated dimension table data.
[0110] In some embodiments of the present application, the processing unit 402 is configured to determine second data from the stream data based on the first time; the event time of the second data is less than the first time;
[0111] Associate the second data with the dimension table data.
[0112] Based on the foregoing embodiments, an embodiment of the present application provides an electronic device. Figure 5 FIG. Figure 5 is a schematic diagram of a hardware structure of the electronic device according to an embodiment of the present invention. The electronic device 500 includes: at least one processor 501, a memory 502. Optionally, the electronic device 500 may further include at least one communication interface 503. Each component in the electronic device 500 is coupled together through a bus system 504. It can be understood that the bus system 504 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 5 all kinds of buses are labeled as the bus system 504.
[0113] Based on the hardware implementation of the foregoing program module, the communication interface 503 is capable of interacting with other communication devices for information.
[0114] The processor 501 is connected to the communication interface 503 to realize information interaction with other communication devices, and is used to execute the method provided by the above one or more technical solutions when running a computer program.
[0115] The memory 502 stores the computer program.
[0116] Specifically, the processor 501 is used to obtain a dimension table; the dimension table is used to store dimension table data.
[0117] Generate a hash table based on the dimension table; the hash table is used to load the dimension table data into the cache.
[0118] Associate the dimension table data with the real-time stream data based on the hash table.
[0119] In some embodiments of the present application, the processor 501 is used to set a first field and a second field; the first field is used to associate the real-time stream data with the dimension table; the second field is used to identify the backfill field in the dimension table.
[0120] Write the dimension table data into the hash table based on the first field, the second field, and the dimension table.
[0121] In some embodiments of the present application, the processor 501 is used to obtain the real-time stream data.
[0122] Generate a real-time wide table based on the hash table and the real-time stream data.
[0123] Associate the dimension table data with the real-time stream data based on the real-time wide table.
[0124] In some embodiments of the present application, the processor 501 is used to update the hash table if the dimension table data changes, to obtain an updated hash table and updated dimension table data.
[0125] Obtain the streaming data and the first time during the hash table update; the first time indicates the time when the dimension table data changes;
[0126] Associate the streaming data with the updated dimension table data based on the first time.
[0127] In some embodiments of the present application, the processor 501 is configured to determine first data from the streaming data based on the first time; the event time of the first data is greater than or equal to the first time;
[0128] Associate the first data with the updated dimension table data.
[0129] In some embodiments of the present application, the processor 501 is configured to determine second data from the streaming data based on the first time; the event time of the second data is less than the first time;
[0130] Associate the second data with the dimension table data.
[0131] It can be understood that the memory 502 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 502 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0132] The memory 502 in the embodiments of the present invention is used to store various types of data to support the operation of the electronic device 500. Examples of such data include: any computer program for operating on the electronic device 500, and the program for implementing the method of the embodiments of the present invention may be included in the memory 502.
[0133] The methods disclosed in the above embodiments of the present invention can be applied to the processor 501 or implemented by the processor 501. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above methods can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the methods disclosed in the embodiments of the present invention, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the foregoing methods.
[0134] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components for performing the above methods.
[0135] Based on the foregoing embodiments, an embodiment of the present application further provides a computer product, including a computer program, which when executed by a processor, implements Figure 1 the steps in the data processing method provided by the corresponding embodiment.
[0136] Based on the foregoing embodiments, an embodiment of the present application further provides a storage medium storing computer-executable instructions configured to execute Figure 1 the data processing method provided by the corresponding embodiment.
[0137] It should be noted that the above computer storage medium may be a memory such as ROM, PROM, EPROM, EEPROM, FRAM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0138] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0139] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0141] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate for implementation in the processFigure 1 means for the functions specified in one or more processes and / or blocks Figure 1 or in one or more blocks.
[0142] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 or more processes and / or blocks Figure 1 or in one or more blocks.
[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one Figure 1 or more processes and / or blocks Figure 1 or in one or more blocks.
[0144] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Obtain a dimension table; the dimension table is used to store dimension table data; Generate a hash table based on the dimension table; the hash table is used to load the dimension table data into a cache; The dimension table data is associated with the real-time stream data based on the hash table.
2. The method according to claim 1, characterized in that The generating a hash table based on the dimension table includes: Setting a first field and a second field; the first field is used to associate the real-time stream data with the dimension table; the second field is used to identify a backfill field in the dimension table; The dimension table data is written into the hash table based on the first field, the second field, and the dimension table.
3. The method according to claim 1, characterized in that The associating the dimension table data with the real-time stream data based on the hash table includes: Acquire the real-time streaming data; Generate a real-time wide table based on the hash table and the real-time stream data; The dimension table data is associated with the real-time stream data based on the real-time wide table.
4. The method according to claim 1, characterized in that: After the hash table is generated based on the dimension table, the method includes: If the dimension table data is changed, update the hash table to obtain an updated hash table and updated dimension table data; The stream data and the first time during the update of the hash table are obtained; the first time indicates the time when the dimension table data is changed; and the stream data is associated with the updated dimension table data based on the first time.
5. The method according to claim 4, characterized in that The associating the stream data with the updated dimension table data based on the first time includes: Determine first data from the stream data based on the first time; the event time of the first data is greater than or equal to the first time; The first data is associated with the updated dimension table data.
6. The method according to claim 4, characterized in that The method further comprises: Determine second data from the stream data based on the first time; the event time of the second data is less than the first time; The second data is associated with the dimension table data.
7. A data processing device, characterized in that: The device comprises: An acquisition unit, used for acquiring a dimension table; the dimension table is used for storing dimension table data; A processing unit, configured to generate a hash table based on the dimension table; the hash table is configured to load the dimension table data into a cache; The processing unit is used to associate the dimension table data with the real-time stream data based on the hash table.
8. An electronic device, characterized in that: include: a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of claims 1 to 6.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.