Data aggregation method and device, electronic equipment and storage medium
By mapping and assembling the physical model in the data warehouse as the target business model, and combining distributed cache and analytical database, the high cost problem caused by frequent changes in business data in the data warehouse is solved, and efficient and accurate data aggregation processing is achieved.
Patent Information
- Application Number
- CN202510257968.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-04
AI Technical Summary
When business data changes frequently in data warehouses, the existing technology cannot efficiently and accurately assemble business models, resulting in data aggregation processing costs that are too high and may cause wrong business models.
By mapping multiple physical models with the same primary key field in the data warehouse into the target business model, and generating data processing event identification, querying whether the same event identification exists, creating a business model aggregation event if it does not exist, assembling the target business model based on the event, and waiting for a preset time when necessary or obtaining data from a distributed cache and analytical database, performing data aggregation operations.
It effectively reduces the cost of real-time data aggregation in data warehouses, improves the accuracy and efficiency of data processing, reduces the frequency of data processing, and ensures the instant availability of data and system stability.
Smart Images

Figure CN120256527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for data aggregation. Background Art
[0002] When the business data in the data warehouse changes, technicians need to assemble the physical model into a business model in near real-time to provide subsequent near real-time data analysis. During the process of real-time data collection and processing in the data warehouse, frequent state changes of business data will be faced in a short period of time. Correspondingly, during the process of converting the physical model into a business model, the problem of frequent processing of the business model caused by frequent changes in business data in a short period of time will be faced.
[0003] In the related art, there is no effective solution to the problem of excessive processing costs caused by frequent changes in business data, and when the business data is abnormal, there is a problem that the correct business model cannot be obtained.
[0004] Therefore, when the business data in the data warehouse changes frequently, how to efficiently and accurately assemble the business model to effectively reduce the cost of real-time data aggregation and processing in the data warehouse is a technical problem to be solved urgently. Summary of the Invention
[0005] Aiming at the above problems existing in the prior art, the present invention provides a method, apparatus, electronic device, and storage medium for data aggregation, which realizes efficiently and accurately assembling the business model when the business data in the data warehouse changes frequently, so as to effectively reduce the cost of real-time data aggregation and processing in the data warehouse.
[0006] The present invention provides a method for data aggregation, including the following steps.
[0007] In response to a change in any physical model in the data warehouse, map the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model to a target business model including the primary key field; wherein, the physical model is a representation of the actual data structure in the data warehouse; generate a data processing event identifier for the target business model according to the primary key field; query whether there is an event identifier the same as the data processing event identifier; if there is no event identifier the same as the data processing event identifier, create a corresponding business model aggregation event; based on the business model aggregation event, assemble the target business model with the primary key field as the primary key according to multiple physical models in the data warehouse having the primary key field; based on the assembled target business model, perform a data aggregation operation.
[0008] A method for data aggregation provided by the present invention further includes: waiting for a first preset duration before assembling the target business model with the primary key fields of multiple physical models in the data warehouse as the primary key.
[0009] A method for data aggregation provided by the present invention further includes: in response to a change in any physical model in the data warehouse, writing the data of the changed physical model into a distributed cache.
[0010] Writing the data of the changed physical model into the distributed cache in the method for data aggregation provided by the present invention includes: according to the primary key field of the changed physical model and in combination with a preset cache duration, writing the data of the changed physical model into the distributed cache.
[0011] A method for data aggregation provided by the present invention further includes: in response to a change in any physical model in the data warehouse, writing the data of the physical model into a preset analytical database.
[0012] Assembling the target business model with the primary key fields of multiple physical models in the data warehouse as the primary key in the method for data aggregation provided by the present invention includes: obtaining the data of the multiple physical models from the distributed cache; and assembling the target business model with the primary key fields as the primary key according to the data of the multiple physical models.
[0013] A method for data aggregation provided by the present invention further includes: in response to the distributed cache missing some or all of the data of the multiple physical models; obtaining the missing part or all of the data of the multiple physical models from the analytical database.
[0014] A method for data aggregation provided by the present invention further includes: before assembling the target business model with the primary key fields of multiple physical models in the data warehouse as the primary key, verifying the data of the multiple physical models; in response to there being data that fails the verification among the data of the multiple physical models, waiting for a second preset duration and then re-executing the step of obtaining the data of the multiple physical models from the distributed cache and subsequent steps until the obtained data of the multiple physical models passes the verification or reaches a preset retry count.
[0015] The present invention also provides a data aggregation device, including the following modules: A mapping module, configured to map, in response to a change in any physical model in a data warehouse, the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model into a target business model including the primary key field; wherein the physical model is a representation of the actual data structure in the data warehouse; a generation module, configured to generate a data processing event identifier of the target business model according to the primary key field; a query module, configured to query whether there is an event identifier identical to the data processing event identifier; a creation module, configured to create a corresponding business model aggregation event when there is no event identifier identical to the data processing event identifier; an assembly module, configured to assemble, based on the business model aggregation event and according to multiple physical models having the primary key field in the data warehouse, with the primary key field as the primary key, to obtain the target business model; an execution module, configured to perform a data aggregation operation based on the assembled target business model.
[0016] According to an apparatus for data aggregation provided by the present invention, the apparatus further includes: a waiting module, configured to wait for a first preset duration.
[0017] According to an apparatus for data aggregation provided by the present invention, the apparatus further includes: a first writing module, configured to write the data of the changed physical model into a distributed cache in response to a change in any physical model in the data warehouse.
[0018] According to an apparatus for data aggregation provided by the present invention, the apparatus further includes: a second writing module, configured to write the data of the physical model into a preset analytical database in response to a change in any physical model in the data warehouse.
[0019] According to an apparatus for data aggregation provided by the present invention, the assembling, according to multiple physical models having the primary key field in the data warehouse, with the primary key field as the primary key, to obtain the target business model includes: obtaining the data of the multiple physical models from the distributed cache; and assembling, according to the data of the multiple physical models, with the primary key field as the primary key, to obtain the target business model.
[0020] According to an apparatus for data aggregation provided by the present invention, the apparatus further includes: an obtaining module, configured to, in response to the distributed cache missing some or all of the data of the multiple physical models, obtain some or all of the missing data of the multiple physical models from the analytical database.
[0021] A device for data aggregation provided by the present invention, the device further includes: a verification module for verifying the data of the multiple physical models; a retry module for, in response to the data of the multiple physical models having data that fails verification, after waiting for a second preset duration, re-executing the steps of obtaining the data of the multiple physical models from the distributed cache and subsequent steps until the data of the multiple physical models obtained passes verification or reaches a preset retry count.
[0022] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the data aggregation method as described in any one of the above.
[0023] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the data aggregation method as described in any one of the above.
[0024] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the data aggregation method as described in any one of the above.
[0025] The data aggregation method provided by the present invention, in response to a change in any physical model in the data warehouse, maps the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model into a target business model including the primary key field; generates a data processing event identifier for the target business model according to the primary key field; if no event identifier identical to the data processing event identifier is found through query, creates a corresponding business model aggregation event; based on the business model aggregation event, assembles the target business model with the primary key field as the primary key according to the multiple physical models having the primary key field in the data warehouse; and performs a data aggregation operation based on the assembled target business model. Since multiple physical models are mapped into one target business model for data aggregation processing, compared with the solution of separately performing data aggregation processing on multiple physical models, the frequency of data processing is effectively reduced. In the case of frequent changes in business data, the business model can be efficiently and accurately assembled, effectively reducing the cost of real-time data aggregation processing in the data warehouse. Description of the Drawings
[0026] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is one of the flow diagrams of the data aggregation method provided by the present invention.
[0028] Figure 2 It is the second of the flow diagrams of the data aggregation method provided by the present invention.
[0029] Figure 3 It is the third of the flow diagrams of the data aggregation method provided by the present invention.
[0030] Figure 4 It is the process diagram of the data aggregation provided by the present invention.
[0031] Figure 5 It is the structural diagram of the data aggregation device provided by the present invention.
[0032] Figure 6 It is the structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0033] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0034] The following combines Figures 1 - 3 to describe the data aggregation method of the present invention.
[0035] Figure 1 It is one of the flow diagrams of the data aggregation method provided by the present invention. This method is executed by a data warehouse processing system. As Figure 1 shown, this method includes the following: Step 101, in response to a change in any physical model in the data warehouse, map the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model into a target business model including the primary key field.
[0036] In the specific implementation process, the change of any physical model (such as the table structure) in the data warehouse can be known by consuming database logs, such as binlog (Binary Log, binary log). Binlog is an important component in the MySQL database. It records all SQL statements that have modified data or may modify data, and these statements include INSERT, UPDATE, DELETE, etc.
[0037] The physical model is a representation of the actual data structure in the data warehouse, which includes the storage structure of data, access paths, indexing methods, etc. The physical model can include various structures such as tables, indexes, views, etc. Among them, the table is the most commonly used physical model and can be used to store business data such as sales data, customer information, inventory data, etc.
[0038] The structure of a table consists of rows and columns. Among them, the columns constitute the schema of the table and describe information such as the data type and length in the table. The table name is used to uniquely identify the table, the field name is used to identify the columns in the table, the field type defines the data type stored in the column, and the constraint conditions are used to ensure the integrity and consistency of the data. The primary key field is one or more fields in the table and is used to uniquely identify each row record in the table.
[0039] Only as an example, there is a table named orders in the data warehouse, which is used to store order information. This table is regarded as a physical model in the data warehouse, and order_id (order identifier) is used as the primary key field of this table.
[0040] The physical model shares the primary key field with at least one other physical model in the data warehouse. In this step, the monitored physical model that has changed and one or more physical models in the data warehouse that have the same primary key field as this physical model are mapped to the target business model. The primary key field shared by multiple physical models is used as the primary key field of the target business model.
[0041] The business model refers to the target business model abstracted from the physical model and used to describe business entities and their relationships. The business model usually corresponds to specific business requirements or business scenarios and is the basis for data analysis, report generation, and decision support.
[0042] Only as an example, the data warehouse contains multiple physical models related to orders, such as the orders table (storing basic order information), the order_items table (storing order item information), the customers table (storing customer information), etc. These physical models each contain different fields and table structures, but are all closely related to the order business and have the same key field (order ID). Mapping the above multiple physical models to a logically coherent target business model, the target business model includes the following information: order ID, customer ID, order status, order time, etc., which can be obtained from the orders table; product ID, quantity, unit price, etc., which can be obtained from the order_items table; customer ID, name, address, etc., which can be obtained from the customers table.
[0043] Step 102: Generate a data processing event identifier for the target business model according to the primary key field.
[0044] The data processing event identifier is used to uniquely identify a data processing event. In real-time data processing, a data processing event is an operation that performs a certain processing (such as aggregation, transformation, filtering, etc.) on a set of data. The data processing event identifier is very important for tracking the processing status of the event, avoiding duplicate processing, and ensuring data consistency. For example, a data processing event that is used to process order data and generate a business model. This event can be defined as a processing flow that includes multiple steps, such as reading data from a physical model (database table), performing data transformation and aggregation, and then writing the result to another table or storage system.
[0045] In the specific implementation process, the data processing event identifier of the target business model can be generated in various ways according to the primary key field.
[0046] For example, the primary key field (e.g., order ID) can be directly used as the data processing event identifier. Another example is that the primary key field can be combined with other information (such as event type, etc.) to form a more complex event identifier.
[0047] Step 103: Query whether there is an event identifier that is the same as the data processing event identifier.
[0048] In the specific implementation process, according to the storage method of the data processing event identifier, a suitable query tool or interface can be selected to query whether there is an event identifier that is the same as the data processing event identifier. The data processing event identifier can be stored in storage locations such as a logging system, an event database, or the state management of a real-time computing framework. The query methods for the data processing event identifier can include SQL query, API call query, log analysis query, etc.
[0049] Step 104: If there is no event identifier that is the same as the data processing event identifier, create its corresponding business model aggregation event.
[0050] A business model aggregation event is a data processing event that performs aggregation processing on a business model.
[0051] When querying the data warehouse processing system, if there is no event identifier that is the same as the current data processing event identifier, then create the corresponding business model aggregation event.
[0052] If there is already an event identifier that is the same as the current data processing event identifier in the data warehouse processing system, then end the current data processing flow and do not execute the subsequent steps.
[0053] In some embodiments, before performing the next step, it is necessary to wait for a first preset duration. The value of the first preset duration can be set according to the data update frequency of the data warehouse.
[0054] In the embodiments provided by the present invention, by setting a waiting period of a first preset duration, it is ensured that the data in the physical model where the current data changes and the relevant physical models with the same key fields are partially or fully updated. After that, the assembly of the target business model is executed, which can effectively reduce the frequency of data processing (because there is no need to immediately assemble the business model every time the data changes), and at the same time effectively improve the accuracy of data aggregation in the data warehouse (because the data used for assembly is the latest and most accurate).
[0055] Only as an example, the data warehouse system of an e-commerce platform includes two physical models: "order model" and "inventory model". These two models share a key field: product ID. During the operation of the e-commerce platform, users will perform frequent operations on a certain product, such as adding to the shopping cart, placing an order, making a payment, etc. These operations will cause the order data associated with the product to change continuously. For example, the order status may change from placed to paid. At the same time, these operations of users will also affect the inventory data in a short period of time. For example, before and after payment, the inventory quantity of the product will decrease accordingly. To reduce the processing cost brought by these frequently changing data to the data aggregation task, a first preset duration is set, such as 5 minutes. Within this duration, the operations of users on this product are partially or fully completed. After the waiting duration ends, the system starts to execute the assembly of the "sales analysis model" mapped from the "order model" and the "inventory model". Thus, the frequency of data aggregation in the data warehouse can be effectively reduced, and the efficiency of data aggregation in the data warehouse can be improved.
[0056] Step 105: Based on the business model aggregation event, assemble the target business model according to multiple physical models in the data warehouse that have a primary key field.
[0057] In the specific implementation process, the data warehouse processing system can select relevant data from multiple physical models (such as "sales order model" and "customer information model") according to the primary key field (such as "customer ID") to assemble the target business model. For example, for customer A, all order information of this customer can be selected from the "sales order model", and the basic information of this customer can be selected from the "customer information model".
[0058] Step 106: Based on the assembled target business model, perform a data aggregation operation.
[0059] The data aggregation operation can include operations on multiple data fields in the target business model. For example, sum the order amounts in the "sales analysis model" to obtain the total sales amount of all products within a certain period of time. Another example is to calculate the inventory data in the "sales analysis model" to obtain key indicators such as the increase or decrease of inventory and the turnover rate.
[0060] The data aggregation operation may also include operations such as converting the data format. For example, the order status represented by numbers in the physical model is converted into a status classification with a more understandable text description (such as ordered, paid, shipped, etc.).
[0061] After the data aggregation operation is completed, it is necessary to clear the signal of the data processing event generated by this data processing.
[0062] Figure 2 It is the second flowchart of the data aggregation method provided by the present invention. As Figure 2 shown, the method includes the following: Step 201, in response to a change in any physical model in the data warehouse, write the data of the changed physical model into the distributed cache, and write the data of the changed physical model into a preset analytical database.
[0063] In the specific implementation process, according to the primary key field of the changed physical model, combined with the preset cache duration, write the data of the changed physical model into the distributed cache. The cache duration can be determined according to the actual data processing situation. For example, a cache duration of 5 minutes can be set, and within 5 minutes after the data is stored in the cache, it can be efficiently read from the cache.
[0064] The preset analytical database may include, but is not limited to, databases such as StarRocks and NoSQL.
[0065] In some embodiments, the StarRocks database can be used. It is a high-performance and scalable analytical database. As a high-performance distributed OLAP (Online Analytical Processing) database, StarRocks is good at storing and querying a large amount of historical data. When subsequently, due to the expiration of the cache duration or other reasons, some or all of the data of the physical model is not found in the distributed cache, the required data can be accurately obtained from the StarRocks database.
[0066] Step 202, map the changed physical model and one or more physical models in the data warehouse that have the same primary key field as this physical model into a target business model including the primary key field.
[0067] For the detailed description of this step, refer to the relevant content in step 101, which will not be elaborated here.
[0068] 203. Generate a data processing event identifier for the target business model according to the primary key field.
[0069] For the detailed description of this step, refer to the relevant content in step 102, which will not be elaborated here.
[0070] Step 204: If no event identifier identical to the data processing event identifier is found through query, create a corresponding business model aggregation event.
[0071] For the detailed description of this step, refer to the relevant content in Steps 103 and 104, which will not be elaborated here.
[0072] Step 205: Based on the business model aggregation event, obtain data of multiple physical models from the distributed cache.
[0073] In the case where some or all of the data of multiple physical models are not found in the distributed cache, obtain some or all of the data of the unhit multiple physical models from the analytical database.
[0074] Step 206: Assemble to obtain a target business model with the primary key field as the primary key according to the data of multiple physical models.
[0075] For the detailed description of this step, refer to the relevant content in Step 104, which will not be elaborated here.
[0076] Step 207: Execute a data aggregation operation based on the assembled target business model.
[0077] For the detailed description of this step, refer to the relevant content in Step 105, which will not be elaborated here.
[0078] In the embodiments provided by the present invention, as Figure 4 shown, when the data of the physical model changes, write it into the preset analytical database and the distributed cache; when assembling the target business model, obtain the data of multiple physical models from the distributed cache, and when some or all of the data of multiple physical models are not found in the distributed cache, obtain some or all of the data of the unhit multiple physical models from the analytical database. Since the distributed cache has a high read and write speed, it can significantly reduce the time delay for obtaining physical model data, thereby effectively improving the efficiency of data aggregation processing in the data warehouse. The combined use of the distributed cache and the analytical database enables the system to handle a large number of concurrent requests. Since data can be obtained from the analytical database when the distributed cache misses, it not only ensures the immediate availability of data but also ensures the stability of the data warehouse processing system.
[0079] Figure 3 is the third flow diagram of the data aggregation method provided by the present invention. As Figure 3 shown, the method includes the following: Step 301: In response to the change of any physical model in the data warehouse, write the data of the changed physical model into the distributed cache and write the data of the changed physical model into the preset analytical database.
[0080] For the detailed description of this step, refer to the relevant content in Step 201, which will not be elaborated here.
[0081] Step 302: Map the changed physical model and one or more physical models in the data warehouse that have the same primary key field as this physical model into a target business model that includes the primary key field.
[0082] For the detailed description of this step, refer to the relevant content in Step 101, which will not be elaborated here.
[0083] Step 303: Generate a data processing event identifier for the target business model according to the primary key field.
[0084] For the detailed description of this step, refer to the relevant content in Step 102, which will not be elaborated here.
[0085] Step 304: If there is no event identifier identical to the data processing event identifier after querying, create the corresponding business model aggregation event.
[0086] For the detailed description of this step, refer to the relevant content in Steps 103 and 104, which will not be elaborated here.
[0087] Step 305: Based on the business model aggregation event, obtain the data of multiple physical models from the distributed cache.
[0088] In the case where some or all of the data of multiple physical models are not found in the distributed cache, obtain some or all of the data of the multiple physical models that are not found from the analytical database.
[0089] Step 306: Check the data of multiple physical models. In response to the existence of data that fails the check in the data of multiple physical models, after waiting for the second preset duration, obtain the data of multiple physical models from the distributed cache again.
[0090] Checking the data of multiple physical models includes checking the data format, integrity, etc. of the physical model data.
[0091] The second preset duration can be flexibly set according to the data change situation of the data warehouse. Since there may be data asynchronization problems in the distributed cache, which may lead to the failure of checking the physical model data read from the cache. Therefore, by setting the second preset duration to wait, it is ensured that the data in the distributed cache has been synchronized and is ready, so that the correct physical model data can be obtained when reading again.
[0092] In the embodiments provided by the present invention, as Figure 4 shown, through the retry mechanism, the accuracy and fault tolerance of data aggregation in the data warehouse can be effectively improved.
[0093] After waiting for the second preset duration, the data warehouse processing system re-executes the steps of obtaining data of multiple physical models from the distributed cache and subsequent steps until the data of the multiple physical models obtained passes the verification or reaches the preset retry times.
[0094] After reaching the preset retry times and still failing to read the correct data of the physical model, the data warehouse processing system can send an alarm message to remind manual intervention for processing.
[0095] Step 307: Assemble the obtained data of multiple physical models with the primary key field as the primary key to obtain the target business model.
[0096] For the detailed description of this step, refer to the relevant content in step 104, which will not be elaborated here.
[0097] Step 308: Perform a data aggregation operation based on the assembled target business model.
[0098] For the detailed description of this step, refer to the relevant content in step 105, which will not be elaborated here.
[0099] Next, the data aggregation device provided by the present invention will be described. The data aggregation device described below can be correspondingly referred to the data aggregation method described above.
[0100] Figure 5 is a schematic structural diagram of the data aggregation device provided by the present invention. As Figure 5 shown, the device 500 includes the following modules.
[0101] The mapping module 510 is configured to map the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model to a target business model including the primary key field in response to a change in any physical model in the data warehouse.
[0102] The generation module 520 is configured to generate a data processing event identifier for the target business model according to the primary key field.
[0103] The query module 530 is configured to query whether there is an event identifier that is the same as the data processing event identifier.
[0104] The creation module 540 is configured to create a corresponding business model aggregation event if it is queried that there is no event identifier that is the same as the data processing event identifier.
[0105] An assembly module 550, configured to aggregate events based on the business model, and assemble the target business model with the primary key field as the primary key according to multiple physical models having the primary key field in the data warehouse.
[0106] An execution module 560, configured to perform a data aggregation operation based on the assembled target business model.
[0107] In some embodiments, the apparatus further includes: a waiting module, configured to wait for a first preset duration.
[0108] In some embodiments, the apparatus further includes: a first writing module, configured to write the data of the physical model that has changed in response to a change in any physical model in the data warehouse into a distributed cache.
[0109] In some embodiments, the apparatus further includes: a second writing module, configured to write the data of the physical model into a preset analytical database in response to a change in any physical model in the data warehouse.
[0110] In some embodiments, the assembling the target business model with the primary key field as the primary key according to multiple physical models having the primary key field in the data warehouse includes: obtaining the data of the multiple physical models from the distributed cache; and assembling the target business model with the primary key field as the primary key according to the data of the multiple physical models.
[0111] In some embodiments, the apparatus further includes: an obtaining module, configured to, in response to the distributed cache missing some or all of the data of the multiple physical models, obtain the missing some or all of the data of the multiple physical models from the analytical database.
[0112] In some embodiments, the apparatus further includes: a verification module, configured to verify the data of the multiple physical models; and a retry module, configured to, in response to there being data that fails verification in the data of the multiple physical models, after waiting for a second preset duration, re-execute the step of obtaining the data of the multiple physical models from the distributed cache and subsequent steps until the obtained data of the multiple physical models passes verification or a preset retry count is reached.
[0113] Figure 6 An entity structure diagram of an electronic device is exemplified, as Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute the method for data aggregation. The method includes: in response to a change in any physical model in the data warehouse, mapping the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model into a target business model including the primary key field; generating a data processing event identifier for the target business model according to the primary key field; querying whether there is an event identifier that is the same as the data processing event identifier; if there is no event identifier that is the same as the data processing event identifier, creating a corresponding business model aggregation event; based on the business model aggregation event, assembling the target business model with the primary key field as the primary key according to multiple physical models having the primary key field in the data warehouse; and performing a data aggregation operation based on the assembled target business model.
[0114] In addition, when the logical instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0115] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data aggregation method provided by each of the above methods. The method includes: in response to a change in any physical model in the data warehouse, mapping the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model into a target business model including the primary key field; generating a data processing event identifier for the target business model according to the primary key field; querying whether there is an event identifier identical to the data processing event identifier; if there is no event identifier identical to the data processing event identifier, creating a corresponding business model aggregation event; based on the business model aggregation event, assembling the target business model with the primary key field as the primary key according to multiple physical models in the data warehouse having the primary key field; and performing a data aggregation operation based on the assembled target business model.
[0116] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the data aggregation method provided by each of the above methods. The method includes: in response to a change in any physical model in the data warehouse, mapping the physical model and one or more physical models in the data warehouse that have the same primary key field as the physical model into a target business model including the primary key field; generating a data processing event identifier for the target business model according to the primary key field; querying whether there is an event identifier identical to the data processing event identifier; if there is no event identifier identical to the data processing event identifier, creating a corresponding business model aggregation event; based on the business model aggregation event, assembling the target business model with the primary key field as the primary key according to multiple physical models in the data warehouse having the primary key field; and performing a data aggregation operation based on the assembled target business model.
[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for data aggregation, characterized in that, Including: In response to a change in any physical model in the data warehouse, mapping the physical model, and one or more physical models in the data warehouse that have the same primary key field as the physical model, to a target business model that includes the primary key field; wherein, the physical model is a representation of the actual data structure in the data warehouse; Generating a data processing event identifier for the target business model according to the primary key field; Querying whether there is an event identifier that is the same as the data processing event identifier; If there is no event identifier that is the same as the data processing event identifier, creating a corresponding business model aggregation event; Based on the business model aggregation event, assembling the target business model with the primary key field as the primary key according to multiple physical models in the data warehouse that have the primary key field; Performing a data aggregation operation based on the assembled target business model.
2. The method for data aggregation according to claim 1, wherein The method further includes: In response to a change in any physical model in the data warehouse, writing the data of the changed physical model to a distributed cache.
3. The method for data aggregation according to claim 2, wherein The writing the data of the changed physical model to the distributed cache includes: According to the primary key field of the changed physical model, and in combination with a preset cache duration, writing the data of the changed physical model to the distributed cache.
4. The method for data aggregation according to claim 3, wherein The method further includes: In response to a change in any physical model in the data warehouse, writing the data of the changed physical model to a preset analytical database.
5. The method for data aggregation according to claim 4, wherein The assembling the target business model with the primary key field as the primary key according to multiple physical models in the data warehouse that have the primary key field includes: Obtaining the data of the multiple physical models from the distributed cache; Assembling the target business model with the primary key field as the primary key according to the data of the multiple physical models.
6. The method for data aggregation according to claim 5, wherein The method further includes: In response to the distributed cache missing some or all of the data of the multiple physical models; Obtaining the missing some or all of the data of the multiple physical models from the analytical database.
7. The method for data aggregation according to claim 6, wherein Before assembling the target business model with the primary key field as the primary key according to multiple physical models in the data warehouse that have the primary key field, the method further includes: Validating the data of the multiple physical models; In response to there being data that fails to pass the validation among the data of the multiple physical models, after waiting for a second preset duration, re-executing the step of obtaining the data of the multiple physical models from the distributed cache and subsequent steps until the obtained data of the multiple physical models passes the validation, or until a preset retry count is reached.
8. The method for data aggregation according to claim 1, wherein Before assembling the target business model with the primary key field as the primary key according to multiple physical models in the data warehouse that have the primary key field, the method further includes: Waiting for a first preset duration.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data aggregation method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the data aggregation method according to any one of claims 1 to 8.