Real-time warehouse counting index calculation system, method, device and equipment and storage medium
By adopting a real-time data warehouse index calculation system with a microservice framework and a distributed system architecture in the data warehouse, the problem of difficulty in realizing non-invasiveness of the business system in the existing technology is solved, and efficient and stable real-time data processing and indicator calculation are achieved.
Patent Information
- Application Number
- CN202510155542.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-06
AI Technical Summary
Existing data warehouses are difficult to achieve non-invasiveness of business systems, poor real-time performance, high business coupling, high development and maintenance costs, and cannot achieve business-based configuration and rapid release and deployment of indicators.
It adopts a microservice framework and distributed system architecture to provide a real-time digital warehouse index calculation system, including a data collection service module, a digital warehouse storage service module, a data calculation service module and a management platform. It stores real-time data calculation indicators through an in-memory database, and supports dynamic horizontal expansion of massive data indicator processing.
It realizes the metric calculation of a real-time event data in hundreds of milliseconds, decouples the business system and real-time warehousing indicator calculation, improves data processing efficiency and system stability, and reduces enterprise operation costs.
Smart Images

Figure CN120104700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a real-time data warehouse index calculation system, method, device, equipment and storage medium. Background Art
[0002] With the advent of the Internet era, the amount of data has exploded, and data warehouses have begun to use big data tools to replace traditional tools in classic data warehouses. This architecture belongs to offline big data architecture. In offline big data architecture, data sources are imported into offline data warehouses in an offline way. Downstream applications choose to read DM directly or add a layer of data services, such as MySQL or Redis, according to business needs. Typical data warehouse storage is HDFS / Hive, and ETL can be MapReduce scripts or HiveSQL.
[0003] In the related technologies, the traditional data warehouse full import: by reading the source table, batch import into the data warehouse table through code or tools; incremental synchronization uses triggers or timestamp tags based on business tables for incremental import, which has obvious disadvantages, poor real-time performance, high business coupling, and poor development and maintenance. In terms of distributed computing, the big data real-time processing architecture based on Flink has become the mainstream architecture for real-time computing, but there are many technical stacks to master, the learning cost is high, the deployment and implementation are difficult, and developers need to perform secondary development, which cannot achieve the business configuration of indicators and rapid release and deployment. In terms of distributed storage, big data storage is mainly based on HDFS Hbase storage, which can achieve efficient write performance, but cannot directly support sql and cannot meet business needs. Supporting sql can add secondary indexes, but it undoubtedly increases the complexity of adding components, and the query efficiency cannot achieve real-time response. It is difficult for existing data warehouses to be non-intrusive to business systems. Summary of the invention
[0004] In view of this, the present invention provides a real-time data warehouse index calculation system, method, device, equipment and storage medium to solve the problem that existing data warehouses are difficult to achieve non-intrusion into business systems.
[0005] In a first aspect, the present invention provides a real-time data warehouse index calculation system, the system includes a data collection service module, a data warehouse storage service module, a data calculation service module and a management platform:
[0006] The data collection service module is used to collect data and perform full data migration or database synchronization by monitoring the cluster;
[0007] Data warehouse service module, used to manage and store data;
[0008] Data computing service module, used to calculate data and generate decision support information;
[0009] Management platform for real-time data monitoring and alarm.
[0010] In the present invention, the microservice framework is used as the basic framework of the data service; a distributed system architecture is adopted to support horizontal expansion based on data nodes of any scale, thereby improving the computing and processing capabilities of the overall cluster. Through the real-time data warehouse indicator calculation system, hundreds of millisecond-level indicator calculations will be performed on a real-time event data. The data service adopts a microservice distributed service structure, and the real-time data calculation indicators are stored internally through an in-memory database. The reading and writing of the in-memory database ensures the rapid update and calculation of real-time data. The dynamic horizontal expansion of massive data indicator processing is supported through a distributed service structure.
[0011] In an optional implementation, the data computing service module includes:
[0012] A data cleaning and preprocessing unit, used for cleaning and preprocessing the data to obtain processed data;
[0013] The feature engineering unit is used to screen relevant variables based on business needs, generate features corresponding to the processed data, and use the features to perform calculations to obtain calculation results;
[0014] The result interpretation and visualization unit is used to visualize the calculation results.
[0015] In this approach, the data computing service module is one of the core components in the data analysis process. It plays a key role in processing raw data, performing complex calculations, and generating information that can be used for further analysis or decision support.
[0016] In an optional implementation, the management platform is used to define a data dictionary, and the data dictionary is used to characterize metadata of the processed data;
[0017] The management platform is also used to define statistical models and entity models, and use the statistical models and entity models to define the processing fields to be calculated corresponding to the configuration data, and use the processing fields to be calculated to calculate the operation results.
[0018] In this way, the management platform can realize data integration and consolidation, real-time data processing, real-time monitoring and alarm, and resource management and optimization.
[0019] In a second aspect, the present invention provides a real-time data warehouse index calculation method, which is applied to a real-time data warehouse index calculation system as described in any one of the first aspects, wherein the system includes a data collection service module, a data warehouse storage service module, a data calculation service module, and a management platform, and the method includes:
[0020] Use the data collection service module and data warehouse service module to configure data collection service and data warehouse service;
[0021] Use the management platform to define the original data to obtain fields, and define statistical models and entity models;
[0022] Execute statistical models and entity models, configure fields, and convert them into model-level data;
[0023] Perform index calculation on the model layer data to obtain updated index values, and update the index values to the target memory.
[0024] In the present invention, it is possible to perform hundreds of millisecond-level index calculations on a piece of real-time event data. The data service adopts a microservice distributed service structure, and stores real-time data calculation indicators internally through a memory database. It overcomes many drawbacks caused by code intrusion into the business system in the prior art, realizes the decoupling of the business system and the real-time data warehouse index calculation, improves data processing efficiency and system stability, and reduces enterprise operating costs.
[0025] In an optional implementation, executing the statistical model and the entity model, configuring the fields, and converting them into model layer data includes:
[0026] Add fields under statistical model and entity model;
[0027] Use the data warehouse service module to suspend the statistical model and the entity model, and then execute the statistical model and the entity model;
[0028] Create data tables corresponding to the statistical model and entity model in distributed storage, start indicator configuration check, configure the fields, and convert them into model layer data.
[0029] In this method, the configured indicator items are published and enabled, and the configuration is synchronized to the data service node. Each statistical model and entity model will be created as a separate data table in the distributed storage. The correctness of the configured model fields is ensured through indicator configuration checking.
[0030] In an optional implementation, performing index calculation on the model layer data to obtain an updated index value, and updating the index value to the target memory includes:
[0031] In the data computing service module, matching is performed on the associated calculable fields of the statistical model or entity model;
[0032] Get the main dimension of the data and read the indicator value in the target memory based on the main dimension;
[0033] Calculate the index value corresponding to the model layer data, and perform index calculation based on the index value in the target memory to obtain the updated index value;
[0034] Update the indicator value to the target distributed memory.
[0035] In this method, hundreds of millisecond-level index calculations are performed on a piece of real-time event data. The data service adopts a microservice distributed service structure, and the real-time data calculation index is stored in the memory database internally. The reading and writing of the memory database ensures the rapid update and calculation of real-time data.
[0036] In a third aspect, the present invention provides a real-time data warehouse index calculation device, the device comprising:
[0037] The data configuration module is used to configure the data collection service and the data warehouse service using the data collection service module and the data warehouse storage service module;
[0038] The model definition module is used to define the original data to obtain fields using the management platform, and to define the statistical model and entity model;
[0039] Data calculation module, used to execute statistical models and entity models, configure fields, and convert them into model layer data;
[0040] The data update module is used to perform index calculation on the model layer data, obtain updated index values, and update the index values to the target memory.
[0041] In a fourth aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the real-time data warehouse indicator calculation method of the above-mentioned second aspect or any corresponding embodiment thereof by executing the computer instructions.
[0042] In a fifth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the real-time data warehouse indicator calculation method of the above-mentioned second aspect or any corresponding embodiment thereof.
[0043] In a sixth aspect, the present invention provides a computer program product, comprising computer instructions, which are used to enable a computer to execute the real-time data warehouse indicator calculation method of the above-mentioned second aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 It is a structural diagram of a real-time data warehouse index calculation system according to an embodiment of the present invention.
[0046] Figure 2 It is a schematic diagram of the architecture of a real-time data warehouse indicator calculation business system that does not invade the business system according to an embodiment of the present invention.
[0047] Figure 3 It is a flowchart of a real-time data warehouse index calculation method according to an embodiment of the present invention.
[0048] Figure 4 It is a schematic diagram of an overall data flow of data collection and processing according to an embodiment of the present invention.
[0049] Figure 5 4 is a structural block diagram of a real-time data warehouse index calculation device according to an embodiment of the present invention.
[0050] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0052] In the related technologies, the traditional data warehouse full import: by reading the source table, batch import into the data warehouse table through code or tools; incremental synchronization uses triggers or timestamp tags based on business tables for incremental import, which has obvious disadvantages, poor real-time performance, high business coupling, and poor development and maintenance. In terms of distributed computing, the big data real-time processing architecture based on Flink has become the mainstream architecture for real-time computing, but there are many technical stacks to master, the learning cost is high, the deployment and implementation are difficult, and developers need to perform secondary development, which cannot achieve the business configuration of indicators and rapid release and deployment. In terms of distributed storage, big data storage is mainly based on HDFS Hbase storage, which can achieve efficient write performance, but cannot directly support sql and cannot meet business needs. Supporting sql can add secondary indexes, but it undoubtedly increases the complexity of adding components, and the query efficiency cannot achieve real-time response. It is difficult for existing data warehouses to be non-intrusive to business systems.
[0053] To solve the above problems, a real-time data warehouse index calculation system is provided in an embodiment of the present invention, which is suitable for use scenarios of real-time data warehouses, data processing fields, and real-time processing model data. The present invention provides a real-time data warehouse index calculation system, and a microservice framework is used as the basic framework of data services; a distributed system architecture is adopted to support horizontal expansion based on data nodes of any scale, thereby improving the computing and processing capabilities of the overall cluster. Through the real-time data warehouse index calculation system, hundreds of millisecond-level index calculations will be performed on a piece of real-time event data. The data service adopts a microservice distributed service structure, and the real-time data calculation indicators are stored internally through a memory database. The reading and writing of the memory database ensures the rapid update and calculation of real-time data. The dynamic horizontal expansion of massive data index processing is supported through a distributed service structure.
[0054] According to an embodiment of the present invention, a real-time data warehouse index calculation system embodiment is provided. In this embodiment, a real-time data warehouse index calculation system is provided. Figure 1 is a structural diagram of a real-time data warehouse index calculation system according to an embodiment of the present invention. Figure 1 As shown, the system includes a data collection service module 1, a data warehouse service module 2, a data computing service module 3 and a management platform 4: the data collection service module 1 is used to collect data and perform full data migration or database synchronization through the monitoring cluster; the data warehouse service module 2 is used to manage and store data; the data computing service module 3 is used to calculate the data and generate decision support information; the management platform 4 is used to monitor and alarm the data in real time.
[0055] In an optional embodiment, the data computing service module 3 includes: a data cleaning and preprocessing unit, which is used to clean and preprocess the data to obtain processed data; a feature engineering unit, which is used to screen relevant variables based on business needs, generate features corresponding to the processed data, and use the features to perform operations to obtain operation results; a result interpretation and visualization unit, which is used to visualize the operation results.
[0056] In this approach, the data computing service module is one of the core components in the data analysis process. It plays a key role in processing raw data, performing complex calculations, and generating information that can be used for further analysis or decision support.
[0057] In an optional embodiment, the management platform 4 is used to define a data dictionary, which is used to characterize the metadata of the processing data; the management platform 4 is also used to define a statistical model and an entity model, and use the statistical model and the entity model to define the processing fields to be calculated corresponding to the configuration data, and use the processing fields to be calculated to calculate the operation results.
[0058] In this way, the management platform can realize data integration and consolidation, real-time data processing, real-time monitoring and alarm, and resource management and optimization.
[0059] In one implementation scenario, Figure 2 is a schematic diagram of the architecture of a real-time data warehouse index calculation business system that does not invade the business system according to an embodiment of the present invention, such as Figure 2 As shown in the figure, the real-time data warehouse indicator calculation business system that does not invade the business system includes:
[0060] Data collection service module 1: Based on the open source Maxwell, it is mainly responsible for collecting, processing and integrating data from various sources, and realizing full data migration or database synchronization by monitoring the slave database of the MySQL cluster. The primary key and primary dimension of the intermediate model are partitioned and pushed to Kafka storage. The data processing sequence and consistency are ensured by consuming a partition data by a single thread inside the data computing service node; the processed data is stored safely and reliably to prepare for the data warehouse service.
[0061] The data warehouse storage service module 2 is mainly responsible for efficient and secure management and storage of massive data. It supports the storage of structured data and unstructured data (such as log files, pictures and videos). By providing flexible data model definition capabilities and powerful query performance optimization mechanisms, the data warehouse storage service module lays the foundation for subsequent data analysis, calculation and other links.
[0062] The data computing service module 3 is one of the core components in the data analysis process. It plays a key role in processing raw data, performing complex operations, and generating information that can be used for further analysis or decision support. The working steps are as follows:
[0063] 1) Data cleaning and preprocessing: remove invalid or inaccurate data records, fill in missing values, and convert data in different formats into a form suitable for the next step of analysis.
[0064] 2) Feature Engineering: Select relevant variables (features) based on business needs, and create new features through mathematical transformations to improve model performance.
[0065] 3) Result interpretation and visualization: Convert complex calculation results into easy-to-understand charts or reports for quick use by business personnel with non-technical backgrounds.
[0066] Management platform 4 can realize data integration and consolidation, real-time data processing, real-time monitoring and alarm, resource management and optimization. In management platform 4, the data source, data table, and data field of processing are defined, and the data dictionary is defined. The data dictionary is the metadata definition of the processing data; in management platform 4, the statistical model and the entity model are defined as the definition of the processing target table, where the entity model is used as a feature wide table, such as a processing target table composed of multiple processing fields such as customers, commodities, and orders. The statistical model is a target table obtained by aggregation calculation based on behavior, flow, and event data. The statistical model summarizes the detailed data based on certain field dimensions and time granularity to obtain the corresponding indicator value.
[0067] Exemplarily, the statistical model and the entity model can be mapped to the indicator data table to be processed, and the association relationship and dependency relationship of the processing fields can be configured through the association configuration. Define and configure the processing fields to be calculated on the management platform 4. The types of processing fields can be defined as: basic fields, function fields, script fields, statistical fields, aggregate fields and conditional fields according to the processing calculation method. The basic field is a field directly taken from the original association table; the function field is the processing result calculated by the system predefined function. The script field can define a function similar to UDF for data processing, and can support multiple types of data scripts such as Java, Python, and Groovy; the statistical field is a summary calculation based on data, cumulative summation based on specific conditions, and indicator calculations such as the number of times and deduplication count; the aggregation function used by the aggregate field is similar to the UDAF function, and the aggregation function consists of two functions: the process processing function and the final processing function; the conditional field is a calculation processing similar to the IF THEN and WHEN CASE of the SQL statement, and the indicator result is obtained based on the conditional cumulative calculation. The statistical model and entity model configured by the management platform 4 are configured and loaded on the data service end, and the indicator configuration field item is loaded to complete the initialization configuration of the model.
[0068] The real-time data warehouse indicator calculation system provided in this embodiment uses a microservice framework as the basic framework for data services; it adopts a distributed system architecture, supports horizontal expansion based on data nodes of any scale, and improves the computing and processing capabilities of the overall cluster. Through the real-time data warehouse indicator calculation system, hundreds of millisecond-level indicator calculations will be performed on a piece of real-time event data. The data service adopts a microservice distributed service structure, and internally stores real-time data calculation indicators through an in-memory database. The reading and writing of the in-memory database ensures the rapid update and calculation of real-time data. The distributed service structure supports the dynamic horizontal expansion of massive data indicator processing.
[0069] According to an embodiment of the present invention, a real-time data warehouse index calculation method embodiment is provided. In this embodiment, a real-time data warehouse index calculation method is provided, which can be used in the above-mentioned real-time data warehouse index calculation system. Figure 3 is a flow chart of a real-time data warehouse index calculation method according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0070] Step S301, using the data collection service module and the data warehouse service module, configure the data collection service and the data warehouse service.
[0071] Step S302: using the management platform, define the original data to obtain fields, and define the statistical model and entity model.
[0072] Step S303, execute the statistical model and the entity model, configure the fields, and convert them into model layer data.
[0073] In an optional implementation, step S303 includes:
[0074] Step a1, add fields under the statistical model and entity model.
[0075] Step a2: Use the data warehouse service module to suspend the statistical model and the entity model, and execute the statistical model and the entity model.
[0076] Step a3: Create data tables corresponding to the statistical model and entity model in the distributed storage, start the indicator configuration check, configure the fields, and convert them into model layer data.
[0077] In this method, the configured indicator items are published and enabled, and the configuration is synchronized to the data service node. Each statistical model and entity model will be created as a separate data table in the distributed storage. The correctness of the configured model fields is ensured through indicator configuration checking.
[0078] Step S304, performing index calculation on the model layer data to obtain updated index values, and updating the index values to the target memory.
[0079] In an optional implementation, step S304 includes:
[0080] Step b1, in the data computing service module, matching the associated calculable fields of the statistical model or entity model.
[0081] Step b2, obtain the main dimension of the data, and read the indicator value in the target memory based on the main dimension.
[0082] Step b3, calculate the index value corresponding to the model layer data, and perform index calculation in combination with the index value in the target memory to obtain the updated index value.
[0083] Step b4: Update the indicator value to the target distributed memory.
[0084] In this method, hundreds of millisecond-level index calculations are performed on a piece of real-time event data. The data service adopts a microservice distributed service structure, and the real-time data calculation index is stored in the memory database internally. The reading and writing of the memory database ensures the rapid update and calculation of real-time data.
[0085] In one implementation scenario, Figure 4 is a schematic diagram of an overall data flow of data collection and processing according to an embodiment of the present invention, such as Figure 4 As shown in the figure, the overall data flow of data collection and processing includes: the first step is to configure the MySQL business library monitored by the data collection service and the data warehouse service and start all services.
[0086] The second step is to execute the program to transfer the full business database to the data warehouse database. After that, the data warehouse ODS layer and the business database will be synchronized in real time.
[0087] The third step is to define the data fields and metadata information of the original data of the ODS layer on the management platform for use as basic data.
[0088] The fourth step is to define the statistical model and entity model on the management platform, and define the association relationships and associated fields of multiple ODS layer original tables for multi-table summary calculations.
[0089] Step 5: Add definition fields under the current statistical model and entity model. Each field is a derived indicator for processing and calculation. Each field can be understood as a UDF or UDAF function to be processed and calculated. The field type refers to the above definition.
[0090] Step 6: Execute the program of the model. Before execution, the data warehouse synchronization service will suspend the consumption of the data table referenced by the model and then execute the program of historical data. After execution, the real-time consumption calculation of all models will be resumed.
[0091] Step 7: Publish and enable the configured indicator items, and synchronize the configuration to the data service node. The data service node will initialize the configured fields and relationships based on the configured statistical model and entity model. Each statistical model and entity model will be created as a separate data table in the distributed storage. Start the indicator configuration check to ensure the correctness of the configured model fields.
[0092] Step 8: Based on different application scenarios, the synchronous call interface and asynchronous interface of indicator calculation are included. The data in the original ODS table of the data to be processed is converted into JSON structure data and sent to the data platform.
[0093] In the ninth step, the statistical model or entity model to which the calculated indicator data belongs will be determined based on the request parameters in the data service, and loaded into the corresponding model configuration data information. The original JSON data will be processed and calculated based on the current data, and the JSON format data will be converted into the internal calculation data format. And determine the source table to which the sent original data belongs, and match the associated calculable fields in the current statistical model or entity model.
[0094] The tenth step is to obtain the value of the main dimension of the data based on the current data, and read the indicator value recorded in the distributed memory based on the main dimension of the model.
[0095] In the eleventh step, the calculated indicator of the current record is aggregated or compared with the indicator value recorded in the current memory according to the configured indicator calculation method.
[0096] In the twelfth step, the indicator values calculated by the aggregation update are updated into the distributed memory.
[0097] The real-time data warehouse index calculation method provided in this embodiment realizes the index calculation of hundreds of milliseconds for a real-time event data. The data service adopts a microservice distributed service structure, and stores real-time data calculation indicators internally through a memory database. It overcomes many disadvantages of the prior art caused by code intrusion into the business system, realizes the decoupling of the business system and the real-time data warehouse index calculation, improves data processing efficiency and system stability, and reduces enterprise operating costs.
[0098] In this embodiment, a real-time data warehouse index calculation device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated here. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0099] This embodiment provides a real-time data warehouse index calculation device, such as Figure 5 As shown, including:
[0100] The data configuration module 501 is used to configure the data collection service and the data warehouse service using the data collection service module and the data warehouse storage service module. Figure 3 Step S301 of the illustrated embodiment will not be described in detail here.
[0101] The model definition module 502 is used to define the original data to obtain fields using the management platform, and to define the statistical model and entity model. Figure 3 Step S302 of the illustrated embodiment will not be described in detail here.
[0102] The data calculation module 503 is used to execute the statistical model and entity model, configure the fields, and convert them into model layer data. Figure 3 Step S303 of the illustrated embodiment will not be described in detail here.
[0103] The data update module 504 is used to calculate the index of the model layer data, obtain the updated index value, and update the index value to the target memory. Figure 3 Step S304 of the illustrated embodiment will not be described in detail here.
[0104] In some optional implementations, the data calculation module 503 includes:
[0105] The field adding unit is used to add fields under the statistical model and entity model.
[0106] The model execution unit is used to utilize the data warehouse service module to suspend the statistical model and the entity model and execute the statistical model and the entity model.
[0107] The configuration check unit is used to create data tables corresponding to the statistical model and the entity model in the distributed storage, start the indicator configuration check, configure the fields, and convert them into model layer data.
[0108] In some optional implementations, the data updating module 504 includes:
[0109] The field matching unit is used to match the associated calculable fields of the statistical model or entity model in the data computing service module.
[0110] The indicator reading unit is used to obtain the main dimension of the data and read the indicator value in the target memory based on the main dimension.
[0111] The indicator value update calculation unit is used to calculate the indicator value corresponding to the model layer data, and perform indicator calculation in combination with the indicator value in the target memory to obtain the updated indicator value.
[0112] The memory update unit is used to update the indicator value to the target distributed memory.
[0113] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0114] The real-time data warehouse index calculation device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0115] The embodiment of the present invention also provides a computer device having the above Figure 5 The real-time data warehouse index calculation device shown.
[0116] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0117] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0118] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0119] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0120] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0121] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.
[0122] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0123] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0124] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0125] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A real-time data warehouse index calculation system, characterized in that: The system includes a data collection service module, a data warehouse service module, a data calculation service module and a management platform: The data collection service module is used to collect data and perform full migration or database synchronization on the data by monitoring the cluster; The data warehouse service module is used to manage and store the data; The data computing service module is used to perform operations on the data to generate decision support information; The management platform is used to monitor and warn the data in real time.
2. The system according to claim 1, characterized in that The data computing service module includes: A data cleaning and preprocessing unit, used to clean and preprocess the data to obtain processed data; A feature engineering unit, used to screen relevant variables based on business requirements, generate features corresponding to the processed data, and perform calculations using the features to obtain calculation results; The result interpretation and visualization unit is used to visualize the operation result.
3. The system according to claim 1, characterized in that The management platform is used to define a data dictionary, and the data dictionary is used to represent metadata of the processed data; The management platform is also used to define a statistical model and a physical model, and use the statistical model and the physical model to define and configure the processing fields to be calculated corresponding to the data, and use the processing fields to be calculated to calculate the operation results.
4. A real-time data warehouse index calculation method, characterized in that: Applied to the real-time data warehouse index calculation system according to any one of claims 1 to 3, the system includes a data acquisition service module, a data warehouse storage service module, a data calculation service module and a management platform, and the method includes: Utilize the data collection service module and the data warehouse service module to configure data collection service and data warehouse service; Using the management platform, the original data is defined to obtain fields, and a statistical model and an entity model are defined; Execute the statistical model and the entity model, configure the fields, and convert them into model layer data; The model layer data is subjected to index calculation to obtain updated index values, and the index values are updated to the target memory.
5. The method according to claim 4, characterized in that The executing the statistical model and the entity model, configuring the fields, and converting them into model layer data includes: Adding the field under the statistical model and the entity model; Using the data warehouse service module, pausing the statistical model and the entity model, and executing the statistical model and the entity model; Create data tables corresponding to the statistical model and the entity model in distributed storage, start indicator configuration check, configure the fields, and convert them into model layer data.
6. The method according to claim 5, characterized in that The step of performing index calculation on the model layer data to obtain an updated index value, and updating the index value to the target memory includes: In the data computing service module, matching the associated calculable fields of the statistical model or the entity model; Obtaining the main dimension of the data, and reading the indicator value in the target memory based on the main dimension; Calculate the index value corresponding to the model layer data, and perform index calculation in combination with the index value in the target memory to obtain an updated index value; Update the indicator value to the target distributed memory.
7. A real-time data warehouse index calculation device, characterized in that: The device comprises: The data configuration module is used to configure the data collection service and the data warehouse service using the data collection service module and the data warehouse storage service module; The model definition module is used to define the original data to obtain fields using the management platform, and to define the statistical model and entity model; A data calculation module, used to execute the statistical model and the entity model, configure the fields, and convert them into model layer data; The data updating module is used to perform index calculation on the model layer data to obtain updated index values, and update the index values to the target memory.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the real-time data warehouse indicator calculation method according to any one of claims 4 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the real-time data warehouse indicator calculation method described in any one of claims 4 to 6.
10. A computer program product, characterized in that It includes computer instructions, which are used to enable a computer to execute the real-time data warehouse indicator calculation method described in any one of claims 4 to 6.