Data processing method based on data warehouse and related device
By adopting the integrated batch flow method in the data warehouse, using the batch layer to process historical data and the speed layer to process real-time data, the problem of insufficient real-time data processing capabilities of the data warehouse is solved, and efficient data processing and rapid response are achieved.
Patent Information
- Application Number
- CN202510009104.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, data warehouses have insufficient processing capabilities for real-time data, making it difficult to meet application scenarios in the financial field that require faster response.
The batch-stream integration method is adopted to process historical data through the batch processing layer and process real-time data through the speed layer, thereby improving the real-timeness of data processing.
It realizes efficient processing of historical data and real-time data, improves the real-timeness of data processing, and can meet the needs of rapid response in the financial field.
Smart Images

Figure CN119938799A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular, to a data processing method and related devices based on a data warehouse. Background Art
[0002] In the financial field, a data warehouse refers to a subject-oriented, integrated, relatively stable data set that reflects historical changes and is used to support the decision-making analysis and processing of financial institutions. Data warehouses have the following four basic characteristics: subject-oriented: data is organized by business subject; integrated: data sources are diverse but have unified standards; stable: data is generally only added and not changed; time-varying: historical data is stored to facilitate trend analysis.
[0003] In related technologies, data processing for data warehouses is usually based on batch processing, which is a data processing method that involves processing data at one time after accumulating a certain amount of data. Batch processing mainly processes historical data corresponding to the data warehouse, but the processing capacity for real-time data is insufficient, making it difficult to meet application scenarios in the financial field such as risk control scenarios that require faster response requirements. Summary of the invention
[0004] The present application at least provides a data processing method and related device based on a data warehouse, which processes historical data based on a batch processing layer and processes real-time data based on a speed layer, that is, batch-stream integration is used to process both historical data and real-time data, thereby improving the real-time performance of data processing.
[0005] In a first aspect, the present application provides a data processing method based on a data warehouse, wherein the architecture layer of the data warehouse includes a data collection layer, a batch processing layer, a speed layer, a service layer and an application layer, and the method includes:
[0006] Through the data collection layer, historical data and real-time data to be processed are collected from the data source;
[0007] Through the batch processing layer, the historical data is processed to obtain and store the corresponding historical data processing results; through the speed layer, the real-time data is processed to obtain and store the corresponding real-time data processing results;
[0008] Through the service layer, historical data processing results and real-time data processing results are integrated to obtain and store the corresponding integrated data processing results;
[0009] When a query request is received, the application layer performs a data query in the integrated data processing results to obtain the query result.
[0010] In a second aspect, the present application further provides a data processing device based on a data warehouse, wherein the architecture layer of the data warehouse includes a data collection layer, a batch processing layer, a speed layer, a service layer and an application layer, and the device includes:
[0011] The collection module is used to collect historical data and real-time data to be processed from the data source through the data collection layer;
[0012] The processing module is used to process the historical data through the batch processing layer, obtain and store the corresponding historical data processing results; process the real-time data through the speed layer, obtain and store the corresponding real-time data processing results;
[0013] The integration module is used to integrate the historical data processing results and the real-time data processing results through the service layer, and obtain and store the corresponding integrated data processing results;
[0014] The query module is used to perform data query in the integrated data processing results through the application layer to obtain query results when a query request is received.
[0015] In a third aspect, the present application also provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the data processing method based on the data warehouse provided in the present application is executed.
[0016] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the data processing method based on the data warehouse provided in the present application is executed.
[0017] In a fifth aspect, the present application also provides a computer program product, including a computer program, which, when executed by a processor, executes the data processing method based on the data warehouse provided by the present application.
[0018] In summary, the present application provides a data processing method and related devices based on a data warehouse. The architecture layer of the data warehouse includes a data collection layer, a batch processing layer, a speed layer, a service layer and an application layer. The method includes: through the data collection layer, historical data and real-time data to be processed are collected from the data source; through the batch processing layer, historical data are processed to obtain and store the corresponding historical data processing results; through the speed layer, real-time data are processed to obtain and store the corresponding real-time data processing results; through the service layer, historical data processing results and real-time data processing results are integrated to obtain and store the corresponding integrated data processing results; when a query request is received, data query is performed in the integrated data processing results through the application layer to obtain the query results. Through the above method, historical data is processed based on the batch processing layer and real-time data is processed based on the speed layer, that is, batch and stream integration is used to process both historical data and real-time data, thereby improving the real-time performance of data processing.
[0019] Other advantages of the present application will be explained in more detail in conjunction with the following description and drawings.
[0020] It should be understood that the above description is only an overview of the technical solution of the present application, so that the technical means of the present application can be generally understood and then implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are specifically described below by way of example. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments that conform to the present application and are used together with the specification to illustrate the technical solutions of the present application. It should be understood that the drawings only illustrate certain embodiments of the present application and should not be regarded as limiting the scope of protection. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work. Moreover, the same reference numerals are used to represent the same components throughout the drawings. In the drawings:
[0022] Figure 1 A method flow chart of a data processing method based on a data warehouse provided in an embodiment of the present application;
[0023] Figure 2 A schematic diagram of a data processing device based on a data warehouse provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0025] In the description of the embodiments of the present application, it should be understood that terms such as "including" or "having" are intended to indicate the presence of disclosed features, numbers, steps, behaviors, components, parts, or a combination thereof in the present specification, and do not exclude the possibility of the presence of one or more other features, numbers, steps, behaviors, components, parts, or a combination thereof.
[0026] Unless otherwise specified, “ / ” means or. For example, A / B can mean A or B. The “and / or” in this article is merely a way to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0027] The terms "first", "second", etc. are used only to distinguish the same or similar technical features for the convenience of description, and should not be understood as indicating or implying the relative importance or quantity of these technical features. Thus, the features defined by "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the term "plurality" means two or more than two.
[0028] In related technologies, data processing for data warehouses is usually based on batch processing, which is a data processing method that involves processing data at one time after accumulating a certain amount of data. Batch processing mainly processes historical data corresponding to the data warehouse, but the processing capacity for real-time data is insufficient, making it difficult to meet application scenarios in the financial field such as risk control scenarios that require faster response requirements.
[0029] In view of this, the present application provides a data processing method and related devices based on a data warehouse, which processes historical data based on a batch processing layer and processes real-time data based on a speed layer, that is, batch and stream integration is used to process both historical data and real-time data, thereby improving the real-time performance of data processing.
[0030] The data processing method based on the data warehouse provided in the embodiment of the present application can be implemented by a computer device, which can be a terminal device or a server, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminal devices include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.
[0031] The data processing method based on the data warehouse provided by the present application is described below through a method embodiment. Figure 1 As shown, Figure 1 A method flow chart of a data processing method based on a data warehouse provided in an embodiment of the present application, wherein the architecture layer of the data warehouse includes a data collection layer, a batch processing layer, a speed layer, a service layer, and an application layer, and the aforementioned computer device may be a server, and the method includes:
[0032] S101. Collect historical data and real-time data to be processed from data sources through the data collection layer.
[0033] Historical data refers to data that has already occurred and been recorded. Real-time data refers to data that is captured immediately at the current moment.
[0034] The data collection layer is used to collect corresponding historical data and real-time data from the data source. During the collection process, the integrity and consistency of the data need to be ensured.
[0035] In practical applications, the data collection layer can use the open source tool Apache Sqoop to achieve synchronous collection of database data, use the distributed, reliable and available service Apache Flume to achieve log data collection, and use the Canal component to capture incremental data.
[0036] S102. Process the historical data through the batch processing layer to obtain and store the corresponding historical data processing results; process the real-time data through the speed layer to obtain and store the corresponding real-time data processing results.
[0037] Specifically, different from the prior art which only performs batch processing, in this embodiment, historical data is processed through a batch processing layer to obtain and store corresponding historical data processing results, which may include basic data views and pre-calculated results. Real-time data is also processed through a speed layer to obtain and store corresponding real-time data processing results, which may include real-time views corresponding to the real-time data.
[0038] In a possible implementation, in S102, the historical data is processed by a batch processing layer to obtain and store corresponding historical data processing results, including:
[0039] Through the distributed computing system in the batch processing layer, the historical data is processed offline to obtain the corresponding historical data processing results;
[0040] The historical data processing results are stored through the storage tools in the batch layer.
[0041] In practical applications, the distributed computing system may be Apache Spark, that is, Apache Spark may be used to perform offline computing and processing on historical data; the storage tool may be Apache Hadoop, that is, Apache Hadoop may be used to store the processing results of historical data.
[0042] In addition, to facilitate subsequent service layer queries, the batch processing layer can also include Apache Hive to provide SQL capabilities, map structured historical data processing results into a database table, and provide SQL-like query functions.
[0043] In a possible implementation, in S102, the real-time data is processed by the speed layer to obtain and store the corresponding real-time data processing result, including:
[0044] The distributed stream processing platform in the speed layer pre-processes the real-time data to obtain the to-be-processed message queue corresponding to the real-time data.
[0045] Through the distributed stream processing framework in the speed layer, real-time computing and processing are performed on the message queue to be processed to obtain the corresponding real-time data processing results;
[0046] The real-time data processing results are stored through the key-value database in the speed layer.
[0047] In practical applications, the distributed stream processing platform can be Apache Kafka, that is, Apache Kafka can be used to preprocess real-time data to obtain a queue of pending messages corresponding to the real-time data; the distributed stream processing framework can be Apache Flink, that is, Apache Flink can be used to perform real-time calculations on the queue of pending messages; the key-value database can be Redis, which stores the real-time data processing results.
[0048] S103. Integrate the historical data processing results and the real-time data processing results through the service layer to obtain and store the corresponding integrated data processing results.
[0049] Through the service layer, historical data processing results and real-time data processing results can be integrated so as to provide unified query services to the outside world.
[0050] In practical applications, the service layer may include Apache Hbase, which can store historical data processing results for integration with real-time data processing results at different times.
[0051] In addition, to facilitate subsequent queries at the application layer, the service layer may include Elasticsearch and GraphQL, where Elasticsearch is used to provide query functions and GraphQL is used to provide API services.
[0052] S104: When a query request is received, a data query is performed in the integrated data processing result through the application layer to obtain a query result.
[0053] When a query request is received, based on determining the corresponding integrated data processing result through the service layer, a data query can be performed in the integrated data processing result through the application layer to obtain the query result.
[0054] In a possible implementation, the architecture layer of the data warehouse may further include a scheduling and monitoring layer, and the method further includes:
[0055] For pending jobs, the scheduling and monitoring layer manages job scheduling and obtains the job running status corresponding to the pending jobs in real time so as to display it when the job running status is abnormal.
[0056] In practical applications, the scheduling and monitoring layer may include Apache Airflow, through which job scheduling is managed. At the same time, the scheduling and monitoring layer may include Prometheus and Grafana, through which the job running status corresponding to the pending jobs is obtained in real time for monitoring through Prometheus, and the Grafana visualization tool is used to display the abnormal job running status when it occurs.
[0057] In a possible implementation, the data layer of the data warehouse may include an original data layer corresponding to the data source, a detailed data layer corresponding to the data collection layer, a summary data layer corresponding to the service layer, and an application data layer corresponding to the application layer.
[0058] Specifically, the original data layer corresponding to the data source can maintain the original appearance of the data of the data source, fully record historical changes, and is the "true source" of the data; it can also support data backtracking and problem tracking to meet audit and compliance requirements; and it can provide atomic data for the upper layer to ensure that the data is not distorted.
[0059] The detailed data layer corresponding to the data collection layer can clean, convert and standardize the original data obtained by the original data layer; it can also realize the integration and unification of cross-source data to solve the problem of inconsistent data caliber; and it can establish unified data quality control standards to provide reliable data for the upper layer.
[0060] The summary data layer corresponding to the service layer can perform data aggregation and light processing; it can also provide cross-dimensional statistical analysis capabilities and support fast query of common indicators; and it can reduce real-time computing pressure and improve query efficiency through pre-calculation.
[0061] The application data layer corresponding to the application layer can provide personalized data services for specific business scenarios; it can also support flexible data presentation to meet the needs of different users; and it can realize data productization and provide standardized API interfaces.
[0062] In a possible implementation, the data layer of the data warehouse also includes a metadata layer. Metadata is data that describes the data and records information such as the definition, structure, lineage, quality rules, and permission strategies of the data.
[0063] The metadata layer corresponds to the original data layer, detailed data layer, summary data layer and application data layer respectively. The metadata layer is used to perform data governance on the original data layer, detailed data layer, summary data layer and application data layer, so as to achieve full life cycle tracking of data from collection, cleaning, conversion to application.
[0064] In a possible implementation, for application scenarios in the financial field, the data warehouse includes at least one of a customer domain, a product domain, a transaction domain, a risk control domain, a performance domain, a market domain, and a regulatory domain.
[0065] The customer domain can include core data such as basic customer information, customer portraits, behavioral characteristics, etc.; build a complete customer insight system through a 360-degree customer view; focus on customer risk tolerance assessment, investment preference analysis and anti-money laundering risk control.
[0066] The product domain can cover the entire life cycle data of financial products, including product design and review, pricing, product issuance and establishment, fundraising, and survival period management; build a product risk rating model and a dynamic monitoring indicator system; achieve product penetration management and track changes in underlying asset risks.
[0067] The transaction domain can record all transaction-related data, including subscription, redemption, conversion and other operation records; establish a transaction risk early warning mechanism to monitor abnormal transaction behavior in real time; support transaction limit management, transaction compliance checks and anti-fraud analysis.
[0068] The risk control domain can integrate multi-dimensional risk control data such as market risk, credit risk, operational risk, and liquidity risk; it can build a risk indicator system to achieve quantitative risk assessment and early warning; and it can provide stress testing and risk exposure analysis capabilities.
[0069] The performance domain can track core performance indicators such as product yield, volatility, and Sharpe ratio; implement multi-dimensional performance attribution analysis and performance evaluation; and support benchmarking analysis and market performance evaluation with similar products.
[0070] The market domain can include various financial market data, such as stock, bond, foreign exchange and other market conditions; establish a market indicator monitoring system to track important market signals; and provide market trend analysis and macroeconomic impact assessment.
[0071] The regulatory domain can integrate various types of regulatory reporting data and compliance management requirements; realize automatic calculation and real-time monitoring of regulatory indicators; and support dynamic regulatory policy adaptation and compliance checks.
[0072] It should be noted that in the embodiments of the present application, Apache Sqoop, Apache Flume, Canal, Apache Spark, Apache Hadoop, Apache Hive, Apache Kafka, Apache Flink, Redis, Apache Hbase, Elasticsearch, GraphQL, Apache Airflow, Prometheus, and Grafana are all open source technology stacks, which can avoid the high cost of purchasing commercial software and thus reduce costs.
[0073] It can be seen that the present application provides a data processing method based on a data warehouse, and the architecture layer of the data warehouse includes a data collection layer, a batch processing layer, a speed layer, a service layer and an application layer. The method includes: through the data collection layer, the historical data and real-time data to be processed are collected from the data source; through the batch processing layer, the historical data is processed to obtain and store the corresponding historical data processing results; through the speed layer, the real-time data is processed to obtain and store the corresponding real-time data processing results; through the service layer, the historical data processing results and the real-time data processing results are integrated to obtain and store the corresponding integrated data processing results; when a query request is received, the data query is performed in the integrated data processing results through the application layer to obtain the query results. Through the above method, historical data is processed based on the batch processing layer and real-time data is processed based on the speed layer, that is, batch and stream integration is used to process both historical data and real-time data, thereby improving the real-time performance of data processing.
[0074] In the description of this specification, the description with reference to the terms "some possible embodiments", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application, and the above terms do not necessarily represent the same embodiment or example. Moreover, the described specific features, structures, materials or characteristics may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0075] About the method flow chart of the present application embodiment, some operations are described as different steps performed in a certain order. Such flow chart belongs to illustrative and non-restrictive. Some steps described in this article can be grouped together and performed in a single operation, or some steps can be divided into multiple sub-steps and can be performed in an order different from that shown in this article. Each step shown in the flow chart can be realized in any way by any circuit structure and / or tangible mechanism (for example, by software, hardware (for example, the logical function realized by processor or chip) etc. running on computer equipment and / or any combination thereof).
[0076] Those skilled in the art will appreciate that, in the method described in the above specific implementation, the writing order of each step does not mean a strict execution order, and the specific execution order of each step should be determined by its function and possible internal logic.
[0077] Based on the aforementioned Figure 1 The data processing device based on the data warehouse provided by the present application is described below through a device embodiment. Figure 2 A schematic diagram of a data processing device based on a data warehouse provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the data processing device 200 based on the data warehouse includes:
[0078] The collection module 201 is used to collect historical data and real-time data to be processed from the data source through the data collection layer;
[0079] The processing module 202 is used to process the historical data through the batch processing layer, obtain and store the corresponding historical data processing results; process the real-time data through the speed layer, obtain and store the corresponding real-time data processing results;
[0080] The integration module 203 is used to integrate the historical data processing results and the real-time data processing results through the service layer, and obtain and store the corresponding integrated data processing results;
[0081] The query module 204 is used to perform data query in the integrated data processing result through the application layer to obtain the query result when receiving the query request.
[0082] In a possible implementation, the processing module 202 is configured to:
[0083] Through the distributed computing system in the batch processing layer, the historical data is processed offline to obtain the corresponding historical data processing results;
[0084] The historical data processing results are stored through the storage tools in the batch layer.
[0085] In a possible implementation, the processing module 202 is configured to:
[0086] The distributed stream processing platform in the speed layer pre-processes the real-time data to obtain the to-be-processed message queue corresponding to the real-time data.
[0087] Through the distributed stream processing framework in the speed layer, real-time computing and processing are performed on the message queue to be processed to obtain the corresponding real-time data processing results;
[0088] The real-time data processing results are stored through the key-value database in the speed layer.
[0089] In a possible implementation, the data warehouse-based data processing device 200 further includes a management unit, which is used to:
[0090] The architecture layer of the data warehouse also includes a scheduling and monitoring layer. For pending jobs, the scheduling and monitoring layer manages job scheduling and obtains the job running status corresponding to the pending jobs in real time so that it can be displayed when the job running status is abnormal.
[0091] In a possible implementation, the data layer of the data warehouse includes an original data layer corresponding to the data source, a detailed data layer corresponding to the data collection layer, a summary data layer corresponding to the service layer, and an application data layer corresponding to the application layer.
[0092] In one possible implementation, the data layer of the data warehouse also includes a metadata layer, which corresponds to the original data layer, detailed data layer, summary data layer and application data layer respectively. The metadata layer is used to perform data governance on the original data layer, detailed data layer, summary data layer and application data layer.
[0093] In one possible implementation, the data warehouse includes at least one of a customer domain, a product domain, a transaction domain, a risk control domain, a performance domain, a market domain, and a regulatory domain.
[0094] It should be noted that the device in the implementation mode of the present application can implement each process of the implementation mode of the aforementioned method and achieve the same effects and functions, which will not be repeated here.
[0095] The embodiment of the present application also provides an electronic device, including: a processor, a memory and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus. When the machine-readable instructions are executed by the processor, the following processing is performed:
[0096] The architecture layers of a data warehouse include data collection layer, batch processing layer, speed layer, service layer, and application layer;
[0097] Through the data collection layer, historical data and real-time data to be processed are collected from the data source;
[0098] Through the batch processing layer, the historical data is processed to obtain and store the corresponding historical data processing results; through the speed layer, the real-time data is processed to obtain and store the corresponding real-time data processing results;
[0099] Through the service layer, historical data processing results and real-time data processing results are integrated to obtain and store the corresponding integrated data processing results;
[0100] When a query request is received, the application layer performs a data query in the integrated data processing result to obtain a query result.
[0101] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the data processing method based on the data warehouse described in the above method embodiment are executed. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0102] An embodiment of the present application also provides a computer program product, including a computer program. The computer program product carries a program code. The instructions included in the program code can be used to execute the steps of the data processing method based on the data warehouse described in the above method embodiment. For details, please refer to the above method embodiment, which will not be repeated here.
[0103] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is implemented as a computer storage medium. In another optional embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0104] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and computer-readable storage medium embodiments, since they are basically similar to the method embodiments, their descriptions are simplified, and the relevant parts can be referred to the partial description of the method embodiments.
[0105] The apparatus, equipment and computer-readable storage medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the apparatus, equipment and computer-readable storage medium also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the apparatus, equipment and computer-readable storage medium will not be repeated here.
[0106] Although the spirit and principle of the present application have been described above with reference to several specific embodiments, it should be understood that the present application is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined. The present application is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.
Claims
1. A data processing method based on a data warehouse, wherein the architecture layer of the data warehouse includes a data collection layer, a batch processing layer, a speed layer, a service layer and an application layer, and the method comprises: Through the data collection layer, historical data and real-time data to be processed are collected from the data source; The historical data is processed through the batch processing layer to obtain and store the corresponding historical data processing results; the real-time data is processed through the speed layer to obtain and store the corresponding real-time data processing results; By means of the service layer, the historical data processing result and the real-time data processing result are integrated to obtain and store the corresponding integrated data processing result; When a query request is received, the application layer performs a data query in the integrated data processing result to obtain a query result.
2. The method according to claim 1, characterized in that The historical data is processed through the batch processing layer to obtain and store corresponding historical data processing results, including: By using the distributed computing system in the batch processing layer, the historical data is processed offline to obtain corresponding historical data processing results; The historical data processing results are stored through the storage tools in the batch processing layer.
3. The method according to claim 2, characterized in that The processing of the real-time data by the speed layer to obtain and store corresponding real-time data processing results includes: The real-time data is preprocessed by the distributed stream processing platform in the speed layer to obtain a to-be-processed message queue corresponding to the real-time data; Through the distributed stream processing framework in the speed layer, the message queue to be processed is processed in real time to obtain corresponding real-time data processing results; The real-time data processing results are stored via the key-value pair database in the speed layer.
4. The method according to claim 1, characterized in that: The architecture layer of the data warehouse also includes a scheduling and monitoring layer, and the method further includes: For pending jobs, the scheduling and monitoring layer manages job scheduling, and obtains the job running status corresponding to the pending jobs in real time so as to display it when an abnormality occurs in the job running status.
5. The method according to claim 1, characterized in that The data layer of the data warehouse includes an original data layer corresponding to the data source, a detailed data layer corresponding to the data collection layer, a summary data layer corresponding to the service layer, and an application data layer corresponding to the application layer.
6. The method according to claim 5, characterized in that The data layer of the data warehouse also includes a metadata layer, which corresponds to the original data layer, the detailed data layer, the summary data layer and the application data layer respectively, and the metadata layer is used to perform data governance on the original data layer, the detailed data layer, the summary data layer and the application data layer.
7. The method according to claim 1, characterized in that The data warehouse includes at least one of a customer domain, a product domain, a transaction domain, a risk control domain, a performance domain, a market domain and a regulatory domain.
8. A data processing device based on a data warehouse, wherein the architecture layer of the data warehouse includes a data collection layer, a batch processing layer, a speed layer, a service layer and an application layer, and the device includes: A collection module, used to collect historical data and real-time data to be processed from the data source through the data collection layer; A processing module, used to process the historical data through the batch processing layer, obtain and store corresponding historical data processing results; and process the real-time data through the speed layer, obtain and store corresponding real-time data processing results; An integration module, used to integrate the historical data processing results and the real-time data processing results through the service layer, and obtain and store corresponding integrated data processing results; The query module is used to perform data query in the integrated data processing result through the application layer to obtain the query result when receiving the query request.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the data processing method based on the data warehouse as described in any one of claims 1 to 7 is performed.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data processing method based on a data warehouse as claimed in any one of claims 1 to 7 is executed.
11. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, executes the data processing method based on a data warehouse as claimed in any one of claims 1 to 7.
Citation Information
Cited By
Intelligent power transaction data processing method, system, medium and equipment
CN121166711A
Industrial big data-oriented batch flow integrated data cleaning system
CN121455941A