Data processing system and method based on multi-layer architecture
Through a data processing system based on a multi-layer architecture, combined with ETL automated processing and a high-performance analysis engine, the problems of slow data processing and single data services in existing technologies have been solved, efficient and diversified data services have been achieved, supporting enterprises' real-time analysis and business decision-making, and enhancing their competitiveness.
Patent Information
- Application Number
- CN202510681379.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Existing data processing systems are slow when faced with massive amounts of data and cannot meet real-time analysis needs. In addition, the data service format is single and cannot meet the diverse needs of different business scenarios, resulting in delayed business decisions and affecting corporate competitiveness.
It adopts a data processing system based on a multi-layer architecture, including a detailed data layer, a summary layer, an application domain mart layer, a report mart layer, a data engine layer, and a data service layer. Through ETL automated processing components and a high-performance analysis engine built with HOLOGRES and STARROCKS, it provides diversified data service functions, supports high-concurrency, low-latency data analysis tasks, and is precisely divided into report mart, cockpit mart, and interface mart based on data resource characteristics and business needs.
It significantly improves data processing efficiency and accuracy, meets data service needs in different business scenarios, improves the utilization and management efficiency of data resources, supports the company's real-time analysis and business decision-making, and enhances the company's competitiveness.
Smart Images

Figure CN120763231A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing and analysis, and in particular to a data processing system and method based on a multi-layer architecture. Background Art
[0002] With the continuous expansion of enterprise scale and the diversification of business, the data processing system needs to handle more and more functional modules. In addition, the business processes of modern enterprises are often very complex, involving the collaborative work between multiple links. Each link will generate a large amount of data, and this data needs to be shared and circulated between different links. The data processing method with a multi-layer architecture can better adapt to this complex business process, centrally manage and process business rules through the business logic layer, and ensure the correct transmission and processing of data between different links.
[0003] In the existing technology, when faced with massive amounts of data, the data processing system has a slow processing speed and cannot meet the needs of real-time analysis. Moreover, the data service provided by the system is single in form and cannot meet the diversified needs in different business scenarios, resulting in delayed business decision-making and affecting the competitiveness of the enterprise. Therefore, how to combine a multi-layer architecture and a high-performance analysis engine to support high-concurrency, low-latency data analysis tasks, and divide data resources into three categories: report market, cockpit market and interface market, and provide diversified data service functions including report service, cockpit display service and API interface service, is the problem to be solved by the present invention. To this end, a data processing system and method based on a multi-layer architecture are proposed. Summary of the Invention
[0004] The present invention aims to provide a data processing system and method based on a multi-layer architecture to solve the problems raised in the above background technology.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] In a first aspect, a data processing system based on a multi-layer architecture includes a data processing platform, wherein the data processing platform is communicatively connected to a detailed data layer, a summary layer, an application domain market layer, a report market layer, a data engine layer, and a data service layer;
[0007] The data processing platform is constructed based on MaxCompute, is divided into multiple logical processing layers, is used for receiving, gathering, processing and analyzing raw data, forms structured and standardized data assets, realizes unified management and efficient processing of data, improves data quality, provides reliable data basis for upper-layer application, introduces an ETL automatic processing component, is used for replacing traditional manual data extraction and processing process, automatically completes scheduling, cleaning, conversion and loading operation on raw data, significantly improves data processing efficiency, reduces manual intervention frequency, enhances system stability and automation capability, reduces human error, improves accuracy and consistency of data processing;
[0008] The detailed data layer (DWD) is used for receiving and gathering raw detailed data, performing atomicity, integrity and standardization processing on data, and ensuring originality and integrity of data;
[0009] The summary layer (DWS) extracts key indicators and information items based on the detailed data of the detailed data layer, and forms a unified data caliber;
[0010] The application domain market layer (ADM) is used for subject aggregation of data of the summary layer, and forms exclusive data sets for specific application scenarios;
[0011] The report market layer is used for dividing data resources into three types of report market, cockpit market and interface market, and calculating market matching coefficients, analyzing the adaptation degree of data resources and different types of markets, and meeting the data service demand in different business scenarios;
[0012] The data engine layer includes a high-performance analysis engine constructed based on HOLOGRES and STARROCKS, and respectively connects data visualization and report tools of SMART BI and FINE REPORT, supports high-concurrency and low-latency data analysis tasks, and realizes unified access, processing and sharing of structured and semi-structured data;
[0013] The data service layer is used for providing data service functions including report service, cockpit display service and API interface service, and realizing end-to-end closed-loop management of data from collection, processing, modeling to service consumption.
[0014] Further improvement of the technical scheme of the application is that the detailed data layer specifically includes:
[0015] The detailed data layer receives raw detailed data from different business systems through various data source interfaces, and performs preliminary data gathering, and gathers scattered data into a centralized data warehouse;
[0016] After receiving and aggregating data, the detailed data layer verifies the atomicity and integrity of the data;
[0017] After completing data reception, aggregation, atomicity, and integrity checks, the detailed data layer normalizes the data. Normalization includes data format unification, data cleaning, and data encoding standardization operations. It also classifies and indexes the data according to its business attributes and time dimensions, thereby converting the original data into structured and standardized data assets.
[0018] A further improvement of the technical solution of the present invention is that the summary layer specifically includes:
[0019] The summary layer extracts normalized data from the detailed data layer and performs data aggregation and cleansing operations on the data. The aggregation operation includes performing data aggregation calculations, summarizing data by time, business type, and region, and generating data views. The cleansing operation includes removing duplicate data, correcting outliers, and filling missing values to ensure data accuracy and consistency, reduce data redundancy, and optimize data structure.
[0020] After data aggregation and cleaning, the aggregation layer extracts and standardizes key indicators;
[0021] After completing data aggregation, key indicator extraction and standardization, the aggregation layer organizes the processed data into a structured data model and stores it in the designated data storage area. The data model is designed based on business needs and analysis objectives, and organizes data according to themes or business dimensions.
[0022] A further improvement of the technical solution of the present invention is that: the application field market layer specifically includes:
[0023] The application domain mart layer obtains aggregated and standardized data from the aggregation layer, aggregates the data based on different business themes, and then classifies and organizes the data according to business needs to form exclusive data sets for specific application scenarios;
[0024] After completing topic aggregation, the application domain mart layer further processes and organizes the data according to business needs, generating indicators and data sets that meet the needs of specific business scenarios;
[0025] After completing the topic aggregation and business demand customization processing, the generated exclusive data sets for specific application scenarios will be organized in a structured manner and stored in the data warehouse.
[0026] A further improvement of the technical solution of the present invention is that the report market layer specifically includes:
[0027] Obtain customized and proprietary data sets from the application domain mart layer, classify data resources into three categories: report mart, cockpit mart, and interface mart. Perform feature analysis on the data resources and extract matching features for data classification, including data complexity, data granularity, data sharing requirements, user interaction requirements, and data update requirements. Then, assign values to each matching feature and determine the baseline value for each matching feature belonging to the three mart categories.
[0028] Based on the matching characteristics of data resources and business needs, the market matching coefficient of each data resource is calculated to evaluate the degree of compatibility between data resources and different types of markets, and then each data resource is accurately divided into the corresponding market type;
[0029] After completing the data classification and calculation of the market matching coefficient, we provide diversified data services based on the classification and matching results of data resources. At the same time, we continuously optimize the form and content of data services based on user feedback and actual usage to ensure the efficiency of data services and user satisfaction.
[0030] A further improvement of the technical solution of the present invention is that the assignment of the matching features and the baseline values of the three types of markets specifically include:
[0031] Data complexity is quantified using the number of fields. Data resources with less than 10 fields are assigned a value of 1, data resources with 10-50 fields are assigned a value of 2, and data resources with more than 50 fields are assigned a value of 3. For the report mart, the baseline value of data complexity is set to 3, for the cockpit mart, the baseline value of data complexity is set to 1, and for the interface mart, the baseline value of data complexity is set to 2.
[0032] Data granularity is quantified using the degree of data aggregation. Detailed data is assigned a value of 1, slightly aggregated data is assigned a value of 2, and highly aggregated data is assigned a value of 3. For the report mart, the baseline value of data granularity is set to 1, for the cockpit mart, the baseline value of data granularity is set to 3, and for the interface mart, the baseline value of data granularity is set to 2.
[0033] Data sharing requirements are quantified using sharing frequency, with no sharing requirement assigned a value of 1, low sharing requirement assigned a value of 2, and high sharing requirement assigned a value of 3. For the report mart, the baseline value of data sharing requirements is set to 1, for the cockpit mart, the baseline value of data sharing requirements is set to 1, and for the interface mart, the baseline value of data sharing requirements can be set to 3;
[0034] For user interaction requirements, we use interaction frequency to quantify them. Low interaction requirements are assigned a value of 1, medium interaction requirements are assigned a value of 2, and high interaction requirements are assigned a value of 3. For the report mart, the baseline value of user interaction requirements is set to 3, for the cockpit mart, the baseline value of user interaction requirements is set to 2, and for the interface mart, the baseline value of user interaction requirements is set to 1.
[0035] For data update requirements, the real-time requirements of the update are used for quantification. The real-time update requirement is assigned a value of 3, the daily update requirement is assigned a value of 2, and the regular update is assigned a value of 1. For the report mart, the baseline value of the data update requirement is set to 1. For the cockpit mart, the baseline value of the data update requirement can be set to 3. For the interface mart, the baseline value of the data update requirement is set to 2.
[0036] A further improvement of the technical solution of the present invention is that the calculation process of the market matching coefficient is:
[0037] Determine the quantitative value of each data resource based on five matching characteristics, including data complexity, data granularity, data sharing requirements, user interaction requirements, and data update requirements. Quantified values are assigned according to pre-set rules. A baseline value for each matching characteristic is also determined for the three types of marts. The baseline value is set based on the business needs and characteristics of the marts.
[0038] The weight of each matching feature is determined based on its importance in marketplace matching. The weight of data complexity is 0.20, the weight of data granularity is 0.25, the weight of data sharing requirements is 0.10, the weight of user interaction requirements is 0.15, and the weight of data update requirements is 0.3.
[0039] For each data resource and each marketplace, calculate the deviation term of each matching feature, sum the deviation terms of all matching features, and take the inverse of the sum to obtain the marketplace matching coefficient;
[0040] For each data resource, calculate its market matching coefficient with the three types of markets, compare the three market matching coefficients, and allocate the data resource to the market with the highest matching coefficient.
[0041] A further improvement of the technical solution of the present invention is that the data engine layer specifically includes:
[0042] A high-performance analytics engine built on HOLOGRES and STARROCKS is integrated with data visualization and reporting tools such as SMART BI and FINE REPORT, providing unified access to structured and semi-structured data.
[0043] After accessing the data, the data is processed by using the computing power of HOLOGRES and STARROCKS, different types of data are processed by using optimized algorithms, the operation of data cleaning, aggregation and association is quickly completed, high concurrency and low delay data analysis tasks are supported, and the query request can be responded in real time;
[0044] After the data processing and analysis are completed, the result data is shared through an interface mode, so that SMART BI and FINEREPORT obtain the required data, and intuitive visual report and analysis result are presented to the user.
[0045] The further improvement of the technical scheme of the application is that the data service layer specifically comprises:
[0046] Various data service functions are integrated, including report service, cockpit display service and API interface service;
[0047] After the functions are integrated, the data service layer provides corresponding data services according to user demand and business scenarios, the report service provides detailed data analysis for users by automatically generating and customizing reports, the cockpit display service displays key business indicators in real time through intuitive dashboards and charts, and the API interface service provides data to other systems and applications through standardized interfaces;
[0048] The corresponding data services are uniformly managed through the data service layer, so that the stability and reliability of the services are ensured, and the form and content of the data services are continuously optimized according to user feedback and actual use.
[0049] The second aspect is a data processing method based on a multi-layer architecture, which is realized based on a data processing system based on a multi-layer architecture, and comprises the following steps:
[0050] S1, receiving and gathering original detailed data from different systems through the detailed data layer, and performing atomicity, integrity and standardization processing;
[0051] S2, based on the data of the detailed data layer, data summarization and key indicator extraction are performed, redundancy is reduced, and a simple data view is provided for the upper layer application;
[0052] S3, using the application domain market layer to perform theme aggregation on the summarized data, forming data sets for specific application scenarios, and meeting the customized needs of different business scenarios;
[0053] S4, performing feature analysis on the data resources of the application domain market layer through the report market layer, calculating the market matching coefficient, and dividing the data resources into three types of report market, cockpit market and interface market;
[0054] S5. Use high-performance analysis engines to process data, support high-concurrency, low-latency analysis tasks, and provide reports, cockpit displays, and API interface services through data visualization tools.
[0055] Due to the adoption of the above technical solution, the present invention has the following technical advancements compared to the prior art:
[0056] 1. The present invention provides a data processing system and method based on a multi-layer architecture. By introducing ETL automated processing components to replace traditional manual data extraction and processing processes, the system can automatically complete data scheduling, cleaning, conversion and loading operations, significantly improving data processing efficiency, reducing the frequency of manual intervention, enhancing the stability and automation capabilities of the system, effectively reducing human errors, and improving the accuracy and consistency of data processing, saving enterprises a lot of time and labor costs.
[0057] 2. The present invention provides a data processing system and method based on a multi-layer architecture, which divides data resources into three categories: report market, cockpit market and interface market. The report market layer provides diversified data service functions including report services, cockpit display services and API interface services, meeting the data service needs in different business scenarios, enabling users to choose appropriate data service forms according to actual needs, improving the pertinence and effectiveness of data services, and accelerating the business decision-making process.
[0058] 3. The present invention provides a data processing system and method based on a multi-layer architecture, introduces a high-performance analysis engine built based on HOLOGRES and STARROCKS, supports high-concurrency, low-latency data analysis tasks, and can achieve unified access, processing and sharing of structured and semi-structured data, significantly improving data processing speed, meeting the needs of real-time analysis, enabling enterprises to obtain data insights more quickly, make business decisions in a timely manner, and improve their competitiveness.
[0059] 4. The present invention provides a data processing system and method based on a multi-layer architecture. Through the market matching coefficient, data resources can be accurately divided into corresponding market types according to the characteristics of data resources and business needs, ensuring that data resources can be classified and managed according to different business scenarios, thereby improving the utilization and management efficiency of data resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0061] Figure 1 Schematic diagram of the system function modules of the present invention;
[0062] Figure 2 Schematic diagram of the system function of the present invention. DETAILED DESCRIPTION
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0064] Example 1, as Figure 1 、 Figure 2 As shown, the present invention provides a data processing system based on a multi-layer architecture, including a data processing platform, which is communicatively connected to a detailed data layer, a summary layer, an application domain mart layer, a report mart layer, a data engine layer, and a data service layer. The detailed data layer (DWD), the summary layer (DWS), and the application domain mart layer (ADM) are the data processing layer;
[0065] The data processing platform, built on MaxCompute, is divided into multiple logical processing layers. It receives, aggregates, processes, and analyzes raw data, forming structured and standardized data assets. This enables unified data management and efficient processing, improves data quality, and provides a reliable data foundation for upper-layer applications. The platform introduces ETL automated processing components to replace traditional manual data extraction and processing processes. It automatically schedules, cleans, converts, and loads raw data, significantly improving data processing efficiency, reducing the frequency of manual intervention, enhancing system stability and automation capabilities, reducing human errors, and improving the accuracy and consistency of data processing. The ETL automated processing components specifically include:
[0066] After the ETL automation processing component is started, the original data is scheduled and extracted, and according to the preset schedule or trigger condition, a data extraction task list is automatically generated, data is extracted from various data sources (databases, file systems, APIs, etc.), and the timestamp and data volume information of data extraction are recorded. After completing data extraction, the ETL component enters the data cleaning and conversion stage, automatically cleans the extracted original data, removes duplicate records, corrects error data, fills in missing values, etc., and converts the cleaned data according to predefined rules, including format unification, data type conversion, field mapping, etc. Through automated processing, human error is reduced, data processing accuracy and consistency are improved, data quality is ensured, and the ETL component loads the cleaned and converted data into the target data warehouse. After loading is completed, detailed logs and reports are generated to record data loading time, data volume, success or failure record information, and at the same time, the entire data processing flow is monitored in real time to discover and handle abnormal situations, ensuring stable operation of the system;
[0067] The detailed data layer (DWD) is used to receive and aggregate raw detailed data, and to perform atomicity, integrity and normalization processing to ensure data originality and integrity. The detailed data layer receives raw detailed data from different business systems (ERP, CRM, log system, etc.) through various data source interfaces, and performs preliminary data aggregation. The scattered data is aggregated into a centralized data warehouse. After receiving and aggregating data, the detailed data layer verifies the atomicity and integrity of the data. The atomicity check ensures that the smallest unit of data meets the business rules, and the integrity check ensures the completeness and accuracy of the data record. If data is missing or abnormal, it is marked as problem data and repaired. After completing data reception, aggregation, atomicity and integrity checks, the detailed data layer performs normalization processing, which includes data format unification, data cleaning (removing duplicate data, correcting error data) and data encoding standardization operations. According to the business attributes and time dimensions of the data, the data is classified and indexed, and the raw data is converted into structured and normalized data assets.
[0068] The summary layer (DWS) extracts key indicators and information items based on the detailed data of the detailed data layer to form a unified data caliber. Through data aggregation and key indicator extraction, it reduces data redundancy, improves data utilization, and provides a concise and clear data view for upper-level applications. The summary layer extracts normalized data from the detailed data layer and performs aggregation and cleaning operations on the data. The aggregation operation includes aggregation calculation of the data, aggregation of data by time, business type, and regional dimensions, and generation of data views. The cleaning operation includes removing duplicate data, correcting outliers, and filling missing values, etc., to ensure the accuracy and consistency of the data, reduce data redundancy, and optimize the data structure. After the data aggregation and cleaning operations, the summary layer performs Perform key indicator extraction and standardization operations. Key indicator extraction refers to identifying and extracting indicators that are critical to business decision-making based on business needs. Key indicators are calculated from multiple data fields. Standardization ensures that the calculation methods and definitions of key indicators remain consistent throughout the system. By extracting and standardizing key indicators, a unified and clear data caliber is provided for upper-level applications, enabling business analysis and decision-making to be based on consistent data standards. After completing data aggregation, key indicator extraction and standardization, the aggregation layer organizes the processed data into a structured data model and stores it in a designated data storage area. The data model is designed based on business needs and analysis objectives, and the data is organized by subject or business dimension.
[0069] The application domain mart layer (ADM) is used to perform thematic aggregation on the data in the summary layer to form exclusive data sets for specific application scenarios. According to different business needs, the data is further processed and organized to provide customized data support for different business scenarios, meet the diverse needs in different business scenarios, improve the pertinence and effectiveness of data services, and accelerate the business decision-making process. The application domain mart layer obtains the aggregated and standardized data from the summary layer, and performs thematic aggregation on the data according to different business themes, and then classifies and organizes the data according to business needs to form exclusive data sets for specific application scenarios. After completing thematic aggregation, the application domain mart layer further processes and organizes the data according to business needs to generate indicators and data sets that meet the needs of specific business scenarios. This not only improves the pertinence of the data, but also enhances the data's support for business decision-making, allowing the data to serve business goals more directly. After completing thematic aggregation and customized processing for business needs, the generated exclusive data sets for specific application scenarios are organized in a structured manner and stored in the data warehouse;
[0070] The report market layer is used for dividing data resources into three types of report market, cockpit market and interface market, and calculating market matching coefficients, analyzing the adaptation degree of data resources and different types of markets, meeting the data service needs in different business scenarios, providing comprehensive data services including report services, cockpit display services and API interface services through diversified data service forms, and meeting the needs of different user groups;
[0071] The data engine layer includes a high-performance analysis engine constructed based on HOLOGRES and STARROCKS, and interfaces data visualization and report tools of SMART BI and FINE REPORT respectively, supports high-concurrency and low-latency data analysis tasks, realizes unified access, processing and sharing of structured and semi-structured data, significantly improves data processing speed, meets the needs of real-time analysis, improves the efficiency and accuracy of data analysis, and meets the needs of different user groups.
[0072] The data service layer is used for providing data service functions including report services, cockpit display services and API interface services, realizing end-to-end closed-loop management of data from collection, processing, modeling to service consumption, and significantly improving the utilization efficiency of data assets and the level of enterprise-level data governance.
[0073] Embodiment 2, as shown in Figure 1 , Figure 2 on the basis of embodiment 1, the application provides a technical solution: preferably, the report market layer specifically includes:
[0074] Obtain customized and processed exclusive data sets from the application domain mart layer, divide data resources into three categories: report mart, cockpit mart and interface mart, perform feature analysis on data resources, extract matching features for data classification, namely data complexity, data granularity, data sharing requirements, user interaction requirements and data update requirements, and then assign values to each matching feature. At the same time, determine the baseline value of each matching feature belonging to the three types of marts. Among them, the report mart is used to provide detailed report data, the cockpit mart is used to provide overview key indicator data, and the interface mart is used to provide API interface data for interaction with other systems, ensuring that data resources can be classified and managed according to different business scenarios, according to the matching characteristics of data resources and business Based on the needs, the market matching coefficient of each data resource is calculated, the degree of adaptability between the data resource and different types of markets is evaluated, and then each data resource is accurately divided into the corresponding market type. After completing the data classification and the calculation of the market matching coefficient, a variety of data services are provided according to the classification and matching results of the data resources. Among them, the report market provides detailed report services to meet users' needs for historical data and detailed information. The cockpit market provides real-time cockpit display services to quickly present key indicators and business overviews. The interface market provides API interface services to facilitate integration with other systems and data sharing. At the same time, based on user feedback and actual usage, the form and content of data services are continuously optimized to ensure the efficiency of data services and user satisfaction.
[0075] The assignment of matching features and baseline values for the three types of markets specifically include:
[0076] For data complexity, the number of fields is used for quantification. Data resources with less than 10 fields are assigned a value of 1, data resources with 10-50 fields are assigned a value of 2, and data resources with more than 50 fields are assigned a value of 3. For the report mart, the baseline value of data complexity is set to 3 (high complexity), for the cockpit mart, the baseline value of data complexity is set to 1 (simple data table), and for the interface mart, the baseline value of data complexity is set to 2 (medium complexity). For data granularity, the degree of data aggregation is used for quantification. Detailed data is assigned a value of 1, and light data is assigned a value of 2. Aggregated data (daily summary) is assigned a value of 2, highly aggregated data (monthly summary) is assigned a value of 3, for the report mart, the baseline value of data granularity is set to 1 (detailed data), for the cockpit mart, the baseline value of data granularity is set to 3 (highly aggregated data), for the interface mart, the baseline value of data granularity is set to 2 (lightly aggregated data); for data sharing requirements, the sharing frequency is used for quantification, no sharing requirement is assigned a value of 1, low sharing requirement (occasional sharing) is assigned a value of 2, high sharing requirement (frequent sharing) is assigned a value of 3, for the report mart, data sharing is assigned a value of 1. The baseline value of the demand is set to 1 (no sharing demand). For the cockpit mart, the baseline value of the data sharing demand is set to 1 (no sharing demand). For the interface mart, the baseline value of the data sharing demand can be set to 3 (high sharing demand). For user interaction demand, the interaction frequency is used for quantification. Low interaction demand (occasional query) is assigned a value of 1, medium interaction demand (regular query) is assigned a value of 2, and high interaction demand (frequent query and filtering) is assigned a value of 3. For the report mart, the baseline value of the user interaction demand is set to 3 (high interaction demand). For the cockpit mart, the baseline value of the user interaction demand is set to 2 (medium interaction demand). For the interface mart, the baseline value of the user interaction demand is set to 1 (low interaction demand). For data update demand, the real-time requirement of the update is used for quantification. Real-time update demand is assigned a value of 3, daily update demand is assigned a value of 2, and regular update is assigned a value of 1. For the report mart, the baseline value of the data update demand is set to 1 (regular update). For the cockpit mart, the baseline value of the data update demand can be set to 3 (real-time update). For the interface mart, the baseline value of the data update demand is set to 2 (daily update).
[0077] In addition, the calculation process of the market matching coefficient is:
[0078] Determine the quantitative value of each data resource on the five matching features. The matching features include data complexity, data granularity, data sharing requirements, user interaction requirements, and data update requirements. The quantitative values are assigned according to preset rules, and the baseline values of the three types of marts on each matching feature are determined. The baseline values are set according to the business needs and characteristics of the marts. The weight of each matching feature is determined according to its importance in mart matching, where the weight of data complexity is 0.20, the weight of data granularity is 0.25, the weight of data sharing requirements is 0.10, the weight of user interaction requirements is 0.15, and the weight of data update requirements is 0.3. For each data resource and each mart, calculate the deviation term of each matching feature, and sum the deviation terms of all matching features and take the inverse to obtain the mart matching coefficient. For each data resource, calculate its mart matching coefficient with the three types of marts, compare the three mart matching coefficients, and allocate the data resource to the mart with the highest matching coefficient.
[0079] The expression of the market matching coefficient is:
[0080]
[0081] Where M ij is the market matching coefficient, which represents the market matching coefficient between the i-th data resource and the j-th market. The closer the market matching coefficient is to 1, the higher the degree of adaptation between the data resource and the market. The closer it is to 0, the lower the degree of adaptation. V ik is the quantitative value of the i-th data resource on the k-th matching feature, k is the type of matching feature, B jk is the baseline value of the j-th market on the k-th matching feature, W k is the weight of the kth matching feature, α is the exponential adjustment parameter, which is used to control the influence of the deviation on the matching coefficient, and α=0.5, when V ik With B jk The closer it is, the ij The closer it is to 1, the ik With B jk When the deviation is large, M ij Significantly reduced, exponential function It will amplify the effect of the deviation, making the matching coefficient more sensitive to larger deviations;
[0082] The data engine layer specifically includes:
[0083] A high-performance analysis engine built on HOLOGRES and STARROCKS is integrated with data visualization and reporting tools such as SMART BI and FINE REPORT. It also provides unified access to structured and semi-structured data, adapting data of different formats to the engine through flexible data parsing and conversion mechanisms. After data access, the engine leverages the computing power of HOLOGRES and STARROCKS to process the data, using optimized algorithms for different data types to quickly complete data cleansing, aggregation, and association operations. It supports high-concurrency, low-latency data analysis tasks and can respond to query requests in real time. After data processing and analysis, the resulting data is shared through interfaces, enabling SMART BI and FINE REPORT to obtain the required data and present intuitive visual reports and analysis results to users.
[0084] The data service layer specifically includes:
[0085] The system integrates multiple data service functions, including reporting services, cockpit display services, and API interface services. The reporting service provides detailed report data to meet users' needs for historical data and detailed information. The cockpit display service provides real-time key indicators and business overviews, helping management quickly grasp business dynamics. The API interface service provides data interfaces for integration with other systems, facilitating data sharing and system integration. This ensures the diversity and flexibility of data services and meets the needs of different user groups. After functional integration is completed, the data service layer provides corresponding data services based on user needs and business scenarios. The reporting service provides users with detailed data analysis through automatic generation and customization of reports. The cockpit display service displays key business indicators in real time through intuitive dashboards and charts. The API interface service provides data to other systems and applications through standardized interfaces. The data service layer provides unified management of corresponding data services to ensure service stability and reliability. Based on user feedback and actual usage, the data service format and content are continuously optimized to continuously improve the quality of data services and user experience, ensuring that data services can better support the company's business decision-making and operations, thereby significantly improving the utilization efficiency of data assets and the level of enterprise-level data governance.
[0086] Example 3, as Figure 1 、 Figure 2 As shown, based on Examples 1-2, the present invention further provides a data processing method based on a multi-layer architecture, which is implemented based on a data processing system of a multi-layer architecture and includes the following steps:
[0087] S1. Receive and aggregate original detailed data from different systems through the detailed data layer, and perform atomicity, integrity, and normalization processing;
[0088] S2. Based on the detailed data layer data, data aggregation and key indicator extraction are performed to reduce redundancy and provide a concise data view for upper-layer applications;
[0089] S3. Use the application domain marketplace layer to perform thematic aggregation on the summarized data to form datasets for specific application scenarios to meet the customized needs of different business scenarios.
[0090] S4. Analyze the characteristics of data resources in the application domain market layer through the report market layer, calculate the market matching coefficient, and divide the data resources into three categories: report market, cockpit market, and interface market;
[0091] S5. Use high-performance analysis engines to process data, support high-concurrency, low-latency analysis tasks, and provide reports, cockpit displays, and API interface services through data visualization tools.
[0092] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data processing system based on a multi-layer architecture, comprising a data processing platform, characterized in that: The data processing platform is communicatively connected to the detailed data layer, the summary layer, the application domain market layer, the report market layer, the data engine layer and the data service layer; The data processing platform is built on MaxCompute and is used to receive, aggregate, process, and analyze raw data; The detailed data layer is used to receive and aggregate original detailed data and perform atomicity, integrity and normalization processing on the data; The summary layer extracts key indicators and information items based on the detailed data of the detailed data layer to form a unified data caliber; The application domain mart layer is used to perform thematic aggregation on the data in the summary layer to form a dedicated data set for a specific application scenario; The report market layer is used to divide data resources into three categories: report market, cockpit market and interface market, calculate the market matching coefficient, and analyze the adaptability of data resources to different types of markets; The data engine layer includes analysis engines built on HOLOGRES and STARROCKS, which are connected to the data visualization and reporting tools of SMART BI and FINE REPORT respectively; The data service layer is used to provide data service functions including report services, cockpit display services and API interface services.
2. The data processing system based on a multi-layer architecture according to claim 1, characterized in that: The detailed data layer specifically includes: The detailed data layer receives original detailed data from different business systems through multiple data source interfaces, performs preliminary data aggregation, and aggregates the scattered data into a centralized data warehouse; After receiving and aggregating data, the detailed data layer verifies the atomicity and integrity of the data; After completing data reception, aggregation, atomicity, and integrity checks, the detailed data layer normalizes the data. Normalization includes data format unification, data cleaning, and data encoding standardization operations. It also classifies and indexes the data according to its business attributes and time dimensions, thereby converting the original data into structured and standardized data assets.
3. The data processing system based on a multi-layer architecture according to claim 1, characterized in that: The aggregation layer specifically includes: The summary layer extracts normalized data from the detailed data layer and performs aggregation and cleansing operations on the data. Aggregation operations include performing aggregation calculations on the data, summarizing data by time, business type, and region, and generating data views. Cleansing operations include removing duplicate data, correcting outliers, and filling missing values. After data aggregation and cleaning, the aggregation layer extracts and standardizes key indicators; After completing data aggregation, key indicator extraction and standardization, the aggregation layer organizes the processed data into a structured data model and stores it in the designated data storage area.
4. The data processing system based on a multi-layer architecture according to claim 1, characterized in that: The application domain marketplace layer specifically includes: The application domain mart layer obtains aggregated and standardized data from the aggregation layer, aggregates the data based on different business themes, and then classifies and organizes the data according to business needs to form exclusive data sets for specific application scenarios; After completing topic aggregation, the application domain mart layer further processes and organizes the data according to business needs, generating indicators and data sets that meet the needs of specific business scenarios; After completing the topic aggregation and business demand customization processing, the generated exclusive data sets for specific application scenarios will be organized in a structured manner and stored in the data warehouse.
5. The data processing system based on a multi-layer architecture according to claim 1, characterized in that: The report mart layer specifically includes: Obtain customized and proprietary data sets from the application domain mart layer, classify data resources into three categories: report mart, cockpit mart, and interface mart. Perform feature analysis on the data resources and extract matching features for data classification, including data complexity, data granularity, data sharing requirements, user interaction requirements, and data update requirements. Then, assign values to each matching feature and determine the baseline value for each matching feature belonging to the three mart categories. Based on the matching characteristics of data resources and business needs, the market matching coefficient of each data resource is calculated to evaluate the degree of compatibility between data resources and different types of markets, and then each data resource is accurately divided into the corresponding market type; After completing the data classification and calculation of the market matching coefficient, we provide diversified data services based on the classification and matching results of data resources. At the same time, we continuously optimize the form and content of data services based on user feedback and actual usage.
6. The data processing system based on a multi-layer architecture according to claim 5, characterized in that: The assignment of matching features and baseline values of the three types of markets specifically include: Data complexity is quantified using the number of fields. Data resources with less than 10 fields are assigned a value of 1, data resources with 10-50 fields are assigned a value of 2, and data resources with more than 50 fields are assigned a value of 3. For the report mart, the baseline value of data complexity is set to 3, for the cockpit mart, the baseline value of data complexity is set to 1, and for the interface mart, the baseline value of data complexity is set to 2. Data granularity is quantified using the degree of data aggregation. Detailed data is assigned a value of 1, slightly aggregated data is assigned a value of 2, and highly aggregated data is assigned a value of 3. For the report mart, the baseline value of data granularity is set to 1, for the cockpit mart, the baseline value of data granularity is set to 3, and for the interface mart, the baseline value of data granularity is set to 2. Data sharing requirements are quantified using sharing frequency, with no sharing requirement assigned a value of 1, low sharing requirement assigned a value of 2, and high sharing requirement assigned a value of 3. For the report mart, the baseline value of data sharing requirements is set to 1, for the cockpit mart, the baseline value of data sharing requirements is set to 1, and for the interface mart, the baseline value of data sharing requirements can be set to 3; For user interaction requirements, we use interaction frequency to quantify them. Low interaction requirements are assigned a value of 1, medium interaction requirements are assigned a value of 2, and high interaction requirements are assigned a value of 3. For the report mart, the baseline value of user interaction requirements is set to 3, for the cockpit mart, the baseline value of user interaction requirements is set to 2, and for the interface mart, the baseline value of user interaction requirements is set to 1. For data update requirements, the real-time requirements of the update are used for quantification. The real-time update requirement is assigned a value of 3, the daily update requirement is assigned a value of 2, and the regular update is assigned a value of 1. For the report mart, the baseline value of the data update requirement is set to 1. For the cockpit mart, the baseline value of the data update requirement can be set to 3. For the interface mart, the baseline value of the data update requirement is set to 2.
7. The data processing system based on a multi-layer architecture according to claim 6, characterized in that: The calculation process of the market matching coefficient is: Determine the quantitative value of each data resource based on five matching characteristics, including data complexity, data granularity, data sharing requirements, user interaction requirements, and data update requirements. Quantified values are assigned according to pre-set rules, and baseline values for each matching characteristic are determined for the three types of marketplaces. The weight of each matching feature is determined based on its importance in marketplace matching. The weight of data complexity is 0.20, the weight of data granularity is 0.25, the weight of data sharing requirements is 0.10, the weight of user interaction requirements is 0.15, and the weight of data update requirements is 0.
3. For each data resource and each marketplace, calculate the deviation term of each matching feature, sum the deviation terms of all matching features, and take the inverse of the sum to obtain the marketplace matching coefficient; For each data resource, calculate its market matching coefficient with the three types of markets, compare the three market matching coefficients, and allocate the data resource to the market with the highest matching coefficient.
8. The data processing system based on a multi-layer architecture according to claim 1, characterized in that: The data engine layer specifically includes: A high-performance analytics engine built on HOLOGRES and STARROCKS, integrated with SMART BI and FINE REPORT data visualization and reporting tools, and unified access to structured and semi-structured data. After the data is accessed, HOLOGRES and STARROCKS are used to process the data, and optimization algorithms are used for different types of data to complete data cleaning, aggregation, and association operations; After completing data processing and analysis, the resulting data will be shared through an interface, allowing SMART BI and FINEREPORT to obtain the required data and present visual reports and analysis results to users.
9. The data processing system based on a multi-layer architecture according to claim 8, characterized in that: The data service layer specifically includes: Integrate multiple data service functions, including reporting services, cockpit display services, and API interface services; After functional integration is completed, the data service layer provides corresponding data services based on user needs and business scenarios. The reporting service provides data analysis for users by automatically generating and customizing reports. The cockpit display service displays key business indicators in real time through intuitive dashboards and charts. The API interface service provides data to other systems and applications through standardized interfaces. The corresponding data services are uniformly managed through the data service layer, and the form and content of data services are continuously optimized based on user feedback and actual usage.
10. A data processing method based on a multi-layer architecture, implemented based on the data processing system based on a multi-layer architecture according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1. Receive and aggregate original detailed data from different systems through the detailed data layer, and perform atomicity, integrity, and normalization processing; S2. Based on the detailed data layer data, perform data aggregation and key indicator extraction; S3. Use the application domain mart layer to thematically aggregate the summarized data to form a data set for specific application scenarios; S4. Analyze the characteristics of data resources in the application domain market layer through the report market layer, calculate the market matching coefficient, and divide the data resources into three categories: report market, cockpit market, and interface market; S5. Use the analysis engine to process data and provide reports, cockpit displays and API interface services through data visualization tools.
Citation Information
Patent Citations
Business visualization method and system based on big data platform
CN112364086A
Power grid data mart construction method and system, terminal equipment and storage medium
CN112988919A
Data analysis method and system and computer program product
CN119537461A
Report generation and optimization method and system based on offline data warehouse
CN119917595A
Intelligent configuration discovery techniques
US20170373933A1