Data warehouse data processing method, device and equipment and storage medium

CN116932664BActive Publication Date: 2026-09-08ZHEJIANG GEELY HLDG GRP CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311008510.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-10
Publication Date
2026-09-08
Estimated Expiration
2043-08-10

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供一种数仓数据处理方法、装置、设备及存储介质,旨在解决如何提高数据处理效率的技术问题

Benefits of technology

[0040] This invention discloses a data warehouse data processing method, apparatus, device, and storage medium. The method includes: acquiring vehicle attribute data and vehicle message data, and processing the vehicle attribute data and vehicle message data according to a data decryption layer; acquiring vehicle bus data, and processing the vehicle bus data according to a data decryption layer; standardizing the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to a data standardization layer; clustering the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to a data clustering layer to obtain target vehicle data; and processing the target vehicle data according to a data configuration table generated by a data processing layer and a data table configuration layer to obtain vehicle data indicators. This invention sets up a data decryption layer, a data standardization layer, a data table configuration layer, a data clustering layer, and a data processing layer in both real-time and offline data warehouses. Different data processing methods are used at different layers to process vehicle attribute data, vehicle message data, and vehicle bus data, thereby improving the processing efficiency of data in the data warehouse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116932664B_ABST
    Figure CN116932664B_ABST
Patent Text Reader

Abstract

The application discloses a kind of data warehouse data processing method, device, equipment and storage medium, the method includes: obtaining vehicle attribute data and vehicle message data, and according to data decryption layer, vehicle attribute data and vehicle message data are handled;Vehicle bus data is obtained, and according to data decryption layer, vehicle bus data is handled;According to data standardization layer, standardization is handled;According to data clustering layer, after standardization vehicle attribute data, after standardization vehicle message data and after standardization vehicle bus data are clustered, obtain target vehicle data;According to the data configuration table generated by data processing layer and data table configuration layer, target vehicle data is handled, and vehicle data index is obtained.The present application is provided with data decryption layer, data standardization layer, data table configuration layer, data clustering layer and data processing layer, according to different levels, different data processing mode is used to process data, to improve the processing efficiency of data in data warehouse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of commercial vehicle data processing technology, and in particular to a data warehouse data processing method, apparatus, equipment and storage medium. Background Technology

[0002] The data processing of commercial vehicles is too fragmented, including aggregation, data quality management, and data cleaning. This makes it impossible to centrally process data from various commercial vehicles, resulting in low data processing efficiency.

[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this invention is to provide a data warehouse data processing method, apparatus, device, and storage medium, aiming to solve the technical problem of how to improve data processing efficiency.

[0005] To achieve the above objectives, the present invention provides a data warehouse data processing method, which is applied to a data acquisition and scheduling platform, the data acquisition and scheduling platform including a real-time data warehouse and an offline data warehouse; both the real-time data warehouse and the offline data warehouse include: a data decryption layer, a data standardization layer, a data table configuration layer, a data clustering layer, and a data processing layer; the data warehouse generation method includes the following steps:

[0006] Obtain vehicle attribute data and vehicle message data, and process the vehicle attribute data and vehicle message data according to the data decryption layer;

[0007] Acquire vehicle bus data and process the vehicle bus data according to the data decryption layer;

[0008] The processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data are standardized according to the data standardization layer.

[0009] Based on the data clustering layer, the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data are clustered to obtain the target vehicle data;

[0010] The target vehicle data is processed based on the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators.

[0011] Optionally, the step of obtaining vehicle attribute data and vehicle message data, and processing the vehicle attribute data and vehicle message data according to the data decryption layer, includes:

[0012] Obtain vehicle attribute data and write the vehicle attribute data into the data decryption layer for storage;

[0013] Vehicle message data is retrieved from the message queue and written to the data decryption layer;

[0014] The vehicle message data is decrypted according to the data decryption layer, and the decrypted vehicle message data is stored.

[0015] The decrypted vehicle message data is pushed down to the message queue of the data standardization layer according to the data decryption layer.

[0016] Optionally, the step of acquiring vehicle bus data and processing the vehicle bus data according to the data decryption layer includes:

[0017] Acquire vehicle bus data and synchronize the vehicle attribute data;

[0018] The vehicle bus data is decrypted through the data decryption layer, and the vehicle attribute data and the decrypted vehicle bus data are stored in the data decryption layer.

[0019] Optionally, the step of standardizing the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to the data standardization layer includes:

[0020] The decrypted vehicle message data is split according to a preset granularity type based on the data standardization layer.

[0021] The data standardization layer is used to standardize the split vehicle message data, the decrypted vehicle bus data, and the vehicle attribute data.

[0022] The standardized vehicle message data is pushed down to the message queue of the data clustering layer, and the standardized vehicle message data, standardized vehicle bus data, and standardized vehicle attribute data are stored in the data standardization layer.

[0023] Optionally, the step of clustering the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to the data clustering layer to obtain the target vehicle data includes:

[0024] Based on the data clustering layer, the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data are clustered to obtain the target vehicle data;

[0025] The target vehicle data is stored in the data clustering layer and then pushed down to the data processing layer.

[0026] Optionally, the step of processing the target vehicle data based on the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators includes:

[0027] Obtain the standardized vehicle message data, the standardized vehicle bus data, the standardized vehicle attribute data, the target vehicle data, and the data configuration table generated by the data table configuration layer;

[0028] The standardized vehicle message data, standardized vehicle bus data, standardized vehicle attribute data, target vehicle data, and data configuration table generated by the data table configuration layer are processed according to the data indicator processing theme, type granularity, statistical period, and data statistical indicators to obtain vehicle data indicators.

[0029] Optionally, after the step of processing the target vehicle data according to the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators, the method further includes...

[0030] The expired vehicle message data and the vehicle message data to be synchronized in the real-time data warehouse are stored in the offline data warehouse;

[0031] The vehicle summary index table from the offline data warehouse is stored in the real-time data warehouse.

[0032] In addition, to achieve the above objectives, the present invention also proposes a data warehouse data processing device, which includes: a data processing module;

[0033] The data processing module is used to acquire vehicle attribute data and vehicle message data, and to process the vehicle attribute data and vehicle message data according to the data decryption layer.

[0034] The data processing module is also used to acquire vehicle bus data and process the vehicle bus data according to the data decryption layer.

[0035] The data processing module is also used to standardize the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to the data standardization layer.

[0036] The data processing module is also used to cluster the standardized vehicle attribute data, standardized vehicle message data and standardized vehicle bus data according to the data clustering layer to obtain target vehicle data.

[0037] The data processing module is further configured to process the target vehicle data according to the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators.

[0038] Furthermore, to achieve the above objectives, the present invention also proposes a data warehouse data processing device, which includes a memory, a processor, and a data warehouse data processing program stored in the memory and capable of running on the processor. The data warehouse data processing program is configured to implement the data warehouse data processing method as described above.

[0039] In addition, to achieve the above objectives, the present invention also proposes a storage medium storing a data warehouse data processing program, which, when executed by a processor, implements the data warehouse data processing method as described above.

[0040] This invention discloses a data warehouse data processing method, apparatus, device, and storage medium. The method includes: acquiring vehicle attribute data and vehicle message data, and processing the vehicle attribute data and vehicle message data according to a data decryption layer; acquiring vehicle bus data, and processing the vehicle bus data according to a data decryption layer; standardizing the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to a data standardization layer; clustering the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to a data clustering layer to obtain target vehicle data; and processing the target vehicle data according to a data configuration table generated by a data processing layer and a data table configuration layer to obtain vehicle data indicators. This invention sets up a data decryption layer, a data standardization layer, a data table configuration layer, a data clustering layer, and a data processing layer in both real-time and offline data warehouses. Different data processing methods are used at different layers to process vehicle attribute data, vehicle message data, and vehicle bus data, thereby improving the processing efficiency of data in the data warehouse. Attached Figure Description

[0041] Figure 1 This is a schematic diagram of the structure of the data warehouse data processing equipment in the hardware operating environment involved in the embodiments of the present invention;

[0042] Figure 2 This is a flowchart illustrating the first embodiment of the data warehouse data processing method of the present invention;

[0043] Figure 3 This is a flowchart illustrating the second embodiment of the data warehouse data processing method of the present invention;

[0044] Figure 4 This is a flowchart illustrating the third embodiment of the data warehouse data processing method of the present invention;

[0045] Figure 5 This is a real-time data warehouse data processing structure diagram of an embodiment of the data warehouse data processing method of the present invention;

[0046] Figure 6 This is an offline data warehouse data processing structure diagram of an embodiment of the data warehouse data processing method of the present invention;

[0047] Figure 7 This is a structural block diagram of the first embodiment of the data warehouse data processing device of the present invention.

[0048] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0049] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0050] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a data warehouse data processing device in the hardware operating environment involved in the embodiments of the present invention.

[0051] like Figure 1 As shown, the data warehouse data processing device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen, and optionally, it may also include a standard wired interface or a wireless interface. In this invention, the wired interface of the user interface 1003 may be a USB interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or non-volatile memory (NVM), such as a disk storage device. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.

[0052] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the data warehouse data processing equipment and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0053] like Figure 1 As shown, the memory 1005, which is identified as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a data warehouse data processing program.

[0054] exist Figure 1 In the data warehouse data processing device shown, the network interface 1004 is mainly used to connect to the backend server and communicate with the backend server; the user interface 1003 is mainly used to connect to the user device; the data warehouse data processing device calls the data warehouse data processing program stored in the memory 1005 through the processor 1001 and executes the data warehouse data processing method provided in the embodiment of the present invention.

[0055] Based on the above hardware structure, an embodiment of the data warehouse data processing method of the present invention is proposed.

[0056] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the data warehouse data processing method of the present invention, which presents the first embodiment of the data warehouse data processing method of the present invention.

[0057] Step S10: Obtain vehicle attribute data and vehicle message data, and process the vehicle attribute data and vehicle message data according to the data decryption layer.

[0058] It should be noted that the execution subject of this embodiment may be a computer software service device with data processing, network communication and program running functions, such as a data acquisition and scheduling platform, or other electronic devices that can achieve the same or similar functions. This embodiment does not limit this.

[0059] It should be noted that this embodiment is applied to a data acquisition and scheduling platform, which includes a real-time data warehouse and an offline data warehouse. Both the real-time and offline data warehouses have the same data decryption layer, data standardization layer, data table configuration layer, data clustering layer, and data processing layer.

[0060] It's important to note that the data acquisition scheduling platform can be Dolphinscheduler. Apache DolphinScheduler is a distributed, scalable, and visual DAG workflow task scheduling platform. It aims to resolve the complex dependencies in data processing workflows, making the scheduling system ready to use out of the box. Compared to traditional task scheduling frameworks, the Dolphinscheduler task scheduling platform provides a visual interface, making it simple and easy to use. It also offers an error retry mechanism to ensure stable workflow operation. The Dolphinscheduler task scheduling platform can be used to schedule and monitor the data acquisition process in distributed and heterogeneous data architectures.

[0061] It is understandable that the same data decryption layer, data standardization layer, data table configuration layer, data clustering layer, and data processing layer are set up in both real-time and offline data warehouses. This is to ensure a clear data structure during data processing, reduce redundant development, make the data more organized, and mitigate the impact of raw data.

[0062] It's understandable that the heterogeneous data architecture could be DataX, an open-source offline synchronization tool for heterogeneous data sources. DataX aims to achieve stable and efficient data synchronization between various heterogeneous data sources, including relational databases (MySQL, Oracle, etc.), HDFS, Hive, HBase, and FTP. For this embodiment, configuring JDBC information and field mapping for relational MySQL data extraction is convenient and easy to use.

[0063] It's important to note that the distributed architecture can be Flink. Flink is a memory-based distributed computing framework suitable for streaming, unbounded data, as well as bounded batch processing. For batch processing, FlinkSQL highly integrates with Hive metadata, offering higher efficiency than HiveSQL. For stream processing, Flink integrates Flink Kafka Consumer with Flink's checkpointing mechanism to provide precise, one-time processing semantics. Flink not only relies on Kafka consumer group offset tracking but also internally tracks and inspects these offsets.

[0064] It is understandable that vehicle attribute data can be obtained through a heterogeneous data architecture, and vehicle message data can be obtained through a distributed architecture.

[0065] It should be noted that the data decryption layer (ODS) is used to decode encrypted vehicle message data and store the decrypted vehicle message data and vehicle attribute data.

[0066] Understandably, vehicle attribute data can include user information, vehicle information, equipment information, basic component information, vehicle sales information, and battery traceability information. This data is typically stored in a relational database and includes two main categories: factual data and dimensional data. The data generation is relatively slow and the volume is small, so it is usually integrated with the data warehouse through batch fetching or change data capture.

[0067] Understandably, vehicle message data can be data reported by commercial vehicles according to protocols, such as GB32960, GB17691, and TSP. This dynamic data describes the current vehicle status; it is generated frequently, is structured, and encrypted, and is interfaced with the data warehouse via message queues.

[0068] Furthermore, in order to improve data processing efficiency, step S10 of this embodiment may include:

[0069] Obtain vehicle attribute data and write the vehicle attribute data into the data decryption layer for storage;

[0070] Vehicle message data is retrieved from the message queue and written to the data decryption layer;

[0071] The vehicle message data is decrypted according to the data decryption layer, and the decrypted vehicle message data is stored.

[0072] The decrypted vehicle message data is pushed down to the message queue of the data standardization layer according to the data decryption layer.

[0073] Understandably, the distributed architecture Flink consumes vehicle message data from the Kafka message queue, decrypts the vehicle message data using Flink according to the decryption rules, and stores the decrypted vehicle message data.

[0074] In the specific implementation, in the real-time data warehouse, after the data decryption layer receives vehicle attribute data transmitted from the heterogeneous data architecture and vehicle message data transmitted from the distributed architecture, it changes the prefix of the data table where the vehicle attribute data and vehicle message data are stored to "ODS_", and decrypts the vehicle message data according to the decryption rules and the distributed architecture. The decrypted vehicle message data is then pushed down to the message queue of the data standardization layer, and the decrypted vehicle message data and vehicle attribute data are stored.

[0075] It's worth noting that the real-time data warehouse uses Doris for storage. Data retention in Kafka has time limitations, such as only keeping the most recent 7 days. Real-time data warehouses inevitably involve recalculation and rollback operations, thus requiring a separate storage technology. The vehicle condition data in the real-time data warehouse is massive, and its popularity decays over time. Doris provides different granularities for hot and cold data. For example, Doris supports dynamic partition management and configuration (TTL), automatic partition addition and deletion, and different lifecycles for hot and cold data to facilitate data recycling. Doris supports multiple storage media (SSD / SATA), and combined with time partitioning, it can store hot data on SSDs and cold data on SATA. For business scenarios involving both hot and cold data, partitions can be selected during the Doris table creation phase to improve performance.

[0076] Step S20: Obtain vehicle bus data and process the vehicle bus data according to the data decryption layer.

[0077] It should be noted that vehicle bus data can be the vehicle's CAN bus data, i.e., unstructured data that collects various vehicle signals. In CAN data, the CAN Serial number reported by each vehicle message may be different; that is, not every CAN Serial number is reported. Furthermore, CAN data is encrypted, reported frequently, and involves large amounts of data, interfacing with the data warehouse via message queues.

[0078] Understandably, in an offline data warehouse, the data decryption layer is used to store the acquired vehicle bus data.

[0079] It's important to note that the tables stored in offline and real-time data warehouses are divided into fact tables and dimension tables. Fact tables record specific events, containing the specific elements of each event and the details of what happened. The fact table is the backbone, concisely describing a fact. Characteristics of fact tables include: each row includes additive numerical measures; foreign keys connecting to dimension tables, typically two or more; a large number of rows; relatively narrow content (few columns, mainly foreign key IDs and measures); and frequent changes, with many new tables added daily. Dimension tables, on the other hand, describe facts. Each dimension table corresponds to an object or concept in the real world. Each dimension table expands upon each column / field in the fact table. Generally, for data warehouses, fact tables are added, deleted, and modified more frequently than dimension tables. After the model is built, the data in dimension tables remains relatively stable. Characteristics of dimension tables include: a wide scope (multiple attributes, many columns); a relatively small number of rows compared to fact tables; and relatively fixed content: a coding table. For example: city administrative code table, etc.

[0080] It's important to note that the offline data warehouse uses Hive on HDFS for storage. Since the offline data warehouse stores massive amounts of historical data, the cost and performance of data storage must be considered. From a cost perspective, regarding the cost and efficiency of storing hot and cold data: HDFS supports different storage media (DISK / SSD / RAM), allowing cold data to be configured in DISK; in Hive, cold data can be configured in a high-compression ORC+Snappy format to reduce storage costs. From a performance perspective, focusing on the response speed of offline business scenarios: hot data is categorized by popularity, utilizing HDFS storage strategies (Lazy_persist|ALL_SSD....) for efficient data loading; hot data is configured in Hive in a high-query-efficiency Parquet+Lzo format; Hive supports efficient query engines such as Presto.

[0081] Furthermore, in order to improve data processing efficiency, step S20 of this embodiment may include:

[0082] Acquire vehicle bus data and synchronize the vehicle attribute data;

[0083] The vehicle bus data is decrypted through the data decryption layer, and the vehicle attribute data and the decrypted vehicle bus data are stored in the data decryption layer.

[0084] It should be noted that offline data warehouses can import vehicle bus data in one go using a heterogeneous data architecture via batch processing or file processing, and can also synchronize offline data from real-time data warehouses using a heterogeneous data architecture via batch processing or file processing.

[0085] It is understandable that the vehicle bus data is CAN bus data. Therefore, it is necessary to decrypt the vehicle bus data according to the heterogeneous architecture and store the decrypted vehicle bus data.

[0086] Step S30: Standardize the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to the data standardization layer.

[0087] It should be noted that standardization of processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data can include data cleaning, field unification, format unification, and data splitting.

[0088] It should be noted that the processed vehicle attribute data is stored separately as dimension tables and fact tables.

[0089] Furthermore, in order to improve data processing efficiency, step S30 of this embodiment may include:

[0090] The decrypted vehicle message data is split according to a preset granularity type based on the data standardization layer.

[0091] The data standardization layer is used to standardize the split vehicle message data, the decrypted vehicle bus data, and the vehicle attribute data.

[0092] The standardized vehicle message data is pushed down to the message queue of the data clustering layer, and the standardized vehicle message data, standardized vehicle bus data, and standardized vehicle attribute data are stored in the data standardization layer.

[0093] Understandably, after understanding the vehicle indicators for commercial vehicles, it was found that the collection of indicator items varies greatly among different models. Furthermore, vehicle model is also an important dimension in business scenarios. While considering storage costs and query efficiency, the final decision was to use vehicle model as the granularity for data splitting at the business level. For example, at the data standardization layer, the granularity of vehicle dynamic data storage (taking the E5 electric vehicle as an example) is as follows: E5_GB32960 protocol data, E5_TSP protocol data, and E5_CAN protocol data.

[0094] Understandably, both the real-time data warehouse and the offline data warehouse's data standardization layer (DWD) perform data standardization processing and storage, with the stored data tables prefixed with "DWD_". In the real-time data warehouse, the standardized vehicle message data is pushed down to the message queue of the data clustering layer.

[0095] It is understandable that the decrypted vehicle message data is split in the real-time data warehouse, and the split vehicle message data and vehicle attribute data are standardized, while the vehicle bus data is standardized in the offline data warehouse.

[0096] It should be noted that data standardization is divided into data type standardization and data structure standardization. For ease of understanding, please refer to Tables 1 and 2. Table 1 is the data type standardization table, and Table 2 is the data structure standardization table. In Table 1, char and varchar represent character data types, text represents text data types, tinyint, smallint, mediumint, int, bigint, uint32, uint64, and int32 represent integer data types, float represents single-precision floating-point type, double represents double-precision floating-point type, decimal represents variable-precision floating-point type, datetime and data represent date and time data types, string represents string data type, and boolean represents boolean data type.

[0097] Table 1 - Data Type Standardization Table

[0098]

[0099]

[0100] Table 2_Data Structure Standardization Table

[0101]

[0102] Step S40: Cluster the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to the data clustering layer to obtain the target vehicle data.

[0103] It's important to note that the Data Warehouse Detail (DWM) clustering layer performs a light aggregation of data, generating a series of general intermediate tables. The DWM layer primarily serves the Data Warehouse Detail (DWS) layer. At the requirement level, there is a significant amount of computation directly between the DWM and DWS layers, and the results of this computation may be reused across multiple DWS topic layers. Therefore, performing light aggregation at the DWM layer beforehand facilitates data reuse within the data warehouse and improves its response efficiency.

[0104] Understandably, in order to make subsequent statistical calculations more convenient and efficient, and to reduce the correlation between large tables, taking the vehicle condition wide table business as an example, the relevant data around the vehicle condition will be integrated into a wide table of the vehicle during the real-time calculation process. Combined with the previous refinement of data granularity to the vehicle type (fuel, new energy), a wide table needs to be formed at the DWM layer based on the trip cycle (from one start to the end of the engine) and the charging cycle (from the start of a charge to the end of the charge).

[0105] Furthermore, in order to improve data processing efficiency, step S40 of this embodiment may include:

[0106] Based on the data clustering layer, the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data are clustered to obtain the target vehicle data;

[0107] The target vehicle data is stored in the data clustering layer and then pushed down to the data processing layer.

[0108] It is understandable that in real-time and offline data warehouses, common vehicle bus data, vehicle message data, and vehicle attribute data in business scenarios are lightly summarized, such as conditional association summarization, time-dimensional summary of monthly tables, yearly tables, etc., and the summarized data is stored in the DWM layer.

[0109] Step S50: Process the target vehicle data according to the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators.

[0110] It should be noted that the business processing of the data table configuration layer involves filtering out dimension tables and storing the dimension table data. Firstly, the confirmation of dimension tables requires manual configuration. Secondly, considering future additions or subtractions of dimension data, the design employs a configurable approach to support data changes at the dimension table level.

[0111] Understandably, the data processing layer DWS (Data Warehouse Service) uses the subject objects of analysis as the modeling driver, and builds a summary indicator table with common granularity based on the indicator requirements of upper-layer applications and products.

[0112] Understandably, the DWS layer will construct a summary index table for the following topics proposed for commercial vehicles: vehicle registration data, vehicle fault data, driving behavior data, fuel vehicle energy consumption data, electric vehicle energy consumption data, engine data, pure electric vehicle BMS battery data, fuel cell data, drive motor data, VCU status data, component version data, DC-DC status data, BDCAC status data, SDCAC status data, HCM status data, RCU status data, autonomous driving data, and OTA data.

[0113] For ease of understanding, please refer to Figure 5 and Figure 6 To explain, Figure 5 This is a diagram of the real-time data warehouse data processing structure. Figure 6 This is a diagram of the offline data warehouse data processing structure. Figure 5 Vehicle attribute data is obtained based on a heterogeneous data architecture, and vehicle message data is obtained based on a distributed architecture. The vehicle attribute data and vehicle message data are processed through a data decryption layer, a data standardization layer, a data clustering layer, and a data processing layer. After processing the vehicle message data, the data decryption layer, the data standardization layer, and the data clustering layer will push down to the next layer. The data table configuration layer will transmit the data configuration table to the data processing layer for processing. Figure 6 The system acquires vehicle bus data and real-time data warehouse synchronization data based on a heterogeneous data architecture. The system processes vehicle attribute data and synchronization data through a data decryption layer, a data standardization layer, a data clustering layer, and a data processing layer. The data table configuration layer transmits the data configuration table to the data processing layer for processing.

[0114] This invention discloses a data warehouse data processing method, apparatus, device, and storage medium. The method includes: acquiring vehicle attribute data and vehicle message data, and processing the vehicle attribute data and vehicle message data according to a data decryption layer; acquiring vehicle bus data, and processing the vehicle bus data according to a data decryption layer; standardizing the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to a data standardization layer; clustering the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to a data clustering layer to obtain target vehicle data; and processing the target vehicle data according to a data configuration table generated by a data processing layer and a data table configuration layer to obtain vehicle data indicators. This invention sets up a data decryption layer, a data standardization layer, a data table configuration layer, a data clustering layer, and a data processing layer in both real-time and offline data warehouses. Different data processing methods are used at different layers to process vehicle attribute data, vehicle message data, and vehicle bus data, thereby improving the processing efficiency of data in the data warehouse.

[0115] Reference Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the data warehouse data processing method of the present invention, based on the above. Figure 2 The first embodiment shown presents a second embodiment of the data warehouse data processing method of the present invention.

[0116] In the second embodiment, step S50 includes:

[0117] Step S501: Obtain the standardized vehicle message data, the standardized vehicle bus data, the standardized vehicle attribute data, the target vehicle data, and the data configuration table generated by the data table configuration layer.

[0118] Understandably, the Data Processing Layer (DWS) obtains standardized vehicle message data, standardized vehicle bus data, and standardized vehicle attribute data from the Data Standardization Layer (DWD), obtains the data configuration table from the Data Table Configuration Layer, and obtains the target vehicle data from the Data Clustering Layer (DWM).

[0119] Step S502: Process the standardized vehicle message data, the standardized vehicle bus data, the standardized vehicle attribute data, the target vehicle data, and the data configuration table generated by the data table configuration layer according to the data indicator processing topic, type granularity, statistical period, and data statistical indicators to obtain vehicle data indicators.

[0120] It should be noted that statistical indicator tables with consistent naming conventions and definitions should be constructed around the theme of vehicle fault data (e.g., to form daily, weekly, and monthly granular summary details, or based on vehicle model dimensions, such as a daily summary table at the vehicle model category level, to facilitate the organization of the report data structure in the next step).

[0121] It should be noted that the data processing layer focuses on the theme of vehicle faults, sorting out the relevant indicator data and clarifying the source of each indicator data. For example, vehicle fault information from the vehicle networking system is categorized as follows: BDCAC fault level, BDCAC fault code, SDCAC fault code, HCM fault level, HCM fault code, EGSM fault level, battery thermal runaway alarm, DC-DC fault level, etc.

[0122] Understandably, the granularity can be based on vehicle model, the statistical period can be 1 day, 7 days, or 30 days, and the data statistical indicators can be derived indicators. Derived indicators refer to indicators generated by combining basic indicators or composite indicators with dimension members, statistical attributes, management attributes, etc., such as year-on-year, month-on-month, and percentage of BDCAC fault code occurrences, and cumulative battery thermal runaway alarm values.

[0123] It should be noted that basic metrics are an indivisible set of concepts that express the atomic quantifiable attributes of a business entity, such as the number of times BDCAC fault codes occur or the number of times battery thermal runaway alarms occur. Composite metrics are a set of calculated metrics formed on top of basic metrics through certain calculation rules, such as the frequency of BDCAC fault codes and the frequency of battery thermal runaway alarms.

[0124] In the specific implementation, using vehicle model as the granularity and 1 day as the cycle, derived indicators are sorted out, and a sample table of vehicle fault topics is created, as shown in the following example:

[0125]

[0126]

[0127] COMMENT 'Vehicle Model Granularity Vehicle Fault Summary Fact Sheet for the Recent Day'

[0128] PARTITIONED BY(ds STRING COMMENT'Partition field YYYYMMDD').

[0129] This embodiment acquires the standardized vehicle message data, standardized vehicle bus data, standardized vehicle attribute data, target vehicle data, and the data configuration table generated by the data table configuration layer. It processes these data according to the data indicator processing theme, type granularity, statistical period, and data statistical indicators to obtain vehicle data indicators. In this embodiment, the data processing layer processes data based on the data indicator processing theme, type granularity, statistical period, and data statistical indicators to obtain vehicle data indicators. These vehicle data indicators are then written into a vehicle summary indicator table to complete the processing of the vehicle data, thereby making the relationships between the data clearer and more organized.

[0130] Reference Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the data warehouse data processing method of the present invention, based on the above. Figure 3 The second embodiment shown presents a third embodiment of the data warehouse data processing method of the present invention.

[0131] In the third embodiment, after step S50, the method further includes:

[0132] Step S601: Store the expired vehicle message data and the vehicle message data to be synchronized in the real-time data warehouse to the offline data warehouse.

[0133] It's important to note that the approach of synchronizing data from a real-time data warehouse to an offline data warehouse primarily targets data backup scenarios. Based on business requirements, the real-time data warehouse stores the most recent six months of data, while the offline data warehouse stores three years of historical data. Therefore, it's necessary to periodically synchronize data from the real-time data warehouse to the offline data warehouse.

[0134] In practice, the real-time data warehouse is processed daily through a distributed architecture computing program to filter out the data that needs to be synchronized to the offline data warehouse. Based on the data acquisition and scheduling platform, the expired data stored in the real-time data warehouse is synchronized to the offline data warehouse through a distributed architecture, thus completing the process of backing up historical data from the real-time data warehouse to the offline data warehouse.

[0135] Step S602: Store the vehicle summary index table from the offline data warehouse to the real-time data warehouse.

[0136] It's important to note that the approach of synchronizing computation results from an offline data warehouse to a real-time data warehouse primarily targets offline business scenarios. Offline business scenarios typically require real-time data processing at the T+1 level, necessitating daily batch processing to generate computation results for offline business data. Technically, the offline data warehouse utilizes Hive. However, since Hive is unsuitable for directly providing data to external applications, a separate storage medium is needed at the ADS layer (application layer) to meet the requirements of direct connection to external applications. The real-time data warehouse, Doris, based on an MPP architecture, offers high-efficiency computing capabilities and supports materialized views to accelerate data processing. It is ideally suited for storing ADS layer data and providing services to external applications; therefore, the ADS layer uniformly stores and services data within Doris.

[0137] In the specific implementation, offline business is processed in batches based on the data acquisition and scheduling platform and the distributed architecture. The vehicle summary index table generated by the batch processing is stored on the offline data warehouse, and the vehicle summary index table stored on the offline data warehouse is synchronized to the real-time data warehouse based on the data acquisition and scheduling platform and the distributed architecture.

[0138] This embodiment stores expired vehicle message data and vehicle message data to be synchronized in the real-time data warehouse to the offline data warehouse; it also stores the vehicle summary index table from the offline data warehouse to the real-time data warehouse. This embodiment synchronizes and backs up data between the offline and real-time data warehouses using a distributed architecture, thereby ensuring data real-time performance and security.

[0139] Furthermore, this embodiment of the invention also proposes a storage medium storing a data warehouse data processing program, which, when executed by a processor, implements the data warehouse data processing method as described above.

[0140] In addition, refer to Figure 7 The present invention also proposes a data warehouse data processing device, which includes: a data processing module 10;

[0141] The data processing module 10 is used to acquire vehicle attribute data and vehicle message data, and process the vehicle attribute data and vehicle message data according to the data decryption layer.

[0142] The data processing module 10 is also used to acquire vehicle bus data and process the vehicle bus data according to the data decryption layer.

[0143] The data processing module 10 is also used to standardize the processed vehicle attribute data, processed vehicle message data and processed vehicle bus data according to the data standardization layer.

[0144] The data processing module 10 is further configured to cluster the standardized vehicle attribute data, standardized vehicle message data and standardized vehicle bus data according to the data clustering layer to obtain target vehicle data.

[0145] The data processing module 10 is further configured to process the target vehicle data according to the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators.

[0146] This invention discloses a data warehouse data processing method, apparatus, device, and storage medium. The method includes: acquiring vehicle attribute data and vehicle message data, and processing the vehicle attribute data and vehicle message data according to a data decryption layer; acquiring vehicle bus data and processing the vehicle bus data according to a data decryption layer; standardizing the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to a data standardization layer; clustering the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to a data clustering layer to obtain target vehicle data; and processing the target vehicle data according to a data configuration table generated by a data processing layer and a data table configuration layer to obtain vehicle data indicators. This invention sets up a data decryption layer, a data standardization layer, a data table configuration layer, a data clustering layer, and a data processing layer in both real-time and offline data warehouses. Different data processing methods are used at different layers to process vehicle attribute data, vehicle message data, and vehicle bus data, thereby improving the processing efficiency of data in the data warehouse.

[0147] Based on the first embodiment of the data warehouse data processing device of the present invention described above, a second embodiment of the data warehouse data processing device of the present invention is proposed.

[0148] In this embodiment, the data processing module 10 is used to acquire vehicle attribute data and write the vehicle attribute data into the data decryption layer for storage.

[0149] Furthermore, the data processing module 10 is also used to obtain vehicle message data from the message queue and write the vehicle message data into the data decryption layer.

[0150] Furthermore, the data processing module 10 is also used to decrypt the vehicle message data according to the data decryption layer and store the decrypted vehicle message data.

[0151] Furthermore, the data processing module 10 is also used to push the decrypted vehicle message data down to the message queue of the data standardization layer according to the data decryption layer.

[0152] Furthermore, the data processing module 10 is also used to acquire vehicle bus data and synchronize the vehicle attribute data.

[0153] Furthermore, the data processing module 10 is also used to decrypt the vehicle bus data through the data decryption layer, and store the vehicle attribute data and the decrypted vehicle bus data in the data decryption layer.

[0154] Furthermore, the data processing module 10 is also used to split the decrypted vehicle message data according to a preset granularity type based on the data standardization layer.

[0155] Furthermore, the data processing module 10 is also used to standardize the split vehicle message data, the decrypted vehicle bus data, and the vehicle attribute data according to the data standardization layer.

[0156] Furthermore, the data processing module 10 is also used to push the standardized vehicle message data down to the message queue of the data clustering layer, and store the standardized vehicle message data, standardized vehicle bus data and standardized vehicle attribute data in the data standardization layer.

[0157] Furthermore, the data processing module 10 is also used to cluster the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to the data clustering layer to obtain target vehicle data.

[0158] Furthermore, the data processing module 10 is also used to store the target vehicle data in the data clustering layer and push the target vehicle data down to the data processing layer.

[0159] Furthermore, the data processing module 10 is also used to acquire the standardized vehicle message data, the standardized vehicle bus data, the standardized vehicle attribute data, the target vehicle data, and the data configuration table generated by the data table configuration layer.

[0160] Furthermore, the data processing module 10 is also used to process the standardized vehicle message data, the standardized vehicle bus data, the standardized vehicle attribute data, the target vehicle data, and the data configuration table generated by the data table configuration layer according to the data indicator processing theme, type granularity, statistical period, and data statistical indicators to obtain vehicle data indicators.

[0161] Furthermore, the data processing module 10 is also used to store the expired vehicle message data and the vehicle message data to be synchronized in the real-time data warehouse to the offline data warehouse.

[0162] Furthermore, the data processing module 10 is also used to store the vehicle summary index table of the offline data warehouse to the real-time data warehouse.

[0163] Other embodiments or specific implementations of the data warehouse data processing device of the present invention can be referred to the above-described method embodiments, and will not be repeated here.

[0164] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0165] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0166] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory image (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0167] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A data warehouse data processing method, characterized in that, The data warehouse data processing method is applied to a data acquisition and scheduling platform, which includes a real-time data warehouse and an offline data warehouse; both the real-time data warehouse and the offline data warehouse include: The data warehouse data processing method comprises a data decryption layer, a data standardization layer, a data table configuration layer, a data clustering layer, and a data processing layer, and includes the following steps: The system acquires vehicle attribute data and vehicle message data, and processes the vehicle attribute data and vehicle message data according to the data decryption layer. The vehicle message data is encrypted structured data reported according to the vehicle communication protocol. The data decryption layer is used to decrypt the data and store the decrypted data. The vehicle bus data is acquired and processed according to the data decryption layer. The vehicle bus data is encrypted unstructured signal data. The processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data are standardized according to the data standardization layer. Based on the data clustering layer, the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data are clustered to obtain the target vehicle data; The target vehicle data is processed based on the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators; The forms stored in the real-time data warehouse and the offline data warehouse are divided into fact tables and dimension tables. The data table configuration layer filters the dimension tables through the generated data configuration table. The real-time data warehouse obtains vehicle attribute data based on a heterogeneous data architecture and vehicle message data based on a distributed architecture. It processes the vehicle attribute data and vehicle message data through a data decryption layer, a data standardization layer, a data clustering layer, and a data processing layer. After processing the vehicle message data, the data decryption layer, the data standardization layer, and the data clustering layer will be pushed down to the next layer. The offline data warehouse acquires vehicle bus data and real-time data warehouse synchronization data based on a heterogeneous data architecture, and processes vehicle attribute data and synchronization data through a data decryption layer, a data standardization layer, a data clustering layer, and a data processing layer.

2. The data warehouse data processing method as described in claim 1, characterized in that, The steps of acquiring vehicle attribute data and vehicle message data, and processing the vehicle attribute data and vehicle message data according to the data decryption layer, include: Obtain vehicle attribute data and write the vehicle attribute data into the data decryption layer for storage; Vehicle message data is retrieved from the message queue and written to the data decryption layer; The vehicle message data is decrypted according to the data decryption layer, and the decrypted vehicle message data is stored. The decrypted vehicle message data is pushed down to the message queue of the data standardization layer according to the data decryption layer.

3. The data warehouse data processing method as described in claim 2, characterized in that, The step of acquiring vehicle bus data and processing the vehicle bus data according to the data decryption layer includes: Acquire vehicle bus data and synchronize the vehicle attribute data; The vehicle bus data is decrypted through the data decryption layer, and the vehicle attribute data and the decrypted vehicle bus data are stored in the data decryption layer.

4. The data warehouse data processing method as described in claim 3, characterized in that, The step of standardizing the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to the data standardization layer includes: The decrypted vehicle message data is split according to a preset granularity type based on the data standardization layer. The data standardization layer is used to standardize the split vehicle message data, the decrypted vehicle bus data, and the vehicle attribute data. The standardized vehicle message data is pushed down to the message queue of the data clustering layer, and the standardized vehicle message data, standardized vehicle bus data, and standardized vehicle attribute data are stored in the data standardization layer.

5. The data warehouse data processing method as described in claim 4, characterized in that, The step of clustering the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data according to the data clustering layer to obtain the target vehicle data includes: Based on the data clustering layer, the standardized vehicle attribute data, standardized vehicle message data, and standardized vehicle bus data are clustered to obtain the target vehicle data; The target vehicle data is stored in the data clustering layer and then pushed down to the data processing layer.

6. The data warehouse data processing method as described in claim 5, characterized in that, The step of processing the target vehicle data based on the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators includes: Obtain the standardized vehicle message data, the standardized vehicle bus data, the standardized vehicle attribute data, the target vehicle data, and the data configuration table generated by the data table configuration layer; The standardized vehicle message data, standardized vehicle bus data, standardized vehicle attribute data, target vehicle data, and data configuration table generated by the data table configuration layer are processed according to the data indicator processing theme, type granularity, statistical period, and data statistical indicators to obtain vehicle data indicators.

7. The data warehouse data processing method according to any one of claims 1 to 6, characterized in that, After the step of processing the target vehicle data according to the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators, the method further includes: The expired vehicle message data and the vehicle message data to be synchronized in the real-time data warehouse are stored in the offline data warehouse; The vehicle summary index table from the offline data warehouse is stored in the real-time data warehouse.

8. A data warehouse data processing device, characterized in that, The data warehouse data processing device includes: a data processing module, which is applied to a data acquisition and scheduling platform, which includes a real-time data warehouse and an offline data warehouse; both the real-time data warehouse and the offline data warehouse include: a data decryption layer, a data standardization layer, a data table configuration layer, a data clustering layer, and a data processing layer. The data processing module is used to acquire vehicle attribute data and vehicle message data, and to process the vehicle attribute data and vehicle message data according to the data decryption layer. The vehicle message data is encrypted structured data reported according to the vehicle communication protocol. The data decryption layer is used to decrypt the data and store the decrypted data. The data processing module is also used to acquire vehicle bus data and process the vehicle bus data according to the data decryption layer, wherein the vehicle bus data is encrypted unstructured signal data. The data processing module is also used to standardize the processed vehicle attribute data, processed vehicle message data, and processed vehicle bus data according to the data standardization layer. The data processing module is also used to cluster the standardized vehicle attribute data, standardized vehicle message data and standardized vehicle bus data according to the data clustering layer to obtain target vehicle data. The data processing module is also used to process the target vehicle data according to the data configuration table generated by the data processing layer and the data table configuration layer to obtain vehicle data indicators; The forms stored in the real-time data warehouse and the offline data warehouse are divided into fact tables and dimension tables. The data table configuration layer filters the dimension tables through the generated data configuration table. The real-time data warehouse obtains vehicle attribute data based on a heterogeneous data architecture and vehicle message data based on a distributed architecture. It processes the vehicle attribute data and vehicle message data through a data decryption layer, a data standardization layer, a data clustering layer, and a data processing layer. After processing the vehicle message data, the data decryption layer, the data standardization layer, and the data clustering layer will be pushed down to the next layer. The offline data warehouse acquires vehicle bus data and real-time data warehouse synchronization data based on a heterogeneous data architecture, and processes vehicle attribute data and synchronization data through a data decryption layer, a data standardization layer, a data clustering layer, and a data processing layer.

9. A data warehouse data processing device, characterized in that, The data warehouse data processing device includes: a memory, a processor, and a data warehouse data processing program stored in the memory and executable on the processor. When the data warehouse data processing program is executed by the processor, it implements the steps of the data warehouse data processing method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a data warehouse data processing program, which, when executed by a processor, implements the steps of the data warehouse data processing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Warehouse counting system based on industrial cloud side service

    CN112380295A

  • Flow and batch integrated counting warehouse integration method and system

    CN115114266A

  • Bin counting system, data processing method and device, medium and equipment

    CN116303814A