System and method for building lithium battery industry data integration platform

By building a data integration platform for the lithium battery industry and using Flink-cdc and DataStream API for merging and storing multi-source data, the problems of data inconsistency and timeliness of data collection between information systems have been solved. This has enabled real-time data analysis and efficient decision support, and improved the collaboration of supply chain management and production efficiency.

CN116821101BActive Publication Date: 2025-10-21HEFEI GUOXUAN HIGH TECH POWER ENERGY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310947789.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-10-21
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

The lack of integration in the information systems of the lithium battery industry in the current technology leads to inconsistent data, information silos, low work efficiency and inaccurate decision-making, and the timeliness of data collection and the degree of standardization of storage are also low.

Method used

We build a data integration platform based on the lithium battery industry. We use Flink-cdc to collect data from different systems, convert the format, and then merge the data from multiple sources through the DataStream API. We use Kafka middleware for resource reuse and store and analyze data through a data warehouse layered module, including data model design and the construction of standard and process layers.

Benefits of technology

It enables real-time collection and analysis of data from different information systems, improves data utilization efficiency, supports data-driven business decisions, enhances supply chain transparency and collaboration, reduces costs, and improves production efficiency and decision-making accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821101B_ABST
    Figure CN116821101B_ABST
Patent Text Reader

Abstract

The application provides a lithium battery industry data set integration platform construction system and method, the system comprises: a data acquisition module collects industry data of different systems by using Flink-cdc and converts the format; a resource reuse module performs multi-source merging operation by using DataStream API and bus Kafka middleware to reuse resources of the multi-source merged data; a source library metadata reading and data migration module reads source library metadata in the multi-source merged data to generate a data migration configuration file; an olap construction module obtains ods layer data to construct an olap and transmit to a hive database; and an olap layering module performs olap layering operation on the olap according to the data migration configuration file, performs Flink real-time calculation on the layered industry data of different systems to obtain real-time display report and algorithm requirement data. The application solves the technical problems of poor data acquisition timeliness, low data storage standardization and standardization degree, information island, data inconsistency and inaccurate decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data management in the lithium battery manufacturing industry, and in particular to a system and method for constructing a data integration platform based on the lithium battery industry. Background Art

[0002] With the continuous deepening of digitalization and informatization, more and more information systems are involved within enterprises, including SRM systems, ERP systems, WMS systems, etc. These systems are responsible for management tasks in different fields, such as procurement management, enterprise resource management, warehouse management, and office automation. The existing invention patent application document with publication number CN111091261A, "Method for Full Life Cycle Management of Lithium Batteries", includes the following methods: the battery basic information system reports the basic information of the lithium battery to the data acquisition module, and the battery information reporting system reports the real-time information of the lithium battery to the data acquisition module; the data acquisition module sends the lithium battery information data to the data storage module to save the basic information and real-time information of the lithium battery, and filters the lithium battery information data before sending it to the lithium battery analysis module; the lithium battery analysis module sets the analysis parameters and analysis conditions of the lithium battery analysis, and analyzes the received lithium battery data information; the display module displays the analysis results completed by the lithium battery analysis module. The battery basic information system reports basic lithium battery information to the data collection module, including: battery serial number, battery model, battery production date, battery cell material number, customer information, design voltage, design capacity, number of battery cells in series, and number of battery cells in parallel. The basic lithium battery information data source is directly imported from the production ERP system, recording the battery information at the time of shipment. The existing invention patent application document, "Logistics System and Management Method for Automated Production Line of Square Lithium Iron Phosphate Batteries," with publication number CN116022487A, includes a finished cell warehouse, encapsulation stations, a pack module line, and an electrical control system. The finished cell warehouse includes a robot, conveyor lines, a buffer area, a stacking buffer area, a lifting and loading device, a stacker, a vertical warehouse, a WMS, and a WCS. Among them, the finished battery cell warehouse includes a manipulator, a conveyor line, a buffer area, a stacked tray buffer area, a lifting and transferring device, a stacker, a vertical warehouse, a WMS (warehouse management system) and a WCS (warehouse control system). The WMS is electrically connected to the manipulator, the conveyor line, the buffer area, the stacked tray buffer area, the lifting and transferring device, the in-and-out handling system and the vertical warehouse, and the WCS is electrically connected to the manipulator, the conveyor line, the buffer area, the stacked tray buffer area, the lifting and transferring device, the stacker and the vertical warehouse. However, the data of the information systems in these existing technologies exist independently and cannot realize information sharing and circulation, resulting in redundancy and duplication of internal enterprise data, and also hindering the value mining and application of data. Therefore, establishing a data integration platform has become an effective means to solve enterprise information islands and data islands.

[0003] In the lithium battery industry, SRM systems are primarily used for managing suppliers and purchase orders, ERP systems are primarily used for managing production, sales, and logistics, and WMS systems are primarily used for warehouse management. Therefore, building a data integration platform requires designing and implementing a technical solution tailored to the specific industry characteristics and business needs.

[0004] Currently, the lithium battery industry's information systems manage a large number of business processes, including production management, supply chain management, and customer relationship management. These businesses are managed by different companies or departments, and the data formats and quality vary. This leads to the following problems:

[0005] Data inconsistencies: Due to the lack of integration between different systems, data sharing and collaboration cannot be achieved, which can easily lead to data inconsistencies. For example, data inconsistencies between the supplier dimension information in the SRM system and the supplier dimension in the ERP system can lead to a mismatch between the supplier of purchase orders and the supplier of sales orders, affecting the execution of procurement plans.

[0006] Information silos: Lack of integration between different systems can lead to the formation of information silos, preventing interoperability between departments. For example, lack of integration between ERP production information and SRM procurement information can lead to poor coordination between production and procurement, hindering the implementation of production plans.

[0007] Inefficient work: Lack of integration between different systems leads to duplication of operations and manual processing, resulting in low work efficiency. For example, lack of integration between material information in the SRM system and the ERP system may result in the need to repeatedly enter the material BOM and material details into the corresponding system, wasting time and manpower.

[0008] Inaccurate decision-making: Lack of integration between different systems can lead to inaccurate and incomplete data used to support business decisions, impacting the quality and effectiveness of these decisions. For example, lack of integration between sales information in an ERP system and inventory information in a WMS system can lead to information asymmetry between sales and inventory, hindering sales strategy formulation and inventory control.

[0009] Meanwhile, the prior art also has the following problems:

[0010] Timeliness of data collection: When building a data integration platform, timely data collection is crucial. Data collection from SRM, ERP, and WMS systems must be timely. Inconsistent data collection times across systems—for example, when the data integration platform collects ERP production order information, but the WMS system's material delivery information has not yet been transmitted to the middle platform—can cause significant errors in inventory demand calculations, impacting data accuracy and integrity.

[0011] Data storage standardization: When data from different information systems are connected to the middle platform, there are various different libraries and different tables, as well as the business of different information systems. The data business logic is cumbersome and the data volume increases greatly. Therefore, careful consideration and reasonable selection are required in choosing a data storage framework.

[0012] Data standardization: Since different information systems are provided by different manufacturers, the definitions of data parameters vary. For example, the same Chinese meaning may be in different fields, the same field may have different types, the same type may have different units, etc.

[0013] In summary, existing technologies have poor timeliness in data collection, low degree of normalization and standardization of data storage, information silos, and inconsistent data, which lead to technical problems such as inaccurate decision-making. Summary of the Invention

[0014] The technical problem to be solved by the present invention is: how to solve the technical problems in the prior art of poor timeliness of data collection, low degree of normalization and standardization of data storage, the existence of information islands, inconsistent data, and inaccurate decision-making.

[0015] The present invention solves the above technical problems by adopting the following technical solutions: a system based on the lithium battery industry data integration platform is constructed, including:

[0016] The data collection module uses Flink-cdc to collect industry data from different systems and convert the formats of industry data from different systems;

[0017] The resource reuse module uses the DataStream API to merge multiple sources based on industry data from different systems. This generates merged data, imports it into the Kafka bus middleware, and performs Flink SQL merge operations to reuse the data. The resource reuse module is connected to the data acquisition module.

[0018] The source library metadata reading and data migration module is used to read the source library metadata in the multi-source merged data and generate a data migration configuration file based on it. The source library metadata reading and data migration module is connected to the resource reuse module;

[0019] The data warehouse construction module is used to obtain ODS layer data, build a data warehouse based on it, and transmit it to the Hive database;

[0020] The data warehouse stratification module is used to stratify the data warehouse according to the data migration configuration file. It performs Flink real-time calculations on the industry data of different stratified systems to obtain real-time display reports and algorithm requirements data. The data warehouse stratification module is connected to the data warehouse construction module and the source library metadata reading and data migration module. The data warehouse stratification module includes:

[0021] Business requirements definition module, used to organize data from different information systems to formulate data warehouse design plans and determine business requirements data;

[0022] The data model design module is used to design a data model based on the business processes and data relationships in the business requirement data. The data model includes data dimensions and indicator data. Based on the data model, data warehouse layer information is constructed to perform layered operations, resulting in the standard layer and process layer. The data model design module is connected to the business requirement definition module.

[0023] The data storage and hierarchical processing module is used to store industry data from different information systems in Hudi using the CopyOnWrite type, partition the industry data according to the insertion time of different information systems, and perform data configuration to deduplicate the entire partition.

[0024] This invention effectively integrates data from various sources, improves data utilization efficiency, and provides data-driven business decision support. It can help companies achieve transparency and collaboration in their supply chains, improve production efficiency, and reduce costs. For example, by integrating and unifying supplier and purchase order information in SRM and ERP systems with inventory and inbound and outbound information in the WMS system, comprehensive supply chain monitoring and collaborative management can be achieved. Furthermore, by analyzing and mining supply chain data, companies can achieve tasks such as demand forecasting and production plan optimization, thereby improving decision-making accuracy.

[0025] This invention stratifies data warehouses based on supply chain business needs. In addition to conventional stratification, it also constructs standard and process layers between different information systems to streamline business flows and commonality across these systems. Through a data integration platform, this invention enables real-time data collection and analysis across different information systems, enabling efficient and accurate execution of production plans, improving production efficiency and product quality.

[0026] In a more specific technical solution, the data acquisition module includes:

[0027] Database type and address confirmation module, used to determine the type and address of ERP business database, SRM business database and WMS business database;

[0028] The format conversion module is used to monitor the underlying logs of the source database through Flink-cdc, collect JSON format strings, and convert the JSON format strings into Debezium-json format data.

[0029] In a more specific technical solution, the data acquisition module also includes:

[0030] Table data synchronization module, used to monitor CRUD operations of business database tables and update data in the target database in real time;

[0031] The change details synchronization module is used to collect and synchronize the process details of specific business table indicators in the business database.

[0032] This invention monitors CRUD operations on business database tables and synchronizes data with the target database, ensuring consistent results between the source and target databases and real-time updates. Based on the synchronization details function, this invention records every data change in the dimension table and sets effective and expiration times. This allows you to pull the latest and historical status of a dimension based on the effective time.

[0033] In a more specific technical solution, the resource reuse module includes:

[0034] The API multi-merge module is used to perform multi-source merge operations based on Debezium-json format data using the DataStream API to obtain multi-source merged data;

[0035] The Kafka middleware import module is used to import the bus Kafka middleware and process multi-source merged data to reuse resources of the multi-source merged data. The Kafka middleware import module is connected to the API multi-merge module.

[0036] In a more specific technical solution, the Kafka middleware import module collects the field types of Oracle metadata to convert the data migration script. The data migration script carries information including: Kafka library table metadata information, Flink-SQL script, and Hudi connector information.

[0037] In a more specific technical solution, the resource reuse module also includes:

[0038] The dynamic flow splitting module uses Flink-stream to perform Flink dynamic flow splitting based on the table name of the Kafka library table to generate side flow splitting;

[0039] The metadata information reading module is used to divert the side stream and read the metadata information of the Kafka library table according to the consumer group ID. The metadata information reading module is connected to the dynamic diversion module.

[0040] The present invention addresses the problems of various different libraries and tables, as well as the services of different information systems, complicated data business logic, and large data volume increments when data from different information systems are accessed to the middle platform. The present invention uses the Datastream-api of flink-cdc to merge multiple sources of multiple libraries and tables into kafka, and subsequently uses flink-hudi to synchronize data, thereby reducing operational complexity and configuration equipment costs. The present invention adopts the function of dynamic diversion of flink. flink-stream generates side stream diversion according to the table name, and only one consumer group id is required to read kakfa. The present invention organizes and manages data elements according to a certain classification system, marks key information, and makes it easy to find and use. The present invention can realize real-time monitoring and interaction of supply chain information, improve supply chain collaboration capabilities, reduce inventory costs, reduce delays and losses, and improve supply chain efficiency and reliability.

[0041] In a more specific technical solution, the source database metadata reading and data migration module includes:

[0042] The collection table metadata information generation module is used to identify Oracle metadata when the synchronization information system log in the source database's underlying log is output to Kafka, and generate metadata information for the corresponding collection table. The metadata information for the corresponding collection table includes: table name, field name, field attributes, and Flink-SQL write HUDI configuration file;

[0043] The raw data table field processing module uses the Flink program to obtain the raw data table field names and field types based on the metadata information configuration file. The raw data table field processing module is connected to the collection table metadata information generation module.

[0044] The OutputTag side stream generation module is used to write Hudi scripts using Flink. According to the original data table field name, it traverses and generates N OutputTag side streams to store different table data. The OutputTag side stream generation module is connected to the original data table field processing module; the OutputTag side stream generation module writes Hudi scripts based on Flink, obtains table metadata information, and converts the data source format into TypeInformation <row>Format, using tableEnv's fromChangelogStream operator, TypeInformation <row>The format is converted to table type;

[0045] The different table data storage module is used to store different table data into a map set according to the key of the original data table field name and the value of the OutputTag side stream. The different table data storage module is connected to the original data table field processing module and the OutputTag side stream generation module;

[0046] The different table data diversion module is used to divert different table data to different OutputTag side streams according to the table names carried by different table data when Kafka flows in real time. It uses the preset real-time stream to generate a single table format side stream according to the table name. The different table data diversion module is connected to the different table data storage module.

[0047] This paper uses Flink data splitting to split Kafka real-time streams and then converts them into Flink-Hudi storage data. This ensures that Oracle's CDC and Kafka connection numbers are not affected by newly added libraries and tables. In building a data integration platform, this paper ensures the timeliness of data collection operations for SRM, ERP, and WMS systems, improving data accuracy and integrity.

[0048] In a more specific technical solution, the data warehouse building modules include:

[0049] Temporary table generation and storage module, used to generate temporary tables based on the createTemporaryView operator and store temporary tables in pre-set memory;

[0050] Store data in the Hudi data lake module and use a Flink script to write the Hudi script to store the data in the temporary table in the corresponding Hudi data lake.

[0051] The present invention aims at the situation that different information systems are provided by different manufacturers and have different data parameter definitions.

[0052] In a more specific technical solution, the data storage and tiered processing module uses the following logic to configure data to perform full partition deduplication:

[0053] index.bootstrap.enabled=true;

[0054] The hierarchical objects in the hierarchical information include: original layer ods, dimension layer dim, detail layer dwd, aggregation layer dws, standard layer standardFloor, process layer processLayer and application layer ads.

[0055] In a more specific technical solution, the method for building a data integration platform for the lithium battery industry includes:

[0056] S1. Use Flink-cdc to collect industry data from different systems and convert the formats of industry data from different systems.

[0057] S2. Use the DataStream API to merge multiple sources based on industry data from different systems to obtain the merged data. This data is then imported into the Kafka bus middleware and merged using Flink SQL to reuse the data.

[0058] S3. Read the source database metadata from the multi-source merged data and generate a data migration configuration file based on it.

[0059] S4. Obtain ODS layer data, build a data warehouse based on it, and transfer it to the Hive database;

[0060] S5. Based on the data migration configuration file, the data warehouse is layered and Flink real-time computing is performed on the industry data of different layered systems to obtain real-time display reports and algorithm required data. Step S5 includes:

[0061] S51. Organize data from different information systems to develop data warehouse design plans and determine business demand data;

[0062] S52. Design a data model based on the business processes and data relationships in the business demand data. The data model includes data dimensions and indicator data. Build data warehouse layer information based on the data model to perform layered operations, thereby constructing a standard layer and a process layer.

[0063] S53. Use the CopyOnWrite type to store industry data from different information systems in Hudi, partition them according to the insertion time of industry data from different information systems, and configure the data to deduplicate all partitions.

[0064] This invention offers the following advantages over existing technologies: It effectively integrates data from various sources, improves data utilization efficiency, and provides data-driven business decision support. It can help enterprises achieve transparency and collaboration in their supply chains, improve production efficiency, and reduce costs. For example, by integrating and unifying supplier and purchase order information in SRM and ERP systems with inventory and inbound and outbound information in WMS systems, comprehensive supply chain monitoring and collaborative management can be achieved. Furthermore, by analyzing and mining supply chain data, enterprises can achieve tasks such as demand forecasting and production plan optimization, thereby improving decision-making accuracy.

[0065] This invention stratifies data warehouses based on supply chain business needs. In addition to conventional stratification, it also constructs standard and process layers between different information systems to streamline business flows and commonality across these systems. Through a data integration platform, this invention enables real-time data collection and analysis across different information systems, enabling efficient and accurate execution of production plans, improving production efficiency and product quality.

[0066] This invention monitors CRUD operations on business database tables and synchronizes data with the target database, ensuring consistent results between the source and target databases and real-time updates. Based on the synchronization details function, this invention records every data change in the dimension table and sets effective and expiration times. This allows you to pull the latest and historical status of a dimension based on the effective time.

[0067] The present invention addresses the problems of various different libraries and tables, as well as the services of different information systems, complicated data business logic, and large data volume increments when data from different information systems are accessed to the middle platform. The present invention uses the Datastream-api of flink-cdc to merge multiple sources of multiple libraries and tables into kafka, and subsequently uses flink-hudi to synchronize data, thereby reducing operational complexity and configuration equipment costs. The present invention adopts the function of dynamic diversion of flink. flink-stream generates side stream diversion according to the table name, and only one consumer group id is required to read kakfa. The present invention organizes and manages data elements according to a certain classification system, marks key information, and makes it easy to find and use. The present invention can realize real-time monitoring and interaction of supply chain information, improve supply chain collaboration capabilities, reduce inventory costs, reduce delays and losses, and improve supply chain efficiency and reliability.

[0068] This paper uses Flink data splitting to split Kafka real-time streams and then converts them into Flink-Hudi storage data. This ensures that Oracle's CDC and Kafka connection numbers are not affected by newly added libraries and tables. In building a data integration platform, this paper ensures the timeliness of data collection operations for SRM, ERP, and WMS systems, improving data accuracy and integrity.

[0069] The present invention solves the technical problems in the prior art of poor timeliness of data collection, low degree of normalization and standardization of data storage, the existence of information islands, inconsistent data, and inaccurate decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 This is a schematic diagram showing the connection of basic modules of a system for building a lithium battery industry data integration platform according to Example 1 of the present invention;

[0071] Figure 2 This is a schematic diagram of a system architecture based on a lithium battery industry data integration platform according to Example 1 of the present invention;

[0072] Figure 3 This is a schematic diagram of the steps of a method for constructing a data integration platform for the lithium battery industry according to Example 1 of the present invention;

[0073] Figure 4 This is a data flow processing diagram of the method for building a data integration platform for the lithium battery industry according to Example 1 of the present invention. DETAILED DESCRIPTION

[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0075] Example 1

[0076] like Figure 1 As shown, the system for building a data integration platform based on the lithium battery industry provided by the present invention includes: a data acquisition module 1, a resource reuse module 2, a source library metadata reading and data migration module 3, a data warehouse construction module 4 and a data warehouse stratification module 5.

[0077] Data acquisition module 1;

[0078] In this example, we first determine the type and address of the ERP / SRM / WMS business database (the information system database in this solution is Oracle). Because the information system requires high real-time performance, offline synchronization solutions are not suitable. The general real-time acquisition architecture, Flink CDC, cannot implement multi-source merging across multiple databases and tables, nor can it synchronize updates to downstream Kafka after multi-source merging. Currently, Flink SQL can only perform single-table Flink CDC operations, which results in an excessive number of database CDC connections. In this example, the information system databases used include, but are not limited to, Oracle.

[0079] like Figure 2 As shown, in this embodiment, the Figure 2 The middle architecture mainly uses Flink-cdc to monitor the underlying logs of the source database to avoid the pressure on the database caused by JDBC direct connection. The JSON format strings defined in the collected schema are converted into Debezium-json format.

[0080] In this embodiment, the resource reuse module 2 uses the DataStream API to merge multiple sources and then imports a bus Kafka middleware, which can realize the multi-source merging problem of Flink SQL and resource reuse. Write code to convert the data migration script by collecting the oracle metadata field type. The script carries the metadata information of the kafka library table and the Flink-sql script and hudi connector information. When Flink consumes kafka, since a topic of kafka stores information of many tables, if Flink-sql is used to read kafka directly, it will cause many consumer group ids to consume a topic at the same time (a new consumer group id needs to be created for each table), resulting in repeated consumption and waste of consumer threads. Therefore, the dynamic diversion function of Flink is adopted. Flink-stream generates side stream diversion according to the table name, and only one consumer group id is needed to read kafka.

[0081] In this embodiment, the source library metadata reading and data migration module 3, when the synchronization information system log is output to kafka, identifies the oracle metadata to generate metadata information of the corresponding collection table, including table name + field name + field attributes and Flink-sql writes the hudi configuration file. The Flink program first obtains the original data table field name + field type according to the configuration file (parameter 1: subsequent execution stream conversion sql), Flink writes the hudi script (parameter 2: subsequent execution Flink-sql into hudi), and according to the table name in parameter 1, traverses and generates N OutputTag side streams (used to store data from different tables), and stores them in the map collection according to the table name key and side stream value. When the kafka real-time stream comes in, the data is diverted to different side streams according to the table name carried by the data. In this way, the real-time stream will generate many side streams according to the table name. Each side stream has only one table data format. According to parameter 2, we get the metadata information of this table and convert the data source format into TypeInformation <row>format, because after Flink parses the Debezium-json log file into a steam stream, it will compare the ROW type, and the fromChangelogStream operator of tableEnv converts the ROW type into the table type.

[0082] In data warehouse construction module 4, a temporary table is generated based on the createTemporaryView operator and stored in memory. Finally, the script with parameter 2 is executed to store the data in the temporary table into the corresponding Hudi.

[0083] In this embodiment, the data collection module 1 includes: a table data synchronization module 11 and a change detail synchronization module 12 .

[0084] In this embodiment, the table data synchronization module 11 mainly monitors the CRUD operations of the business database table, synchronizes the data in the target database, ensures that the results of the source database and the target database are consistent, and are updated in real time.

[0085] In this embodiment, in the change detail synchronization module 12, some result data about inventory or price will be continuously updated in the business database, but the final data is retained. The analysis of some process indicators is not intuitive, and only the final results are obtained after synchronizing the table. In order to collect the process details of some business table indicators, a synchronization detail data function is developed.

[0086] In this example, Flink-cdc collects data results marked with C / U / D / R. The primary key update operation deletes and adds data to the target table. When the Flink-cdc DataStream-API implements the DebeziumDeserializationSchema class, all operations are changed to insert operations based on the type type.

[0087] In this embodiment, when the original type is new, the original json string after data is taken out, and the after data effective time field (current time) + data expiration time field (maximum timestamp) + data operation type field (new) are added.

[0088] In this embodiment, when the original type is modification, the original JSON string "after" and "before" data are extracted, and two new data items are added: "before" data effective time field (empty) + data expiration time field (current time) + data operation type field (delete) / "after" data effective time field (current time) + data expiration time field (maximum timestamp) + data operation type field (add). These are used to mark which operation the process data performs and record the details.

[0089] In this embodiment, when the original type is deletion, the original json string before data is taken out, and a new before data effective time field (empty) + data expiration time field (current time) + data operation type field (newly added) is added.

[0090] In this embodiment, the result table is newly added with a data effective time field, a data expiration time field, and a data operation type field, which are used to mark which operation the process data performs and record the details.

[0091] In this embodiment, the data warehouse layering module 5 includes: a business requirement definition module 51 , a data model design module 52 and a data storage module 53 .

[0092] In this embodiment, the business requirements definition module 51 first needs to understand the business needs of the lithium battery industry and the details of information system sources. This involves organizing data from various information systems, such as those related to production management, supply chain management, sales management, and financial management. To determine the data required for the data warehouse, it is necessary to collaborate with the business team to understand business needs and goals and develop a data warehouse design plan.

[0093] In this embodiment, once the data model design module 52 has determined the business requirements, it then designs the data model. The data model should be designed based on business processes and data relationships, and include appropriate dimensions and metrics. For example, in the lithium battery production process, key dimensions might include production date, production line, and product specifications, while key metrics might include output, quality, and production efficiency. Once the data model is defined, data warehouse layer information is constructed based on the model. After performing conventional layering, the standard layer and process layer are then comprehensively constructed.

[0094] In this embodiment, the standard layer: based on the data dictionary of the business configuration, the fields of the same business in different systems / different manufacturers can be standardized, and the same business meaning in different tables of different information systems can be found, and the differences can be found to sort out the standard table.

[0095] In this embodiment, at the process level, different information system tables correspond to different business meanings. Relationship fields are mined, and process diagrams for different system business indicators are organized to map the data flow relationships between different tables. (For example: SRM system procurement demand table -> SRM system procurement table -> ERP procurement information flow table -> ERP procurement incoming information table -> WMS incoming information table -> WMS outgoing information table -> ERP work order material usage table).

[0096] In this embodiment, the data storage and hierarchical processing module 53 uses the CopyOnWrite type to store data in hudi, partitions the data according to the data insertion time, configures index.bootstrap.enabled = true, and ensures that all partitions of the data are deduplicated. The data warehouse design specification uses the star model, snowflake model, and constellation model as templates, and the data warehouse table naming specification is subject domain + hierarchical information + business type + (time).

[0097] In this embodiment, the layered objects include but are not limited to: original layer (ods), dimension layer (dim), detail layer (dwd), aggregation layer (dws), standard layer (standardFloor), process layer (processLayer), and application layer (ads). The data warehouse layer design details are as follows:

[0098] In this example, the ODS layer synchronizes various original business tables with the information system database through the Flink program. These tables use data warehouse naming conventions and are partitioned by write time. These tables primarily contain header and detail information for purchase, sales, inventory, and work orders.

[0099] In this embodiment, the dim layer synchronizes some dimension information in the business system, including supplier information, material information, address information, subsidiary information, production line information, and so on. During the synchronization process, the data acquisition function for synchronizing change details is required. Because after many dimensions are updated, the historical status no longer exists, and the dimension information in the historical wide table cannot be found in the latest dimension table. Therefore, according to the synchronization details function, each data change in the dimension table is recorded, and the effective time and expiration time are set. In this way, we can pull the latest and historical status of a dimension according to the effective time.

[0100] In this embodiment, the DWD layer mainly joins the header table and detail table of the information system, and associates dimensional information such as materials, suppliers, and purchasing companies. For example, the production work order detail table (including dimensional information such as production line and materials) and the purchase demand detail table (including dimensional information such as purchaser, purchasing company, supplier, and materials)

[0101] In this embodiment, the DWS layer: performs light aggregation of key indicators according to dimensional information, for example: calculates business indicators based on hourly / daily / other time intervals or supplier dimensions, groups the DWS layer wide table dimensions, and aggregates the indicators.

[0102] In this embodiment, the SF layer: the standard layer function is to define standardized specifications: first, it is necessary to define a standardized specification, including the structure of the table, column name, data type, data format and other aspects. This specification will become a standardized template for adjusting and merging various tables. This part requires manual editing of the data dictionary as a specification. Then compare the differences between the tables: compare the differences between the tables, and list the commonalities and differences between them. Example: There is a material information dimension in both the SRM system and the ERP system, but the two tables are composed of their own business personnel dimensions. The two tables have different numbers of fields, but different fields have the same meaning. Each table contains its own unique fields, and the total number of maintained materials is also different. When data from different information systems are aggregated to the platform, it is difficult for us to choose which dimension table.

[0103] In this embodiment, dimension tables with the same meaning in different systems are imported into the data warehouse. This allows the metadata information of each of the two tables to be obtained, and the original layer data dictionary (input table, two dimension tables in the original layer) and the standard layer data dictionary (output table, standard layer dimension table) are designed. The original layer dictionary is directly generated from the metadata information of the original table, while the standard layer dictionary requires manual design of data element definitions: including table name, definition, data type, length, valid values, description, and other information. Data element relationships: including the relationships between data elements, such as parent-child relationships, dependency relationships, reference relationships, etc. Data element classification: data elements are organized and managed according to a certain classification system, and key information is marked to make them easy to find and use.

[0104] Based on the edited data dictionary, identify the standard fields corresponding to the same meaning but different fields in the two tables. Join them based on the material's primary key, preserving the unique fields of the two tables. For example, the material name, specification, and manufacturer information can be used as the material dimension primary key. Different dimension tables are linked based on the logical primary key. The largest set of material types is selected. Useful fields are extracted based on business requirements, and redundant fields are removed. This creates a standard material dimension table.

[0105] In this embodiment, the PL layer: the process layer is mainly used to design the relationship between different supply chain systems.

[0106] For example, copper foil, a core material in lithium battery production, is in high demand and used in large quantities. Furthermore, as copper prices are a commodity, price fluctuations and demand analysis are even more critical. Therefore, a full-process information table for copper foil is necessary. In the standard-layer dimension table, the material code of copper foil is defined as the primary key for different information systems. The key is to use a unified time interval across all information systems. Table A cannot use the shipping time and table B to use the arrival time. Instead, we use the copper foil arrival time as the only time dimension. First, we obtain the purchase demand table from the SRM system to obtain the purchase demand and actual purchase quantity of copper foil. Then, based on the material logistics information from the ERP system, we retrieve the in-transit quantity of copper foil. From the production work order, we retrieve the line-side inventory and the total usage on the work order. From the WMS system, we obtain the incoming and outgoing quantities of copper foil, the actual inventory, and the safety stock (values ​​calculated and filled in manually). Using the material code and arrival time dimension information, we can unify the different indicators from different systems into a single table, analyze each link in the supply chain, and identify supply chain optimization solutions.

[0107] In this embodiment, the ads layer: calculates results based on business indicators.

[0108] In this embodiment, unit product cost refers to the production cost of each product, which is calculated as follows: production cost ÷ production quantity.

[0109] In this embodiment, equipment utilization refers to the ratio of the actual use time of the production equipment to the total available time. The calculation formula is:

[0110] Equipment utilization rate = actual equipment usage time ÷ total equipment available time × 100%.

[0111] In this embodiment, the production capacity utilization rate refers to the ratio of the enterprise's actual production capacity to its maximum production capacity. The calculation formula is:

[0112] Production capacity utilization rate = actual production capacity ÷ maximum production capacity × 100%.

[0113] Example 2

[0114] like Figure 3 and Figure 4 As shown, in this embodiment, the method for building a data integration platform based on the lithium battery industry includes:

[0115] S1. Use Flink-cdc to collect industry data from different systems and convert the format of the industry data.

[0116] In this embodiment, supplier management data and purchase order data are obtained from the SRM system, production, sales, and logistics management data are obtained from the ERP system, and warehouse management data is obtained from the WMS system. Logs are collected and converted into Debezlum format through Flink-cdc.

[0117] S2. Use the DataStream API to merge multiple sources and then import them into a bus Kafka middleware to perform multi-source merging in FlinkSQL to achieve resource reuse.

[0118] S3: Read the source database metadata and generate a data migration configuration file.

[0119] In this embodiment, the read files include the Flink program + migration configuration file;

[0120] S4. Import the ODS layer data into the Hudi data lake, build a data warehouse and transfer it to the Hive database;

[0121] S5: Perform data warehouse stratification operations and Flink real-time computing to obtain real-time display reports and algorithm requirement data.

[0122] In summary, the present invention effectively integrates data from various sources, improves data utilization efficiency, and provides data-driven business decision support. This can help enterprises achieve transparency and collaboration in their supply chains, improve production efficiency, and reduce costs. For example, by integrating and unifying supplier and purchase order information in SRM and ERP systems with inventory and inbound and outbound information in the WMS system, comprehensive monitoring and collaborative management of the supply chain can be achieved. Furthermore, by analyzing and mining supply chain data, enterprises can achieve tasks such as demand forecasting and production plan optimization, which helps improve decision-making accuracy.

[0123] This invention stratifies data warehouses based on supply chain business needs. In addition to conventional stratification, it also constructs standard and process layers between different information systems to streamline business flows and commonality across these systems. Through a data integration platform, this invention enables real-time data collection and analysis across different information systems, enabling efficient and accurate execution of production plans, improving production efficiency and product quality.

[0124] This invention monitors CRUD operations on business database tables and synchronizes data with the target database, ensuring consistent results between the source and target databases and real-time updates. Based on the synchronization details function, this invention records every data change in the dimension table and sets effective and expiration times. This allows you to pull the latest and historical status of a dimension based on the effective time.

[0125] The present invention addresses the problems of various different libraries and tables, as well as the services of different information systems, complicated data business logic, and large data volume increments when data from different information systems are accessed to the middle platform. The present invention uses the Datastream-api of flink-cdc to merge multiple sources of multiple libraries and tables into kafka, and subsequently uses flink-hudi to synchronize data, thereby reducing operational complexity and configuration equipment costs. The present invention adopts the function of dynamic diversion of flink. flink-stream generates side stream diversion according to the table name, and only one consumer group id is required to read kakfa. The present invention organizes and manages data elements according to a certain classification system, marks key information, and makes it easy to find and use. The present invention can realize real-time monitoring and interaction of supply chain information, improve supply chain collaboration capabilities, reduce inventory costs, reduce delays and losses, and improve supply chain efficiency and reliability.

[0126] This paper uses Flink data splitting to split Kafka real-time streams and then converts them into Flink-Hudi storage data. This ensures that Oracle's CDC and Kafka connection numbers are not affected by newly added libraries and tables. In building a data integration platform, this paper ensures the timeliness of data collection operations for SRM, ERP, and WMS systems, improving data accuracy and integrity.

[0127] The present invention solves the technical problems in the prior art of poor timeliness of data collection, low degree of normalization and standardization of data storage, the existence of information islands, inconsistent data, and inaccurate decision-making.

[0128] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.< / row> < / row> < / row>

Claims

1. Build a system based on the lithium battery industry data integration platform, characterized by: The system comprises: The data collection module is used to use Flink-cdc to collect industry data from different systems and convert the formats of the industry data from different systems; A resource reuse module is configured to perform a multi-source merging operation using the DataStream API based on the industry data of the different systems to obtain multi-source merged data, import the data into the bus Kafka middleware, and perform the multi-source merging operation of Flink SQL to reuse the multi-source merged data. The resource reuse module is connected to the data acquisition module. A source library metadata reading and data migration module is used to read the source library metadata in the multi-source merged data and generate a data migration configuration file based on the source library metadata. The source library metadata reading and data migration module is connected to the resource reuse module; The data warehouse construction module is used to obtain ODS layer data, build a data warehouse based on it, and transmit it to the Hive database; The data warehouse stratification module is used to perform data warehouse stratification operations on the data warehouse according to the data migration configuration file, and perform Flink real-time calculations on the stratified industry data of different systems to obtain real-time display reports and algorithm requirement data. The data warehouse stratification module is connected to the data warehouse construction module and the source library metadata reading and data migration module. The data warehouse stratification module includes: Business requirements definition module, used to organize data from different information systems to formulate data warehouse design plans and determine business requirements data; A data model design module is used to design a data model based on the business processes and data relationships in the business requirements data. The data model includes data dimensions and indicator data. Data warehouse layer information is constructed based on the data model to perform layered operations to obtain the standard layer and process layer. The data model design module is connected to the business requirements definition module. The data storage and hierarchical processing module is used to store the industry data of the different information systems in hudi using the CopyOnWrite type, partition the industry data according to the insertion time of the industry data of the different information systems, and perform data configuration to deduplicate the entire partition of the data.

2. The system for building a data integration platform for the lithium battery industry according to claim 1 is characterized in that: The data acquisition module includes: Database type and address confirmation module, used to determine the type and address of ERP business database, SRM business database and WMS business database; The format conversion module is used to monitor the underlying logs of the source database through Flink-cdc, collect JSON format strings, and convert the JSON format strings into Debezium-json format data.

3. The system for building a data integration platform for the lithium battery industry according to claim 1 is characterized in that: The data acquisition module also includes: Table data synchronization module, used to monitor CRUD operations of business database tables and update data in the target database in real time; The change details synchronization module is used to collect and synchronize the process details of specific business table indicators in the business database.

4. The system for building a data integration platform for the lithium battery industry according to claim 1 is characterized in that: The resource reuse module includes: An API multi-merge module is used to perform a multi-source merge operation using the DataStream API based on the Debezium-json format data to obtain the multi-source merged data; A Kafka middleware import module is used to import the bus Kafka middleware and process the multi-source merged data to perform resource reuse on the multi-source merged data. The Kafka middleware import module is connected to the API multi-merge module.

5. The system for building a data integration platform for the lithium battery industry according to claim 4 is characterized in that: In the Kafka middleware import module, the field types of the Oracle metadata are collected to convert the data migration script, wherein the information carried by the data migration script includes: Kafka library table metadata information, Flink-SQL script and Hudi connector information.

6. The system for building a data integration platform for the lithium battery industry according to claim 1 is characterized in that: The resource reuse module also includes: The dynamic flow splitting module uses Flink-stream to perform Flink dynamic flow splitting based on the table name of the Kafka library table to generate side flow splitting; The metadata information reading module is used to divert the side stream and read the metadata information of the Kafka library table according to the consumer group ID. The metadata information reading module is connected to the dynamic diversion module.

7. The system for building a data integration platform for the lithium battery industry according to claim 1 is characterized in that: The source library metadata reading and data migration module includes: The collection table metadata information generation module is used to identify Oracle metadata when the synchronization information system log in the source library's underlying log is output to Kafka, and generate metadata information for the corresponding collection table, where the metadata information for the corresponding collection table includes: table name, field name, field attributes, and Flink-SQL write hudi configuration file; The original data table field processing module is used to obtain the original data table field name and field type according to the configuration file of the metadata information using the Flink program. The original data table field processing module is connected to the collection table metadata information generation module; The OutputTag side stream generation module is used to write the Hudi script using Flink. According to the original data table field name, it traverses and generates N OutputTag side streams to store different table data. The OutputTag side stream generation module is connected to the original data table field processing module. The OutputTag side stream generation module writes the Hudi script according to Flink, obtains the table metadata information, and converts the data source format into TypeInformation <row>format, using tableEnv's fromChangelogStream operator to convert the TypeInformation <row> The format is converted to table type;< / row> < / row> A different table data storage module is used to store the different table data into a map set according to the key of the original data table field name and the value of the OutputTag side stream. The different table data storage module is connected to the original data table field processing module and the OutputTag side stream generation module; The different table data diversion module is used to divert the different table data to different OutputTag side streams according to the table names carried by the different table data when the Kafka flows in real time, and use the preset real-time stream to generate a single table format side stream according to the table name. The different table data diversion module is connected to the different table data storage module.

8. The system for building a data integration platform for the lithium battery industry according to claim 1 is characterized in that: The data warehouse construction module includes: A temporary table generation storage module is used to generate a temporary table according to the createTemporaryView operator and store the temporary table in a preset memory; Store the data in the Hudi data lake module and use a Flink script to write the Hudi script to store the data in the temporary table in the corresponding Hudi data lake.

9. The system for building a data integration platform for the lithium battery industry according to claim 1, characterized in that: In the data storage and layer processing module, the following logic is used to configure data to perform full partition deduplication: index.bootstrap.enabled=true; The hierarchical objects in the hierarchical information include: original layer ods, dimension layer dim, detail layer dwd, aggregation layer dws, standard layer standardFloor, process layer processLayer and application layer ads.

10. A method for constructing a data integration platform for the lithium battery industry, characterized in that: The method comprises: S1. Use Flink-cdc to collect industry data from different systems and convert the formats of the industry data from different systems; S2. Use the DataStream API to perform a multi-source merge operation based on the industry data from different systems to obtain multi-source merged data, import the data into the bus Kafka middleware, and perform the multi-source merge operation of Flink SQL to reuse the multi-source merged data. S3. Read source library metadata from the multi-source merged data and generate a data migration configuration file based on the metadata. S4. Obtain ODS layer data, build a data warehouse based on it, and transfer it to the Hive database; S5. Perform data warehouse stratification on the data warehouse according to the data migration configuration file, and perform Flink real-time calculation on the stratified industry data of different systems to obtain real-time display reports and algorithm requirement data. Step S5 includes: S51. Organize data from different information systems to develop data warehouse design plans and determine business demand data; S52. Design a data model based on the business processes and data relationships in the business demand data. The data model includes data dimensions and indicator data. Build data warehouse layer information based on the data model to perform layered operations to obtain a standard layer and a process layer. S53: Use the CopyOnWrite type to store the industry data of the different information systems in Hudi, partition the data according to the insertion time of the industry data of the different information systems, and configure the data to deduplicate the entire partition.

Citation Information

Patent Citations

  • Lithium battery full life cycle management method

    CN111091261A

  • Logistics system and management method for automatic production line of square lithium iron phosphate batteries

    CN116022487A

  • Game engine data lake hierarchical extensible design and stream batch combination ad hoc analysis system

    CN115794768A

  • Operation system based on industrial internet platform

    CN116301825A