Battery data processing method and device
By setting up an inverted index structure in the distributed columnar storage database system, the problems of low maintenance efficiency and query delay in existing battery data processing methods are solved, and efficient data query and maintenance are achieved.
Patent Information
- Application Number
- CN202510357181.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
The existing battery data processing method uses dataStream to implement the logic of flink, resulting in low subsequent maintenance efficiency, high development difficulty, and the problems of delay and low query efficiency for pulling data from mysql.
Correlate battery behavior data and battery dimension data, store it in a distributed columnar storage database system, and set up an inverted index data structure in the system to achieve efficient query and maintenance.
The data query and maintenance efficiency have been improved, and the query speed and subsequent data processing efficiency have been significantly improved through the inverted indexing technology.
Smart Images

Figure CN120256431A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of battery management, and particularly to a method and device for processing battery data. Background Art
[0002] Currently, the processing of battery data usually uses dataStream to implement the logic of Flink, and at the same time, general dimension tables are placed in MySQL. Dimension table data is simply retrieved from MySQL, and finally, wide tables are placed in MySQL or Hive.
[0003] However, using dataStream has problems such as low subsequent maintenance efficiency and high development difficulty. There is also a delay in retrieving data from MySQL, and putting the results in MySQL results in low query efficiency. Summary of the Invention
[0004] Embodiments of this application provide a method and device for processing battery data to achieve the processing of battery data and improve data query and maintenance efficiency.
[0005] To solve the above technical problems, embodiments of this application disclose the following technical solutions:
[0006] In a first aspect, a method for processing battery data is provided, including:
[0007] Obtain battery behavior data and battery dimension data;
[0008] Associate the battery behavior data and the battery dimension data to obtain a first table;
[0009] Store the first table in a distributed columnar storage database system;
[0010] Set up an inverted index data structure in the distributed columnar storage database system to query the data in the first table.
[0011] In combination with the first aspect, setting up the inverted index data structure in the distributed columnar storage database system includes:
[0012] Write an inverted index file while writing a data file in the writing stage. For each row number of the written data, it corresponds one-to-one with the row number in the inverted index;
[0013] In the query stage, if the query condition contains columns for which an inverted index has been established, the distributed columnar storage database system will automatically query the index file, return a list of row numbers that meet the conditions, and then use the common row number filtering mechanism in the distributed columnar storage database system to skip unnecessary rows and pages and only read the rows that meet the preset conditions.
[0014] Combined with the first aspect, the obtaining of the battery behavior data and the battery dimension data includes:
[0015] Obtain the battery behavior data from the message middleware;
[0016] Obtain the battery dimension data from the storage medium.
[0017] Combined with the first aspect, before associating the battery behavior data and the battery dimension data, it includes:
[0018] Store the battery behavior data into a pre-created user log table;
[0019] Create a user profile index table and a battery static information index table;
[0020] Obtain the first battery dimension data from the storage medium according to the user profile index table and store it into a pre-created user profile table;
[0021] Obtain the second battery dimension data from the storage medium according to the battery static information index table and store it into a pre-constructed battery static information table.
[0022] Combined with the first aspect, the associating of the battery behavior data and the battery dimension data to obtain a first table includes:
[0023] Create a table execution environment;
[0024] Left-join the user log table, the user profile index table, the user profile table, the battery static information index table, and the battery static information table according to the table execution environment to form the first table.
[0025] Combined with the first aspect, the first table at least includes: vehicle identification code, timestamp, state of charge of the battery, current, maximum temperature, minimum temperature, temperature difference, maximum cell voltage, minimum cell voltage, maximum cell voltage number, minimum cell voltage number, battery pack number, vehicle model, mileage, historical disposal information id, historical warning information, historical disposal information content, and pressure difference growth rate.
[0026] Combined with the first aspect, before obtaining the battery dimension data, it further includes:
[0027] Create Guava to cache the queried battery dimension data.
[0028] Combined with the first aspect, the battery data processing method further includes: saving the data in the first table into a cache system.
[0029] In combination with the first aspect, the battery behavior data at least includes: vehicle identification code, state of charge of the battery, cell voltage, insulation resistance value, mileage, vehicle speed, timestamp, highest cell voltage, lowest cell voltage, highest cell voltage number, and lowest cell voltage number.
[0030] In a second aspect, a battery data processing device is provided, including:
[0031] An acquisition module, configured to acquire battery behavior data and battery dimension data;
[0032] An association module, configured to associate the battery behavior data and the battery dimension data through a message middleware and a storage medium to obtain a first table;
[0033] A storage module, configured to store the first table in a distributed columnar storage database system;
[0034] A query module, configured to set an inverted index data structure in the distributed columnar storage database system to query the data in the first table.
[0035] One of the above technical solutions has the following advantages or beneficial effects:
[0036] Compared with the prior art, a battery data processing method of the present application includes: acquiring battery behavior data and battery dimension data; associating the battery behavior data and the battery dimension data to obtain a first table; storing the first table in a distributed columnar storage database system; setting an inverted index data structure in the distributed columnar storage database system to query the data in the first table. The battery data processing method provided by the present application can improve data query efficiency and subsequent maintenance efficiency by associating battery behavior data and battery dimension data, putting the associated data into DorisDB, and setting an inverted index data structure in DorisDB to query or calculate the associated data.
[0037] A battery data processing device of the present application can improve data query efficiency and subsequent maintenance efficiency by associating battery behavior data and battery dimension data, putting the associated data into DorisDB, and setting an inverted index data structure in DorisDB to query or calculate the associated data. Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0039] Figure 1 It is a flowchart of a battery data processing method provided in an embodiment of the present application;
[0040] Figure 2 It is a schematic diagram of an internal inverted index file provided in an embodiment of the present application;
[0041] Figure 3 It is a schematic diagram of an inverted index provided in an embodiment of the present application;
[0042] Figure 4 It is a schematic flowchart of battery data processing provided in an embodiment of the present application;
[0043] Figure 5 It is a principle block diagram of a battery data processing device provided in an embodiment of the present application.
[0044] Reference numerals:
[0045] 100 - Battery data processing device; 101 - Acquisition module; 102 - Association module; 103 - Storage module; 104 - Query module. Specific embodiments
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0047] In the description of the present application, it should be understood that in the description of the present application, the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present application, "a plurality" means two or more, and at least one means one, two or more, unless otherwise specifically defined.
[0048] Figure 1 It is a flowchart of a battery data processing method provided in an embodiment of the present application. This method is applicable to the process of efficiently querying and maintaining battery data in a battery management platform. This method can be executed by a battery data processing device, which can be implemented in a software and / or hardware manner, and the device can be configured in the processor of the battery management platform. Please refer to Figure 1 and the method includes the following steps:
[0049] Step 110: Obtain battery behavior data and battery dimension data.
[0050] Among them, the battery behavior data can be obtained from a message middleware (such as ksfka).
[0051] Among them, the battery behavior data includes field information such as Vehicle Identification Number (VIN), State of Charge (SOC), cell voltage, insulation resistance value, mileage, vehicle speed, timestamp, maximum cell voltage, minimum cell voltage, maximum cell voltage number, and minimum cell voltage number.
[0052] Among them, the battery dimension data can be obtained from a storage medium (such as hbase).
[0053] Among them, the battery dimension data may include information such as province, city, district, reason for repair, number of repairs, whether it has been disposed of, reason for disposal, historical disposal information id, historical disposal information content, battery pack id, battery supplier, battery series-parallel number, vehicle model, battery capacity, etc.
[0054] Step 120: Associate the battery behavior data and the battery dimension data to obtain a first table.
[0055] Among them, the first table is a wide table. The data in the first table includes: vehicle identification number, timestamp, state of charge of the battery, current, maximum temperature, minimum temperature, temperature difference, maximum cell voltage, minimum cell voltage, maximum cell voltage number, minimum cell voltage number, battery pack number, vehicle model, mileage, historical disposal information id, historical warning information, content of historical disposal information, and differential pressure growth rate, etc.
[0056] Step 130: Store the first table into the distributed columnar storage database system.
[0057] Among them, the distributed columnar storage database system (DorisDB) adopts a columnar storage method, stores the data in the first table by columns to improve data query performance and compression ratio. And it supports complex query operations, such as aggregation, join, and window functions, etc., as well as data import and export. DorisDB also provides a mechanism for real-time data synchronization and data consistency guarantee to support applications such as real-time data analysis and report generation. Therefore, associate the battery behavior data and battery dimension data, and store the associated data into DorisDB to facilitate subsequent query and maintenance and improve the efficiency of data query and maintenance.
[0058] Step 140: Set up an inverted index data structure in the distributed columnar storage database system to query the data in the first table.
[0059] Among them, setting up an inverted index data structure in DorisDB is beneficial to improving the query efficiency of the associated data.
[0060] In the technical solution of this embodiment, the working principle of this battery data processing method is as follows: Refer to Figure 1 , first, obtain the battery behavior data and battery dimension data. Then, associate the battery behavior data and battery dimension data to obtain the first table. Secondly, store the first table into the distributed columnar storage database system. Finally, set up an inverted index data structure in the distributed columnar storage database system to query the data in the first table. It can be seen that by associating the battery behavior data and battery dimension data, putting the associated data into DorisDB at the same time, and setting up an inverted index data structure in DorisDB to query or calculate the associated data, the data query efficiency and subsequent maintenance efficiency can be improved.
[0061] In some embodiments, an inverted index data structure is set up in a distributed columnar storage database system, including: during the write phase, while writing a data file, an inverted index file is written at the same time. For each row number of the written data, it corresponds one by one to the row number in the inverted index; during the query phase, if the query condition contains a column for which an inverted index has been established, the distributed columnar storage database system will automatically query the index file, return a list of row numbers that meet the conditions, and then use the row number filtering mechanism common in the distributed columnar storage database system to skip unnecessary rows and pages and only read the rows that meet the preset conditions.
[0062] Figure 2 It is a schematic diagram of an internal inverted index file provided in an embodiment of the present application. Specifically, an inverted index is built into the DorisDB database. In DorisDB, from a logical perspective, the inverted index is applied at the column level of the table, while from the perspective of physical storage and implementation, the inverted index is actually built at the data file level. Please refer to Figure 2 , in the write phase: while writing a data file (segment file), an inverted index file (invertedindex file) is synchronously written. For each row number of the written data, it corresponds one by one to the row number in the inverted index. In the query phase: if the query WHERE condition contains a column for which an inverted index has been established, DorisDB will automatically query the index file, return a list of row numbers that meet the conditions, and then use the common row number filtering mechanism of DorisDB to skip unnecessary rows and pages and only read the rows that meet the conditions to achieve the effect of query acceleration.
[0063] Refer to Figure 2 , Header represents the header; Dala Region represents the data region; Index Region represents the index region; Footer represents the footer. Among them, the Dala Region contains multiple data pages, such as Page,..., Page, and each page contains several data records. Among them, the size of the page is fixed, supporting fast block reading. Among them, the data is stored in sequence.
[0064] Among them, the Index Region contains various index structures for accelerating queries. The specific index types include: (1) Ordinal Filter Page, which is a lightweight filter based on order, such as range queries (e.g., WHERE id>100). It is used to implement: recording the maximum and minimum values of the key values in the data page to quickly exclude irrelevant pages. (2) Zonemap Page, whose functions include: similar to Ordinal Filter, but with finer granularity, partitioning data by "zones" (such as time, space). If the data is partitioned by date, the Zonemap can record the date range of each data page. (3) BitmapIndex Page, which marks whether a specific value exists in the data page through a bitmap. A bitmap index is established for the gender field, with one bitmap corresponding to male and another corresponding to female. (4) Bloomfilter Index Page, whose functions include: using a Bloom filter to quickly determine whether a value does not exist in the data page, reducing invalid queries. Its characteristics include: there is a false positive rate, but it has high space efficiency. (5) Fooler (should be Footer, the tail), whose functions include: storing the check information of the index region, the end-of-file flag, and the statistical information of the index structure (such as the total number of terms, index distribution).
[0065] Among them, the core functions of the Inverted Index File include: mapping terms to the documents or records that contain them. Its structure includes: Term Dictionary: storing all unique terms and pointers to the posting lists. Posting List: recording the locations (such as page numbers, offsets) of the data pages or documents that contain the term.
[0066] Among them, in the writing stage: the data is stored in the pages of the Dala Region. The various indexes in the Index Region are synchronously updated (such as the inverted index records the locations of new terms).
[0067] Among them, in the query stage: according to the query conditions (such as keywords, ranges), the index region is preferentially used to filter irrelevant pages. For example: when querying "vin = xxxx", the inverted index directly locates the data page that contains the term, avoiding a full table scan.
[0068] Specifically, add an inverted index: Add an inverted index to the battery pack ID column of the first table. This index uses English word segmentation and supports Phrase queries. That is, when performing text searches, the order of the words after word segmentation will affect the search results. Create an index for historical data: Build an index for historical data according to the newly added index information, so that historical data can also be queried using the inverted index. Command 1: ALTER TABLE batteryType_realTime_table ADD INDEX battery_pack_id_inverted_idx('battery_pack_id') USING INVERTED PROPERTIES("perser" = "english", "support_phrase" = "true"); Command 2: BUILD INDEX battery_pack_id_inverted_idx ON batteryType_realTime_table. After that, for data of the same battery pack, MATCH_PHRASE can be directly used for query. For queries with inverted indexes set, the query rate is nearly 40 times faster than that without inverted indexes enabled.
[0069] Among them, Command 1 is used to: Create an inverted index. Create an inverted index named battery_pack_id_inverted_idx on the field battery_pack_id of the table batteryType_realTime_table. USING INVERTED: Specify the index type as an inverted index, which is suitable for text or word-segmented fields. PROPERTIES: parser = "english": Use an English word tokenizer (such as splitting by spaces, handling stop words, stemming, etc.) to split the field content into tokens. support_phrase = "true": Support phrase matching (exactly match consecutive terms, such as "pack_123" is regarded as an overall query).
[0070] Among them, Command 2 is used to: Build an index, immediately build the defined inverted index, generate index data and make it effective (instead of building it later).
[0071] In DoriDB, the core logic of the inverted index includes: Term mapping: Tokenize field values (such as battery_pack_id) into terms (e.g., ["pack", "123"]), and record the record positions where each term is located. Inverted list: Maintain a list (Posting List) for each term, storing references (such as line numbers or page numbers) of all records containing that term. Query acceleration: Locate relevant records directly through terms, avoiding full table scans.
[0072] Among them, the function of parser = "english": includes performing English word segmentation on the value of battery_pack_id. For example, if the field value is "battery_pack_123", it may be split into ["battery", "pack", "123"]. Applicable scenarios: When fuzzy query by term is required (such as WHERE battery_pack_id LIKE '%pack%').
[0073] Among them, the function of support_phrase = "true": includes supporting exact matching of consecutive terms. For example, when querying "battery_pack", only consecutive occurrences of battery and pack are matched. Applicable scenarios: Scenarios where similar terms need to be distinguished (such as "battery_pack" vs "battery_cell").
[0074] Figure 3 It is a schematic diagram of an inverted index provided in an embodiment of the present application. Refer to Figure 3 , the inverted index is created by decomposing the text into words and establishing a mapping from words to a list of line numbers. These mapping relationships are sorted by words and a skip list index is constructed. When querying a specific word, methods such as skip list index and binary search can be used to quickly locate the corresponding list of line numbers in the ordered mapping, and then obtain the content of the line. This query method avoids line-by-line matching, reducing the algorithm complexity from O(n) to O(logn), especially significantly improving query performance when dealing with large-scale data.
[0075] Specifically, data preparation and index creation include: Suppose there are three documents stored in the DorisDB table, with the following content: Document 1: "Winter is coming.", Document 2: "Ours is the fury.", Document 3: "The choice is yours.". When creating an inverted index for a text field (such as content), DorisDB will perform the following operations:
[0076] First, perform word segmentation: Use the default English word segmenter (or specify a word segmenter) to split the sentence into tokens and standardize them (such as converting to lowercase and removing punctuation). For example: Document 1 → ["winter", "is", "coming"]; Document 2 → ["ours", "is", "the", "fury"]; Document 3 → ["the", "choice", "is", "yours"].
[0077] Then, build an inverted index: Generate a term dictionary (Dictionary) and a postings list (Postings List), as shown in Table 1 specifically.
[0078] Table 1 Term Dictionary (Dictionary) and Postings List
[0079] Term Freq Documents choice 1 3 coming 1 1 fury 1 2 is 3 1,2,3 ours 1 2 the 2 2,3 winter 1 1 yours 1 3
[0080] Secondly, build the core components of the inverted index. Term dictionary (Dictionary): Store all unique terms and their global statistical information (such as total frequency Freq). Example: The total frequency of "is" is 3, indicating that it appears 3 times in all documents. Postings list (Postings List): Each term corresponds to a list that records the document IDs containing the term and additional information (such as term frequency, position offset). Example: The postings list of "the" is [2, 3], indicating that it appears in Document 2 and Document 3.
[0081] Exemplarily, an example of the query process is as follows:
[0082] Scenario: Query the documents containing the term "fury". Parse the query: Segment and standardize the query term "fury" (such as converting to lowercase).
[0083] Retrieve the term dictionary: Look up "fury" in the dictionary and obtain its postings list. Locate the document: According to the postings list [2], directly locate Document 2. Return the result: Without scanning the entire table, only need to read the content of Document 2, significantly improving the efficiency.
[0084] In some embodiments, obtain battery behavior data and battery dimension data, including: Obtain battery behavior data from a message middleware; Obtain battery dimension data from a storage medium.
[0085] Among them, the message middleware can be Kafka, and the storage medium can be HBase.
[0086] In some embodiments, before correlating battery behavior data and battery dimension data, it includes: storing the battery behavior data into a pre-created user log table; creating a user profile index table and a battery static information index table; obtaining first battery dimension data from a storage medium according to the user profile index table and storing it into a pre-created user profile table; obtaining second battery dimension data from the storage medium according to the battery static information index table and storing it into a pre-constructed battery static information table.
[0087] Among them, the first battery dimension data includes information such as province, city, district, reason for repair, number of repairs, whether it has been disposed of, reason for disposal, historical disposal information id, historical disposal information content, etc. The second battery dimension data includes static information such as battery pack id, battery supplier, number of battery series and parallel connections, vehicle model, battery capacity, etc.
[0088] Specifically, obtain the battery behavior data from kafka (message middleware), and correspondingly create a user log table show_log to store the battery behavior data obtained from Kafka into a user log table for subsequent data processing, analysis, and query operations. This can facilitate the statistics, calculation of metrics, generation of reports, etc. of the battery behavior data. And set the watermark to 10s using the timestamp.
[0089] First, create a user profile index table user_profile_index. Then, obtain the first battery dimension data from hbase (storage medium) according to the user profile index table and parse it. And store the parsed data into the pre-created user profile table user_profile. Among them, for the user profile index table: rowkey STRING, INFO ROW <vinstring>, PRIMARY KEY (rowkey) NOT ENFORCED. Among them, the user profile table: uniquely identifies vin, dimensional variables (province, city, district, reason for repair, number of repairs, whether it has been disposed of, reason for disposal, historical disposal information id, historical disposal information content).
[0090] Among them, rowkey is used as the primary key:
[0091] rowkey is the unique identifier of user profile data (such as user ID), used to quickly locate the records of a specific user. Even if marked as NOT ENFORCED, its uniqueness still needs to be guaranteed in design (controlled by the application layer) to avoid data conflicts.
[0092] Optimizing query performance: The primary key is usually automatically indexed by the database (such as a clustered index), significantly accelerating point queries based on rowkey (such as WHERE rowkey = 'user123').
[0093] INFO is used as a composite field (ROW type): The nested field VIN (Vehicle Identification Number) is part of the user profile, representing the association information between the user and the vehicle. Using the ROW type can encapsulate multiple attributes (such as adding phone and address in the future) in an extensible manner, avoiding frequent modification of the table structure.
[0094] The functions of NOT ENFORCED include: The database does not enforce the uniqueness check of rowkey, allowing the application layer to manage it by itself (for example: deduplication logic, temporary duplicates during batch import). It is suitable for high-throughput write scenarios, reducing the constraint check overhead of the database. Queries based on rowkey directly locate data through the primary key index without full table scanning.
[0095] For example, to find the vehicle information (INFO.VIN) of user rowkey = 'user123', the response time is extremely short. The INFO nested field makes the table structure more readable. The relevant program code is as follows:
[0096] {
[0097] "rowkey": "user123", -- The primary key information with rowkey as user123 for quick location;
[0098] "INFO": {
[0099] "VIN": "ABCD1234567890" -- Which contains the user's vin information;
[0100] }
[0101] }
[0102] Support direct access to nested fields (e.g., SELECT INFO.VIN FROM table WHERE rowkey = 'user123'). -- The query can retrieve the corresponding VIN information based on the primary key 'user123'.
[0103] -- Create an HBase dimension table (based on Flink SQL syntax)
[0104] CREATE TABLE user_profile(
[0105] -- Primary key field, mapping to the rowkey in HBase
[0106] rowkey STRING, -- The ROW type defines the INFO column family structure, and the nested field VIN corresponds to the column qualifier;
[0107] INFO ROW<
[0108] VIN STRING -- Vehicle identification number field, stored as INFO:VIN
[0109] >,
[0110] -- Explicitly declare the primary key (must be defined), NOT ENFORCED means not enforcing constraints;
[0111] PRIMARY KEY(rowkey)NOT ENFORCED
[0112] )WITH( -- Connector parameter configuration;
[0113] 'connector' = 'hbase', -- Specify the HBase connector;
[0114] 'table-name' = 'user_profile', -- HBase physical table name;
[0115] 'zookeeper.quorum' = 'zk1:2181,zk2:2181', -- Zookeeper cluster address; 'zookeeper.znode.parent' = ' / hbase', -- HBase root node;
[0116] 'lookup.cache.max-rows' = '1000', -- Maximum number of rows in the dimension table cache;
[0117] 'lookup.cache.ttl' = '60s' -- Cache expiration time; );
[0119] It should be noted that when adding new user attributes, only sub - fields need to be added in INFO (such as INFO ROW<VINSTRING, phone STRING>), and there is no need to modify the existing business logic.
[0120] Then, create a battery static information index table battery_static_message_index. Then, according to the battery static information index table, obtain the second - battery - dimension data from hbase for parsing. And store the parsed data into the pre - created battery static information table battery_static_message. Among them, the battery static information table: uniquely identifies vin, and dimension variables (static information such as battery pack id, battery supplier, number of battery series - parallel connections, vehicle model, battery capacity, etc.).
[0121] In some embodiments, before obtaining the battery - dimension data, it further includes: creating Guava to cache the queried battery - dimension data.
[0122] Specifically, before querying the relevant dimension data from hbase, add Google Guava, set the initial capacity to 10000, the maximum capacity to 20000, use expireAfterAccess (one hour) (specify how long the data expires after it has not been accessed) to set the key as vin, and use getIfPresent(). If there is a value, take the value; if there is no value, put the value in. Two caches can be set corresponding to two dimension tables. Thus, by adding Google Guava before querying the relevant dimension data from hbase, it means using the cache function in the Guava library during the query process to improve the query performance and response speed.
[0123] In some embodiments, associating the battery behavior data and the battery - dimension data to obtain a first table includes: creating a table execution environment; performing a left - join on the user log table, user portrait index table, user portrait table, battery static information index table, and battery static information table according to the table execution environment to form the first table.
[0124] Specifically, using the created table execution environment StreamTableEnvironment, the id of the index table, and the unique identifier vin, perform a left join on the user log table, user portrait index table, user portrait table, battery static information index table, and battery static information table directly to form a wide table (i.e., the first table). Among them, the wide table contains fields such as VIN, timestamp, soc, current, maximum temperature, minimum temperature, temperature difference, maximum cell voltage, minimum cell voltage, maximum cell voltage number, minimum cell voltage number, battery pack number, supplier, vehicle model, mileage, battery pack number, supplier, historical disposal information id, historical warning information, historical disposal information content, pressure difference growth rate, etc.
[0125] After associating the battery behavior data and battery dimension data, create a table batteryType_realTime_table and store the associated data (i.e., the wide table) in DorisDB. Among them, for the batteryType_realTime_table table, use the UNIQUE model (a table model that automatically removes duplicate data with the same key) (key: VIN, timestamp).
[0126] Furthermore, in DorisDB for this wide table, first establish appropriate indexes: For the query requirements, create a prefix index for the commonly used columns (VIN + pressure difference + pressure difference growth rate), and create a BitMap index using (VIN + historical warning information id) to improve the query speed. Then, optimize the Sql query statement: For the pressure difference exceeding 300mV, the historical disposal information id is the pressure difference warning id (e.g., 1), the historical disposal information content (the disposed x# cell), the disposed (x# cell), and whether the minimum cell voltage number corresponding to the pressure difference exceeding 300mV in the current data is consistent, select the inconsistent results (if they are consistent, no processing is required), the historical calculated pressure difference growth rate is greater than 2, and use the index to reduce the full table retrieval. Finally, optimize the partition and bucket strategies: Reasonably set the partition and bucket strategies according to the business characteristics to improve the query efficiency. According to the formula Tablet quantity = partition quantity * Bucket quantity * replication factor, considering the actual situation when the data is >50G, set the partition (timestamp, 30 days) and the bucket (VIN, 64 shards).
[0127] In some embodiments, the battery data processing method further includes: saving the data in the first table to the cache system.
[0128] Specifically, while storing the data of the first table into DorisDB, the data in the first table is also saved into Redis with a set TTL (one week). So that when using the relevant results from DorisDB later, it is possible to directly query from Redis first (adding a Bloom filter). If the query fails, then query from DorsDB, and at the same time save the result into Redis. Thus, by putting the associated data into DorisDB and Redis, it is convenient to query the associated results later, or perform further operations on the associated results, such as applying them to the vehicle real-time big data scenario.
[0129] In some embodiments, the battery behavior data at least includes: vehicle identification code, state of charge of the battery, cell voltage, insulation resistance value, mileage, vehicle speed, timestamp, maximum cell voltage, minimum cell voltage, maximum cell voltage number, and minimum cell voltage number.
[0130] Figure 4 It is a schematic flowchart of a battery data processing provided in an embodiment of the present application. Exemplarily, please refer to Figure 4 The overall process of this battery data processing is as follows: First, obtain battery behavior data from Kafka. Then, create a user portrait index table, and obtain the first battery dimension data from HBase according to the user portrait index table and store it in the HBase dimension table 1 (i.e., the user portrait table). At the same time, create a battery static information index table, and obtain the second battery dimension data from HBase according to the battery static information index table and store it in the HBase dimension table 2 (i.e., the battery static information table). Secondly, use the created table execution environment, the id of the index table, and the unique identifier vin to directly perform a left join on the user log table, user portrait index table, user portrait table, battery static information index table, and battery static information table to form a large wide table. And save the large wide table into DorisDB for caching at the same time. Specifically, if the Redis cache misses, then query DorisDB, and save the query result into Redis (so that the same query can directly return the result from Redis next time). The data related to the cache miss is returned from DorisDB to the client. If the Redis cache misses, it is returned to the client. If the client request does not exist, it is directly rejected. If the client request exists, then release the Bloom filter to point to Redis. Thus, using the highly abstract FlinkSql to process real-time battery data, combining the advantages of multiple big data components, explaining the data processing method, data storage medium, and constructing a real-time data warehouse for the entire real-time processing architecture.
[0131] Figure 5 It is a schematic block diagram of the principle of a battery data processing device provided in an embodiment of the present application. Correspondingly, an embodiment of the present application also provides a battery data processing device. Please refer to Figure 5 , the battery data processing device 100 includes: an acquisition module 101 for acquiring battery behavior data and battery dimension data; an association module 102 for associating the battery behavior data and the battery dimension data through a message middleware and a storage medium to obtain a first table; a storage module 103 for storing the first table into a distributed columnar storage database system; and a query module 104 for setting an inverted index data structure in the distributed columnar storage database system to query the data in the first table.
[0132] In the technical solution of this embodiment, by providing a battery data processing device, the battery data processing device includes: an acquisition module for acquiring battery behavior data and battery dimension data; an association module for associating the battery behavior data and the battery dimension data through a message middleware and a storage medium to obtain a first table; a storage module for storing the first table into a distributed columnar storage database system; and a query module for setting an inverted index data structure in the distributed columnar storage database system to query the data in the first table. It can be seen that by associating the battery behavior data and the battery dimension data, and at the same time putting the associated data into DorisDB and setting an inverted index data structure in DorisDB to query or calculate the associated data, the data query efficiency and subsequent maintenance efficiency can be improved.
[0133] In some embodiments, the query module 104 is further configured to: write an inverted index file while writing a data file in the writing stage, and for each row number of the written data, it corresponds one by one to the row number in the inverted index; in the query stage, if the query condition contains a column for which an inverted index has been established, the distributed columnar storage database system will automatically query the index file, return a list of row numbers that meet the conditions, and then use the row number filtering mechanism common in the distributed columnar storage database system to skip unnecessary rows and pages and only read the rows that meet the preset conditions.
[0134] In some embodiments, the acquisition module is further configured to: acquire battery behavior data from the message middleware; and acquire battery dimension data from the storage medium.
[0135] In some embodiments, the battery data processing device further includes: a first storage unit for storing battery behavior data into a pre-created user log table; a first creation unit for creating a user profile index table and a battery static information index table; a first acquisition unit for acquiring first battery dimension data from a storage medium according to the user profile index table; a second storage unit for storing the first battery dimension data into a pre-created user profile table; a second acquisition unit for acquiring second battery dimension data from the storage medium according to the battery static information index table; and a third storage unit for storing the second battery dimension data into a pre-constructed battery static information table.
[0136] The above has introduced in detail a battery data processing method and device provided by an embodiment of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the technical solution and its core idea of the present application. Those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.< / vinstring>
Claims
1. A battery data processing method, characterized in that, Including: Obtain battery behavior data and battery dimension data; Associate the battery behavior data and the battery dimension data to obtain a first table; Store the first table in a distributed columnar storage database system; Set up an inverted index data structure in the distributed columnar storage database system to query the data in the first table.
2. The battery data processing method according to claim 1, wherein The setting up of the inverted index data structure in the distributed columnar storage database system includes: Write an inverted index file while writing a data file during the write phase. For each row number of the written data, it corresponds one by one to the row number in the inverted index; During the query phase, if the query condition contains columns for which an inverted index has been established, the distributed columnar storage database system will automatically query the index file, return a list of row numbers that meet the conditions, and then use the common row number filtering mechanism in the distributed columnar storage database system to skip unnecessary rows and pages and only read the rows that meet the preset conditions.
3. The battery data processing method according to claim 1, characterized in that, The obtaining of the battery behavior data and the battery dimension data includes: Obtain the battery behavior data from a message middleware; Obtain the battery dimension data from a storage medium.
4. The battery data processing method according to claim 3, wherein, Before associating the battery behavior data and the battery dimension data, it includes: Store the battery behavior data in a pre-created user log table; Create a user portrait index table and a battery static information index table; Obtain first battery dimension data from the storage medium according to the user portrait index table and store it in a pre-created user portrait table; Obtain second battery dimension data from the storage medium according to the battery static information index table and store it in a pre-constructed battery static information table.
5. The battery data processing method according to claim 4, wherein The associating of the battery behavior data and the battery dimension data to obtain a first table includes: Create a table execution environment; According to the table execution environment, perform a left join on the user log table, the user portrait index table, the user portrait table, the battery static information index table, and the battery static information table to form the first table.
6. The battery data processing method according to claim 1, wherein The first table at least includes: vehicle identification code, timestamp, state of charge of the battery, current, maximum temperature, minimum temperature, temperature difference, maximum cell voltage, minimum cell voltage, maximum cell voltage number, minimum cell voltage number, battery pack number, vehicle model, mileage, historical disposal information id, historical warning information, historical disposal information content, and differential pressure growth rate.
7. The battery data processing method according to claim 1, wherein Before obtaining the battery dimension data, it further includes: Create Guava to cache the queried battery dimension data.
8. The battery data processing method according to claim 1, characterized in that, It further includes: Save the data in the first table to a cache system.
9. The battery data processing method according to claim 1, characterized in that The battery behavior data at least includes: vehicle identification code, state of charge of the battery, cell voltage, insulation resistance value, mileage, vehicle speed, timestamp, maximum cell voltage, minimum cell voltage, maximum cell voltage number, and minimum cell voltage number.
10. A battery data processing device, characterized in that, Including: An obtaining module for obtaining battery behavior data and battery dimension data; An associating module for associating the battery behavior data and the battery dimension data through a message middleware and a storage medium to obtain a first table; A storage module for storing the first table into a distributed columnar storage database system; A query module for setting up an inverted index data structure in the distributed columnar storage database system to query the data in the first table.