Distributed storage method for vehicle-mounted OBD data
By dividing vehicle OBD data into key levels and timestamps, constructing multi-dimensional data feature vectors, and dynamically adjusting data fragmentation and redundant coding, the problems of storage resource waste and slow access speed caused by improper data management in existing technologies are solved, achieving efficient data storage and access.
Patent Information
- Application Number
- CN202511003016.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Existing technologies fail to effectively manage the differences between data content and access status in distributed data storage, resulting in non-critical data occupying storage space for a long time, slow data access speed, increased difficulty in data retrieval, and reduced performance of the storage system.
By parsing the data blocks uploaded by the vehicle's OBD terminal, and combining DTC fault codes, vehicle speed, engine speed, and gear information, the data blocks are assigned a critical level and divided into hot or cold data according to timestamps. A data fidelity and criticality score set is established, a multi-dimensional data feature vector is constructed, and data sharding and redundant coding are dynamically adjusted to achieve dynamic updates of the data hierarchical composite index and storage structure.
It effectively avoids the waste of storage space caused by blindly storing data, improves data access efficiency and response speed, flexibly manages different data states, and avoids the loss of important data and performance degradation.
Smart Images

Figure CN120994661A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of distributed storage, in particular to a distributed storage method for vehicle OBD data. BACKGROUND
[0002] Distributed storage technology refers to storing data on multiple independent devices connected through a network, achieving flexible expansion and efficient utilization of data storage resources, and improving data reliability and access performance by dividing, redundancy encoding and managing multiple copies of data.
[0003] The prior art adopts static redundancy sharding and copy management when performing distributed storage of data, ignores the differences in data content and access state, and causes non-important data and expired data to occupy storage space for a long time, resulting in waste of storage resources. In the data positioning and retrieval process, only simple static data sharding location is used for data addressing, and the data access speed is slow. As the data size continues to increase, the data searching difficulty increases, and the overall performance of the storage system is reduced. Therefore, improvement is needed. SUMMARY
[0004] The purpose of the present application is to solve the shortcomings in the prior art and provide a distributed storage method for vehicle OBD data.
[0005] In order to achieve the above purpose, the present application adopts the following technical scheme, a distributed storage method for vehicle OBD data, comprising the following steps: Receiving the original data block uploaded by the vehicle OBD terminal, parsing the DTC fault code, vehicle speed, engine speed, gear information and timestamp, assigning the data block associated with the sudden acceleration and deceleration event and the DTC fault code to a high key level identifier, and assigning the hot data or cold data identifier according to the difference between the timestamp and the current time, and establishing a data fidelity and key level score set; Based on the data fidelity and key level score set, the key level identifier, data temperature identifier and fidelity score of each data block are extracted, combined into a standardized multi-dimensional data feature vector, and then mapped to a predefined and hierarchical index space address according to the preset fidelity threshold, key level type and temperature type, to establish a data hierarchical composite index; According to the data hierarchical composite index, obtain the data block coding strategy selection instruction, dispatch the distributed storage system to execute the data block coding strategy selection instruction, perform data sharding and redundancy encoding on the data block, and write the encoded shards into the distributed storage nodes in different storage pools to generate a distributed storage node layout table; Start a background scheduled task to scan the timestamp information in the data hierarchical composite index, filter the index entries that have been cooled from hot data to cold data, summarize and generate a list of data blocks to be migrated, and establish an updated storage structure index based on the list of data blocks to be migrated and the distributed storage node layout table.
[0006] Preferably, the steps for obtaining the data fidelity and key score set are as follows: Extract the corresponding vehicle speed value, engine speed value, gear information, sensor self-test status value and data source reputation score from each raw data block. Input the engine speed value and gear information into the transmission ratio function and calculate the theoretical vehicle speed. Perform difference analysis between the theoretical vehicle speed and the vehicle speed value. At the same time, accumulate the abnormal count of the sensor self-test status value. Assign high criticality level identification to the data blocks associated with rapid acceleration and deceleration events and DTC fault codes to generate a set of parameters to be scored. Calculate the fidelity score for each data block based on the set of parameters to be scored; Based on the fidelity score, the criticality identifier of each data block, and the data temperature identifier, a data item identifier set is formed by combining them in sequence, and then grouped according to the index number of the original data block to generate a data fidelity and criticality score set.
[0007] Preferably, the steps for obtaining the multidimensional data feature vector are as follows: Based on the index number of each data item in the data fidelity and criticality score set, the criticality identifier, data temperature identifier, and fidelity score of the data item corresponding to the index number are extracted one by one. The criticality identifier is converted through binary encoding, and the data temperature identifier is converted into a standard value through the timeliness weight table to form an initial set of data feature values. Based on the initial set of data feature values, the fidelity threshold, key level type classification table, and temperature type classification table are called respectively. Classification is performed based on the difference between the fidelity score and the fidelity threshold. The key level identifier value is matched with the index position of the key level type classification table, and the data temperature identifier value is matched with the index position of the temperature type classification table. The matching results are combined to form a multidimensional data feature vector.
[0008] Preferably, the steps for obtaining the hierarchical composite index are as follows: Based on the multidimensional data feature vectors, the corresponding index space addresses are retrieved item by item according to the preset mapping rule table. The index space is located step by step using a step-by-step mapping method. The multidimensional data feature vectors are mapped and located to the addresses in the index space one by one, thus establishing a hierarchical composite index for data.
[0009] Preferably, the step of obtaining the data block encoding strategy selection instruction is as follows: Based on the index space address of each index entry in the data hierarchical composite index, the data items corresponding to the address are retrieved one by one, the key level identifier and data temperature identifier contained in the data item are extracted, and the index matching conditions are formed based on the encoded value of the key level identifier and the classification value of the data temperature identifier, thereby generating a set of encoding strategy query conditions. Based on the set of query conditions for the coding strategy, the preset erasure coding strategy library is called one by one. The applicable key level identifier range and data temperature identifier range of each strategy in the erasure coding strategy library are compared one by one. Erasure coding strategies that meet the query conditions are filtered out, and the strategy index number of the erasure coding strategy that meets the conditions is determined to form an erasure coding strategy index set. Based on the erasure coding policy index set, erasure coding policy entries with corresponding policy index numbers are retrieved one by one from the erasure coding policy library. The encoding rules in the erasure coding policy entries are parsed, and the corresponding erasure coding instructions are output one by one according to the encoding rules to obtain the data block encoding policy selection instructions.
[0010] Preferably, the step of obtaining the distributed storage node layout table is as follows: According to the data block encoding strategy, select instructions, perform equal-length splitting of data blocks according to the number of erasure coding fragments, map and bind each fragment identifier to the original data block identifier, and extract the total capacity value, available capacity value, current write rate value, historical write rate standard deviation value and current write task number of each candidate distributed storage node to generate a set of resource status of nodes to be scheduled. Based on the set of resource statuses of the nodes to be scheduled, calculate the node placement score for each storage node; Based on the node placement score of each storage node, all candidate nodes are sorted in descending order of node placement score. The nodes with the highest node placement scores are selected in sequence according to the data shard numbers. Each data shard is assigned to a matching node, and the binding relationship between the data shard and the storage node is recorded to generate a distributed storage node layout table.
[0011] Preferably, the step of obtaining the list of data blocks to be migrated is as follows: Scan each index entry according to the data hierarchical composite index, extract the data temperature identifier and corresponding data timestamp value from each index entry, calculate the difference between the data timestamp value and the current system time value, and determine if the data temperature identifier is hot data and the time difference is greater than the preset cooling threshold. Then mark the entry as an index entry to be migrated and generate a list of data blocks to be migrated.
[0012] Preferably, the step of obtaining the updated storage structure index is as follows: Based on the list of data blocks to be migrated, the index numbers of the data blocks in the list of data blocks to be migrated are extracted one by one. The distributed storage node layout table is queried one by one according to the index numbers of the data blocks. The binding relationship between the data shards and storage nodes corresponding to each data block index number is determined. The storage nodes are located one by one through the binding relationship and all data shard contents are read in sequence. The complete data content is reorganized according to the data shard order numbering to generate a complete original data block. Based on the complete original data block, according to the preset cold data encoding strategy, each original data block is re-divided into multiple data shards. Distributed storage nodes with the storage pool category of cold data storage pool are selected one by one to write the data shard content in sequence, and the binding mapping relationship between each data shard and the storage node written is recorded to generate an updated storage structure index.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention analyzes raw data blocks uploaded by the vehicle's OBD terminal, correlates DTC fault codes with rapid acceleration / deceleration events, and further assigns criticality levels to the data by combining vehicle speed, engine speed, and gear information. Simultaneously, it classifies data based on the difference between the data's timestamp and the current time, forming a data fidelity and criticality score set. Based on the score set, it constructs a standardized multi-dimensional data feature vector and establishes a hierarchical index space address according to preset thresholds and classification mapping rules, realizing a hierarchical composite index. Based on the index information, it further selects the most suitable data encoding strategy, performs data sharding and redundant encoding processing on data blocks, dynamically adjusts the shard storage location, and executes periodic tasks in the background to monitor the data's hot / cold status in real time, achieving automatic data migration and dynamic updates to the storage structure index. This avoids the waste of storage space caused by blind data storage, improves data access efficiency and response speed through data temperature management, and flexibly executes encoding strategies for different data states, effectively preventing the loss of important data and performance degradation. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0016] Please see Figure 1 This invention provides a technical solution: a distributed storage method for vehicle-mounted OBD data, comprising the following steps: Receive raw data blocks uploaded by the vehicle OBD terminal, parse DTC fault codes, vehicle speed, engine speed, gear information and timestamp, assign high criticality level identifiers to data blocks associated with rapid acceleration and deceleration events and DTC fault codes, and assign hot data or cold data identifiers based on the difference between the timestamp and the current time, and establish a data fidelity and criticality score set. Based on data fidelity and criticality score sets, the criticality identifier, data temperature identifier and fidelity score of each data block are extracted and combined into a standardized multidimensional data feature vector. Then, according to the preset fidelity threshold, criticality type and temperature type, the multidimensional data feature vector is mapped to the predefined and hierarchical index space address to establish a hierarchical composite index. Based on the data hierarchical composite index, obtain the data block encoding strategy selection instruction, schedule the distributed storage system to execute the data block encoding strategy selection instruction, perform data sharding and redundant encoding on the data block, and write the encoded shards into the distributed storage nodes in different storage pools to generate a distributed storage node layout table. Start a background scheduled task to scan the timestamp information in the data hierarchical composite index, filter the index entries that have been cooled from hot data to cold data, summarize and generate a list of data blocks to be migrated, and establish the updated storage structure index based on the list of data blocks to be migrated and the distributed storage node layout table.
[0017] The steps for obtaining data fidelity and the critical score set are as follows: Extract the corresponding vehicle speed value, engine speed value, gear information, sensor self-test status value and data source reputation score from each raw data block. Input the engine speed value and gear information into the transmission ratio function and calculate the theoretical vehicle speed. Perform difference analysis between the theoretical vehicle speed and the vehicle speed value. At the same time, accumulate the abnormal count of the sensor self-test status value. Assign high criticality level identification to the data blocks associated with rapid acceleration and deceleration events and DTC fault codes to generate a set of parameters to be scored. Based on the set of parameters to be scored, calculate the fidelity score for each data block using the following formula: ; in, Indicates the first The fidelity score of each data block. Indicates the first The normalized result of the data source reputation score for each data block. Indicates the first The sensor self-test status anomaly count value in each data block, This represents the impact coefficient of self-test failure. Indicates the first The vehicle speed values for each data block. Indicates the first The engine speed values for each data block. Indicates the first The gear information of each data block, Indicates gear position The corresponding total gear ratio function result, Indicates the reference vehicle speed constant. This represents the logical consistency sensitivity coefficient. Based on the fidelity score, the criticality identifier of each data block, and the data temperature identifier, a set of data item identifiers is formed by combining them in sequence, and then grouped according to the index number of the original data blocks to generate a set of data fidelity and criticality scores.
[0018] Specifically, the process begins by extracting vehicle speed, engine speed, gear position, sensor self-test status, and data source reputation score from each raw data block uploaded by the vehicle's OBD terminal. The vehicle speed is directly read from the OBD-II standard PID 0D, the engine speed from PID 0C, the gear position information is obtained by parsing the vehicle's proprietary CAN bus protocol, and the sensor self-test status is read from the PID... 01. The data source reputation score is obtained by querying a reputation rating database that maintains the historical reliability of each terminal, based on the terminal ID. Next, the system inputs the engine speed and gear information into a preset transmission ratio function for processing. This function contains a lookup table extracted from the vehicle manufacturer's technical manual and embedded in the program. This table maps gear information (e.g., 'P', 'R', 'N', 'D', '1', '2') to specific transmission ratio values. Combining the final transmission ratio and tire dynamic circumference, the theoretical vehicle speed is calculated. The calculation method is: theoretical vehicle speed equals engine speed divided by the total transmission ratio, multiplied by the tire dynamic circumference, with the unit converted from meters per minute to kilometers per hour. Then, the difference between the calculated theoretical vehicle speed and the actual vehicle speed extracted from the data block is calculated, and the absolute value of the difference is recorded. Simultaneously, the system parses the sensor self-check status value, which is a multi-bit binary number, with each bit corresponding to the health status of a specific sensor. The system iterates through a preset list of key sensors (including but not limited to oxygen sensors). The system checks the status bits of sensors (such as the air flow sensor and throttle position sensor) to see if they are abnormal (e.g., "1"). It then accumulates the number of abnormal flags to obtain the sensor self-test anomaly count for that data block. Furthermore, the system analyzes the vehicle speed and timestamp sequence in the data block, calculating the acceleration between consecutive data points (acceleration equals the speed change divided by the time change) and comparing it to a preset rapid acceleration / deceleration threshold. This threshold is set after statistical analysis of over 1000 hours of driving data under various road conditions, including urban, highway, and rural areas. For example, a rapid acceleration / deceleration event is defined as an acceleration greater than 3.5 meters per second or an absolute deceleration greater than 4.0 meters per second. If an event occurs, or if the data block contains a non-empty DTC fault code list, the system assigns a high-criticality flag to the data block; otherwise, it assigns a low-criticality flag. Finally, the raw data, calculated theoretical vehicle speed, speed difference, sensor self-test anomaly count, and criticality flags are integrated into a structured dataset, forming a set of parameters to be scored.
[0019] formula: The advantage of this formula lies in that it doesn't rely solely on a single data validity check, but rather achieves precise quantification of OBD data quality by organically combining the credibility of the data source itself, the health status of the vehicle's sensor hardware, and the logical consistency between multiple key dynamic data points. Specifically, this is achieved by introducing a data source reputation score. This enables the system to make prior judgments about data sources based on historical performance, suppressing contaminated data from inferior or faulty terminals; and it integrates sensor self-test anomaly counts. With influence coefficient It directly incorporates the health status of the hardware into its considerations, penalizing data distortion caused by sensor malfunctions; most importantly, it uses an exponential decay term... By examining the physical correlation between the three core dynamic data points—vehicle speed, engine speed, and gear position—abnormal data that contradicts the vehicle dynamics model can be identified. This design makes the fidelity score highly sensitive to logical inconsistencies, thereby greatly improving the accuracy and intelligence of data cleaning and providing a high-quality decision-making basis for subsequent distributed storage strategies.
[0020] Indicates the first The data source reputation score for each data block is obtained through the following steps: The system maintains a reputation profile for each registered OBD terminal on the backend server, with an initial score of 100. The data center's backend analysis program periodically samples and verifies the stored data, for example, cross-validating specific trip data uploaded by a terminal with high-precision GPS tracks and vehicle video recordings. If a data block is found to have a serious deviation (such as a speed error consistently exceeding 20%), the terminal's reputation score is deducted, calculated using the following formula: Conversely, if a terminal uploads 10,000 data blocks consecutively without any logical inconsistencies or conflicts with the verification data, it will be rewarded. The calculation formula is as follows: The score has a lower limit of 0 and an upper limit of 100. For example, a device with terminal ID "SN202308" has a current reputation score of 92. .
[0021] Indicates the first The normalized result of the data source reputation score of each data block is obtained through the following steps: In order to normalize the range within... Reputation score mapped to Within the interval, to facilitate multiplication with other factors, the max-min normalization method is used. The normalization formula is: .in The maximum credit score is set at 100. The minimum reputation score is set to 0. This normalization process ensures that the weight of the data source reputation in the formula is on the same scale as other factors. For example, for the terminal with a reputation score of 92 above, the normalization result is: .
[0022] This represents the impact coefficient of a self-test failure, obtained through the following steps: This coefficient is used to adjust the penalty for sensor failure on the fidelity score. Its value was set based on statistical analysis of over 500,000 historical data records containing DTC fault codes. The analysis showed that when a vehicle experiences a failure of one critical sensor (such as the crankshaft position sensor), there is approximately a 40% probability that related data streams (such as engine speed) will become completely unreliable. To quantify this impact, a target was set: when one sensor fails (i.e.,...)... When considering this, the desired penalty factor in the fidelity score should be... The value is approximately 0.6. Calculations are performed based on this objective: Solving for Therefore, setting It is 0.67.
[0023] Indicates the first The abnormal count value of sensor self-test status in each data block is obtained by directly calculating the value when generating the "set of parameters to be scored" in the previous step. The status of specific bits associated with core sensors (such as engine coolant temperature sensor, intake manifold absolute pressure sensor, vehicle speed sensor, etc.) is checked, and for each bit with a "failed" status, the count value is... Add 1. For example, if the system detects that both the intake manifold pressure sensor and the oxygen sensor report a self-test failure, then... .
[0024] The logical consistency sensitivity coefficient is determined by the following steps: This coefficient reflects the sensitivity of the fidelity score to logical deviations in vehicle speed. To determine its value, OBD data from 10 mainstream vehicle models were collected under standard test cycles (such as NEDC), totaling 100 hours. The deviation rate between theoretical and actual vehicle speeds in this data was analyzed. It was found that under fault-free conditions, 95% of the deviation rate was within 0.05. The design goal is to reduce the decay factor of the exponential part when the deviation rate reaches the critical value of 0.05. It should decrease significantly to To form an effective punishment. Calculations are performed based on this objective: Solving for Therefore, a logical consistency sensitivity coefficient is set. It is 20.
[0025] Indicates the first The vehicle speed value for each data block, in km / h. This value is directly parsed from the OBD-II PID 0D (Vehicle Speed) contained in the original data block. For example, if the vehicle speed parsed from the current data block is 80 km / h, then... .
[0026] Indicates the first The engine speed value for each data block, in RPM. This value is directly parsed from the OBD-II PID 0C (Engine RPM) contained in the original data block. For example, if the engine speed parsed from the current data block is 2500 RPM, then... .
[0027] Indicates gear position The corresponding total gear ratio function result. This function encapsulates the vehicle's transmission system parameters. First, the gear information is parsed from the data block. For example, 'D4' (4th gear). Then, a pre-set vehicle parameter library is consulted to find the gearbox gear ratio (e.g., 1.0), final drive ratio (e.g., 3.2), and tire dynamic circumference (e.g., 2.1 meters, or 0.0021 kilometers) that match the vehicle model and gear 'D4'. The overall gear ratio function integrates these parameters to calculate and outputs a coefficient that converts RPM to km / h. The calculation formula is: .For example, .
[0028] This represents the reference vehicle speed constant. This constant is used to normalize speed deviations, making them dimensionless relative values. Its value is set at 120 km / h, a generally representative value for high-speed vehicle operation, ensuring that deviation assessments are within a reasonable range. Therefore, it is set... .
[0029] Calculation process: The fidelity score is calculated for a specific OBD data block. The parameter values for this data block are as follows: Data source reputation score After normalization .
[0030] Sensor self-test anomaly count value .
[0031] Actual vehicle speed km / h.
[0032] Engine speed RPM.
[0033] Gear information For 'D4', the corresponding total gear ratio function result .
[0034] The coefficients are set as follows: , , .
[0035] First, calculate the theoretical vehicle speed: ; Then, calculate the absolute deviation of vehicle speed: ; Next, substitute all parameters into the fidelity score formula. middle: ; ; Perform the calculation step by step: Calculate the sensor penalty factor: ; Calculate the logical consistency index term: ; Finally, multiplying the components together yields the final fidelity score: ; The result indicates that the fidelity score of this data block is approximately 0.018, which is extremely low (scoring range 0 to 1). This low score is primarily due to a significant speed logic error (actual speed 80 km / h, while theoretical speed is as high as 98.4 km / h) and two sensor malfunctions. In fidelity grading, a score below 0.3 is typically classified as "unreliable," meaning the data block contains serious errors and its internal values should not be used for any accurate analysis or model training.
[0036] Based on the fidelity score of each data block calculated in the previous step, and the key identifiers and data temperature identifiers extracted or determined from the data blocks, the system initiates a data integration and structuring process. First, for determining the data temperature identifier, the system extracts the UTC timestamp recorded within each data block and compares it with the current system's standard time, calculating the time difference. The system has a data cooling time threshold set at 2,592,000 seconds (30 days). This value is based on statistical analysis of data query logs over the past year. The analysis shows that over 95% of data access requests are concentrated within 30 days of data generation. Therefore, if the time difference is less than or equal to this threshold, the data block is designated as "hot data." The system assigns a "high" critical level identifier to data blocks with a high value and a "cold" identifier to data blocks with a low value. Then, it creates a structured data item for each raw data block. This data item contains four core metadata fields: the data block's unique index number, the calculated fidelity score, the critical level identifier ('high' or 'low'), and the data temperature identifier ('hot' or 'cold'). For example, for a data block with index number "BLK_20240520_1130_001", its calculated fidelity score is 0.018. Because it is associated with a DTC fault code, it is assigned a "high" critical level identifier, and because its timestamp is within 30 days, it is assigned a "hot" data identifier. The system will then generate a corresponding data item, represented in memory as a data block containing {index: For records like "BLK_20240520_1130_001", fidelity: 0.018, criticality: "high", temperature: "hot", after all data blocks are processed, a set of data item identifiers containing tens of thousands of such records will be formed. Finally, to facilitate subsequent fast retrieval and index construction, the system transforms this flattened set of records into a hash map or similar key-value data structure with the original data block index number as the key. This process involves traversing the set of data item identifiers, using the index number in each record as the key, and storing the remaining part, which includes the fidelity score, criticality identifier, and data temperature identifier, as a composite value in the hash map, ultimately generating a set of data fidelity and criticality scores.
[0037] The steps for obtaining feature vectors from multidimensional data are as follows: Based on the index number of each data item in the data fidelity and criticality scoring set, the criticality identifier, data temperature identifier, and fidelity score are extracted from the data item corresponding to the index number. The criticality identifier is converted through binary encoding, and the data temperature identifier is converted into a standard value through the timeliness weight table to form an initial set of data feature values. Based on the initial set of data feature values, the fidelity threshold, key-level type classification table, and temperature type classification table are called respectively. Classification is performed based on the difference between the fidelity score and the fidelity threshold. The key-level identifier value is matched with the index position of the key-level type classification table, and the data temperature identifier value is matched with the index position of the temperature type classification table. The matching results are combined to form a multi-dimensional data feature vector.
[0038] Specifically, based on the data fidelity and criticality score set generated in the aforementioned steps, the system initiates an iterative processing flow, processing each data item in the score set one by one. Specifically, the system first extracts the corresponding critical level identifier (e.g., "high"), data temperature identifier (e.g., "hot"), and fidelity score (e.g., 0.018) from the score set using the data item's index number (e.g., "BLK_20240520_1130_001). Next, the system quantizes the extracted non-numerical identifiers. For critical level identifiers, a fixed binary encoding rule is used for conversion: "high" critical level identifiers are mapped to the value 1, and "low" critical level identifiers are mapped to the value 0. This encoding directly reflects the data priority. For data temperature identifiers, the system calls a pre-set "timeliness weight table" for conversion. This table is constructed based on the statistical analysis of data access logs over the past year. Analysis revealed that data generated within 30 days (i.e., "hot" data) had an average access frequency 18.5 times higher than data generated more than 30 days ago (i.e., "cold" data). To quantify this difference, the timeliness weight table mapped the "hot" identifier to a standard value of 1.0 and the "cold" identifier to a standard value of 0.2. This value setting logic was based on the relative relationship of access frequency and underwent normalization processing. For example, for data identified as "hot," its timeliness weight was 1.0. This transformation process converted discrete temperature classifications into continuous numerical weights. Finally, the original fidelity score, the key-level value after binary encoding conversion, and the standard temperature value of the data after conversion through the timeliness weight table were combined as a tuple. For example, for the data item with index number "BLK_20240520_1130_001," the processed numerical tuple was (0.018, 1, 1.0). The system repeated this process for all data items in the scoring set to form an initial set of data feature values.
[0039] Based on the initial set of data feature values formed in the previous process, the system performs multi-dimensional classification and matching of the numerical tuples in each data block. First, the system calls a set of preset fidelity thresholds to discretize and classify the fidelity scores. This threshold set includes three values: 0.3, 0.6, and 0.9. These thresholds are set based on statistical analysis of the fidelity score distribution of over one million historical data blocks. Scores below 0.3 are considered unreliable, accounting for 5% of the total, while scores between 0.3 and 0.6 are considered... Low fidelity (20%), medium fidelity (0.6-0.9%), and high fidelity (15%) are classified as follows: 0.3 or less is classified as level 0; 0.6-0.6 as level 1; 0.6-0.9 as level 2; and greater than 0.9 as level 3. The system then calls a "key-level type classification table," a simple key-value pair mapping structure: {0: The system uses the key-level identifier value (0 or 1) from the initial data feature value set as the key to query the table to obtain the corresponding classification index position. For example, if the key-level value is 1, the matched index position is 1. Then, the system calls the "Temperature Type Classification Table," which is also based on threshold-based classification rules. The threshold is set to 0.5, which is determined by the midpoint between the "hot" data weight of 1.0 and the "cold" data weight of 0.2 determined in the previous process. The classification rule is: if the data temperature identifier value is greater than 0.5, it matches index position 1 (representing hot data); if the value is less than or equal to 0.5, it matches index position 0 (representing cold data). Finally, the system combines the fidelity classification level, key-level type index position, and temperature type index position obtained through the above three steps into a three-dimensional integer vector in sequence. For example, if the initial data feature values of a data block are (0.85, 1, 1), the system will use the key to query the table to obtain the corresponding classification index position ... 1.0), its fidelity score of 0.85 is between 0.6 and 0.9, and it is classified as level 2. The key level value 1 matches index 1, and the data temperature value 1.0 is greater than 0.5 and matches index 1. Therefore, the final multidimensional data feature vector is (2, 1, 1).
[0040] The steps to obtain a hierarchical composite index are as follows: Based on multidimensional data feature vectors, the corresponding index space addresses are retrieved item by item according to the preset mapping rule table. The index space is located step by step using a hierarchical mapping method. The multidimensional data feature vectors are mapped and located to the addresses in the index space one by one, thus establishing a hierarchical composite index.
[0041] Specifically, based on the multidimensional data feature vectors corresponding to each data block generated in the previous steps, the system executes a deterministic, hierarchical address mapping process. This process relies on an embedded, structured "mapping rule table." This rule table is not a physical table, but rather a set of path construction rules implemented through program logic. Its core idea is to parse the three-dimensional feature vectors dimension by dimension and map them to an index space address with a hierarchical structure. The index space is designed as / Temperature / Criticism / Fidelity, forming a three-level directory structure. The mapping process adopts a step-by-step positioning method. Specifically, for a given multidimensional data feature vector, such as (2, 1, ... 1) The system first parses the third element of the vector (temperature type index), which maps to the first-level directory: index 1 corresponds to "hot_data", index 0 corresponds to "cold_data", thus determining the first-level path as / hot_data / . Next, the system parses the second element of the vector (criticality level index), which maps to the second-level directory: index 1 corresponds to "high_criticality", index 0 corresponds to "low_criticality", concatenating them under the determined first-level path to form the path / hot_data / high_criticality / . Finally, the system parses the first element of the vector (fidelity classification level), which maps to the third-level directory: level 3 corresponds to "fidelity_high", level 2 corresponds to "fidelity_medium", and so on. Level 1 corresponds to "fidelity_low", and Level 0 corresponds to "fidelity_untrusted". The system continues to concatenate the data along the established secondary paths to form the final index space address / hot_data / high_criticality / fidelity_medium / . The system associates the original index number of the currently processed data block (e.g., "BLK_20240520_1130_001") with this generated index space address and adds this association to a global index data structure. This data structure is typically implemented as a nested hash table or tree structure, where the key path is the generated index space address and the value is a list storing the index numbers of all data blocks mapped to that address. By repeatedly performing this mapping process on the multidimensional data feature vectors of all data blocks, a hierarchical composite index is finally established.
[0042] The steps for obtaining the data block encoding strategy selection instruction are as follows: Based on the index space address of each index entry in the hierarchical composite index, the data items corresponding to the address are retrieved one by one, the key level identifier and data temperature identifier contained in the data item are extracted, and the index matching conditions are formed based on the coded value of the key level identifier and the classification value of the data temperature identifier, generating a set of coded strategy query conditions. Based on the set of query conditions for coding strategies, the preset erasure coding strategy library is called one by one. The applicable key level identifier range and data temperature identifier range of each strategy in the erasure coding strategy library are compared one by one. The erasure coding strategies that meet the query conditions are filtered out, and the strategy index number of the erasure coding strategies that meet the conditions is determined to form the erasure coding strategy index set. Based on the erasure coding policy index set, the erasure coding policy entries with the corresponding policy index numbers in the erasure coding policy library are retrieved one by one. The encoding rules in the erasure coding policy entries are parsed, and the corresponding erasure coding instructions are output one by one according to the encoding rules to obtain the data block encoding policy selection instructions.
[0043] Specifically, based on the data hierarchical composite index established in the aforementioned steps, the system initiates a traversal program. This program scans all entries in the index using either depth-first or breadth-first search until reaching the leaf nodes. For each leaf node, the path itself contains the criticality level and data temperature classification of the corresponding data block. For example, for an index entry with the path / hot_data / high_criticality / fidelity_medium / , the system directly parses the data temperature classification as "hot data" and the criticality level classification as "high criticality." The system then converts this classification information into predefined values. This conversion follows a fixed mapping rule: for criticality levels, "high" is converted to the value 1, and "low" is converted to the value 0; for data temperature, "hot" is converted to the value 1, and "cold" is converted to the value 0. This process does not require re-extraction. Instead of using the original data items, the system directly utilizes the information from the index structure itself. Therefore, for the above index path, the system obtains a tuple (1, 1) consisting of the key-level coding value and the data temperature classification value. This tuple constitutes the index matching condition. The system binds this matching condition to the index numbers of all data blocks associated with this index path (e.g., "BLK_20240520_1130_001", "BLK_20240520_1132_005", etc.) to form a temporary query task. This process is repeated for all unique path combinations in the index. For example, another path / cold_data / low_criticality / fidelity_high / will generate a matching condition (0, 0). Finally, the system summarizes all these temporary query tasks to generate a set of coding strategy query conditions.
[0044] Based on the set of coding strategy query conditions generated in the previous process, the system initiates a strategy matching process. This process processes each query condition in the set and calls a preset erasure coding strategy library for comparison. This erasure coding strategy library is a static data structure configured during system initialization. Its construction is based on a comprehensive consideration of the importance of the data, access frequency, storage cost, and system reliability requirements. Specifically, the strategy library contains four core strategies: Strategy 1 (strategy index number "EC_HP_HT"), which is for data with a criticality level of 1 and a temperature of 1 (high criticality, hot data). Strategy 1 employs highly reliable RS(6,3) encoding, consisting of 6 data shards and 3 parity shards, providing high redundancy to handle potential failures in frequent read / write environments. This strategy is based on the fact that this type of data is the most valuable and accessed most frequently, necessitating guaranteed availability. Strategy 2 (strategy index number "EC_LP_HT") uses the more cost-effective RS(10,2) encoding for data with a criticality level of 0 and a temperature of 1 (low-criticality, hot data), consisting of 10 data shards and 2 parity shards, reducing storage overhead while maintaining a certain level of reliability. Strategy 3 (strategy index number "... For data with a criticality level of 1 and a temperature of 0 (high criticality, cold data), strategy 4 (EC_HP_CT) uses RS(4,4) encoding, which is optimal for long-term archiving, providing extremely high redundancy to combat silent errors that may occur during long-term storage. Strategy 5 (EC_LP_CT) uses RS(12,2) encoding, which has the highest storage efficiency, for data with a criticality level of 0 and a temperature of 0 (low criticality, cold data), to preserve it for the long term at the lowest cost. For each query condition in the set of query conditions for the encoding strategy, such as (1,1), the system will iterate through... All erasure coding strategies in the strategy library are strictly compared with the applicable critical level identifier range and data temperature identifier range of each strategy. When the query condition completely matches the applicable range of a certain strategy, for example, (1,1) matches the applicable range of strategy one (critical level 1, temperature 1), the system selects the strategy and extracts its strategy index number "EC_HP_HT". This number is bound to the index numbers of all data blocks associated with the query condition to form a strategy mapping entry. This process is repeated for all query conditions, and finally the erasure coding strategy index set is formed.
[0045] Based on the erasure coding policy index set formed in the previous process, the system starts the instruction generation program to process each policy mapping entry in the set one by one. For an entry, for example, containing a policy index number "EC_HP_HT" and a list of associated data block index numbers, the system first uses the policy index number to query the preset erasure coding policy library to retrieve the complete policy details. The details clearly record the encoding rule as "RS(6,3)". Then, the system's internal instruction parser parses the string "RS(6,3)", identifies the encoding algorithm as "Reed-Solomon", and extracts two core parameters: the number of data fragments is 6, and the number of redundant (checksum) fragments is 3. After obtaining these parameters, the system generates a structured encoding instruction for each data block index number contained in the policy mapping entry according to these specific encoding rules. This instruction is a data object containing multiple key-value pairs. For example, for the data block "BLK_20240520_1130_001", the specific content of the generated instruction is {"target_block_id": The instruction is generated as follows: “BLK_20240520_1130_001”, “encoding_algorithm”: “Reed-Solomon”, “data_shards_k”: 6, “parity_shards_m”: 3, “chunk_size”: 1048576}, where chunk_size is a preset chunk size used uniformly for all data blocks, such as 1MB. The system repeats this instruction generation process for all data block index numbers in the list to ensure that each data block obtains a clear and executable encoding task description. After all policy mapping entries have been processed, all generated individual encoding instructions are aggregated into a queue or list, ultimately yielding the data block encoding policy selection instruction.
[0046] The steps to obtain the distributed storage node layout table are as follows: Based on the data block encoding strategy, select instructions, perform equal-length splitting of data blocks according to the number of erasure coding fragments, map and bind each fragment identifier to the original data block identifier, and extract the total capacity value, available capacity value, current write rate value, historical write rate standard deviation value and current write task number of each candidate distributed storage node to generate a set of resource status of nodes to be scheduled. Based on the set of resource statuses of the nodes to be scheduled, calculate the node placement score for each storage node using the following formula: ; in, Indicates candidate storage nodes Node placement score Indicates storage node The current available capacity, Indicates storage node Total capacity Indicates storage node Current I / O write rate, This represents the maximum I / O write rate across all nodes in the system. This is the coefficient for the percentage of capacity score. This is the I / O performance score ratio coefficient. Indicates storage node Current number of concurrent write tasks, As the load penalty factor, Indicates storage node Historical I / O write rate standard deviation As a stability penalty factor, To prevent division by zero for extremely small constants; Based on the node placement score of each storage node, all candidate nodes are sorted in descending order of node placement score. The nodes with the highest node placement scores are selected in sequence according to the data shard numbers. Each data shard is assigned to a matching node, and the binding relationship between the data shard and the storage node is recorded to generate a distributed storage node layout table.
[0047] Specifically, based on the data block encoding strategy, the system selects instructions and initiates a data distribution preprocessing process. This process first retrieves instructions one by one from the instruction queue for parsing. For example, for the instruction {“target_block_id”: “BLK_20240520_1130_001”, “encoding_algorithm”: “Reed-Solomon”, “data_shards_k”: 6, “parity_shards_m”: 3, “chunk_size”: 1048576}, the system locates and reads the content of the original data block based on the target_block_id. Then, according to the number of data fragments k=6 and the block size chunk_size=1MB in the instruction, the data block is divided into 6 equal-length 1MB data fragments from beginning to end. If the original data block size is not an integer multiple of 6MB, the last fragment will be padded with zero bytes to reach 1MB. Subsequently, the system creates a unique global fragment identifier for each generated data fragment. The rule is to concatenate underscores and fragment indexes to the original data block identifier. For example, generate 6 data fragment identifiers from BLK_20240520_1130_001_ds_0 to BLK_20240520_1130_001_ds_5, while reserving 3 redundant fragment identifiers from BLK_20240520_1130_001_ps_0 to BLK_20240520_1130_001_ps_2, and then link these fragment identifiers with their common original data block identifier. The system uses BLK_20240520_1130_001 to establish a temporary mapping table in memory. Simultaneously, a parallel monitoring agent broadcasts a status query request to all candidate storage nodes in the distributed storage cluster. Upon receiving the request, each node immediately returns its current resource status information, including: the node's total storage capacity (a static configuration value, e.g., 10TB); the node's current available capacity (obtained in real-time by querying the file system, e.g., 4.2TB); the node's current I / O write rate (calculated by collecting the number of bytes written per second from the disk over the past 5 seconds, e.g., 85MB / s); the node's historical I / O write rate standard deviation (obtained by calculating the standard deviation of the I / O write rate sampling sequence per minute over the past hour, e.g., 12.5MB / s); and the node's current number of concurrent write tasks (obtained by querying the node's internal task scheduling queue length, e.g., 5 tasks). All these status information returned by the nodes are aggregated to form a set of resource statuses for nodes awaiting scheduling.
[0048] formula: The advantage of this formula lies in its ability to intelligently select the optimal storage location for data sharding. Its core advantage is reflected in its comprehensive consideration of multiple dimensions, through linear weighting terms. It balances the two fundamental resource metrics of node storage capacity and real-time I / O performance, avoiding resource bottlenecks caused by focusing on only one. More importantly, it introduces two non-linear penalty factors: load penalty factor and load penalty factor. It can effectively penalize nodes with excessive load, preventing a surge in write latency due to task backlog, and has a stability penalty factor. This approach focuses on the long-term performance of nodes and downgrades nodes with unstable I / O rates. This design allows the scoring system to consider not only the current state of a node but also its historical stability and future predictability. As a result, data shards are prioritized for placement on nodes with sufficient resources, balanced load, and stable performance, thereby improving the write throughput and reliability of the entire distributed storage system.
[0049] Indicates storage node The current available capacity, in TB, is obtained in real time by monitoring agents deployed on each storage node by calling underlying operating system interfaces (such as the statfs system call in Linux) to query the available space of the file system. It is a dynamic field in the resource status set of the nodes to be scheduled. For example, a node... The available capacity is 4TB, then .
[0050] Indicates storage node The total capacity, measured in TB, is a static configuration value set by the administrator when a node joins the storage cluster and recorded in the cluster's metadata service. It represents the total size of the node's physical storage media. For example, a node... If the total capacity is configured to be 10TB, then .
[0051] Indicates storage node The current I / O write rate, measured in MB / s, is calculated by the monitoring agent through continuous collection of disk I / O statistics from the node (e.g., the sectors_written field in the / proc / diskstats file on Linux systems). Specifically, it is calculated by multiplying the total number of sectors written in the last 5 seconds by the sector size and then dividing by 5, to reflect the node's instantaneous write capability. For example, node... If the current write speed is 50 MB / s, then .
[0052] This represents the maximum I / O write rate across all nodes in the system, measured in MB / s. This value is calculated during each scheduling calculation by iterating through the resource status set of all nodes to be scheduled. The value is dynamically determined by taking the maximum value, and is used to normalize the I / O performance of each node. For example, if the I / O performance of three nodes in the cluster is... The speeds are 50 MB / s, 80 MB / s, and 65 MB / s respectively. .
[0053] and These are the capacity score weighting coefficient and the I / O performance score weighting coefficient, respectively. They are two dimensionless weighting parameters that satisfy... The setting of these two coefficients aims to balance the importance of storage capacity and write performance in node selection. Their values are determined based on regression analysis of historical system data. By simulating data placement strategies under different weight combinations and evaluating their impact on average write latency and storage balance, the optimal weight combination is selected. For scenarios like in-vehicle OBD data, which involve continuous inflow and are highly sensitive to write performance, analysis suggests setting the I / O performance weight slightly higher (i.e.,...). This will result in a better system response, therefore it is set to... and .
[0054] Indicates storage node The current number of concurrent write tasks is a dimensionless integer, calculated in real-time by the central scheduler or the task queue maintained by the node itself, representing the node's current load pressure. For example, a node... There are currently 3 write tasks being processed. .
[0055] This represents the load penalty factor, a dimensionless coefficient used to adjust the negative impact of concurrent tasks on node scores. Its value is set based on benchmark testing, by gradually increasing the number of concurrent write tasks on a single node. The I / O throughput changes were recorded. When the number of concurrent tasks reached 5, the throughput dropped to 50% of its peak, and a penalty target was established based on this. By solving this equation, ,have to Calculate .
[0056] Indicates storage node The historical I / O write rate standard deviation, measured in MB / s, measures the stability of node performance. It is calculated by the monitoring agent as the standard deviation of a sequence of I / O write rate data points collected every minute over the past hour. For example, node... If the standard deviation of the write rate sample over the past hour is 5 MB / s, then... .
[0057] This represents the stability penalty factor, a dimensionless coefficient used to adjust the severity of the penalty imposed on the score by I / O performance fluctuations. Its value is determined by analyzing historical failure data and... The correlation is used to determine the target, which is set as: when the relative volatility of the nodes... When the stability penalty term reaches 0.2 (a relatively high volatility level determined based on operational experience), the value should decay to... (Approximately 0.368). Based on this, the equation can be derived as follows: Solving for .
[0058] It is a very small constant used to prevent... A division-by-zero error occurs when the value is 0, as it is much smaller than the normal I / O rate. It is then set to... .
[0059] Calculation process: Currently targeting candidate storage nodes Calculate its node placement score. The parameters of the node are obtained from the resource status set of nodes to be scheduled as follows: TB, TB, MB / s, , MB / s.
[0060] Obtain from the system's global state: MB / s.
[0061] Use preset coefficients: , , , , .
[0062] Substitute the above values into the formula: ; Perform the calculation step by step: Calculate the basic resource scoring items: ; Calculate the load penalty factor: ; Calculate the stability penalty factor: ; Finally, multiply the three results together to obtain the final score: ; This result indicates that candidate storage nodes The node placement score is 0.203. This value is a relative, dimensionless suitability measure used to rank all candidate nodes. The higher the score, the more suitable the node is to receive new data shards at the current moment. The score takes into account the node's remaining space, write performance, current load, and historical stability. A lower score (as in this example) may mean that the node is performing poorly in one or more dimensions (e.g., although the basic resources are acceptable, the load and stability penalties are large).
[0063] Based on the node placement score calculated for each storage node in the preceding steps, the system initiates a deterministic data shard placement scheduling process. First, the system organizes all candidate storage nodes and their corresponding node placement scores into a list, and then sorts this list in descending order of score value, thus obtaining a node priority queue from best to worst. For example, if there are five nodes with scores {node-02: 0.85, node-04: 0.76, node-01: 0.51, node-05: 0.33, node-03: 0.20}, then the sorted priority queue is [node-02, node-04, node-01, node-05, ... Next, the system retrieves a batch from the queue of data shards to be placed. This batch contains all data shards and redundant shards belonging to the same original data block, for example, 9 shards (6 data shards and 3 redundant shards). Then, the system allocates storage nodes to each shard sequentially according to its shard numbering (from _ds_0 to _ds_5, then to _ps_0 to _ps_2). During allocation, the system retrieves a node from the head of the sorted node priority queue and assigns the current shard to it. Furthermore, to ensure high data availability, a key constraint is that different shards of the same original data block must be stored on different... On physical nodes, once a node is assigned to a shard, it is removed from the candidate node list for this batch of allocations. For example, after BLK_..._ds_0 is assigned to node-02, node-02 will no longer be used to allocate the remaining 8 shards of that data block. Then, BLK_..._ds_1 will be assigned to the next node in the queue, node-04, and so on, until all 9 shards are allocated to 9 different nodes with the highest scores. After each allocation, the system immediately records the binding relationship between the shard identifier and the identifier of the assigned node, for example, {"shard_id": "BLK_..._ds_0", "node_id": "node-02"}. When all shards in a batch have been allocated, these binding relationships are committed and persisted together, forming part of the distributed storage node layout table.
[0064] The steps to obtain the list of data blocks to be migrated are as follows: Scan each index entry in the hierarchical composite index, extract the data temperature identifier and corresponding data timestamp value from each index entry, calculate the difference between the data timestamp value and the current system time value, and determine if the data temperature identifier is hot data and the time difference is greater than the preset cooling threshold. Then mark the entry as an index entry to be migrated and generate a list of data blocks to be migrated.
[0065] Specifically, based on the hierarchical composite index, a scheduled task in the system background is periodically triggered, for example, every hour. After starting, this task systematically traverses all index entries, beginning with the root node. Specifically, it focuses on scanning entries whose paths contain " / hot_data / ". For each such entry, the task extracts its associated data block index number and queries a separate metadata database based on this number to obtain the timestamp value of when the data block was initially generated. This timestamp is a Unix timestamp accurate to milliseconds. Subsequently, the task obtains the current system's standard UTC time value and calculates the difference between these two timestamps, resulting in a time span in seconds. Finally, the task compares this calculated time difference with a... The system compares the data with a preset cooling threshold, which is set at 2,592,000 seconds, or 30 days. This threshold is based on the statistical analysis of historical data access patterns mentioned earlier. The analysis clearly indicates that the data is accessed most frequently within 30 days of its generation, and then drops sharply. Therefore, if the temperature identifier of the data parsed in the current index entry is indeed "hot data", and the calculated time difference is strictly greater than 2,592,000 seconds, the system determines that the data block has met the cooling conditions. Subsequently, the task marks the index entry (i.e., the index number of the data block) as an object to be migrated and adds it to a temporary list stored in memory. After the scheduled task has traversed all relevant index entries, this temporary list is solidified, forming a list of data blocks to be migrated.
[0066] The steps for obtaining the updated storage structure index are as follows: Based on the list of data blocks to be migrated, the index numbers of the data blocks in the list are extracted one by one. The distributed storage node layout table is queried one by one according to the index numbers of the data blocks to be migrated to determine the binding relationship between the data shards and storage nodes corresponding to each data block index number. The storage nodes are located one by one through the binding relationship and all data shard contents are read in sequence. The complete data content is reorganized according to the data shard order numbering to generate a complete original data block. Based on the complete original data block, according to the preset cold data encoding strategy, each original data block is re-divided into multiple data shards. Distributed storage nodes with the storage pool category of cold data storage pool are selected one by one to write the data shard content in sequence, and the binding mapping relationship between each data shard and the storage node written is recorded to generate the updated storage structure index.
[0067] Specifically, based on the list of data blocks to be migrated generated in the previous process, the system initiates a data reorganization and recovery process. This process processes each data block index number in the list one by one. For a number, such as "BLK_20231201_1000_001", the system first uses this number as the query key to initiate a query request to the distributed storage node layout table. This layout table is a key-value pair store, where the key is the data block index number and the value is a list containing all shards (including data shards and redundant shards) of the data block and their corresponding storage node identifiers. The query operation will return a list similar to [{"shard_id": "...", "node_id": "..."}, The system parses the structure to determine the locations of all the fragments needed to reconstruct the data block. For example, it determines that data fragments _ds_0 to _ds_5 are stored on nodes node-01 to node-06, and redundant fragments _ps_0 to _ps_2 are stored on nodes node-07 to node-09. Then, the system sends data read requests to these nodes in parallel. Each request contains a specific fragment identifier. After receiving the request, each storage node reads the corresponding fragment content from its local disk and transmits it back to the data reassembly service over the network. This service waits for all fragment data to arrive, or, in the event of a partial node failure, uses the redundant fragments to recover the lost data fragments using erasure coding decoding algorithms. Once all six original data fragments are collected, the system strictly follows the fragment number order (from 0 to 5) to concatenate the fragment content in memory. For example, the content of _ds_0 is placed at the beginning of the buffer, followed by the content of _ds_1, and so on, ultimately seamlessly combining them into a complete original data block that is completely consistent with the original data block before migration.
[0068] Based on the complete original data block generated in the previous process, the system initiates a data re-encoding and cold storage archiving process. This process first processes the recovered original data block according to a preset cold data encoding strategy. This cold data encoding strategy is specified as RS(12,2) in the system configuration, which uses the Reed-Solomon algorithm to divide the data block into 12 data fragments and 2 redundant fragments. This strategy is chosen based on cost-effectiveness considerations, because cold data is accessed very infrequently, and appropriately reducing redundancy can significantly save storage space. The system then follows this strategy to re-divide the complete original data block into 14 equal-length fragments (12 data fragments and 2 redundant fragments) and generate new data for them. Next, the system retrieves a list of nodes marked as "cold data storage pool" from the cluster's node manager. This storage pool consists of low-cost, high-capacity, but relatively slow I / O performance storage hardware. The system uses round-robin to select storage nodes for these 14 new shards one by one from this cold storage pool, ensuring that different shards of the same data block are placed on different physical nodes. Then, the system sends write commands to the selected cold storage nodes in sequence, writing the content of each newly generated shard to the target node's disk. After each successful write operation, the system records the binding mapping relationship between the new identifier of the shard and the identifier of the cold storage node being written, for example, {"shard_id": "...", "node_id": "cold_node_01"}. When all 14 shards of an original data block have been successfully written to the cold storage pool, all these newly generated binding mapping relationships are collected as a whole and updated into the system's metadata, forming the updated storage structure index.
[0069] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A distributed storage method for vehicle-mounted OBD data, characterized in that, Includes the following steps: Receive raw data blocks uploaded by the vehicle OBD terminal, parse DTC fault codes, vehicle speed, engine speed, gear information and timestamp, assign high criticality level identifiers to data blocks associated with rapid acceleration and deceleration events and DTC fault codes, and assign hot data or cold data identifiers based on the difference between the timestamp and the current time, and establish a data fidelity and criticality score set. Based on the data fidelity and criticality score set, the criticality level identifier, data temperature identifier and fidelity score of each data block are extracted and combined into a standardized multidimensional data feature vector. Then, according to the preset fidelity threshold, criticality level type and temperature type, the multidimensional data feature vector is mapped to the predefined and hierarchical index space address to establish a data hierarchical composite index. Based on the data hierarchical composite index, the data block encoding strategy selection instruction is obtained, the distributed storage system is scheduled to execute the data block encoding strategy selection instruction, the data block is sharded and redundantly encoded, and the encoded shards are written to the distributed storage nodes in different storage pools to generate a distributed storage node layout table. Start a background scheduled task to scan the timestamp information in the data hierarchical composite index, filter the index entries that have been cooled from hot data to cold data, summarize and generate a list of data blocks to be migrated, and establish an updated storage structure index based on the list of data blocks to be migrated and the distributed storage node layout table.
2. The distributed storage method for vehicle OBD data according to claim 1, characterized in that, The steps for obtaining the data fidelity and key score set are as follows: Extract the corresponding vehicle speed value, engine speed value, gear information, sensor self-test status value and data source reputation score from each raw data block. Input the engine speed value and gear information into the transmission ratio function and calculate the theoretical vehicle speed. Perform difference analysis between the theoretical vehicle speed and the vehicle speed value. At the same time, accumulate the abnormal count of the sensor self-test status value. Assign high criticality level identification to the data blocks associated with rapid acceleration and deceleration events and DTC fault codes to generate a set of parameters to be scored. Calculate the fidelity score for each data block based on the set of parameters to be scored; Based on the fidelity score, the criticality identifier of each data block, and the data temperature identifier, a data item identifier set is formed by combining them in sequence, and then grouped according to the index number of the original data block to generate a data fidelity and criticality score set.
3. The distributed storage method for vehicle-mounted OBD data according to claim 1, characterized in that, The steps for obtaining the multidimensional data feature vector are as follows: Based on the index number of each data item in the data fidelity and criticality score set, the criticality identifier, data temperature identifier, and fidelity score of the data item corresponding to the index number are extracted one by one. The criticality identifier is converted through binary encoding, and the data temperature identifier is converted into a standard value through the timeliness weight table to form an initial set of data feature values. Based on the initial set of data feature values, the fidelity threshold, key level type classification table, and temperature type classification table are called respectively. Classification is performed based on the difference between the fidelity score and the fidelity threshold. The key level identifier value is matched with the index position of the key level type classification table, and the data temperature identifier value is matched with the index position of the temperature type classification table. The matching results are combined to form a multidimensional data feature vector.
4. The distributed storage method for vehicle OBD data according to claim 1, characterized in that, The steps for obtaining the hierarchical composite index are as follows: Based on the multidimensional data feature vectors, the corresponding index space addresses are retrieved item by item according to the preset mapping rule table. The index space is located step by step using a step-by-step mapping method. The multidimensional data feature vectors are mapped and located to the addresses in the index space one by one, thus establishing a hierarchical composite index for data.
5. The distributed storage method for vehicle-mounted OBD data according to claim 1, characterized in that, The steps for obtaining the data block encoding strategy selection instruction are as follows: Based on the index space address of each index entry in the data hierarchical composite index, the data items corresponding to the address are retrieved one by one, the key level identifier and data temperature identifier contained in the data item are extracted, and the index matching conditions are formed based on the encoded value of the key level identifier and the classification value of the data temperature identifier, thereby generating a set of encoding strategy query conditions. Based on the set of query conditions for the coding strategy, the preset erasure coding strategy library is called one by one. The applicable key level identifier range and data temperature identifier range of each strategy in the erasure coding strategy library are compared one by one. Erasure coding strategies that meet the query conditions are filtered out, and the strategy index number of the erasure coding strategy that meets the conditions is determined to form an erasure coding strategy index set. Based on the erasure coding policy index set, erasure coding policy entries with corresponding policy index numbers are retrieved one by one from the erasure coding policy library. The encoding rules in the erasure coding policy entries are parsed, and the corresponding erasure coding instructions are output one by one according to the encoding rules to obtain the data block encoding policy selection instructions.
6. The distributed storage method for vehicle-mounted OBD data according to claim 1, characterized in that, The steps for obtaining the distributed storage node layout table are as follows: According to the data block encoding strategy, select instructions, perform equal-length splitting of data blocks according to the number of erasure coding fragments, map and bind each fragment identifier to the original data block identifier, and extract the total capacity value, available capacity value, current write rate value, historical write rate standard deviation value and current write task number of each candidate distributed storage node to generate a set of resource status of nodes to be scheduled. Based on the set of resource statuses of the nodes to be scheduled, calculate the node placement score for each storage node; Based on the node placement score of each storage node, all candidate nodes are sorted in descending order of node placement score. The nodes with the highest node placement scores are selected in sequence according to the data shard numbers. Each data shard is assigned to a matching node, and the binding relationship between the data shard and the storage node is recorded to generate a distributed storage node layout table.
7. The distributed storage method for vehicle OBD data according to claim 1, characterized in that, The steps for obtaining the list of data blocks to be migrated are as follows: Scan each index entry according to the data hierarchical composite index, extract the data temperature identifier and corresponding data timestamp value from each index entry, calculate the difference between the data timestamp value and the current system time value, and determine if the data temperature identifier is hot data and the time difference is greater than the preset cooling threshold. Then mark the entry as an index entry to be migrated and generate a list of data blocks to be migrated.
8. The distributed storage method for vehicle-mounted OBD data according to claim 1, characterized in that, The steps for obtaining the updated storage structure index are as follows: Based on the list of data blocks to be migrated, the index numbers of the data blocks in the list of data blocks to be migrated are extracted one by one. The distributed storage node layout table is queried one by one according to the index numbers of the data blocks. The binding relationship between the data shards and storage nodes corresponding to each data block index number is determined. The storage nodes are located one by one through the binding relationship and all data shard contents are read in sequence. The complete data content is reorganized according to the data shard order numbering to generate a complete original data block. Based on the complete original data block, according to the preset cold data encoding strategy, each original data block is re-divided into multiple data shards. Distributed storage nodes with the storage pool category of cold data storage pool are selected one by one to write the data shard content in sequence, and the binding mapping relationship between each data shard and the storage node written is recorded to generate an updated storage structure index.
Citation Information
Patent Citations
An elastic multi-dimensional redundancy method in a distributed storage system
CN109783016A
Data storage method and system based on legal knowledge service platform and storage medium
CN118964496A
Meteorological metadata storage method and system based on machine learning
CN120104579A
Index establishment method, and electronic device and computer-readable storage medium
WO2023241246A1
Cited By
Cloud computing platform and method based on data mining and big data analysis
CN121542326A