Vehicle and road cloud mass data safety management method and system based on distributed storage
By constructing a distributed data storage cluster in the vehicle-road-cloud system, data collection, layered encryption, and multi-node parallel verification are performed, solving the scalability and security issues of the vehicle-road-cloud data storage system and achieving stable and reliable storage and secure management of massive amounts of data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTELLIGENT INTER CONNECTION TECH CO LTD
- Filing Date
- 2026-03-04
- Publication Date
- 2026-05-12
AI Technical Summary
Existing vehicle-road-cloud data storage systems are inadequate in terms of scalability, security, and access verification efficiency, making it difficult to balance efficient storage and secure management of massive amounts of data, and they also face the risks of data silos and privacy leaks.
Data is collected by connecting to roadside sensing devices and vehicle terminals to generate datasets to be processed. A distributed addressing mechanism is used to determine the mapping relationship of storage nodes, a data storage cluster is constructed, and layered encryption processing is performed to achieve parallel verification and dynamic adjustment of multiple nodes and optimize data distribution.
It achieves stable and reliable storage of massive vehicle-road-cloud data, improves security and scalability, ensures data confidentiality and access controllability, and dynamically optimizes data distribution management.
Smart Images

Figure CN122027313A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data security technology, specifically to a method and system for managing massive amounts of vehicle-road-cloud data based on distributed storage. Background Technology
[0002] With the accelerated deployment of vehicle-road-cloud integration and intelligent connected vehicles, roadside sensing devices, vehicle terminals, and cloud platforms continuously generate massive amounts of heterogeneous data from multiple sources, including video, radar, trajectory, signal control, vehicle-to-everything (V2X) interaction, and user privacy data. This data exhibits characteristics of high-concurrency access, strong real-time performance, cross-domain sharing, and long-term retention. Existing centralized or traditional distributed storage systems face significant challenges in terms of petabyte-scale expansion, cross-node consistency, fault tolerance, and access performance. Furthermore, insufficient industry standards and multi-entity collaboration can easily lead to data silos. More importantly, vehicle-road-cloud data involves public safety and personal privacy, facing risks such as leakage, tampering, unauthorized access, key abuse, and decreased availability due to node failure. Therefore, there is an urgent need for a vehicle-road-cloud data security management solution that can simultaneously address massive high-concurrency storage, hierarchical security protection, efficient access verification, and dynamic optimization of storage distribution. This solution aims to solve the technical problems of poor storage scalability, rudimentary security protection, low access verification efficiency, and insufficient adaptability to data distribution and access patterns in existing technologies. Summary of the Invention
[0003] This application provides a method and system for secure management of massive vehicle-road cloud data based on distributed storage, which solves the technical problems in existing vehicle-road cloud data storage that make it difficult to balance efficient expansion of massive data, security and privacy protection, and storage cost optimization.
[0004] The first aspect of this application provides a method for secure management of massive vehicle-road-cloud data based on distributed storage, the method comprising:
[0005] The system integrates roadside sensing devices and vehicle-mounted terminals for real-time data collection, obtains multi-source traffic data, identifies its type, and generates a dataset to be processed. This dataset is then distributed and addressed to determine the mapping relationship between storage nodes. The dataset is distributed to multiple storage nodes according to this mapping relationship, constructing a data storage cluster. The data storage cluster undergoes layered encryption processing to generate a first encrypted storage fragment. Data access requests are received and parsed. Based on the parsing results, the first encrypted storage fragment is verified in parallel across multiple nodes, and a second encrypted storage fragment is read. The second encrypted storage fragment is then verified, and based on the verification results, the storage node mapping relationship is dynamically adjusted to generate a data distribution optimization management scheme.
[0006] A second aspect of this application provides a vehicle-road-cloud-based massive data security management system based on distributed storage, the system comprising:
[0007] Data Acquisition Unit: Connects to roadside sensing devices and vehicle-mounted terminals for real-time data acquisition, obtains multi-source traffic data, identifies its type, and generates a dataset to be processed; Data Distribution Unit: Distributes the dataset to be processed using distributed addressing, determines the storage node mapping relationship, and distributes the dataset to multiple storage nodes according to the storage node mapping relationship to construct a data storage cluster; Layered Encryption Unit: Performs layered encryption processing on the data storage cluster to generate a first encrypted storage fragment; Parallel Verification Unit: Receives and parses data access requests, performs multi-node parallel verification on the first encrypted storage fragment based on the parsing results, and reads a second encrypted storage fragment; Dynamic Adjustment Unit: Verifies the second encrypted storage fragment, and dynamically adjusts the storage node mapping relationship based on the verification results to generate a data distribution optimization management scheme.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0009] First, multi-source traffic data generated by roadside sensing devices and vehicle-mounted terminals are collected and classified in real time to form a unified dataset for processing. Then, a distributed addressing mechanism is used to determine the mapping relationship between data and storage nodes, distributing data fragments to multiple storage nodes to construct a distributed data storage cluster. Next, during the storage phase, the data is encrypted in a hierarchical manner, generating encrypted storage fragments to enhance security. When a data access request is received, the request is parsed, and the encrypted data is verified and read through a multi-node parallel verification mechanism. Finally, the integrity of the read data fragments is verified, and the storage node mapping relationship is dynamically optimized based on the verification results, achieving adaptive adjustment of data distribution and secure and efficient management. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of the process for a vehicle-road-cloud-based massive data security management method based on distributed storage, provided in an embodiment of this application.
[0012] Figure 2 This is a schematic diagram of the structure of a vehicle-road-cloud-based massive data security management system based on distributed storage, provided in an embodiment of this application.
[0013] Explanation of reference numerals in the attached figures: Data acquisition unit 11, data distribution unit 12, hierarchical encryption unit 13, parallel verification unit 14, dynamic adjustment unit 15. Detailed Implementation
[0014] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0015] Example 1, as Figure 1 As shown, this application provides a method for secure management of massive vehicle-road-cloud data based on distributed storage, wherein the method includes:
[0016] The system connects to roadside sensing devices and vehicle terminals for real-time data collection, obtains multi-source traffic data, identifies the types of traffic data, and generates a dataset to be processed.
[0017] In this embodiment, during the data acquisition phase, a vehicle-road-cloud data access gateway establishes communication connections with roadside sensing devices and vehicle-mounted terminals. These roadside sensing devices include, but are not limited to, cameras, millimeter-wave radar, lidar, and RSU roadside units, while the vehicle-mounted terminals include, but are not limited to, OBUs, T-Boxes, and vehicle-mounted cameras. Access methods can include Ethernet, fiber optics, 4G, 5G, and C-V2X, with protocol adaptation and data unpacking performed uniformly at the gateway. Subsequently, the system receives data packets reported by each device in real time according to a preset sampling period or event-triggered method, treating them as multi-source traffic data. Necessary descriptive fields are extracted from each data record in the multi-source traffic data as type identification information. These descriptive fields include at least the data source identifier, data category, timestamp, geographical location, data structure version number, and business scenario label. When the original data lacks a timestamp or coordinates, the access gateway's time source is used to supplement a unified time reference, and the location field is normalized based on the roadside deployment coordinate system or map reference. Next, basic quality processing is performed on the data records, including deduplication, format normalization, outlier filtering, field integrity verification, and necessary anonymization, such as converting sensitive information like license plates and faces into hash or mask features. The processed records are then aggregated and encapsulated according to data category, road segment / intersection, and time window. For structured data, such as target lists, trajectories, and traffic control data, they are merged into batches according to fixed time windows. For unstructured data, such as video clips, images, and point cloud files, SHA hashes are generated as content digests, and file indexes and metadata are recorded. Finally, the system generates a dataset to be processed in a unified data object format. This dataset includes at least the data body, type identifier metadata, integrity digest, source and permission markers, and the sequence number required for subsequent distribution, thus providing directly executable input for subsequent distributed addressing, fragmented storage, and encrypted management.
[0018] The dataset to be processed is distributed and addressed to determine the storage node mapping relationship. The dataset to be processed is then distributed to multiple storage nodes according to the storage node mapping relationship to build a data storage cluster.
[0019] In one embodiment, after generating the dataset to be processed, the system first obtains information about the currently online storage nodes, including parameters such as node identifier, network address, storage capacity, load status, and health status. The cluster management module then registers and performs heartbeat checks on the online nodes to form a list of available nodes. Subsequently, a distributed addressing mechanism is constructed by introducing consistent hashing or an improved hash algorithm. Each storage node is mapped to a logical hash ring according to its node identifier, generating a node distribution ring. Weights are assigned to different nodes based on their performance and capacity to achieve load balancing. Then, for each data record to be processed, the system performs a hash operation based on its unique data identifier to obtain the corresponding data hash value. Based on this hash value, the system searches for the first matching node clockwise on the hash ring to determine the primary storage node. Following a preset replica count strategy, the system continues to traverse subsequent nodes on the hash ring to determine several replica storage nodes, thus forming a set of primary and secondary nodes. Then, the system writes the correspondence between primary and replica storage nodes into a storage node mapping table, logically segments the dataset to be processed according to this mapping table, and sends the generated data shards to the corresponding primary and replica storage nodes for storage via a parallel transmission mechanism. Through the above-mentioned distributed addressing, primary-replica mapping, and parallel distribution mechanisms, a data storage cluster with horizontal scalability, load balancing, and high availability redundancy is constructed, achieving stable and reliable storage of massive traffic data in a multi-node environment.
[0020] Furthermore, the method involves distributing the dataset to be processed using distributed addressing to determine the storage node mapping relationship, and then distributing the dataset to be processed to multiple storage nodes according to the storage node mapping relationship to construct a data storage cluster.
[0021] A distributed file traversal approach is introduced to traverse online storage nodes, extracting multiple node identifiers and multiple network addresses. These node identifiers are then mapped to hash rings according to the network addresses, constructing a node distribution ring. The dataset to be processed is hashed based on the node distribution ring to obtain multiple data hash values. The hash ring is traversed according to these hash values to determine the primary storage node. The hash ring is then continuously traversed based on the primary storage node to determine the replica storage nodes. Storage node mapping analysis is performed based on the primary and replica storage nodes to construct a storage node mapping relationship table. The dataset to be processed is then segmented according to the storage node mapping relationship table, generating multiple data fragments. These multiple data fragments are simultaneously sent to the primary and replica storage nodes for storage, constructing the data storage cluster.
[0022] Preferably, the distributed file management module first periodically sends probe requests to each storage node, confirming the node's online status through a node registration service or heartbeat detection mechanism. When a node returns a response, the system records information such as the node's ID, device name, data center, network address, current remaining storage capacity, read / write load ratio, and most recent response time. Nodes in abnormal or overloaded states are marked as unavailable for allocation, thus forming a list of available nodes. After obtaining the list of available nodes, the system constructs a logical distribution structure. This involves concatenating each node's ID with its network address to form a unique node identifier string, which is then hashed to obtain an integer value. The system pre-defines a numerical range as the logical address space, for example, from zero to a maximum integer value. The calculated values for each node are then arranged in ascending order within this range, forming a closed hash ring, which serves as the node distribution ring. To prevent some nodes from handling excessive data, weighted allocation can be performed based on the node's remaining capacity. For example, nodes with larger remaining capacity can be repeatedly mapped multiple times on the node distribution ring, increasing their number of positions on the ring and allowing them to handle more data allocation tasks.
[0023] During data allocation, the system reads data records one by one from the dataset to be processed, extracting the unique identifier information of each record. For example, the device number, collection time, and data sequence number are concatenated into a string in a fixed order, and then hashed to obtain a data location value. The system places this value into the aforementioned node distribution ring and searches clockwise. That is, starting from the value position, it searches backward to find the first node position and determines this node as the primary storage node. If the value is greater than the values of all node positions on the ring, it searches again from the beginning of the logical ring to find the first node as the primary storage node. After determining the primary storage node, the system continues to search for subsequent nodes clockwise along the logical ring, selecting several different nodes as replica storage nodes in sequence. The number of replicas can be determined according to the system settings. For example, two different nodes after the primary storage node can be selected as replica storage nodes. If a node marked as unusable is encountered during the traversal, it is automatically skipped and the search continues forward until the replica number requirement is met. After the primary storage node and replica storage nodes are determined, the system generates a storage node mapping record, which clearly records the unique data identifier, the network address of the primary storage node, the list of network addresses of the replica storage nodes, the shard numbering rules, and the current version number. This mapping record is stored in the storage node mapping relationship table for subsequent querying and access.
[0024] The system then segments the data to be processed according to preset segmentation rules. If the data volume is small, it can be treated as a single data block; if the data volume is large, such as video clips or point cloud files, it is segmented according to a fixed size, for example, by dividing the total number of bytes by the set segment size to obtain the number of segments, and any remaining insufficient parts are formed into separate segments. Each segment generates a segment number and calculates an integrity check value for subsequent verification. During the data transmission phase, the system establishes parallel transmission connections with the primary storage node and replica storage nodes, simultaneously sending the corresponding segments to these nodes. After the primary storage node completes the write, it returns a write success confirmation message, and the replica storage nodes also return confirmation messages respectively; if a replica storage node fails to write, the system automatically searches for a new available node in the logical ring to supplement the replica and resends the segment. After the primary storage node and all replica storage nodes have confirmed the write success, the system marks the data as stored successfully. Through the specific processes of node traversal, logical ring construction, data location and lookup, primary replica selection, shard generation and parallel writing described above, the data is evenly distributed and redundantly backed up across multiple storage nodes. Ultimately, a distributed data storage cluster with horizontal scalability, automatic fault tolerance and high reliability is constructed, thereby improving the stability, scalability and data security of massive vehicle-road-cloud data storage.
[0025] The data storage cluster is subjected to hierarchical encryption processing to generate the first encrypted storage fragment.
[0026] In one embodiment, after data is written to the storage cluster, the system reads data type identification information from the dataset to be processed and uses this information to classify the data into security levels, determining data security level labels. These labels are categorized into three types: basic perception data, collaborative control data, and privacy-sensitive data. After classification, the system calls an encryption policy library for matching, determines the corresponding encrypted data blocks based on the current data security level labels, and combines and encapsulates these encrypted data blocks to construct a first encrypted storage shard. This first encrypted storage shard includes at least a shard header, a hierarchical ciphertext block area, and a metadata area. The shard header records the shard number, version number, and the identifier of the data record to which it belongs; the hierarchical ciphertext block area is arranged in the order of basic level, collaborative level, and privacy level; and the metadata area records the algorithm identifier, key index, access control condition digest, and integrity verification information. After this hierarchical encryption and encapsulation process, the first encrypted storage shard can be securely distributed to storage nodes and used for subsequent multi-node parallel verification and access reading, thereby improving the confidentiality and access controllability of massive amounts of vehicle-road-cloud data while ensuring performance.
[0027] Furthermore, the data storage cluster is subjected to layered encryption processing to generate a first encrypted storage shard, the method of which includes:
[0028] Extract data type identifiers from data records in the dataset to be processed; perform data parsing based on the data type identifiers to determine data security level labels, which include basic perception data level, collaborative control data level, and privacy data level; introduce an encryption policy library, and match the encryption policy library according to the basic perception data level, the collaborative control data level, and the privacy data level to determine multiple encrypted data blocks; combine and encapsulate the multiple encrypted data blocks to construct the first encrypted storage fragment.
[0029] Preferably, before performing layered encryption on the data, the data type identifier is first read from each data record in the dataset to be processed. This data type identifier is stored in the header or metadata area of the data record and is used to describe the source and content attributes of the data. For example, it records whether the data comes from a roadside sensing device or an on-board terminal, and whether it belongs to the category of video images, radar point clouds, target lists, trajectory information, signal control parameters, cooperative instructions, or event alarms. It may also include the collection time, collection location, data structure version number, and prompts indicating whether it contains sensitive fields. The system extracts the above identifiers in sequence and uses them as the basis for subsequent parsing and classification. After extracting the data type identifier, the system parses the data record, that is, it determines its business attributes according to the data category. For example, raw sensing data such as video images and point clouds are generally classified as sensing data; signal timing, vehicle-road cooperative control parameters, and issued instructions belong to the control category, which has a direct impact on traffic operation; and content that can point to the identity of an individual or vehicle, such as license plates, faces, precise trajectories, in-vehicle identity information, and account payment information, belongs to the privacy-sensitive category. The system further validates data by combining the source, scenario tags, and field list. For example, even with trajectory data, if it's only road segment statistical trajectory, it can be classified as basic perception level; if it's precise trajectory of a single vehicle and the travel pattern can be deduced, it's classified as privacy data level. Based on these rules, the system generates a unique data security level label for each data record. This label must include at least one of the following: basic perception data level, collaborative control data level, or privacy data level. When the same data record contains multiple field types, the system can split it at the field level and assign corresponding level labels to the different fields.
[0030] After determining the hierarchical labels, the system imports an encryption policy library and performs matching. This library pre-stores multiple encryption policy entries, each corresponding to at least one data security level and including encryption method selection, key strength requirements, key update cycle, whether access control conditions need to be bound, and encapsulation format requirements. The system traverses the policy library using the data security level label as the search condition. When the label is for the basic awareness data level, a low-computation, high-throughput encryption policy is matched; when the label is for the collaborative control data level, a higher-key-strength, more frequent-key-rotation encryption policy is matched; when the label is for the privacy data level, an encryption policy with access control conditions is matched, ensuring that the ciphertext can only be decrypted when the permission conditions are met. After matching, the system splits the data records into several encrypted data blocks according to the policy requirements. These encrypted data blocks refer to the smallest encrypted units after being divided according to the security level. For example, basic fields form basic-level encrypted data blocks, control parameters form collaborative-level encrypted data blocks, and sensitive fields form privacy-level encrypted data blocks. Each encrypted data block is accompanied by corresponding encrypted metadata after encryption, including the algorithm type marker, key index marker, data block number, generation time, and integrity verification information, for subsequent verification and decryption.
[0031] The system then combines and encapsulates multiple encrypted data blocks. In this process, the system first creates a fragment header, writing the data record identifier, fragment number, version number, and number of blocks to which the fragment belongs. Then, it writes the encrypted data blocks into the fragment content area in a preset order; for example, writing the basic-level encrypted data blocks first, then the collaborative-level encrypted data blocks, and finally the privacy-level encrypted data blocks, or arranging them in ascending order of block number. At the fragment tail or in a separate metadata area, the system summarizes and writes the algorithm identifier, key index, access control condition digest, and overall integrity check value for each encrypted data block. The first encrypted storage fragment formed after encapsulation can then be used as the object for subsequent storage and access verification, ensuring that data of different sensitivity levels are protected with different strengths, and guaranteeing the verifiability and traceability of the fragment during transmission and storage between distributed nodes.
[0032] Furthermore, an encryption policy library is introduced. Multiple encrypted data blocks are determined by matching the encryption policy library according to the basic perception data level, the collaborative control data level, and the privacy data level. The method includes:
[0033] When the data security level label is the basic perception data level, a symmetric encryption algorithm is used to traverse the encryption policy library according to the first key length for encryption processing, generating a basic-level encrypted data block; when the data security level label is the collaborative control data level, a symmetric encryption algorithm is used to traverse the encryption policy library according to the second key length for encryption processing, generating a collaborative-level encrypted data block, wherein the second key length is greater than the first key length; when the data security level label is the privacy data level, an attribute-based encryption algorithm is used to traverse the encryption policy library according to the access control policy tree for encryption processing, generating a privacy-level encrypted data block, wherein the access control policy tree contains multiple attribute condition combinations.
[0034] Optionally, after determining the data security level label, the system will proceed to select an encryption strategy based on the level and generate encrypted data blocks. During this process, the system pre-establishes an encryption strategy library. Each strategy in the library records information such as the applicable data level, recommended encryption method, key strength requirements, key acquisition method, key update cycle, encapsulation field rules, and required access control condition templates. The system compares each strategy in the library using the data security level label as an index to find a strategy entry that matches the current data level and is in an enabled state. When the data security level label is the basic perception data level, the system first extracts basic perception data blocks from the content to be encrypted. Examples include aggregated information such as target quantity statistics, vehicle speed, and traffic flow, as well as fields that do not directly point to personal identity, such as non-sensitive target bounding box coordinates. Subsequently, the system obtains a symmetric key of the first key length according to the requirements of the corresponding entry in the strategy library. This key can be generated by the key management service or retrieved from the key pool. The system checks whether the key is valid; if it is about to expire, an update is triggered. During encryption, the basic perception data block is organized into the content to be encrypted according to the byte sequence, and processed according to the rules of symmetric encryption algorithms. That is, the data block is first padded to meet the block length requirements of the encryption algorithm, and then the data block is encrypted segment by segment with the first key to generate ciphertext. At the same time, a random initialization vector or random perturbation parameter is generated and bound to the corresponding ciphertext. After encryption, the system forms a basic-level encrypted data block, which at least contains the ciphertext content, algorithm identifier, initialization vector information, first key index marker, and checksum for integrity verification.
[0035] When the data security level label is "cooperative control data level," the system extracts cooperative control-related content as the object to be encrypted, such as signal timing parameters, vehicle-road cooperative control instructions, priority passage control information, and hazard warning control fields. Because this type of data is more sensitive to tampering and leakage, the system still uses symmetric encryption to ensure real-time performance. However, during policy matching, it prioritizes higher-strength policy entries and obtains a symmetric key according to the second key length. The second key length is greater than the first key length, resulting in a longer key and a larger combination space, thus improving resistance to brute-force attacks. The system also obtains or generates the second key through the key management service and checks its validity and rotation cycle, forcibly rotating to a new key when necessary. Next, symmetric encryption is performed on the cooperative control data block. The overall steps are similar to those for basic sensing data, but the encryption parameters are more stringent, such as stronger encryption mode requirements, higher random parameter update frequency, and more comprehensive verification fields. After encryption, a cooperative-level encrypted data block is generated. This data block includes at least the cooperative control ciphertext, algorithm identifier, initialization vector information, second key index marker, checksum, and version number or sequence number of the control data for consistency verification during subsequent access.
[0036] When the data security level label is set to privacy data level, the system extracts privacy-sensitive fields to form privacy data blocks, such as license plate information, facial features, precise trajectory point sequences, vehicle identity association fields, and account identifiers. This type of data not only requires confidentiality but also strict decryption according to permissions. Therefore, the system uses attribute-based encryption for processing and introduces an access control policy tree. This access control policy tree describes which attributes are satisfied to allow decryption. The system generates or selects a policy tree template based on business rules and specifies it as a combination of multiple attribute conditions, such as: the access subject is a traffic management department with evidence collection authority; the access time is within the authorized time window and the access area matches the scope of responsibility; the request operation type is a query and the purpose is law enforcement auditing, etc. Then, the system traverses the encryption policy library, filters out policy entries that support attribute-based encryption and match the privacy data category, and writes the access control policy tree into the parameters of the policy entry as the control conditions for this encryption. When performing privacy encryption, the system first generates or obtains the public parameters required for attribute-based encryption, and the key management or permission management module generates the corresponding attribute private key for qualified legitimate users. The system takes privacy-preserving data blocks as plaintext input and performs encryption operations constrained by public parameters and an access control policy tree. This ensures that the resulting ciphertext is inherently bound to the policy tree conditions. Even if the ciphertext is copied to another location, only the private key of a user whose attribute set satisfies the policy tree can decrypt it to obtain the plaintext. After encryption, a privacy-level encrypted data block is generated. This block includes at least the privacy-preserving ciphertext, policy tree description information or its digest, algorithm identifier, public parameter index or version number, and integrity check value. To prevent policy tree tampering, the system also generates a check value for the policy tree digest and the entire ciphertext and writes it into the metadata. Through this hierarchical policy matching and encryption execution process, the system achieves low-overhead symmetric encryption protection for basic perception data, higher-strength symmetric encryption protection for collaborative control data, and attribute-based encryption protection for privacy data bound to access conditions. This ensures differentiated security protection and access controllable by permissions while maintaining the throughput and real-time performance of massive vehicle-road-cloud data.
[0037] The system receives and parses data access requests, performs multi-node parallel verification on the first encrypted storage fragment based on the parsing results, and then reads the second encrypted storage fragment.
[0038] In one embodiment, upon receiving a data access request, the system first parses the request content, extracting key fields such as the visitor's identity information, digital certificate, requested data identifier, and operation type, and performs basic legality verification. After successful parsing, the system queries the metadata management node for the corresponding storage node mapping relationship and the location of the first encrypted storage fragment based on the data identifier. Subsequently, the system simultaneously sends verification requests to multiple storage nodes to perform parallel verification of the corresponding first encrypted storage fragment. After each node returns its verification results, the system summarizes and judges them. When the preset verification pass conditions are met, the system selects the node with a normal status and fastest response to read the data and obtain the corresponding second encrypted storage fragment, thereby completing a secure and reliable data access process.
[0039] Furthermore, the method for receiving and parsing data access requests, performing multi-node parallel verification on the first encrypted storage fragment based on the parsing results, and reading the second encrypted storage fragment includes:
[0040] The system receives data access requests through the access interface of the distributed storage cluster; it parses the data access requests to determine the request header information and request body information. The request header information includes the visitor's digital certificate and identity identifier, and the request body information includes a list of data record identifiers to be accessed and the request operation type. Based on the data record identifier list, it initiates a batch query request to the metadata management node to obtain the data storage node mapping relationship and encrypted metadata. It sends the visitor's digital certificate to the blockchain verification network for parallel verification to generate a signature verification result. Based on the signature verification result, it performs verification return statistics to determine the number of verification nodes. Based on the number of verification nodes, it performs judgment analysis and performs multi-node parallel verification on the first encrypted storage shard according to the analysis results, and reads the second encrypted storage shard.
[0041] Optionally, the system sets up a unified access interface outside the data storage cluster as the entry point. This interface can be deployed on a gateway server or in a cluster front-end service to receive data access requests from traffic management platforms, operation platforms, or authorized applications. Upon arrival of an access request, the entry point first performs basic verification, such as checking the completeness of the request format, the presence of necessary fields, and whether the request body has been truncated. It also records the arrival time and source address for auditing purposes. After successful verification, the system proceeds to the parsing process. During the parsing phase, the system splits the data access request into request header information and request body information. The request header information at least extracts the visitor's digital certificate, visitor identity identifier, and request signature information. The digital certificate proves the visitor's identity; the identity identifier can be an organization number, user number, or system account. The signature information proves that the request has not been tampered with. The request body information at least includes a list of data record identifiers to be accessed and a request operation type. The data record identifier list describes which data records need to be accessed, such as a set of record numbers corresponding to data collected during a specific time period, at a specific intersection, or by a specific device. The request operation type describes the nature of the access behavior, such as query, download, playback, or statistical analysis. The system writes the above parsing results to the audit log and performs preliminary compliance checks on the request operation type. For example, it prohibits unauthorized entities from performing batch downloads or sensitive playback operations.
[0042] Subsequently, a batch query request is initiated to the metadata management node based on the list of data record identifiers. This query request can carry multiple data record identifiers at once to reduce frequent network interactions. After receiving the query request, the metadata management node returns storage location information and security information corresponding to each data record identifier, including at least the data storage node mapping relationship and encrypted metadata. Based on this, the system obtains a list of shards to be accessed, clarifying which nodes to access and which first encrypted storage shards to read, while simultaneously obtaining the metadata information required for decryption and verification. After completing the location, the system performs trusted verification of the visitor's identity. Specifically, the visitor's digital certificate in the request header is sent to the blockchain verification network for parallel verification. This blockchain verification network consists of multiple verification nodes, each of which stores trusted records or hash digests related to certificate issuance, revocation, and renewal, and can independently verify the authenticity and validity of certificates. The system broadcasts or distributes the same certificate verification task to multiple verification nodes simultaneously. The verification nodes respectively perform certificate chain verification, certificate validity check, revocation status check, and signature correctness check, and return their respective verification conclusions to the system, forming a signature verification result set.
[0043] Afterwards, the system performs statistical analysis on the returned signature verification results, calculating the number of nodes that passed verification as the verification node count. For example, nodes marked as having valid certificates and correct signatures in the returned results are counted as passed nodes, while nodes marked as having invalid certificates, revoked certificates, mismatched signatures, or unverifiable signatures are counted as failed nodes. The number of nodes with timeout responses can also be recorded as an anomaly reference. After obtaining the verification node count, the system enters the judgment and analysis phase. If the number of passed verification nodes reaches or exceeds the preset verification threshold, the digital certificate verification is deemed successful, allowing continued data reading; if the threshold is not reached, verification is deemed a failure, and the access request is directly rejected. A rejection reason is generated and returned to the requester, and the event is written to the audit log. When the certificate verification is successful, the system initiates multi-node parallel verification of the first encrypted storage shard based on the aforementioned list of shards to be accessed. During this process, the system simultaneously sends verification commands to the corresponding primary storage node and replica storage node, requesting the nodes to return the shard version number, integrity verification comparison result, and readability status. Upon receiving the instruction, each storage node reads the metadata area of the first encrypted storage fragment from its local machine, extracts the original integrity check value, recalculates the real-time check value for the stored ciphertext content, and then compares the two. If the comparison matches, verification passes; otherwise, verification fails. After aggregating the results from multiple nodes, the system selects the node that has passed verification, responds faster, has lower load, and is updated in version as the actual reading node, establishes a reading connection, and pulls the corresponding data content to obtain the second encrypted storage fragment. If the primary node fails verification or reading fails, the system automatically switches to a verified replica node to continue reading, ensuring uninterrupted access. Through the above process, a closed-loop mechanism is achieved: first, location; then, certificate on-chain parallel verification; then, threshold-based access; and finally, multi-node parallel verification and reading. This ensures that the visitor's identity is trustworthy, the request has not been tampered with, and that the encrypted fragment read from the distributed cluster has integrity and availability in a multi-replica environment, thereby securely obtaining the second encrypted storage fragment for subsequent verification and decryption.
[0044] Furthermore, the method for determining and analyzing based on the number of verification nodes includes:
[0045] If the number of verification nodes exceeds a preset verification pass threshold, the digital certificate verification is deemed successful, and a first judgment result is generated. Based on the first judgment result, the metadata management nodes are traversed to associate data record identifiers and determine the access control policy. The identity identifier and the request operation type are matched and compared with the access control policy. When the access permission matches successfully, the storage node mapping relationship is extracted and matched with the data to be accessed, generating an access node list. Data reading connections are made based on the access node list and multiple storage nodes to generate the analysis result.
[0046] Optionally, after receiving multiple certificate verification conclusions from the blockchain verification network, the system first summarizes and counts the results of the nodes that have passed verification. The system categorizes the conclusions returned by each verification node into two types: pass and fail. A node is considered pass if it returns a valid certificate and a correct signature; it is considered fail if it returns an expired, revoked, inconsistent, or unverifiable certificate. Nodes that do not return a conclusion within the timeout period are treated as fail or abnormal nodes. The total number of verification nodes is obtained after the statistics are completed. If this number exceeds a preset verification pass threshold, the system determines that the digital certificate verification is successful and generates a first judgment result. This first judgment result includes a verification success identifier, the number of pass nodes, the verification time, and the tracking number of this request, and is written to the audit log as the basis for subsequent authorization judgments. After obtaining the initial judgment result, the system does not directly read the data. Instead, it enters the permission policy determination stage. In this stage, based on the list of data record identifiers carried in the request subject, the system initiates a correlation query to the metadata management node, either individually or in batches, to find the metadata entry corresponding to each data record identifier. From this entry, it extracts access control-related information, such as the data category, region or road segment, data generating unit, sensitivity level label, allowed access subject type, allowed operation type range, access validity period, and whether dual approval or logging is required. The system merges and summarizes this information to form an access control policy for this request. If a single request contains multiple records with different sensitivity levels, the system will integrate them using the strictest policy as the priority. For example, if the request contains privacy-level data, the access policy for privacy data will be applied to prevent permissions from being amplified by lower-level policies.
[0047] Subsequently, the system matches the identity identifier parsed from the request header with the request operation type in the request body against the determined access control policy. This process first checks whether the entity category to which the identity identifier belongs is within the allowed range (e.g., whether it is a traffic management department, road operator, or authorized service provider). Then, it checks whether the request operation type is permitted. For example, if the policy only allows querying or statistics but not downloading or playback, the disallowed operation will be directly rejected. Additional conditions are also checked, such as whether the access time is within the authorized time window, whether the access area is within the scope of responsibility, and whether the user possesses a specific position or permission level. Only when all necessary conditions are met does the system determine that the access permission has been successfully matched; otherwise, it generates a rejection reason and terminates subsequent readings, while simultaneously writing the rejection event to the audit log.
[0048] Once access permissions are successfully matched, the system begins determining the specific read nodes. At this point, it extracts the storage node mapping relationships from the information returned by the metadata management node, and then matches the identifier of the record data to be accessed with the mapping relationship table to obtain all shard locations involved in this request. Based on this, an access node list is generated. This access node list not only contains node addresses but also priority information; for example, it prioritizes primary nodes, and if the primary node is overloaded, it selects replica nodes. It also records the shard number, version number, and required verification information to facilitate subsequent parallel reading and verification. Then, the system sends read requests in parallel to multiple nodes in the access node list. Each read request carries the shard number and version information, and the node returns the corresponding encrypted shard or encrypted data block. The system summarizes the results returned by each node and combines the shards according to the order in the metadata to restore the data set required for this request. If a node fails to connect or returns a timeout during the reading process, the system automatically switches to another replica node of the same shard to continue reading. If a shard has multiple versions, the system selects the one with the higher version number and successful verification as the final valid shard. Finally, the system compiles the information such as the successfully read shard set, read node status, failure and switching status into analysis results and outputs them to the upper-layer business module. At the same time, it records the complete link information in the audit log to meet compliance traceability and security audit requirements.
[0049] The second encrypted storage shard is verified, and the storage node mapping relationship is dynamically adjusted based on the verification result to generate a data distribution optimization management scheme.
[0050] In one embodiment, after reading the second encrypted storage fragment, the system first verifies the fragment's authenticity and integrity. This is achieved by integrating the target encryption algorithm type and key index field read from the metadata area of the second encrypted storage fragment, and then using the integration result to index the key management service to determine the required target encryption key. Subsequently, based on this target encryption key, the system obtains the original integrity verification value through a secure channel, obtains the real-time integrity verification value through hash calculation, and then compares these two integrity verification values to generate a verification result. After obtaining the verification result, the system enters a reverse dynamic adjustment phase. Specifically, the system analyzes the current storage node mapping relationship by combining verification success or failure, node response latency, read success rate, node load status, and other operational metrics. If the verification result is successful and the node is running stably, the node is recorded as a healthy node, maintaining its weight in the mapping relationship. If a primary storage node or replica node repeatedly experiences verification failures, read anomalies, or response timeouts, it is marked as a risk node or abnormal node, its assigned weight is reduced, or it is temporarily removed from the list of available nodes. When adjustments are needed, the system recalculates the mapping relationship of some data shards based on the existing hash ring or node distribution structure. For example, for an abnormal master node, the system searches for the next healthy node clockwise in the hash ring and promotes it to the new master storage node. Simultaneously, to ensure the number of replicas remains unchanged, it continues searching for suitable nodes to supplement the new replica nodes. Afterward, a data migration process is triggered, copying the valid shards from the original abnormal node to the newly determined node and updating the storage node mapping table in the metadata management node to reflect the new master-slave node correspondence. Furthermore, the system can also formulate data distribution optimization management schemes based on statistical analysis results. For example, when some nodes are chronically overloaded while other nodes are idle, the number of virtual nodes can be adjusted or the hash ring distribution can be rebalanced to make the data more evenly distributed across nodes. When the data access frequency in a certain area increases significantly, replica nodes can be added near that area to shorten the access path. When a type of highly sensitive data frequently triggers anomaly verification, the number of replicas can be increased or the security level of the storage node can be improved. Ultimately, the system generates a data distribution optimization management scheme that includes node status evaluation results, mapping relationship adjustment records, data migration plans, and load balancing optimization strategies, and writes it into the system configuration and management log. This enables adaptive dynamic optimization driven by verification results, thereby improving the security, reliability, and operational efficiency of the distributed vehicle-road-cloud data storage system.
[0051] Furthermore, the method for verifying the second encrypted storage fragment includes:
[0052] The following steps are taken: First, the encryption algorithm identifier field is read from the metadata area of the second encrypted storage segment to identify the target encryption algorithm type. Then, the key index field is read from the metadata area of the second encrypted storage segment. Based on the target encryption algorithm type and the key index field, a key acquisition request is determined. The key acquisition request is sent to the key management service via a secure communication channel for index lookup to determine the target encryption key. The target encryption key is returned to the verification module via a secure channel for decryption to generate the original data segment. The original integrity verification value is extracted from the second encrypted storage segment. A hash operation is performed on the original data segment to generate a real-time integrity verification value. The real-time integrity verification value is compared with the original integrity verification value to generate the verification result.
[0053] Preferably, after successfully reading the second encrypted storage fragment from the storage node, the system first reads the encryption algorithm identifier field from the metadata area of the second encrypted storage fragment. This field indicates the encryption method used by the fragment, such as whether it is a symmetric encryption method or an encryption method with access control conditions, and further specifies the specific algorithm category or mode. Based on this identifier field, the system determines the target encryption algorithm type to be used and loads the corresponding decryption process and parameter parsing rules accordingly. Then, the system continues to read the key index field from the metadata area. The key index field is not the key itself, but is used to locate the target key's number or index information in the key management service. For example, it may include the key number, key version number, key pool flag, and key validity period. Next, the target encryption algorithm type and the key index field are integrated to form a key acquisition request. If necessary, the system may attach the identity credentials or authorization ticket for this access so that the key management service can perform permission verification. After generating the key acquisition request, the system sends the request to the key management service through a secure communication channel. The secure communication channel can be a two-way authenticated encrypted channel to ensure that the request content is not eavesdropped on or tampered with during transmission. Upon receiving a key retrieval request, the key management service first verifies the caller's identity and permissions to confirm their eligibility to request the shard key. If verification is successful, it performs an index lookup in the key index table using the key index field to locate the target encryption key. If multiple versions of the key exist, the key management service selects the matching version based on the version number or validity period requirement specified in the request. If the request specifies that only the currently valid version must be used, the latest and valid key is returned.
[0054] Afterwards, the key management service returns the target encryption key through the same secure channel, along with the key's version and validity period information, for system recording and auditing. Upon obtaining the target encryption key, decryption is performed on the ciphertext content area of the second encrypted storage fragment according to the previously determined target encryption algorithm type, yielding the original data fragment. If key mismatch, abnormal ciphertext structure, or inconsistent algorithm parameters occur during decryption, an abnormal verification result indicating decryption failure is generated directly, and the reason for failure is recorded for subsequent tracking. After successfully decrypting and obtaining the original data fragment, integrity verification begins. During this process, the original integrity verification value is extracted from the metadata area of the second encrypted storage fragment. This original integrity verification value is calculated during fragment generation or encryption encapsulation and represents the characteristic value of the original data content at that time; it is typically the result of a hash calculation on the original data. Next, the original data fragments obtained from decryption are re-hashed. This involves inputting the content of each fragment byte-by-byte into the hash calculation process according to fixed rules. After multiple rounds of mixing and compression calculations, a fixed-length real-time integrity check value is output. This check value is extremely sensitive to the data content; any change in the data will result in a significantly different output value. Finally, the real-time integrity check value is compared with the original integrity check value. If they match perfectly, it means the decrypted data is consistent with the data generated, and the fragments have not been tampered with or damaged during storage and transmission, resulting in a successful verification result. If they do not match, it indicates that the fragment content may have been tampered with, some data may be missing, or a storage error may have occurred, resulting in a failed verification result, along with a marker indicating the reason for the difference. The system writes this verification result to the audit log and uses it as input for subsequent dynamic adjustments to storage node mapping relationships and data distribution optimization.
[0055] In summary, the embodiments of this application have at least the following technical effects:
[0056] First, roadside sensing devices and vehicle-mounted terminals collect multi-source traffic data in real time, identify the data types, and generate a dataset to be processed. Then, the dataset is distributed and addressed to determine the storage node mapping relationship. The dataset is then distributed to multiple storage nodes according to this mapping relationship, constructing a data storage cluster. Next, the data storage cluster undergoes layered encryption processing to generate a first encrypted storage shard. Then, data access requests are received and parsed. Based on the parsing results, the first encrypted storage shard is verified in parallel across multiple nodes, and a second encrypted storage shard is read. Finally, the second encrypted storage shard is verified, and based on the verification results, the storage node mapping relationship is dynamically adjusted to generate a data distribution optimization management scheme. This solves the technical problems in existing vehicle-road-cloud data storage that make it difficult to simultaneously achieve efficient expansion of massive amounts of data, security and privacy protection, and storage cost optimization. It achieves high availability and elastic expansion of distributed data storage, ensuring data security throughout its entire lifecycle through layered encryption, and dynamically optimizing data distribution based on access verification to reduce storage costs and improve access efficiency.
[0057] Example 2 is based on the same inventive concept as the vehicle-road-cloud massive data security management method based on distributed storage in the previous examples, such as... Figure 2 As shown, this application provides a vehicle-road-cloud-based massive data security management system based on distributed storage, wherein the system includes:
[0058] Data acquisition unit 11: Connects to roadside sensing devices and vehicle terminals for real-time data acquisition, obtains multi-source traffic data, identifies its type, and generates a dataset to be processed; Data distribution unit 12: Distributes the dataset to be processed using distributed addressing, determines the storage node mapping relationship, and distributes the dataset to multiple storage nodes according to the storage node mapping relationship to construct a data storage cluster; Layered encryption unit 13: Performs layered encryption processing on the data storage cluster to generate a first encrypted storage fragment; Parallel verification unit 14: Receives and parses data access requests, performs multi-node parallel verification on the first encrypted storage fragment based on the parsing results, and reads a second encrypted storage fragment; Dynamic adjustment unit 15: Verifies the second encrypted storage fragment, and dynamically adjusts the storage node mapping relationship based on the verification results to generate a data distribution optimization management scheme.
[0059] Furthermore, the data distribution unit 12 is configured to perform the following method:
[0060] A distributed file traversal approach is introduced to traverse online storage nodes, extracting multiple node identifiers and multiple network addresses. These node identifiers are then mapped to hash rings according to the network addresses, constructing a node distribution ring. The dataset to be processed is hashed based on the node distribution ring to obtain multiple data hash values. The hash ring is traversed according to these hash values to determine the primary storage node. The hash ring is then continuously traversed based on the primary storage node to determine the replica storage nodes. Storage node mapping analysis is performed based on the primary and replica storage nodes to construct a storage node mapping relationship table. The dataset to be processed is then segmented according to the storage node mapping relationship table, generating multiple data fragments. These multiple data fragments are simultaneously sent to the primary and replica storage nodes for storage, constructing the data storage cluster.
[0061] Furthermore, the layered encryption unit 13 is used to perform the following method:
[0062] Extract data type identifiers from data records in the dataset to be processed; perform data parsing based on the data type identifiers to determine data security level labels, which include basic perception data level, collaborative control data level, and privacy data level; introduce an encryption policy library, and match the encryption policy library according to the basic perception data level, the collaborative control data level, and the privacy data level to determine multiple encrypted data blocks; combine and encapsulate the multiple encrypted data blocks to construct the first encrypted storage fragment.
[0063] Furthermore, the layered encryption unit 13 is used to perform the following method:
[0064] When the data security level label is the basic perception data level, a symmetric encryption algorithm is used to traverse the encryption policy library according to the first key length for encryption processing, generating a basic-level encrypted data block; when the data security level label is the collaborative control data level, a symmetric encryption algorithm is used to traverse the encryption policy library according to the second key length for encryption processing, generating a collaborative-level encrypted data block, wherein the second key length is greater than the first key length; when the data security level label is the privacy data level, an attribute-based encryption algorithm is used to traverse the encryption policy library according to the access control policy tree for encryption processing, generating a privacy-level encrypted data block, wherein the access control policy tree contains multiple attribute condition combinations.
[0065] Furthermore, the parallel verification unit 14 is used to perform the following method:
[0066] The system receives data access requests through the access interface of the distributed storage cluster; it parses the data access requests to determine the request header information and request body information. The request header information includes the visitor's digital certificate and identity identifier, and the request body information includes a list of data record identifiers to be accessed and the request operation type. Based on the data record identifier list, it initiates a batch query request to the metadata management node to obtain the data storage node mapping relationship and encrypted metadata. It sends the visitor's digital certificate to the blockchain verification network for parallel verification to generate a signature verification result. Based on the signature verification result, it performs verification return statistics to determine the number of verification nodes. Based on the number of verification nodes, it performs judgment analysis and performs multi-node parallel verification on the first encrypted storage shard according to the analysis results, and reads the second encrypted storage shard.
[0067] Furthermore, the parallel verification unit 14 is used to perform the following method:
[0068] If the number of verification nodes exceeds a preset verification pass threshold, the digital certificate verification is deemed successful, and a first judgment result is generated. Based on the first judgment result, the metadata management nodes are traversed to associate data record identifiers and determine the access control policy. The identity identifier and the request operation type are matched and compared with the access control policy. When the access permission matches successfully, the storage node mapping relationship is extracted and matched with the data to be accessed, generating an access node list. Data reading connections are made based on the access node list and multiple storage nodes to generate the analysis result.
[0069] Furthermore, the dynamic adjustment unit 15 is used to perform the following method:
[0070] The following steps are taken: First, the encryption algorithm identifier field is read from the metadata area of the second encrypted storage segment to identify the target encryption algorithm type. Then, the key index field is read from the metadata area of the second encrypted storage segment. Based on the target encryption algorithm type and the key index field, a key acquisition request is determined. The key acquisition request is sent to the key management service via a secure communication channel for index lookup to determine the target encryption key. The target encryption key is returned to the verification module via a secure channel for decryption to generate the original data segment. The original integrity verification value is extracted from the second encrypted storage segment. A hash operation is performed on the original data segment to generate a real-time integrity verification value. The real-time integrity verification value is compared with the original integrity verification value to generate the verification result.
[0071] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for secure management of massive vehicle-road-cloud data based on distributed storage, characterized in that: The method includes: The system connects to roadside sensing devices and vehicle terminals to collect data in real time, obtains multi-source traffic data, identifies the types of traffic data, and generates a dataset to be processed. The dataset to be processed is distributed and addressed to determine the storage node mapping relationship. The dataset to be processed is then distributed to multiple storage nodes according to the storage node mapping relationship to build a data storage cluster. The data storage cluster is subjected to layered encryption processing to generate the first encrypted storage fragment; The system receives and parses data access requests, performs multi-node parallel verification on the first encrypted storage fragment based on the parsing results, and then reads the second encrypted storage fragment. The second encrypted storage shard is verified, and the storage node mapping relationship is dynamically adjusted based on the verification result to generate a data distribution optimization management scheme.
2. The method for secure management of massive vehicle-road-cloud data based on distributed storage as described in claim 1, characterized in that, The method involves distributing the dataset to be processed using distributed addressing to determine the storage node mapping relationship, and distributing the dataset to be processed to multiple storage nodes according to the storage node mapping relationship to construct a data storage cluster. A distributed file traversal method is introduced to traverse online storage nodes and extract multiple node identifiers and multiple network addresses. Map multiple node identifiers to a hash ring based on the multiple network addresses to construct a node distribution ring; Based on the node distribution ring, perform a hash operation on the dataset to be processed to obtain multiple data hash values; The hash ring is traversed based on the multiple data hash values to determine the primary storage node. The hash ring is then traversed based on the primary storage node to determine the replica storage node. Based on the primary storage node and the replica storage node, a storage node mapping analysis is performed to construct a storage node mapping relationship table; The dataset to be processed is segmented according to the storage node mapping table to generate multiple data fragments. The multiple data shards are simultaneously sent to the primary storage node and the replica storage node for storage, thereby constructing the data storage cluster.
3. The method for secure management of massive vehicle-road-cloud data based on distributed storage as described in claim 1, characterized in that, The method for performing layered encryption processing on the data storage cluster to generate a first encrypted storage shard includes: Extract data type identifiers from data records in the dataset to be processed; Data is parsed based on the data type identifier to determine the data security level label, which includes basic perception data level, collaborative control data level, and privacy data level. An encryption policy library is introduced, and multiple encrypted data blocks are determined by matching the basic perception data level, the collaborative control data level, and the privacy data level through the encryption policy library. The multiple encrypted data blocks are combined and encapsulated to construct the first encrypted storage fragment.
4. The method for secure management of massive vehicle-road-cloud data based on distributed storage as described in claim 3, characterized in that, An encryption policy library is introduced, and multiple encrypted data blocks are determined by matching the encryption policy library according to the basic perception data level, the collaborative control data level, and the privacy data level. The method includes: When the data security level label is the basic perception data level, a symmetric encryption algorithm is used to traverse the encryption strategy library according to the first key length to perform encryption processing and generate a basic level encrypted data block. When the data security level label is the collaborative control data level, a symmetric encryption algorithm is used to traverse the encryption strategy library according to the second key length to perform encryption processing and generate a collaborative-level encrypted data block, wherein the second key length is greater than the first key length. When the data security level label is the privacy data level, the attribute-based encryption algorithm is used to traverse the encryption policy library according to the access control policy tree to perform encryption processing and generate a privacy-level encrypted data block. The access control policy tree contains multiple attribute condition combinations.
5. The method for secure management of massive vehicle-road-cloud data based on distributed storage as described in claim 1, characterized in that, The method includes receiving and parsing data access requests, performing multi-node parallel verification on the first encrypted storage fragment based on the parsing result, and reading the second encrypted storage fragment. Receive data access requests through the access interface of the distributed storage cluster; The data access request is parsed to determine the request header information and request body information. The request header information includes the visitor's digital certificate and identity identifier. The request body information includes the list of data record identifiers to be accessed and the request operation type. Based on the data record identifier list, a batch query request is initiated to the metadata management node to obtain the data storage node mapping relationship and encrypted metadata; The visitor's digital certificate is sent to the blockchain verification network for parallel verification, generating a signature verification result. Based on the signature verification results, perform verification return statistics to determine the number of verification nodes; The first encrypted storage fragment is then subjected to multi-node parallel verification based on the number of verification nodes, and the second encrypted storage fragment is read based on the analysis results.
6. The method for secure management of massive vehicle-road-cloud data based on distributed storage as described in claim 5, characterized in that, The determination and analysis are based on the number of verification nodes, and the method includes: If the number of verification nodes exceeds a preset verification pass threshold, the digital certificate verification is deemed successful, and a first judgment result is generated. Based on the first determination result, the metadata management nodes are traversed to associate data record identifiers and determine the access control policy; The identity identifier and the request operation type are matched and compared with the access control policy. When the access permission matches, the storage node mapping relationship is extracted and matched with the data of the record to be accessed to generate an access node list. The analysis results are generated by connecting to multiple storage nodes based on the access node list.
7. The method for secure management of massive vehicle-road-cloud data based on distributed storage as described in claim 1, characterized in that, The method for verifying the second encrypted storage fragment includes: Read the encryption algorithm identifier field from the metadata area of the second encrypted storage segment to identify the target encryption algorithm type; Read the key index field from the metadata area of the second encrypted storage segment; Based on the target encryption algorithm type and the key index field, the key acquisition request is determined; The key acquisition request is sent to the key management service via a secure communication channel for index lookup to determine the target encryption key. The target encryption key is returned to the verification module through a secure channel for decryption, generating original data fragments. Extract the original integrity verification value based on the second encrypted storage fragment; The original data fragments are hashed to generate real-time integrity verification values; The real-time integrity check value is compared with the original integrity check value to generate the verification result.
8. A vehicle-road-cloud-based massive data security management system based on distributed storage, characterized in that: The system is used to implement the vehicle-road-cloud-based massive data security management method based on distributed storage as described in any one of claims 1-7, the system comprising: Data acquisition unit: Connects to roadside sensing devices and vehicle terminals for real-time data acquisition, obtains multi-source traffic data, identifies the types of traffic data, and generates a dataset to be processed; Data distribution unit: Distributes the dataset to be processed through distributed addressing, determines the storage node mapping relationship, and distributes the dataset to be processed to multiple storage nodes according to the storage node mapping relationship to build a data storage cluster; Layered encryption unit: performs layered encryption processing on the data storage cluster to generate the first encrypted storage fragment; Parallel verification unit: receives data access requests, parses them, performs multi-node parallel verification on the first encrypted storage fragment based on the parsing results, and reads the second encrypted storage fragment; Dynamic adjustment unit: verifies the second encrypted storage shard, and dynamically adjusts the storage node mapping relationship based on the verification result to generate a data distribution optimization management scheme.