Vehicle-mounted social network data storage method based on block chain
Through the blockchain-based on-vehicle social network data storage method, social grouping and multi-dimensional pruning mechanism are adopted to solve the problem of data redundancy and storage pressure in the Internet of Vehicles, and efficient and secure data storage and transmission are achieved, which improves the scalability and reliability of the Internet of Vehicles.
Patent Information
- Application Number
- CN202510447993.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-01
AI Technical Summary
There are problems with data redundancy and storage pressure in the Internet of Vehicles, vehicle resources are limited, data transmission is unstable, security and credibility are difficult to guarantee, and the centralized cloud computing architecture has high latency and single point of failure risks.
Through the blockchain-based on-vehicle social network data storage method, social grouping and multi-dimensional pruning mechanism are adopted to divide vehicle nodes into full nodes and ordinary nodes, stacked automatic encoder is used to optimize packet accuracy, pruning strategy eliminates redundant data, and combines distributed storage and consensus mechanism to ensure data integrity and security.
It effectively reduces the storage pressure of on-board nodes, improves data sharing efficiency and security, enhances the scalability and reliability of the Internet of Vehicles, and reduces communication overhead.
Smart Images

Figure CN120406835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of blockchain technology and vehicle networking data storage, and particularly relates to a method for storing vehicle-mounted social network data based on blockchain. Background Art
[0002] With the continuous growth of vehicle ownership, the road traffic pressure has increased. The traditional traffic management system is difficult to cope with the complex traffic environment, and there is an urgent need to introduce innovative technologies to improve road safety and traffic efficiency. As an important part of intelligent transportation, the Internet of Vehicles (IoV) realizes real-time data perception and sharing through the efficient interconnection between vehicles (V2V) and between vehicles and infrastructure (V2I). Based on the low-latency and high-bandwidth characteristics of 5G communication, the IoV can support advanced applications such as cooperative driving and autonomous driving, and reduce traffic accidents and optimize the overall traffic flow through data dynamic analysis and prediction.
[0003] Although the IoV has shown significant advantages, it still faces many challenges in large-scale applications. First, due to the high-speed movement of vehicles, the network topology changes frequently, resulting in easy interruption of communication connections and affecting the stability and real-time nature of data transmission. Second, the openness of the IoV increases the data security risk, and malicious nodes may spread false information, misleading drivers and causing traffic chaos. Moreover, the IoV relies on the collection and sharing of massive amounts of data, but due to the lack of an effective trust management mechanism, the credibility of the data cannot be guaranteed. Some vehicles may even manipulate the system by creating multiple false identities (such as Sybil attacks), disrupting the traffic order. In addition, the computing and storage resources of vehicles are limited, and many IoV tasks still rely on a centralized cloud computing architecture, which has problems such as high latency, high cost, and single-point failure, and is difficult to meet the requirements of the IoV for efficiency, reliability, and real-time nature.
[0004] To solve the above problems, Vehicular Social Network (VSN) provides a new solution. VSN introduces the concept of social network into the Internet of Vehicles, constructs virtual communities based on vehicle behavior and social needs, enabling vehicles with similar interests or needs to cooperate efficiently. Compared with traditional Internet of Vehicles, VSN adopts a distributed architecture, reduces the dependence on centralized systems, lowers system costs and enhances robustness. At the same time, VSN constructs a trust mechanism through social relationships, filters false information, and improves data credibility. In addition, based on social grouping methods, VSN divides vehicles with similar interests into social groups and combines a data similarity deduplication strategy to significantly reduce redundant storage within the groups. Management nodes store complete data to ensure data security and integrity. In recent years, the introduction of blockchain technology has provided strong protection for the data security of the Internet of Vehicles. The distributed ledger mechanism of blockchain eliminates the risk of single-point failure and ensures the immutability and traceability of data through a consensus mechanism, fundamentally enhancing the security and credibility of data. At the same time, the smart contract function can realize automated data interaction and verification, reduce human intervention, and improve system efficiency.
[0005] However, although vehicular social networks and blockchain show great potential in solving Internet of Vehicles problems, the issue of limited vehicle resources still needs to be considered in practical applications. Therefore, it is particularly important to design a lightweight vehicular social network. While reducing resource occupancy, this architecture needs to ensure data transmission efficiency and security, and adapt to the dynamic environment of the Internet of Vehicles through distributed computing and optimized protocols. This not only helps to promote the popularization and application of the Internet of Vehicles, but also further improves the reliability and scalability of intelligent transportation systems, laying a technical foundation for the development of future smart transportation. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a blockchain-based data storage method for vehicular social networks, which can effectively solve the problems of data redundancy and storage pressure in Vehicular Social Network (VSN), and improve the efficiency and security of data sharing. By performing social grouping on vehicle nodes through multi-dimensional similarity calculation and adopting a grouping election and consensus algorithm, the nodes are divided into ordinary nodes and full nodes. Ordinary nodes store group-related data, and full nodes store complete data and are responsible for deduplication; using a stacked autoencoder to optimize the grouping accuracy and clearing redundant data through a pruning strategy to achieve storage efficiency and data integrity.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for storing data in a vehicle-mounted social network based on blockchain. In the process of blockchain data management in the vehicle-mounted social network, efficient data storage and transmission optimization are realized through a social grouping and multi-dimensional pruning mechanism. The social grouping mechanism dynamically divides vehicle nodes based on the physical location, network topology, social behavior, and data compatibility of the vehicles, effectively reducing the storage and transmission of redundant data. The multi-dimensional pruning mechanism performs data deduplication according to the timeliness, similarity, and access frequency of the data, ensuring the efficiency and accuracy of data storage. The method specifically includes the following steps: vehicle social grouping election stage, consensus stage, social grouping stage based on physical location similarity, structural similarity, modular similarity, and data compatibility, and data deduplication stage based on social grouping. The method combines blockchain technology with the vehicle-mounted social network, and through distributed storage and consensus mechanism, ensures the integrity and immutability of the data, effectively reduces the storage pressure of vehicle nodes, and solves the problem of data redundancy through a deduplication strategy. In addition, a lightweight blockchain consensus mechanism is adopted, reducing storage and communication overhead, while improving the scalability and security of the vehicle-mounted social network, and promoting the data sharing and collaboration efficiency in the vehicle networking environment.
[0008] The vehicle social grouping election stage:
[0009] Step 1-1: The vehicle nodes in each group automatically become initial candidates, that is, each node has the opportunity to be elected as a full node. The preliminary election qualification is evaluated by computing power and activity to determine which nodes have strong candidate qualifications.
[0010] Step 1-2: Each node can vote for itself to ensure the basic weight of self-recommendation.
[0011] Step 1-3: Nodes can canvass votes within the group, inviting other nodes to support themselves, and the nodes can canvass votes based on factors such as computing power and activity.
[0012] Step 1-4: After the canvassing ends, the voting results will be counted by all group nodes according to the number of votes obtained and ranked according to the number of votes. The node with the most votes ranks first.
[0013] Step 1-5: According to the ranking of the number of votes, the nodes ranked in the top 1 / 3 of the votes within the group will be elected as full nodes, responsible for data management and network consensus maintenance within the group. The nodes that fail to enter the top 1 / 3 will be ordinary nodes, continuing to participate in data transmission and collaboration tasks, but not responsible for core data management.
[0014] The vehicle social grouping consensus stage:
[0015] Step 2-1: Transaction initiation: When G kThe vehicle in it obtains enough data to initiate a new transaction request, and transaction T is broadcast to all full nodes in the group. Each transaction T in the blockchain is the basic unit of cross-group or intra-group data interaction.
[0016] Step 2-2: After receiving the transaction request, all full nodes in the group perform the following review operations:
[0017] Verify the identity of the initiating node: Confirm whether the identity of the transaction initiating node is true and valid through Verief.
[0018] Check the transaction content: Verify the authenticity and validity of the transaction content by calculating hash′ and comparing hash′ = hash.
[0019] After the review is passed, the transaction will enter the next stage of the voting session.
[0020] Step 2-3: The transaction requests that pass the review will be broadcast to all full nodes. Each node votes for or against the transaction within a predetermined time, and the votes are recorded through the aggregate signature method Σ. The voting weight of each node is equal, and the voting result directly determines whether the transaction can enter the blockchain.
[0021] Step 2-4: If the number of supporting votes exceeds the set threshold, the transaction initiating node obtains the right to record, and can package the transaction into a block and record it on the blockchain. If the number of supporting votes does not reach the threshold, the transaction will be rejected.
[0022] Step 2-5: The node S that obtains the right to record χ Packages the transaction together with other pending transactions and the aggregate signature Σ into a block B k , and submits it to the blockchain network and broadcasts the record within the entire chain range.
[0023] The stages of the social grouping:
[0024] Step 3-1: Physical location similarity is used to measure the geographical proximity between vehicles. Based on the geographical coordinates of vehicle v i and v j being (x i , y j ) and (x i , y j ), calculate the physical distance d ij between them. If the distance between two nodes is less than the threshold τ, they are adjacent in physical space, and the physical location similarity is 1, otherwise it is 0. Obtain the physical location similarity matrix Sim dist
[0025] Step 3-2: Structural similarity measures the connection characteristics of vehicles in the network topology, mainly considering the degree of neighbor overlap and the direct connection state. For vehicle vi and v j , calculate their neighbor overlap ratio through their respective neighbor sets N(v i ) and N(v j ) to obtain the structure matrix Sim str .
[0026] Step 3-3: Modular similarity measures the community structure stability in the network and is mainly calculated based on the actual connection status between vehicles and the degrees of vehicle nodes. For vehicles v i and v j , if they are directly connected, the modular similarity Sim mod considers their adjacency relationship A ij and their respective degrees κ i , κ j .
[0027] Step 3-4: Data similarity is mainly used to measure the closeness of vehicle data features. For nodes v i and v j , collect their feature vectors X i and X j , project them into a low-dimensional space through PCA to obtain the feature vectors after dimensionality reduction Measure the data similarity by calculating the Euclidean distance dist pca (i, j) between the feature vectors, and calculate the data similarity matrix Sim data .
[0028] Step 3-5: To obtain a more comprehensive similarity measure, first calculate the physical location similarity matrix, structure similarity matrix, modular similarity matrix, and data similarity matrix. Then, fuse these four matrices in a weighted manner to generate a comprehensive similarity matrix Sim total . The weights w1, w2, w3, w4 of each similarity matrix can be adjusted according to the specific application scenario to optimize the grouping accuracy.
[0029] Step 3-6: To optimize the similarity matrix and reduce the computational complexity, we use Stacked Autoencoders to perform dimensionality reduction on the comprehensive similarity matrix Sim total .
[0030] Encoding stage: Through multiple non-linear transformations, compress Sim total into a low-dimensional embedded representation Enc N .
[0031] Decoding stage: Regenerate from Enc N And minimize the reconstruction error to ensure that as much information as possible is retained.
[0032] Final output: The matrix after dimensionality reduction is used for subsequent clustering, reducing the consumption of computing resources and improving the efficiency of social grouping.
[0033] Steps 3-7: Use a density-based clustering algorithm to perform social grouping on vehicle nodes. Calculate the local density of each node, and select the nodes with high density as seed nodes. Then, include the neighborhood nodes (with high density and high similarity) of the seed nodes into the same community and expand recursively. Isolated nodes with low density are regarded as noise and do not participate in social grouping. Multiple highly similar social groups are formed to improve the efficiency and accuracy of vehicle data sharing.
[0034] The social grouping deduplication stage:
[0035] Step 4-1: After the grouping process is completed, the system selects full nodes M i and ordinary nodes V i through a full-node election mechanism based on computing power C i and network activity A i . Full nodes are responsible for storing the complete blockchain data D full , performing data pruning, managing data D cross synchronization, and verification, while ordinary nodes only store their own relevant data D topic , obtain key information from full nodes, reduce the storage burden, and improve the query efficiency. After the election, full nodes perform the data deduplication task and ensure that the pruned data is traceable and verifiable.
[0036] Step 4-2: Full nodes make pruning decisions on the stored data according to the rules of data timeliness Prune(data), data similarity Prune sim (data p , data q ), and access frequency Prune freq (data). If the timestamp t of the data data is earlier than the set threshold t now of the current time t τ , it is marked as prunable data to reduce the storage of expired data. If the similarity between the data data p and other data data q in the group is higher than the threshold ι on the feature vectors X p and X q , only the representative data is retained, and the remaining data is marked as prunable data. If the access frequency f of the data data within the set time is lower than the threshold f τ , it is marked as prunable data to release storage space. Any data that meets any pruning condition is marked as prunable data and enters the next stage of processing.
[0037] Step 4-3: For the data marked as prunable, the full node performs data storage optimization. Data deletion enables the full node to remove expired, redundant, or low-frequency data from local storage, releasing storage resources. Data compression is applied to the data that still has some value, and hash indexing and Merkle trees are used for compressed storage. Ensure that critical data is still stored in the full node, while ordinary nodes only retain the data fragments they need, reducing storage pressure and improving the storage efficiency of the system.
[0038] Step 4-4: To ensure the integrity of the data after pruning, the full node synchronizes the data through the blockchain consensus mechanism. After the pruning operation is completed, the full node broadcasts update information to the network to notify all participating nodes. The full node verifies the signature Sign χ to ensure the authenticity of the pruned data. Ordinary nodes re-pull the latest data through trusted full nodes to ensure that their own data is consistent with the whole network. After consensus confirmation, the pruning status is officially recorded on the blockchain to ensure the immutability of the data, further enhancing the security and storage optimization capabilities of the system.
[0039] The privacy protection method described includes three entities: vehicle nodes, roadside units (RSUs), and cloud service providers (CSPs);
[0040] The vehicle nodes, which are the main source of data in the system, are responsible for real-time collection of driving data and environmental information, and data exchange with other nearby vehicle nodes. In addition, vehicle nodes undertake basic communication tasks in the network and interact with roadside units to provide more extensive traffic information.
[0041] The roadside units, as edge computing nodes, are deployed in road infrastructure and act as a bridge between vehicle nodes and cloud service providers. The RSU receives data from multiple vehicle nodes and performs data preprocessing, consensus verification, and storage optimization. The RSU can also coordinate the social grouping mechanism in the vehicular social network to ensure efficient sharing of data within the same social group.
[0042] The cloud service provider, which is the core manager of the entire system, is responsible for storing global data and performing data analysis. The CSP receives vehicle data collected from the RSU and calculates four similarity matrices based on the vehicle data according to the community grouping mechanism, and then groups the vehicle nodes. The CSP also serves as the centralized processing and storage of global data.
[0043] The advantages of the present invention are as follows: 1. Grouping method based on social forest: Combining physical location, network structure, modularity, and data compatibility similarity, a grouping algorithm is designed, and the stacking autoencoder is used to optimize the grouping result, thereby improving the intra-group cooperation efficiency and reducing the communication overhead.
[0044] 2. Efficient data pruning strategy: A multi-dimensional pruning method based on timeliness, similarity, and access frequency is proposed, significantly reducing the storage of ordinary nodes.
[0045] 3. Guarantee of data integrity and security: By storing complete data in full nodes and a verification mechanism based on authentication and intra-group consensus, the reliability of the pruned data in terms of authenticity and security is ensured.
[0046] 4. Lightweight blockchain consensus mechanism: A distributed consensus mechanism based on social grouping is designed. Through the cooperation between full nodes and ordinary nodes, the scalability and reliability of the system are improved while reducing communication volume. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The content expressed in each drawing of the specification of the present invention and the marks in the drawings are briefly described below:
[0048] Figure 1 It is a schematic diagram of the overall architecture of the in-vehicle social network system of the present invention;
[0049] Figure 2 It is a flowchart of social grouping of the in-vehicle social network of the present invention.
[0050] Figure 3 It is a schematic diagram of data deduplication of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0051] The specific embodiments of the present invention are further described in detail below with reference to the accompanying drawings through the description of the optimal embodiments.
[0052] The present invention proposes a method for storing vehicle social network data based on blockchain. In the process of blockchain data management in vehicle social networks, efficient data storage and transmission optimization are achieved through social grouping and multi-dimensional pruning mechanisms. The social grouping mechanism dynamically divides vehicle nodes based on the physical location, network topology, social behavior, and data compatibility of vehicles, effectively reducing the storage and transmission of redundant data. The multi-dimensional pruning mechanism removes duplicate data according to the timeliness, similarity, and access frequency of data, ensuring the efficiency and accuracy of data storage. The method specifically includes the following steps: vehicle social grouping election stage, consensus stage, social grouping stage based on physical location similarity, structural similarity, modular similarity, and data compatibility, and data deduplication stage based on social grouping. The method combines blockchain technology with vehicle social networks, and through distributed storage and consensus mechanisms, ensures the integrity and immutability of data, effectively reducing the storage pressure of vehicle nodes, and solves the problem of data redundancy through a deduplication strategy. In addition, a lightweight blockchain consensus mechanism is adopted, reducing storage and communication overhead, while improving the scalability and security of vehicle social networks, and promoting data sharing and collaboration efficiency in the vehicle networking environment.
[0053] As Figure 1 shown, the system involved in the method for storing data in the vehicle social network includes three entities: vehicle nodes, roadside units (RSUs), and cloud service providers (CSPs);
[0054] The vehicle nodes, which are the main source of data in the system, are responsible for real-time collection of driving data and environmental information, and for data exchange with other nearby vehicle nodes. In addition, vehicle nodes undertake basic communication tasks in the network and interact with roadside units to provide more extensive traffic information.
[0055] The roadside units, as edge computing nodes, are deployed in road infrastructure and act as a bridge between vehicle nodes and cloud service providers. RSUs receive data from multiple vehicle nodes and perform data preprocessing, consensus verification, and storage optimization. RSUs can also coordinate the social grouping mechanism in the vehicle social network to ensure efficient sharing of data within the same social group formed by aggregating vehicle nodes through social similarity. A social group refers to a group formed by aggregating vehicle nodes through social similarity.
[0056] The cloud service provider, which is the core manager of the entire system, is responsible for storing global data and performing data analysis. The CSP receives vehicle data collected from RSUs and calculates four similarity matrices based on the vehicle data according to the community grouping mechanism, and then groups the vehicle nodes. The CSP also serves as the centralized processing and storage of global data.
[0057] For vehicle nodes in a vehicular social network, social grouping is performed through multi-dimensional similarity calculation. The social grouping includes a vehicle social grouping election stage based on blockchain, a vehicle social grouping consensus stage, and a social grouping stage based on physical location similarity, structural similarity, modular similarity, and data compatibility. After grouping, data storage and deduplication are performed according to the divided ordinary nodes and full nodes. Among them, ordinary nodes store grouping-related data, and full nodes store complete data and are responsible for deduplication. The stacked autoencoder is used to optimize the grouping accuracy, and redundant data is cleared through pruning strategies to achieve storage efficiency and data integrity.
[0058] The following is a detailed description of each stage in social grouping and data storage deduplication as follows:
[0059] 1. The vehicle social grouping election stage based on blockchain is specifically as follows
[0060] Step 1-1: When a full node commits an illegal act or is inactive for a long time and is removed from the full node identity, the smart contract in the blockchain automatically runs a voting mechanism. Each non-full node vehicle node in the social group formed by social similarity aggregation automatically becomes an initial candidate, that is, each node has the opportunity to be elected as a full node. The node attribute set is defined as P i ={C i , A i}, where C i represents the computing power of the node, and A i represents the node activity. The computing power measures the ability of the vehicle node to process data and execute tasks, and the activity reflects the active degree and participation frequency of the node in the network. The preliminary election eligibility is evaluated through these two attributes to determine which nodes have strong candidacy.
[0061] Step 1-2: The blockchain smart contract stipulates that each candidate node can vote for itself to ensure the basic weight of self-recommendation. Self-voting provides an opportunity for the node to self-recommend and gives it basic preliminary influence.
[0062] Step 1-3: Since more benefits can be obtained after being elected as a full node, candidate nodes will send group messages containing factors such as their own computing power and activity to canvass votes within the group to invite other nodes to support themselves. At this time, nodes with stronger computing power and higher activity are more likely to obtain the support of other nodes.
[0063] Step 1-4: After the canvassing ends, the voting results will be counted by the remaining full nodes (RSU must be a full node) for the number of votes of all group nodes, and ranked according to the number of votes. The node with the most votes ranks at the top.
[0064] Step 1-5: According to the ranking of the number of votes, the nodes in the top 1 / 3 of the number of votes within the group will be elected as full nodes, responsible for data management within the group and network consensus maintenance. These nodes have higher computing power, stronger activity, and stronger sense of responsibility, and can ensure the integrity and security of data. The nodes that fail to enter the top 1 / 3 will be ordinary nodes and continue to participate in data transmission and collaboration tasks, but are not responsible for core data management.
[0065] 2. The vehicle social grouping consensus stage:
[0066] Step 2-1: Transaction initiation: Set the blockchain as the data distributed interaction platform for vehicle nodes. Vehicle nodes submit data transaction requests in the platform as blockchain nodes. Only after the transaction is verified by the consensus mechanism will it be stored in the blockchain, using the characteristics of the blockchain to ensure the immutability and verifiability of data. When the vehicle nodes in the kth group G k acquire data that can generate a block, this vehicle node can initiate a new transaction request, and the transaction T is broadcast to the full nodes within the group. Each transaction T in the blockchain is the basic unit of cross-group or intra-group data interaction, and its format is defined as: T = {tx id , t, data, hash, S χ , R γ , sign χ} where tx id is the unique transaction identifier, t is the timestamp of the transaction, indicating the time when the transaction is generated. data represents the content of the transaction data. hash represents the transaction hash value, used to verify the integrity of the data. S χ represents the sending node full node M i or vehicle node V j , R γ represents the receiving node full node M i or vehicle node V j . sign χ represents the transaction signature, used to verify the identity of the initiating node.
[0067] Step 2-2: After receiving the transaction request, the full nodes within the group perform the following review operations:
[0068] Verify the identity of the initiating node: Confirm whether the identity of the transaction initiating node is true and valid through Verief(sign χ , hash, key p ). Among them, sign χ represents the transaction signature, used to verify the identity of the initiating node, hash represents the transaction hash value, used to verify the integrity of the data, and key p represents the public key of the sending node.
[0069] Check the transaction content: By calculating hash′ = Hash(t, data, S χ , R γ ), compare the hash hash in the transaction message with the calculated hash hash′ above. If hash′ = hash, it is verified that the transaction content is true and valid. Where t is the timestamp of the transaction, data represents the transaction data content, and S χ represents the full node M i or vehicle node V j of the sending node of the transaction, and R γ represents the full node M i or vehicle node V j of the receiving node of the transaction.
[0070] After the review is passed, the transaction will enter the voting stage of the next phase.
[0071] Step 2-3: The transaction requests that pass the review will be broadcast to all full nodes, and the blockchain smart contract will automatically run the voting mechanism. Each node votes within a predetermined time based on its verification result of the transaction message, supporting or opposing the transaction. The votes are recorded through the aggregate signature method Σ = Aggregate(sign1, sign2,…, sign χ ). Each node has an equal voting weight, and the voting result directly determines whether the transaction can enter the blockchain. Where sign1~sign n represents the signature of the supporting nodes.
[0072] Step 2-4: If the number of supporting votes exceeds the set threshold, the transaction initiating node obtains the right to record and can package the transaction into a block and record it on the blockchain. If the number of supporting votes does not reach the threshold, the transaction will be rejected.
[0073] Step 2-5: The node S χ that obtains the right to record will package the transaction together with other pending transactions T1~T n that have not been confirmed by the blockchain and need to wait for the verification and voting of the full nodes, as well as the aggregate signature Σ, into a block B k = {T1, T2,…, T n , Σ}, submit it to the blockchain network, and broadcast and record it within the entire chain.
[0074] As Figure 2 shown, the social grouping is divided into three levels, namely the data collection layer, the data processing layer, and the social grouping layer. The vehicle nodes collect data and upload it to the RSU for data preprocessing, and finally submit it to the CSP for social grouping based on the aggregated similarity matrix.
[0075] 3. The stages of the social grouping described above:
[0076] Step 3-1: Physical location similarity is calculated based on the geographical location information of vehicles (such as GPS coordinates). It measures the spatial distance between vehicles. Assuming that vehicles closer in distance have a greater potential to share data. Let vehicle v i and v j have geographical coordinates (x i , y j ) and (x i , y j ). The physical location similarity between the two is calculated by the following formula:
[0077]
[0078] where is the geographical distance between vehicle v i and v j , calculated using the Euclidean distance of physical coordinates, and τ is a preset distance threshold. If the distance between two nodes is less than the threshold τ, they are adjacent in physical space and the physical location similarity is 1; otherwise, it is 0. The design of this function is inspired by the "proximity effect" in physical networks, that is, vehicles closer in location are more likely to share data. Therefore, this similarity helps to identify groups of vehicles with close locations, especially in urban traffic scenarios.
[0079] Step 3-2: Structural similarity is measured based on the connection characteristics of vehicle nodes in the network topology, reflecting the direct connection relationship between vehicle nodes. The structural similarity between node v i and v j is defined as the degree of neighbor overlap, and the formula is as follows:
[0080]
[0081] where N(v i ) and N(v j ) respectively represent the sets of neighbor nodes that have established direct data transmission links with nodes v i and v j through wireless communication or other network protocols, and |·| represents the size of the set. This similarity measures the connection density between nodes. The more overlapping neighbor nodes there are, the greater the structural similarity. This similarity is widely used in both social networks and vehicle-to-vehicle networks and is suitable for analyzing the possibility of direct data transmission between vehicles.
[0082] Step 3-3: Modularity similarity measures the degree of connectivity between nodes within a community, where vehicle nodes are clustered together due to similar attributes or behaviors (e.g., similar physical locations, data similarity, etc.). This is primarily used to identify similarities within a group. We define modularity similarity by combining an improved modularity function with the community modularity structure. The improved modularity function formula is as follows:
[0083]
[0084] Among them, a ij For vehicle v i and v j Whether a direct data transmission link is established between them through wireless communication or other network protocols, if so, it is 1, otherwise it is 0. i and κ j is the degree of nodes i and j, that is, the number of other nodes that the node has direct links to. The degree indicates the degree of connectivity of a node. The higher the degree, the more connections the node has and the greater its influence in the network. is node v i and node v j The geometric mean of the degree is used to normalize the similarity. If there is a direct connection between two nodes and their degrees are relatively small, their similarity will be higher. The modular similarity matrix formula is as follows:
[0085]
[0086] Where i, j represents the node v i and node v j , where N represents the number of nodes. Modular similarity is widely used for community detection in complex networks. By highlighting the dense connections within communities, it can effectively divide networks into groups with similar functions or interests. This approach is particularly suitable for identifying vehicle groupings in vehicular social networks (VSNs). By quantifying the connection density between nodes, it can more accurately optimize the division of vehicles into communities, thereby improving data sharing efficiency and collaborative performance in the Internet of Vehicles.
[0087] Steps 3-5: PCA (Principal Component Analysis) is a commonly used data dimensionality reduction method. Its core idea is to map high-dimensional data to a low-dimensional space through linear transformation while preserving the main features of the original data as much as possible. The purpose of PCA is to simplify the data representation by extracting the main eigenvectors of the data and removing redundant information. The mathematical process is as follows:
[0088] First, center the given data matrix X: in is the mean of the data. Then calculate the data covariance matrix C: where n is the number of samples. Perform eigenvalue decomposition on the covariance matrix C to obtain the projection matrix U: CU = ΛU, where Λ is a diagonal matrix containing the eigenvalues of the covariance matrix, and U is the corresponding eigenvector matrix. Use the projection matrix U formed by the eigenvectors corresponding to the d largest eigenvalues to d perform data dimensionality reduction and map it to a low-dimensional space:
[0089]
[0090] After performing dimensionality reduction on the data of vehicle nodes through PCA, for node v i and v j , respectively collect their eigenvectors X i and X j . The eigenvector refers to the representation of each vehicle node in the reduced-dimensional space. Project it to the low-dimensional space through PCA to obtain the reduced-dimensional eigenvector:
[0091]
[0092] Node v i and v j The data similarity between them is measured by the Euclidean distance between the reduced-dimensional eigenvectors, and is calculated through , where ||·||2 represents the two-norm of the vector. According to the calculated Euclidean distance dist pca (i,j), combined with the preset similarity threshold to judge the data compatibility between nodes. The data compatibility similarity matrix Sim data is defined as:
[0093]
[0094] The purpose of this matrix design is to ensure a high degree of consistency in data types and formats for vehicles sharing data. Through this similarity matrix, the efficiency and accuracy of data sharing can be effectively improved, while reducing communication and storage overhead. In addition, combined with PCA dimensionality reduction processing, the computational complexity can also be reduced, which is especially suitable for the dynamic environment of large-scale vehicle networks.
[0095] Step 3 - 6: After obtaining the above four similarity matrices, in order to comprehensively measure these similarity metrics, we use the weighted sum method to fuse them to generate a comprehensive similarity matrix. The specific formula is as follows:
[0096] Sim total = w1·Sim str + w2·Sim mod + w3·Simdata + w4·Sim dist
[0097] Among them, w1, w2, w3, and w4 are the weights of four similarity measures respectively. These weights can be flexibly adjusted according to the requirements of specific application scenarios to achieve the best community division effect. Among them, Sim str represents the structural similarity matrix, Sim mod represents the modular similarity matrix, Sim data represents the data compatibility similarity matrix, Sim dist represents the physical location similarity matrix.). In order to further optimize the feature expression of the comprehensive similarity matrix and reduce the computational complexity, Stacked Autoencoders are introduced to reduce the dimension of the high-dimensional similarity matrix. As Figure 2 shown, Stacked Autoencoders is an unsupervised deep learning model that can compress high-dimensional data into a low-dimensional embedding space by learning the key features of the high-dimensional data, providing a more efficient and accurate input for subsequent community division.
[0098] Stacked Autoencoders mainly consists of an input layer, a hidden layer (encoding layer), a decoding layer, and an output layer. Its core goal is to minimize the dimension of the data to the greatest extent while retaining the main features of the original data. First, the similarity matrix Sim total is used as the input and is compressed into a low-dimensional space through the non-linear mapping of the encoder layer. It is expressed as:
[0099] Enc N = σ(ω1·Sim total + b1)
[0100] Among them, Sim total is the input composed of the physical location, structure, modular, and data compatibility similarity matrices. ω1 and b1 are the weight and bias respectively, and σ(·) is the activation function used to introduce non-linear features. The encoder compresses these matrices into a low-dimensional embedding representation, reducing the dimension of the matrix while retaining the key similarity information.
[0101] Then, the decoder attempts to reconstruct the original matrix from the low-dimensional embedding representation Enc N . Its reconstruction formula is:
[0102]
[0103] Among them, is the reconstructed similarity matrix, ω2 and b2 are the weight and bias respectively, and σ(·) is the activation function used to introduce non-linear features. The encoder and decoder continuously adjust the weights and biases to make the reconstruction error Minimize to ensure that the compressed low - dimensional representation retains the information of the original matrix as completely as possible.
[0104] To further optimize the distribution characteristics of the encoder output and improve the generalization ability of the model, we introduce a sparsity constraint. The difference between the activation distribution of hidden layer units and the target sparse distribution is measured by the Kullback - Leibler divergence (KL divergence), which is defined as follows:
[0105]
[0106] where ρ is the preset sparsity target value, is the average activation value of the m - th hidden layer unit. The final objective function combines the reconstruction error and the sparsity constraint, which is defined as follows:
[0107]
[0108] where represents the reconstruction error, β is the weight coefficient of the sparse regularization term, used to balance the reconstruction error and the sparsity requirement, 1 - M represents the number of hidden units, ρ is the preset sparsity target value, is the average activation value of the m - th hidden layer unit. Through the low - dimensional embedding representation generated by stacking autoencoders, the similarity relationship between vehicle nodes is further optimized and quantified. Subsequently, the density - based clustering algorithm can utilize these low - dimensional features to more accurately identify and divide communities, forming vehicle groups with strong functionality and high collaboration.
[0109] Steps 3 - 7: After obtaining the low - dimensional embedding representation, we use the Kullback - Leibler (KL) divergence in information theory to calculate the "information distance" between each vehicle node. The KL divergence can measure the difference between two probability distributions, and the formula is as follows:
[0110]
[0111] where v i (n) and v j (n) are the embedding representations of vehicle v i and v j in the low - dimensional space. By calculating the KL divergence, we can further quantify the similarity between vehicle nodes. The smaller the information distance, the higher the similarity between nodes.
[0112] To divide vehicles with similarity into the same community, we adopt a density - based clustering method. The density - clustering method can automatically identify different communities according to the distribution density of nodes without the need to preset the number of communities. The specific steps are as follows:
[0113] Define the neighborhood range of each node using the distance threshold ε, and calculate the neighborhood density π i of node v i :
[0114]
[0115] where Ⅱ is the indicator function, v i and v j represent vehicle nodes, N represents the total number of nodes, d(v i , v j ) represents the Euclidean distance between vehicle nodes, and ε represents the distance threshold. When d(v i , v j ) ≤ ε, the value is 1, otherwise it is 0. The local density of each node can be calculated through this formula. If the density π i of a certain node exceeds the set density threshold π min , then it is marked as a core point. These core points are the seed nodes for community division. For each core point, all nodes within its neighborhood are grouped into the same community. If there are other core points in its neighborhood, then recursively expand the neighborhoods of these core points to gradually form a complete community structure. The specific formula is as follows:
[0116]
[0117] where G k represents the k-th community, N ε (v i ) is the set of neighborhood nodes of node v i , π i represents the neighborhood density, and π min represents the density threshold. Nodes that do not reach the core point density threshold but belong to the neighborhood of a certain core point are marked as boundary points; nodes that do not belong to the neighborhood of any core point are regarded as noise points. Through the above steps, the density clustering algorithm can automatically identify high-density regions and divide them into communities without the need to preset the number of communities.
[0118] As Figure 3 shown, in the social grouping and duplicate removal stage, we divide the nodes into two types. Among them, full nodes store the complete data blocks of the entire blockchain, while intra-group nodes only need to store the intra-group blocks and the content they are interested in.
[0119] The social grouping and duplicate removal stage described above:
[0120] Step 4-1: After the grouping process is completed, the system selects full nodes M i and ordinary nodes V i through the full node election mechanism based on the computing power C i and network activity A iThe full node needs to store the complete history of all blockchain transactions of group G, including detailed information of each transaction. The storage set is defined as:
[0121]
[0122] Where T x represents the xth transaction. n is the total number of transactions. These data are used to verify transactions in the blockchain and provide data integrity verification for other nodes. Full nodes also need to manage different social topic groups G k The full node not only stores the data of a specific topic group, but also is responsible for querying and verifying cross-group data. The storage set is defined as:
[0123]
[0124] in Represents the set of all data in the pth group. Unlike full nodes, ordinary nodes in a group do not need to store all data, but selectively store specific topic group data based on interest points. This selective storage can significantly reduce the storage pressure of ordinary nodes. Ordinary nodes only need to store the data related to their social group G. k Related topic data. For example, if vehicle A is only interested in traffic information for a specific route, then the vehicle only needs to store the topic group data related to that route, without having to store other irrelevant data in the entire blockchain network. Its storage set is defined as:
[0125]
[0126] in Group G where ordinary nodes are located k At the same time, due to the high dynamics of the vehicle network, ordinary nodes tend to store the latest and more real-time data. This type of data is more helpful for vehicle path planning and decision-making, while historical data is relatively less valuable. The real-time data storage collection is defined as:
[0127] D real (V j )={T|(t now -t) <t τ}
[0128] where t now is the current time, t is the timestamp of transaction T, t τ It is the real-time threshold.
[0129] Step 4-2: In the design of the in-group node data pruning method, we mainly perform pruning based on two core dimensions: data timeliness and data relevance. The following is a detailed description of the pruning strategy:
[0130] Pruning based on timeliness: Vehicle nodes tend to be more concerned about the latest data. Especially in vehicular networks, real-time performance is crucial. For example, vehicles' requirements for data such as road conditions, weather, and traffic accidents are often limited to a short period. Therefore, outdated data will be preferentially pruned. The timeliness pruning formula is as follows:
[0131]
[0132] where t now is the current time, t is the timestamp when the data data was generated, and t τ is the set time threshold. If the data exceeds the time threshold, it is marked as 1, indicating that the data can be pruned.
[0133] Pruning based on relevance: Through the grouping mechanism, vehicle nodes are divided into multiple groups with similarities. Within the group, due to the high similarity of vehicle data, there is redundancy. By calculating the similarity of the data within the group, redundant data can be deleted and representative data can be retained. Use cosine similarity to calculate the similarity between data data p and data q :
[0134]
[0135] where X p and X q are the feature vectors of data data p and data q , ||X p || and ||X q || are the two-norms of the feature vectors, and Sim(data p , data q ) is the cosine similarity, ranging from [0,1]. Let the similarity threshold be ι. If the similarity between the data is higher than the threshold, it is considered that there is redundancy, and only one representative data is retained. The formula is as follows:
[0136]
[0137] If the newly uploaded data is too similar to the previously stored data, it is marked as 1, indicating that the data can be pruned.
[0138] Pruning Based on Access Frequency: The access frequency of data is another important factor in determining whether it needs to be retained. If certain data has not been accessed or has very few access records within a certain period of time, pruning can be considered, and data with high access frequency should be retained preferentially.
[0139] Access Frequency Pruning Rule: Record the access times and the most recent access time for each piece of data, and set a minimum access frequency threshold F. min If the access frequency of a certain piece of data is lower than this threshold, it is marked as prunable.
[0140] The formula is as follows:
[0141]
[0142] where f represents the access frequency of data data, and f τ is the set minimum access frequency threshold. If the data access frequency is lower than this threshold, it is marked as 1, indicating that this data can be pruned.
[0143] Step 4-3: For the data marked as prunable, the full node performs data storage optimization. Data deletion enables the full node to remove expired, redundant, or low-frequency data from local storage, releasing storage resources. Data compression is applied to data that still has some value, using hash indexing and Merkle trees for compressed storage. Ensure that critical data is still stored on the full node, while ordinary nodes only retain the data fragments they need, reducing storage pressure and improving the storage efficiency of the system.
[0144] Step 4-4: To ensure that the data after pruning is still complete, consistent, and traceable, the full node synchronizes data through the blockchain consensus mechanism. To ensure the immutability and security of transaction data, the system introduces a two-layer mechanism based on authentication and intra-group consensus voting. The identity of each node is bound through the encryption key mechanism in the blockchain (such as based on public-private key pairs). When a node V j submits a data pruning or data update request, it needs to provide its encrypted signature to verify the legitimacy of its identity:
[0145] Sign x =Hash(data||key s ),
[0146] where data is the transaction data and key s is the private key of the submitting node. By verifying the signature Sign χ it is ensured that the data request comes from a legitimate node and prevents malicious nodes from forging identities to tamper with data. After verifying the legitimacy of the node's identity, the data pruning or update operation of the ordinary node V j needs to be approved by all full nodes M within the group iThe consensus voting mechanism. Each full node approves or vetoes a request according to its voting weight, and an operation can be performed only after the votes of full nodes exceed the threshold. Ordinary nodes re-pull the latest data through trusted full nodes to ensure that their own data is consistent with the whole network.
[0147] Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention, they are all within the protection scope of the present invention.
Claims
1. A blockchain-based method for storing vehicle social network data, characterized in that: In an in-vehicle social network, vehicle nodes are socially grouped through multi-dimensional similarity calculation, and a grouping election and consensus algorithm is adopted to divide vehicle nodes into ordinary nodes and full nodes. Ordinary nodes are used to store grouping-related data, and full nodes are used to store complete data and deduplicate the data.
2. The method for storing vehicle social network data based on blockchain according to claim 1, characterized in that: During the social grouping process, a stacked autoencoder is used to optimize the grouping accuracy.
3. The method for storing vehicle social network data based on blockchain according to claim 2, wherein: A pruning strategy is adopted for the data stored in full nodes to perform redundant data cleaning operations.
4. A method for storing vehicle social network data based on blockchain according to any one of claims 1-3, characterized in that: During the social grouping process, full nodes store complete chain data and act as consensus management nodes to ensure data integrity and security. Full nodes perform pruning operations based on timestamps, access frequencies, and similarity calculations; ordinary nodes participate in consensus and store deduplicated data to reduce storage overhead.
5. A method for storing vehicle social network data based on blockchain according to any one of claims 1-3, characterized in that: The storage method includes a vehicle social grouping election stage based on blockchain, a vehicle social grouping consensus stage, a social grouping stage based on physical location similarity, structural similarity, modular similarity, and data compatibility, and a data deduplication stage based on social grouping.
6. A method for storing data in an in-vehicle social network based on blockchain according to claim 5, characterized in that: The vehicle social grouping election stage includes: Vehicle nodes in each group automatically become initial candidates, and the preliminary election qualifications are evaluated by computing power and activity to determine the strength of the node's candidacy; Each node votes for itself or other nodes; nodes canvass for votes within the group, inviting other nodes to support themselves, and the nodes canvass for votes based on factors such as computing power and activity; After the canvassing ends, the voting results will be statistically counted by all group nodes according to the number of votes obtained, and ranked according to the number of votes obtained. The node with the most votes ranks first. According to the ranking of the number of votes, the nodes ranked in the top 1 / 3 of the votes within the group will be elected as full nodes, responsible for data management and network consensus maintenance within the group; the nodes that fail to enter the top 1 / 3 will be used as ordinary nodes and continue to participate in data transmission and collaboration tasks.
7. A method for storing vehicle social network data based on blockchain according to claim 5, characterized in that: The vehicle social grouping consensus stage includes: Step 2-1: Transaction Initiation: When the vehicle in G k obtains sufficient data, a new transaction request is initiated, and transaction T is broadcast to all nodes in the group; each transaction T in the blockchain is the basic unit of cross-group or intra-group data interaction; Step 2-2: After receiving a transaction request, the full nodes within the group perform the following review operations: Verify the identity of the initiating node: Confirm whether the identity of the transaction initiating node is true and valid through Verief; Check the transaction content: Verify the authenticity and validity of the transaction content by calculating hash′; After the review is passed, the transaction will enter the next stage of the voting session; Step 2-3: The approved transaction request will be broadcast to all full nodes. Each node votes within a predetermined time to support or oppose the transaction, and the votes are recorded by means of aggregated signatures; the voting weight of each node is equal, and the voting result directly determines whether the transaction can enter the blockchain; Step 2-4: If the number of supporting votes exceeds the set threshold, the transaction initiating node will obtain the accounting right, package the transaction into a block and record it on the blockchain; if the number of supporting votes does not reach the threshold, the transaction will be rejected; Step 2-5: The node S that obtains the bookkeeping right χ Pack the transaction together with other pending transactions and the aggregated signature into a block B k , and submit it to the blockchain network and broadcast and record it within the entire chain 8. The method for storing data of a vehicle-mounted social network based on a blockchain according to claim 5, characterized in that: The social grouping stage includes: Step 3-1: Physical location similarity is used to measure the geographical proximity between vehicles. Based on the vehicle v i and v j with geographical coordinates (x i , y j ) and (x i , y j ), calculate the physical distance d ij . If the distance between two nodes is less than the threshold τ, they are adjacent in physical space and the physical location similarity is 1; otherwise, it is 0. Obtain the physical location similarity matrix Sim dist Step 3-2: The structural similarity measures the connection characteristics of vehicles in the network topology. For vehicles vi i and vj, calculate their neighbor overlap ratio through their respective neighbor sets N(vi i ) and N(vj j ) to obtain the structure matrix Sim str ; Step 3-3: Modular similarity measures the stability of the community structure in the network and is calculated based on the actual connection status between vehicles and the degrees of vehicle nodes; for vehicle v i and v j , if they are directly connected, the modular similarity Sim mod considers their adjacency relationship A ij and their respective degrees κ i , κ j ; Step 3-4: Data similarity is used to measure the proximity of vehicle data features; for nodes v i and v j , their feature vectors X i and X j are collected respectively, projected into a low-dimensional space through PCA, and the dimensionality-reduced feature vectors are obtained. The Euclidean distance dist pca (i,j) between the feature vectors is calculated to measure the data similarity, and the data similarity matrix Sim data is calculated; Step 3-5: Calculate the physical location similarity matrix, the structural similarity matrix, the modular similarity matrix, and the data similarity matrix; then, fuse these four matrices in a weighted manner to generate a comprehensive similarity matrix Sim total ; The weights w1, w2, w3, and w4 of each similarity matrix are adjusted according to the specific application scenario; Step 3-6: Use a stacked autoencoder to perform dimensionality reduction on the comprehensive similarity matrix Sim total ; Encoding stage: Through multiple non-linear transformations, Sim total is compressed into a low-dimensional embedded representation Enc N ; Decoding stage: From Enc N Regenerate and minimize the reconstruction error to ensure that as much information as possible is retained; Final output: Output the matrix after dimensionality reduction for subsequent clustering; Step 3-7: Use the density-based clustering algorithm to perform social grouping on vehicle nodes; calculate the local density of each node, and select the nodes with high density as seed nodes; then incorporate the neighborhood nodes of the seed nodes into the same community and expand recursively; the isolated nodes with low density are regarded as noise and do not participate in social grouping; multiple social groupings with high similarity are formed to improve the efficiency and accuracy of vehicle data sharing.
9. A method for storing data in a vehicle-mounted social network based on blockchain according to claim 5, wherein: The social grouping deduplication stage includes: After the grouping process is completed, the system selects full nodes M i and ordinary nodes V i through a full node election mechanism based on computing power C i ; Full nodes are responsible for storing the complete blockchain data D i , performing data pruning, managing data D full synchronization and verification, while ordinary nodes only store their own relevant data D cross topic and obtain key information from full nodes to reduce the storage burden and improve the query efficiency; After the election, full nodes perform the data deduplication task and ensure that the pruned data is traceable and verifiable; Full nodes perform pruning determination on the stored data according to the rules of Prune(data) based on data timeliness, Prune sim (data p , data q ), and Prune based on access frequency freq (data); if the timestamp t of the data data is earlier than the set threshold t now of the current time t τ , it is marked as data to be pruned, reducing the storage of expired data; if the data data p and other data data q within the group are more similar than the threshold ι on the feature vectors X p and X q , only the representative data is retained, and the rest of the data is marked as data to be pruned; if the access frequency f of the data data within the set time is lower than the threshold f τ , it is marked as data to be pruned, releasing storage space; data that meets any pruning condition is marked as data to be pruned and enters the next stage of processing. For the data marked as prunable, the full node performs data storage optimization; data deletion causes the full node to remove expired, redundant or low-frequency data from local storage, releasing storage resources; data compression is performed on the data that still has some value, and hash indexing and Merkle trees are used for compressed storage; Full nodes perform data synchronization through the blockchain consensus mechanism; after the pruning operation is completed through data broadcasting, full nodes broadcast update information to the network to notify all participating nodes; full nodes verify signatures Sign χ to ensure the authenticity of the pruned data; ordinary nodes re-pull the latest data through trusted full nodes to ensure that their own data is consistent with the entire network; after consensus confirmation, the pruning status is officially recorded on the blockchain to ensure the immutability of the data.