A cloud computing-based user travel data security management system and method

By employing multi-source data acquisition and encrypted transmission, privacy protection, distributed storage, and access control, the system addresses the issues of lag and privacy threats in traditional systems handling massive amounts of data, achieving secure and efficient data management and analysis.

CN120408577BActive Publication Date: 2026-01-02BEIJING TRAVEL INT TOURISM TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510359647.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-01-02
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Traditional data security management systems are prone to lag and data loss when massive amounts of data flood in. They suffer from inefficient processing, lax access control, and are unable to meet the needs of real-time scheduling and precision marketing, while also threatening user privacy.

Method used

It employs a multi-source data acquisition and adaptation module, a hybrid encrypted transmission channel module, a privacy verification and desensitization module, a distributed storage architecture module, an access control and auditing module, and a data backup and recovery scheduling module. Combined with quantum key distribution, AES encryption, deep learning models, and blockchain technology, it achieves secure data acquisition, transmission, storage, analysis, and access control.

Benefits of technology

Ensure the confidentiality and integrity of data transmission, accurately identify privacy-sensitive content, guarantee high availability and security of data, achieve transparent access and efficient recovery of data, and enhance data value mining and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408577B_ABST
    Figure CN120408577B_ABST
Patent Text Reader

Abstract

The application provides a kind of cloud computing-based user travel data security management system and method, system includes: multi-source data acquisition adaptation module, for from multiple types of travel equipment User travel data acquisition;Mixed encryption transmission channel module, for between data acquisition terminal and cloud computing platform Establish secure transmission link;Privacy discrimination and desensitization module, for identifying and processing the privacy sensitive content in the data received by cloud computing platform;Distributed storage architecture module, for reliably storing user travel data on cloud computing platform;Authority control and audit module, for controlling the calling authority of external subject to user travel data and auditing data access situation;Data backup and recovery scheduling module, restore data when there is data risk;Data visualization and analysis auxiliary module, for presenting data in visualized mode.The application ensures the confidentiality, integrity and availability of travel data transmission, the transparency and traceability of data access.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data management, in particular to a user travel data security management system and method based on cloud computing. BACKGROUND

[0002] The user travel data security management system based on cloud computing is a system that collects, stores, processes and analyzes user travel data using cloud computing technology. The system provides powerful computing and storage capabilities through a cloud platform to ensure the security, privacy and efficiency of data. It uses encryption, desensitization and access control technologies to protect user data security, and uses intelligent algorithms to analyze travel data to provide personalized services.

[0003] In today's digital age, intelligent travel is booming, and various travel APPs have become an indispensable tool for people's daily travel. Massive travel data, including real-time geographic location, travel preferences, frequently visited locations, and accurate travel time, are collected and stored. Once these data fall into the hands of criminals, they can easily pinpoint users through data mining, correlation analysis, and other means, and implement harassment, fraud, and even more serious personal threats.

[0004] However, the storage architecture of traditional data security management systems lacks flexibility and is difficult to adapt to the influx of massive data during peak periods, frequently causing lag and even data loss. The data processing process is lengthy and inefficient, and valuable travel insights cannot be timely feedback, making it difficult to meet the needs of real-time traffic operation scheduling, precise marketing, and other business needs. The authority control is extensive, and internal personnel abuse of authority and external illegal access continue to occur, seriously threatening user privacy. Therefore, it is urgent to develop innovative security management systems and methods. SUMMARY

[0005] The technical solutions provided by the embodiments of the present application are as follows:

[0006] First aspect:

[0007] The multi-source data collection and adaptation module is used to collect user travel data from various types of travel devices;

[0008] The mixed encryption transmission channel module is used to establish a transmission link between the data collection terminal and the cloud computing platform;

[0009] The privacy identification and desensitization module is used to identify and process privacy sensitive content in the data received by the cloud computing platform;

[0010] The distributed storage architecture module is used to store user travel data on the cloud computing platform;

[0011] A permission control and audit module is configured to control the calling permission of an external subject to user travel data and audit data access;

[0012] A data backup and recovery scheduling module is configured to recover data when there is a data risk through a backup strategy, a backup task execution, a monitoring process, and a recovery process;

[0013] A data visualization and analysis assistance module is configured to present user travel data in a visualized manner, perform in-depth analysis on the user travel data by combining a data analysis algorithm, and arrange analysis results into a report.

[0014] The second aspect is:

[0015] The user travel data security management method based on cloud computing provided by the embodiment of the application is applied to the user travel data security management system based on cloud computing of the first aspect, and includes the following steps.

[0016] S1: Real-time monitoring of user travel state changes is performed at each access device end, and when it is detected that a user performs a travel-related behavior, a data collection process of the corresponding device is activated, user travel data is accurately collected according to an adaptive strategy, and the user travel data is subjected to preliminary format standardization and error checking;

[0017] S2: A symmetric key is obtained by using a quantum key distribution mechanism, and according to the symmetric key, the user travel data that has passed the checking is subjected to transmission segmentation and parallel encryption by using an AES algorithm, a plurality of encrypted data groups are obtained, a message authentication code is attached to each of the encrypted data groups, after the encryption and packaging are completed, the encrypted data groups are transmitted at a high speed to a cloud computing platform through a special network channel;

[0018] S3: After the cloud computing platform receives the encrypted data groups, the privacy screening and desensitization module scans each of the encrypted data groups row by row by using a deep neural network model, identifies the privacy sensitive content of each of the encrypted data groups according to a training and learning result, determines privacy data, and performs real-time desensitization processing on the privacy data according to a desensitization rule;

[0019] S4: The distributed storage architecture module receives the privacy data that has been desensitized, cuts the privacy data that has been desensitized into a plurality of data segments in parallel according to a data sharding strategy, uniformly stores the data segments to storage nodes in each geographical region, and synchronously generates redundant copies to guarantee data security, based on high-frequency access requirements, uses edge cache nodes to intelligently cache popular data segments related to the high-frequency access requirements;

[0020] S5: The data backup and recovery scheduling module formulates multiple backup strategies according to the importance, update frequency and regulatory requirements of the door data segments, and dynamically adjusts the priority of the backup tasks; utilizes a distributed task scheduling framework to distribute the backup tasks to multiple computing nodes and storage resources for parallel processing; when data recovery is required, specific data at a specific time point in the past is selected for recovery according to the requirements, and specific geographic area data or user travel data in the specific data is recovered;

[0021] S6: The data visualization and analysis assistance module selects appropriate visualization chart types according to the characteristics of the user travel data and the analysis requirements of the user; starts built-in data analysis algorithms to run on the desensitized data, mines the potential value of the data on the premise of protecting the privacy of the user; generates a standardized analysis report template automatically according to the analysis task of the user and the visualization result;

[0022] S7: When an external subject initiates a data calling request, the permission control and audit module receives the request information, and the smart contract strictly examines the identity, permission range and access purpose elements of the requestor based on the permission ledger of the blockchain; if the verification is passed, the corresponding data is retrieved and extracted from the storage node, and after decryption, the data is provided to the requestor; if the verification is not passed, the request is rejected and detailed audit logs are recorded for subsequent traceability analysis.

[0023] The technical scheme provided by the embodiment of the application has at least the following beneficial effects:

[0024] In the embodiment of the application, the mixed encryption transmission channel module ensures the confidentiality, integrity and availability of data transmission through quantum key distribution and AES encryption, effectively prevents hacker attacks, data theft and tampering, the privacy discrimination and desensitization module accurately identifies privacy sensitive content and protects personal identity information and high-precision geographic location while retaining the analysis value of the data, the distributed storage architecture module adopts multi-dimensional sharding and redundancy backup technology to ensure the high availability, security and efficient access of data, the permission control and audit module establishes an unalterable permission ledger through blockchain technology and monitors data abuse risks to ensure the transparency and traceability of data access, the data backup and recovery scheduling module ensures the safe backup of key data and supports efficient and flexible data recovery to guarantee the integrity and consistency of data, the data visualization and analysis assistance module provides rich charts and custom layouts, supports deep data analysis and user classification, and improves data value mining and user experience. BRIEF DESCRIPTION OF DRAWINGS

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a block diagram illustrating the principle of the cloud-based user travel data security management system of the present invention.

[0027] Figure 2 This is a block diagram illustrating the principle of the multi-source data acquisition and adaptation module of the present invention.

[0028] Figure 3 This is a schematic diagram of the hybrid encrypted transmission channel module of the present invention;

[0029] Figure 4 This is a block diagram illustrating the principle of the privacy verification and desensitization module of the present invention.

[0030] Figure 5 This is a block diagram illustrating the principle of the distributed storage architecture module of the present invention;

[0031] Figure 6 This is a block diagram illustrating the principle of the access control and auditing module of the present invention. Detailed Implementation

[0032] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0033] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0034] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0035] Reference manual attached Figure 1 The diagram illustrates the principle block diagram of a cloud-based user travel data security management system provided by the present invention.

[0036] like Figure 1, multi-source data collection adaptation module 101, for collecting user travel data from various types of travel equipment, mixed encryption transmission channel module 102, for establishing a secure transmission link between the data collection terminal and the cloud computing platform, privacy discrimination and desensitization module 103, for identifying and processing privacy sensitive content in the data received by the cloud computing platform, distributed storage architecture module 104, for reliable storage of user travel data on the cloud computing platform, permission control and audit module 105, for controlling the calling permission of external subjects to user travel data and auditing data access, data backup and recovery scheduling module 106, through backup strategy, backup task execution and monitoring and recovery process, restore data when there is data risk, data visualization and analysis assistance module 107, for presenting data in a visualized manner, combining data analysis algorithms to perform in-depth analysis on travel data, and organizing analysis results into reports.

[0037] Reference the accompanying drawings Figure 2 , the multi-source data collection adaptation module principle block diagram provided by the embodiment of the application is shown.

[0038] As Figure 2 , device interface adaptation sub-module 1011, for configuring exclusive interface drivers for not limited to smart phones, vehicle intelligent terminals, wearable devices, for smart phones, adapting iOS and Android system underlying various types of sensor interfaces, such as using Core Motion framework (iOS) and SensorManager class (Android) to accurately interface GPS, accelerometer, gyroscope and other sensors, to ensure stable data acquisition; for vehicle intelligent terminals, develop conversion programs to adapt to different vehicle OBD interface protocols, covering common protocols such as CAN, LIN, to realize lossless reading of vehicle operating parameters; for wearable devices, customize communication protocols based on Bluetooth Low Energy (BLE), to ensure stable pairing and data transmission with mobile phone APP.

[0039] Collection frequency optimization sub-module 1012, dynamically adjusts the collection frequency according to device power, network status and user travel state, through the built-in intelligent power monitoring algorithm, when the device power is lower than 20%, the non-critical sensor collection frequency is automatically reduced by 50%; combined with network signal strength detection, in weak network environment (signal strength lower than -90dBm), reduce data upload frequency, preferentially cache locally, and transmit in batches after network recovery; through machine learning model, analyze user travel behavior pattern in real time, if judge user is in static or regular commuting state, appropriately reduce high-frequency sensor collection frequency, once identify that user enters emergency travel (such as catch a plane, train) or special travel scene (such as travel and exploration), immediately improve the key data collection frequency.

[0040] The data preprocessing submodule 1013 is configured to perform denoising and time synchronization calibration on the collected data at the device end, and encapsulate the data into a unified data frame in a preset format, and remove the noise of GPS positioning data by using a Kalman filtering algorithm.

[0041] Referring to the accompanying drawings Figure 3 , a principle block diagram of the mixed encryption transmission channel module is shown.

[0042] As Figure 3 , the quantum key distribution submodule 1021 is configured to deploy quantum key distribution (QKD) terminal equipment in data collection intensive areas (such as urban traffic hubs and large parking lots) and cloud computing data centers, and is matched with a professional operation and maintenance system. The submodule is responsible for the initialization configuration of the QKD equipment, the adjustment of the key generation parameters, and the management of the key storage and update. The error rate of the quantum channel is monitored regularly, and when the error rate exceeds 1%, the channel calibration program is automatically triggered. At the same time, according to the data transmission demand, the quantum key resources are dynamically allocated to ensure that high-priority data (such as real-time location information) can obtain key encryption preferentially.

[0043] The AES encryption execution submodule 1022 is configured to select a high-performance encryption chip with hardware acceleration capability or use a GPU acceleration module of a cloud computing platform to build a parallel encryption processing architecture. The round function calculation process of the AES algorithm is optimized, the pipeline technology is used to reduce the calculation delay, and for large data block transmission, the data is segmented into multiple fixed-size subblocks (such as 128 bits) for parallel encryption to improve the encryption efficiency. According to the real-time encryption task quantity, the cloud computing resources are dynamically scheduled to ensure that the encryption processing does not become a bottleneck of data transmission, and the security of data in public network transmission is ensured.

[0044] The authentication code processing submodule 1023 is configured to generate and verify a message authentication code (MAC) by using an HMAC-SHA256 algorithm. At the data sending end, a 256-bit authentication tag is calculated and generated for each data packet, and the related key information and time stamp of the tag generation are recorded. The receiving end verifies the received data by using the same key and algorithm, and if the tags are inconsistent, the data retransmission mechanism is triggered immediately, and the abnormal information is fed back to the security audit center, so that the data integrity is ensured and the data in the transmission process is prevented from being tampered with.

[0045] Referring to the accompanying drawings Figure 4 , a principle block diagram of the privacy screening and desensitization module is shown.

[0046] As Figure 4Data sample management submodule 1031, jointly with big data companies in the travel industry and research institutions, widely collects massive travel data samples in different regions, cultures and industry backgrounds worldwide. A sample database management system is constructed to classify, store, version manage and regularly update the samples. According to the travel mode (such as public transportation, subway, self-driving, cycling, etc.), travel scenario (commuting, tourism, business travel, etc.), user group (age, occupation, gender, etc.) and other multi-dimensional classifications, the samples can be accurately selected for subsequent model training; new samples are regularly obtained from internet public data, industry reports and other channels to update the sample library, ensuring the timeliness and diversity of the model training data.

[0047] Deep learning model construction submodule 1032, selects a long short-term memory network (LSTM) suitable for processing sequence data as the basic architecture, and builds a multi-layer hybrid neural network model combining the feature extraction advantages of a convolutional neural network (CNN). The spatial features (such as geographic location distribution, regional travel hotspots) in the travel data are extracted using CNN, and then the feature map is input into the LSTM network to learn the time sequence features, accurately judging the privacy sensitive content. During model training, an adaptive learning rate adjustment strategy is adopted to dynamically optimize the learning rate according to the training loss change, combined with the early stopping method (Early Stopping) to prevent overfitting, and continuously improve the model recognition accuracy and desensitization precision.

[0048] Desensitization rule execution submodule 1033, formulates detailed desensitization rule library according to different privacy levels and data types, and embeds it into the model output end. For high-precision geographic location information, region fuzzing, coordinate offset and other technologies are used to convert accurate coordinates into geographic area description within a certain range or coordinate values after random offset; for personal identity information associated travel data such as name, ID number, etc., anonymization processing is used to replace the real information with a unique identifier; for sensitive place access records, the specific place name is fuzzed, only the place category information (such as medical institutions, religious places, etc.) is retained, which retains part of the analysis value of the data and effectively protects the core content of user privacy.

[0049] Reference the attached drawings Figure 5 , which shows the principle block diagram of the distributed storage architecture module provided by the application.

[0050] As Figure 5The data slice management submodule 1041 designs a composite slice function based on a user ID, a travel timestamp, a geographical area code and multiple dimensions. The user travel data is uniformly sliced by using a consistent hash algorithm, for example, the national travel data is divided into the East China, South China and the like large areas according to the geographical code, and each large area is further subdivided according to the user ID, so that the data slice is convenient for management and can balance the storage load; meanwhile, a slice index management system is established to track and record the storage location, replica distribution and the like information of each slice in real time, so as to facilitate fast data retrieval and recovery.

[0051] The redundancy backup control submodule 1042 adopts a Reed-Solomon error correction code technology, sets different redundancies according to the data importance and reliability requirements. For general data, the redundancy is 2 times, and for key core data, the redundancy is 3 times. A redundancy backup scheduling program is developed to store the redundant copies in a distributed manner among the storage nodes. The consistency of the redundant copies is checked regularly. When it is found that the data of the copies is inconsistent, the error correction code algorithm is used to recover the error copy from other normal copies, so as to guarantee the high availability of the data and ensure that the original data can be quickly recovered even if part of the nodes are physically damaged or data is lost.

[0052] The distributed cache scheduling submodule 1043 selects an open source distributed cache system (such as Redis) to configure a cache server on an edge node close to the user request end. An intelligent cache scheduling algorithm is developed to dynamically manage the cache space by using a cache eviction algorithm (such as LRU), intelligently decides the cache content according to the data access heat, time limit and the like factors, combines the data prefetching technology, and according to the historical access behavior of the user and the real-time travel trend, the data that can be accessed is cached to the edge node in advance, so as to improve the access speed of the hot data, optimize the user experience, and reduce the delay of data back source reading.

[0053] The accompanying drawings Figure 6 illustrate the principle block diagram of the permission control and audit module provided by the application.

[0054] As Figure 6 The blockchain platform building submodule 1051 selects a mainstream blockchain open source framework (such as Hyperledger Fabric) to build an enterprise-level blockchain platform. It is responsible for the deployment, network configuration and initialization of the genesis block generation of the blockchain node. An independent identity certificate is created for each data access subject, and the security of the certificate is ensured by using the asymmetric encryption technology. Detailed permission information is registered on the blockchain, including the key information such as the accessible data range, the access time limit, the use purpose restriction and the like, and the distributed ledger characteristics of the blockchain are used to ensure that the permission record is tamper-proof.

[0055] The permission check execution submodule 1052 automatically obtains the identity and permission information of the requestor from the blockchain when there is a data call request, and compares the request details. The comparison content covers key indicators such as accessible data time period, geographic area, data type, etc. If it is completely matched, the data retrieval instruction is sent to the storage module; if it is not matched, the alarm notification system administrator is triggered through the blockchain event mechanism, and the request details are recorded in the audit log area on the blockchain. The identity of the requestor, the scope of the permission, the purpose of the access and other elements are strictly reviewed.

[0056] The audit backtracking analysis submodule 1053 develops an audit data analysis system using a blockchain browser tool. The administrator can view all data access history records at any time, including requestor, access time, obtained data content and other detailed information; the audit data is analyzed through a data mining algorithm to mine potential data abuse risks such as frequent access to data beyond business needs, concentrated access in abnormal time periods; the audit report is generated regularly to provide a basis for system optimization and compliance management, and the transparency and traceability of data call history are realized.

[0057] The embodiment of the application provides a user travel data security management system based on cloud computing, which comprises:

[0058] The multi-source data collection adaptation module 101 is used for collecting user travel data from various types of travel equipment;

[0059] The mixed encryption transmission channel module 102 is used for establishing a transmission link between the data collection terminal and the cloud computing platform;

[0060] The privacy discrimination and desensitization module 103 is used for identifying and processing privacy sensitive content in the data received by the cloud computing platform;

[0061] The distributed storage architecture module 104 is used for storing user travel data on the cloud computing platform;

[0062] The permission control and audit module 105 is used for controlling the calling permission of external subjects to user travel data, and auditing data access;

[0063] The data backup and recovery scheduling module 106 recovers data when there is data risk through backup strategy, backup task execution, monitoring process and recovery process;

[0064] The data visualization and analysis auxiliary module 107 is used for presenting user travel data in a visualized manner, combining data analysis algorithm to deeply analyze user travel data, and arranging analysis results into a report.

[0065] It should be noted that by device interface adaptation, dynamic adjustment of collection frequency and data preprocessing (such as denoising, time synchronization and Kalman filtering), the data collection process is optimized, the device load and network pressure are reduced, the accuracy and efficiency of the collected data are ensured, and thus the performance and data quality of the system are improved.

[0066] In a possible implementation, the multi-source data collection adaptation module comprises:

[0067] The device interface adaptation sub-module 1011 is configured to configure exclusive interface drivers for smartphones, vehicle-mounted intelligent terminals and wearable devices.

[0068] The collection frequency optimization sub-module 1012 is configured to dynamically adjust the collection frequency according to the device power, network status and user travel state.

[0069] The data preprocessing sub-module 1013 is configured to perform denoising and time synchronization calibration on the collected data at the device end, encapsulate the data into a unified data frame according to a preset format, and remove the noise of GPS positioning data by using a Kalman filtering algorithm.

[0070] In a possible implementation, the mixed encryption transmission channel module 102 comprises:

[0071] The quantum key distribution sub-module 1021 is configured to deploy quantum key distribution terminal devices in data collection intensive areas and cloud computing data centers, and is matched with an operation and maintenance system, which is responsible for initialization configuration of the QKD device, adjustment of key generation parameters, key storage and key update management.

[0072] The AES encryption execution sub-module 1022 is configured to build a parallel encryption processing architecture, optimize the round function calculation process of the AES algorithm, perform transmission segmentation and parallel encryption on large data blocks, restore the data after completing the verification of the cloud computing user's verifiable static data storage integrity, and dynamically schedule cloud computing resources according to real-time encryption task quantity.

[0073] The authentication code processing sub-module 1023 is configured to calculate and generate a 256-bit authentication tag for each data packet at the data sending end by using the HMAC-SHA256 algorithm, record the related key information and time stamp of the authentication tag generation, and verify the authentication tag at the receiving end by using the same key and algorithm. If the tags are inconsistent, the data retransmission mechanism is triggered immediately, and abnormal information is fed back to the security audit center.

[0074] Among them, the QKD device is an encryption technology based on quantum physics principles, used to securely distribute encryption keys, transmit keys through quantum bits (qubits), and use quantum mechanics properties (such as quantum superposition, quantum entanglement) to ensure the security of the keys, AES (Advanced Encryption Standard) is a symmetric encryption algorithm, one of the most widely used encryption standards, uses the same key for data encryption and decryption, supports different key lengths (128 bits, 192 bits, 256 bits), HMAC-SHA256 (Hash-based Message Authentication Code with SHA-256) is a message authentication code algorithm based on the SHA-256 hash function.

[0075] It should be noted that through quantum key distribution (QKD), AES encryption and authentication code processing, multi-level encryption and authentication protection is provided to ensure the confidentiality and integrity of data during transmission. Quantum key distribution ensures the security of key transmission, AES encryption optimizes the encryption efficiency of large data blocks, and HMAC-SHA256 authentication code processing further enhances data verification and reliability, preventing data tampering and loss, and improving the security and efficiency of the entire system.

[0076] In one possible implementation, large data blocks are transmitted, segmented and encrypted in parallel, and after verifying the integrity of the cloud computing user's verifiable static data storage, the data is restored, specifically including:

[0077] According to the internal logical structure or predetermined rules of the data, the large data block is segmented into multiple fixed-size sentence blocks, and the AES algorithm is used to assign independent encryption keys to each data sub-block, wherein the encryption keys are dynamically generated and distributed by a key management system.

[0078] After the data is stored on the cloud computing platform, a verifiable storage integrity verification mechanism is adopted, based on the hash function message authentication code HMAC technology. When the data is stored, the hash value of each data sub-block is calculated and stored together with the data sub-block. When the data needs to be restored, the hash value of the stored data sub-block is recalculated and compared with the previously stored hash value. If the hash values are consistent, the segmented sub-blocks are reassembled in the original order by the data restoration program to restore the complete data file. If the hash values are not consistent, the data repair or alarm mechanism is triggered to notify the administrator that the data may be damaged or tampered with.

[0079] In the present application, large data blocks are divided into multiple fixed-size sub-blocks according to the internal logical structure of the data or predefined rules, for example, a data file with a capacity of 1 GB is divided into units of 128 MB. For each sub-block, a multi-core processor or GPU acceleration resource provided by a cloud computing platform is used to start the encryption process in parallel. Selecting an algorithm such as the Advanced Encryption Standard (AES), each sub-block is assigned an independent encryption key, which can be dynamically generated and distributed by a key management system to ensure the security of the encryption process. Through parallel encryption, the powerful computing power of cloud computing is fully utilized, greatly improving the encryption efficiency, shortening the transmission preparation time of large data blocks, and meeting the application scenarios with high real-time requirements.

[0080] After the data is stored in the cloud computing platform, in order to ensure the integrity of the data, a verifiable storage integrity verification mechanism is adopted. For example, using the Hash-based Message Authentication Code (HMAC) technology, when the data is stored, the hash value of each data sub-block is calculated and stored together with the data sub-block. When data recovery is needed, first calculate the hash value of the stored data sub-block, and compare it with the previously stored hash value. If the hash values are consistent, it means that the data has not been tampered with during storage, and the segmented sub-blocks are reassembled in the original order through the data recovery program to restore the complete data file. If the hash values are inconsistent, the data repair or alarm mechanism is triggered, notifying the administrator that the data may be damaged or tampered with, and taking appropriate measures in time to ensure the reliability and availability of user data.

[0081] It should be noted that by dividing large data blocks and assigning independent encryption keys to each sub-block, the encryption processing efficiency and security are improved, combined with the Hash-based Message Authentication Code (HMAC) technology, ensuring the integrity and verifiability of data storage, and when data recovery, through hash value verification, ensures that the data has not been tampered with during transmission or storage. If found to be inconsistent, the repair mechanism can be triggered in time to effectively prevent data damage or loss, ensuring the high availability and reliability of the data.

[0082] In one possible implementation, the privacy screening and desensitization module 103 includes:

[0083] The data sample management sub-module 1031 is used to collect travel data samples in different regions, cultures and industry backgrounds, and build a sample database management system based on the travel data samples. The user travel data samples are stored, version managed and regularly updated, and classified according to travel mode, travel scenario and user group in multiple dimensions.

[0084] The deep learning model construction submodule 1032 is configured to select a long short-term memory network combined with a convolutional neural network to build a multi-layer hybrid neural network model, use the CNN to extract spatial features of the user travel data, determine a feature map, and input the feature map into the LSTM network to learn time sequence features to determine the privacy-sensitive content.

[0085] The CNN is configured to extract spatial features of the user travel data, determine a feature map, and input the feature map into the LSTM network to learn time sequence features to determine the privacy-sensitive content, and specifically includes the following steps.

[0086] The CNN is configured to extract spatial features of the user travel data, determine a feature map, and input the feature map into the LSTM network to learn time sequence features to determine the privacy-sensitive content, and specifically includes the following steps.

[0087]

[0088] wherein a i,j represents an activation value at the feature map (i, j) position after the convolution operation, f represents an activation function, w m,n represents a corresponding weight at the convolution kernel position (m, n), x i+m,j+n represents a value at the position (i+m, j+n) in the input data, m=1, 2,..., M, M represents the total number of rows of the convolution kernel, n=1, 2,..., N, N represents the total number of columns of the convolution kernel, and b represents a bias.

[0089] The feature map is flattened.

[0090] x t =flatten(a i,j )

[0091] wherein x t represents the flattened feature map, and flatten represents a flattening operation.

[0092] The flattened feature map is input into the LSTM network to determine a hidden state.

[0093] i t =σ(w xi x t +w hi h t-1 +b i )

[0094] f t =σ(w xf x t +w hf h t-1 +b f )

[0095] o t =σ(w xo x t +who h t-1 +b o )

[0096]

[0097] h t =o t ·tanh(c t )

[0098] where i t denotes the input gate, f t denotes the forget gate, o t denotes the output gate, c t denotes the candidate memory state at time t, c t denotes the decided memory cell state at time t, σ denotes the Sigmoid function, w xi denotes the weight matrix of the input data x t to the input gate, w hi denotes the weight matrix of the hidden state h t-1 to the input gate, b i denotes the bias vector of the input gate, w xf denotes the weight matrix of the input data x t to the forget gate, w hf denotes the weight matrix of the hidden state h t-1 to the forget gate, b f denotes the bias vector of the forget gate, w xo denotes the weight matrix of the input data x t to the output gate, w ho denotes the weight matrix of the hidden state h t-1 to the output gate, b o denotes the bias vector of the output gate, w xc denotes the weight matrix of the input data x t to the candidate memory state, w hc denotes the weight matrix of the hidden state h t-1 to the candidate memory state, b c denotes the bias vector of the candidate memory state, tanh denotes the tanh activation function.

[0099] determining a predicted score of the privacy-sensitive content according to the hidden state:

[0100] y = W f · h t + b s

[0101] where y denotes the predicted score of the privacy-sensitive content, W f denotes the weight matrix of the fully connected layer, b sbias of the fully connected layer.

[0102] When the prediction score is greater than the preset prediction score, it is determined that the travel data has privacy-sensitive content.

[0103] The desensitization rule execution submodule 1033 is configured to formulate a desensitization rule library according to different privacy levels and data types, and embed the desensitization rule library into a model output end. For high-precision geographic location information, regional fuzzification and coordinate offset are adopted. For personal identity information associated travel data, anonymization processing is adopted. For sensitive place entry and exit records, specific place names are fuzzified.

[0104] The fuzzification specifically includes: dividing a regional map into geographic grid cells of uniform size according to certain rules, determining a grid cell to which a high-precision geographic location coordinate point belongs when the high-precision geographic location coordinate point is acquired, and fuzzily describing location information of a user as a grid region in which the user is located. For high-precision longitude and latitude coordinates, coordinate offset is performed by adding randomly generated direction and position noise values.

[0105] The desensitization rule execution submodule adopts regional fuzzification and coordinate offset for high-precision geographic location information, and specifically includes:

[0106] The map is divided into geographic grid cells of uniform size according to certain rules. For example, the map can be divided according to street division in city planning, subdivision of administrative regions, or longitude and latitude intervals (for example, each 0.01 degree x 0.01 degree is a grid cell). When a high-precision geographic location coordinate point is acquired, the grid cell to which the coordinate point belongs is determined, and location information of a user is fuzzily described as a grid region in which the user is located. For example, a coordinate point located near Wangfujing Street in Beijing is classified into a “Wangfujing commercial district grid cell” after geographic grid division, and the location information displayed to the outside world becomes “the user is located in the Wangfujing commercial district” instead of being accurate to a specific house number or store coordinate. The advantage of this method is that it can hide precise location details while retaining certain regional characteristics. Moreover, the more precise the grid division, the more accurate the description of regional characteristics, which can meet the application requirements of some regional analysis, such as statistical analysis of the trend of the number of people in a commercial district, and protect the privacy of users.

[0107] For high-precision latitude and longitude coordinates, random direction and position noise values are added to perform coordinate offset. For example, let the original coordinates be (x, y), the offset of the longitude is generated by a random number generator, and the offset of the latitude is generated by a random number generator, and satisfies a certain distribution rule, such as normal distribution, the mean is 0, and the standard deviation is set according to the required privacy protection level and the positioning accuracy requirement of the application scene (if the application does not require high positioning accuracy and requires strong privacy, the standard deviation can be set relatively large). The new coordinates become (x', y'), so when data is stored or shared externally, the offset coordinates are used, and even if the data is leaked, it is difficult for attackers to restore the real accurate position. For example, a random angle between 0 and 360 degrees is selected, and a pre-set offset distance (also set according to privacy protection requirements, such as 100 meters, 500 meters, etc.) is used to calculate the specific offset of the longitude and latitude through a trigonometric function.

[0108] It should be noted that by combining convolutional neural networks (CNN) and long short-term memory networks (LSTM), the spatial and temporal features of the trip data are efficiently extracted, thereby accurately identifying privacy-sensitive content. At the same time, using desensitization techniques such as regional fuzzing, coordinate offsetting, and anonymization processing, the security of user privacy data is protected, and through multi-dimensional classification and regular sample management, the accuracy and adaptability of privacy identification are further improved, effectively preventing privacy leakage and ensuring data security compliance.

[0109] In one possible implementation, the distributed storage architecture module 104 includes:

[0110] The data sharding management submodule 1041 is configured to design a composite sharding function in multiple dimensions based on user ID, trip timestamp, and geographic region code, uniformly shard user trip data using a consistent hashing algorithm, and establish a sharding index management system to track shard storage locations and replica distribution information.

[0111] The redundancy backup control submodule 1042 is configured to use Reed-Solomon error correction code technology to develop a redundancy backup scheduling program for distributed storage of redundant replicas, regularly perform consistency checking on the redundant replicas, and use Reed-Solomon error correction code technology to recover erroneous replicas.

[0112] The Reed-Solomon error correction code technology is a widely used error correction coding technology in data storage and transmission, mainly used to solve the problems of data loss and damage.

[0113] The distributed cache scheduling submodule 1043 is configured to select an open-source distributed cache system, configure cache servers on edge nodes close to user request ends, develop intelligent cache scheduling algorithms, dynamically manage cache space using intelligent cache eviction algorithms, and combine data prefetching technology to cache data that may be accessed in advance.

[0114] It should be noted that the efficient storage and high availability of user travel data are ensured by data sharding management, consistent hashing and redundant backup technology, the recovery and consistency check of data copies are realized by Reed-Solomon error correction code technology, the reliability of data is ensured, the distributed cache scheduling optimizes cache space management and data prefetching, improves the data access speed, reduces the main storage pressure, and further improves the performance and response efficiency of the system.

[0115] In a possible implementation, the permission control and audit module 105 specifically includes:

[0116] The blockchain platform building submodule 1051 is configured to select a blockchain open source framework to build an enterprise-level blockchain platform, is responsible for deployment of a blockchain node, network configuration and initialization of a genesis block generation, creates an independent identity certificate for each data access subject, and registers detailed permission information on the blockchain.

[0117] The permission verification execution submodule 1052 is configured to, in the case of a data call request, automatically obtain identity information and permission information of a requestor from the blockchain through a smart contract, and compare the identity information and the permission information with request details, covering key indicators such as accessible data time period, geographic area and data type. If the match is matched, a retrieval instruction is sent to the storage module, and if the match is not matched, an alarm is triggered and the request information is recorded to an audit log area.

[0118] The audit backtracking analysis submodule 1053 is configured to develop an audit data analysis system through a blockchain browser tool. An administrator can view data access history records, mine data abuse through a data mining algorithm, and generate an audit report.

[0119] Generating the audit report specifically includes:

[0120] A correlation rule mining algorithm is selected to find frequently occurring request patterns, and a clustering analysis algorithm is used to group data access requests according to similarity. When a sudden concentrated access to sensitive travel data occurs from a strange IP address that has never had data access permission, it indicates a potential data abuse risk.

[0121] The execution result of the data mining algorithm is taken as input to sort out potential data abuse clues, verify whether the real identity, permission range and access purpose of the relevant requestor are consistent with the records, and generate an audit report according to the verification result according to a predetermined report template. The audit report includes a summary of potential problems discovered by data mining, detailed problem clue descriptions, problem verification situation descriptions, and suggested measures for discovered problems.

[0122] Among them, the data abuse clues include abnormal frequent access patterns and suspicious access clusters obtained by clustering analysis.

[0123] Need to be explained, through the blockchain technology for each data access subject to provide independent identity certificate, and use smart contract to check the right automatically, ensure the transparency and traceability of data access, combined with audit backtracking and data mining algorithm, the system can real-time monitoring of potential data abuse risk, generate detailed audit report, help administrator to identify abnormal access mode and abuse behavior, enhance data security and compliance, prevent illegal access and abuse of power.

[0124] In the present application, the association rule mining algorithm is selected, such as Apriori algorithm or its improved version, to analyze the data access log. These logs record all data call requests, including requesters, access time, obtained data content and other detailed information. Through association rule mining, it tries to find out the frequently occurring request patterns, such as a certain specific institution frequently accessing the travel data of users in a certain area within a short time, or the travel data of a certain user group is always called by certain specific third parties within a similar time period.

[0125] At the same time, the clustering analysis algorithm is used to group the data access requests according to similarity. It can be clustered according to the type of requester (such as enterprise, government department, scientific research institution, etc.), the type of data accessed (such as high-precision geographic location data, travel trajectory data, travel habit data, etc.) and the distribution of access time, etc. This helps to find abnormal access clusters, such as a group of strangers IP addresses that have never had data access rights suddenly access a certain type of sensitive travel data, which may indicate potential data abuse risk.

[0126] Periodically (e.g., weekly or monthly), an audit report generation task is initiated. First, the execution results of the data mining algorithms are taken as input, and potential data misuse clues are sorted out. These clues include abnormal frequent access patterns discovered through association rule mining, suspicious access clusters obtained through clustering analysis, etc. Each clue is investigated and verified in detail. Interact with the data access permission management system to verify whether the real identity, permission range, and access purpose of the relevant requestor are consistent with the records. For example, for the discovered institution that frequently accesses a certain type of data, check whether its originally applied data access permission exceeds the authorized range of operation. According to the verification results, generate an audit report according to the predetermined report template. The report content should include an overview of the potential problems discovered by data mining, detailed description of the problem clues (such as the requestor involved, access time, data type, etc.), problem verification situation explanation, and proposed measures for the discovered problems, such as suspending the data access permission of the violating requestor, strengthening the access control of specific data types, etc. The audit report is presented in a clear and easy-to-understand format for reference by system administrators, compliance departments, and top decision-makers, so as to take timely measures to prevent data misuse risks and ensure the safe and compliant operation of the system.

[0127] In a possible implementation, the data backup and recovery scheduling module specifically comprises:

[0128] The backup strategy formulation submodule is configured to formulate different backup strategies according to the importance, update frequency, and regulatory requirements of the data, and dynamically adjust the priority of backup according to the storage time and access heat of the data.

[0129] The way of dynamically adjusting the priority of backup according to the storage time and access heat of the data specifically comprises:

[0130] Determine the priority of backup:

[0131] CP(x) = W1 · UF(x) + W2 · k(x) + W3 · PC(x) + W4 · evi(x) + W5 · hotness(x)

[0132] Wherein, CP() represents the priority of backup, W1, W2, W3, W4 and W5 all represent weights, UF() represents the update frequency of data, k() represents the criticality of data, PC() represents the access priority of data, evi() represents the validity period of data, and hotness() represents the heat of data.

[0133] Dynamically adjust the priority of backup according to the storage time and access heat:

[0134] CP(x)' = (α · CP(x) · e -λt ) + (1-α) · H

[0135] Wherein, CP()'indicates the dynamically adjusted priority, a indicates the weight coefficient, l indicates the attenuation coefficient, t indicates the storage time, e indicates the index, H indicates the access times of data.

[0136] The backup task execution and monitoring submodule utilizes a distributed task scheduling framework to distribute backup tasks to multiple computing nodes and storage resources for parallel processing, and maintains real-time communication with the storage nodes through a heartbeat detection mechanism to monitor backup progress and status.

[0137] The manner of distributing backup tasks to multiple computing nodes and storage resources for parallel processing by utilizing a distributed task scheduling framework is specifically:

[0138] A target function and constraint conditions related to the information span of the backup task are established.

[0139] Under the constraint of the constraint condition, a particle swarm optimization algorithm is adopted to determine an optimal scheduling scheme with the goal of minimizing the target function.

[0140] The backup tasks are distributed to multiple computing nodes and storage resources for parallel processing through the optimal scheduling scheme.

[0141] The target function is specifically:

[0142] f(x)=MinimizeS,M low ≤S≤M up

[0143] M low =max(d1,d2,d3)

[0144]

[0145] d1=max(D i )

[0146]

[0147] Wherein, f(x) indicates the target function, Minimize indicates minimization, S indicates the information span of the backup task, M low indicates the lower limit of the information span, M up indicates the upper limit of the information span, max indicates maximization, d1 indicates the longest duration required by the backup task with the longest execution time, d2 indicates the shortest total processing duration when all backup tasks are processed according to the maximum computing throughput, d3 indicates the shortest total processing duration when all backup tasks are processed according to the maximum number of backup tasks that can be processed synchronously by the server, D i indicates the duration of the i-th backup task, and T irepresents the computing requirement of the i-th backup task, i = 1, 2, …, n, n represents the total number of backup tasks, m represents the total number of servers, sumT represents the maximum computing throughput of the servers, and mnsDA represents the maximum number of backup tasks processed synchronously by each server.

[0148] The constraint condition specifically includes:

[0149]

[0150] R ijt ≥ P ij + Q it - 1

[0151] wherein, P ij represents whether the i-th backup task is assigned to the j-th server, i = 1, 2, …, n, n represents the total number of backup tasks, j = 1, 2, …, m, m represents the total number of servers, and Q it represents whether the i-th backup task is started at time t, R ijt′ represents whether the i-th backup task is processed by the j-th server at time t', t' represents the time at which the i-th backup task starts to be executed, and Q it′ represents whether the i-th backup task is started at time t'.

[0152] The method for determining the optimal scheduling scheme by using the particle swarm optimization algorithm specifically includes:

[0153] Initialize the particle swarm, set the inertia weight range, learning factor, maximum iteration number, population size, particle position, particle velocity, individual optimal value, and population global optimal value.

[0154] Calculate the nonlinear weight of each particle:

[0155]

[0156] wherein, ω(t) represents the inertia weight at time t, ω max represents the initial maximum value of the inertia weight, ω min represents the initial minimum value of the inertia weight, T represents the maximum iteration number, and t represents the current iteration number.

[0157] It should be noted that by dynamically adjusting the inertia weight, the global search and local search capabilities can be effectively balanced. In the early stage of iteration, the larger inertia weight helps the particles to perform more extensive global search and cover the entire solution space. In the later stage of iteration, the smaller inertia weight enhances the local search capability of the particles and accelerates the convergence to the optimal solution. The nonlinear decreasing mode is more flexible than the linear decreasing mode, and can improve the convergence efficiency and optimization accuracy of the algorithm in complex optimization problems.

[0158] According to the nonlinear weight, the speed and position of each particle are updated:

[0159] v i (t+1) = ω(t) · v i (t) + c1r1(P i -x i ) + c2r2(G - x i )

[0160] x i (t+1) = x i (t) + v i (t+1)

[0161] Wherein, v i (t+1) represents the speed of the i-th particle at the t+1 iteration, v i (t) represents the speed of the i-th particle at the t iteration, c1 represents the individual learning factor, r1 and r2 represent random numbers, P i represents the historical optimal position of the i-th particle, x i represents the position of the i-th particle at the t iteration, c2 represents the group learning factor, G represents the global optimal position, x i (t+1) represents the position of the i-th particle at the t+1 iteration.

[0162] According to the updated position and speed, the fitness value of each particle is calculated:

[0163] F(X) = wf(x)

[0164] Wherein, F(X) represents the fitness value, w represents the weight coefficient, and f(X) represents the objective function.

[0165] According to the fitness value, it is judged whether the particle position meets the constraint condition; if yes, the individual optimal value and the population global optimal value of the particle are updated, otherwise, the particle swarm is reinitialized.

[0166] It is judged whether the maximum iteration number is reached; if yes, the optimal position of the particle and the global optimal position of the population are output; otherwise, the speed and position of each particle are updated again.

[0167] The recovery process management submodule is selected according to the demand to restore the specific data at a specific time point in the past, and the specific geographic area data or user travel data of the specific data is restored, and the integrity and consistency of the restored specific data are checked.

[0168] It should be noted that by dynamically adjusting the backup priority, allocating backup tasks and processing backup tasks in parallel, the backup efficiency and resource utilization are significantly improved, at the same time, the recovery process management submodule can restore the data of a specific time point or region on demand, ensuring the integrity and consistency of data recovery, reducing the risk of data loss, and improving the reliability and flexibility of the system.

[0169] In a possible implementation, the data visualization and analysis assistance module specifically comprises:

[0170] The data visualization tool integration submodule is configured to provide visualization chart types according to the characteristics of travel data and the analysis requirements of users.

[0171] The data analysis algorithm library embedding submodule is configured to embed commonly used data analysis algorithms, including clustering analysis algorithms and path planning algorithms; the clustering analysis algorithms classify user groups with similar travel behaviors, and the path planning algorithms provide optimal travel route suggestions for users according to historical travel data of the users and real-time traffic information.

[0172] In the present application, when the clustering analysis algorithms classify user groups with similar travel behaviors, multiple travel behavior characteristics are considered, such as travel frequency, travel time regularity, travel distance, travel destination, etc. When the path planning algorithms use historical travel data of users and real-time traffic information, more traffic factors are considered, such as traffic signal light time, traffic control information, etc. The algorithms plan paths according to the travel time of users, traffic rules and characteristics of different traffic modes (such as bus stop sites and timetables).

[0173] The analysis report generation submodule is configured to automatically generate a standardized analysis report template according to the analysis tasks of users and the visualization results, including data overview, key indicator analysis, visualization chart display, data insight conclusion and suggestion measures based on the analysis results, and output the report in the form of PDF and HTML.

[0174] It should be noted that by integrating data visualization tools and commonly used analysis algorithms, intuitive travel data display and in-depth analysis are provided, which can effectively identify user travel patterns and optimize route planning, automatically generate standardized analysis reports, facilitate users to quickly obtain key data insights and suggestions, improve decision-making efficiency and user experience, and at the same time ensure the convenience and systematization of the data analysis process.

[0175] A user travel data security management method based on cloud computing is applied to the user travel data security management system based on cloud computing, and the method comprises the following steps:

[0176] S1: Real-time monitoring of user travel state changes at each access device, when detecting user travel-related behavior, activating the data collection process of the corresponding device, accurately collecting user travel data according to the adaptation strategy, and performing preliminary format standardization and error checking on user travel data.

[0177] It should be noted that by real-time monitoring of user travel state changes, the data collection process is activated in time when the user starts traveling, improving the accuracy and timeliness of data collection. Combined with the adaptation strategy and error checking, the quality and consistency of the collected data are ensured, reducing data omission and errors, and optimizing the subsequent data processing process.

[0178] S2: Obtain symmetric key using quantum key distribution mechanism, use AES algorithm to transmit, segment and parallel encrypt the checked user travel data according to the symmetric key, obtain multiple encrypted data packets, attach message authentication code to each encrypted data packet, and complete encryption and packaging. After high-speed transmission of encrypted data packets to the cloud computing platform through a dedicated network channel.

[0179] It should be noted that combined with quantum key distribution mechanism and AES encryption algorithm, strong security is provided for data transmission. Quantum key distribution ensures the unbreakability of key transmission, while AES encryption and message authentication code enhance the confidentiality and integrity of data. Through a dedicated network channel, data is transmitted efficiently and securely, effectively preventing data leakage and tampering risks.

[0180] S3: After the cloud computing platform receives the encrypted data packets, the privacy screening and desensitization module scans each encrypted data packet row by row through a deep neural network model, identifies the privacy sensitive content of each encrypted data packet based on the training and learning results, determines the privacy data, and performs real-time desensitization processing on the privacy data according to the desensitization rules.

[0181] It should be noted that through the deep neural network model, privacy sensitive content is automatically identified and processed, ensuring that data meets privacy protection requirements during transmission and storage. Real-time desensitization processing ensures user privacy security while preserving data analysis value, improving data compliance and security, and reducing privacy leakage risks.

[0182] S4: The distributed storage architecture module receives desensitized privacy data, parallel cuts the desensitized privacy data into multiple data segments according to the data sharding strategy, uniformly stores them in storage nodes in each geographic region, and synchronously generates redundant copies to ensure data security. Based on high-frequency access requirements, use edge cache nodes to intelligently cache popular data segments related to high-frequency access requirements.

[0183] It should be noted that by cutting and storing data in parallel to multiple storage nodes in multiple geographic regions, the storage efficiency and security of the data are improved, the generation of redundant copies ensures the availability of data in the event of any node failure, and the edge cache node optimizes the data read speed of high-frequency access, improves the response speed and overall performance of the system, and ensures the high availability and fast access of data.

[0184] S5: The data backup and recovery scheduling module formulates multiple backup strategies according to the importance, update frequency and regulatory requirements of the door data segments, and dynamically adjusts the priority of the backup tasks; using a distributed task scheduling framework, the backup tasks are distributed to multiple computing nodes and storage resources for parallel processing; when data recovery is needed, the specific data at a specific time point is selected for recovery, and the specific geographic area data or user travel data in the specific data is recovered.

[0185] It should be noted that by dynamically adjusting the backup priority according to the importance and demand of the data, the key data is prioritized for backup, reducing the risk of data loss, and using distributed task scheduling to achieve efficient parallel processing of backup tasks, improving the backup efficiency and flexibility of the system, in addition, supporting on-demand recovery of data at a specific time point or region, ensuring the accuracy and pertinence of data recovery.

[0186] S6: The data visualization and analysis assistance module selects appropriate visualization chart types according to the characteristics of user travel data and user analysis needs; starts the built-in data analysis algorithm to run on the desensitized data, and mines the potential value of the data on the premise of protecting user privacy; according to the user's analysis task and visualization result, a standardized analysis report template is automatically generated.

[0187] It should be noted that by selecting appropriate visualization charts according to the characteristics of travel data and user needs, combining common data analysis algorithms, ensuring the protection of user privacy while deeply mining data value, automatically generating standardized reports, improving the efficiency and accuracy of data analysis, and facilitating users to intuitively understand data insights and provide support for decision-making.

[0188] S7: When an external subject initiates a data call request, the permission control and audit module receives the request information, and the smart contract strictly checks the identity, permission range and access purpose of the request party based on the permission ledger of the blockchain, if the verification is passed, the corresponding data is retrieved and extracted from the storage node, and after decryption, it is provided to the request party; if the verification fails, the request is rejected and detailed audit logs are recorded for subsequent trace analysis.

[0189] It should be noted that the permission control is realized through the blockchain technology, the security and transparency of data access are ensured, the identity, permission and access purpose of the requester are automatically verified by the smart contract, unauthorized access is avoided, the audit log is recorded for traceability, the protection and traceability of data are enhanced, compliance is ensured and data abuse is effectively prevented.

[0190] The technical scheme provided by the embodiment of the application brings at least the following beneficial effects:

[0191] In the embodiment of the application, the mixed encryption transmission channel module is distributed through quantum key distribution and AES encryption, ensuring the confidentiality, integrity and availability of data transmission, effectively preventing hacker attacks, data theft and tampering, the privacy screening and desensitization module accurately identifies privacy sensitive content using a deep learning model and desensitization rules, protects personal identity information and high-precision geographic location, while retaining the analytical value of the data, the distributed storage architecture module uses multi-dimensional sharding and redundant backup technology to ensure high availability, security and efficient access of data, the permission control and audit module establishes an unalterable permission ledger through blockchain technology and monitors data abuse risks to ensure the transparency and traceability of data access, the data backup and recovery scheduling module ensures the secure backup of critical data and supports efficient and flexible data recovery, ensuring the integrity and consistency of data, the data visualization and analysis assistance module provides rich charts and custom layouts, supports deep data analysis and user classification, and improves data value mining and user experience.

[0192] It should be understood that the processor in the embodiment of the application can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0193] It should also be understood that the memory in the embodiments of the present application can be volatile or nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM) used as external cache. By way of example, and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0194] The above-described embodiments can be implemented in part or in whole through software, hardware (e.g., circuitry), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When loaded and executed by a computer, the computer instructions or computer programs can generate the flow or function according to the embodiments of the present application in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, such as from one website site, computer, server, or data center to another website site, computer, server, or data center through a wired (e.g., infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing a set of one or more available media. The available media can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.

[0195] It should be understood that the term "and / or" used herein is merely an association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but it can also represent an "and / or" relationship, which can be understood in the context before and after it.

[0196] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0197] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined by their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0198] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0199] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0200] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0201] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0202] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.

[0203] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0204] The embodiment of the present application provides a computer readable storage medium, which stores a computer program. The program is executed by a processor to realize the cloud computing based user travel data security management method according to the method embodiment.

[0205] The computer readable storage medium provided by the present application can realize the steps and effects of the cloud computing based user travel data security management method according to the method embodiment. To avoid repetition, the present application will not be described again.

[0206] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0207] The following points need to be explained:

[0208] (1) The drawings of the embodiments of the present application only relate to the structures involved in the embodiments of the present application, and other structures can be referred to the general design.

[0209] (2) For the sake of clarity, the thickness of the layers or regions is exaggerated or reduced in the drawings used to describe the embodiments of the present application, that is, the drawings are not drawn according to the actual proportion. It can be understood that when an element such as a layer, a film, a region or a substrate is referred to as being located "on" or "under" another element, the element can be "directly" located on or under another element or there can be an intermediate element.

[0210] (3) In the case of no conflict, the embodiments of the present application and the features in the embodiments can be combined with each other to obtain new embodiments.

[0211] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1.A cloud computing-based user travel data security management system, characterized by, The application relates to a multi-source data collection and privacy protection method and system. The application comprises the following modules: A multi-source data collection adaptation module for collecting user travel data from various types of travel equipment; A hybrid encryption transmission channel module for establishing a transmission link between a data collection terminal and a cloud computing platform; A privacy identification and desensitization module for identifying and processing privacy-sensitive content in data received by the cloud computing platform; A distributed storage architecture module for storing user travel data on the cloud computing platform; A permission control and audit module for controlling the calling permission of external subjects to user travel data and auditing data access; A data backup and recovery scheduling module for recovering data when there is a data risk through a backup strategy, a backup task execution, a monitoring process and a recovery process; The data backup and recovery scheduling module specifically comprises: A backup strategy formulation sub-module for formulating different backup strategies according to the importance, update frequency and regulatory requirements of data, and dynamically adjusting the priority of backup according to the storage time and access heat of data; The way of dynamically adjusting the priority of backup according to the storage time and access heat of data specifically comprises: Determining the priority of backup: CP(x) = W1*UF(x) + W2*k(x) + W3*PC(x) + W4*evi(x) + W5*hotness(x) Where CP() represents the priority of backup, W1, W2, W3, W4 and W5 all represent weights, UF() represents the update frequency of data, k() represents the criticality of data, PC() represents the access priority of data, evi() represents the validity period of data, and hotness() represents the heat of data; CP(x)' = (a - CP(x) - e -λt ) + (1 - a) - H Dynamically adjusting the priority of backup according to the storage time and the access heat: Where CP() represents the dynamically adjusted priority, alpha represents a weight coefficient, lambda represents a decay coefficient, t represents storage time, e represents an index, and H represents the number of data access times; A backup task execution and monitoring sub-module for distributing backup tasks to multiple computing nodes and storage resources for parallel processing by using a distributed task scheduling framework, maintaining real-time communication with storage nodes through a heartbeat detection mechanism, and monitoring backup progress and state; The way of distributing backup tasks to multiple computing nodes and storage resources for parallel processing by using a distributed task scheduling framework specifically comprises: Establishing a target function and constraint condition about the span of backup task information; Under the constraint of the constraint condition, an optimal scheduling scheme is determined by using a particle swarm optimization algorithm to minimize the target function; Backup tasks are distributed to multiple computing nodes and storage resources for parallel processing through the optimal scheduling scheme; f(x) = Minimize S, M low ≤ S ≤ M up M low = max(d1, d2, d3) d1 = max(D i ) wherein f(x) represents an objective function, Minimize represents minimization, S represents an information span of backup tasks, M low represents a lower limit of the information span, M up represents an upper limit of the information span, max represents maximization, d1 represents a duration required for a backup task with the longest execution time, d2 represents a shortest total processing duration when all backup tasks are processed at a maximum computing throughput, d3 represents a shortest total processing duration when all backup tasks are processed at a maximum number of backup tasks that can be synchronously processed by a server, D i represents a duration of an i-th backup task, T i represents a computing demand of an i-th backup task, i = 1, 2, …, n, n represents a total number of backup tasks, m represents a total number of servers, sumT represents a maximum computing throughput of a server, and mnsDA represents a maximum number of backup tasks that are synchronously processed by each server. The target function is specifically: wherein P ij represents whether the i-th backup task is assigned to the j-th server, i = 1, 2,..., n, n representing the total number of backup tasks, j = 1, 2,..., m, m representing the total number of servers, Q it represents whether the i-th backup task is started at time t, R ijt′ represents whether the i-th backup task is processed by the j-th server at time t', t' representing the time at which the i-th backup task starts to be executed, Q it′ represents whether the i-th backup task is started at time t'; The constraint condition specifically comprises: A recovery process management sub-module for selecting specific data recovered to a specific time point according to requirements, and recovering specific geographic area data or user travel data of the specific data, and performing integrity and consistency verification on the recovered specific data. The data visualization and analysis auxiliary module is used for presenting user travel data in a visualized manner, performing deep analysis on the user travel data in combination with a data analysis algorithm, and arranging analysis results into a report. 2.The cloud computing based user travel data security management system according to claim 1, wherein, The multi-source data collection adaptation module comprises: A device interface adaptation sub-module is configured to configure exclusive interface drivers for smartphones, vehicle-mounted intelligent terminals, and wearable devices; A collection frequency optimization sub-module is configured to dynamically adjust the collection frequency according to the device power, network status, and user travel state; A data preprocessing sub-module is configured to perform noise removal and time synchronization calibration on the collected data at the device end, encapsulate the data into uniform data frames in a preset format, and remove GPS positioning data noise by using a Kalman filtering algorithm. 3.The cloud computing based user travel data security management system according to claim 1, wherein, The mixed encryption transmission channel module comprises: A quantum key distribution sub-module is configured to deploy quantum key distribution terminal devices in data collection intensive areas and cloud computing data centers, and is matched with an operation and maintenance management system, which is responsible for the initialization configuration, key generation parameter adjustment, key storage, and key update management of the QKD devices; An AES encryption execution sub-module is configured to build a parallel encryption processing architecture, optimize the round function calculation process of the AES algorithm, perform transmission segmentation and parallel encryption on large data blocks, restore the data after completing the verification of the cloud computing user verifiable static data storage integrity, and dynamically schedule cloud computing resources according to real-time encryption task quantity; An authentication code processing sub-module is configured to calculate and generate a 256-bit authentication tag for each data packet at the data sending end by using the HMAC-SHA256 algorithm, record the related key information and time stamp of the authentication tag generation, and verify the authentication tag at the receiving end by using the same key and algorithm. If the authentication tags are inconsistent, a data retransmission mechanism is triggered immediately, and abnormal information is fed back to the security audit center. 4.The cloud computing based user travel data security management system according to claim 3, characterized in that, The transmission segmentation and parallel encryption on large data blocks, and the data restoration after completing the verification of the cloud computing user verifiable static data storage integrity, specifically comprise: According to the internal logical structure or predetermined rules of the data, the large data block is segmented into a plurality of fixed-size data sub-blocks, and the AES algorithm is used to allocate independent encryption keys to each data sub-block, wherein the encryption keys are dynamically generated and distributed by a key management system. After the data is stored on the cloud computing platform, a verifiable storage integrity verification mechanism is adopted, and the message authentication code HMAC technology based on the hash function is used to calculate the hash value of each data sub-block during data storage, and the hash value is stored together with the data sub-block. When data needs to be restored, the hash value of the stored data sub-block is recalculated and compared with the previously stored hash value. If the hash values are consistent, the segmented data sub-blocks are reassembled in the original order by a data restoration program to restore the complete data file. If the hash values are inconsistent, a data repair or alarm mechanism is triggered to notify the administrator that the data has been damaged or tampered with. 5.The cloud computing based user travel data security management system according to claim 1, wherein, The privacy discrimination and desensitization module comprises: The data sample management submodule is configured to collect user travel data samples in different regions, cultures, and industry backgrounds, and construct a sample database management system based on the user travel data samples, so as to classify, store, version manage, and regularly update the user travel data samples in multiple dimensions according to travel modes, travel scenarios, and user groups. The learning model construction submodule is configured to select a long short-term memory network combined with a convolutional neural network to construct a multi-layer hybrid neural network model, use the CNN to extract spatial features of the user travel data, determine a feature map, and input the feature map into the LSTM network to learn time sequence features to determine the privacy-sensitive content. The method for determining the privacy-sensitive content by using the CNN to extract the spatial features of the user travel data, determining the feature map, and inputting the feature map into the LSTM network to learn the time sequence features to determine the privacy-sensitive content specifically includes the following steps. The method for determining the privacy-sensitive content by using the CNN to extract the spatial features of the user travel data, determining the feature map, and inputting the feature map into the LSTM network to learn the time sequence features to determine the privacy-sensitive content specifically includes the following steps. wherein a i,j represents the activation value at the feature map (i, j) position after convolution operation, f represents an activation function, w m,n represents the corresponding weight at the convolution kernel position (m, n), x i+m,j+n represents the value at the position (i+m, j+n) in the input data, m=1, 2, …, M, M represents the total number of rows of the convolution kernel, n=1, 2, …, N, N represents the total number of columns of the convolution kernel, and b represents a bias. The method for determining the privacy-sensitive content by using the CNN to extract the spatial features of the user travel data, determining the feature map, and inputting the feature map into the LSTM network to learn the time sequence features to determine the privacy-sensitive content specifically includes the following steps. x t = flatten(a i,j ) wherein x t represents the flattened feature map, and flatten represents the flattening operation. The method for determining the privacy-sensitive content by using the CNN to extract the spatial features of the user travel data, determining the feature map, and inputting the feature map into the LSTM network to learn the time sequence features to determine the privacy-sensitive content specifically includes the following steps. i t = σ(w xi x t + w hi h t-1 + b i ) f t = σ(w xf x t + w hf h t-1 + b f ) o t = σ(w xo x t + w ho h t-1 + b o ) h t = o t tanh(c t ) where i t denotes the input gate, f t denotes the forget gate, o t denotes the output gate, denotes the candidate memory state at time t, c t denotes the decision memory cell state at time t, σ denotes the Sigmoid function, w xi denotes the input data x t to the input gate, w hi denotes the hidden state h t-1 to the input gate, b i denotes the bias vector for the input gate, w xf denotes the input data x t to the forget gate, w hf denotes the hidden state h t-1 to the forget gate, b f denotes the bias vector for the forget gate, w xo denotes the input data x t to the output gate, w ho denotes the hidden state h t-1 to the output gate, b o denotes the bias vector for the output gate, w xc denotes the input data x t to the candidate memory state, w hc denotes the hidden state h t-1 to the candidate memory state, b c denotes the bias vector for the candidate memory state, tanh denotes the tanh activation function; When the prediction score is greater than a preset prediction score, the privacy-sensitive content of the user travel data is determined. y = W f • h t + b s where y denotes the predicted score of the privacy sensitive content, W f denotes the weight matrix of the fully connected layer, b s denotes the bias of the fully connected layer; The desensitization rule execution submodule is configured to formulate a desensitization rule library according to different privacy levels and data types, and embed the desensitization rule library into a model output end, so as to use regional fuzzification and coordinate offset for high-precision geographic location information, anonymize personal identity information associated travel data, and fuzzify specific place names for sensitive place entry and exit records. The fuzzification specifically includes the following steps: dividing a regional map into geographic grids of uniform size according to certain rules, determining a grid unit to which a high-precision geographic location coordinate point belongs when the high-precision geographic location coordinate point is obtained, and fuzzily describing the location information of a user as a grid area in which the user is located; and adding randomly generated direction and position noise values to perform coordinate offset for high-precision longitude and latitude coordinates. The distributed storage architecture module includes the following modules. 6.The cloud computing based user travel data security management system according to claim 1, wherein, The data sharding management submodule is configured to design a composite sharding function in multiple dimensions based on user IDs, travel timestamps, and geographic region codes, uniformly shard user travel data by using a consistent hashing algorithm, establish a sharding index management system, and track and record sharding storage locations and replica distribution information. The redundancy backup control submodule is configured to use Reed-Solomon error correction code technology to develop a redundancy backup scheduling program for distributed storage of redundant replicas, regularly perform consistency checking on the redundant replicas, and use the Reed-Solomon error correction code technology to recover error replicas. The distributed cache scheduling submodule is configured to select an open-source distributed cache system, configure a cache server on an edge node close to a user request end, develop an intelligent cache scheduling algorithm, dynamically manage cache space by using the intelligent cache scheduling algorithm, and combine a data pre-fetching technology to cache data that is likely to be accessed in advance. The permission control and audit module specifically includes the following modules. 7.The cloud computing based user travel data security management system according to claim 1, wherein, ​ The blockchain platform building submodule is configured to select a blockchain open source framework to build an enterprise-level blockchain platform, and is responsible for deployment of a blockchain node, network configuration, and generation of a genesis block, creation of an independent identity certificate for each data access subject, and registration of detailed permission information on the blockchain; The permission verification execution submodule is configured to, in the case of a data call request, automatically obtain identity information and permission information of a requestor from the blockchain through an intelligent contract, compare the identity information and the permission information with request details, cover key indicators of an accessible data time period, a geographic area, and a data type, and if the comparison is matched, send a search instruction to the storage module, and if the comparison is not matched, trigger an alarm and record request information to an audit log area; The audit backtracking analysis submodule is configured to develop an audit data analysis system through a blockchain browser tool, and an administrator can view data access history records, mine data abuse through a data mining algorithm, and generate an audit report. 8.The cloud computing based user travel data security management system according to claim 1, wherein, The data visualization and analysis auxiliary module specifically includes: A data visualization tool integration submodule is configured to provide a visualization chart type according to characteristics of user travel data and analysis requirements of a user; A data analysis algorithm library embedding submodule is configured to embed common data analysis algorithms, including a clustering analysis algorithm and a path planning algorithm; the clustering analysis algorithm classifies user groups with similar travel behaviors, and the path planning algorithm can provide an optimal travel route suggestion for a user according to historical travel data of the user and real-time traffic information; An analysis report generation submodule is configured to automatically generate a standardized analysis report template according to an analysis task of a user and a visualization result, including a data overview, key indicator analysis, visualization chart display, data insight conclusions, and suggested measures based on analysis results, and output the report in the form of PDF and HTML. 9.A cloud computing-based user travel data security management method applied to the cloud computing-based user travel data security management system of any one of claims 1-8, characterized in that, The method comprises the following steps: S1: Real-time monitoring of user travel state changes is performed at each access device end, and when a user performs a travel-related behavior, a data collection process of the corresponding device is activated, user travel data is accurately collected according to an adaptive strategy, and the user travel data is subjected to preliminary format standardization and error checking; S2: A symmetric key is obtained by using a quantum key distribution mechanism, and the user travel data that passes the checking is subjected to transmission segmentation and parallel encryption by using an AES algorithm according to the symmetric key, to obtain a plurality of encrypted data groups, a message authentication code is attached to each of the encrypted data groups, and after the encrypted data groups are packaged, the encrypted data groups are transmitted at a high speed to a cloud computing platform through a special network channel; S3: After the cloud computing platform receives the encrypted data groups, a privacy identification and desensitization module scans each of the encrypted data groups row by row by using a deep neural network model, identifies privacy sensitive content of each of the encrypted data groups according to training and learning results, determines privacy data, and performs real-time desensitization processing on the privacy data according to a desensitization rule; S4: The distributed storage architecture module receives the desensitized privacy data, cuts the desensitized privacy data into multiple data segments in parallel according to a data sharding strategy, uniformly stores the data segments in storage nodes in each geographic region, synchronously generates redundant copies to ensure data security, uses edge cache nodes to intelligently cache popular data segments related to high-frequency access demands based on high-frequency access demands; S5: The data backup and recovery scheduling module formulates multiple backup strategies according to the importance, update frequency and regulatory requirements of the popular data segments, and dynamically adjusts the priority of the backup tasks; The backup tasks are allocated to multiple computing nodes and storage resources for parallel processing using a distributed task scheduling framework; when data recovery is needed, the specific data at a specific time point is selected for recovery according to the requirements, and the specific geographic area data or user travel data in the specific data is recovered; S6: The data visualization and analysis assistance module selects appropriate visualization chart types according to the characteristics of the user travel data and the analysis requirements of the user; Start the built-in data analysis algorithm to run on the desensitized data, and mine the potential value of the data on the premise of protecting the privacy of the user; generate a standardized analysis report template automatically according to the analysis task and the visualization result of the user; S7: When an external subject initiates a data call request, the permission control and audit module receives the request information, and the smart contract strictly checks the identity, permission range and access purpose of the requestor based on the permission ledger of the blockchain; if the verification is passed, the corresponding data is retrieved and extracted from the storage node, and after decryption, it is provided to the requestor; If the verification fails, the request is rejected and detailed audit logs are recorded for subsequent trace analysis.

Citation Information

Patent Citations

  • Highway section-level data middle station system

    CN112687097A

  • Data security and privacy protection method and system

    CN119646838A