Rail transit data anti-leakage processing method and device based on matrix transformation

By converting rail transit data into matrix form and using reversible random matrix for obfuscation, combining quantum entangled states and blockchain technology, the risk of leakage of rail transit data in complex environments is solved, and the data is safely stored and legally used.

CN120509054AActive Publication Date: 2025-08-19CHENGDU BITTRUST TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510622439.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-19
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

When facing complex data environments, existing rail transit data protection methods have problems such as risk of leakage during data decryption, difficulty in preventing leakage caused by internal personnel or system vulnerabilities, and difficulty in dealing with structured and unstructured data in a unified manner.

Method used

Transit data is converted into matrix data, a reversible random matrix is ​​generated for obfuscation, and data is restored by evaluating the privacy protection intensity and meeting preset standards. It is used to encrypt and store it in combination with quantum entangled states and blockchain technology.

Benefits of technology

It significantly reduces the risk of data leakage during storage and transmission, ensures the integrity and availability of data when used legally, and achieves a balance between privacy protection and data availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509054A_ABST
    Figure CN120509054A_ABST
Patent Text Reader

Abstract

The invention provides a rail transit data anti-leakage processing method and device based on matrix transformation, and relates to the technical field of rail transit data security, and the method comprises the steps: converting rail transit data into matrix data; each row of the matrix data represents a data record, and each column represents a data attribute; generating a reversible random matrix according to the column number of the matrix data; the row number and the column number of the random matrix are matched with the column number of the matrix data; multiplying the matrix data by a random matrix to obtain confusion matrix data; evaluating the privacy protection intensity of the confusion matrix data until the privacy protection intensity meets a preset standard; when a use instruction of the rail transit data is received, the inverse matrix of the random matrix is determined, and the inverse matrix is multiplied by the confusion matrix data to obtain the restored rail transit data, so that the leakage risk of the data in the storage and transmission process is remarkably reduced, the integrity of the data in legal use is ensured, and the use efficiency of the rail transit data is improved. And the balance between privacy protection and data availability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rail transit data security technology, and in particular to a rail transit data anti-leakage processing method and device based on matrix transformation. Background Art

[0002] With the rapid development of the rail transit industry, a large amount of rail transit data is generated and transmitted during system operation, maintenance, and management. This data often contains sensitive information, such as passenger personal information, train operation status, and dispatch instructions. Leakage of this data can lead to serious consequences, including the loss of passenger privacy, increased operational safety risks, and damage to corporate reputation.

[0003] Existing data protection methods primarily focus on data encryption and access control, but these methods may present the following problems when faced with the complex rail transit data environment. For example, Problem 1: While traditional encryption methods can encrypt data, they require decryption during data use, which can lead to data leakage risks during transmission and processing after decryption. Problem 2: Relying solely on access control cannot completely prevent data leakage caused by internal personnel or system vulnerabilities. Problem 3: Rail transit data is a combination of structured and unstructured data, making it difficult for traditional protection methods to uniformly handle different types of data. Summary of the Invention

[0004] The present invention provides a method and device for anti-leakage processing of rail transit data based on matrix transformation, which are used to solve the technical problem that rail transit data cannot be efficiently protected in the prior art.

[0005] In one aspect, the present invention provides a method for preventing rail transit data from being leaked based on matrix transformation, comprising: Converting rail transit data into matrix data; wherein each row of the matrix data represents a data record and each column represents a data attribute; Generate a reversible random matrix according to the number of columns of the matrix data; wherein the number of rows and columns of the random matrix matches the number of columns of the matrix data; Multiplying the matrix data by the random matrix to obtain confusion matrix data; Evaluating the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard; When an instruction to use the rail transit data is received, the inverse matrix of the random matrix is determined, and the inverse matrix is multiplied by the confusion matrix data to obtain restored rail transit data.

[0006] According to the present invention, a method for preventing rail transit data from being leaked based on matrix transformation is provided, which converts rail transit data into matrix data, including: Determine the data type of each data attribute of the rail transit data, and map each data attribute to a column of the matrix data using a mapping rule corresponding to the data type; Arrange each data record according to its relevance to form a row of matrix data; The step of determining the data type of each data attribute of the rail transit data and mapping each data attribute to a column of matrix data using a mapping rule corresponding to the data type includes: When the data attribute is numerical data, its value is directly used as the matrix element; When the data attribute is categorical data, one-hot encoding is used to convert it into a numerical type as a matrix element; The association includes a temporal order or a spatial order.

[0007] According to the present invention, a method for preventing rail transit data from being leaked based on matrix transformation is provided, which generates a reversible random matrix according to the number of columns of the matrix data, including: Determine the dimensions of the random matrix so that the number of its rows and columns matches the number of columns of the matrix data; determining a distribution characteristic of rail transit data, and generating elements of a random matrix using a random generation algorithm corresponding to the distribution characteristic; Check whether the random matrix is invertible by calculating the determinant value or matrix rank. If it is not invertible, regenerate it until a reversible random matrix is obtained; The determining of the distribution characteristics of the rail transit data and the use of a random generation algorithm corresponding to the distribution characteristics to generate elements of a random matrix include: When the distribution characteristics of rail transit data belong to normal distribution, the normal distribution method is used to generate the element values of the random matrix; When the distribution characteristics of rail transit data belong to uniform distribution, the uniform distribution method is used to generate the element values of the random matrix.

[0008] According to a rail transit data anti-leakage processing method based on matrix transformation provided by the present invention, the matrix data is multiplied by the random matrix to obtain confusion matrix data, including: Determining a matrix scale according to the matrix data and the number of rows and columns of the random matrix; When the matrix size is smaller than the minimum value of the preset size range, directly multiplying the matrix data by the random matrix; When the matrix size falls within the preset size range, multiplying the matrix data by the random matrix using the Strassen algorithm or the Winograd algorithm; When the matrix size is greater than the maximum value of the preset size range, dividing the matrix data and the random matrix into a plurality of sub-matrix blocks respectively; Perform matrix multiplication operations on each sub-matrix block to obtain the corresponding sub-result block; All sub-result blocks are spliced according to their positions in the original matrix to form a complete confusion matrix data.

[0009] According to a rail transit data anti-leakage processing method based on matrix transformation provided by the present invention, the privacy protection strength of the confusion matrix data is evaluated until the privacy protection strength meets a preset standard, including: Determine the evaluation indicators of privacy protection strength based on the characteristics of rail transit data and privacy protection requirements; Calculate the privacy protection strength value of the confusion matrix data based on the evaluation index; Compare the privacy protection strength value with the preset standard; If the privacy protection strength value does not meet the preset standard, the random matrix is regenerated, the complexity of the random matrix is increased, or noise is added to the confusion matrix data to enhance the privacy protection strength value until the privacy protection strength meets the preset standard.

[0010] According to the present invention, a method for preventing leakage of rail transit data based on matrix transformation is provided, which determines an evaluation index of privacy protection strength according to the characteristics of rail transit data and privacy protection requirements, including: Evaluate the sensitivity of each data attribute in rail transit data and determine the corresponding evaluation index weight according to the sensitivity level; For rail transit data where the proportion of categorical data or numerical data exceeds a preset ratio, information entropy is selected as the evaluation indicator. The probability distribution of each element value in the confusion matrix data is calculated to quantify the privacy protection strength. The higher the information entropy, the stronger the privacy protection strength. When it is necessary to evaluate the correlation between confusion matrix data and matrix data, mutual information is selected as the evaluation indicator; the lower the mutual information, the higher the privacy protection strength; For rail transit data that needs to meet the requirements of differential privacy protection, the privacy budget value of differential privacy is selected as the evaluation indicator; among them, the smaller the privacy budget value, the higher the privacy protection strength.

[0011] According to a rail transit data anti-leakage processing method based on matrix transformation provided by the present invention, after evaluating the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard, the method further includes: Slice the confusion matrix data according to preset rules to generate multiple data segments; Store each data fragment in a different storage node or storage medium to ensure isolation between data fragments; When receiving a rail transit data usage instruction, it also includes: Collect all data fragments and perform restoration processing.

[0012] According to a method for preventing leakage of rail transit data based on matrix transformation provided by the present invention, before converting the rail transit data into matrix data, the method further includes: Distinguish between structured and unstructured data in rail transit data; combining the structured data with the unstructured data; The combination method includes at least one of the following: Feature concatenation, weighted averaging, or feature extraction and fusion of deep learning models.

[0013] According to the present invention, a method for preventing rail transit data from being leaked based on matrix transformation further comprises, after generating a reversible random matrix: Using a quantum entangled state generator to inject quantum entangled state information into each element of the random matrix; wherein the quantum entangled state information is represented by a quantum bit, and the state of each quantum bit is determined by the element value of the random matrix; Before multiplying the matrix data with the random matrix, the method further includes: Quantum state verification is performed on the random matrix injected with quantum entangled state to ensure that the quantum state information of each element is complete and has not been tampered with.

[0014] On the other hand, the present invention also provides a rail transit data anti-leakage processing device based on matrix transformation, comprising: A matrix conversion module, configured to convert rail transit data into matrix data, wherein each row of the matrix data represents a data record and each column represents a data attribute; A random matrix module, configured to generate a reversible random matrix according to the number of columns of the matrix data; wherein the number of rows and columns of the random matrix matches the number of columns of the matrix data; a matrix mixing module, configured to multiply the matrix data by the random matrix to obtain confusion matrix data; A matrix evaluation module, configured to evaluate the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard; The matrix restoration module is used to determine the inverse matrix of the random matrix when receiving an instruction to use the rail transit data, and multiply the inverse matrix by the confusion matrix data to obtain restored rail transit data.

[0015] The matrix-transformation-based rail transit data anti-leakage processing method and device provided by this invention effectively conceals the characteristics of the original data by converting it into a matrix form and obfuscating it using a reversible random matrix. This significantly reduces the risk of data leakage during storage and transmission. Furthermore, data restoration using the inverse of the random matrix ensures the integrity and availability of the data for legitimate use, achieving a balance between privacy protection and data availability, and providing an efficient and reliable method for the secure processing of rail transit data. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 1 is a flow chart of a method for preventing rail transit data from being leaked based on matrix transformation provided by an embodiment of the present invention; Figure 2 1 is a schematic structural diagram of a rail transit data anti-leakage processing device based on matrix transformation provided by an embodiment of the present invention; Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0019] Figure 1 The figure is a flow chart of a method for preventing rail transit data leakage based on matrix transformation provided by an embodiment of the present invention. The method can be executed by a computer, a tablet computer, a smart wearable device, etc.

[0020] See also Figure 1 The rail transit data anti-leakage processing method based on matrix transformation may include the following steps 101 to 105.

[0021] Step 101: Convert rail transit data into matrix data; wherein each row of the matrix data represents a data record, and each column represents a data attribute.

[0022] In this step, rail transit data may include passenger personal information, train operation status data, dispatch instruction data, ticketing system data, equipment operation data, and operational management data. Passenger personal information may include name, ID number, contact information, ticketing information, and payment information. Train operation status data may include train location information, train speed, train status, car status, and train timetable. Dispatching instruction data may include dispatching orders and dispatch plans. Ticketing system data may include ticket sales data, ticket inspection data, and discount information. Equipment operation data may include track equipment status, signaling system data, vehicle equipment status, and station equipment status. Operation management data may include passenger flow data, operating revenue data, operating cost data, and safety inspection data.

[0023] For example, a data record is Zhang San’s ticketing information, and the data attributes can be name (Zhang San), ID number, train number, seat number, etc.; a data record is the passenger flow statistics of the station, and the data attributes can be the number of people entering the station, the number of people leaving the station, the statistical time, etc.

[0024] Step 101 may specifically include: Determine the data type of each data attribute of the rail transit data, and map each data attribute to a column of the matrix data using a mapping rule corresponding to the data type; Arrange each data record according to its relevance to form a row of matrix data; The data type of each data attribute of the rail transit data is determined, and each data attribute is mapped to a column of the matrix data using a mapping rule corresponding to the data type, including: When the data attribute is numerical data, its value is directly used as the matrix element; When the data attribute is categorical data, one-hot encoding is used to convert it into a numerical type as a matrix element; Relevance includes temporal or spatial order.

[0025] Time sequence: Data records are arranged in chronological order. For example, train operation status data can be arranged in chronological order according to the train's operation time; passenger ticket information can be arranged in chronological order according to the ticket purchase time.

[0026] Spatial order: Data records are arranged in the order of spatial location. For example, passenger flow statistics for a station can be arranged in the order of the station's geographical location; seat information for a train can be arranged in the order of the seat's physical location.

[0027] This step defines the specific steps for converting rail transit data into matrix data, including how to process different data types (direct conversion for numerical data and one-hot encoding for categorical data) and how to arrange data records (in temporal or spatial order). This processing method not only unifies complex data structures into a matrix form, facilitating subsequent matrix transformation operations, but also preserves the original characteristics and relevance of the data, providing a solid foundation for subsequent privacy protection processing. It also enhances the applicability and flexibility of the entire method, making it adaptable to the diverse data types and application scenarios in the rail transit field.

[0028] Step 102: Generate a reversible random matrix according to the number of columns of the matrix data; wherein the number of rows and columns of the random matrix matches the number of columns of the matrix data.

[0029] Step 102 may specifically include: Determine the dimensions of the random matrix so that its number of rows and columns matches the number of columns of the matrix data; Determine the distribution characteristics of rail transit data, and use a random generation algorithm corresponding to the distribution characteristics to generate elements of a random matrix; Check whether the random matrix is invertible by calculating the determinant value (for example, non-zero means it is invertible) or the matrix rank (for example, equal to the number of rows or columns means it is invertible). If it is not invertible, regenerate it until a reversible random matrix is obtained; The distribution characteristics of rail transit data are determined, and random generation algorithms corresponding to the distribution characteristics are used to generate elements of a random matrix, including: When the distribution characteristics of rail transit data belong to normal distribution, the normal distribution method is used to generate the element values of the random matrix; Specifically, you need to determine the mean and standard deviation of the normal distribution, and then use a normal distribution random number generator to generate random numbers.

[0030] The generated random numbers will be used as element values of the random matrix.

[0031] When the distribution characteristics of rail transit data belong to uniform distribution, the uniform distribution method is used to generate the element values of the random matrix; Specifically, it is necessary to determine the uniform distribution interval range and then use a uniform distribution random number generator to generate random numbers.

[0032] The generated random numbers will be used as element values of the random matrix.

[0033] This step generates a reversible random matrix that matches the number of columns in the matrix data. It also selects an appropriate random generation algorithm (such as a normal or uniform distribution method) based on the distribution characteristics of rail transit data. This allows for a more targeted and adaptable random matrix. By calculating the determinant or matrix rank to check the reversibility of the random matrix, it ensures its reliability and validity, avoiding the issue of data restoration due to the irreversibility of the random matrix. This approach not only improves the efficiency of random matrix generation but also enhances the data obfuscation effect, further enhancing the strength and reliability of data privacy protection.

[0034] Step 103: Multiply the matrix data by the random matrix to obtain confusion matrix data.

[0035] Step 103 may specifically include: Determine the matrix size based on the matrix data and the number of rows and columns of the random matrix; When the matrix size is smaller than the minimum value of the preset size range, the matrix data is directly multiplied with the random matrix; When the matrix size falls within the preset size range, the matrix data is multiplied with a random matrix using the Strassen algorithm or the Winograd algorithm; When the matrix size is greater than the maximum value of the preset size range, the matrix data and the random matrix are divided into multiple sub-matrix blocks respectively; Perform matrix multiplication operations on each sub-matrix block to obtain the corresponding sub-result block; All sub-result blocks are spliced according to their positions in the original matrix to form a complete confusion matrix data.

[0036] This step uses direct multiplication, optimization algorithms (such as the Strassen algorithm or the Winograd algorithm), or block processing to perform matrix multiplication operations on matrix data of different sizes, allowing for flexible selection of the most appropriate calculation method based on the data size. This not only improves the efficiency of matrix operations and reduces the consumption of computing resources, but also ensures the efficiency and stability of the entire data processing process, making this method suitable for privacy-preserving large-scale rail transit data, enhancing its practicality and scalability, and providing strong support for the secure processing of rail transit data in a big data environment.

[0037] Step 104: Evaluate the privacy protection strength of the confusion matrix data until the privacy protection strength meets the preset standard.

[0038] Step 104 may specifically include: Determine the evaluation indicators of privacy protection strength based on the characteristics of rail transit data and privacy protection requirements; Calculate the privacy protection strength value of the confusion matrix data based on the evaluation index; Compare the privacy protection strength value with the preset standard; If the privacy protection strength value does not meet the preset standard, the random matrix is regenerated, the complexity of the random matrix is increased, or noise is added to the confusion matrix data to enhance the privacy protection strength value until the privacy protection strength meets the preset standard.

[0039] This step introduces a privacy protection strength assessment mechanism, which quantitatively evaluates the privacy protection effectiveness of confusion matrix data by setting evaluation indicators and preset standards. This method can monitor the level of data privacy protection in real time, ensuring that it meets the privacy protection requirements of practical applications. When the privacy protection strength does not meet the preset standard, adjustments can be made by regenerating the random matrix, increasing the complexity of the random matrix, or adding noise to the confusion matrix data. This can dynamically optimize privacy protection measures, further enhancing the data privacy protection effect. This improves the reliability and adaptability of the entire method, enabling it to be flexibly adjusted according to different application scenarios and privacy requirements, effectively addressing various data leakage risks.

[0040] Step 105: When receiving a rail transit data usage instruction, determine the inverse matrix of the random matrix, and multiply the inverse matrix by the confusion matrix data to obtain restored rail transit data.

[0041] In this embodiment, by converting rail transit data into a matrix form and obfuscating it using a reversible random matrix, the characteristics of the original data can be effectively hidden, significantly reducing the risk of data leakage during storage and transmission. Furthermore, using the inverse of the random matrix for data restoration ensures the integrity and availability of the data for legitimate use, achieving a balance between privacy protection and data availability, and providing an efficient and reliable method for the secure processing of rail transit data.

[0042] In one embodiment of this specification, the evaluation index of the privacy protection strength is determined based on the characteristics of rail transit data and the privacy protection requirements, including: Step 1: Evaluate the sensitivity of each data attribute in the rail transit data and determine the corresponding evaluation index weight according to the sensitivity.

[0043] Suppose rail transit data contains passenger personal information (such as ID numbers and mobile phone numbers) and train status data (such as speed and location). ID numbers and mobile phone numbers are highly sensitive, while train speed and location, while also important, are less sensitive. Therefore, a higher weight (such as 0.8) can be assigned to ID numbers and mobile phone numbers, while a lower weight (such as 0.2) can be assigned to train speed and location.

[0044] Step 2: For rail transit data where the proportion of categorical data or numerical data exceeds a preset ratio, information entropy is selected as the evaluation indicator. The privacy protection strength is quantified by calculating the probability distribution of each element value in the confusion matrix data. The higher the information entropy, the stronger the privacy protection strength.

[0045] Categorical data refers to data types with finite, fixed categories (or labels). Unlike numerical data (continuous values like age and income), categorical data typically represents discrete, non-numeric attributes. Even when numerical values are present, they primarily identify categories rather than represent magnitude or quantity. For example, the range of values for a train's operating status might include: normal operation, delayed, or suspended. The range of values for a passenger's ticket type might include: one-way, round-trip, monthly, or annual.

[0046] Step 3: When it is necessary to evaluate the correlation between confusion matrix data and matrix data, select mutual information as the evaluation indicator; the lower the mutual information, the higher the privacy protection strength.

[0047] Step 4: For rail transit data that needs to meet the requirements of differential privacy protection, select the privacy budget value of differential privacy as the evaluation indicator; the smaller the privacy budget value, the higher the privacy protection strength.

[0048] In this embodiment, a variety of privacy protection strength assessment metrics are provided, including information entropy, mutual information, and differential privacy privacy budget. These metrics can quantify the degree of data privacy protection from different perspectives. For example, information entropy is used to measure data uncertainty, mutual information is used to assess the correlation between data, and differential privacy privacy budget is used to meet specific privacy protection requirements. By flexibly selecting evaluation metrics based on different data characteristics and privacy requirements, the strength of privacy protection can be more accurately assessed, providing a more scientific basis for optimizing privacy protection measures. This improves the accuracy and adaptability of the entire method, enabling it to better meet the privacy protection needs of rail transit data in different scenarios.

[0049] In one embodiment of the present specification, after evaluating the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard, the method further includes: Slice the confusion matrix data according to preset rules to generate multiple data segments; Specifically, determine the sharding rules, such as sharding by row, by column, or by matrix block. For example, a large confusion matrix can be split into several small matrices by row, each of which contains a portion of the rows of the original matrix; According to the sharding rule, the confusion matrix data is divided into multiple data fragments. For example, if the confusion matrix is a 100×100 matrix, it can be divided into 10 small matrix fragments of 10×100 by row.

[0050] Store each data fragment in a different storage node or storage medium to ensure isolation between data fragments; Specifically, multiple storage nodes or storage media are selected, such as different servers, cloud storage services, or physical storage devices; Store each data fragment in a different storage node or storage medium. For example, store the first data fragment in server A, the second data fragment in server B, and so on. Ensure the security of each storage node or storage medium, such as protecting the security of data fragments through encryption, access control, and other measures.

[0051] When receiving a rail transit data usage instruction, it also includes: Collect all data fragments and perform restoration processing; Specifically, after receiving the usage instruction, the data restoration process is triggered; Collect all data fragments stored in different storage nodes or storage media; According to the rules of sharding, all data fragments are spliced together to restore the complete confusion matrix data; Determine the inverse matrix of the random matrix, and multiply the inverse matrix with the restored confusion matrix data to obtain the restored rail transit data.

[0052] In this embodiment, after the privacy protection strength meets the preset criteria, the confusion matrix data is further fragmented and stored in different storage nodes or storage media, achieving distributed data storage. This storage method increases the isolation of data storage. Even if some storage nodes or storage media are compromised, it is difficult for attackers to obtain complete data fragments, further reducing the risk of data leakage. Furthermore, by collecting all data fragments and restoring them when the data is needed, the integrity and availability of the data are ensured, providing more reliable protection for the secure storage and use of rail transit data and enhancing the data security of the entire method.

[0053] In one embodiment of the present specification, before converting the rail transit data into matrix data, the method further includes: Distinguish between structured and unstructured data in rail transit data; Specifically, structured data: usually stored in a relational database with a clear table structure, such as passenger information (ID number, mobile phone number, ticket purchase time, etc.), train operation status (train number, speed, location, etc.); Unstructured data: Usually stored in the file system and without a fixed table structure, such as surveillance videos, log files, passenger feedback, etc. Distinguishing method: Distinguishing by data storage format, data type, and data source. For example, data extracted from a relational database is structured data, while data extracted from a file system is unstructured data.

[0054] Combine structured and unstructured data; The combination method includes at least one of the following: Feature concatenation, weighted averaging, or feature extraction and fusion of deep learning models; Specifically, feature concatenation: directly concatenates the features of structured and unstructured data. For example, a passenger's ID number (structured data) and the corresponding surveillance video features (unstructured data) can be concatenated into a single feature vector. Weighted average: Take a weighted average of the features of structured and unstructured data. For example, assign a weight of 0.7 to structured data and a weight of 0.3 to unstructured data, and then calculate the weighted average. Feature extraction and fusion of deep learning models: Deep learning models (such as convolutional neural networks and recurrent neural networks) are used to extract features from structured and unstructured data, and then the extracted features are fused. For example, a convolutional neural network can be used to extract features from surveillance videos, and a recurrent neural network can be used to extract features from time series data. The two features are then fused through a fully connected layer.

[0055] In this embodiment, before converting rail transit data into matrix data, the structured and unstructured data in the data are distinguished and combined using methods such as feature concatenation, weighted averaging, or feature extraction and fusion using deep learning models. This processing method can fully utilize the clear structure of structured data and the rich information of unstructured data, organically integrating different types of data to form more complete and representative matrix data. This not only enriches the content of the data and improves its usability, but also enhances the adaptability of the entire method to complex data environments, enabling it to better handle the mixed structured and unstructured data commonly found in the rail transit field, thereby improving the flexibility and practicality of data processing.

[0056] In one embodiment of the present specification, after generating a reversible random matrix, the following steps are further included: Using a quantum entangled state generator, quantum entangled state information is injected into each element of the random matrix. The quantum entangled state information is represented by quantum bits (qubits), and the state of each qubit is determined by the element value of the random matrix. Before multiplying the matrix data with the random matrix, also include: Quantum state verification is performed on the random matrix injected with quantum entangled state to ensure that the quantum state information of each element is complete and has not been tampered with.

[0057] In this embodiment, a quantum entangled state generator is used to inject quantum entangled state information into each element of the random matrix, and quantum state verification is used to ensure the integrity and non-tampering of the information. The introduction of quantum entangled state information adds quantum-level security to the random matrix. Due to the non-cloning and non-tampering properties of the quantum state, the random matrix has higher security during storage and transmission, further enhancing the data obfuscation effect and privacy protection strength. At the same time, the quantum state verification mechanism can effectively detect the integrity of the random matrix during processing, preventing the data from being maliciously tampered with or damaged during the obfuscation process, providing a higher level of security for the privacy protection of rail transit data, improving the security and reliability of the entire method, and making it more advantageous in the face of high-level security threats.

[0058] In some other embodiments of this specification, after generating a reversible random matrix, the method further includes: Perform quantum encryption on the random matrix, and use the non-cloning and non-tampering properties of the quantum state to generate a unique quantum encryption identifier for each element of the random matrix; Before multiplying the matrix data with the random matrix, quantum state encoding is performed on each row of the matrix data so that each row of the matrix data corresponds to a quantum state; Perform quantum matrix multiplication on the matrix data after quantum state encoding and the random matrix after quantum encryption to obtain quantum confusion matrix data; When evaluating the privacy protection strength of confusion matrix data, quantum entanglement is introduced as an evaluation indicator to quantify the privacy protection strength by measuring the degree of entanglement between quantum states in the quantum confusion matrix data; among them, the higher the quantum entanglement, the higher the privacy protection strength.

[0059] In some other embodiments of the present specification, after multiplying the matrix data with the random matrix, the method further includes: Blockchain technology is used to distribute and record confusion matrix data in an unalterable manner. The confusion matrix data is divided into multiple data blocks, and each data block generates a hash value and is stored on each node of the blockchain. A smart contract is set up on the blockchain. When a command to use rail transit data is received, the smart contract automatically triggers the data restoration process. The data restoration operation is allowed only after the integrity of the hash value of each data block on the blockchain is verified; During the restoration process, the blockchain consensus mechanism is combined to perform multi-node verification on the restored rail transit data to ensure the authenticity and integrity of the data and prevent the data from being tampered with or forged during the restoration process.

[0060] In some other embodiments of this specification, after multiplying the matrix data with the random matrix, the method further includes: Step 1: Perform multimodal feature extraction on the confusion matrix data to extract the numerical features, structural features and time series features of the confusion matrix data to form a multimodal feature vector; Specifically, numerical feature extraction: calculate the statistical characteristics of each element in the confusion matrix, such as mean, variance, maximum, minimum, etc. These features can reflect the distribution of the data; Structural feature extraction: Analyze the structural characteristics of the confusion matrix, such as the sparsity, symmetry, and row-column correlation of the matrix. For example, the linear relationship between data can be evaluated by calculating the row-column correlation coefficient matrix of the matrix; Time series feature extraction: If the confusion matrix data has a time sequence (for example, rail transit data is recorded by time), then extract time series features, such as the data's changing trend, periodicity, autocorrelation, etc. Time series analysis methods, such as ARIMA models or wavelet transforms, can be used to extract these features; Multimodal feature vector fusion: Combine the three features mentioned above into a multimodal feature vector for subsequent anomaly detection. For example, numerical features, structural features, and time series features are weighted together to form a high-dimensional feature vector.

[0061] Step 2: Based on the multimodal feature vector, a deep neural network model is constructed to dynamically monitor abnormal behavior of confusion matrix data during storage and transmission; Specifically, model architecture design: Design a deep neural network model, such as using a convolutional neural network (CNN) to process numerical features, a graph neural network (GNN) to process structural features, and a recurrent neural network (RNN) or long short-term memory network (LSTM) to process time series features. The output features of these networks are then integrated to form the final anomaly detection model. Training data preparation: Prepare sample data of normal and abnormal behavior for training deep neural network models. Normal behavior samples can be privacy-preserving data, while abnormal behavior samples can be generated by simulating data leakage or tampering. Model training and optimization: Use training data to train a deep neural network model, and adjust model performance through cross-validation and hyperparameter optimization to ensure that the model can accurately detect abnormal behavior.

[0062] Step 3: Introduce an attention mechanism into the deep neural network model to weight the key features in the multimodal feature vector to improve the model's detection accuracy for potential data leakage risks; Specifically, attention mechanism design: Introducing an attention mechanism into the feature fusion layer of a deep neural network, such as using a self-attention mechanism or a channel attention mechanism. The attention mechanism can automatically learn the importance weight of each feature in the feature vector; Feature weighting: Based on the weights calculated by the attention mechanism, each feature in the multimodal feature vector is weighted to highlight key features (weights higher than the first preset weight) and suppress unimportant features (weights lower than the second preset weight); Step 4: Improve detection accuracy: Input the weighted feature vector into the anomaly detection model to improve the model's detection accuracy for abnormal behavior; Specifically, anomaly detection threshold setting: a threshold is set based on the anomaly score output by the model. When the anomaly score exceeds the threshold, it is determined to be abnormal behavior; Alarm mechanism triggering: When abnormal behavior is detected, the alarm mechanism is automatically triggered and the system administrator is notified via email, SMS or system notification; Abnormal information recording: Records detailed information of abnormal behavior, including occurrence time, data segment location, abnormal characteristic value, etc., to facilitate subsequent analysis and processing.

[0063] Step 5: When abnormal behavior is detected, the alarm mechanism is automatically triggered and detailed information about the abnormal behavior is recorded, including the time of occurrence, data segment location, and abnormal feature value; According to the characteristics of abnormal behavior, the privacy protection strategy is dynamically adjusted, such as regenerating the random matrix, increasing the noise intensity of the confusion matrix data, or adjusting the data storage path to enhance the privacy protection strength of the data.

[0064] Specifically, abnormal behavior analysis: analyzing the characteristics of abnormal behavior, such as abnormal type, frequency of occurrence, scope of impact, etc. Privacy protection policy adjustment: Select appropriate privacy protection policies and adjust them based on the characteristics of abnormal behavior; If data tampering is detected, regenerate the random matrix and re-obfuscate the data; If it is detected that the data is frequently accessed, increase the noise intensity of the confusion matrix data; If it is detected that the data storage path is attacked, the data storage path is adjusted to store the data fragments in a more secure node.

[0065] In some other embodiments of this specification, after evaluating the privacy protection strength of the confusion matrix data, the method further includes: Step 1: Build a data association graph based on a graph neural network (GNN). Each data segment in the confusion matrix data is used as a node in the graph, and the associations between data segments are used as edges to form a complex data association network. Specifically, the data segment node definition is as follows: the confusion matrix data is divided into multiple data segments, each data segment serves as a node; Relevance edge definition: Define edges based on the similarity or correlation between data segments. For example, evaluate relevance by calculating cosine similarity or mutual information between data segments. Graph construction: Use graph neural networks (such as GCN or GAT) to build a data association graph. The nodes in the graph represent data fragments, and the edges represent the associations between data fragments.

[0066] Step 2: Use graph neural networks to conduct in-depth analysis of the data association graph, explore the potential correlations between data segments, and evaluate the importance of data segments in the graph.

[0067] In the data association graph, association relationships refer to the logical or semantic connections between data fragments. These connections can be direct or indirect, reflecting the mutual dependence or similarity between data fragments. Direct association: For example, two data fragments may come from the same data source (such as the operation data and scheduling data of the same train), or there may be a clear causal relationship between them (such as equipment failure data and maintenance records). Indirect association: A relationship that is indirectly connected through other data fragments or nodes. For example, two data fragments may be related to each other through multiple data fragments in the middle (such as passenger flow data at multiple stations).

[0068] In graph neural networks, importance refers to the criticality of a particular data segment (node) within the entire graph, reflecting its role and value within the overall data structure. The norm of the node embedding vector is measured by the size (norm) of the node embedding vector learned by the graph neural network. A larger norm indicates a more important node within the graph. Node centrality metrics, such as degree centrality, betweenness centrality, and closeness centrality, measure a node's connectivity within the network and its position on critical paths.

[0069] Specifically, graph neural network training: Use graph neural networks to train data association graphs and learn the association relationships between data segments; Importance assessment: Evaluate the importance of each data segment in the graph using the output of the graph neural network. For example, the norm of the node's embedding vector or the node's centrality index can be used to evaluate importance. Potential association mining: Mining potential associations between data segments through the attention weight or edge prediction module of the graph neural network.

[0070] Step 3: Based on the importance assessment results of the data association graph, the confusion matrix data is encrypted in layers, using a higher-level encryption algorithm for more important data segments and a lighter-weight encryption algorithm for less important data segments; Specifically, encryption algorithm selection: select an appropriate encryption algorithm based on the importance of the data segment. For example, use the AES-256 encryption algorithm for highly important data segments and the AES-128 encryption algorithm for less important data segments. Layered encryption implementation: Based on the importance assessment results, each data fragment is encrypted in layers; Encryption key management: Generate and manage encryption keys to ensure key security.

[0071] Step 4: During data storage and transmission, the layered encrypted data fragments are dynamically reorganized. Based on the position and relevance of the data fragments in the graph, the storage order and transmission path of the data fragments are randomly adjusted to further reduce the risk of data leakage. Specifically, the dynamic reorganization strategy is designed as follows: Storage order adjustment: Based on the topological structure of nodes (data fragments) in the data association graph, the storage order of data fragments is randomly adjusted. For example, for data fragments with higher importance, they can be stored on more secure storage nodes, and their security can be enhanced through encryption and access control; Transmission path adjustment: During data transmission, the transmission path is dynamically selected based on the correlation between data fragments. For example, for data fragments with strong correlation, they can be sent through different transmission paths to avoid being intercepted at the same time; Introduction of randomness: Introducing randomness makes the order and path of storage and transmission different each time, increasing the difficulty for attackers to analyze and crack.

[0072] The storage node selection is as follows: Security assessment: Perform security assessments on all available storage nodes. Assessment indicators include physical security, network security, access control policies, etc. Dynamic allocation: Storage nodes are dynamically selected based on the importance and security assessment results of data fragments. For example, highly important data fragments are stored on highly secure nodes, while less important data fragments are stored on ordinary nodes.

[0073] The transmission path planning is as follows: Path diversity: Design multiple transmission paths to ensure that data fragments can be transmitted through different paths. Path selection can be based on factors such as network topology, bandwidth, and latency. Random path selection: A random path is selected during each transmission, making the transmission path of the data fragment unpredictable.

[0074] Step 5. When receiving the instruction to use rail transit data, quickly restore the layered encrypted data fragments according to the structural information of the data association map to ensure the integrity and availability of the data.

[0075] Based on the same general inventive concept, the present invention also protects a rail transit data anti-leakage processing device based on matrix transformation, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the structure of a rail transit data anti-leakage processing device based on matrix transformation provided by an embodiment of the present invention. The rail transit data anti-leakage processing device based on matrix transformation provided by the present invention is described below. The rail transit data anti-leakage processing device based on matrix transformation described below and the rail transit data anti-leakage processing method based on matrix transformation described above can be used in conjunction with each other.

[0076] The rail transit data anti-leakage processing device based on matrix transformation includes a matrix conversion module 201, a random matrix module 202, a matrix mixing module 203, a matrix evaluation module 204 and a matrix restoration module 205.

[0077] The matrix conversion module 201 is used to convert rail transit data into matrix data; wherein each row of the matrix data represents a data record, and each column represents a data attribute; The random matrix module 202 is used to generate a reversible random matrix according to the number of columns of the matrix data; wherein the number of rows and columns of the random matrix matches the number of columns of the matrix data; The matrix mixing module 203 is used to multiply the matrix data with the random matrix to obtain confusion matrix data; The matrix evaluation module 204 is used to evaluate the privacy protection strength of the confusion matrix data until the privacy protection strength meets the preset standard; The matrix restoration module 205 is used to determine the inverse matrix of the random matrix when receiving a use instruction of the rail transit data, and multiply the inverse matrix by the confusion matrix data to obtain the restored rail transit data.

[0078] Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention.

[0079] like Figure 3 As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the rail transit data anti-leakage processing method based on matrix transformation.

[0080] Furthermore, the logic instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0081] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the rail transit data anti-leakage processing method based on matrix transformation provided by the above methods.

[0082] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the rail transit data anti-leakage processing method based on matrix transformation provided by the above methods.

[0083] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0084] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for preventing rail transit data from being leaked based on matrix transformation, characterized in that: include: Converting rail transit data into matrix data; wherein each row of the matrix data represents a data record and each column represents a data attribute; Generate a reversible random matrix according to the number of columns of the matrix data; wherein the number of rows and columns of the random matrix matches the number of columns of the matrix data; Multiplying the matrix data by the random matrix to obtain confusion matrix data; Evaluating the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard; When an instruction to use the rail transit data is received, the inverse matrix of the random matrix is determined, and the inverse matrix is multiplied by the confusion matrix data to obtain restored rail transit data.

2. The method for preventing rail transit data leakage based on matrix transformation according to claim 1 is characterized in that: Convert rail transit data into matrix data, including: Determine the data type of each data attribute of the rail transit data, and map each data attribute to a column of the matrix data using a mapping rule corresponding to the data type; Arrange each data record according to its relevance to form a row of matrix data; The step of determining the data type of each data attribute of the rail transit data and mapping each data attribute to a column of matrix data using a mapping rule corresponding to the data type includes: When the data attribute is numerical data, its value is directly used as the matrix element; When the data attribute is categorical data, one-hot encoding is used to convert it into a numerical type as a matrix element; The association includes a temporal order or a spatial order.

3. The method for preventing rail transit data leakage based on matrix transformation according to claim 1, characterized in that: Generate a reversible random matrix according to the number of columns of the matrix data, including: Determine the dimensions of the random matrix so that the number of its rows and columns matches the number of columns of the matrix data; determining a distribution characteristic of rail transit data, and generating elements of a random matrix using a random generation algorithm corresponding to the distribution characteristic; Check whether the random matrix is invertible by calculating the determinant value or matrix rank. If it is not invertible, regenerate it until a reversible random matrix is obtained; The determining of the distribution characteristics of the rail transit data and the use of a random generation algorithm corresponding to the distribution characteristics to generate elements of a random matrix include: When the distribution characteristics of rail transit data belong to normal distribution, the normal distribution method is used to generate the element values of the random matrix; When the distribution characteristics of rail transit data belong to uniform distribution, the uniform distribution method is used to generate the element values of the random matrix.

4. The method for preventing rail transit data leakage based on matrix transformation according to claim 1, characterized in that: Multiplying the matrix data with the random matrix to obtain confusion matrix data, including: Determining a matrix scale according to the matrix data and the number of rows and columns of the random matrix; When the matrix size is smaller than the minimum value of the preset size range, directly multiplying the matrix data by the random matrix; When the matrix size falls within the preset size range, multiplying the matrix data by the random matrix using the Strassen algorithm or the Winograd algorithm; When the matrix size is greater than the maximum value of the preset size range, dividing the matrix data and the random matrix into a plurality of sub-matrix blocks respectively; Perform matrix multiplication operations on each sub-matrix block to obtain the corresponding sub-result block; All sub-result blocks are spliced according to their positions in the original matrix to form a complete confusion matrix data.

5. The method for preventing rail transit data leakage based on matrix transformation according to claim 1, characterized in that: Evaluating the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard, including: Determine the evaluation indicators of privacy protection strength based on the characteristics of rail transit data and privacy protection requirements; Calculate the privacy protection strength value of the confusion matrix data based on the evaluation index; Compare the privacy protection strength value with the preset standard; If the privacy protection strength value does not meet the preset standard, the random matrix is regenerated, the complexity of the random matrix is increased, or noise is added to the confusion matrix data to enhance the privacy protection strength value until the privacy protection strength meets the preset standard.

6. The method for preventing rail transit data from being leaked based on matrix transformation according to claim 5, characterized in that: The evaluation indicators for privacy protection strength are determined based on the characteristics of rail transit data and privacy protection requirements, including: Evaluate the sensitivity of each data attribute in rail transit data and determine the corresponding evaluation index weight according to the sensitivity level; For rail transit data where the proportion of categorical data or numerical data exceeds a preset ratio, information entropy is selected as the evaluation indicator. The probability distribution of each element value in the confusion matrix data is calculated to quantify the privacy protection strength. The higher the information entropy, the stronger the privacy protection strength. When it is necessary to evaluate the correlation between confusion matrix data and matrix data, mutual information is selected as the evaluation indicator; the lower the mutual information, the higher the privacy protection strength; For rail transit data that needs to meet the requirements of differential privacy protection, the privacy budget value of differential privacy is selected as the evaluation indicator; among them, the smaller the privacy budget value, the higher the privacy protection strength.

7. The method for preventing rail transit data leakage based on matrix transformation according to claim 1, characterized in that: After evaluating the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard, the method further includes: Slice the confusion matrix data according to preset rules to generate multiple data segments; Store each data fragment in a different storage node or storage medium to ensure isolation between data fragments; When receiving a rail transit data usage instruction, it also includes: Collect all data fragments and perform restoration processing.

8. The method for preventing rail transit data leakage based on matrix transformation according to claim 1, characterized in that: Before converting rail transit data into matrix data, it also includes: Distinguish between structured and unstructured data in rail transit data; combining the structured data with the unstructured data; The combination method includes at least one of the following: Feature concatenation, weighted averaging, or feature extraction and fusion of deep learning models.

9. The method for preventing rail transit data leakage based on matrix transformation according to claim 1, characterized in that: After generating a reversible random matrix, also include: Using a quantum entangled state generator to inject quantum entangled state information into each element of the random matrix; wherein the quantum entangled state information is represented by a quantum bit, and the state of each quantum bit is determined by the element value of the random matrix; Before multiplying the matrix data with the random matrix, the method further includes: Quantum state verification is performed on the random matrix injected with quantum entangled state to ensure that the quantum state information of each element is complete and has not been tampered with.

10. A rail transit data anti-leakage processing device based on matrix transformation, characterized in that: include: A matrix conversion module, configured to convert rail transit data into matrix data, wherein each row of the matrix data represents a data record and each column represents a data attribute; A random matrix module, configured to generate a reversible random matrix according to the number of columns of the matrix data; wherein the number of rows and columns of the random matrix matches the number of columns of the matrix data; a matrix mixing module, configured to multiply the matrix data by the random matrix to obtain confusion matrix data; A matrix evaluation module, configured to evaluate the privacy protection strength of the confusion matrix data until the privacy protection strength meets a preset standard; The matrix restoration module is used to determine the inverse matrix of the random matrix when receiving an instruction to use the rail transit data, and multiply the inverse matrix by the confusion matrix data to obtain restored rail transit data.

Citation Information

Patent Citations

  • SO file security reinforcement / calling method and system and reversible matrix generation server

    CN113849780A

  • Data encryption method and device, electronic equipment, storage medium and computer program

    CN118612355A

  • Private data protection method and system based on homomorphic encryption and federated learning

    CN119513919A

  • Data processing method and device, equipment, storage medium and program product

    CN119903291A

  • Privacy-Preserving Aggregated Data Mining

    US20140040172A1