Intelligent networked automobile data desensitization method and device based on machine learning
Through the intelligent connected vehicle data desensitization method based on machine learning, the problem that the existing technology is difficult to flexibly adjust the data desensitization strategy in complex scenarios is solved, and efficient and intelligent data desensitization effects are achieved to meet the needs of advanced analysis.
Patent Information
- Application Number
- CN202411971867.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing technology is difficult to flexibly adjust data desensitization strategies in complex scenarios, and real-time desensitization technology cannot meet the needs of advanced analysis in some scenarios.
Using the intelligent connected vehicle data desensitization method based on machine learning, efficient desensitization of intelligent connected vehicle data is achieved by collecting original data, identifying sensitive information, building desensitization strategies, building target data sets and training desensitization models.
It improves the performance and effect of data desensitization, can flexibly adjust the desensitization strategy in complex scenarios, meet the needs of advanced analysis, and trains the data desensitization model through deep learning models to improve the intelligence level of desensitization.
Smart Images

Figure CN119939651A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data desensitization, and in particular to a method and device for desensitizing data of an intelligent connected vehicle based on machine learning. Background Art
[0002] With the development of Internet of Vehicles technology, the amount of data generated by vehicles has increased dramatically, including a large amount of data involving personal privacy, such as vehicle positioning information, on-board camera records, etc. In order to protect the privacy rights and interests of users, these sensitive data need to be desensitized. Currently, there are intelligent desensitization technologies, real-time desensitization technologies, etc., but these technologies lack an in-depth understanding of the data context, resulting in the inability to flexibly adjust desensitization strategies in complex scenarios. Intelligent desensitization technology can adaptively adjust desensitization strategies based on data usage scenarios and user behaviors, but this method may require more manual intervention to adjust desensitization rules. Although real-time desensitization technology saves storage space, in some scenarios, the desensitized data may not meet the needs of advanced analysis. Summary of the invention
[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a method and device for data desensitization of intelligent connected vehicles based on machine learning, in order to solve at least one of the problems of the prior art. The present invention can improve the performance and effect of data desensitization.
[0004] To achieve the above-mentioned purpose, an embodiment of the present invention provides a method for desensitizing data of intelligent connected vehicles based on machine learning, comprising:
[0005] Collect raw data of vehicle information;
[0006] Based on preset rules, sensitive information of the original data is identified to obtain sensitive data;
[0007] Constructing a desensitization strategy based on the sensitive data;
[0008] According to the desensitization strategy, construct a target data set;
[0009] According to the target data set, the initial data desensitization model is trained to obtain a target data desensitization model;
[0010] Among them, the target data desensitization model is used to desensitize the intelligent connected vehicle data.
[0011] In some embodiments, a method for desensitizing data of a smart connected vehicle based on machine learning further includes:
[0012] Evaluate the target data desensitization model to obtain an evaluation result;
[0013] If the evaluation result is that the desensitization does not meet the standard, then return to the step of training the initial data desensitization model according to the target data set to obtain the target data desensitization model. If the evaluation result is that the desensitization meets the standard, then deploy and monitor the target data desensitization model.
[0014] In some embodiments, the collecting of raw data of vehicle information includes the following steps:
[0015] The original data of the vehicle information is obtained through active collection of on-board equipment and passive collection in the cloud.
[0016] In some embodiments, the step of identifying sensitive information of the original data based on preset rules to obtain sensitive data includes the following steps:
[0017] Presetting the data type of the sensitive data;
[0018] Presetting identification rules corresponding to each of the data types;
[0019] According to the identification rule, the sensitive information of the original data is identified to obtain the sensitive data.
[0020] In some embodiments, constructing a desensitization strategy based on the sensitive data includes the following steps:
[0021] Obtaining the sensitivity and application scenarios of the sensitive data;
[0022] Constructing the desensitization strategy according to the sensitivity and the application scenario;
[0023] Among them, the desensitization strategy includes data replacement, data masking, data encryption and data perturbation.
[0024] In some embodiments, constructing a target data set according to the desensitization strategy comprises the following steps:
[0025] Desensitizing the original data according to the desensitization strategy to obtain a first desensitized text;
[0026] Obtaining an initial data set according to the original data and the corresponding first desensitized text;
[0027] The initial data set is preprocessed to obtain the target data set.
[0028] In some embodiments, the training of the initial data desensitization model according to the target data set to obtain the target data desensitization model includes the following steps:
[0029] Inputting the target data set into the encoder of the initial data desensitization model to obtain an intermediate vector;
[0030] Inputting the intermediate vector into the decoder of the initial data desensitization model to obtain a second desensitized text;
[0031] Constructing a cross entropy loss function according to the first desensitized text and the second desensitized text;
[0032] Through the teacher forcing technology, the next word of the first desensitized text is input into the decoder, and the initial data desensitization model is trained in combination with the cross entropy loss function to obtain the target data desensitization model.
[0033] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides a data desensitization device for intelligent connected vehicles based on machine learning, the device comprising:
[0034] The first module is used to collect raw data of vehicle information;
[0035] The second module is used to identify the sensitive information of the original data based on preset rules to obtain sensitive data;
[0036] The third module is used to build a desensitization strategy based on the sensitive data;
[0037] The fourth module is used to construct a target data set according to the desensitization strategy;
[0038] The fifth module is used to train the initial data desensitization model according to the target data set to obtain a target data desensitization model;
[0039] Among them, the target data desensitization model is used to desensitize the intelligent connected vehicle data.
[0040] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned method for desensitizing data of intelligent connected vehicles based on machine learning.
[0041] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for desensitizing data of intelligent connected vehicles based on machine learning.
[0042] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the aforementioned method for desensitizing data of an intelligent connected vehicle based on machine learning.
[0043] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method and device for desensitizing data of intelligent networked vehicles based on machine learning, which collects raw data of vehicle information; identifies sensitive information of the raw data based on preset rules to obtain sensitive data; constructs a desensitization strategy based on the sensitive data; constructs a target data set based on the desensitization strategy; trains an initial data desensitization model based on the target data set to obtain a target data desensitization model; wherein the target data desensitization model is used to desensitize data of intelligent networked vehicles. The present invention uses a deep learning model to train a data desensitization model, allowing the model to learn complex desensitization rules, thereby improving the performance and desensitization effect of data desensitization. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0045] Figure 1 It is a flow chart of a method for desensitizing data of an intelligent connected vehicle based on machine learning provided by an embodiment of the present invention;
[0046] Figure 2 It is a schematic diagram of an optional implementation process of intelligent connected vehicle data desensitization based on machine learning provided in an embodiment of the present invention;
[0047] Figure 3 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention, and they are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.
[0049] It should be noted that, although the functional modules are divided in the system schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the flow chart. The terms "first / S100" and "second / S200" in the specification and claims and the above-mentioned drawings may be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiment of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determination".
[0050] The terms "at least one", "multiple", "each", "any", etc. used in the present invention, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0052] With the development of Internet of Vehicles technology, the amount of data generated by vehicles has increased dramatically, including a large amount of data involving personal privacy, such as vehicle positioning information, on-board camera records, etc. In order to protect the privacy rights and interests of users, these sensitive data need to be desensitized. Internet of Vehicles systems face a variety of security threats, including data leakage, tampering, and unauthorized access. Desensitization technology can be used as part of security protection to reduce security risks by reducing the sensitivity of data. Currently, there are intelligent desensitization technology, real-time desensitization technology, etc., but these technologies lack in-depth understanding of data context, resulting in the inability to flexibly adjust desensitization strategies in complex scenarios. Intelligent desensitization technology can adaptively adjust desensitization strategies based on data usage scenarios and user behaviors, but this method may require more manual intervention to adjust desensitization rules. Although real-time desensitization technology saves storage space, in some scenarios, the desensitized data may not meet the needs of advanced analysis.
[0053] In view of this, if Figure 1 As shown, an embodiment of the present invention provides a method for desensitizing data of an intelligent connected vehicle based on machine learning, which may include but is not limited to steps S100 to S500:
[0054] Step S100, collecting raw data of vehicle information;
[0055] Step S200, based on preset rules, identifying sensitive information of the original data to obtain sensitive data;
[0056] Step S300, constructing a desensitization strategy based on the sensitive data;
[0057] Step S400, constructing a target data set according to the desensitization strategy;
[0058] Step S500, training the initial data desensitization model according to the target data set to obtain a target data desensitization model;
[0059] Among them, the target data desensitization model is used to desensitize the intelligent connected vehicle data.
[0060] In step S100 of some embodiments, the raw data of the vehicle information is obtained through active collection by the vehicle-mounted device and passive collection by the cloud. Optionally, the real-time data of the vehicle is collected by the vehicle-mounted device to realize real-time monitoring of the vehicle's operation status, and the vehicle information can also be collected by the cloud server to realize passive data collection. Among them, the vehicle information may include but is not limited to numerical information such as the vehicle's location, speed, driving route, fuel consumption, shock absorption, as well as the vehicle's status information, fault information, etc. By collecting the raw data of the vehicle information, preparations are made for the generation of a large amount of sensitive data required for subsequent machine learning model training.
[0061] In some embodiments, step S200 may include but is not limited to steps S210 to S230:
[0062] Step S210, presetting the data type of the sensitive data;
[0063] Step S220, presetting an identification rule corresponding to each data type;
[0064] Step S230: identifying the sensitive information of the original data according to the identification rule to obtain the sensitive data.
[0065] In step S210 of some embodiments, the data type of sensitive data is defined to clarify which types of data are considered sensitive. For example, for Internet of Vehicles data, sensitive data may include: personal identity information (such as name, ID number, driver's license number); vehicle information (such as license plate number, vehicle identification number VIN); location information (such as GPS coordinates); communication information (such as phone number, email address), etc., but not limited to this.
[0066] In steps S220 to S230 of some embodiments, corresponding identification rules and patterns may be formulated for each type of sensitive data. Exemplarily, these identification rules may be implemented based on regular expressions (Regex), for example:
[0067] The regular expression for the ID number can be:
[0068] \d{6}(18|19|20)? \d{2}(0[1-9]|1[0-2])(0[1-9]|
[12] \d|3
[01] )\d{3}(\d|X);
[0069] The regular expression for the license plate number can be:
[0070] [Beijing, Tianjin, Shanghai, Chongqing, Hebei, Henan, Yunnan, Liaoning, Heilongjiang, Hunan, Anhui, Shandong, Xinjiang, Jiangsu, Zhejiang, Jiangxi, Hubei, Guangxi, Gansu, Shanxi, Inner Mongolia, Shaanxi, Jilin, Fujian, Guizhou, Guangdong, Qinghai, Tibet, Sichuan, Ningxia, Hainan][AZ]{1}[A-HJ-NPQRTUWXY0-9]{5};
[0071] A regular expression for a phone number could be:
[0072] (\+\d{1,3}[-]?)?\d{10}.
[0073] A rule-based approach is used to identify and process sensitive information in raw data. Through the above-mentioned identification rules, sensitive information in raw data can be identified to obtain sensitive data.
[0074] In some embodiments, step S300 may include but is not limited to steps S310 to S320:
[0075] Step S310, obtaining the sensitivity and application scenario of the sensitive data;
[0076] Step S320: construct the desensitization strategy according to the sensitivity and the application scenario.
[0077] In steps S310 to S320 of some embodiments, a personalized desensitization strategy is formulated by obtaining the sensitivity and application scenarios of sensitive data. Among them, the desensitization strategy may include methods such as data replacement, data masking, data encryption and data perturbation. For sensitive data with high sensitivity (such as personal identity information), encryption or data perturbation may be required to provide stronger protection. For application scenarios where data format and some features need to be maintained (such as credit card number desensitization retaining the first six digits and the last four digits), a masking method can be selected. If the desensitized data still needs to be used for statistical analysis, the overall distribution and trend of the data can be kept unchanged using the data perturbation method.
[0078] In some embodiments, the desensitization strategy constructed according to sensitivity and application scenarios includes the following methods:
[0079] 1. Replacement: When the exact value of the data is not necessary for analysis, or the true value of the data needs to be completely hidden, data replacement can be used. For example, the name is replaced with the user ID. The specific implementation process is to create a mapping table to replace the original data value with a non-sensitive value.
[0080] 2. Masking: When you need to retain some information of the data, such as format or part of the numbers, for easy identification or further processing, you can use masking. For example, mask the credit card number to only display the last four digits. The specific implementation process is to define masking rules, such as retaining the last four digits and replacing the rest with asterisks (*).
[0081] 3. Encryption: Encryption can be used when data needs to be transmitted or stored in a secure environment and can only be accessed by authorized users, or when the data is used for verification. The specific real-time process is to select a powerful encryption algorithm (such as AES, RSA, MD5, etc.) to encrypt the data. For example, for personal information data that needs to be transmitted, asymmetric encryption methods such as AES and RSA can be used; for data such as passwords that only need to be verified, the MD5 hash function can be used for encryption to maximize the security of user accounts.
[0082] 4. Data perturbation: When you need to publish a data set for research but do not want to disclose individual information, you can use data perturbation. For example, to protect the privacy of participants when publishing statistical data. Introducing random noise into the data makes it impossible to accurately restore the original data, but the overall statistical characteristics of the data are preserved. For example, according to differential privacy technology, Laplace noise can be added: Among them, x is the original data; is the data after adding Laplace noise; λ is the noise parameter, which is determined according to the privacy protection level.
[0083] In some embodiments, step S400 may include but is not limited to steps S410 to S430:
[0084] Step S410, desensitizing the original data according to the desensitization strategy to obtain a first desensitized text;
[0085] Step S420, obtaining an initial data set according to the original data and the corresponding first desensitized text;
[0086] Step S430: preprocess the initial data set to obtain the target data set.
[0087] In steps S410 to S420 of some embodiments, the original data may be desensitized according to the desensitization strategy to obtain a first desensitized text. An initial data set including the original data and its corresponding first desensitized text is constructed. The original data contains sensitive information, and the desensitized text is data that has been desensitized.
[0088] In step S430 of some embodiments, the data quality of the initial data set needs to be evaluated and preprocessed. Optionally, the preprocessing may include but is not limited to noise removal, normalization, filling in missing values, abnormal data processing, erroneous data elimination, etc. These preprocessing steps can improve data accuracy and ensure data security.
[0089] In some embodiments, the method for preprocessing the initial data set is as follows:
[0090] 1. Denoising: Moving average filtering can be used, which is a simple filtering method used to smooth data. For time series data x1, x2, ..., x n Moving average y t The calculation formula at time point t is:
[0091]
[0092] Where M is the window size.
[0093] 2. Normalization: The minimum-maximum normalization method can be used to scale the data to the [0,1] interval. The formula is:
[0094]
[0095] 3. Outlier processing: Common missing value processing methods include mean filling and median filling. There are generally few outliers or extreme values in the real-time data of vehicles, and the data is approximately normally distributed, so the mean filling method can be used. The formula is:
[0096] x′=u
[0097] Among them, x′ is the filled value; u is the mean of the data column.
[0098] 4. Error data elimination: Failures of sensors and vehicle components or errors in reading data will generate some abnormal data. The interquartile range (IQR) method can be used to remove these erroneous data, that is, sort the data from small to large, and then find the middle value (median) and the two boundaries above and below the middle value (quartiles).
[0099] 1) Lower limit: median - 1.5 × IQR (interquartile range);
[0100] 2) Upper bound: median + 1.5 × IQR;
[0101] If a data point is below the lower bound or above the upper bound, it is considered an anomaly and is removed or replaced.
[0102] In some embodiments, step S500 may include but is not limited to steps S510 to S540:
[0103] Step S510, inputting the target data set into the encoder of the initial data desensitization model to obtain an intermediate vector;
[0104] Step S520, inputting the intermediate vector into the decoder of the initial data desensitization model to obtain a second desensitized text;
[0105] Step S530, constructing a cross entropy loss function according to the first desensitized text and the second desensitized text;
[0106] Step S540: Input the next word of the first desensitized text into the decoder through the teacher forcing technique, and train the initial data desensitization model in combination with the cross entropy loss function to obtain the target data desensitization model.
[0107] In steps S510 to S540 of some embodiments, a Transformer-based model (such as BERT) is selected to understand the context of sensitive data, and then the Seq2Seq model architecture is used to desensitize the data. Transformer models such as BERT can capture contextual information when processing text data and are suitable for identifying and processing sensitive information in text.
[0108] In step S510 of some embodiments, the collected target data set containing sensitive information is input into the encoder of the Transformer-based initial data desensitization model, and BERT is used to understand the context of the input sensitive data. The encoder processes the input sequence time step by time and outputs a vector at each time step. These vectors together constitute an intermediate representation of the input text, namely, the intermediate vector.
[0109] In step S520 of some embodiments, the decoder and the encoder use the same Transformer architecture to generate the second desensitized text based on the intermediate vector of the encoder. The initial state of the decoder is usually set to the final hidden state of the encoder, so that the context information of the input sequence can be passed. The decoder starts with an empty input sequence and gradually generates each word of the second desensitized text, and the output of each step is used as the input of the next step.
[0110] Exemplarily, the decoder is initialized with an empty input sequence or a special start marker (such as <start>) as the initial input, and sets the initial state of the decoder to the final hidden state of the encoder in order to pass the contextual information of the input sequence. At each time step, the decoder uses not only its own hidden state but also the intermediate vector of the encoder to generate the next word of the desensitized text. Optionally, the next word of the desensitized text can be generated by an attention mechanism: the decoder calculates the attention score between the current hidden state and each hidden state of the encoder, and then weights the hidden state of the encoder according to these scores to generate a context vector; the context vector is combined with the current hidden state of the decoder to form a new hidden state for predicting the next word. The output layer of the decoder then predicts the probability distribution of the next word based on the hidden state combined with the context vector. According to this probability distribution, the word with the highest probability is selected as the next word output, or the next word output is selected through a sampling method (such as greedy sampling, beam search). The decoder updates its hidden state at each step and uses the output of the previous step as the input of the current step, repeating the above process until a special end marker (such as <end>) Or when the preset maximum number of steps is reached, the generation process ends, and the sequence output by the decoder is the desensitized text, i.e., the second desensitized text.
[0111] In step S530 of some embodiments, the cross-entropy loss function is constructed from the first desensitized text and the second desensitized text. The cross-entropy loss function is used to measure the difference between the predicted desensitized text (the second desensitized text) and the true desensitized text (the first desensitized text), thereby improving the accuracy and quality of the desensitized text.
[0112] In step S540 of some embodiments, in the initial stage of training, the Teacher Forcing technique is used, that is, the next word of the true desensitized text (the first desensitized text) is directly used as the input of the decoder to stabilize the training process, and the model parameters are updated through backpropagation and gradient descent algorithms to minimize the loss function. Among them, Teacher Forcing is a training strategy used to accelerate and stabilize the training process of the Seq2Seq model. Under Teacher Forcing, the input of the decoder of the model at each time step is not the output of the previous step, but the next word in the actual target output (i.e., the true desensitized text). Exemplarily, assume there is a true desensitized text sequence "[START]Hello "; during the training process, the decoder may receive the special start token "[START]" as the input at the first time step, and based on this start token and the intermediate vector of the encoder, predict the next word; at the second time step, according to the Teacher Forcing strategy, the input of the decoder is not the output of the first time step, but the next word "Hello" in the true desensitized text, and then the decoder combines the intermediate vector of the encoder and the input "Hello" to predict the next word, and so on, until the decoder generates the special end token " " or reaches the preset maximum number of steps.
[0113] In some embodiments, after obtaining the target data desensitization model, an independent validation set is used to evaluate the performance of the model to ensure that the model is not overfitting. And the hyperparameters of the model, such as the learning rate, the size of the hidden layer, the number of attention heads, etc., are adjusted according to the performance on the validation set.
[0114] In some embodiments, a machine learning-based intelligent connected vehicle data desensitization method may further include, but is not limited to, steps S600 to S700:
[0115] Step S600, evaluate the target data desensitization model to obtain an evaluation result;
[0116] Step S700, if the evaluation result is that the desensitization does not meet the standard, then return to the step of training the initial data desensitization model according to the target data set to obtain the target data desensitization model; if the evaluation result is that the desensitization meets the standard, then deploy and monitor the target data desensitization model.
[0117] In step S600 of some embodiments, the target data desensitization model is used to desensitize the data, and the desensitized data is evaluated, thereby completing the evaluation of the target data desensitization model to ensure the availability and security of the data. The evaluation indicators used may include but are not limited to K-anonymity, l-diversity and t-closeness to quantify the desensitization effect. Exemplarily, there are the following evaluation indicators:
[0118] 1. K-anonymity
[0119] K-anonymity ensures that in the published dataset, each record is indistinguishable from fewer than K other records. This means that an attacker cannot associate an individual with any record, as long as there are at least K-1 individuals whose records are identical to the target individual's record.
[0120] Calculation process:
[0121] 1) For each record, identify all records with the same attribute values (except the identifier).
[0122] 2) If the number of records with the same attribute value is at least K, the record satisfies K-anonymity.
[0123] 3) Repeat steps 1) and 2) until all records meet K-anonymity.
[0124] 2. l-Diversity
[0125] l-diversity ensures that in the published dataset, the sensitive attribute value of each record appears at least l times in each equivalence class. This means that an attacker cannot determine the sensitive attribute value of an individual because each value is common enough.
[0126] Calculation process:
[0127] 1) For each equivalence class (a set of records with the same non-sensitive attribute values), identify the sensitive attribute values.
[0128] 2) If each sensitive attribute value appears at least l times in an equivalence class, then the equivalence class satisfies l-diversity.
[0129] 3) Repeat steps 1 and 2 until all equivalence classes satisfy l-diversity.
[0130] 3. t-proximity
[0131] t-closeness ensures that the distribution of sensitive attribute values of each record in the published dataset is close enough to the distribution in the original dataset. This means that an attacker cannot infer the distribution of the original data from the published data.
[0132] Calculation process:
[0133] 1) Calculate the frequency of each sensitive attribute value in the original dataset.
[0134] 2) Calculate the frequency of each sensitive attribute value in the published dataset.
[0135] 3) Use a statistical test (such as the chi-squared test) to compare the difference between the two distributions. If the difference is within an acceptable threshold (usually defined by t), then the published dataset satisfies t-closeness.
[0136] Through these evaluation indicators and algorithms, the effect of the desensitization algorithm can be quantified, thereby ensuring the availability and security of the data.
[0137] In step S700 of some embodiments, if the evaluation result is that the desensitization is not up to standard, that is, the desensitization effect is poor, the training parameters of the model are adjusted and the model training step is returned. If the evaluation result is that the desensitization is up to standard, that is, the desensitization effect is good, the target data desensitization model is deployed and monitored. Exemplarily, the deployment strategy includes: encapsulating the trained target data desensitization model as a service, such as using the F l ask framework, so that it can be easily integrated into the Internet of Vehicles platform, and designing a standardized API interface, such as RESTful API, to ensure compatibility and interoperability between different vehicles and systems. At the same time, ensure that all data transmissions are encrypted through SSL / TLS to prevent data from being intercepted during transmission; implement API authentication and authorization mechanisms, such as OAuth 2.0, to ensure that only authorized users and systems can access the desensitization service. The monitoring strategy includes: real-time monitoring of system performance indicators, such as response time, throughput and error rate, to ensure the stability and response speed of the service. Record detailed access logs and system logs, including request details, response results and exception information, to facilitate problem tracking and performance analysis. Regularly evaluate the desensitization effect to ensure that the desensitized data meets the privacy protection standards, such as the above-mentioned K-anonymity, l-diversity, t-closeness and other indicators. During the monitoring process, if the desensitization effect does not meet expectations, adjust the model training parameters and return to the step of training the model.
[0138] refer to Figure 2 , Figure 2 An optional implementation process of intelligent connected vehicle data desensitization based on machine learning provided by an embodiment of the present invention is illustrated as follows:
[0139] Step 1: Collect raw vehicle information data through active collection of on-board equipment and passive collection in the cloud.
[0140] Step 2: Use a recognition rule-based method to identify and process sensitive information in the original data to obtain sensitive data.
[0141] Step 3: Develop a personalized desensitization strategy based on the sensitivity and usage scenarios of sensitive data.
[0142] Step 4: Train the desensitization model
[0143] 1) Build an initial dataset containing the original data and its corresponding real desensitized text;
[0144] 2) Preprocess the initial data set to obtain the target data set;
[0145] 3) Input the target dataset containing sensitive information into the encoder, use BERT to understand the context of the input sensitive data, and process the input sequence through the encoder time step by time. The encoder outputs a vector at each time step, which together constitute the intermediate representation of the input text.
[0146] 4) The decoder uses the same Transformer architecture as the encoder and generates predictive desensitized text based on the encoder’s intermediate representation;
[0147] 5) During the training process, the cross entropy loss function is used to measure the difference between the predicted desensitized text and the real desensitized text. In the early stage of training, the teacher forcing technique is used, that is, the next word of the real desensitized text is directly used as the input of the decoder to stabilize the training process, and the model parameters are updated through back propagation and gradient descent algorithms to minimize the loss function;
[0148] 6) Evaluate the performance of the model on an independent validation set to ensure that the model is not overfitting, and adjust the model's hyperparameters such as learning rate, hidden layer size, number of attention heads, etc. based on the performance on the validation set.
[0149] Step 5: Desensitize the data using the target data desensitization model obtained after training, and evaluate the desensitized data using evaluation indicators (including but not limited to K-anonymity, l-diversity, and t-closeness). If the desensitization effect is poor, adjust the training parameters of the model and return to step 4. If the desensitization effect is good, proceed to the next step.
[0150] Step 6: Deploy and monitor the target data desensitization model, that is, encapsulate the trained target data desensitization model so that it can be easily integrated into the Internet of Vehicles platform, and monitor the system performance indicators in real time. If it is monitored that the desensitization effect does not meet expectations, adjust the training parameters of the model and return to step 4.
[0151] It should be noted that in each specific embodiment of the present invention, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0152] The embodiment of the present invention further provides a device for desensitizing data of a smart connected vehicle based on machine learning, which can implement the above-mentioned method for desensitizing data of a smart connected vehicle based on machine learning. The device includes:
[0153] The first module is used to collect raw data of vehicle information;
[0154] The second module is used to identify the sensitive information of the original data based on preset rules to obtain sensitive data;
[0155] The third module is used to build a desensitization strategy based on the sensitive data;
[0156] The fourth module is used to construct a target data set according to the desensitization strategy;
[0157] The fifth module is used to train the initial data desensitization model according to the target data set to obtain a target data desensitization model;
[0158] Among them, the target data desensitization model is used to desensitize the intelligent connected vehicle data.
[0159] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0160] The embodiment of the present invention further provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the above-mentioned method for desensitizing data of a smart connected vehicle based on machine learning is implemented. The electronic device can be any smart terminal including a tablet computer, a car computer, etc.
[0161] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0162] refer to Figure 3 , Figure 3 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0163] The processor 801 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0164] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store an operating system and other applications. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 802, and the processor 801 calls and executes a method for desensitizing data of a smart connected vehicle based on machine learning in an embodiment of the present invention;
[0165] Input / output interface 803, used to implement information input and output;
[0166] The communication interface 804 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WI FI, Bluetooth, etc.);
[0167] A bus 805 that transmits information between the various components of the device (e.g., the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);
[0168] The processor 801 , the memory 802 , the input / output interface 803 and the communication interface 804 are connected to each other in communication within the device via a bus 805 .
[0169] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for desensitizing data of intelligent connected vehicles based on machine learning.
[0170] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0171] The embodiment of the present invention further provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the aforementioned method for desensitizing data of an intelligent connected vehicle based on machine learning.
[0172] In summary, the method and device for desensitizing data of intelligent connected vehicles based on machine learning according to the embodiment of the present invention have the following advantages:
[0173] 1. The embodiment of the present invention uses a deep learning model to train a desensitization algorithm, so that the model can learn how to apply different desensitization strategies to different types of sensitive data. This method can use a large amount of training data to allow the model to learn complex desensitization rules and improve the intelligence level of desensitization.
[0174] 2. The embodiment of the present invention uses evaluation indicators such as K-anonymity, l-diversity and t-closeness to quantify the desensitization effect. This quantitative evaluation method can provide a more objective evaluation of the desensitization effect and help further optimize the desensitization strategy.
[0175] 3. The embodiment of the present invention deploys the trained desensitization model to the Internet of Vehicles platform and monitors the performance and effect of the desensitization system in real time. This real-time monitoring mechanism can ensure the continuity and consistency of data desensitization and timely discover and handle potential problems.
[0176] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0177] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0178] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0179] The logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, apparatus or device and execute the instructions), or in conjunction with such instruction execution system, apparatus or device. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, apparatus or device, or in conjunction with such instruction execution system, apparatus or device.
[0180] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0181] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0182] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0183] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0184] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.< / end> < / start>
Claims
1. A method for desensitizing data of intelligent connected vehicles based on machine learning, characterized in that: The following steps are involved: Collect raw data of vehicle information; Based on preset rules, sensitive information of the original data is identified to obtain sensitive data; Constructing a desensitization strategy based on the sensitive data; According to the desensitization strategy, construct a target data set; According to the target data set, the initial data desensitization model is trained to obtain a target data desensitization model; Among them, the target data desensitization model is used to desensitize the intelligent connected vehicle data.
2. The method for desensitizing data of intelligent connected vehicles based on machine learning according to claim 1, characterized in that: The following steps are also included: Evaluate the target data desensitization model to obtain an evaluation result; If the evaluation result is that the desensitization does not meet the standard, then return to the step of training the initial data desensitization model according to the target data set to obtain the target data desensitization model; if the evaluation result is that the desensitization meets the standard, then deploy and monitor the target data desensitization model.
3. The method for desensitizing data of intelligent connected vehicles based on machine learning according to claim 1, characterized in that: The method of collecting raw data of vehicle information comprises the following steps: The original data of the vehicle information is obtained through active collection of on-board equipment and passive collection in the cloud.
4. The method for desensitizing data of intelligent connected vehicles based on machine learning according to claim 1, characterized in that: The step of identifying the sensitive information of the original data based on the preset rules to obtain the sensitive data includes the following steps: Presetting the data type of the sensitive data; Presetting identification rules corresponding to each of the data types; According to the identification rule, the sensitive information of the original data is identified to obtain the sensitive data.
5. The method for desensitizing data of intelligent connected vehicles based on machine learning according to claim 1, characterized in that: The desensitization strategy is constructed according to the sensitive data, including the following steps: Obtaining the sensitivity and application scenarios of the sensitive data; Constructing the desensitization strategy according to the sensitivity and the application scenario; Among them, the desensitization strategy includes data replacement, data masking, data encryption and data perturbation.
6. The method for desensitizing data of intelligent connected vehicles based on machine learning according to claim 1, characterized in that: According to the desensitization strategy, constructing a target data set includes the following steps: Desensitizing the original data according to the desensitization strategy to obtain a first desensitized text; Obtaining an initial data set according to the original data and the corresponding first desensitized text; The initial data set is preprocessed to obtain the target data set.
7. The method for desensitizing data of intelligent connected vehicles based on machine learning according to claim 6, characterized in that: The initial data desensitization model is trained according to the target data set to obtain a target data desensitization model, comprising the following steps: Inputting the target data set into the encoder of the initial data desensitization model to obtain an intermediate vector; Inputting the intermediate vector into the decoder of the initial data desensitization model to obtain a second desensitized text; Constructing a cross entropy loss function according to the first desensitized text and the second desensitized text; Through the teacher forcing technology, the next word of the first desensitized text is input into the decoder, and the initial data desensitization model is trained in combination with the cross entropy loss function to obtain the target data desensitization model.
8. A data desensitization device for intelligent connected vehicles based on machine learning, characterized in that: include: The first module is used to collect raw data of vehicle information; The second module is used to identify the sensitive information of the original data based on preset rules to obtain sensitive data; The third module is used to build a desensitization strategy based on the sensitive data; The fourth module is used to construct a target data set according to the desensitization strategy; The fifth module is used to train the initial data desensitization model according to the target data set to obtain a target data desensitization model; Among them, the target data desensitization model is used to desensitize the intelligent connected vehicle data.
9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent multi-modal data desensitization method and device in vertical field
CN111625858A
Data desensitization method and device for Internet of Vehicles data, equipment and medium
CN117150549A
Multi-party privacy data joint pixel level labeling method and system based on federal learning
CN117409270A
Data personalized desensitization method and system
CN117744150A
Vehicle data desensitization detection method, device and equipment and storage medium
CN117874824A
Cited By
De-identification deep packet inspection and compliance verification method and device
CN120812156A