Vehicle log data desensitization method and vehicle log data processing system

By performing migration training and incremental training on the pre-trained BERT model, the problem of inefficient desensitization processing of vehicle log data is solved, efficient and accurate identification and desensitization of sensitive data is achieved, and the risk of sensitive data leakage is reduced.

CN120429893APending Publication Date: 2025-08-05ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510582490.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the prior art, the desensitization processing of vehicle log data is inefficient and there is a risk of omission, especially in manual code modification and regular matching methods, it is difficult to achieve centralized and unified management.

Method used

A sensitive data recognition model based on a pre-trained bidirectional coded transformer model (BERT) is used for transfer training and incremental training. Combined with supervised learning, adversarial training and data augmentation, it identifies and classifies sensitive data in vehicle log data, and desensitizes according to preset strategies.

Benefits of technology

It reduces the investment in manually identifying sensitive data, improves the efficiency and accuracy of desensitization processing, avoids the omission of sensitive data, and realizes centralized and unified desensitization management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429893A_ABST
    Figure CN120429893A_ABST
Patent Text Reader

Abstract

The invention relates to a vehicle log data desensitization method and a vehicle log data processing system, and the method comprises the steps: carrying out the sensitive data recognition of target vehicle log data based on a trained sensitive data recognition model, and obtaining a sensitive data classification result for the target vehicle log data; wherein the sensitive data identification model is obtained by performing migration training and incremental training on a pre-trained bidirectional coding converter model based on a preset log database; and desensitizing the target vehicle log data according to the sensitive data classification result and a preset desensitization strategy to obtain target desensitized data. According to the method, sensitive data in the vehicle log data can be identified and classified based on the natural language large model, so that the labor input for privacy data identification is reduced, the risk of sensitive data leakage is reduced, the efficiency and accuracy of vehicle log data desensitization processing are improved, and the missing of the sensitive data is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle log data processing, and in particular to a vehicle log data desensitization method and a vehicle log data processing system. Background Art

[0002] With the development of intelligent and connected vehicles, the entertainment functions of in-car cabins are becoming increasingly rich, which means that some key electronic control units (ECUs) of the vehicle, such as the entertainment host system (IVI) and the vehicle communication module (T-box), are able to obtain and store more and more information. This information may include a lot of sensitive information related to consumers' personal data and the vehicle itself. This information is stored in the vehicle's system memory in the form of system logs (for ease of description, referred to as vehicle log data below), which poses a risk of leakage.

[0003] Currently, during the development of key ECU applications, domestic vehicle OEMs typically require various applications (APPs) to write logs to memory using a unified logging class, printing them during development and debugging for troubleshooting. To desensitize vehicle log data, developers often manually modify code to identify and desensitize sensitive data written by the corresponding module at the source. These developers then independently configure sensitive data identification and processing logic for each module at the source. This decentralized management of each module consumes significant human resources for the entire system, preventing centralized management and making it prone to omissions. Furthermore, the log data printer, independently managed by a non-APP developer, uses regular expression matching to identify private data, perform targeted desensitization, and output the desensitized logs. This regular expression matching approach relies on a large number of predefined log formats and matching rules. In actual development, it is extremely difficult to exhaustively consider all possible log formats, resulting in a significant log desensitization workload and a high risk of missing sensitive data.

[0004] Therefore, the current desensitization processing of vehicle log data still has problems such as low processing efficiency and omission risks that need to be solved. Summary of the Invention

[0005] In this embodiment, a vehicle log data desensitization method and a vehicle log data processing system are provided to solve the problems of low efficiency and omission risk in the desensitization processing of vehicle log data in related technologies.

[0006] First, in this embodiment, a method for desensitizing vehicle log data is provided, including:

[0007] Obtain target vehicle log data;

[0008] Based on the trained sensitive data identification model, sensitive data identification is performed on the target vehicle log data to obtain a sensitive data classification result for the target vehicle log data; wherein the sensitive data identification model is obtained by performing transfer training and incremental training on a pre-trained bidirectional transcoder model based on a preset log database;

[0009] According to the sensitive data classification result and the preset desensitization strategy, the target vehicle log data is desensitized to obtain target desensitized data.

[0010] In some embodiments, the training process of the sensitive data identification model includes:

[0011] Constructing a training set; the training set includes data labels for various types of sensitive data that have been annotated; wherein the data labels include: personal sensitive labels, location sensitive labels, electronic control unit sensitive labels, and vehicle sensitive labels;

[0012] Performing transfer training on the pre-trained bidirectional encoder transformer model according to the training set to obtain an initial recognition model;

[0013] Inputting a preset test set into the initial recognition model to classify sensitive data and obtain a test classification result;

[0014] According to the test classification result, the initial recognition model is incrementally trained to obtain the sensitive data recognition model.

[0015] In some embodiments, transfer training is performed on a pre-trained bidirectional transcoder model based on the training set to obtain an initial recognition model, including:

[0016] The data-augmented training set is used as the input of the bidirectional encoder transformer model, and the bidirectional encoder transformer model is trained based on the cross loss function and the adversarial loss function to obtain an initial recognition model; wherein, the output of the bidirectional encoder transformer model passes through a fully connected layer and then is connected to a normalization layer to output the predicted probability of the sensitive type to which each sample in the training set belongs.

[0017] In some embodiments, incremental training is performed on the initial recognition model based on the test classification result to obtain the sensitive data recognition model, including:

[0018] Annotating the false positive data and the missed negative data of the test classification results, and constructing the annotated test classification results into an incremental training set;

[0019] Incremental training is performed on the initial recognition model according to the incremental training set.

[0020] In some embodiments, constructing a training set includes:

[0021] Remove invalid data from the preset original data set, and remove duplicate data based on continuous text decomposition and nearest neighbor search to obtain a cleaned data set;

[0022] The cleaned data set is annotated with sensitive types to obtain various data labels.

[0023] In some embodiments, the original data set includes an Android intelligent log analysis general public data set and an electronic control unit system log history data set.

[0024] In some embodiments, the target vehicle log data is desensitized according to the sensitive data classification result and a preset desensitization strategy to obtain target desensitized data, including:

[0025] Desensitizing the personal sensitive data in the target vehicle log data based on invalidation processing;

[0026] Desensitizing electronic control unit sensitive data and vehicle sensitive data in the target vehicle log data based on encryption processing;

[0027] Desensitizing the location-sensitive data in the target vehicle log data based on data biasing processing;

[0028] Desensitizing the numerically sensitive data in the target vehicle log data based on data generalization.

[0029] In a second aspect, a vehicle log data processing system is provided in this embodiment, including a log debugging terminal and a log centralized processing module, wherein:

[0030] The log debugging terminal is used to call the log processing interface encapsulated by the log centralized processing module to initiate a log processing request;

[0031] The log centralized processing module receives target vehicle log data sent by the log source, executes the vehicle log data desensitization method described in the first aspect above, and outputs the target desensitized data to the log debugging end.

[0032] In a third aspect, a vehicle host is provided in this embodiment, comprising the vehicle log data processing system described in the second aspect.

[0033] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored. When the program is executed by a processor, the vehicle log data desensitization method described in the first aspect is implemented.

[0034] Compared with the related art, the present embodiment provides a vehicle log data desensitization method and a vehicle log data processing system. The vehicle log data desensitization method obtains target vehicle log data; based on the trained sensitive data identification model, the target vehicle log data is subjected to sensitive data identification to obtain sensitive data classification results for the target vehicle log data; wherein, the sensitive data identification model is obtained by performing migration training and incremental training on the pre-trained bidirectional encoder transformer model based on a preset log database; according to the sensitive data classification results and the preset desensitization strategy, the target vehicle log data is desensitized to obtain target desensitized data. It can identify and classify sensitive data in vehicle log data based on a large natural language model, thereby reducing the manual input for privacy data identification, reducing the risk of sensitive data leakage, improving the efficiency and accuracy of vehicle log data desensitization processing, and avoiding the omission of sensitive data.

[0035] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0037] Figure 1 This is a hardware structure block diagram of a terminal of the vehicle log data desensitization method of this embodiment;

[0038] Figure 2 is a flow chart of the vehicle log data desensitization method of this embodiment;

[0039] Figure 3 is a flowchart of a method for training a sensitive data recognition model provided by some embodiments;

[0040] Figure 4 Schematic diagram of the structure of the vehicle log data processing system of this embodiment;

[0041] Figure 5 The following is a schematic diagram of log desensitization processing interactions in some embodiments. DETAILED DESCRIPTION

[0042] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0043] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "the," "these," and similar expressions in this application do not denote limitations on quantity and may be singular or plural. The terms "comprise," "include," "have," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include unlisted steps or modules (units) or other steps or modules (units) inherent to the process, method, product, or device. The terms "connected," "connected," "coupled," and similar expressions used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used in this application, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone; A and B exist simultaneously; or B exists alone. Generally, the character " / " indicates that the objects in the preceding and following relationship are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0044] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 This is a hardware block diagram of the terminal of the vehicle log data desensitization method of this embodiment. Figure 1 As shown, the terminal may include one or more ( Figure 1 The processor 102 (only one is shown) and a memory 104 for storing data, wherein the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The terminal may also include a transmission device 106 for communication functions and an input / output device 108. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0045] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the vehicle log data desensitization method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0046] Transmission device 106 is used to receive or transmit data via a network. This network may include a wireless network provided by the terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0047] In this embodiment, a method for desensitizing vehicle log data is provided. Figure 2 This is a flow chart of the vehicle log data desensitization method of this embodiment. Figure 2 As shown, the process includes the following steps:

[0048] Step S210: Obtain target vehicle log data.

[0049] In electric vehicles, the ECU carries functions such as entertainment, networking, and data connectivity, and ECU systems typically host numerous applications and services. The target vehicle log data can specifically be historical data generated during the operation of the vehicle's internal ECU system, requiring desensitization. This data may include, for example, vehicle operating status data, sensor data, battery management system data, and autonomous driving system data.

[0050] Step S220, based on the trained sensitive data identification model, sensitive data identification is performed on the target vehicle log data to obtain a sensitive data classification result for the target vehicle log data; wherein, the sensitive data identification model is obtained by performing migration training and incremental training on the pre-trained bidirectional encoder transformer model based on a preset log database.

[0051] The target vehicle log data can be input into a trained sensitive data recognition model. The trained sensitive data recognition model then identifies sensitive data from the target vehicle log data and classifies it into sensitive categories, thereby obtaining a sensitive data classification result. The sensitive data classification result can specifically be a data label of a certain sensitive category corresponding to each data item in the target vehicle log data.

[0052] Specifically, the sensitive data recognition model described above can be implemented based on a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model. This model is then transfer-trained using a pre-built training set related to vehicle logs. After initial model training, the model is tested using new ECU system log data using a feedback loop optimization mechanism. The model is then incrementally trained based on the model's output. This results in a sensitive data recognition model that meets preset standards for accuracy and completeness after multiple rounds of optimization and incremental training.

[0053] Step S230: Desensitize the target vehicle log data according to the sensitive data classification result and the preset desensitization strategy to obtain target desensitized data.

[0054] After obtaining the sensitive data classification results of the target vehicle log data, each type of sensitive data in the target vehicle log data can be desensitized according to pre-set desensitization methods for different sensitivity types. For example, the desensitization methods may include anonymization, encryption, etc.

[0055] In related technologies, log data desensitization often relies on application developers to manage their own log content and desensitize sensitive information. This approach requires each developer to adopt different rules to match and modify sensitive data based on the format characteristics of their own logs. This requires manual code modifications, resulting in high human resource consumption, fragmented and inconsistent desensitization processing, and the risk of omissions.

[0056] In this embodiment, instead of requiring each developer to individually encapsulate a log processing class to process logs, a centralized desensitization processing interface can be provided based on the vehicle log data desensitization method provided in this embodiment. When printing logs, the encapsulated desensitization processing interface can be directly called, replacing the native log processing mechanism. This centralized processing of sensitive log information replaces the separate processing of sensitive log data by various applications in the ECU system, reducing the consumption of human resources, forming a centralized, unified and standardized log desensitization processing, and avoiding omissions. This not only reduces the workload of developers, but also reduces the difficulty of managing sensitive log data, and reduces the risk of sensitive data leakage from a systemic perspective.

[0057] In addition, the current processing method for vehicle log data in related technologies still relies on the printing end to identify sensitive log data based on regular matching methods. This regular matching-based method requires reliance on a large number of predefined log formats and matching rules. For actual development tasks, it is difficult to exhaust all possible formats of log printing content, which leads to a large workload for actual log desensitization and the risk of missing sensitive data processing.

[0058] To this end, this embodiment uses the BERT model to complete the log analysis task. This embodiment takes into account the BERT model's excellent performance in natural language understanding tasks and its suitability for tasks such as named entity recognition and text classification. Therefore, it is more suitable for understanding log content and identifying and labeling sensitive data. This embodiment specifically utilizes a training set related to vehicle logs to perform transfer training on a pre-trained BERT model. In some embodiments, the training process can incorporate supervised learning, adversarial training, and data augmentation to improve the model's performance and recognition accuracy. Ultimately, a sensitive data recognition model suitable for log analysis tasks is obtained. Based on its understanding of vehicle log data, this sensitive data recognition model can identify, classify, and label the sensitive data contained therein. Finally, the sensitive data classification results of the target vehicle log data, combined with the preset desensitization strategy, are used to perform corresponding desensitization processing on different types of sensitive data. This reduces the manual effort required to identify sensitive data, improves recognition efficiency and accuracy, enhances the compliance and usability of log data, and reduces the risk of omissions in sensitive data processing.

[0059] In steps S210 to S230, target vehicle log data is acquired; based on the trained sensitive data identification model, sensitive data is identified on the target vehicle log data to obtain sensitive data classification results for the target vehicle log data; wherein the sensitive data identification model is obtained by transfer training and incremental training of a pre-trained bidirectional transcoder model based on a preset log database; based on the sensitive data classification results and a preset desensitization strategy, the target vehicle log data is desensitized to obtain target desensitized data. This method can identify and classify sensitive data in vehicle log data based on a large natural language model, thereby reducing the manual effort required to identify private data, lowering the risk of sensitive data leakage, improving the efficiency and accuracy of vehicle log data desensitization processing, and avoiding the omission of sensitive data.

[0060] In one embodiment, the training process of the sensitive data identification model may specifically include:

[0061] Construct a training set; the training set includes data labels for various types of labeled sensitive data; among them, the data labels include: personal sensitive labels, location sensitive labels, electronic control unit sensitive labels, and vehicle sensitive labels; based on the training set, transfer training is performed on the pre-trained bidirectional encoder converter model to obtain an initial recognition model; a preset test set is input into the initial recognition model to classify sensitive data and obtain a test classification result; based on the test classification result, incremental training is performed on the initial recognition model to obtain a sensitive data recognition model.

[0062] When constructing the training set, sensitive information in each log in the original dataset can be categorized through manual labeling. For example, based on the log content that a vehicle ECU may contain, labels for sensitive information can be categorized into the following types: personal sensitive labels, location sensitive labels, electronic control unit sensitive labels, and vehicle sensitive labels. Sensitive data corresponding to personal sensitive labels can include personal identity information (e.g., personal ID information, raw biometric information), personal property information (e.g., financial account information, personal transaction information, personal asset information), and personal communication information (e.g., address books, text messages). Sensitive data corresponding to location sensitive labels can include location information such as GPS positioning information; sensitive data corresponding to electronic control unit sensitive labels can include the ECU hardware serial number (SN) and Internet Protocol (IP) address; and vehicle sensitive labels can include basic vehicle information, such as the vehicle identification number (VIN) and vehicle hardware and software information.

[0063] The BERT model is selected as the base model, and a pre-trained BERT model is loaded for transfer learning. The pre-training process can be based on a large-scale set of general language text data, training the BERT model to enable it to master general language knowledge, such as grammar, semantics, and common sense. The pre-trained BERT model is then loaded and transferred using the constructed training set to adapt the pre-trained BERT model to the system log analysis task. During the training process, supervised learning, adversarial training, and data augmentation methods are combined to improve the model's performance and recognition accuracy.

[0064] After completing the initial training of the BERT model, the model optimization phase begins. New ECU system log data is used to construct a test set based on the initial recognition model. The test set data is processed based on the initial recognition model to obtain test classification results. These test classification results are then compared with the actual results annotated by human reviewers. This allows adjustments to the initial recognition model's training strategy and alignment for incremental training to optimize model performance.

[0065] This embodiment can realize the migration training of the BERT model in the field of vehicle log data desensitization processing, thereby obtaining a sensitive data recognition model with the ability to accurately identify and classify sensitive data.

[0066] Specifically, in one embodiment, based on the training set, transfer training is performed on the pre-trained bidirectional transcoder model to obtain an initial recognition model, which may specifically include:

[0067] The data-augmented training set is used as the input of the bidirectional encoder transformer model. The bidirectional encoder transformer model is trained based on the cross loss function and the adversarial loss function to obtain the initial recognition model. The output of the bidirectional encoder transformer model passes through the fully connected layer and then connects to the normalization layer to output the predicted probability of the sensitive type to which each sample in the training set belongs.

[0068] In this embodiment, transfer training is performed on a pre-trained BERT model using supervised learning. Supervised learning aims to adjust a pre-trained natural language processing model through labeled data to better suit a specific task. During the supervised learning phase, this embodiment introduces cross-entropy loss to optimize and improve model performance. Furthermore, for the training of the sensitive data labeling task in this embodiment, the output of the BERT model passes through a fully connected layer followed by a normalization (softmax) layer to output a predicted probability for each sensitive data type. The data label corresponding to a particular sensitive data type can then be determined based on the predicted probability. Specifically, a cross-entropy function can be defined as the loss function used to calculate the loss measurement error. By using the Adaptive Moment Estimation optimizer (ADAM), the model parameters are adaptively and continuously adjusted to minimize the cross-entropy loss, thereby improving model performance.

[0069] Additionally, adversarial training can be introduced during training to improve model robustness and enhance anti-interference capabilities. First, adversarial samples are generated using the Fast Gradient Sign Method (FGSM), Project Gradient Descent (PGD), or Generative Adversarial Networks (GANs) algorithm and mixed with real samples for training. An adversarial loss function, such as Tradeoff-inspired Adversarial Defense (TRADES), is added to the loss function, and model parameters are dynamically adjusted to minimize the adversarial loss function, enhancing model robustness.

[0070] In particular, the data in the training set has undergone data augmentation. This data augmentation includes, but is not limited to, synonym replacement, random string replacement or deletion, and the addition of random noise. By performing data augmentation on the data in the training set, the data content in the training set can be enriched.

[0071] This embodiment can combine supervised learning, adversarial training, data enhancement and other methods to improve the performance and recognition accuracy of the model in the desensitization processing of vehicle log data.

[0072] Additionally, in one embodiment, incremental training is performed on the initial recognition model based on the test classification results to obtain a sensitive data recognition model, which may specifically include:

[0073] The false positive data and missed negative data of the test classification results are labeled, and the labeled test classification results are constructed as an incremental training set; the initial recognition model is incrementally trained based on the incremental training set.

[0074] In this embodiment, a feedback loop optimization mechanism is introduced. After initial model training is completed based on the aforementioned transfer training process, the initial recognition model obtained after initial training is used to identify new ECU system log data to obtain test classification results. Manual review can be used to identify and annotate the false positive and false negative logs in the test classification results output by the initial recognition model. Subsequently, the false positive and false negative log data, along with the annotations, are used as an incremental training set to perform incremental training on the initial recognition model. During the incremental training process, the training strategy is adjusted based on the model's output to optimize model performance.

[0075] Incremental training further refines and optimizes the recognition performance of the initial recognition model. Specifically, the trained initial recognition model is used to identify new log data. False positives and false negatives are identified through manual review and annotated. The manually reviewed data, including the false positive and false negative log entries, is then constructed into a new dataset. Using this newly constructed dataset to train the model, model parameters, learning strategies, and training methods can be continuously adjusted during actual verification to improve the model's recognition accuracy and reduce false positive and false negative rates.

[0076] In one embodiment, constructing a training set may include:

[0077] Invalid data is removed from the preset original data set, and duplicate data is removed based on continuous text decomposition and nearest neighbor search to obtain a cleaned data set; sensitive types are annotated on the cleaned data set to obtain various data labels.

[0078] The quality of the sensitive data identification model is strongly dependent on the quality of the training set. Therefore, the quality of data cleansing largely determines the effectiveness and reliability of the model. During the training set construction phase, data cleansing can be performed on the raw data set. Within the raw data set, log entries with missing values are identified as invalid and removed. Additionally, log content with duplicate data is deleted to improve data quality and usability. In particular, in some embodiments, given the characteristics of log data sets, duplicate data fields may comprise a high proportion. Therefore, during the data cleansing phase, statistical natural language processing (n-grams), similarity search (minhash), and local hashing (LSH)-sensitive methods can be used to deduplicate the data set, decompose continuous text, and perform efficient nearest neighbor searches to deduplicate large text volumes. This embodiment can improve the quality of the training set through pre-processed data cleansing, thereby improving the quality of the trained model.

[0079] Additionally, in one embodiment, the original data set may include: an Android smart log analysis general public data set and an electronic control unit system log history data set.

[0080] Data is read from the Android Intelligent Log Analysis Public Database (Loghub) and from the log history dataset corresponding to the electronic control unit, forming the original dataset. The log dataset maintained by Loghub totals over 77GB, including both production and laboratory data. Data types can be categorized as distributed system logs, supercomputer cluster logs, operating system logs, mobile application logs, server application logs, and standalone software logs. This embodiment uses Loghub as the data source for the training set, which covers the required data requirements and enables the trained model to identify different types of sensitive data.

[0081] In particular, in one embodiment, based on the sensitive data classification result and the preset desensitization strategy, the target vehicle log data is desensitized to obtain the target desensitized data, which may specifically include:

[0082] Personal sensitive data in the target vehicle log data is desensitized based on invalidation processing; electronic control unit sensitive data and vehicle sensitive data in the target vehicle log data are desensitized based on encryption processing; positioning sensitive data in the target vehicle log data is desensitized based on data biasing processing; numerical sensitive data in the target vehicle log data is desensitized based on data generalization.

[0083] Desensitization methods can be pre-set for different types of sensitive data based on actual application scenarios or experience. In this embodiment, considering that personal property, personal communication information, and other personal information must be protected and cannot be made public, log data identified as personal sensitive data is invalidated by replacing parts of the information such as names, ID numbers, contacts, and phone numbers with asterisks. For example, "Zhang San" is desensitized to "Zhang※".

[0084] Furthermore, symmetric encryption can be used to process sensitive ECU data, such as ECU device information and basic vehicle information, to ensure the reversibility of data desensitization, allowing decryption of this data when the original data is needed. For example, a symmetric encryption algorithm (AES) can be used to encrypt the original data to produce the desensitized data. Authorized personnel will then have access to the encryption algorithm's key and can decrypt the original data when needed, thereby improving data security.

[0085] In addition, for information such as geographic location, a biased method can be adopted to conceal the real geographic location and replace it with false geographic coordinates, thereby achieving the concealment and protection of real information while retaining the basic coordinate location format.

[0086] In addition, for some numerically sensitive data, such as user age, IP address, and physical address, a generalized rule can be adopted to use a range segment instead of a specific numerical value for desensitization. For example, a data field with an age of "28" can be desensitized to "20+." This embodiment uses different processing methods to desensitize different types of sensitive data, thereby improving data security while ensuring data availability and preventing the leakage of sensitive data.

[0087] In one embodiment, to identify and classify sensitive data in logs, a BERT model suitable for natural language processing is introduced. Targeted transfer training is performed on the pre-trained BERT model to obtain a sensitive data identification model for processing system logs. The training set is constructed using the LogHub, a general public dataset for Android intelligent log analysis, superimposed with the OEM's internal ECU system log historical dataset as the original dataset. The original dataset is cleaned to remove invalid log entries, and sensitive information in the logs is classified through manual labeling to obtain the training set. Based on the training set, the trained BERT model is selected as the base model for transfer learning, thereby improving model training performance under limited training data conditions and specifically adapting to log analysis tasks. Supervised learning, adversarial training, and data augmentation methods are combined to improve the model's performance and recognition accuracy, allowing the model to classify and label sensitive data based on its understanding of the log content. Based on the initially trained model, manual review and feedback loop optimization are used to statistically analyze the model's false positive and false negative rates, which serve as a basis for fine-tuning. Use new log data to identify the initial model, identify false positives and missed negatives through manual review and add annotations, use the manually reviewed data to build a new data set, and use it for incremental training of the model, while adjusting the training strategy to optimize the performance of the model. Finally, use the sensitive data recognition model obtained after training to identify and classify sensitive data in the target vehicle log data, and adopt corresponding desensitization strategies such as anonymization, desensitization, encryption, etc. to desensitize the data based on the sensitive data classification results. In this way, this embodiment realizes the centralized and unified processing of sensitive log data instead of the separate processing of sensitive data by various developers, eliminates the current manpower dependence of desensitization on developers to manually modify the code, reduces labor costs, reduces the workload of desensitization processing, improves the efficiency and accuracy of desensitization processing, and avoids the risk of missing sensitive data.

[0088] Figure 3 A flowchart of a method for training a sensitive data recognition model is provided for some of the embodiments. Figure 3 As shown, the training method of the sensitive data recognition model includes the following steps:

[0089] Step S301: Select an original data set, wherein the Android smart log analysis general public data set and the electronic control unit system log history data set can be selected as the original data set.

[0090] Step S302: Data cleaning is performed on the original data set, wherein invalid data and duplicate data can be removed from the original data set. The specific removal method can refer to the above embodiment and will not be repeated here.

[0091] Step S303: Select a base model, wherein the pre-trained BERT model can be used as the base model.

[0092] Step S304: Perform migration training on the model. In order to improve the performance of the model, the model can be trained using cross entropy loss and adversarial loss.

[0093] Step S305: Manually review the trained current recognition model. Specifically, a new data set can be used as input for the current recognition model to obtain the test classification results output by the current recognition model, from which missed and false positive log data can be determined.

[0094] Step S306 determines whether the current recognition model has met the preset training termination condition. If so, step S309 is executed; otherwise, step S307 is executed. For example, the accuracy reaches a preset threshold, or the number of training times reaches a preset number, etc. The specific conditions can be determined according to the needs of the actual application scenario.

[0095] Step S307: Construct an incremental training set based on the manual review results. Manual review is used to identify false positives and false negatives in the log data, add annotations, and construct an incremental data set. Execute step S308.

[0096] Step S308: Perform incremental training on the model based on the incremental training set. Execute step S304.

[0097] Step S309: output the sensitive data recognition model finally obtained by training.

[0098] The above steps S301 to S309 implement supervised training of the BERT model, which can improve the performance of the BERT model in identifying and classifying sensitive log data.

[0099] In this embodiment, a vehicle log data processing system is provided. Figure 4 FIG. 4 is a schematic diagram of the structure of the vehicle log data processing system 40 of this embodiment. Figure 4 As shown, the vehicle log data processing system 40 can specifically include a log debugging terminal 41 and a log centralized processing module 42, wherein: the log debugging terminal 41 is used to call the log processing interface encapsulated by the log centralized processing module 42 to initiate a log processing request; the log centralized processing module 42 receives the target vehicle log data sent by the log source, executes the vehicle log data desensitization method provided by any of the above embodiments, and outputs the target desensitized data to the log debugging terminal.

[0100] In the present embodiment, the desensitization processing of log data is centralized in the log centralized processing module 42, thereby realizing the centralized and unified processing of log sensitive data. When the log debugging terminal 41 needs to use the system log of the capture and viewing device or simulator, for example, when the command of the logcat class is used to output the log, the library used by the native log mechanism of the operating system is no longer directly called, but the log centralized processing module 42 that encapsulates the log desensitization processing is used to perform the desensitization processing and print the log data after the desensitization processing. Among them, the command interface for printing logs such as logcat can be re-encapsulated, and the original logcat call to the operating system library is changed to call the log centralized processing module of the present embodiment, and the log centralized processing module calls the operating system library to obtain log data, and outputs the log data after desensitization.

[0101] Figure 5 The following is a log desensitization processing interactive diagram for some of the embodiments. Figure 5 As shown, the example of an Android-based in-vehicle entertainment system (IVI) is used. When developers and debuggers need to output logs using commands like logcat, they no longer call the native Android log library. Instead, they call the centralized log processing module to implement the log output function. First, the debugger issues a command to print sensitive data, calling the encapsulated log processing interface (API) of the centralized log processing module. When the centralized log processing API is executed, it sends a log output request (log request) to the logd service. Upon receiving the request, logd sends the complete target vehicle log data to the centralized log processing module. After receiving the target vehicle log data, the centralized log processing module calls the trained sensitive data recognition model to identify and classify sensitive fields, generating a sensitive data classification result. The centralized log processing module then desensitizes the sensitive fields and finally prints the desensitized target data to the log debugging terminal.

[0102] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0103] It should be noted that, for specific examples in this embodiment, reference may be made to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.

[0104] This embodiment also provides a vehicle host, which includes the vehicle log data processing system provided in the above embodiment.

[0105] In addition, in conjunction with the vehicle log data desensitization method provided in the above embodiment, this embodiment may also provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, it implements any of the vehicle log data desensitization methods in the above embodiment.

[0106] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0107] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0108] Obviously, the accompanying drawings are merely examples or embodiments of the present application. A person skilled in the art can also apply the present application to other similar situations based on these drawings without inventive effort. Furthermore, it is understandable that, although the work involved in this development process may be complex and lengthy, certain design, manufacturing, or production changes based on the technical content disclosed in this application are merely routine technical means for a person skilled in the art and should not be considered to constitute a deficiency in the disclosure of the present application.

[0109] The term "embodiment" as used in this application refers to specific features, structures, or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily mean that the embodiment is the same, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is understood, either explicitly or implicitly, by those skilled in the art that the embodiments described in this application can be combined with other embodiments when there is no conflict.

[0110] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A vehicle log data desensitization method, characterized in that: include: Obtain target vehicle log data; Based on the trained sensitive data identification model, sensitive data identification is performed on the target vehicle log data to obtain a sensitive data classification result for the target vehicle log data; wherein the sensitive data identification model is obtained by performing transfer training and incremental training on a pre-trained bidirectional transcoder model based on a preset log database; According to the sensitive data classification result and the preset desensitization strategy, the target vehicle log data is desensitized to obtain target desensitized data.

2. The vehicle log data desensitization method according to claim 1, characterized in that: The training process of the sensitive data identification model includes: Constructing a training set; the training set includes data labels for various types of sensitive data that have been annotated; wherein the data labels include: personal sensitive labels, location sensitive labels, electronic control unit sensitive labels, and vehicle sensitive labels; Performing transfer training on the pre-trained bidirectional encoder transformer model according to the training set to obtain an initial recognition model; Inputting a preset test set into the initial recognition model to classify sensitive data and obtain a test classification result; According to the test classification result, the initial recognition model is incrementally trained to obtain the sensitive data recognition model.

3. The vehicle log data desensitization method according to claim 2, characterized in that: Based on the training set, transfer training is performed on the pre-trained bidirectional encoder transformer model to obtain an initial recognition model, including: The data-augmented training set is used as the input of the bidirectional encoder transformer model, and the bidirectional encoder transformer model is trained based on the cross loss function and the adversarial loss function to obtain an initial recognition model; wherein, the output of the bidirectional encoder transformer model passes through a fully connected layer and then is connected to a normalization layer to output the predicted probability of the sensitive type to which each sample in the training set belongs.

4. The vehicle log data desensitization method according to claim 2, characterized in that: Incrementally training the initial recognition model according to the test classification result to obtain the sensitive data recognition model, including: Annotating the false positive data and the missed negative data of the test classification results, and constructing the annotated test classification results into an incremental training set; Incremental training is performed on the initial recognition model according to the incremental training set.

5. The vehicle log data desensitization method according to claim 2, characterized in that: Construct a training set, including: Remove invalid data from the preset original data set, and remove duplicate data based on continuous text decomposition and nearest neighbor search to obtain a cleaned data set; The cleaned data set is annotated with sensitive types to obtain various data labels.

6. The vehicle log data desensitization method according to claim 5, characterized in that: The original data sets include the Android intelligent log analysis general public data set and the electronic control unit system log history data set.

7. The vehicle log data desensitization method according to any one of claims 1 to 6, characterized in that: According to the sensitive data classification result and the preset desensitization strategy, the target vehicle log data is desensitized to obtain target desensitized data, including: Desensitizing the personal sensitive data in the target vehicle log data based on invalidation processing; Desensitizing electronic control unit sensitive data and vehicle sensitive data in the target vehicle log data based on encryption processing; Desensitizing the location-sensitive data in the target vehicle log data based on data biasing processing; Desensitizing the numerically sensitive data in the target vehicle log data based on data generalization.

8. A vehicle log data processing system, characterized in that: It includes the log debugging terminal and the log centralized processing module, including: The log debugging terminal is used to call the log processing interface encapsulated by the log centralized processing module to initiate a log processing request; The log centralized processing module receives target vehicle log data sent by the log source, executes the vehicle log data desensitization method according to any one of claims 1 to 7, and outputs the target desensitized data to the log debugging terminal.

9. A vehicle host, characterized in that: The vehicle log data processing system includes the vehicle log data processing system according to claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the vehicle log data desensitization method according to any one of claims 1 to 7 are implemented.