A sensitive information processing method and related device
By matching the pending text in the training data with the target training data, trace the original data of the target and generate an impact range report, the problem of difficult identification of sensitive information in the training data is solved, and model performance and data security are guaranteed.
Patent Information
- Application Number
- CN202510287266.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-12
AI Technical Summary
In large-scale data processing, sensitive information implicit in training data is difficult to be identified and filtered in time, resulting in sensitive information contained in model generation results, affecting model performance and may cause social impact.
By obtaining the pending text containing sensitive information generated by the model, matching with the target training data, tracing the corresponding target original data, and generating an impact range report of the sensitive information to indicate the associated training version of the model and preprocessing process.
It realizes traceability and impact investigation of sensitive information in the model generation results, ensures model performance, and improves the security and compliance of training data.
Smart Images

Figure CN119808162B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a sensitive information processing method and related devices. Background Art
[0002] In the wave of big data and artificial intelligence in the information age, large models (such as deep learning models, generative adversarial networks, etc.) have demonstrated powerful capabilities in image recognition, natural language processing, recommendation systems, etc. However, the success of large models depends on massive training data, and the quality of this data directly affects the performance and reliability of the model. In actual operations, training data usually comes from the Internet or other large data sources, which inevitably contain some sensitive information. The presence of sensitive information not only weakens the accuracy of the model, but may also cause serious problems. For example, extreme speech content contained in training data may cause the model to generate extreme content, causing adverse social impacts. Therefore, it is particularly necessary to avoid sensitive information in training data.
[0003] The current methods of avoiding sensitive information in training data have made significant progress in large-scale data processing and sensitive information filtering. Advanced machine learning models and filtering technologies can effectively identify and filter out obvious sensitive information in most cases, thereby improving the security of training data. However, due to the large scale and high complexity of training data, it is almost impossible to completely avoid sensitive information in training data. Some obscure expressions or newly emerging sensitive content cannot be identified in time, resulting in sensitive information being included in the model generation results, which not only affects the performance of the model, but may also cause unforeseen problems.
[0004] Therefore, how to provide a sensitive information processing method to ensure the performance of the model has become an urgent problem to be solved by those skilled in the art. Summary of the invention
[0005] In view of the above problems, this application provides a sensitive information processing method and related devices to achieve the purpose of tracing the source and investigating the impact of sensitive information when sensitive information is included in the model generation results, thereby ensuring the performance of the model. The specific scheme is as follows:
[0006] The first aspect of the present application provides a sensitive information processing method, comprising:
[0007] Obtaining a text to be processed; the text to be processed is a text containing sensitive information generated by a model of a target training version;
[0008] Matching the text to be processed with the training data of the target training version to obtain target training data, wherein the target training data is the training data in the training data of the target training version that matches the text to be processed;
[0009] Obtaining corresponding target original data based on the meta information of the target training data, wherein the target original data and the corresponding target training data contain the same meta information;
[0010] Based on the target original data, an impact scope report of the sensitive information is generated, where the impact scope report is used to indicate a training version model associated with the target original data and a preprocessing flow of the target original data.
[0011] In a possible implementation, the method for determining the training data of the target training version includes:
[0012] Determining a data processing version associated with the target training version;
[0013] Determine the target original data corresponding to the data processing version, each target original data carries the meta information of the target original data;
[0014] Preprocessing the target original data corresponding to the data processing version to obtain preprocessed data corresponding to the data processing version;
[0015] Based on the preprocessed data corresponding to the data processing version, the training data of the target training version is determined.
[0016] In a possible implementation, determining the target original data corresponding to the data processing version includes:
[0017] Obtaining pre-constructed raw data management information, wherein the raw data management information includes multiple data sets, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of raw data;
[0018] Determining a target data set and a target data batch in the target data set from the original data management information;
[0019] According to the storage directory of the original data corresponding to each of the target data batches, the original data corresponding to each of the target data batches is obtained as the target original data corresponding to the data processing version.
[0020] In a possible implementation, the original data management information is constructed as follows:
[0021] Determine the storage directory for the raw data required for model training;
[0022] Dividing the original data required for the model training to obtain at least one data set, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of original data;
[0023] For each data batch, directory scanning, directory authorization and data backup operations are performed on the data batch to obtain the original data management information.
[0024] In a possible implementation, preprocessing the target original data corresponding to the data processing version to obtain the preprocessed data corresponding to the data processing version includes:
[0025] Cleaning the target original data corresponding to the data processing version to obtain cleaned data corresponding to the data processing version;
[0026] Determine the cleaned data corresponding to the data processing version as the pre-processed data corresponding to the data processing version;
[0027] Alternatively, further manually reviewing the cleaned data corresponding to the data processing version to obtain manually reviewed data corresponding to the data processing version;
[0028] The data corresponding to the data processing version and having passed manual review is determined as the pre-processed data corresponding to the data processing version.
[0029] In a possible implementation, the cleaning the target original data corresponding to the data processing version to obtain the cleaned data corresponding to the data processing version includes:
[0030] For each of the target data batches, the original data corresponding to the target data batch is cleaned to obtain a new data batch; each new data batch is the cleaned data corresponding to the data processing version.
[0031] In a possible implementation, determining the training data of the target training version based on the preprocessed data corresponding to the data processing version includes:
[0032] Using the preprocessed data corresponding to the data processing version as training data of the target training version;
[0033] Alternatively, determine the target sampling version;
[0034] Sampling is performed from the preprocessed data corresponding to the data processing version to obtain training data corresponding to the target sampling version as training data for the target training version.
[0035] In a possible implementation, matching the text to be processed with the training data of the target training version to obtain the target training data includes:
[0036] Determine a sensitive word list based on the text to be processed; the sensitive word list contains at least one sensitive word;
[0037] The sensitive word list is matched with the training data of the target training version to obtain target training data.
[0038] In a possible implementation, determining a sensitive word list based on the text to be processed includes:
[0039] Determining sensitive words in the text to be processed;
[0040] Determine a vocabulary consisting of sensitive words in the text to be processed as a sensitive word vocabulary;
[0041] or,
[0042] Obtaining related words corresponding to the sensitive words in the text to be processed, wherein the related words include any one or more of synonyms, antonyms, and words with similar forms;
[0043] A vocabulary consisting of sensitive words in the text to be processed and the related words is determined as a sensitive word vocabulary.
[0044] A second aspect of the present application provides a sensitive information processing device, including:
[0045] An acquisition unit, used for acquiring a text to be processed; the text to be processed is a text containing sensitive information generated by a model of a target training version;
[0046] A matching unit, used for matching the text to be processed with the training data of the target training version to obtain target training data, wherein the target training data is the training data in the training data of the target training version that matches the text to be processed;
[0047] A tracing unit, configured to obtain corresponding target original data based on the meta information of the target training data, wherein the target original data and the corresponding target training data contain the same meta information;
[0048] An impact scope assessment unit is used to generate an impact scope report of the sensitive information based on the target original data, wherein the impact scope report is used to indicate a training version model associated with the target original data and a preprocessing process of the target original data.
[0049] In a possible implementation, the apparatus further includes a training data determination unit for a target training version, and the training data determination unit for the target training version includes:
[0050] A data processing version determining unit, used to determine the data processing version associated with the target training version;
[0051] A target original data determination unit, used to determine the target original data corresponding to the data processing version, each target original data carries the meta information of the target original data;
[0052] A preprocessing unit, used to preprocess the target original data corresponding to the data processing version to obtain preprocessed data corresponding to the data processing version;
[0053] The training data determination unit is used to determine the training data of the target training version based on the preprocessed data corresponding to the data processing version.
[0054] In a possible implementation, the target original data determination unit is specifically configured to:
[0055] Obtaining pre-constructed raw data management information, wherein the raw data management information includes multiple data sets, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of raw data;
[0056] Determining a target data set and a target data batch in the target data set from the original data management information;
[0057] According to the storage directory of the original data corresponding to each of the target data batches, the original data corresponding to each of the target data batches is obtained as the target original data corresponding to the data processing version.
[0058] In a possible implementation, the device further includes: an original data management information construction unit, wherein the original data management information construction unit is specifically configured to:
[0059] Determine the storage directory for the raw data required for model training;
[0060] Dividing the original data required for the model training to obtain at least one data set, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of original data;
[0061] For each data batch, directory scanning, directory authorization and data backup operations are performed on the data batch to obtain the original data management information.
[0062] In a possible implementation, the preprocessing unit is specifically used to:
[0063] Cleaning the target original data corresponding to the data processing version to obtain cleaned data corresponding to the data processing version;
[0064] Determine the cleaned data corresponding to the data processing version as the pre-processed data corresponding to the data processing version;
[0065] Alternatively, further manually reviewing the cleaned data corresponding to the data processing version to obtain manually reviewed data corresponding to the data processing version;
[0066] The data corresponding to the data processing version and having passed manual review is determined as the pre-processed data corresponding to the data processing version.
[0067] In a possible implementation, the preprocessing unit is specifically used to:
[0068] For each of the target data batches, the original data corresponding to the target data batch is cleaned to obtain a new data batch; each new data batch is the cleaned data corresponding to the data processing version.
[0069] In a possible implementation, the training data determination unit is specifically configured to:
[0070] Using the preprocessed data corresponding to the data processing version as training data of the target training version;
[0071] Alternatively, determine the target sampling version;
[0072] Sampling is performed from the preprocessed data corresponding to the data processing version to obtain training data corresponding to the target sampling version as training data for the target training version.
[0073] In a possible implementation, the matching unit includes:
[0074] A sensitive word list determining unit, configured to determine a sensitive word list based on the text to be processed; the sensitive word list includes at least one sensitive word;
[0075] The matching subunit is used to match the sensitive word list with the training data of the target training version to obtain target training data.
[0076] In a possible implementation, the sensitive word list determination unit is specifically used to:
[0077] Determining sensitive words in the text to be processed;
[0078] Determine a vocabulary consisting of sensitive words in the text to be processed as a sensitive word vocabulary;
[0079] or,
[0080] Obtaining related words corresponding to the sensitive words in the text to be processed, wherein the related words include any one or more of synonyms, antonyms, and words with similar forms;
[0081] A vocabulary consisting of sensitive words in the text to be processed and the related words is determined as a sensitive word vocabulary.
[0082] The third aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the sensitive information processing method of the first aspect or any implementation method of the first aspect.
[0083] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0084] The memory is used to store computer programs;
[0085] The processor is used to execute the computer program so that the electronic device can implement the sensitive information processing method of the above-mentioned first aspect or any implementation method of the first aspect.
[0086] A fifth aspect of the present application provides a computer-readable storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the sensitive information processing method of the above-mentioned first aspect or any implementation of the first aspect.
[0087] By means of the above technical solution, the present application provides a sensitive information processing method and related device, after obtaining the to-be-processed text containing sensitive information generated by the model of the target training version; matching the to-be-processed text with the training data of the target training version, obtaining the target training data matching the to-be-processed text in the training data of the target training version, and tracing the target original data containing the same meta-information as the target training data based on the meta-information of the target training data, and finally generating a sensitive information impact range report based on the target original data to assist relevant personnel in remedying the training version model associated with the target original data and the pre-processing process of the target original data. This solution can trace the sensitive information and investigate the impact when the model generation result contains sensitive information, thereby ensuring the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0089] Figure 1 A flowchart of a sensitive information processing method provided in an embodiment of the present application;
[0090] Figure 2 A flowchart of a method for constructing raw data management information provided in an embodiment of the present application;
[0091] Figure 3 A flowchart of a method for determining training data of a target training version provided in an embodiment of the present application;
[0092] Figure 4 A flowchart of a method for matching a text to be processed with training data of a target training version to obtain target training data provided in an embodiment of the present application;
[0093] Figure 5 A schematic diagram of the structure of a sensitive information processing device provided in an embodiment of the present application;
[0094] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0095] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation mode of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0096] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0097] In the wave of big data and artificial intelligence in the information age, large models (such as deep learning models, generative adversarial networks, etc.) have demonstrated powerful capabilities in image recognition, natural language processing, recommendation systems, etc. However, the success of large models depends on massive training data, and the quality of this data directly affects the performance and reliability of the model. In actual operations, training data usually comes from the Internet or other large data sources, which inevitably contain some sensitive information. The presence of sensitive information not only weakens the accuracy of the model, but may also cause serious problems. For example, extreme speech contained in training data may cause the model to generate extreme content, causing adverse social impact. Therefore, it is particularly necessary to avoid sensitive information in training data.
[0098] Currently, the methods for avoiding sensitive information in training data can be divided into the following steps:
[0099] Step 1: Data collection;
[0100] Raw data is collected from various data sources (e.g., open datasets, internet scraping, internal data warehouses, etc.).
[0101] Step 2: Data preprocessing;
[0102] Perform preliminary cleaning of the raw data, such as removing null values, correcting data formats, processing noisy data, etc.
[0103] Step 3: Data cleaning;
[0104] The quality issues of the preprocessed data are further processed, including removing duplicate data, correcting outliers, normalizing data values, etc., to ensure the consistency and accuracy of the data.
[0105] Step 4: Filter sensitive information;
[0106] Use a variety of filtering techniques (such as regular expression matching, dictionary filtering, machine learning classifiers, pre-trained sensitive information detection models, etc.) to screen and remove sensitive information.
[0107] Step 5: Data review;
[0108] Review the data after sensitive information is filtered to ensure that the data quality meets the standards and complies with relevant laws and ethical standards.
[0109] Step 6: Feedback and correction;
[0110] Based on the audit results, if there are any problems with the data, return to the sensitive information filtering or data cleansing stage for correction.
[0111] Current methods for avoiding sensitive information in training data have made significant progress in large-scale data processing and sensitive information filtering. Advanced machine learning models and filtering technologies can effectively identify and filter out obvious sensitive information in most cases, thereby improving the security and compliance of training data.
[0112] However, due to the large scale and high complexity of training data, it is almost impossible to completely avoid sensitive information in training data. Certain obscure expressions or newly emerging sensitive content cannot be identified in a timely manner, resulting in sensitive information being included in the model generation results. This not only affects the effectiveness of the model, but may also cause unforeseen problems.
[0113] Therefore, how to provide a sensitive information processing method to achieve the traceability and impact investigation of sensitive information, and then ensure the effectiveness of the model, has become an urgent problem to be solved by technical personnel in this field.
[0114] In order to solve the above problems, the present application embodiment provides a sensitive information processing method. The training data processing method of the present application embodiment is described in detail below with reference to the accompanying drawings.
[0115] Reference Figure 1 , Figure 1 A flowchart of a sensitive information processing method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, a sensitive information processing method provided in an embodiment of the present application may include the following steps, and these steps are described in detail below.
[0116] S101: Obtain a text to be processed; the text to be processed is a text containing sensitive information generated by a model of a target training version;
[0117] The current methods of avoiding sensitive information in training data have made significant progress in large-scale data processing and sensitive information filtering. Advanced machine learning models and filtering technologies can effectively identify and filter out obvious sensitive information in most cases, thereby improving the security of training data. However, due to the large scale and high complexity of training data, it is almost impossible to completely avoid sensitive information in training data. Therefore, sensitive information may still be found when evaluating the model training effect.
[0118] In a possible implementation, the text to be processed may be text containing sensitive information generated by the model corresponding to the target training version during the training effect evaluation. The text to be processed may be a word-level text or a sentence-level text. The word-level text may contain one or more sensitive words, and the sentence-level text may contain one or more sentences. This application does not impose any limitation on this.
[0119] S102: Matching the text to be processed with the training data of the target training version to obtain target training data, where the target training data is the training data in the target training version that matches the text to be processed;
[0120] In this application, the training data of the target training version is the training data used to train the model of the target training version. The target training data is data that is used for model training and has not been effectively preprocessed, which has an impact on the performance of the model and causes the model to generate text containing sensitive information.
[0121] S103: tracing the source of the meta information of the target training data to obtain the corresponding target original data, wherein the target original data and the corresponding target training data contain the same meta information;
[0122] In order to ensure the relevance of a data in the entire model training process, in this application, for each original data, the metadata of the original data can be generated, and the metadata of the original data can identify the uniqueness of the data. The metadata of the original data will run through the processing of the original data in the entire model training process, which is convenient for data traceability. In one achievable method, the metadata of the original data contains data fingerprint information; for ease of understanding, this application provides an example of metadata in json format, as follows:
[0123] {
[0124] "meta": {
[0125] "title": "file name or web page title",
[0126] "path": "file path or web URL",
[0127] "id": "data fingerprint, i.e. hash summary of path"
[0128] }
[0129] }.
[0130] In this application, the source of each piece of data can be transparent and traceable based on the metadata. When problematic data is found, its source can be quickly located and targeted processing can be performed.
[0131] S104: Generate an impact scope report of the sensitive information based on the target original data, where the impact scope report is used to indicate a training version model associated with the target original data and a preprocessing flow of the target original data.
[0132] In the present application, after determining the target original data, the training version model associated with the target original data and the preprocessing process of the target original data can be obtained, and based on the training version model associated with the target original data and the preprocessing process of the target original data, an impact scope report of the sensitive information is generated to assist relevant personnel in taking corresponding remedial measures (such as retraining, adjusting parameters or deleting problematic data, etc.) to remedy the training version model associated with the target original data and the preprocessing process of the target original data.
[0133] The present embodiment provides a sensitive information processing method, after obtaining the to-be-processed text containing sensitive information generated by the model of the target training version; matching the to-be-processed text with the training data of the target training version, obtaining the target training data that matches the to-be-processed text in the training data of the target training version, and tracing the target original data containing the same meta-information as its meta-information based on the meta-information of the target training data, and finally generating a sensitive information impact range report based on the target original data to assist relevant personnel in remediating the training version model associated with the target original data and the pre-processing process of the target original data. This solution can trace the source and check the impact of sensitive information when sensitive information is included in the model generation results, so as to ensure the performance of the model. This method can significantly improve the security and reliability of the model training process, reduce the potential risks brought by sensitive information, ensure that the output results of the final model are compliant and credible, and promote the healthy development of artificial intelligence technology.
[0134] In the present application, the original data management information can be pre-constructed, so as to determine the training data of the target training version by using the pre-constructed original data management information. Next, the construction method of the original data management information is described in detail.
[0135] Reference Figure 2 , Figure 2 A flowchart of a method for constructing raw data management information provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the method comprises the following steps:
[0136] S201: Determine the storage directory of the original data required for model training;
[0137] In this application, the original data are mostly PDF, e-books or web pages, which are stored in advance.
[0138] Common storage methods include file storage and HDFS (Hadoop Distributed File System). In this application, the original data can be stored in advance using file storage or HDFS storage. The original data stored using file storage or HDFS storage can be directly managed, that is, the storage directory of the original data can be managed.
[0139] Among them, the advantages of file storage are as follows:
[0140] Efficient file access: File systems are usually more efficient at reading and writing small files.
[0141] File Hierarchy: File systems support directory and folder structures for easy organization and management of files.
[0142] POSIX compatibility: Many large model training frameworks (such as TensorFlow and PyTorch) natively support the POSIX file system, which facilitates direct reading and writing of data.
[0143] Low latency: Local or close-range file storage generally has lower access latency.
[0144] The advantages of HDFS storage are as follows:
[0145] High scalability: It was designed to handle large-scale data sets and distributed storage, and can be easily expanded to PB-level data volumes.
[0146] Data redundancy: Ensure high data availability and fault tolerance through replication mechanism.
[0147] Deep integration with the big data ecosystem: natively supports the Hadoop ecosystem (MapReduce, Spark, etc.) and is suitable for large-scale distributed computing tasks.
[0148] Improve efficiency by separating computing and storage: Combined with resource scheduling systems such as YARN, computing resources can be effectively scheduled and managed.
[0149] In the present application, determining the original data required for model training may specifically be determining a storage directory of the original data required for model training.
[0150] S202: Divide the original data required for the model training to obtain at least one data set, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of original data;
[0151] In this application, the raw data required for model training can be used to establish different data sets according to different data types or uses. Considering that processing the entire data set may cause insufficient memory and slow computing speed during model training, in this application, for each data set, the data set can be divided into several small blocks, each of which is called a data batch, and each data batch corresponds to a storage directory for a batch of raw data. Compared with processing the entire data set, processing smaller data batches can avoid insufficient memory problems and facilitate memory management. Moreover, smaller data batches can also be processed faster on computing devices (such as GPUs or TPUs) to speed up computing. Considering that there is a lot of data in a data batch, in this application, for a data batch, the data batch can also be divided into several folders, each folder containing at least one file, and each file corresponds to a raw data.
[0152] In the present application, a data management system may be pre-built, and permission levels of different users of the data management system may be configured.
[0153] The permission levels are divided into Owner permission level, Modify permission level, and ReadOnly permission level (read-only). Except for Owner permission level, Modify permission level and ReadOnly permission level both support validity period setting. Owner permission level and Modify permission level only open data set and data batch operation behaviors. Owner permission level is only open to data owners, Modify permission level is generally open to collaborators, and ReadOnly permission level is open to data users. Opening different permission levels for different users means determining whether the permission level of different users is Owner permission level, Modify permission level, or ReadOnly permission level (read-only).
[0154] In this application, users with Owner and Modify permissions can operate the data management system to divide the raw data required for model training to obtain at least one data set, and determine the data batches contained in the data set and the storage directory of the raw data corresponding to each data batch. These users can also edit data sets and data batches (for example, add data sets, delete data sets, add data batches, delete data batches, etc.).
[0155] S203: For each data batch, perform directory scanning, directory authorization and data backup operations on the data batch to obtain the original data management information.
[0156] In this application, after performing directory scanning, directory rights collection and data backup operations on each data batch, the original data management information can be obtained. The operations of directory scanning, directory rights collection and data backup are described in detail below, as follows:
[0157] First, perform directory scanning on the data batch, which means traversing the storage directory and folder of the original data corresponding to the data batch, and collecting batch information (such as batch size, number of folders, number of files in the folder, number of lines of text files, etc.).
[0158] Second, taking control of the directory for the data batch refers to setting permissions for the storage directory of the original data corresponding to the data batch.
[0159] Directory reclaiming is an important measure to improve the security of data management, especially during model training. Currently, most raw data for model training lacks effective management, and many cleaning scripts read data directly from the storage directory of the raw data, resulting in confusion in directory hierarchy and permission management. In this case, reclaiming all write permissions for the raw data and ensuring that the cleaned data is stored in a dedicated directory can greatly reduce the security risk of the raw data. By systematically reclaiming the write permissions for the raw data, the raw data can be effectively prevented from being accidentally or maliciously modified, ensuring the integrity and security of the raw data. At the same time, the cleaned data is stored in a new directory with clear permission settings, making management more orderly and secure. This not only reduces the risk of data leakage and tampering, but also provides a more reliable data foundation for model training. Overall, directory reclaiming is a key security measure that can significantly improve the robustness and security of data management.
[0160] In the present application, reclaiming the directory rights for the data batch means reclaiming all write permissions for the storage directory of the original data corresponding to the data batch. That is, users of three permission levels are only given read-only permissions for the storage directory of the original data corresponding to the data batch. Reclaiming the directory rights for the data batch can greatly reduce the security risk of the original data. By systematically reclaiming the write permissions for the storage directory of the original data corresponding to the data batch, the original data can be effectively prevented from being accidentally or maliciously modified, thereby ensuring the integrity and security of the original data.
[0161] In one feasible way, the directory rights of the data batch can be taken by combining the chmod permission and the ACL (Access Control Lists) permission, where:
[0162] The chmod permission is the system permission. The ACL permission is a set of rules that define which users or system processes have which permissions on file system objects (such as files and directories). ACL provides more fine-grained control than traditional chmod permission-based control. Starting with Hadoop 2.4.0, HDFS introduced support for POSIX ACLs, which enables HDFS to provide more fine-grained permission management.
[0163] In one feasible method, the chmod permission and ACL permission are combined to recover the directory rights of the data batch, including chmod permission recovery and ACL permission setting. Among them, the chmod permission recovery is as follows: establish a platform management user and user group, change the storage directory and folder owner to the platform management user and user group, change the storage directory permission to 750, tighten the entrance, and other users except the platform management user have no right to access the directory to ensure data security. At the same time, only the entrance directory is restricted, which greatly reduces the complexity of permission configuration operations; modify the folder and file permissions in the storage directory to 755, which are readable and executable by all users (directory cd permissions require x). ACL permission settings include granting storage directory ACL rx permissions to users with Owner permission level, and users with other permission levels (Modify permission level and ReadOnly permission level) will be granted ACL rx permissions to the storage directory when they are added. When the permission level of these users is deleted or expired, their ACL rx permissions will be automatically deleted.
[0164] Third, backing up the data batch means storing the original data of the data batch in the backup storage space to add security insurance for the data batch. In this application, the backup option of the data batch can be supported when the original data is managed.
[0165] In another embodiment of the present application, a method for determining the training data of the target training version is described in detail. Figure 3 , Figure 3 A flowchart of a method for determining training data of a target training version provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the method comprises the following steps:
[0166] S301: Determine a data processing version associated with the target training version;
[0167] In this application, in order to effectively manage the data cleaning process, corresponding data processing versions can be created at different stages of the cleaning process. Each data processing version can include multiple data processing tasks, which are respectively associated with multiple data sets and data batches as original data.
[0168] When training different types of models, such as general models and industry models, their data requirements are different. Therefore, it is necessary to reasonably plan the model training tasks before model training. When designing the training version, it is necessary to accurately grasp the data requirements and determine the data processing version associated with the training version.
[0169] S302: Determine the target original data corresponding to the data processing version, each target original data carries the meta information of the target original data;
[0170] In one achievable manner, determining the target original data corresponding to the data processing version includes: obtaining pre-constructed original data management information, the original data management information including multiple data sets, each data set including at least one data batch, and each data batch corresponding to a storage directory of a batch of original data; determining the target data set and the target data batch in the target data set from the original data management information; and obtaining the original data corresponding to each target data batch based on the storage directory of the original data corresponding to each target data batch as the target original data corresponding to the data processing version.
[0171] S303: preprocessing the target original data corresponding to the data processing version to obtain preprocessed data corresponding to the data processing version;
[0172] Considering that the raw data required for model training usually comes from the Internet or other large data sources, these data sources inevitably contain some sensitive information. Different types of data have their own unique preprocessing requirements. By performing targeted preprocessing on different types of data, data quality can be improved.
[0173] In a possible implementation, the preprocessing of the target original data corresponding to the data processing version to obtain the preprocessed data corresponding to the data processing version includes: cleaning the target original data corresponding to the data processing version to obtain the cleaned data corresponding to the data processing version; and determining the cleaned data corresponding to the data processing version as the preprocessed data corresponding to the data processing version.
[0174] In this application, automated tools can be used to clean the original data, and the specific implementation method will not be described in detail in this application.
[0175] The cleaning process is performed on the target original data corresponding to the data processing version to obtain the cleaned data corresponding to the data processing version, including: for each target data batch, the original data corresponding to the target data batch is cleaned to obtain a new data batch; each new data batch is the cleaned data corresponding to the data processing version.
[0176] That is, in this application, after the cleaning process, a new data set and data batch can be obtained. It should be noted that the new data set and data batch can be stored in a new directory with clear permission settings using a file storage method or an HDFS storage method.
[0177] It should be noted that in this application, automated tools can be used to clean the raw data. For example, a variety of filtering technologies can be introduced, including keyword matching, machine learning classifiers, and pre-trained models, to achieve multiple filtering and conduct strict screening of various sensitive information. Using advanced artificial intelligence and natural language processing technology, the data is contextually analyzed and the implicit meaning is intelligently screened, greatly improving the accuracy of detection. The specific implementation method will not be described in detail in this application.
[0178] Although automated tools can efficiently clean data, they may not be sufficient to capture all errors or complex contextual details. Therefore, manual review is still necessary after data cleaning, and manual review can provide further assurance in ensuring data accuracy and quality. Especially in some fields such as medicine or finance, data quality is crucial, and manual review can help identify and correct subtle errors or anomalies and ensure that data meets professional standards and legal regulations. Therefore, manual review is an important step to improve data consistency and reliability.
[0179] Therefore, in another possible implementation, preprocessing the target original data corresponding to the data processing version to obtain the preprocessed data corresponding to the data processing version includes:
[0180] The target original data corresponding to the data processing version is cleansed to obtain the cleansed data corresponding to the data processing version; the cleansed data corresponding to the data processing version is further manually reviewed to obtain the manually reviewed data corresponding to the data processing version; the manually reviewed data corresponding to the data processing version is determined as the pre-processed data corresponding to the data processing version.
[0181] In this application, the manual review of the cleaned data corresponding to the data processing version may be to sample the cleaned data according to the total amount of data and the number of samples of the data processing version, generate sample data, and manually review the sample data. The total amount of data of the data processing version is the sum of the number of text lines in the output file of the data processing version.
[0182] In a possible implementation, the total amount of data T may be calculated using the following formula:
[0183] .
[0184] Where T is the total amount of data, n is the number of files in the data processing version, and L is the number of text lines in a single file.
[0185] Since the single data processing version is basically for the same type of data, in this application, a simple random sampling algorithm without replacement can be used for sampling to generate sample data. In a possible implementation, the sampling probability formula for a single sample data is:
[0186] .
[0187] Among them, i is the number of sample data, is the sampling probability of the i-th sample data, and the average sampling probability P(1) of the sample data is:
[0188] .
[0189] The manual review results of sample data are either passed or failed. If the manual review results of all sample data are passed, the data of the data processing version can be used for model training. During the model training process, the corresponding data processing version of the data can be selected when creating the data sampling version.
[0190] If the manual review result of some sample data is failed, it is necessary to further determine whether the sample data is not cleaned thoroughly and needs to be re-cleaned or there are some non-compliant data in the original data and it cannot be cleaned. If the cleaning is not thorough and needs to be re-cleaned, it is only necessary to invalidate the data batch corresponding to the sample data. If there are some non-compliant data in the original data and it cannot be cleaned, the entire data set corresponding to the sample data will be invalidated.
[0191] S304: Determine the training data of the target training version based on the preprocessed data corresponding to the data processing version.
[0192] In one achievable manner, determining the training data of the target training version based on the preprocessed data corresponding to the data processing version includes: using the preprocessed data corresponding to the data processing version as the training data of the target training version.
[0193] Considering that the preprocessed data corresponding to the data processing version is huge, the efficiency of model training is greatly affected. Data sampling in model training is to select representative data from huge data, and its purpose is to ensure that the data used for training is diverse and extensive, thereby improving the generalization ability and performance of the model. Through reasonable data sampling, various data types can be balanced to avoid the model from being biased against data of certain specific categories or sources. This method helps to create a more robust and accurate model.
[0194] In another achievable manner, determining the training data of the target training version based on the preprocessed data corresponding to the data processing version includes: sampling from the preprocessed data corresponding to the data processing version to obtain the training data corresponding to the target sampling version as the training data of the target training version.
[0195] It should be noted that when training different types of models, such as general models and industry models, their data requirements and sampling ratios are different. Therefore, it is necessary to reasonably plan the model training tasks and sampling rates before model training. When designing the training version, it is necessary to accurately grasp the data requirements and sampling strategies to ensure that the performance and generalization ability of the model are optimal.
[0196] In another embodiment of the present application, a specific implementation method of matching the text to be processed with the training data of the target training version to obtain the target training data is described as follows:
[0197] In a possible implementation, the text to be processed may be accurately matched with the training data of the target training version. As long as a certain training data is retrieved and hits the text to be processed, the training data is determined to be the target training data.
[0198] Considering the poor matching accuracy of the above method, another specific implementation method for matching the text to be processed with the training data of the target training version to obtain the target training data is proposed in this application, referring to Figure 4 , Figure 4 A flowchart of a method for matching a text to be processed with training data of a target training version to obtain target training data is provided in an embodiment of the present application. The method may include the following steps:
[0199] S401: Determine a sensitive word list based on the text to be processed; the sensitive word list contains at least one sensitive word;
[0200] In a possible implementation, determining a sensitive word vocabulary based on the text to be processed includes: determining sensitive words in the text to be processed; and determining a vocabulary consisting of sensitive words in the text to be processed as the sensitive word vocabulary.
[0201] In the present application, the text to be processed can be segmented to obtain sensitive words in the text to be processed, and the segmentation methods include but are not limited to manual segmentation and automatic segmentation. Among them, manual segmentation can be performed by manually cutting the text to be processed according to delimiters (such as ", ", etc.). Automatic segmentation can be performed by using a segmentation engine (such as jieba) to segment the text to be processed.
[0202] Taking into account that there may be text errors in the text to be processed, and the sensitive words in the text to be processed may also have problems, which will affect the matching accuracy. Therefore, in another possible implementation, based on the text to be processed, determining the sensitive word list includes: determining the sensitive words in the text to be processed; obtaining related words corresponding to the sensitive words in the text to be processed, and the related words include any one or more of synonyms, antonyms, and similar words; determining the sensitive words in the text to be processed and the word list composed of the related words as the sensitive word list.
[0203] The examples of related words corresponding to the sensitive words in the text to be processed are as follows:
[0204] Mongolian and Chinese medicine → Menghan medicine (homophone);
[0205] The next meeting will be without reason → will not be without reason (similar words);
[0206] If you don't pay, it's a shame → Just forget it if you don't pay (similar word + homonym).
[0207] S402: Match the sensitive word list with the training data of the target training version to obtain target training data.
[0208] In the present application, the sensitive word list is matched with the training data of the target training version. Specifically, full-text search can be performed on the training data of the target training version based on the sensitive word list. Usually, full-text search uses solutions such as Elasticsearch or Lucene. However, for huge raw data, even after preprocessing, the data volume may still reach the PB (Petabyte, i.e., one trillion bytes) level. There are the following problems with using these tools: slow index establishment speed, significant data expansion (such as the Elasticsearch data expansion rate is between 2-10 times), and high storage pressure. In addition, the construction and maintenance of a PB-level full-text retrieval system will consume huge resources and be costly. Therefore, in the present application, the data matching range can be preset, and the data processing version associated with the target training version or the training data corresponding to the target sampling version can be selected for matching, which greatly reduces the scope of the retrieved files, and the file reading efficiency can be greatly improved by using a multi-threaded method, and it is easy to expand without additional storage overhead.
[0209] In this application, the target training data can be determined based on the matching degree between the sensitive word list and the training data of the target training version. In this application, the maximum threshold of the target training data can be preset (the default is 10,000). If the search result exceeds this threshold, the search will fail directly. This design can effectively block invalid queries, reduce system memory overhead, and trade time for space.
[0210] The target training data is the training data of the target training version whose matching degree with the sensitive word list is greater than a preset threshold.
[0211] For each training data, the calculation method of its matching degree with the sensitive word list depends on the word segmentation method, as follows:
[0212] In manual word segmentation, if homophone, synonym, or similar word search is not enabled, the matching degree of a training data The calculation formula is:
[0213] .
[0214] Among them, K is the number of hit sensitive words, and N is the total number of sensitive words in the text to be processed.
[0215] If you enable homophone, synonym, or similar word search, the matching degree The calculation formula is:
[0216] .
[0217] Among them, the subscripts r1, y, x, and m represent manual word segmentation, homophones, similar words, and synonyms respectively, K is the number of hit sensitive words, N is the total number of sensitive words in manual word segmentation, homophones, similar words, and synonyms, and W is the weight of each item in manual word segmentation, homophones, similar words, and synonyms.
[0218] Automatic word segmentation method, a matching calculation method for training data, uses the TD-IDF algorithm as the basic algorithm, regularizes the TD-IDF results, and increases the weight of low-frequency sensitive words through the weighted average algorithm design to calculate the final score.
[0219] TF-IDF consists of two parts, TF (Term Frequency) and IDF (Inverse Document Frequency). The importance of a word increases in direct proportion to the number of times it appears in a document, but decreases in inverse proportion to the frequency of its appearance in the corpus. Low-frequency words have higher values and are more in line with the needs of sensitive word retrieval.
[0220] If homophone, synonym, or similar word search is not enabled, the matching degree calculation steps and formula for a training data set are as follows:
[0221] Step 1: Calculate the TD-IDF value of a single term T , the formula is as follows:
[0222] .
[0223] Where T is a term, D is a document, f(T,D) represents the number of times term T appears in document D, S(D) is the total number of terms contained in document D, df(T) is the total number of documents in the document set that contain term T, and N d Is the number of all documents in the retrieved directory.
[0224] Step 2: Normalize the TF-IDF value. After calculating the TF-IDF value, you need to normalize the TF-IDF value to a value range of [0-1]. The formula is as follows:
[0225] .
[0226] in, is the TF-IDF value of the regularized term T, is the TF-IDF value of term T, and max is the maximum value.
[0227] Step 3: Calculate the weights using simple linear weights. The higher the value, the greater the weight, and vice versa.
[0228] The values are sorted from small to large, n is the total number of terms, W is the term weight, k is the step coefficient (used for the step amplitude between conditional weights), and i is the sorted Subscript, then the i-th Weight The calculation formula is:
[0229] .
[0230] Step 4: Calculate the matching score , the calculation formula is:
[0231] .
[0232] in, After the values are sorted from small to large, t (T,n) Indicates the nth , Indicates the nth The weight of .
[0233] If you enable homophone, synonym, or similar word search, the matching degree The calculation formula is:
[0234] .
[0235] Among them, the subscripts r2, y, x, and m represent automatic word segmentation, homophones, similar words, and synonyms respectively, and W is the weight of each item in automatic word segmentation, homophones, similar words, and synonyms.
[0236] A sensitive information processing method provided in an embodiment of the present application is introduced above, and a device for executing the above sensitive information processing method will be introduced below.
[0237] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of a sensitive information processing device provided in an embodiment of the present application. Figure 5 As shown, the sensitive information processing device includes:
[0238] An acquisition unit 11 is used to acquire a text to be processed; the text to be processed is a text containing sensitive information generated by a model of a target training version;
[0239] A matching unit 12 is used to match the text to be processed with the training data of the target training version to obtain target training data, where the target training data is the training data in the target training version that matches the text to be processed;
[0240] A tracing unit 13, configured to obtain corresponding target original data based on the meta information of the target training data, wherein the target original data and the corresponding target training data contain the same meta information;
[0241] The impact scope assessment unit 14 is used to generate an impact scope report of the sensitive information based on the target original data, wherein the impact scope report is used to indicate the training version model associated with the target original data and the preprocessing process of the target original data.
[0242] In a possible implementation, the apparatus further includes a training data determination unit for a target training version, and the training data determination unit for the target training version includes:
[0243] A data processing version determining unit, used to determine the data processing version associated with the target training version;
[0244] A target original data determination unit, used to determine the target original data corresponding to the data processing version, each target original data carries the meta information of the target original data;
[0245] A preprocessing unit, used to preprocess the target original data corresponding to the data processing version to obtain preprocessed data corresponding to the data processing version;
[0246] The training data determination unit is used to determine the training data of the target training version based on the preprocessed data corresponding to the data processing version.
[0247] In a possible implementation, the target original data determination unit is specifically configured to:
[0248] Obtaining pre-constructed raw data management information, wherein the raw data management information includes multiple data sets, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of raw data;
[0249] Determining a target data set and a target data batch in the target data set from the original data management information;
[0250] According to the storage directory of the original data corresponding to each of the target data batches, the original data corresponding to each of the target data batches is obtained as the target original data corresponding to the data processing version.
[0251] In a possible implementation, the device further includes: an original data management information construction unit, wherein the original data management information construction unit is specifically configured to:
[0252] Determine the storage directory for the raw data required for model training;
[0253] Dividing the original data required for the model training to obtain at least one data set, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of original data;
[0254] For each data batch, directory scanning, directory authorization and data backup operations are performed on the data batch to obtain the original data management information.
[0255] In a possible implementation, the preprocessing unit is specifically used to:
[0256] Cleaning the target original data corresponding to the data processing version to obtain cleaned data corresponding to the data processing version;
[0257] Determine the cleaned data corresponding to the data processing version as the pre-processed data corresponding to the data processing version;
[0258] Alternatively, further manually reviewing the cleaned data corresponding to the data processing version to obtain manually reviewed data corresponding to the data processing version;
[0259] The data corresponding to the data processing version and having passed manual review is determined as the pre-processed data corresponding to the data processing version.
[0260] In a possible implementation, the preprocessing unit is specifically used to:
[0261] For each of the target data batches, the original data corresponding to the target data batch is cleaned to obtain a new data batch; each new data batch is the cleaned data corresponding to the data processing version.
[0262] In a possible implementation, the training data determination unit is specifically configured to:
[0263] Using the preprocessed data corresponding to the data processing version as training data of the target training version;
[0264] Alternatively, determine the target sampling version;
[0265] Sampling is performed from the preprocessed data corresponding to the data processing version to obtain training data corresponding to the target sampling version as training data for the target training version.
[0266] In a possible implementation, the matching unit includes:
[0267] A sensitive word list determining unit, configured to determine a sensitive word list based on the text to be processed; the sensitive word list includes at least one sensitive word;
[0268] The matching subunit is used to match the sensitive word list with the training data of the target training version to obtain target training data.
[0269] In a possible implementation, the sensitive word list determination unit is specifically used to:
[0270] Determining sensitive words in the text to be processed;
[0271] Determine a vocabulary consisting of sensitive words in the text to be processed as a sensitive word vocabulary;
[0272] or,
[0273] Obtaining related words corresponding to the sensitive words in the text to be processed, wherein the related words include any one or more of synonyms, antonyms, and words with similar forms;
[0274] A vocabulary consisting of sensitive words in the text to be processed and the related words is determined as a sensitive word vocabulary.
[0275] The present application also provides an electronic device in an embodiment. Figure 6 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0276] like Figure 6 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0277] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0278] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any sensitive information processing method provided in the embodiment of the present application.
[0279] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When one or more computer programs are executed by an electronic device, the electronic device can implement any sensitive information processing method provided in the embodiment of the present application.
[0280] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0281] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0282] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0283] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
Claims
1. A sensitive information processing method, characterized in that: include: Get the text to be processed; The text to be processed is text containing sensitive information generated by the model of the target training version; Matching the text to be processed with the training data of the target training version to obtain target training data, wherein the target training data is the training data in the training data of the target training version that matches the text to be processed; Obtaining corresponding target original data based on the meta information of the target training data, wherein the target original data and the corresponding target training data contain the same meta information; Based on the target original data, an impact scope report of the sensitive information is generated, where the impact scope report is used to indicate a training version model associated with the target original data and a preprocessing flow of the target original data.
2. The method according to claim 1, characterized in that The method for determining the training data of the target training version includes: Determining a data processing version associated with the target training version; Determine the target original data corresponding to the data processing version, each target original data carries the meta information of the target original data; Preprocessing the target original data corresponding to the data processing version to obtain preprocessed data corresponding to the data processing version; Based on the preprocessed data corresponding to the data processing version, the training data of the target training version is determined.
3. The method according to claim 2, characterized in that The determining the target original data corresponding to the data processing version includes: Obtaining pre-constructed raw data management information, wherein the raw data management information includes multiple data sets, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of raw data; Determining a target data set and a target data batch in the target data set from the original data management information; According to the storage directory of the original data corresponding to each of the target data batches, the original data corresponding to each of the target data batches is obtained as the target original data corresponding to the data processing version.
4. The method according to claim 3, characterized in that: The original data management information is constructed as follows: Determine the storage directory for the raw data required for model training; Dividing the original data required for the model training to obtain at least one data set, each data set includes at least one data batch, and each data batch corresponds to a storage directory of a batch of original data; For each data batch, directory scanning, directory authorization and data backup operations are performed on the data batch to obtain the original data management information.
5. The method according to claim 3, characterized in that: The preprocessing of the target original data corresponding to the data processing version to obtain the preprocessed data corresponding to the data processing version includes: Cleaning the target original data corresponding to the data processing version to obtain cleaned data corresponding to the data processing version; Determine the cleaned data corresponding to the data processing version as the pre-processed data corresponding to the data processing version; Alternatively, further manually reviewing the cleaned data corresponding to the data processing version to obtain manually reviewed data corresponding to the data processing version; The data corresponding to the data processing version and having passed manual review is determined as the pre-processed data corresponding to the data processing version.
6. The method according to claim 5, characterized in that The cleaning process of the target original data corresponding to the data processing version to obtain the cleaned data corresponding to the data processing version includes: For each of the target data batches, the original data corresponding to the target data batch is cleaned to obtain a new data batch; each new data batch is the cleaned data corresponding to the data processing version.
7. The method according to claim 2, characterized in that The determining the training data of the target training version based on the pre-processed data corresponding to the data processing version includes: Using the preprocessed data corresponding to the data processing version as training data of the target training version; Alternatively, determine the target sampling version; Sampling is performed from the preprocessed data corresponding to the data processing version to obtain training data corresponding to the target sampling version as training data for the target training version.
8. The method according to claim 1, characterized in that: The step of matching the text to be processed with the training data of the target training version to obtain the target training data includes: Based on the text to be processed, determining a sensitive word list; the sensitive word list contains at least one sensitive word; The sensitive word list is matched with the training data of the target training version to obtain target training data.
9. The method according to claim 8, characterized in that The step of determining a sensitive word list based on the text to be processed includes: Determining sensitive words in the text to be processed; Determine a vocabulary consisting of sensitive words in the text to be processed as a sensitive word vocabulary; or, Obtaining related words corresponding to the sensitive words in the text to be processed, wherein the related words include any one or more of homophones, synonyms, and similar words; A vocabulary consisting of sensitive words in the text to be processed and the related words is determined as a sensitive word vocabulary.
10. A sensitive information processing device, characterized in that: include: An acquisition unit, used for acquiring the text to be processed; The text to be processed is text containing sensitive information generated by the model of the target training version; A matching unit, used for matching the text to be processed with the training data of the target training version to obtain target training data, wherein the target training data is the training data in the training data of the target training version that matches the text to be processed; A tracing unit, configured to obtain corresponding target original data based on the meta information of the target training data, wherein the target original data and the corresponding target training data contain the same meta information; An impact scope assessment unit is used to generate an impact scope report of the sensitive information based on the target original data, wherein the impact scope report is used to indicate a training version model associated with the target original data and a preprocessing process of the target original data.
11. A computer program product, characterized in that It includes computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the sensitive information processing method as described in any one of claims 1 to 9.
12. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the sensitive information processing method as described in any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that: The storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the sensitive information processing method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Text auditing method, text auditing model, equipment and storage medium
CN112487149A
Sensitive word auditing method based on large language model, storage medium and electronic equipment
CN116720515A