Data encryption and desensitization method and system based on AI and rule features
By introducing a data platform and AI models, combined with rule-based feature technology, the ETL process achieves automated data encryption and desensitization, solving the problems of complex configuration and insufficient accuracy in existing technologies, and improving data processing efficiency and operational efficiency.
Patent Information
- Application Number
- CN202511470267.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-11
AI Technical Summary
Existing ETL processes are complex and cumbersome to configure, manual configuration is outdated, AI prediction accuracy is insufficient, and traditional rules have limited adaptability in unknown knowledge scenarios, resulting in low data processing efficiency and high costs.
By introducing data standard logic into the data platform and utilizing AI models and rule feature technologies, automated data encryption and desensitization are achieved through data metadata, ETL jobs, AI models, and rule matching processes. Combined with data quality and security management, this improves operational efficiency.
It enables convenient and intelligent data encryption and desensitization processes, reduces operation and maintenance costs, improves data processing efficiency and accuracy, and adapts to changes in the data environment.
Smart Images

Figure CN120930185A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of data processing, and specifically relates to a data encryption and desensitization method and system based on AI and rule features. Background Technology
[0002] In the digital economy era, data has become a core production factor for enterprises and institutions, providing crucial support for business decision-making and innovation. However, data silos formed by heterogeneous systems such as CRM and ERP between departments severely restrict the full release of the value of government data. While ETL technology, as a core means of data integration, can achieve cross-system connectivity through unified access to data from various systems, it faces multiple challenges in practical applications: the diverse and massive types of data sources that enterprises and institutions need to handle lead to a surge in the number of ETL tasks, significantly increasing scheduling and maintenance pressure; to ensure secure and compliant data sharing, sensitive information needs to be encrypted and de-identified to prevent data leakage; simultaneously, different tasks and different data tables require separate parameter configurations, and readjustments are necessary when business changes or new data sources are added, resulting in not only a large amount of repetitive work and high labor costs but also a high risk of configuration errors. These problems urgently necessitate the automation and intelligent upgrading of the ETL process.
[0003] The existing technology has the following problems: 1. The existing ETL process configuration is cumbersome and complex, requiring manual setting of encryption and desensitization rules for each field, resulting in low overall operation efficiency and severely restricting data processing timeliness; 2. The traditional manual ETL security configuration mode is significantly outdated, failing to keep up with technological development trends and has not yet been optimized and upgraded by integrating the currently widely used rule base technology with continuously iterating AI technology, resulting in extremely poor efficiency. 3. While AI prediction technology has advantages in reducing labor costs and improving processing efficiency, its limitations are also quite prominent: although it has a strong ability to learn unknown knowledge, the accuracy of its prediction results still needs to be improved and has not yet reached an ideal level. 4. While traditional rules such as regular expressions can achieve accurate data matching, they lack the ability to effectively handle scenarios with unknown knowledge, resulting in significant limitations in the adaptability of the technology. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a data encryption and desensitization method and system based on AI and rule features, which solves the technical problems in the prior art.
[0005] In a first aspect, the present invention provides the following technical solution: a data encryption and desensitization method based on AI and rule features, comprising: Create data elements and configure data element information to obtain a data element library; Acquire training data and label the data features of the training data. Input the labeled training data into the preset AI model for training to output the target AI model. Create an ETL job, obtain matching information based on the ETL job, perform initial matching in the data metadata database based on the matching information, and if the initial matching fails, execute the feature matching process; if the feature matching fails, execute the rule matching process. If the rule matching fails, the information to be matched is input into the target AI model for prediction to obtain field features, and model matching is performed based on the field features; If the initial matching, feature matching, rule matching, or model matching is successful, the corresponding data element information is output and the encryption and desensitization rules associated with the data element information are obtained. If the model matching fails, the information to be matched is marked and the marked information to be matched is input into the target AI model for the next round of training.
[0006] Compared with existing technologies, the beneficial effects of this invention are as follows: By introducing the data standard logic of the data platform, this invention uses data elements as the driving force, leverages emerging and gradually maturing AI model technologies, and combines them with mature rule recognition technologies to integrate information technologies such as data quality, data security, and metadata. This allows users to complete user-set goals more conveniently, intelligently, and automatically during data governance and data exchange, freeing users from tedious task configuration, improving the work efficiency of operation and maintenance personnel, and reducing operation and maintenance costs.
[0007] Preferably, the data metadata includes field standard naming, field logical type, field length, field precision, data dictionary, field name matching rules, field data security level, field security processing mode, encryption and desensitization algorithm corresponding to the field, quality inspection rules, and AI feature code.
[0008] Preferably, the step of inputting the labeled training data into a preset AI model for training to output a target AI model includes: The labeled training data is input into a preset AI model, and the preset data processing algorithm stored in the preset AI model is used to preprocess the labeled training data to obtain processed data. The processed data is subjected to text recognition, and the recognized text data is converted into a data vector; The data vector is augmented and then input into several basic models in the preset AI model for classification training to obtain an initial model; Obtain a test dataset, input the test dataset into the base model in the preset AI model for prediction, and output several prediction results. Combine the several prediction results through a probability fusion model to obtain a combined result. Optimize the model parameters based on the combined result to obtain the target AI model.
[0009] Preferably, the steps of creating an ETL job, obtaining matching information based on the ETL job, performing an initial match on the metadata database based on the matching information, and executing a feature matching process if the initial match fails, and executing a rule matching process if the feature matching fails, include: Create an ETL job and set the extraction nodes and extraction node information in the preset job configuration template. The extraction node information includes binding data source, extraction range, filtering conditions, timestamp field, and sorting field. Based on the extraction node, information to be matched is extracted, and based on the field standard naming, field logical type, field length, and field precision in the information to be matched, precise matching is performed in the data element database to complete the initial matching process. If no matching data element is found, the initial matching fails and sample row value distribution statistics are performed on the fields in the information to be matched to output statistical results. The similarity between the output statistical results and the data fields corresponding to each data element information in the data element library is calculated. If the similarity is less than a preset value, feature matching fails and a portion of the data in the information to be matched is extracted. This portion of the data is then matched against the data elements using a quality rule base and according to the quality rules. If no matching data element is found, rule matching fails.
[0010] Preferably, if rule matching fails, the information to be matched is input into the target AI model for prediction to obtain field features. The specific steps for model matching based on the field features are as follows: If the rule matching fails, the field information in the information to be matched will be extracted and input into the target AI model for two-layer prediction to output field features. The field features will be matched with data elements. If no corresponding data element is matched, the model matching will fail.
[0011] Secondly, the present invention provides the following technical solution: a data encryption and desensitization system based on AI and rule features, the system comprising: The creation module is used to create data elements and configure data element information to obtain a data element library; The training module is used to acquire training data and perform data feature labeling on the training data. The labeled training data is then input into a preset AI model for training to output the target AI model. The first matching module is used to create an ETL job, obtain matching information based on the ETL job, perform an initial matching in the data metadata library based on the matching information, and execute a feature matching process if the initial matching fails, and execute a rule matching process if the feature matching fails. The second matching module is used to input the information to be matched into the target AI model for prediction if the rule matching fails, so as to obtain field features and perform model matching based on the field features. The output module is used to output the corresponding data element information and obtain the encryption and desensitization rules associated with the data element information if the initial matching, feature matching, rule matching, or model matching is successful. If the model matching fails, the information to be matched is marked and the marked information to be matched is input into the target AI model for the next round of training.
[0012] Preferably, the first matching module includes: The job creation submodule is used to create ETL jobs and set extraction nodes and extraction node information in a preset job configuration template. The extraction node information includes data source binding, extraction range, filtering conditions, timestamp field, and sorting field. The extraction submodule is used to extract information to be matched based on the extraction node, and to perform precise matching in the data element library based on the field standard naming, field logical type, field length and field precision in the information to be matched, so as to complete the initial matching process. The statistics submodule is used to perform sample row value distribution statistics on the fields in the information to be matched if no corresponding data element is matched, so as to output the statistical results and calculate the similarity between the output statistical results and the data fields corresponding to each data element information in the data element library. The rules submodule is used to extract a portion of the data from the information to be matched if the similarity is less than a preset value. The data is then matched against the data elements in the quality rule library according to the quality rules. If no matching data element is found, the rule matching fails.
[0013] Preferably, the second matching module is specifically used for: If the rule matching fails, the field information in the information to be matched will be extracted and input into the target AI model for two-layer prediction to output field features. The field features will be matched with data elements. If no corresponding data element is matched, the model matching will fail.
[0014] Thirdly, the present invention provides the following technical solution: a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data encryption and desensitization method based on AI and rule features as described above.
[0015] Fourthly, the present invention provides the following technical solution: a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the data encryption and desensitization method based on AI and rule features as described above. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of a data encryption and desensitization method based on AI and rule features provided in Embodiment 1 of the present invention; Figure 2 This is a structural block diagram of the data encryption and desensitization system based on AI and rule features provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the hardware structure of a computer provided for another embodiment of the present invention.
[0018] The embodiments of the present invention will be further described below with reference to the accompanying drawings. Detailed Implementation
[0019] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain embodiments of the present invention, and should not be construed as limiting the present invention.
[0020] Example 1 In Embodiment 1 of the present invention, as Figure 1 As shown, a data encryption and desensitization method based on AI and rule features includes: S1. Create data elements and configure data element information to obtain the data element library; The data metadata includes field standard naming, field logical type, field length, field precision, data dictionary, field name matching rules, field data security level, field security processing mode, encryption and desensitization algorithm corresponding to the field, quality inspection rules and AI feature code; Specifically, when creating a data element, multi-dimensional configuration is required to ensure its standardization and functionality: First, define the standard naming, logical type, length, and precision of the fields, and associate them with the corresponding standard data dictionary to lay the foundation for the data element's attribute framework; second, set field name matching rules and clarify the word matching logic; at the same time, bind the data security level, security processing mode (de-identification or encryption), and the corresponding encryption and de-identification algorithms; in addition, it is also necessary to associate quality inspection rules and AI feature codes to achieve the linkage of data standards, data security, data quality control, and intelligent recognition. After completing the above configuration, submit the data element. After approval, the system will automatically complete the release process to ensure that the data element can be quickly put into application.
[0021] S2. Acquire training data and label the data features of the training data. Input the labeled training data into the preset AI model for training to output the target AI model. Specifically, step S2 includes: S21. Input the labeled training data into the preset AI model and perform data preprocessing on the labeled training data using the preset data processing algorithm stored in the preset AI model to obtain processed data. Specifically, the training data with completed feature identification is first input into the model, and the training process is started. The AI model algorithm will preprocess the data by cleaning, transforming, and normalizing the original data to remove noise and compensate for defects, making it meet the needs of subsequent training or analysis, while improving data quality and processing efficiency. The AI model algorithm here is a commonly used data processing algorithm in existing technology, such as cleaning algorithm, normalization algorithm, and filtering algorithm. The preset AI model here is a supervised deep learning model. For example, it can be trained by uploading labeled data such as ID cards, names, and addresses. When making a prediction, a certain amount of data from a certain column in the table is taken to predict what data element this column is (ID card or name).
[0022] S22. Perform text recognition on the processed data and convert the recognized text data into a data vector; Specifically, the next step is text tagging, which converts unstructured text (such as words, sentences, and paragraphs) into numerical vectors (or matrices) that computers can understand, so that the semantic information of the text can be captured and calculated by the model.
[0023] S23. The data vector is augmented and the augmented data vector is input into several basic models in the preset AI model for classification training to obtain an initial model; Specifically, data augmentation is then carried out by reasonably transforming existing data or generating new samples to expand the dataset size and enrich data diversity, thereby alleviating the data scarcity problem and improving the model's generalization ability (reducing overfitting).
[0024] S24. Obtain a test dataset, input the test dataset into the base model in the preset AI model for prediction, and output several prediction results. Combine the several prediction results through a probability fusion model to obtain a combined result. Optimize the model parameters based on the combined result to obtain the target AI model. Specifically, classification training is performed first, and then the prediction results (usually probability distributions) of multiple basic models are combined through a probability fusion model to combine the advantages of each model to reduce the bias or noise of a single model and improve the overall prediction performance (such as accuracy and robustness). After training, test set data is submitted for data feature prediction. The model parameters are dynamically adjusted to optimize the model based on the prediction results and the final result is confirmed. At the same time, the newly added manually adjusted data elements in the table column are scanned and incorporated into the model iteration system as incremental training data to continuously improve the model's adaptability and accuracy. The newly added manually adjusted data elements in the table column can be data that the model did not match in subsequent steps or updated training data.
[0025] S3. Create an ETL job, obtain matching information based on the ETL job, perform initial matching in the data metadata database based on the matching information, and execute feature matching process if the initial matching fails, and execute rule matching process if feature matching fails. Step S3 includes: S31. Create an ETL job and set the extraction node and extraction node information in the preset job configuration template. The extraction node information includes binding data source, extraction range, filtering conditions, timestamp field, and sorting field. Specifically, when creating an ETL job, the extraction node configuration needs to be completed in the preset job workflow configuration template: First, bind the target data source to the extraction node to clarify the data source; then select the database table to be extracted, or precisely limit the extraction range through custom SQL statements to ensure the targeting and accuracy of data extraction, laying the foundation for subsequent data processing.
[0026] S32. Extract the information to be matched based on the extraction node, and perform precise matching in the data element database based on the field standard naming, field logical type, field length, and field precision in the information to be matched, so as to complete the initial matching process. Specifically, during the initial matching process, an initial match is performed based on some field information in the information to be matched, such as field standard naming, field logical type, field length, and field precision. If the initial match is successful, the encryption and desensitization rules associated with the data element are obtained, a calculation expression is generated according to the rules, and then the processor fills in the expression.
[0027] S33. If no corresponding data element is matched, the initial matching fails and sample row value distribution statistics are performed on the fields in the information to be matched to output the statistical results and calculate the similarity between the output statistical results and the data fields corresponding to each data element information in the data element library. Specifically, if the initial match fails, the similarity between the output statistical result and the data fields corresponding to each data element in the data element library is calculated. If the similarity is not less than 80%, the match is considered successful. The encryption and desensitization rules associated with the data element are then obtained, a calculation expression is generated according to the rules, and the processor fills in the expression. Regarding the similarity mentioned above, if it is considered that for two data to be calculated, if 80% or more of their values belong to the same dictionary, the match is considered successful. If the similarity is less than 80%, the match is considered failed.
[0028] S34. If the similarity is less than the preset value, feature matching fails and a portion of the data in the information to be matched is extracted. The portion of the data is then matched with data elements through the quality rule base and according to the quality rules. If no corresponding data element is matched, the rule matching fails. Specifically, if feature matching fails, some data will be extracted and data elements will be matched according to quality rules through the quality rule base. If the matching is successful, the process is the same as above.
[0029] S4. If the rule matching fails, the information to be matched is input into the target AI model for prediction to obtain field features, and model matching is performed based on the field features. Specifically, step S4 is as follows: If the rule matching fails, the field information in the information to be matched will be extracted and input into the target AI model for two-layer prediction to output field features. The field features will be matched with data elements. If no corresponding data element is matched, the model matching will fail. Specifically, in the two-layer prediction process, the model is built with the text representation results as input, and the network structure is designed in combination with the task requirements. The model parameters are optimized through training to enable it to have a deep understanding of text semantics and the ability to predict task objectives. Then, the raw results output by the model are parsed, filtered and post-processed to generate intuitive and reliable final conclusions in combination with business scenarios. The fields adopt a two-layer prediction strategy, that is, first perform feature prediction of the major category, and then perform feature prediction of the precise category. The precise field meaning positioning effect is achieved through two-layer prediction.
[0030] S5. If the initial matching is successful, or the feature matching is successful, or the rule matching is successful, or the model matching is successful, the corresponding data element information is output and the encryption and desensitization rules associated with the data element information are obtained. If the model matching fails, the information to be matched is marked and the marked information to be matched is input into the target AI model for the next round of training. Specifically, if the initial matching, feature matching, rule matching, or model matching is successful, that is, if any matching is successful, the encryption and desensitization rule associated with the data element is obtained, a calculation expression is generated according to the rule, and then the processor fills in the expression. Finally, these encryption and desensitization rules are automatically bound to the target data column to complete the automated configuration of encryption and desensitization. After the configuration is completed, the automated encryption and desensitization process can be executed. If the last model matching also fails, the information to be matched is marked, and the marked information to be matched is input into the target AI model for the next round of training. Specifically, the above steps aim to continuously improve the accuracy of model predictions using newly generated data, while avoiding repeated training on historical data. In practical applications, after initial training, the model is deployed to business scenarios, accumulating a large amount of new labeled data or feedback information over time. Incremental training, through a well-designed training mechanism, combines this new data with key historical data, selectively learning patterns and rules in the new data while retaining the model's existing knowledge. This approach not only effectively reduces the computational cost and time consumption of repeated training but also enables the model to adapt to changes in data distribution (such as concept drift, the emergence of new scenarios, etc.), continuously optimizing prediction accuracy or decision-making capabilities. During implementation, it is usually necessary to balance the weights of new and old data, handle class imbalance issues, and prevent the model from overfitting to new data through regularization and other means, ultimately achieving robust evolution of the model in dynamic environments and ensuring its long-term adaptability to changes in business needs.
[0031] The data encryption and desensitization method based on AI and rule features provided in Embodiment 1 of this invention introduces the data standard logic of a data platform, uses data elements as the driving force, leverages emerging and gradually maturing AI model technologies, and combines mature rule recognition technologies to integrate information technologies such as data quality, data security, and metadata. This allows users to more conveniently, intelligently, and automatically complete user-set goals during data governance and data exchange, freeing users from tedious task configuration, improving the work efficiency of operation and maintenance personnel, and reducing operation and maintenance costs.
[0032] Example 2 like Figure 2 As shown, in Embodiment 2 of the present invention, a data encryption and desensitization system based on AI and rule features is provided. The system includes: Create module 1, which is used to create data elements and configure data element information to obtain the data element library; Training module 2 is used to acquire training data and perform data feature labeling on the training data. The labeled training data is then input into a preset AI model for training to output the target AI model. The first matching module 3 is used to create an ETL job, obtain matching information based on the ETL job, perform initial matching in the data element library based on the matching information, and execute a feature matching process if the initial matching fails, and execute a rule matching process if the feature matching fails. The second matching module 4 is used to input the information to be matched into the target AI model for prediction if the rule matching fails, so as to obtain field features and perform model matching based on the field features; Output module 5 is used to output the corresponding data element information and obtain the encryption and desensitization rules associated with the data element information if the initial matching, feature matching, rule matching, or model matching is successful. If the model matching fails, the information to be matched is marked and the marked information to be matched is input into the target AI model for the next round of training.
[0033] The training module 2 includes: The preprocessing submodule is used to input the labeled training data into the preset AI model and perform data preprocessing on the labeled training data through the preset data processing algorithm stored in the preset AI model to obtain processed data. The text submodule is used to perform text recognition on the processed data and convert the recognized text data into a data vector. An enhancement submodule is used to enhance the data vector and input the enhanced data vector into several basic models in the preset AI model for classification training to obtain an initial model; The parameter submodule is used to acquire a test dataset, input the test dataset into the base model in the preset AI model for prediction, output several prediction results, combine the several prediction results through a probability fusion model to obtain a combined result, and optimize the model parameters based on the combined result to obtain the target AI model.
[0034] The first matching module 3 includes: The job creation submodule is used to create ETL jobs and set extraction nodes and extraction node information in a preset job configuration template. The extraction node information includes data source binding, extraction range, filtering conditions, timestamp field, and sorting field. The extraction submodule is used to extract information to be matched based on the extraction node, and to perform precise matching in the data element library based on the field standard naming, field logical type, field length and field precision in the information to be matched, so as to complete the initial matching process. The statistics submodule is used to perform sample row value distribution statistics on the fields in the information to be matched if no corresponding data element is matched, so as to output the statistical results and calculate the similarity between the output statistical results and the data fields corresponding to each data element information in the data element library. The rules submodule is used to extract a portion of the data from the information to be matched if the similarity is less than a preset value. The data is then matched against the data elements in the quality rule library according to the quality rules. If no matching data element is found, the rule matching fails.
[0035] Specifically, the second matching module 4 is used for: If the rule matching fails, the field information in the information to be matched will be extracted and input into the target AI model for two-layer prediction to output field features. The field features will be matched with data elements. If no corresponding data element is matched, the model matching will fail.
[0036] In other embodiments of the present invention, the present invention provides the following technical solution: a computer, including a memory 102, a processor 101, and a computer program stored on the memory 102 and executable on the processor 101, wherein the processor 101 executes the computer program to implement the data encryption and desensitization method based on AI and rule features as described above.
[0037] Specifically, the processor 101 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of the present invention.
[0038] The memory 102 may include a large-capacity memory for data or instructions. For example, and not limitingly, the memory 102 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 102 may include removable or non-removable (or fixed) media. Where appropriate, the memory 102 may be internal or external to a data processing device. In a particular embodiment, the memory 102 is non-volatile memory. In a particular embodiment, the memory 102 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random Access Memory (FPMDRAM), Extended Data Out Dynamic Random Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.
[0039] The memory 102 can be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 101.
[0040] The processor 101 reads and executes the computer program instructions stored in the memory 102 to implement the above-mentioned data encryption and desensitization method based on AI and rule features.
[0041] In some embodiments, the computer may further include a communication interface 103 and a bus 100. For example, Figure 3 As shown, the processor 101, memory 102, and communication interface 103 are connected through bus 100 and complete communication with each other.
[0042] The communication interface 103 is used to enable communication between the various modules, devices, units, and / or equipment in the embodiments of the present invention. The communication interface 103 can also enable data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.
[0043] Bus 100 includes hardware, software, or both, that couples components of a computer device together. Bus 100 includes, but is not limited to, at least one of the following: data bus, address bus, control bus, expansion bus, and local bus. For example, and not as a limitation, bus 100 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 100 may include one or more buses. Although specific buses are described and illustrated in the embodiments of the present invention, the present invention is contemplated by any suitable bus or interconnect.
[0044] The computer can execute the AI-based and rule-based data encryption and desensitization method of the present invention based on the acquired AI-based and rule-based data encryption and desensitization system, thereby realizing AI-based and rule-based data encryption and desensitization.
[0045] In some further embodiments of the present invention, in conjunction with the above-described data encryption and desensitization method based on AI and rule features, the present invention provides the following technical solution: a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the above-described data encryption and desensitization method based on AI and rule features.
[0046] Those skilled in the art will understand that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0047] More specific examples of readable media (a non-exhaustive list) include: electrical connections (electronic devices) with one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0048] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0049] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0050] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A data encryption and desensitization method based on AI and rule features, characterized in that, include: Create data elements and configure data element information to obtain a data element library; Acquire training data and label the data features of the training data. Input the labeled training data into the preset AI model for training to output the target AI model. Create an ETL job, obtain matching information based on the ETL job, perform initial matching in the data metadata database based on the matching information, and if the initial matching fails, execute the feature matching process; if the feature matching fails, execute the rule matching process. If the rule matching fails, the information to be matched is input into the target AI model for prediction to obtain field features, and model matching is performed based on the field features; If the initial matching, feature matching, rule matching, or model matching is successful, the corresponding data element information is output and the encryption and desensitization rules associated with the data element information are obtained. If the model matching fails, the information to be matched is marked and the marked information to be matched is input into the target AI model for the next round of training.
2. The data encryption and desensitization method based on AI and rule features according to claim 1, characterized in that, The data metadata includes field standard naming, field logical type, field length, field precision, data dictionary, field name matching rules, field data security level, field security processing mode, encryption and desensitization algorithm corresponding to the field, quality inspection rules, and AI feature code.
3. The data encryption and desensitization method based on AI and rule features according to claim 1, characterized in that, The step of inputting the labeled training data into a preset AI model for training to output a target AI model includes: The labeled training data is input into a preset AI model, and the preset data processing algorithm stored in the preset AI model is used to preprocess the labeled training data to obtain processed data. The processed data is subjected to text recognition, and the recognized text data is converted into a data vector; The data vector is augmented and then input into several basic models in the preset AI model for classification training to obtain an initial model; Obtain a test dataset, input the test dataset into the base model in the preset AI model for prediction, and output several prediction results. Combine the several prediction results through a probability fusion model to obtain a combined result. Optimize the model parameters based on the combined result to obtain the target AI model.
4. The data encryption and desensitization method based on AI and rule features according to claim 1, characterized in that, The steps of creating an ETL job, obtaining matching information based on the ETL job, performing an initial match on the data metadata based on the matching information, and executing a feature matching process if the initial match fails, and then executing a rule matching process if the feature matching fails, include: Create an ETL job and set the extraction nodes and extraction node information in the preset job configuration template. The extraction node information includes binding data source, extraction range, filtering conditions, timestamp field, and sorting field. Based on the extraction node, information to be matched is extracted, and based on the field standard naming, field logical type, field length, and field precision in the information to be matched, precise matching is performed in the data element database to complete the initial matching process. If no matching data element is found, the initial matching fails and sample row value distribution statistics are performed on the fields in the information to be matched to output statistical results. The similarity between the output statistical results and the data fields corresponding to each data element information in the data element library is calculated. If the similarity is less than a preset value, feature matching fails and a portion of the data in the information to be matched is extracted. This portion of the data is then matched against the data elements using a quality rule base and according to the quality rules. If no matching data element is found, rule matching fails.
5. The data encryption and desensitization method based on AI and rule features according to claim 1, characterized in that, If the rule matching fails, the information to be matched is input into the target AI model for prediction to obtain field features. The specific steps for model matching based on the field features are as follows: If the rule matching fails, the field information in the information to be matched will be extracted and input into the target AI model for two-layer prediction to output field features. The field features will be matched with data elements. If no corresponding data element is matched, the model matching will fail.
6. A data encryption and desensitization system based on AI and rule features, characterized in that, The system includes: The creation module is used to create data elements and configure data element information to obtain a data element library; The training module is used to acquire training data and perform data feature labeling on the training data. The labeled training data is then input into a preset AI model for training to output the target AI model. The first matching module is used to create an ETL job, obtain matching information based on the ETL job, perform an initial matching in the data metadata library based on the matching information, and execute a feature matching process if the initial matching fails, and execute a rule matching process if the feature matching fails. The second matching module is used to input the information to be matched into the target AI model for prediction if the rule matching fails, so as to obtain field features and perform model matching based on the field features. The output module is used to output the corresponding data element information and obtain the encryption and desensitization rules associated with the data element information if the initial matching, feature matching, rule matching, or model matching is successful. If the model matching fails, the information to be matched is marked and the marked information to be matched is input into the target AI model for the next round of training.
7. The data encryption and desensitization system based on AI and rule features according to claim 6, characterized in that, The first matching module includes: The job creation submodule is used to create ETL jobs and set extraction nodes and extraction node information in a preset job configuration template. The extraction node information includes data source binding, extraction range, filtering conditions, timestamp field, and sorting field. The extraction submodule is used to extract information to be matched based on the extraction node, and to perform precise matching in the data element library based on the field standard naming, field logical type, field length and field precision in the information to be matched, so as to complete the initial matching process. The statistics submodule is used to perform sample row value distribution statistics on the fields in the information to be matched if no corresponding data element is matched, so as to output the statistical results and calculate the similarity between the output statistical results and the data fields corresponding to each data element information in the data element library. The rules submodule is used to extract a portion of the data from the information to be matched if the similarity is less than a preset value. The data is then matched against the data elements in the quality rule library according to the quality rules. If no matching data element is found, the rule matching fails.
8. The data encryption and desensitization system based on AI and rule features according to claim 6, characterized in that, The second matching module is specifically used for: If the rule matching fails, the field information in the information to be matched will be extracted and input into the target AI model for two-layer prediction to output field features. The field features will be matched with data elements. If no corresponding data element is matched, the model matching will fail.
9. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data encryption and desensitization method based on AI and rule features as described in any one of claims 1 to 5.
10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the data encryption and desensitization method based on AI and rule features as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Mobile service data desensitization rule generation method based on width learning
CN112989414A
Automatic quality inspection rule configuration method and device based on rule and semantic analysis
CN115221893A
Sample data processing method and system based on sensitive data identification and desensitization
CN117574427A
Data desensitization method and device, electronic equipment and storage medium
CN117668908A
Sensitive data identification method and apparatus, device, and computer storage medium
WO2024109619A1