Data security management method and device based on natural language transformation technology

By converting natural language data request information into SQL statements, and combining data classification and protection strategies, the limitations of automation and refined management and control in the data query and extraction process in the existing technology are solved, and efficient data security management is achieved.

CN120067138AActive Publication Date: 2025-05-30LIAONING BRANCH OF CHINA UNITED NETWORK COMM CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510541730.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing data security management methods have limitations in the process of data query and extraction, and cannot realize automated data classification and grading and refined management, and lack control and traceability of user operation behavior.

Method used

The natural language conversion technology is used to convert the natural language data request information into SQL statements, and the data is evaluated through a preset classification and grading model, and the target protection strategy is determined based on the level, so as to protect the data security.

Benefits of technology

It realizes the automation of data query and extraction, improves data acquisition efficiency, avoids the risk of human data leakage, and improves the effectiveness of data security management through refined management and traceability capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067138A_ABST
    Figure CN120067138A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data security management method and equipment implemented based on a natural language conversion technology, and relates to the technical field of data security, the method comprises the following steps: through a preset model, the preset model is used for converting a natural language into an SQL statement, and converting approved target data request information into an SQL statement; querying to-be-processed data corresponding to the data request information in a database according to the SQL statement, and exporting the to-be-processed data to a data security management and control platform; classifying and grading the to-be-processed data through a preset classifying and grading model to obtain a sensitivity grade of the to-be-processed data; and determining a target protection strategy of the to-be-processed data according to the sensitivity level and a preset multi-level data protection strategy, and controlling the data security management and control platform to perform security protection on the to-be-processed data according to the target protection strategy. The data query and data acquisition efficiency is improved, and the risk of artificial data leakage is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security, and in particular to a data security management method and device implemented based on natural language conversion technology. Background Art

[0002] As the value of data is gradually mined and emphasized, data has become the target of illegal acquisition by some people. In scenarios such as background operation and maintenance, front-end business access, and terminal file office, data is prone to various internal and external data leakage risks during the processes of storage, processing, processing, circulation, and distribution.

[0003] Currently, the existing data security control means is to achieve the data security control ability through terminal anti-leakage technology. However, in actual use, the terminal anti-leakage technology has limitations. For example, in the query stage, automated query cannot be achieved, and it is impossible to control whether the user's query requirements are reasonable and whether the query scope is compliant. The personnel responsible for querying data are also risk points for data leakage. For another example, during the process of user querying and extracting data, there is a lack of automated classification and grading ability for data, especially unable to automatically classify and grade unstructured data. At the same time, there is a lack of refined control ability based on classification and grading and post-audit ability for the entire process of data extraction and use, and it is impossible to control, trace, and implement the user operation behavior to the relevant responsible persons. Summary of the Invention

[0004] In view of this, the present invention provides a data security management method and device implemented based on natural language conversion technology to solve the problem of limitations in existing data security management.

[0005] In a first aspect, an embodiment of the present invention provides a data security management method implemented based on natural language conversion technology, and the method includes: By means of a preset model, which is used to convert natural language into SQL statements, converting the approved target data request information into SQL statements, including: Performing feature extraction on the approved target data request information through a feature extraction algorithm in the preset model to obtain natural language features, where the natural language features include context features, semantic features, and syntactic features; Generating an SQL statement corresponding to the target data request information according to the context features, the semantic features, and the syntactic features through an SQL statement generation algorithm in the preset model; Querying the data to be processed corresponding to the target data request information in the database according to the SQL statement, and exporting the data to be processed to a data security control platform; Classifying and grading the data to be processed through a preset classification and grading model to obtain the sensitivity level of the data to be processed; The target protection strategy for the data to be processed is determined according to the sensitivity level and the preset multi-level data protection strategy, and the data to be processed is securely protected according to the target protection strategy through the data security management and control platform.

[0006] Optionally, before the step of converting the approved target data request information into SQL statements through a preset model, the preset model is used to convert the natural language into SQL statements, the step further includes: When there is data request information for data query and / or data reading, obtaining the identity information of the user who initiated the data request information and the required data content of the data required by the data request information; Determine the queryable data content corresponding to the identity identification information based on the identity identification information and the preset authority information; If the queryable data content matches the required data content, the data request information is used as the approved target data request information; If the queryable data content does not match the required data content, an approval request is generated based on the identity identification information and the required data content, and the approval request is sent to a preset client, and feedback information for the approval request is obtained through the preset client; When the feedback information is used to indicate that the approval request is passed, the data request information is used as the target data request information that has passed the approval.

[0007] Optionally, the method further includes: Building a data set based on the target data request information and a SQL statement corresponding to the target data request information; Based on the computing power resource pool, the preset model is trained using the data set.

[0008] Optionally, the step of training the preset model through the data set based on the computing resource pool includes: Integrate the computing resources of each computing device in the computing system where the preset model is located to obtain a computing resource pool; When a training task for training the preset model is obtained, the preset model is trained using the data set based on the idle computing power of the computing power resource pool to complete the training task.

[0009] Optionally, the step of classifying and grading the data to be processed by using a preset classification and grading model to obtain the sensitivity level of the data to be processed includes: Using the NLP-based classification and grading model as the preset classification and grading model, and extracting labels from the data to be processed by an unsupervised learning algorithm in the NLP-based classification and grading model to obtain data labels for the data to be processed; The data labels of the data to be processed are identified and determined by a supervised learning algorithm in a classification and grading model based on NLP, so as to obtain the sensitivity level of the data to be processed.

[0010] Optionally, the step of determining a target protection strategy for the data to be processed according to the sensitivity level and a preset multi-level data protection strategy, and performing security protection on the data to be processed according to the target protection strategy through the data security management and control platform includes: Determine a target protection strategy for the data to be processed according to the sensitivity level and a preset multi-level data protection strategy; The data security management and control platform performs data protection processing on the data to be processed according to the target protection strategy to complete the security protection of the data to be processed, including: The data security management and control platform sets a functional component bound to the data to be processed according to the target protection strategy to perform data protection processing on the data to be processed, and the functional component is used to control the number of views and / or browsing duration of the data to be processed; And / or, obtaining operation information performed on the data to be processed, and generating watermark information according to the operation information and the identity identification information of the user who initiated the data request information, and setting a dark watermark for the data to be processed according to the watermark information through the data security management and control platform in accordance with the target protection strategy, so as to complete the watermarking of the data to be processed; and / or, identifying the data to be encrypted in the data to be processed, and encrypting the data to be encrypted in the data to be processed according to the target protection strategy setting through the data security management and control platform, so as to complete the encryption processing of the data to be processed; And / or, identifying sensitive data in the data to be processed, and performing synonymous replacement / masking on the sensitive data in the data to be processed according to the target protection strategy setting through the data security management and control platform to complete desensitization processing of the data to be processed; The data protection processing includes at least one of data desensitization processing, data encryption processing, setting the number of browsing times and / or browsing time for the data to be processed, and setting a watermark for the data to be processed.

[0011] Optionally, the method further includes: When a data download instruction for downloading the data to be processed is obtained, the data to be processed is approved based on a preset approval matrix; If the approval is passed, download the data to be processed after security protection in response to the data download instruction; If the approval fails, stop responding to the data download instruction.

[0012] Optionally, the method further includes: When the data in the database is read and / or queried, obtain the log information of each management system corresponding to the database; According to the log information, perform an audit process on the actions corresponding to the log information, including: Match the preset rules with the content data of the log information to determine whether the content data conforms to the preset rules, so as to complete the audit process of the actions corresponding to the log information; And / or, perform an association analysis on the log information of different management systems to complete the audit process of the actions corresponding to the log information.

[0013] In a second aspect, an embodiment of the present invention provides a data security management device implemented based on natural language conversion technology, characterized in that the device includes: A conversion module, configured to convert the approved target data request information into an SQL statement through a preset model, where the preset model is used to convert natural language into an SQL statement, including: extracting features of the approved target data request information through a feature extraction algorithm in the preset model to obtain natural language features, where the natural language features include context features, semantic features, and syntactic features; generating an SQL statement corresponding to the target data request information through an SQL statement generation algorithm in the preset model according to the context features, the semantic features, and the syntactic features; A query module, configured to query the data to be processed corresponding to the target data request information in the database according to the SQL statement and export the data to be processed to a data security control platform; A classification and grading module, configured to classify and grade the data to be processed through a preset classification and grading model to obtain the sensitivity level of the data to be processed; A protection module, configured to determine the target protection strategy of the data to be processed according to the sensitivity level and a preset multi-level data protection strategy, and control the data security control platform to perform security protection on the data to be processed according to the target protection strategy.

[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, where the electronic device includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data security management method based on natural language conversion technology in any of the embodiments of the present invention.

[0015] In a fourth aspect, an embodiment of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the data security management method based on natural language conversion technology in any of the embodiments of the present invention when executed by a computer processor.

[0016] The technical solution of the embodiment of the present invention, through a preset model, the preset model is used to convert natural language into SQL statements, and convert the target data request information passed through approval into SQL statements; query the data to be processed corresponding to the data request information in the database according to the SQL statements, and export the data to be processed to the data security control platform; classify and grade the data to be processed through a preset classification and grading model to obtain the sensitivity level of the data to be processed; determine the target protection strategy of the data to be processed according to the sensitivity level and the preset multi-level data protection strategy, and perform security protection on the data to be processed through the data security control platform according to the target protection strategy, improving data query and data acquisition efficiency, and avoiding the risk of human data leakage. Exporting the data to be processed to the data security control platform facilitates the security management / control of the data to be processed, and classifying and grading the data to be processed can flexibly perform targeted security protection on data with different sensitivity levels. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Among them: Figure 1 is a schematic flowchart of a data security management method based on natural language conversion technology in an embodiment of the present invention; Figure 2 is a schematic structural diagram of a data security management device based on natural language conversion technology in an embodiment of the present invention; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention; Figure 4 is a schematic structural diagram of a computer-readable storage provided by an embodiment of the present invention. Detailed implementation manners

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] The embodiment of the present invention provides a data security management method implemented based on natural language conversion technology. Through a preset model, the preset model is used to convert natural language into SQL statements, and the approved target data request information is converted into SQL statements, including: extracting features of the approved target data request information through a feature extraction algorithm in the preset model to obtain natural language features, where the natural language features include context features, semantic features, and syntactic features; generating an SQL statement corresponding to the target data request information according to the context features, the semantic features, and the syntactic features through an SQL statement generation algorithm in the preset model; querying the to-be-processed data corresponding to the data request information in the database according to the SQL statement, and exporting the to-be-processed data to a data security control platform; classifying and grading the to-be-processed data through a preset classification and grading model to obtain the sensitivity level of the to-be-processed data; determining the target protection strategy of the to-be-processed data according to the sensitivity level and a preset multi-level data protection strategy, and performing security protection on the to-be-processed data through the data security control platform according to the target protection strategy. Converting natural language requirements such as data request information into SQL statements through a preset model can improve the efficiency of data query and data acquisition and avoid the risk of human data leakage. Exporting the to-be-processed data to a data security control platform is convenient for the security management / control of the to-be-processed data. Classifying and grading the to-be-processed data can flexibly perform targeted security protection on data with different sensitivity levels.

[0021] In an embodiment, the embodiment of the present invention provides a data security management method implemented based on natural language conversion technology. The data security management method implemented based on natural language conversion technology in the embodiment of the present invention can be executed by a data security management device implemented based on natural language conversion technology. The data security management device implemented based on natural language conversion technology can be implemented by software and / or hardware.

[0022] As Figure 1 shown, the data security management method implemented based on natural language conversion technology in the embodiment of the present invention specifically includes the following steps: S110. Using a preset model for converting natural language into SQL statements, convert the approved target data request information into an SQL statement, including: extracting features of the approved target data request information through a feature extraction algorithm in the preset model to obtain natural language features, where the natural language features include context features, semantic features, and syntactic features; generating an SQL statement corresponding to the target data request information through an SQL statement generation algorithm in the preset model according to the context features, the semantic features, and the syntactic features; In a possible implementation manner, the step of using the preset model for converting natural language into SQL statements and converting the approved target data request information into an SQL statement includes: extracting features of the approved target data request information through a feature extraction algorithm in the preset model to obtain natural language features, where the natural language features include context features, semantic features, and syntactic features; generating an SQL statement corresponding to the target data request information through an SQL statement generation algorithm in the preset model according to the context features, the semantic features, and the syntactic features.

[0023] Exemplarily, the preset model takes the T5 model as an example. By combining synthetic data with the original labeled data, the T5 model is supervised and fine-tuned, taking into account both the data generation efficiency (reducing the cost of manual annotation) and the instruction fine-tuning adaptability of T5, which is suitable for the rapid iteration of the model. Specifically, the original labeled data is parsed to obtain label data, where the label data refers to <natural language description, executable SQL script> data pairs. Based on the original labeled data, the synthetic data is generated through data generation techniques, including but not limited to: synonym replacement: replacing words in the text with a synonym dictionary while keeping the sentence meaning unchanged. Lexical replacement: using near synonyms or related words to replace the original words to increase diversity. Inserting / deleting words: randomly inserting or deleting words in the sentence to simulate different expressions. Syntactic transformation: changing the syntactic structure of the sentence, such as changing from the passive voice to the active voice. At the same time, transformations are performed on SQL, including but not limited to: SQL statement restructuring: changing the expression of the SQL query, such as using different aggregation functions or condition combinations. Adding / deleting conditions: adding or deleting certain conditions in the query to generate new queries. Rewriting the SQL statement: expressing the same meaning using different SQL syntactic structures. The label data is obtained through the generation of synthetic samples. Specifically, the process of generating synthetic samples includes but not limited to: mixed generation: combining the natural language descriptions and SQL queries in multiple samples to form new pairs. Template filling: using a preset template and filling in different words into the template to generate new samples. Rule-driven generation: generating new samples according to predefined rules, such as changing the table names, field names, etc. in the SQL query.

[0024] Use the label data as the dataset for training the T5 model. Before training the T5 model with this dataset, perform noise removal processing on the dataset to delete irrelevant or invalid data records, such as duplicate records or obviously incorrect data; handle the missing values in the dataset to fill or delete the records containing missing values, which can be done by interpolation, deletion, etc.; perform standardized data format processing on the dataset to unify the data format, such as dates, currency units, etc., to ensure data consistency; remove the punctuation marks in the dataset to delete or replace the punctuation marks in the text and reduce interference factors; convert the case of the text in the dataset to unify the case of the text, usually converting to lowercase to reduce lexical diversity; perform stop word removal processing on the dataset to remove common words in natural language, such as 'de', 'shi', 'zai', etc., to reduce redundant information, and perform text tokenization on the dataset to split the sentence into individual words or tokens.

[0025] Exemplarily, the step of extracting natural language features from the approved target data request information through the feature extraction algorithm in the preset model, where the natural language features include context features, semantic features, and syntactic features, specifically includes: the process of context feature extraction, the process of semantic feature extraction, and the process of syntactic feature extraction; Among them, the process of context feature extraction is as follows: identifying keywords, that is, identifying the key vocabulary in the sentence through word segmentation technology, such as table names, field names, aggregation functions, etc.; extracting context information: analyzing the context information in the sentence, that is, identifying the relationships between keywords, such as conjunctions, sequence words, demonstrative words, etc.; context window, that is, constructing a context window for each keyword to capture the surrounding lexical information to better understand its contextual meaning; entity recognition, that is, using named entity recognition (NER) technology to identify the entity names in the sentence, such as table names, field names, etc.

[0026] The process of syntactic feature extraction is as follows: sentence structure analysis, that is, using dependency parsing or constituency parsing technology to extract the syntactic structure of the sentence; part-of-speech tagging, that is, tagging the part of speech of each word in the sentence (such as noun, verb, adjective, etc.) to help understand the role of the word in the sentence; syntactic pattern matching, that is, identifying specific syntactic patterns in the sentence, such as passive voice, comparative degree, etc.; syntactic error detection, that is, checking the syntactic errors in the sentence and trying to correct or prompt the errors.

[0027] The process of semantic feature extraction is as follows: semantic role labeling, that is, identifying the predicates and their arguments in the sentence, that is, the actions (predicates) and participants (arguments), to help understand the core meaning of the sentence; semantic similarity calculation, that is, using word vectors (such as Word2Vec, GloVe) or context-sensitive embeddings (such as BERT, RoBERTa) to calculate the semantic similarity between words; logical form conversion, that is, converting the natural language description into a logical form (such as first-order logic) for subsequent SQL query generation; semantic relationship analysis, that is, analyzing the semantic relationships between words in the sentence, such as causal relationship, parallel relationship, turning relationship, etc.

[0028] Exemplarily, perform syntax checking, security checking, and performance checking on the generated SQL statement corresponding to the target data request information. Among them, the syntax checking specifically includes: syntax verification, that is, checking whether the SQL script conforms to the SQL standard syntax, including the correct use of keywords, the integrity of the syntax structure, etc.; syntax tree construction, that is, using an abstract syntax tree (AST) to represent the structure of the SQL statement to ensure that all syntax elements are correctly parsed; error location and reporting, that is, when a syntax error is found, accurately pointing out the location of the error and providing detailed error information for debugging; version compatibility checking, that is, ensuring that the SQL script can be correctly executed in different database management systems (DBMS).

[0029] The security checking specifically includes: SQL injection prevention, that is, detecting whether there is a risk of SQL injection in the SQL script, such as generating SQL statements by string concatenation; input validation, that is, ensuring that all inputs are properly validated to prevent security issues caused by malicious inputs; permission checking, that is, checking whether the operations involved in the SQL script exceed the user's due permission scope; sensitive data protection, that is, ensuring that the SQL script does not expose or leak sensitive data, such as usernames, passwords, etc.

[0030] The performance checking specifically includes: index usage checking, that is, evaluating whether the SQL query makes full use of the existing indexes to speed up the query; query optimization, that is, analyzing the execution plan of the SQL query, finding possible performance bottlenecks, and proposing improvement suggestions; statistical information utilization, that is, ensuring that the SQL query reasonably uses the statistical information of the database for optimization; avoiding full table scans, that is, detecting whether full table scans can be avoided and instead using a more efficient query strategy.

[0031] S120. Query the data to be processed corresponding to the target data request information in the database and export the data to be processed to the data security control platform; Exemplarily, when exporting the data to be processed to the data security control platform, encrypt the data to be processed and store it in the data security control platform. Specifically, use the transparent encryption method based on the national cryptographic algorithm SM4 for data storage encryption, that is, before data storage, encrypt the data through the encryption layer to achieve data disk encryption and ensure the confidentiality of the data.

[0032] S130. Classify and grade the data to be processed through a preset classification and grading model to obtain the sensitivity level of the data to be processed; In a possible implementation manner, the step of classifying and grading the data to be processed through a preset classification and grading model to obtain the sensitivity level of the data to be processed includes: Use the NLP-based classification and grading model as the preset classification and grading model, and extract labels for the data to be processed through the unsupervised learning algorithm in the NLP-based classification and grading model to obtain the data labels of the data to be processed; Determine the sensitivity level of the data to be processed by performing identity determination on the data labels of the data to be processed through the supervised learning algorithm in the NLP-based classification and grading model.

[0033] Exemplarily, after the data is stored in the data security control platform, the file content is actively scanned, and according to the sensitive data label characteristics, the sensitive data in the document is automatically discovered, and the sensitive level of the file is classified and graded.

[0034] Specifically, by using keywords and regular expressions to scan the file data content, the characteristics of the data to be processed are identified, fully ensuring compliance requirements. By setting more than 80 data characteristic rules, 90% of personal information can be automatically identified. The sensitive content scanning of the document supports xls(x), doc(x), ppt(x), PDF, csv, text files and compressed files (without password) of the above files.

[0035] Among them, the process of keyword matching can be understood as follows: taking name, country, province, and city as representatives, when the data can accurately match the corresponding keywords, the data characteristics are determined. Therefore, by establishing a supporting data dictionary, such as the hundred family names, countries, province names, city names, etc. It supports adding sensitive control by directly matching keywords, and can support adding sensitive keywords in the OR and AND ways. For example, if "customer information OR ID number OR mobile phone number" is added to the sensitive keywords, then when a user uploads a file to the data security control platform, the system will perform sensitive scanning on the document according to the sensitive scanning policy. If any of the words "customer information", "ID number", or "mobile phone number" appears in the file, the corresponding sensitive attribute will be added to the document, and when the user performs document operations on the document, the system will perform document permission control according to the control policy issued for the sensitive label.

[0036] The process of regular expression matching can be understood as follows: taking mobile phone number, ID number, and license number as representatives, such data has a fixed length, fixed format, and fixed data range. Through regular expression configuration and parsing, the data characteristics are determined. The sensitive rules matched by regular expressions can wildcard sensitive information such as name, ID number, and mobile phone number. For example, the regular expression for defining the ID number matching rule can be: (?<=\D{1}|\s|^)(1[1-5]|2[1-3]|3[1-7]|4[1-6]|5[0-4]|6[1-5])[0-9]{15}([0-9]|X|x)(?=\D{1}|\s|$) The ID number can be identified through the regular expression of the above ID number matching rule, so as to add the sensitive attribute of the document, so that the system can follow up the sensitive attribute of the document for document permission control in the later stage.

[0037] Exemplarily, for the scenario of unstructured data flow, an unstructured data classification and grading model based on NLP is adopted, and an unsupervised learning and supervised learning combined method is used to realize the identification of unstructured data.

[0038] Specifically, when facing long texts of unstructured data, considering the model accuracy and processing efficiency, it is necessary to convert them into short texts. Therefore, the unstructured data is parsed to extract important text information as the theme summary.

[0039] Secondly, unstructured data has no label information and a large quantity, and it is difficult to manually label. Therefore, it is considered to give priority to using unsupervised learning methods for label extraction. In this process, the prediction results of the unsupervised learning model are used to gradually accumulate label data. When the label data accumulates to a certain amount, a supervised learning model is used to further optimize the algorithm. By comprehensively applying a variety of NLP technologies, a classification and grading model with stronger scalability and higher accuracy is constructed. Among them, the unsupervised learning model can be understood as an algorithm model constructed based on the unsupervised learning algorithm, and the supervised learning model can be understood as an algorithm model constructed based on the supervised learning algorithm.

[0040] The specific method of unsupervised learning is as follows: The contrast learning model extracts the text vectors of unstructured data, calculates the similarity between the text vectors, and divides the data with high similarity into the same classification data, that is, the same class of label data; Perform part-of-speech analysis and keyword matching on unstructured data, and divide the data with successful matching into the corresponding label data; Multiple unsupervised learning models are fused to obtain the comprehensive label data result.

[0041] The specific method of supervised learning is as follows: The text retrieval algorithm model gives the label data most similar to the current unstructured data, and uses these label data as the basis for classifying the current unstructured data; Perform text classification on the current unstructured data according to the label data to obtain a fine-grained classification result; Multiple supervised learning models are fused to obtain a multi-level classification result.

[0042] Exemplarily, the (global) and sensitive scanning strategies for scanning can be defined. The definition of the sensitive scanning strategy includes the document sensitive level, sensitive rules, the number of sensitive data items, etc.

[0043] Exemplarily, according to the formulated sensitivity rules and scanning strategies, the system identifies sensitive content and defines sensitivity levels for stored personal documents. The sensitivity levels can be set to extremely sensitive, sensitive, relatively sensitive, low sensitivity, and other file sensitivity levels; Exemplarily, the documents may be classified according to data content based on the identification of the document content, and the classifications may be set such as: user identity related data, user service content data, user service derived data, and enterprise operation management data.

[0044] S140. Determine a target protection strategy for the data to be processed according to the sensitivity level and a preset multi-level data protection strategy, and perform security protection on the data to be processed according to the target protection strategy through the data security management and control platform.

[0045] In a possible implementation, the step of determining a target protection strategy for the data to be processed according to the sensitivity level and the preset multi-level data protection strategy, and performing security protection on the data to be processed according to the target protection strategy through the data security management and control platform includes: Determine a target protection strategy for the data to be processed according to the sensitivity level and a preset multi-level data protection strategy; The data security management and control platform performs data protection processing on the data to be processed according to the target protection strategy to complete the security protection of the data to be processed, including: The data security management and control platform sets a functional component bound to the data to be processed according to the target protection strategy to perform data protection processing on the data to be processed, and the functional component is used to control the number of views and / or browsing duration of the data to be processed; And / or, obtaining operation information performed on the data to be processed, and generating watermark information according to the operation information and the identity identification information of the user who initiated the data request information, and setting a dark watermark for the data to be processed according to the watermark information through the data security management and control platform in accordance with the target protection strategy, so as to complete the watermarking of the data to be processed; And / or, identifying the data to be encrypted in the data to be processed, and encrypting the data to be encrypted in the data to be processed according to the target protection strategy setting through the data security management and control platform to complete the encryption processing of the data to be processed; And / or, identifying sensitive data in the data to be processed, and performing synonymous replacement / masking on the sensitive data in the data to be processed according to the target protection strategy setting through the data security management and control platform to complete desensitization processing of the data to be processed; The data protection processing includes at least one of data desensitization processing, data encryption processing, setting the number of browsing times and / or browsing time for the data to be processed, and setting a watermark for the data to be processed.

[0046] Exemplarily, the sensitive data desensitization function module of the data security management and control platform is called to scan the file content for sensitive keywords. After the sensitive information is found, the scanned specified content is desensitized, supporting synonym replacement, data masking, etc.

[0047] The data security management and control platform is called upon to provide watermark technology embedded in the file carrier to display the sensitive content in the file with a watermark, that is, to add a watermark to the display page of the data to be processed. The watermark content can be configured according to the preset watermark strategy, including watermark content, font, transparency, tilt, etc.

[0048] At the same time, the data security management and control platform can also be called to provide dark watermark (invisible watermark) technology embedded in the frequency domain layer, which ensures the robustness of the watermark. By generating a watermark layer and loading it on the frequency domain layer of the file, it does not affect the normal opening, reading and writing of the file, and is not easy to delete. At the same time, the watermark information is generated by the operation information and the identity information of the user who initiated the data request information, so that the source can be accurately traced after the data leak occurs.

[0049] In a possible implementation manner, before the step of converting the natural language into SQL statements by using a preset model, and converting the approved target data request information into SQL statements, the step further includes: When there is data request information for data query and / or data reading, obtaining the identity information of the user who initiated the data request information and the required data content of the data required by the data request information; Determine the queryable data content corresponding to the identity identification information based on the identity identification information and the preset authority information; If the queryable data content matches the required data content, the data request information is used as the approved target data request information; If the queryable data content does not match the required data content, an approval request is generated based on the identity identification information and the required data content, and the approval request is sent to a preset client, and feedback information for the approval request is obtained through the preset client; When the feedback information is used to indicate that the approval request is passed, the data request information is used as the target data request information that has passed the approval.

[0050] Exemplarily, the data request information for data query and / or data reading can be understood as a work order application. The user initiates a data query and / or data reading application requirement through the work order system, selects the business system to be queried in the requirement application, and details the data query and / or data reading requirements (e.g., count the number of users passing through base station CGI 123456789 and 112345678 from 10:30 on December 27, 2024 to 14:50 on December 29, 2024), and provides relevant attachments. After completing the requirement approval on the work order system, if the approval result is "passed", the work order requirement information will be sent to the preset model. Initiating requirement approval through the online work order system can effectively control the rationalization and standardization of data query and / or data reading requirements, and avoid the risk of data leakage caused by illegally obtaining data.

[0051] Exemplarily, the preset permission information includes different identity identification information and the accessible data content corresponding to different identity identification information. For example, the data content corresponding to the identity identification information a of type A is X and Y. When A initiates a request to extract data X and Y, it is automatically determined that the request is supported (approval passed); when A initiates a request to extract data Z, this request is automatically rejected. If A continues to initiate a query request for the data with data content Z after the request of type A is rejected, then the query request initiated by type A for the data with data content Z is sent as an approval request to the preset client, and feedback information for the approval request is obtained through the preset client. Through the preset permission information, the approval efficiency and accuracy can be improved, that is, compliant data requests do not require manual approval.

[0052] In a possible implementation manner, the method further includes: Construct a data set based on the target data request information and the SQL statement corresponding to the target data request information; Based on the computing power resource pool, train the preset model through the data set.

[0053] Exemplarily, continuously training the preset model is beneficial to ensuring the accuracy of the output result of the preset model.

[0054] In a possible implementation manner, the step of training the preset model through the data set based on the computing power resource pool, where the preset model is used to convert natural language into an SQL statement, includes: Integrate the computing power resources of each computing device in the computing system where the preset model is located to obtain a computing power resource pool; When obtaining a training task for training the preset model, based on the idle computing power of the computing power resource pool, train the preset model through the data set to complete the training task.

[0055] Exemplarily, the process of model training often requires a large amount of computing power support, and the preset model is often used during the day (working hours). If the preset model is trained during the day (working hours), it will affect user usage. Therefore, by combining the change data of the idle resource data in the computing power resource pool and the working rules of the preset model, a target time period is selected to train the preset model.

[0056] Exemplarily, the GPU resources of multiple machines are integrated into a large resource pool (computing power resource pool). According to the actual task requirements, resources are intelligently allocated. For "large tasks", multiple GPUs are called across machines for parallel computing. For "small tasks", they are run whenever there is an opportunity to avoid resource idleness. For example, if the model training task is regarded as a large task and the model inference task is regarded as a small task, then the computing power of the computing power resource pool is dynamically allocated according to the task volume of the model training task and the task volume of the model inference.

[0057] In a possible implementation manner, the method further includes: When the data in the database is read and / or queried, obtain the log information of each management system corresponding to the database; Perform audit processing on the actions corresponding to the log information according to the log information, including: Match the preset rules with the content data of the log information to determine whether the content data conforms to the preset rules, so as to complete the audit processing of the actions corresponding to the log information; And / or, perform correlation analysis on the log information of different management systems to complete the audit processing of the actions corresponding to the log information.

[0058] The audit includes but is not limited to: operation log audit, exception audit, and bypass audit.

[0059] Exemplarily, provide the full-link log records of all operations such as uploading, accessing, browsing, sharing, downloading, and exporting the data to be processed to the log audit platform. On the log audit platform, according to the association with the document as the object or the user as the object, realize the query and traceability of the transmission operation track of the sensitive data document for evidence collection.

[0060] Operation log audit: The system records all user logs such as export, preview, edit, share, download, etc. Through the strong association between people and data, realize the traceability audit ability of "tracking by numbers" and "tracking numbers by people".

[0061] Exception audit: Analyze and count behaviors such as high-frequency access, high user download volume, high turnover volume, and external network access.

[0062] Bypass audit: By correlating and analyzing the logs of the work order system, 4A system, and data security control platform, audit the bypass of the work order system and the anti-bypass behavior of the 4A system.

[0063] Exemplarily, when the preset rule specifies a data path, for example, it is specified that data is transmitted from the first address (IP1) to the second address (IP2), then identify whether the data sending address in the content data of the log information is the first address (IP1), and whether the data receiving address in the content data of the log information is the second address (IP2), so as to realize the audit processing for the action corresponding to the log information.

[0064] Exemplarily, perform correlation analysis on the log information of different management systems, that is, when there is target log information indicating the reading of the data N to be processed in any management system, query whether there is log information corresponding to the target log information in other management systems, so as to realize the audit processing for the action corresponding to the log information.

[0065] In a possible implementation manner, the method further includes: When obtaining a data download instruction for downloading the data to be processed, approve the data to be processed based on a preset approval matrix; If the approval is passed, respond to the data download instruction to download the data to be processed after completing security protection; If the approval is not passed, stop responding to the data download instruction.

[0066] Exemplarily, in the face of a download request (data download instruction), approve the data to be processed according to a preset approval matrix, further prevent data leakage, and improve data security.

[0067] In a second aspect, as Figure 2 shown, an embodiment of the present invention provides a data security management device implemented based on natural language conversion technology, characterized in that the device includes: A conversion module 201, configured to convert the approved target data request information into an SQL statement through a preset model, where the preset model is used to convert natural language into an SQL statement, including: extracting features of the approved target data request information through a feature extraction algorithm in the preset model to obtain natural language features, where the natural language features include context features, semantic features, and syntactic features; generating an SQL statement corresponding to the target data request information through an SQL statement generation algorithm in the preset model according to the context features, the semantic features, and the syntactic features; A query module 202, configured to query the to-be-processed data corresponding to the target data request information in a database according to the SQL statement, and export the to-be-processed data to a data security control platform; A classification and grading module 203, configured to classify and grade the to-be-processed data through a preset classification and grading model to obtain the sensitivity level of the to-be-processed data; A protection module 204, configured to determine a target protection policy for the to-be-processed data according to the sensitivity level and a preset multi-level data protection policy, and perform security protection on the to-be-processed data through the data security control platform according to the target protection policy.

[0068] It should be noted that the various modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional modules are only for the convenience of mutual distinction and are not used to limit the protection scope of the embodiments of the present invention.

[0069] In another embodiment of the present invention, an electronic device is further provided. Figure 3 The block diagram of an exemplary electronic device 50 suitable for implementing the embodiment mode of the embodiment of the present invention is shown. Figure 3 The shown electronic device 50 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0070] As Figure 3 shown, the electronic device 50 is presented in the form of a general-purpose computing device. The components of the electronic device 50 may include, but are not limited to: one or more processors or processing units 501, a system memory 502, and a bus 503 connecting different system components (including the system memory 502 and the processing unit 501).

[0071] The bus 503 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0072] The electronic device 50 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 50, including volatile and non-volatile media, removable and non-removable media.

[0073] System memory 502 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 504 and / or cache memory 505. Electronic device 50 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 506 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 3 not shown and typically called a "hard disk drive"). Although Figure 3 not shown in, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to bus 503 through one or more data media interfaces. Memory 502 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0074] A program / utility 508 having a set (at least one) of program modules 507 can be stored, for example, in memory 502. Such program modules 507 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. An implementation of a network environment may be included in each or some combination of these examples. Program modules 507 generally execute the functions and / or methods in the embodiments described in the present invention.

[0075] Electronic device 50 may also communicate with one or more external devices 509 (such as a keyboard, a pointing device, a display 510, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 50, and / or communicate with any device that enables the electronic device 50 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 511. And, electronic device 50 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 512. As shown in the figure, network adapter 512 communicates with other modules of electronic device 50 through bus 503. It should be understood that although Figure 3 not shown in, other hardware and / or software modules can be used in conjunction with electronic device 50, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0076] The processing unit 501 executes various functional applications and data processing by running the programs stored in the system memory 502, for example, implementing the data security management method provided by the embodiments of the present invention based on natural language conversion technology.

[0077] In another embodiment of the present invention, as Figure 4 shown, there is also provided a storage medium 400 containing a computer program 411, and the computer program 411 is used to execute a data security management method implemented based on natural language conversion technology when executed by a computer processor. The method includes: through a preset model, the preset model is used to convert natural language into SQL statements, and convert the approved target data request information into SQL statements, including: Performing feature extraction on the approved target data request information through a feature extraction algorithm in the preset model to obtain natural language features, where the natural language features include context features, semantic features, and syntactic features; Generating an SQL statement corresponding to the target data request information according to the context features, the semantic features, and the syntactic features through an SQL statement generation algorithm in the preset model; Querying the data to be processed corresponding to the target data request information in the database according to the SQL statement, and exporting the data to be processed to a data security control platform; Classifying and grading the data to be processed through a preset classification and grading model to obtain the sensitivity level of the data to be processed; Determining the target protection strategy for the data to be processed according to the sensitivity level and a preset multi-level data protection strategy, and performing security protection on the data to be processed according to the target protection strategy through the data security control platform.

[0078] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0079] The computer-readable signal media may include data signals propagated in a baseband or as part of a carrier wave, which carry computer-readable program codes. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, and this computer-readable media can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0080] The program codes contained on the computer-readable media can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0081] The computer program codes for performing the operations of the embodiments of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program codes can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0082] The above-disclosed is only the preferred embodiment of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A data security management method based on natural language conversion technology, characterized in that: include: By using a preset model, the preset model is used to convert natural language into SQL statements, and the approved target data request information is converted into SQL statements, including: Extract features from the approved target data request information using a feature extraction algorithm in a preset model to obtain natural language features, where the natural language features include context features, semantic features, and grammatical features; Generate an SQL statement corresponding to the target data request information through an SQL statement generation algorithm in the preset model according to the context feature, the semantic feature and the grammatical feature; According to the SQL statement, query the database for the data to be processed corresponding to the target data request information, and export the data to be processed to the data security management and control platform; Classify and grade the data to be processed by using a preset classification and grading model to obtain a sensitivity level of the data to be processed; The target protection strategy for the data to be processed is determined according to the sensitivity level and the preset multi-level data protection strategy, and the data to be processed is securely protected according to the target protection strategy through the data security management and control platform.

2. The data security management method based on natural language conversion technology according to claim 1 is characterized in that: Before the step of converting the natural language into SQL statements by using a preset model and converting the approved target data request information into SQL statements, the method further includes: When there is data request information for data query and / or data reading, obtaining the identity information of the user who initiated the data request information and the required data content of the data required by the data request information; Determine the queryable data content corresponding to the identity identification information based on the identity identification information and the preset authority information; If the queryable data content matches the required data content, the data request information is used as the approved target data request information; If the queryable data content does not match the required data content, an approval request is generated based on the identity identification information and the required data content, and the approval request is sent to a preset client, and feedback information for the approval request is obtained through the preset client; When the feedback information is used to indicate that the approval request is passed, the data request information is used as the target data request information that has passed the approval.

3. The data security management method based on natural language conversion technology as claimed in claim 1, characterized in that: The method further comprises: Building a data set based on the target data request information and a SQL statement corresponding to the target data request information; Based on the computing power resource pool, the preset model is trained using the data set.

4. The data security management method based on natural language conversion technology as claimed in claim 3 is characterized in that: The step of training the preset model through the data set based on the computing resource pool includes: Integrate the computing resources of each computing device in the computing system where the preset model is located to obtain a computing resource pool; When a training task for training the preset model is obtained, the preset model is trained using the data set based on the idle computing power of the computing power resource pool to complete the training task.

5. The data security management method based on natural language conversion technology as claimed in claim 1, characterized in that: The step of classifying and grading the data to be processed by using a preset classification and grading model to obtain the sensitivity level of the data to be processed includes: Using the NLP-based classification and grading model as the preset classification and grading model, and extracting labels from the data to be processed by an unsupervised learning algorithm in the NLP-based classification and grading model to obtain data labels for the data to be processed; The data labels of the data to be processed are identified and determined by a supervised learning algorithm in a classification and grading model based on NLP, so as to obtain the sensitivity level of the data to be processed.

6. The data security management method based on natural language conversion technology as claimed in claim 1, characterized in that: The step of determining the target protection strategy for the data to be processed according to the sensitivity level and the preset multi-level data protection strategy, and performing security protection on the data to be processed according to the target protection strategy through the data security management and control platform includes: Determine a target protection strategy for the data to be processed according to the sensitivity level and a preset multi-level data protection strategy; The data security management and control platform performs data protection processing on the data to be processed according to the target protection strategy to complete the security protection of the data to be processed, including: The data security management and control platform sets a functional component bound to the data to be processed according to the target protection strategy to perform data protection processing on the data to be processed, and the functional component is used to control the number of views and / or the browsing time of the data to be processed; and / or, obtaining operation information performed on the data to be processed, and generating watermark information according to the operation information and the identity information of the user who initiated the data request information, and setting a dark watermark for the data to be processed according to the watermark information through the data security management and control platform in accordance with the target protection strategy, so as to complete setting of a watermark for the data to be processed; and / or, identifying the data to be encrypted in the data to be processed, and encrypting the data to be encrypted in the data to be processed according to the target protection strategy setting through the data security management and control platform, so as to complete the encryption processing of the data to be processed; And / or, identifying sensitive data in the data to be processed, and performing synonymous replacement / masking on the sensitive data in the data to be processed according to the target protection strategy setting through the data security management and control platform to complete desensitization processing of the data to be processed; The data protection processing includes at least one of data desensitization processing, data encryption processing, setting the number of browsing times and / or browsing time for the data to be processed, and setting a watermark for the data to be processed.

7. The data security management method based on natural language conversion technology as claimed in claim 6 is characterized in that: The method further comprises: When a data download instruction for downloading the data to be processed is obtained, the data to be processed is approved based on a preset approval matrix; If approved, respond to the data download instruction to download the data to be processed after completing security protection; If the approval is not passed, stop responding to the data download instruction.

8. The data security management method based on natural language conversion technology as claimed in claim 1, characterized in that: The method further comprises: When the data in the database is read and / or queried, log information of each management system corresponding to the database is obtained; According to the log information, auditing the action corresponding to the log information includes: Matching the preset rules with the content data of the log information to determine whether the content data conforms to the preset rules, so as to complete the audit processing of the action corresponding to the log information; And / or, log information of different management systems is correlated and analyzed to complete audit processing of actions corresponding to the log information.

9. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to execute the data security management method based on natural language conversion technology as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • SQL statement generation method, apparatus, computer device, and storage medium

    CN109408526A

  • Data processing method and device

    CN114218607A

  • Data outgoing method based on data security automatic classification and grading result

    CN116383693A

  • SQL (Structured Query Language) statement generation method and device based on large language model, equipment and medium

    CN119201966A

  • Query system, method and equipment based on conversion from natural language to SQL (Structured Query Language) and medium

    CN119293072A