A data security compliance assessment and identification method and system

By collecting and preprocessing regulatory text data, using word embedding and named entity recognition to generate feature vectors, and evaluate and identify based on machine learning algorithms, solving the flexibility and stability problems of traditional evaluation models and detection methods, and achieving efficient and accurate data security compliance evaluation and identification.

CN119577120BActive Publication Date: 2025-05-27BEIJING SINOBANG DIGITAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510143066.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-27
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The traditional data security compliance evaluation model relies on fixed threshold settings and is difficult to adapt to changes in different scenarios, resulting in inflexible and accurate evaluation results. The existing anomaly detection methods are difficult to cope with complex and changeable regulatory text environments, resulting in instability in the detection results.

Method used

Regulatory text data is collected through the API interface, and the natural language processing library NLTK is used for pre-processing. Pre-calculated word embedding method and named entity recognition NER method are used to generate numeric vectors and key concept texts. An evaluation model and recognition model are established based on machine learning algorithms, and weighted average method and clustering algorithm are used for evaluation and recognition.

Benefits of technology

It realizes flexible and accurate evaluation of the text data of the regulations, can quickly respond to changes in laws and regulations, improves the timeliness and compliance of the system, and improves the effectiveness and practicality of the identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577120B_ABST
    Figure CN119577120B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for data security compliance assessment and identification, which relates to the technical field of data security compliance assessment and identification. It includes collecting regulatory text data through an API interface, preprocessing the regulatory text data using the natural language processing library NLTK to obtain formatted regulatory text data; using a pre-computed word embedding method to map each word in the formatted regulatory text data to a fixed-length vector and generate a digital vector; using the named entity recognition NER method to mark important terms in the digital vector to obtain key concept text; establishing an evaluation model based on a machine learning algorithm, inputting the digital vector into the evaluation model, calculating the output evaluation value using the weighted average method, setting an evaluation threshold and comparing it with the evaluation value to determine whether the digital vector is compliant and obtaining an evaluation result; constructing an identification model based on an anomaly detection algorithm and inputting the key concept text into the identification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security compliance evaluation and identification, and particularly to a data security compliance evaluation and identification method and system. Background Art

[0002] Data security compliance evaluation and identification technology refers to the technology of automatically conducting compliance evaluation and anomaly detection on various regulatory texts, enterprise internal policies, and operation processes by applying advanced data analysis and machine learning algorithms. Its core goal is to ensure that enterprises and organizations strictly comply with relevant laws and regulations during operation and avoid legal risks and economic losses caused by violations.

[0003] In the field of data security compliance evaluation and identification technology, traditional evaluation models rely on fixed threshold settings, making it difficult to adapt to changes in different scenarios, resulting in inflexible and inaccurate evaluation results. Moreover, existing anomaly detection methods mostly rely on a single algorithm, making it difficult to cope with the complex and changing regulatory text environment, resulting in unstable detection results. At the same time, in the existing technology, compliance evaluation and vulnerability information are usually isolated, making it difficult to form an effective linkage mechanism, resulting in potential security risks not being discovered and repaired in a timely manner. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a data security compliance evaluation and identification method to solve the problem that traditional evaluation models rely on fixed threshold settings, making it difficult to adapt to changes in different scenarios, resulting in inflexible and inaccurate evaluation results.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a data security compliance evaluation and identification method, which includes:

[0008] Collect regulatory text data through an API interface, and preprocess the regulatory text data using the natural language processing library NLTK to obtain formatted regulatory text data;

[0009] Adopt a pre-computed word embedding method to map each word in the formatted regulatory text data to a fixed-length vector and generate a digital vector;

[0010] Adopt the named entity recognition NER method to mark important terms in the digital vector to obtain key concept text;

[0011] Build an evaluation model based on a machine learning algorithm, input the digital vector into the evaluation model, and calculate the output evaluation value using the weighted average method;

[0012] Set the evaluation threshold and compare it with the evaluation value to determine whether the digital vector is compliant and obtain the evaluation result;

[0013] Build an identification model based on the anomaly detection algorithm, input the key concept text into the identification model, and use the clustering algorithm to output the identification result.

[0014] As a preferred solution of the data security compliance evaluation and identification method described in the present invention, wherein: the regulatory text data is collected through the API interface, and the natural language processing library NLTK is used to preprocess the regulatory text data to obtain formatted regulatory text data. The specific steps are as follows:

[0015] Use the requests library to send an HTTP protocol request to the API endpoint to obtain regulatory text data;

[0016] Use the natural language processing library NLTK to filter out meaningless words in the regulatory text data and reduce noise;

[0017] Restore the words in the regulatory text data and split the text in the regulatory text data into individual words to obtain formatted regulatory text data.

[0018] As a preferred solution of the data security compliance evaluation and identification method described in the present invention, wherein: the pre-computed word embedding method is used to map each word in the formatted regulatory text data to a fixed-length vector and generate a digital vector. The specific steps are as follows:

[0019] Use the word embedding library Word2Vec to perform word segmentation on the formatted regulatory text data, collect all the independent words after word segmentation, and form a non-repeating vocabulary;

[0020] Create an empty set and fill it with the vocabulary, count the occurrence frequency of the independent words, sort the vocabulary according to the occurrence frequency of the independent words, and form the final vocabulary;

[0021] For the independent words in the final vocabulary, find the corresponding fixed-length vector in the word embedding library Word2Vec, and use the found corresponding fixed-length vector as the representation of the independent word to form a word-level digital vector;

[0022] Create a zero vector with the same dimension as the word-level digital vector as an accumulator, accumulate all the word-level digital vectors, and for each independent word in the final vocabulary, obtain its corresponding word-level digital vector;

[0023] Add each word-level digital vector to the accumulator one by one to obtain the average word vector;

[0024] The obtained average word vectors are used as the representation of the document-level digital vectors, obtaining the document-level digital vectors and marking them as digital vectors.

[0025] As a preferred solution of the data security compliance evaluation and identification method described in the present invention, wherein: the important terms in the digital vectors are marked by using the named entity recognition NER method to obtain the key concept text, and the specific steps are as follows:

[0026] Using the named entity recognition NER model, the formatted regulatory text data is input sentence by sentence or paragraph by paragraph into the named entity recognition NER model;

[0027] The named entity recognition NER model scans the formatted regulatory text data and identifies the entities;

[0028] For the identified entities, the named entity recognition NER model assigns different category labels to them and outputs the text information with annotations;

[0029] The output text information with annotations is marked as the important terms in the text and forms the key concept text.

[0030] As a preferred solution of the data security compliance evaluation and identification method described in the present invention, wherein: an evaluation model is established based on the machine learning algorithm, the digital vectors are input into the evaluation model, and the weighted average method is used to calculate the output evaluation value. The specific steps are as follows:

[0031] The formatted regulatory text data is divided into a training set and a validation set;

[0032] Using the scikit-learn library, import the LogisticRegression class to create a logistic regression instance, and combine it with the formatted regulatory text data to construct an evaluation model and set the initial parameter weights of the evaluation model ;

[0033] The digital vectors in the formatted regulatory text data are input into the evaluation model, and the weighted average method is used to calculate the evaluation value , and the expression is:

[0034] ;

[0035] wherein, is the word weight, is the weight of the th word, is the th word vector, is the Sigmoid activation function, is the evaluation value, is the formatted regulatory text data, is the digital vector in the formatted regulatory text data;

[0036] As a preferred solution of the data security compliance assessment and identification method described in the present invention, wherein: comparing the set evaluation threshold with the evaluation value to determine whether the digital vector is compliant and obtaining an evaluation result, the specific steps are as follows:

[0037] Set the threshold based on the formatted regulatory text data that has never participated in training in the validation set ;

[0038] When ≥ , it indicates that the digital vector in the formatted regulatory text data is lower than the set threshold , then it is determined that the regulatory text is compliant;

[0039] When < , it is determined that the digital vector in the formatted regulatory text data is non-compliant, then the formatted regulatory text data is modified. After modification, the digital vector in the formatted regulatory text data is recalculated according to the digital vector expression, and the recalculated digital vector in the formatted regulatory text data is re-input into the evaluation model to calculate the evaluation value. Set the number of iterations to , until ≥ .

[0040] As a preferred solution of the data security compliance assessment and identification method described in the present invention, wherein: constructing an identification model based on the anomaly detection algorithm, inputting the key concept text into the identification model, and outputting the identification result by using the clustering algorithm, the specific steps are as follows:

[0041] Construct an identification model based on the Isolation Forest anomaly detection algorithm and the formatted regulatory text data;

[0042] Use the Isolation Forest algorithm to isolate observation points by constructing multiple decision trees, and initialize the parameters of the identification model by using the IsolationForest class in the Python library;

[0043] Convert the key concept text in the formatted regulatory text data into TF-IDF values, calculate the TF-IDF vector values of the key concept text in the formatted regulatory text data, and perform standardization processing on the key concept text in the formatted regulatory text data represented by the TF-IDF vector values. The expression is:​

[0044] ;

[0045] Among them, is a word, is the key concept text in the formatted regulatory text data in it, is all the key concept texts, is the total number of documents in all the key concept texts, is the number of documents containing the word ; is the word in the formatted regulatory text data in the key concept text frequency, is the TF-IDF vector value;

[0046] Input the TF-IDF vector value into the recognition model, and use the K-means clustering algorithm to classify the output of the recognition model. The expression is:

[0047] ;

[0048] Among them, is the clustering label, is the anomaly score matrix, represents the anomaly score matrix in the th sample anomaly score vector, is the th cluster center point, is the set number of clusters, is the formatted regulatory text data in the key concept text, represents the key text concept anomaly score, is the number of trees in the isolation forest algorithm, is the anomaly score matrix in the element, is the th cluster, is the isolation forest algorithm for path length, is the maximum path length of all trees in the isolation forest, is expected value, that is, the average value of the path length;

[0049] Set the anomaly score threshold and the anomaly score matrix for comparison;

[0050] When ≥ it indicates that the key text concepts in the formatted regulatory text data are non - compliant;

[0051] When < it indicates that the key text concepts in the formatted regulatory text data are compliant.

[0052] In a second aspect, the present invention provides a data security compliance evaluation and identification system, including: a data processing module, a word embedding generation module, a key concept extraction module, an evaluation model module, and an anomaly detection module;

[0053] The data processing module is used to collect regulatory text data through an API interface and pre - process the regulatory text data using the natural language processing library NLTK to obtain formatted regulatory text data;

[0054] The word embedding generation module is used to map each word in the formatted regulatory text data into a fixed - length vector using a pre - calculated word embedding method and generate a digital vector;

[0055] The key concept extraction module is used to mark important terms in the digital vector using a named entity recognition method to obtain key concept text;

[0056] The evaluation model module is used to establish an evaluation model based on a machine learning algorithm, input the digital vector into the evaluation model, calculate the output evaluation value using the weighted average method, compare the set evaluation threshold with the evaluation value to determine whether the digital vector is compliant, and obtain an evaluation result;

[0057] The anomaly detection module is used to construct an identification model based on an anomaly detection algorithm, input the key concept text into the identification model, and output an identification result using a clustering algorithm.

[0058] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the data security compliance evaluation and identification method as described in the first aspect of the present invention is implemented.

[0059] In a fourth aspect, the present invention provides a computer - readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the data security compliance evaluation and identification method as described in the first aspect of the present invention is implemented.

[0060] The beneficial effects of the present invention are as follows: By collecting regulatory text data through the API interface and using the requests library to send HTTP protocol requests to the API endpoint to obtain regulatory text data, automation and real-time updates are achieved, ensuring the timeliness and accuracy of the regulatory text data, enabling quick response to changes in laws and regulations, thereby enhancing the timeliness and compliance of the system. By adopting a pre-computed word embedding method, especially using the Word2Vec model to perform word segmentation on the formatted regulatory text data, collecting all the independent words after word segmentation to form a non-repeating vocabulary, not only improves the consistency and stability of text representation, but also enhances the recognition model's ability to understand semantic information, making subsequent evaluation and recognition more accurate. By adopting a named entity recognition NER model to scan the formatted regulatory text data, identifying entities and assigning different category labels to them, outputting text information with annotations, and forming key concept text, the process accurately extracts the key terms in the regulatory text, ensuring the pertinence and accuracy of subsequent anomaly detection, enabling the recognition method to focus on important regulatory clauses and concepts, and improving the effectiveness and practicality of the recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0062] Figure 1 It is a flowchart of the data security compliance evaluation and recognition method in Embodiment 1;

[0063] Figure 2 It is a schematic diagram of the data security compliance evaluation and recognition system in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention with reference to the drawings in the specification.

[0065] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar promotions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0066] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.

[0067] Example 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides a method for data security compliance assessment and identification, including the following steps:

[0068] S1. Collect regulatory text data through the API interface, and preprocess the regulatory text data using the natural language processing library NLTK to obtain formatted regulatory text data;

[0069] Furthermore, use the requests library to send an HTTP protocol request to the API endpoint to obtain regulatory text data;

[0070] Use the natural language processing library NLTK to filter out meaningless words in the regulatory text data and reduce noise;

[0071] Use the built-in stopword list in NLTK to check each word. If it belongs to the stopwords, remove it from the regulatory text data;

[0072] If the regulatory text data contains HTML-formatted content, remove the HTML tags to retain only the plain text;

[0073] Restore the words in the regulatory text data and split the text in the regulatory text data into individual words;

[0074] Use the WordNetLemmatizer in NLTK for word lemmatization and split the continuous text into individual words or tokens to obtain formatted regulatory text data;

[0075] It should be noted that in the process of collecting regulatory text data through the API interface and preprocessing the regulatory text data using the natural language processing library NLTK, using the requests library to send an HTTP protocol request to the API endpoint to obtain regulatory text data, this process can not only ensure the reliability and timeliness of the data source, but also effectively remove irrelevant information by denoising and standardizing the regulatory text data, improve the efficiency and accuracy of subsequent processing, ensure the quality and consistency of the regulatory text data. In particular, steps such as stopword filtering, HTML tag removal, and word lemmatization help reduce noise and retain key information, providing a high-quality data basis for subsequent analysis.

[0076] S2. Adopt the pre-computed word embedding method to map each word in the formatted legal text data into a fixed-length vector and generate a digital vector;

[0077] Furthermore, use the word embedding library Word2Vec to perform word segmentation on the formatted legal text data, collect all the independent words after word segmentation, and form a non-repeating vocabulary;

[0078] Create an empty set and fill it with the vocabulary, count the occurrence frequencies of the independent words, sort the vocabulary in descending order of frequency, and form the final vocabulary;

[0079] For the independent words in the final vocabulary, look up the corresponding fixed-length vectors in the word embedding library Word2Vec, and use the found corresponding fixed-length vectors as the representation of the independent words to form word-level digital vectors;

[0080] Create a zero vector with the same dimension as the word-level digital vectors as an accumulator, accumulate all the word-level digital vectors, and for each independent word in the final vocabulary, obtain its corresponding word-level digital vector;

[0081] Ensure there is a final vocabulary containing all independent words, and load the word embedding library Word2Vec;

[0082] Traverse each independent word in the final vocabulary, check if it exists in the word embedding library Word2Vec. If it exists, obtain its vector; if not, replace it with a zero vector, and create a dictionary to store the word-level digital vectors of each word;

[0083] Add each word-level digital vector to the accumulator one by one, and divide each word-level digital vector in the accumulator by the total number of independent words to obtain the average word vector. If the vocabulary is empty, return a zero vector;

[0084] Use the obtained average word vector as the representation of the document-level digital vector to obtain the document-level digital vector, and mark it as the digital vector. The expression is:

[0085] ;

[0086] where, is the number of words in the formatted legal text data in, is the formatted legal text data in the word, is the th word vector of the word, is the formatted legal text data of the digital vector, Is the formatted regulatory text data;

[0087] It should be noted that using the pre - calculated word embedding method to map each word in the formatted regulatory text data into a fixed - length vector can not only capture the semantic relationships between words, but also represent the core features of the entire regulatory text by generating document - level digital vectors, ensuring the numerical operability of text information and the effectiveness of model training. When creating the final vocabulary, sorting according to the occurrence frequency of independent words can preferentially retain high - frequency words to further optimize the quality of word vectors. The calculation of the average word vector enables texts of different lengths to be compared in a unified vector space, thus improving the understanding ability of the recognition model for texts of different scales.

[0088] S3. Use the named entity recognition NER method to mark important terms in the digital vector to obtain the key concept text;

[0089] Furthermore, use the named entity recognition NER model to input the formatted regulatory text data sentence - by - sentence or paragraph - by - paragraph into the named entity recognition NER model;

[0090] The named entity recognition NER model scans the formatted regulatory text data and identifies entities;

[0091] For the identified entities, the named entity recognition NER model assigns different category labels to them and outputs the text information with annotations and entity types ;

[0092] Create a new text string that contains the original text and annotation information, and mark the original text with annotations as important terms in the text;

[0093] Mark the output text information with annotations as important terms in the text, create a list or dictionary to store all entities and their category labels as important terms in the text, and form the key concept text. The expression is:

[0094] ;

[0095] Where, Is the set of words after word segmentation, Is the formatted regulatory text data In the key concept text, Is the Entity type corresponding to the Word,

[0096] It should be noted that by using the named entity recognition (NER) method to mark important terms in digital vectors, key concepts in regulatory texts, such as legal names, institutional names, dates, etc., can be accurately extracted. Scanning the text through the NER model and assigning category labels not only enhances the structural degree of text information but also facilitates subsequent analysis and application. The text information with annotations not only intuitively displays the important terms in the text but also enables users to quickly locate and understand key content, improving the accuracy and efficiency of regulatory text interpretation. The formed key concept text, as the basis for subsequent processing, can better support compliance assessment and other advanced analysis tasks.

[0097] S4. Establish an evaluation model based on machine learning algorithms, input the digital vector into the evaluation model, and calculate the output evaluation value using the weighted average method;

[0098] Furthermore, divide the formatted regulatory text data into a training set and a validation set;

[0099] Use the scikit - learn library, import the LogisticRegression class to create a logistic regression instance, and combine it with the formatted regulatory text data to construct an evaluation model, and set the initial parameter weights of the evaluation model ;

[0100] Input the digital vector in the formatted regulatory text data into the evaluation model, and calculate the evaluation value using the weighted average method , and the expression is:

[0101] ;

[0102] Among them, is the word weight, is the weight of the th word, is the th word vector, is the Sigmoid activation function, is the evaluation value, is the formatted regulatory text data, is the digital vector in the formatted regulatory text data;

[0103] It should be noted that by establishing an evaluation model based on machine learning algorithms, calculating the output evaluation value through the weighted average method, and setting an evaluation threshold to compare with the evaluation value, the automatic determination of the compliance of regulatory text data can be achieved. The logistic regression algorithm, combined with the formatted regulatory text data, constructs an efficient binary classification model that can distinguish compliant and non - compliant texts.

[0104] S5. Set the evaluation threshold and compare it with the evaluation value to determine whether the digital vector complies with the regulations and obtain the evaluation result;

[0105] Furthermore, set the threshold based on the formatted regulatory text data that has never participated in the training in the validation set , including selecting the formatted regulatory text data that has never participated in the training as the validation set, inputting the digital vectors of the validation set into the trained evaluation model to obtain the evaluation value of each sample, performing k-fold cross-validation on the validation set, and determining the threshold ;

[0106] When ≥ , it indicates that the digital vector in the formatted regulatory text data is lower than the set threshold , then it is determined that the regulatory text complies with the regulations;

[0107] When < , it is determined that the digital vector in the formatted regulatory text data is non-compliant, then the formatted regulatory text data is modified. After modification, the digital vector in the formatted regulatory text data is recalculated according to the digital vector expression, and then re-input into the evaluation model to calculate the evaluation value. Set the number of iterations to , until ≥ ; ;

[0108] It should be noted that dynamically adjusting the evaluation threshold ensures the adaptability and robustness of the evaluation model. Especially when facing new regulatory texts or changing compliance standards, it can respond flexibly. The process of recalculating the evaluation value ensures the accuracy and reliability of the evaluation result, making the compliance evaluation of regulatory texts more scientific and reasonable.

[0109] S6. Build an identification model based on the anomaly detection algorithm, input the key concept text into the identification model, and output the identification result using the clustering algorithm;

[0110] Furthermore, build an identification model based on the isolation forest anomaly detection algorithm and the formatted regulatory text data, including importing the scikit-learn library for machine learning in Python, ensuring that the key concept text is converted into a numerical feature vector through TF-IDF, performing standardization processing on the numerical feature vector, creating an isolation forest model instance and setting parameters, and training the isolation forest model using the formatted regulatory text data to obtain the identification model;

[0111] The isolation forest algorithm is used to isolate observation points by constructing multiple decision trees, and the IsolationForest class in the Python library is used to initialize the recognition model parameters;

[0112] Convert the key concept text in the formatted regulation text data into TF-IDF values, calculate the TF-IDF vector values of the key concept text in the formatted regulation text data, and standardize the key concept text in the formatted regulation text data represented by the TF-IDF vector values. The expression is:

[0113] ;

[0114] Among them, is a word, is the formatted regulation text data in the key concept text, is all the key concept texts, is the total number of documents in all the key concept texts, is the document containing the word number, is the word in the formatted regulation text data in the frequency of the key concept text, is the TF-IDF vector value;

[0115] Input the TF-IDF vector value into the recognition model, and use the K-means clustering algorithm to classify the output of the recognition model. The expression is:

[0116] ;

[0117] Among them, is the clustering label, is the anomaly score matrix, represents the anomaly score matrix in the th sample anomaly score vector, is the th cluster center point, is the set number of clusters, is the formatted regulation text data in the key concept text, represents the key text concept anomaly score, is the number of trees in the isolation forest algorithm, is the anomaly score matrix in the element, is the a cluster is the path length in the isolation forest algorithm for the path length is the maximum path length of all trees in the isolation forest is the expected value of, that is, the average value of the path length;

[0118] Set the anomaly score threshold and compare it with the anomaly score matrix for comparison;

[0119] When ≥ it means that the key text concepts in the formatted regulatory text data are non-compliant; non-compliant;

[0120] When < it means that the key text concepts in the formatted regulatory text data are compliant; compliant;

[0121] It should be noted that by constructing an identification model based on the anomaly detection algorithm and analyzing the key concept text through the isolation forest and K-means clustering algorithms, the abnormal situations in the regulatory text can be effectively identified. The isolation forest algorithm can efficiently detect potential abnormal points by constructing multiple decision trees to isolate the observation points, and the K-means clustering algorithm classifies these abnormal points to help determine which key concept texts do not meet the regulations, providing a clear judgment standard for setting the anomaly score threshold, ensuring the accuracy and reliability of the identification results. It can not only discover the non-compliant parts in the regulatory text but also provide valuable feedback for text improvement, enhancing the overall quality of the regulatory text.

[0122] This embodiment also provides a data security compliance evaluation and identification system, including:

[0123] a data processing module, a word embedding generation module, a key concept extraction module, an evaluation model module, and an anomaly detection module;

[0124] The data processing module is used to collect regulatory text data through the API interface and preprocess the regulatory text data using the natural language processing library NLTK to obtain formatted regulatory text data;

[0125] The word embedding generation module is used to map each word in the formatted regulatory text data to a fixed-length vector using a pre-computed word embedding method and generate a digital vector;

[0126] A key concept extraction module, which is used to mark important terms in the digital vector by using the named entity recognition method to obtain the key concept text;

[0127] An evaluation model module, which is used to establish an evaluation model based on the machine learning algorithm, input the digital vector into the evaluation model, calculate the output evaluation value by using the weighted average method, set an evaluation threshold to compare with the evaluation value, judge whether the digital vector is compliant, and obtain the evaluation result;

[0128] An anomaly detection module, which is used to construct an identification model based on the anomaly detection algorithm, input the key concept text into the identification model, and output the identification result by using the clustering algorithm.

[0129] This embodiment also provides a computer device, which is applicable to the case of the data security compliance evaluation and identification method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the data security compliance evaluation and identification method as proposed in the above embodiment.

[0130] This computer device can be a terminal. This computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of this computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through Wi-Fi, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of this computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of this computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0131] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for realizing data security compliance evaluation and identification proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, a magnetic disk or an optical disc.

[0132] In summary, the present invention collects regulatory text data through an API interface, and uses the requests library to send an HTTP protocol request to the API endpoint to obtain the regulatory text data, achieving automation and real-time updates, ensuring the timeliness and accuracy of the regulatory text data, being able to quickly respond to changes in laws and regulations, thereby improving the timeliness and compliance of the system. By adopting a pre-computed word embedding method, especially using the Word2Vec model to perform word segmentation on the formatted regulatory text data, collecting all the independent words after word segmentation to form a non-repeating vocabulary, it not only improves the consistency and stability of text representation, but also enhances the recognition model's ability to understand semantic information, making subsequent evaluation and identification more accurate. By adopting a named entity recognition NER model to scan the formatted regulatory text data, identifying entities and assigning different category labels to them, outputting the text information with annotations, and forming key concept text, the process accurately extracts the key terms in the regulatory text, ensuring the pertinence and accuracy of subsequent anomaly detection, enabling the recognition method to focus on important regulatory clauses and concepts, and improving the effectiveness and practicality of the recognition results.

[0133] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A data security compliance assessment and identification method, characterized by: include: Collect regulatory text data through the API interface, and use the natural language processing library NLTK to pre-process the regulatory text data to obtain formatted regulatory text data; Using a pre-computed word embedding method, each word in the formatted regulatory text data is mapped into a fixed-length vector and a digital vector is generated; The named entity recognition (NER) method is used to mark the important terms in the digital vector and obtain the key concept text; An evaluation model is established based on a machine learning algorithm, the digital vector is input into the evaluation model, and the output evaluation value is calculated using the weighted average method; Set the evaluation threshold and compare it with the evaluation value to determine whether the digital vector is compliant and obtain the evaluation result; Build a recognition model based on the anomaly detection algorithm, input the key concept text into the recognition model, and use the clustering algorithm to output the recognition results. The specific steps are: Building a recognition model based on the Isolation Forest Anomaly Detection Algorithm and formatted regulatory text data; The isolation forest algorithm is used to isolate observation points by building multiple decision trees, and the IsolationForest class in the Python library is used to initialize the recognition model parameters; The key concept text in the formatted legal text data is converted into a TF-IDF value, the TF-IDF vector value of the key concept text in the formatted legal text data is calculated, and the key concept text in the formatted legal text data represented by the TF-IDF vector value is standardized. The expression is: ; in, It's words. Formatted regulatory text data Key concepts in the text, is all the key concept text, is the total number of documents in all key concept texts, Contains words The number of documents, It's a word In formatting regulatory text data The frequency of key concepts in the text, for TF-IDF vector value of ; The TF-IDF vector value is input into the recognition model, and the K-means clustering algorithm is used to classify the output of the recognition model. The expression is: ; in, is the cluster label, is the anomaly score matrix, Represents the anomaly score matrix Middle The anomaly score vector of samples, It is The center point of the cluster, is the set number of clusters, Formatted regulatory text data Key concepts in the text, Represents key text concepts The anomaly score, is the number of trees in the isolation forest algorithm, is the anomaly score matrix The elements in It is Clusters, is the isolation forest algorithm The path length, is the maximum path length of all trees in the isolation forest, yes The expected value of , that is, the average value of the path length; Setting anomaly score threshold With the anomaly score matrix Make a comparison; when ≥ , indicates formatted regulatory text data Key text concepts in Non-compliance; when < , indicates formatted regulatory text data Key text concepts in Compliance.

2. The data security compliance assessment and identification method according to claim 1, characterized in that: The regulatory text data is collected through the API interface, and the regulatory text data is preprocessed using the natural language processing library NLTK to obtain formatted regulatory text data. The specific steps are: Use the requests library to send HTTP protocol requests to the API endpoint to obtain regulatory text data; Use the natural language processing library NLTK to filter meaningless words in regulatory text data and reduce noise The words in the regulatory text data are restored, and the text in the regulatory text data is segmented into separate words to obtain formatted regulatory text data.

3. The data security compliance assessment and identification method according to claim 2, characterized in that: The pre-calculated word embedding method is used to map each word in the formatted regulatory text data into a fixed-length vector and generate a digital vector. The specific steps are as follows: The word embedding library Word2Vec is used to perform word segmentation on the formatted regulatory text data, and all independent words after word segmentation are collected to form a non-repetitive vocabulary; Create an empty collection and fill it with a vocabulary, count the frequency of occurrence of independent words, sort the vocabulary according to the frequency of occurrence of independent words, and form a final vocabulary; For independent words in the final vocabulary, the corresponding fixed-length vector is searched in the word embedding library Word2Vec, and the corresponding fixed-length vector found is used as the representation of the independent word to form a word-level digital vector; Create a zero vector of the same dimension as the word-level digital vector as an accumulator, accumulate all word-level digital vectors, and obtain the corresponding word-level digital vector for each independent word in the final vocabulary; Add each word-level digital vector to the accumulator one by one to get the average word vector; The obtained average word vector is used as the representation of the document-level digital vector to obtain the document-level digital vector and mark it as a digital vector.

4. The data security compliance assessment and identification method according to claim 3, characterized in that: The named entity recognition (NER) method is used to mark important terms in the digital vector to obtain key concept texts. The specific steps are as follows: Using a named entity recognition (NER) model, input the formatted regulatory text data sentence by sentence or paragraph by paragraph into the named entity recognition (NER) model; The named entity recognition (NER) model scans the formatted regulatory text data and identifies entities; For the identified entities, the named entity recognition (NER) model assigns different category labels to them and outputs the annotated text information; The output annotated text information is marked as important terms in the text and forms a key concept text.

5. The data security compliance assessment and identification method according to claim 4, characterized in that: The evaluation model is established based on the machine learning algorithm, the digital vector is input into the evaluation model, and the weighted average method is used to calculate the output evaluation value. The specific steps are as follows: The regulatory text data to be formatted Divide into training set and validation set; Use the scikit-learn library to import the LogisticRegression class to create a logistic regression instance and combine it with the formatted regulatory text data Build an evaluation model and set the initial parameter weights of the evaluation model ; A numeric vector containing the regulatory text data to be formatted Input into the evaluation model and calculate the evaluation value using the weighted average method , the expression is: ; in, is the word weight, It is The weight of the word, It is The word vector of each word, is the Sigmoid activation function, is the evaluation value, is the formatted regulatory text data, A numeric vector containing the formatted regulatory text data.

6. The data security compliance assessment and identification method according to claim 5, characterized in that: The evaluation threshold is set to be compared with the evaluation value to determine whether the digital vector is compliant and obtain the evaluation result. The specific steps are as follows: Thresholds are set based on formatted regulatory text data in the validation set that was never used in training ; when ≥ , representing formatted regulatory text data A numeric vector in Below the set threshold , then the regulatory text is judged to be compliant; when < , determine the formatted regulatory text data A numeric vector in If it is not compliant, the formatted regulatory text data Make the modification, and after the modification, recalculate the formatted regulatory text data according to the digital vector expression A numeric vector in , re-enter it into the evaluation model to calculate the evaluation value, and set the number of iterations to , until ≥ .

7. A data security compliance assessment and identification system, based on the data security compliance assessment and identification method according to any one of claims 1 to 6, characterized in that: include: Data processing module, word embedding generation module, key concept extraction module, evaluation model module and anomaly detection module; The data processing module is used to collect regulatory text data through an API interface, and pre-process the regulatory text data using a natural language processing library NLTK to obtain formatted regulatory text data; The word embedding generation module is used to map each word in the formatted regulatory text data into a fixed-length vector using a pre-calculated word embedding method and generate a digital vector; The key concept extraction module is used to mark important terms in the digital vector using a named entity recognition method to obtain a key concept text; The evaluation model module is used to establish an evaluation model based on a machine learning algorithm, input a digital vector into the evaluation model, calculate an output evaluation value using a weighted average method, set an evaluation threshold and compare it with the evaluation value, determine whether the digital vector is compliant, and obtain an evaluation result; The anomaly detection module is used to build a recognition model based on the anomaly detection algorithm, input the key concept text into the recognition model, and output the recognition result using the clustering algorithm.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the data security compliance assessment and identification method described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the data security compliance assessment and identification method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Unsupervised text multi-label marking method based on entity word influence area evaluation standard

    CN116644182A

  • Data processing method and device, electronic equipment, storage medium and program product

    CN118038214A