Data anti-leakage processing system and method based on natural language and storage medium

By segmenting data files and processing large language models, sensitive content is automatically identified and marked, and the problems of high technical requirements and workload of confidential personnel in the existing technology are solved, and the automated identification and marking of sensitive content of data files are realized.

CN120296776APending Publication Date: 2025-07-11HEFEI SAINI TENGLONG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150765.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, data leakage prevention products require confidential personnel to be familiar with the principles of identification policy identification technology, and each time a new data file is identified, the policy content needs to be adjusted, resulting in high technical requirements and high workload.

Method used

By dividing the data file into multiple texts, using a large language model to extract abstracts and eliminate unimportant abstracts, identify and mark sensitive content based on the integrated content, and realize automated data file sensitive content recognition.

Benefits of technology

It reduces the technical requirements and workload of confidential personnel and improves the automation of data file sensitive content recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296776A_ABST
    Figure CN120296776A_ABST
Patent Text Reader

Abstract

The invention discloses a data anti-leakage processing system based on a natural language. The data anti-leakage processing system comprises an information definition module, a data matching module and an analysis module, relates to the technical field of data desensitization, and solves the technical problems that the technical requirements on confidentiality personnel are relatively high and the processing workload is relatively large due to the fact that strategy contents need to be added or modified when a new data file is identified every time. The method comprises the following steps: dividing a data file into a plurality of multi-copy texts, extracting abstracts from the divided contents through a large language model, removing unimportant abstracts in a plurality of abstracts, summarizing reserved key abstracts by using the large language model, and determining a screening subject; according to the method, the screened integrated content is determined for the screening subject according to the large language model, and then the wool fabric in the data file is matched and marked according to the integrated content, so that the sensitive content in the data file is automatically identified, a worker does not need to understand the data file in detail, and the workload of secrecy personnel is reduced to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data desensitization, and specifically relates to a data leakage prevention processing system based on natural language. Background Art

[0002] In current data leakage prevention products with content recognition as the core, when users need to know which files contain sensitive and non-disclosable information or data, they need to tell the detection system some inspection policies in advance. The system needs to know which data to be inspected contains sensitive content. For this purpose, the traditional method is to tell the device: keywords, regular expressions, and some more complex indexes, etc. These methods require users to be familiar with their own specialties and master the content recognition principles implemented in computer logic languages, and there are some thresholds in production, which greatly limits the use of data leakage prevention products.

[0003] The current existing method is that the security personnel write the policy content after viewing the content to be screened, and then desensitize the file content according to the policy content. And for each different content, the policy content to be implemented will change, and it needs to be adjusted before use. This method requires the security personnel to not only understand the principle of the recognition strategy recognition technology, but also add or modify the policy content every time they identify a new data file, resulting in high technical requirements for the security personnel and a large amount of workload to be processed. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; for this reason, the present invention proposes a data leakage prevention processing system based on natural language, which is used to solve the technical problems that every time a new data file is identified, the policy content needs to be added or modified, resulting in high technical requirements for security personnel and a large amount of workload to be processed. The present invention divides the data file into several multi-connected texts, extracts summaries of the divided content through a large language model, eliminates unimportant summaries among multiple summaries, and then summarizes the remaining key summaries through a large language model to determine the screening theme; determines the integrated content to be screened according to the screening theme by the large language model, and then matches and marks the content in the data file according to the integrated content, so as to automatically identify sensitive content in the data file, thereby solving the above problems.

[0005] To achieve the above object, the first aspect of the present invention provides a data leakage prevention processing system based on natural language, including: an information definition module, a data matching module, an analysis module, and a database;

[0006] The information definition module: sets the content of sensitive types; wherein, the sensitive types include personal information, work secrets, and sensitive data;

[0007] The content of the sensitive type can be selected and set according to the needs of each industry;

[0008] Data matching module: Preprocess the data file to obtain the screening topic, determine the desensitized content in the screening text according to the large language model and the screening topic, compare the content of the sensitive type with the desensitized content and then integrate them to obtain the integrated content;

[0009] Analysis module: Analyze and process the data file according to the integrated content to obtain the marked content; Desensitize the data file according to the marked content;

[0010] Database: Used to store analysis prompt words, abstract prompt words, summary prompt words and amplification prompt words.

[0011] Preferably, the preprocessing of the data file to obtain the screening topic includes:

[0012] Convert the data file into text data, segment the text data through a word segmentation tool, split the sentences based on punctuation marks, and remove the punctuation marks in the text data to obtain a number of multi-connected texts;

[0013] Input a number of multi-connected texts / abstract prompt words into the large language model to obtain a number of topic abstracts ZYi; Determine the important ranking of the number of topic abstracts through TextRank, and mark the number of topic abstracts accounting for the top n in the important ranking as key abstracts; where i = 1, 2,... a, a is the total number of topic abstracts, n ∈ (0, 100%); where, / is the paragraph symbol;

[0014] Input a number of key abstracts / summary prompt words into the large language model to obtain the screening topic.

[0015] Through the above method, the content in the data file that is relatively deviated from the topic can be removed, avoiding interference to the large language model during the subsequent content processing.

[0016] Preferably, the obtaining method of n includes:

[0017] Use a pre-trained sentence embedding model to convert a number of topic abstracts ZYi into vector representations of a specified length; Calculate the average value of the cosine similarity between the topic abstract ZYj and the remaining topic abstracts respectively, and mark it as the similarity average value YZj; where j ∈ i;

[0018] Arrange the similarity average values ZYj of a number of topic abstracts in descending order and perform linear fitting to obtain a fitting curve. After taking the derivative of the fitting curve, obtain the curve f’(ZYi), and obtain the derivative value Kj of a number of topic abstracts ZYj in the curve f’(ZYi);

[0019] Extract the maximum value Kjmax among the derivative values Kj corresponding to several topic abstracts ZYj, and obtain the proportion position of Kjmax in the curve f’(ZYi);

[0020] Compare the proportion position with a preset threshold; when the proportion position is less than the preset threshold, assign the proportion position to n; otherwise, assign the preset threshold to n.

[0021] The content of the topic abstracts varies greatly starting from the maximum value Kjmax among the derivative values Kj corresponding to the above-mentioned topic abstracts ZYj. Therefore, using this point as a demarcation point to classify the topic abstracts into two categories for screening can better fit the specific content of the data file, thus ensuring the accuracy during the subsequent processing by the large language model.

[0022] Preferably, the method for determining the desensitized content in the screened text according to the large language model and the screened topics includes:

[0023] Extract amplification prompt words from the prompt word database and input them into the large language model to generate several amplification prompt phrases TSx; combine and input the several amplification prompt phrases TSx into the large language model respectively / data file to obtain the desensitized content.

[0024] Preferably, the method for integrating the sensitive type content and the desensitized content after comparison includes:

[0025] Extract the sensitive type content and the desensitized content; mark the subset where the sensitive type content is the same as the desensitized content as the integrated content.

[0026] Preferably, the method for integrating the sensitive type content and the desensitized content after comparison includes:

[0027] Extract the sensitive type content and the desensitized content; mark the collection of the sensitive type content and the desensitized content as the integrated content.

[0028] Determining the integrated content through the above two methods respectively can determine the sensitivity of desensitization according to requirements, so as to achieve desensitization with two different sensitivity levels.

[0029] Preferably, the method for analyzing and processing the data file according to the integrated content includes:

[0030] Extract the sub-data of the integrated content and the analysis prompt words, input the analysis prompt words / sub-data of the integrated data into the large language model to obtain the output content; perform feature marking on the output content in the data file to obtain the marked content.

[0031] Preferably, the method for desensitizing the data file according to the marked content includes:

[0032] Identify the feature marks in the data file, and uniformly replace the content covered by the feature marks with feature symbols.

[0033] In a second aspect, the present invention also discloses a data anti-leakage processing system based on natural language, comprising the following steps:

[0034] Step 1: Set the content of sensitive types;

[0035] Step 2: Preprocess the data file to obtain a screening theme, and determine the desensitized content in the screened text according to the large language model and the screening theme;

[0036] Step 3: Compare the content of sensitive types with the desensitized content and then integrate them to obtain the integrated content;

[0037] Step 4: Analyze and process the data file according to the integrated content to obtain the marked content; desensitize the data file according to the marked content.

[0038] In a third aspect, the present invention also discloses a storage medium storing a computer program, which, when executed, runs the content of the first aspect.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows: The data file is divided into several multi-connected texts, and the large language model is used to extract the abstracts of the divided content, and the unimportant abstracts among the multiple abstracts are removed, and then the large language model is used to summarize the remaining key abstracts to determine the screening theme; according to the large language model, the integrated content to be screened is determined, and then the content in the data file is matched and marked according to the integrated content, so as to automatically identify the sensitive content in the data file, without the need for staff to have a detailed understanding of the data file, which reduces the workload of the confidentiality personnel to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic flow chart of the data anti-leakage processing based on natural language in the present invention;

[0042] Figure 2 It is a schematic flow chart of the n acquisition in the present invention;

[0043] Figure 3 It is a schematic structural diagram of the data anti-leakage processing system based on natural language in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.

[0045] Please refer to Figure 1 and Figure 3 , an embodiment of the first aspect of the present invention provides a data leakage prevention processing system based on natural language, including: an information definition module, a data matching module, an analysis module, and a database;

[0046] Information definition module: Set the content of sensitive types; among them, the sensitive types include personal information, work secrets, and sensitive data;

[0047] The content of sensitive types can be selected and set according to the needs of each industry;

[0048] Data matching module: Preprocess the data file to obtain a screening theme, determine the desensitized content in the screening text according to the large language model and the screening theme, compare the content of the sensitive type with the desensitized content, and integrate them to obtain integrated content;

[0049] Analysis module: Analyze and process the data file according to the integrated content to obtain marked content; desensitize the data file according to the marked content;

[0050] Database: Used to store analysis prompt words, summary prompt words, summary prompt words, and amplification prompt words.

[0051] In order to enable the large language model to better identify the content that needs to be desensitized in the data file, the preprocessing of the data file to obtain the screening theme in this embodiment includes the following steps:

[0052] Convert the data file into text data, segment the text data through a word segmentation tool, split the sentences based on punctuation marks, and remove the punctuation marks in the text data to obtain a number of multi-connected texts;

[0053] Input a number of multi-connected texts / summary prompt words into the large language model to obtain a number of topic summaries ZYi; determine the important ranking of the number of topic summaries through TextRank, and mark the number of topic summaries that account for the top n in the important ranking as key summaries; where i = 1, 2,... a, a is the total number of topic summaries, n ∈ (0, 100%); where / is a paragraph symbol;

[0054] Input a number of key summaries / summary prompt words into the large language model to obtain the screening theme.

[0055] It should be noted that both the abstract prompt and the summary prompt are pre-set contents. For example: Key Abstract: I need to desensitize the data file. Now you are required to summarize the above content into one sentence; Summary Prompt: I need to desensitize the data file. Now you are required to summarize the above key abstract into one sentence for screening the theme.

[0056] Please refer to Figure 2 as shown, the method for obtaining n in the above content includes:

[0057] Use a pre-trained sentence embedding model to convert several topic abstracts ZYi into vector representations of a specified length; calculate the average value of the cosine similarity between the topic abstract ZYj and the remaining topic abstracts respectively, and mark it as the similarity average value YZj; where, j ∈ i;

[0058] Arrange the similarity average values ZYj of several topic abstracts from large to small, and perform linear fitting to obtain a fitting curve. After taking the derivative of the fitting curve, obtain the curve f’(ZYi), and obtain the derivative value Kj of several topic abstracts ZYj in the curve f’(ZYi);

[0059] Extract the maximum value Kjmax of the derivative values Kj corresponding to several topic abstracts ZYj, and obtain the proportion position of Kjmax in the curve f’(ZYi);

[0060] Compare the proportion position with a preset threshold; when the proportion position is less than the preset threshold, assign the proportion position to n; otherwise, assign the preset threshold to n; the preset threshold can be set according to the total number of topic abstracts.

[0061] The content of the topic abstract starts to be quite different from this point for the maximum value Kjmax of the derivative value Kj corresponding to the above topic abstract ZYj. Therefore, using this point as a demarcation point to classify the topic abstracts into two categories for screening can be more in line with the specific content of the data file, so as to ensure the accuracy during the subsequent processing by the large language model.

[0062] Exemplarily, there is the following fictional article:

[0063] In the brand-new 2024, Mr. Li's work and life picture scroll unfolds both full of vitality and challenges. At the age of 37, he weaves his own wonderful chapter in the warm little nest at Room 502, Unit 3, Building 16 in a certain community in Binjiang District, Hangzhou City, Zhejiang Province. As a senior software engineer, he drives his car with the license plate number Zhe A88888 every day, shuttling between home and the company located on Wen San Road, weaving dreams with code and driving the future with technology.

[0064] Mr. Li's life circle is rich and diverse, and his family is his strongest backing. His wife, Ms. Wang, a gentle primary school teacher, joins hands with him to build a loving haven. The couple has a son and a daughter, and the happiness of the family is palpable. The eldest son, Xiaoming, at the age of 9, is in the prime of childhood innocence. He studies at the First Experimental Primary School in Hangzhou. He swims in the ocean of knowledge, and every step of his growth embodies the expectations and pride of his parents. The youngest daughter, Xiaohong, always wears an innocent smile on her tender 3-year-old face. Her kindergarten life is full of exploration and discovery.

[0065] In addition to his dual roles in the workplace and at home, Mr. Li is also an amateur photography enthusiast. On weekends, he often picks up his camera and joins the ranks of photography lovers, capturing the beautiful moments in life with his lens and sharing the joy and insights of creation with like-minded friends.

[0066] In terms of financial management, Mr. Li shows a rigorous and meticulous side. Every month, he carefully reviews the bank statements to ensure that every expense is clear. This is an important line of defense for safeguarding property security.

[0067] Health is a wealth that Mr. Li particularly cherishes. He adheres to the habit of working out three times a week. When night falls, from 7 pm to 9 pm, his sweating figure can always be seen in the "Healthy New Life" gym near his home. In addition, he has developed the good habit of having a comprehensive physical examination every year. The most recent physical examination was successfully completed on January 15, 2024, at the First Affiliated Hospital of Zhejiang University School of Medicine. The professionalism and attentiveness of the attending physician, Zhang Hua, made him feel at ease and became another guarantee for his health protection.

[0068] In today's era when digital payment is becoming increasingly popular, Mr. Li also keeps up with the trend and uses digital currency, which not only facilitates daily shopping but also becomes a small window for him to connect to the digital world. The binding of the bank card makes payment easy and fast, and every payment operation is his recognition and enjoyment of the convenience of modern life.

[0069] Thus, Mr. Li's life is not only a deep embrace of technology but also a deep love for his family and a persistent pursuit of personal interests and health. Every day in 2024, he writes his own wonderful story with full enthusiasm.

[0070] The following are several topic summaries generated based on the above content: Several topic summaries:

[0071] 1. Mr. Li's life picture and professional background demonstrate his profound accumulation and professional pursuit in the technical field.

[0072] 2. Mr. Li's family life highlights his wife, children, and the happy times they spent together.

[0073] 3. Mr. Li's hobbies include photography and sharing the joy of creation with like-minded friends.

[0074] 4. Mr. Li's financial management and health habits reflect his emphasis on the safety of family property and personal health.

[0075] 5. Mr. Li's digital payment habits reflect his keeping up with the trend of modern technology and the transformation of lifestyle.

[0076] And after calculating the average similarity corresponding to the above several topic summaries: YZ1 = 0.83 (Career and living conditions), YZ2 = 0.80 (Family life and parent-child relationship), YZ3 = 0.80 (Health management and financial planning), YZ4 = 0.82 (Hobbies and photography), YZ5 = 0.76 (Digital payment habits);

[0077] After calculation, YZ5 is the maximum value Kjmax in life and the content is less, and the preset threshold is 1. Therefore, taking the content corresponding to YZ5 as the demarcation point for screening, the screened topic is summarized as: The situation of Mr. Li's daily life and simple work.

[0078] Specifically, in this embodiment, determining the desensitized content in the screened text according to the large language model and the screened topic includes:

[0079] Extract amplification prompt words from the prompt word database and input them into the large language model to generate several amplification prompt phrases TSx; Combine and input the several amplification prompt phrases TSx into the large language model respectively / data files to obtain the desensitized content.

[0080] Exemplarily, the amplification prompt words: Please polish the following sentences into 2 different sentence descriptions; And the above desensitized content is: address information, license plate number, physical examination information;

[0081] As in the embodiment of the present invention, integrating the sensitive type of content after comparing it with the desensitized content includes:

[0082] Extract the sensitive type of content and the desensitized content; Mark the subset where the sensitive type of content is the same as the desensitized content as the integrated content.

[0083] Through the above method, the calculation amount can be reduced to quickly complete the desensitization of data to adapt to occasions with lower sensitivity. Obviously, the above content can be desensitized in this way.

[0084] As another preferred embodiment to adapt to occasions with lower sensitivity, integrating the sensitive type of content after comparing it with the desensitized content in this embodiment includes:

[0085] Extract the content of sensitive types and the desensitized content; mark the collection of the content of sensitive types and the desensitized content as the integrated content.

[0086] Determine respectively through the above two methods that the integrated content can determine the sensitivity of desensitization according to requirements, so as to achieve desensitization of two different sensitive degrees.

[0087] In this embodiment, analyzing and processing the data file according to the integrated content includes:

[0088] Extract the sub-data of the integrated content and the analysis prompt words, input the analysis prompt words / sub-data of the integrated data into the large language model to obtain the output content; mark the features of the output content in the data file to obtain the marked content.

[0089] In this embodiment, desensitizing the data file according to the marked content includes:

[0090] Identify the feature marks in the data file, and uniformly replace the content covered by the feature marks with feature symbols.

[0091] Please refer to Figure 1 As shown, in the second aspect, the present invention also discloses a data leakage prevention processing system based on natural language, including the following steps:

[0092] Step 1: Set the content of sensitive types;

[0093] Step 2: Preprocess the data file to obtain the screening theme, and determine the desensitized content in the screening text according to the large language model and the screening theme;

[0094] Step 3: Integrate the content of sensitive types and the desensitized content after comparison to obtain the integrated content;

[0095] Step 4: Analyze and process the data file according to the integrated content to obtain the marked content; desensitize the data file according to the marked content.

[0096] In the third aspect, the present invention also discloses a storage medium, which stores a computer program, and when the computer program is executed, the content of the first aspect runs.

[0097] Some of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is obtained by software simulation of a large amount of collected data to get a formula closest to the actual situation; the preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0098] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A data leakage prevention processing system based on natural language, characterized in that include: Information definition module, data matching module and analysis module; Information definition module: set sensitive content; Data matching module: pre-processes the data file to obtain the screening topic, determines the desensitized content in the screening text based on the large language model and the screening topic, compares the sensitive content with the desensitized content, and integrates them to obtain the integrated content; Analysis module: Analyze and process data files according to the integrated content to obtain marked content; Desensitize data files based on the tag content.

2. The data leakage prevention processing system based on natural language according to claim 1, wherein The step of preprocessing the data file to obtain the screening subject comprises: Convert the data file into text data, segment the text data using a word segmentation tool (SnowNLP), segment sentences based on punctuation marks, and remove punctuation marks from the text data to obtain several multi-linked texts; Input several multi-linked text / abstract prompt words into the large language model to obtain several topic summaries ZYi; determine the importance ranking of several topic summaries through TextRank, and take several topic summaries with the top n proportions in the importance ranking as key summaries; where i = 1, 2, ... a, a is the total number of topic summaries, n∈(0, 100%); where / is a segmentation symbol; Several key abstract / summary prompt words are input into the large language model to obtain the screening topics.

3. The data leakage prevention processing system based on natural language according to claim 2, characterized in that, The method for obtaining n includes: Use the pre-trained sentence embedding model to convert several topic summaries ZYi into vector representations of specified length; calculate the average cosine similarity between topic summary ZYj and the remaining topic summaries, and mark them as similarity mean YZj; where j∈i; Arrange the similar mean values ​​ZYj of several subject abstracts from large to small, and perform linear fitting to obtain a fitting curve. After derivation of the fitting curve, obtain the curve f'(ZYi), and obtain the derivative values ​​Kj of several subject abstracts ZYj in the curve f'(ZYi); Extract the maximum value Kjmax among the derivative values ​​Kj corresponding to several topic summaries ZYj, and obtain the proportion position of Kjmax in the curve f'(ZYi); The proportion position is compared with a preset threshold; when the proportion position is less than the preset threshold, the proportion position is assigned to n; otherwise, the preset threshold is assigned to n.

4. The data leakage prevention processing system based on natural language according to claim 1, wherein Determining the desensitized content in the screening text according to the large language model and the screening topic includes: Extract the amplified prompt words and input them into the large language model to generate several amplified prompt words TSx; combine the several amplified prompt words TSx respectively / data files and input them into the large language model to obtain desensitized content.

5. The data leakage prevention processing system based on natural language according to claim 4, characterized in that, The step of comparing the sensitive content with the desensitized content and integrating them includes: Extract sensitive content and desensitized content; and mark the same subset of sensitive content and desensitized content as integrated content.

6. The data leakage prevention processing system based on natural language according to claim 4, wherein The step of comparing the sensitive content with the desensitized content and integrating them includes: Extract sensitive content and desensitized content; and mark the collection of sensitive content and desensitized content as integrated content.

7. The data leakage prevention processing system based on natural language according to claim 1, wherein The analyzing and processing of the data files according to the integrated content includes: Extract the sub-data of the integrated content and the analysis prompt words, input the analysis prompt words / sub-data of the integrated data into the large language model to obtain the output content; perform feature marking on the output content in the data file to obtain the marked content.

8. The data leakage prevention processing system based on natural language according to claim 1, characterized in that The desensitization of the data file according to the marked content includes: Identify the feature marks in the data file, and uniformly replace the content covered by the feature marks with feature symbols.

9. A data leakage prevention processing method based on natural language, applied to the data leakage prevention processing method based on natural language described in any one of claims 1-8, characterized in that, The following steps are included: Step 1: Set the content of the sensitive type; Step 2: Preprocess the data file to obtain the screened theme, and determine the desensitized content in the screened text according to the large language model and the screened theme; Step 3: Compare the content of the sensitive type with the desensitized content and then integrate them to obtain the integrated content; Step 4: Analyze and process the data file according to the integrated content to obtain the marked content; Desensitize the data file according to the marked content.

10. A storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed, it runs the system according to any one of claims 1-8.