Data leakage prevention method and apparatus
By acquiring and evaluating sensitive information in document data, generating key data, and implementing dynamic protection measures, the problem of leakage caused by key loss in data encryption technology is solved, and the accurate identification and effective protection of sensitive information is achieved.
Patent Information
- Application Number
- CN202510476032.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing data encryption technologies may lead to data loss and leakage if the key is lost or the encrypted data is corrupted, as there is a lack of effective leakage prevention mechanisms.
By acquiring sensitive information from initial document data, generating key data based on fingerprint information and access permissions, monitoring access to and operation of sensitive information in real time, and using a comprehensive evaluation method to identify key data and implement dynamic protection measures.
It enables accurate identification and protection of sensitive information, avoids misjudgment, ensures that only authorized users can access it, promptly detects and prevents potential leakage risks, and improves data security and privacy.
Smart Images

Figure CN119989426B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to data processing technology, and more particularly to a method and apparatus for preventing data leakage. Background Technology
[0002] With the rapid development of information technology, data has become a core asset for enterprises and organizations. However, frequent data breaches have caused significant losses to businesses and organizations, including customer loss, reputational damage, loss of core technologies, legal issues, and financial compensation. Therefore, data loss prevention technology has become a crucial means of ensuring data security.
[0003] Currently, data loss prevention typically involves encrypting data using encryption algorithms to ensure data security during storage and transmission. Common encryption technologies include disk encryption, file encryption, and transparent document encryption / decryption.
[0004] However, if the key is lost or the encrypted data is corrupted in the above methods, the data may be unrecoverable, leading to data leakage. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed to provide a data leakage prevention method and apparatus that overcomes or at least partially solves the above problems.
[0006] According to one aspect of the present invention, a data leakage prevention method is provided, comprising the following steps:
[0007] Obtain initial document data;
[0008] Sensitive information in the initial document data is obtained based on fingerprint information. The initial document data is configured with different levels, and the sensitive information in the initial document data is data of the sensitive level.
[0009] Based on the sensitive information and access permissions of the initial document data, generate key data corresponding to the initial document data;
[0010] Target document data is generated based on the initial document data, the key data and sensitive information operation data corresponding to the initial document data.
[0011] Optionally, based on sensitive information and access permissions of the initial document data, key data corresponding to the initial document data is generated, including:
[0012] Extract keywords from sensitive information in the initial document data;
[0013] Calculate the contribution of keywords in sensitive information within the initial document data;
[0014] The frequency of accessing initial document data;
[0015] Based on the keywords in the sensitive information of the initial document data, the contribution of the keywords in the sensitive information of the initial document data, the access permissions of the initial document data, and the access frequency of the initial document data, key data corresponding to the initial document data is generated.
[0016] Optionally, the contribution of keywords in sensitive information within the initial document data is calculated, including:
[0017] The first product is obtained by multiplying the squared frequency term of the keywords in the sensitive information of the initial document data with the access data amount of the keywords in the sensitive information of the initial document data;
[0018] The second product is obtained by multiplying the score of the keywords in the sensitive information of the initial document data by the length influence factor of the keywords in the sensitive information of the initial document data;
[0019] Calculate the logarithmic transformation of the amount of accessed data for keywords in the sensitive information of the initial document data;
[0020] Perform word segmentation on the initial document data and count the total number of words in the initial document data to obtain the total number of words in the initial document data;
[0021] Based on the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data, the first product, the second product, and the total number of words in the initial document data, the contribution of keywords in the sensitive information of the initial document data is determined.
[0022] Optionally, based on the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data (first product, second product, and the total number of words in the initial document data), the contribution of keywords in the sensitive information of the initial document data is determined, including:
[0023] Add the first product to the second product to get the first sum;
[0024] The first difference is obtained by subtracting the first sum from the logarithmic transformation of the amount of accessed data of keywords in the sensitive information of the initial document data;
[0025] Obtain the first preset value and the second preset value;
[0026] Calculate the ratio of the total number of words in the initial document data to the first preset value to obtain the first ratio;
[0027] Calculate the sum of the first ratio and the second preset value to obtain the second sum value;
[0028] Calculate the ratio of the first difference to the second sum to obtain the contribution of keywords in the sensitive information of the initial document data.
[0029] Optionally, based on keywords in sensitive information within the initial document data, the contribution of those keywords, access permissions, and access frequency of the initial document data, key data corresponding to the initial document data is generated, including:
[0030] Based on the contribution of keywords in the sensitive information of the initial document data and the logarithmic transformation of keywords in the sensitive information of the initial document data, assign corresponding keyword scores to keywords in the sensitive information of the initial document data;
[0031] Configure the corresponding permission score based on the access permissions of the initial document data at the permission level;
[0032] A frequency score is obtained based on the proportion of the access frequency of the initial document data to all document data.
[0033] The total word count score is obtained based on the proportion of the total word count in the initial document data to all document data.
[0034] The keyness score of keywords in sensitive information in the initial document data is obtained by using keyword score, permission score, frequency score and total word count score;
[0035] Sort the keywords in the sensitive information of the initial document data in descending order according to their keyness scores, and output the key data corresponding to the initial document data.
[0036] Optionally, the method further includes:
[0037] Calculate the length impact factor of keywords in sensitive information within the initial document data;
[0038] Calculate the length impact factor of keywords in sensitive information within the initial document data, including:
[0039] Extract the length and third preset value of keywords from sensitive information in the initial document data;
[0040] The second difference is obtained by calculating the difference between the length of the third preset value and the length of the keywords in the sensitive information of the initial document data;
[0041] The length of keywords in the sensitive information of the initial document data is calculated as the ratio of the length of the keywords to the second difference, thus obtaining the keyword length influence factor in the sensitive information of the initial document data.
[0042] Optionally, target document data is generated based on the initial document data, the key data corresponding to the initial document data, and the operation data of sensitive information, including:
[0043] Based on the key data corresponding to the initial document data, obtain the access type and index information of the key data corresponding to the initial document data;
[0044] The target operation item in the operation data for obtaining sensitive information;
[0045] Based on the initial document data, the target operation items in the sensitive information operation data, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data, the target document data is generated.
[0046] Optionally, based on the initial document data, the target operation items in the sensitive information operation data, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data, target document data is generated, including:
[0047] The position of the key data in the initial document data is obtained by using the index position in the index information of the key data corresponding to the initial document data;
[0048] The target operation items and the corresponding sensitive information processing rules in the operation data are configured by defining a dictionary.
[0049] By analyzing the access type of the key data corresponding to the initial document data and the location of the key data in the initial document data, the detection of the key data is performed to obtain sensitive information in the key data.
[0050] The target operation item in the operation data is obtained from the key data, and the sensitive information is processed using the sensitive information processing rules corresponding to the target operation item to obtain the processed key data.
[0051] The target document data is generated by updating the key data corresponding to the initial document data with the processed key data.
[0052] Optionally, sensitive information in the initial document data can be obtained based on fingerprint information, including:
[0053] Obtain the user ID through fingerprint information;
[0054] The initial document data is traversed by user ID, and a preset function is used to find sensitive information in the initial document data.
[0055] According to another aspect of the present invention, a data leakage prevention device is provided, comprising:
[0056] The data acquisition module is used to acquire initial document data;
[0057] The information extraction module is used to obtain sensitive information from the initial document data based on fingerprint information. The initial document data is configured with different levels, and the sensitive information in the initial document data is data of the sensitive level.
[0058] The key data generation module is used to generate key data corresponding to the initial document data based on sensitive information in the initial document data and the access permissions of the initial document data;
[0059] The target document generation module is used to generate target document data based on the initial document data, the key data and sensitive information operation data corresponding to the initial document data.
[0060] According to the present invention, initial document data is first acquired, and then sensitive information within the initial document data is obtained based on fingerprint information. By acquiring sensitive information from the initial document data based on fingerprint information, key data in the document can be accurately located and identified, ensuring that only truly sensitive information is protected and avoiding omissions or misjudgments. Furthermore, based on the sensitive information in the initial document data and the access permissions of the initial document data, key data corresponding to the initial document data is generated. This generation of key data based on the access permissions of the initial document data ensures that only authorized users can access sensitive information. This access control mechanism effectively prevents unauthorized access and data leakage. Therefore, based on the initial document data, the corresponding key data, and the operation data of sensitive information, target document data is generated. Through the operation data of sensitive information, access and operation behaviors to sensitive information can be monitored in real time, promptly detecting and preventing potential leakage risks, effectively protecting data security and privacy. Attached Figure Description
[0061] Figure 1 A flowchart of a data leakage prevention method according to an embodiment of the present invention is shown;
[0062] Figure 2 A structural block diagram of a data leakage prevention device according to an embodiment of the present invention is shown. Detailed Implementation
[0063] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0064] like Figure 1 As shown in the figure, the data leakage prevention method proposed in this embodiment includes the following steps:
[0065] Step S101: Obtain initial document data.
[0066] Initial document data refers to document data stored outside of the enterprise's information system on media such as personal computers, file servers, email, WeChat, and QQ, including Word, Excel, PPT, and PDF files. This document data may contain sensitive information about the enterprise, such as design documents, design drawings, source code, marketing plans, financial statements, and other content involving state secrets and corporate trade secrets.
[0067] In data loss prevention management, initial document data needs to be classified and graded. This is generally done in 3-5 levels, with different protection strategies applied to different levels. For example, sensitive documents are prohibited from being sent externally, while publicly accessible documents are not subject to any management measures.
[0068] Initial documentation typically consists of a leadership team, an operations team, and a business contact person team. The leadership team is responsible for the overall advancement of data leakage prevention management, the operations team is responsible for coordinating and implementing data leakage prevention work, and the business contact persons are the main force in promoting the implementation of data leakage prevention.
[0069] Step S102: Obtain sensitive information from the initial document data based on fingerprint information.
[0070] To obtain sensitive information from the initial document data based on fingerprint information, the user ID must first be obtained through the fingerprint information. Then, the initial document data is traversed through the user ID, and a preset function is used to find the sensitive information in the initial document data.
[0071] Specifically, when a user visits the website, the front-end uses JavaScript to obtain the browser fingerprint and generate a user ID, which is then sent to the server for storage. The server retrieves all document data for that user (i.e., the initial document data) based on the user ID. The client then iterates through the initial document data and uses a pre-defined sensitive information detection function to look for sensitive information.
[0072] Suppose a company needs to perform sensitive information detection on employee document data to prevent the leakage of sensitive content such as trade secrets and customer information. The company uses a document management system where employee document data is stored. Simultaneously, the company has deployed a sensitive information detection system to identify sensitive information within the documents.
[0073] The implementation steps are as follows:
[0074] Client retrieves document data
[0075] User Login: Employees log in to the document management system through the client. After the system verifies the user's identity, it allows the user to access document data within the scope of their permissions.
[0076] Document list retrieval: The client sends a request to the document management system to retrieve a list of all documents for that user. The document management system retrieves all of the user's documents from the database or file storage based on the user ID and returns the document list to the client.
[0077] Client-side document data traversal
[0078] Document list display: After receiving the document list, the client displays it to the user on the interface. The user can see all their documents, including document name, creation time, modification time, and other information.
[0079] Document Traversal: The client begins traversing each document in the document list one by one. For each document, the client requests the document management system to retrieve the document's detailed content. The document management system reads the document content from storage based on the document ID and returns it to the client.
[0080] Use the preset sensitive information detection function to find sensitive information.
[0081] Sensitive Information Detection Function: The client has a built-in preset sensitive information detection function that identifies sensitive information based on a series of predefined rules. These rules may include keyword matching, regular expressions, and data format checks. For example, the detection function might look for keywords such as "confidential," "financial data," "customer information," or "account password" in a document, or check for data in a format that matches ID card numbers, bank card numbers, or email addresses.
[0082] Document content analysis: The client passes the acquired document content to a sensitive information detection function. The detection function analyzes the document content line by line or paragraph by paragraph, checking for the presence of sensitive information according to preset rules.
[0083] Sensitive Information Tagging: If the detection function finds sensitive information in a document, it will tag this sensitive information and record relevant information, such as the document name, the content of the sensitive information, and its location. The client stores this tagged sensitive information in a separate list or report for later processing.
[0084] Sensitive Information Processing and Feedback
[0085] Sensitive Information Report Generation: Based on the detection results, the client generates a sensitive information report, which details all documents containing sensitive information and their related information. The report can be presented to the user in tables, lists, or other visual formats.
[0086] User Feedback and Handling: The client displays the sensitive information report to the user, who can review the report content to confirm whether there are any false alarms. If the user confirms that the information in the document is indeed sensitive, the client provides options for further processing of the document, such as encryption, deletion, or marking it as processed. Simultaneously, the client can also send the sensitive information report to the enterprise's security administrator for subsequent security audits and handling.
[0087] This application embodiment enables enterprises to promptly identify sensitive information in employee documents by traversing document data and using a preset sensitive information detection function, effectively reducing the risk of data leakage. Simultaneously, this process can be automated, improving work efficiency and reducing the workload of manual review.
[0088] Step S103: Based on the sensitive information in the initial document data and the access permissions of the initial document data, generate the key data corresponding to the initial document data.
[0089] To generate key data corresponding to the initial document data based on sensitive information and access permissions of the initial document data, it is necessary to first obtain the keywords in the sensitive information of the initial document data, then calculate the contribution of the keywords in the sensitive information of the initial document data, and then obtain the access frequency of the initial document data. Thus, based on the keywords in the sensitive information of the initial document data, the contribution of the keywords in the sensitive information of the initial document data, the access permissions of the initial document data, and the access frequency of the initial document data, the key data corresponding to the initial document data is generated.
[0090] This application's embodiments, by comprehensively considering keywords, keyword contribution, access frequency, and access permissions of sensitive information, can more accurately identify which documents are critical data, avoiding misjudgments caused by relying on a single factor. Furthermore, it can optimize resource allocation, helping enterprises concentrate limited security resources on critical data, improving the efficiency and effectiveness of data protection. For highly critical documents, stricter security measures can be adopted, such as encryption, access control, and auditing; while for less critical documents, the strength of security measures can be appropriately reduced, saving resources.
[0091] Specifically, to obtain keywords from sensitive information in the initial document data, sensitive information detection must first be performed. This can be done using pre-set sensitive information detection tools (such as keyword matching, regular expressions, or machine learning models) to scan the document data and identify sensitive information within it.
[0092] Then, keywords are extracted, that is, keywords are extracted from the identified sensitive information. For example, if the document contains sensitive information such as "financial statements," "customer list," or "contract terms," these words are extracted as keywords.
[0093] To calculate the contribution of keywords in sensitive information within the initial document data, the following steps are taken: First, calculate the product of the frequency squared term of the keywords in the sensitive information of the initial document data and the access data volume of the keywords in the sensitive information of the initial document data to obtain the first product. Then, calculate the product of the score of the keywords in the sensitive information of the initial document data and the length influence factor of the keywords in the sensitive information of the initial document data to obtain the second product. Next, calculate the logarithmic transformation of the access data volume of the keywords in the sensitive information of the initial document data. Then, perform word segmentation on the initial document data and count the total number of all words in the initial document data to obtain the total number of words in the initial document data. Based on the first product, the second product, the logarithmic transformation of the access data volume of the keywords in the sensitive information of the initial document data, and the total number of words in the initial document data, the contribution of keywords in sensitive information within the initial document data is determined.
[0094] The contribution of keywords in the sensitive information of the initial document data is determined based on the first product, the second product, the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data, and the total number of words in the initial document data. This includes: adding the first product and the second product to obtain a first sum; subtracting the first sum from the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data to obtain a first difference; obtaining a first preset value and a second preset value; calculating the ratio of the total number of words in the initial document data to the first preset value to obtain a first ratio; calculating the sum of the first ratio and the second preset value to obtain a second sum; and calculating the ratio of the first difference to the second sum to obtain the contribution of keywords in the sensitive information of the initial document data.
[0095] The formula for calculating the contribution of keywords in sensitive information within the initial document data is as follows:
[0096] ,
[0097] in, This indicates keywords in sensitive information within the initial document data. This indicates the frequency of keywords in sensitive information within the initial document data. This indicates the amount of data accessed for keywords in sensitive information within the initial document data. This represents the score of keywords in the sensitive information within the initial document data. This represents the influence factor of the length of keywords in sensitive information within the initial document data. This represents the total number of words in the initial document data. This represents the first preset value. This indicates the second preset value.
[0098] This application's embodiments comprehensively assess the contribution of keywords to document sensitivity by integrating multiple dimensions such as keyword frequency, access data volume, keyword score, length impact factor, and total document word count. Compared to single-dimensional assessment methods, this comprehensive assessment approach more accurately reflects the importance and sensitivity of keywords in a document, avoiding misjudgments due to overlooking certain key factors. Furthermore, by accurately calculating the contribution of each keyword, truly critical sensitive information in a document can be quickly located. For example, in a document containing a large amount of general information and a small amount of sensitive information, this method can accurately identify keywords that play a decisive role in document sensitivity, such as "financial statements" and "customer lists," thereby helping enterprises or organizations focus their attention and resources on this key information and improve the efficiency of data management and protection.
[0099] Furthermore, the introduction of access data volume in this embodiment allows the calculation of keyword contribution to dynamically reflect the real-time status of the document. A higher document access frequency indicates greater importance in actual use, potentially involving more business activities and sensitive information interaction. By incorporating access data volume into the calculation, keywords frequently accessed in actual use and containing sensitive information can be identified promptly, thus better addressing the dynamically changing data environment. As document content is updated and access patterns change, keyword contribution will adjust accordingly. This dynamic adaptability enables enterprises or organizations to continuously and accurately grasp changes in document sensitivity, taking timely and appropriate data protection measures to ensure critical data remains under effective control. Moreover, by calculating keyword contribution, this embodiment allows enterprises to clearly identify which keywords contribute most to document sensitivity, thus focusing limited security resources and management efforts on these critical data. Based on accurately calculated keyword contribution, enterprises can more effectively protect critical data and prevent sensitive information leakage. By identifying and protecting keywords that contribute most to document sensitivity, the risk of data leakage can be effectively reduced, protecting important assets such as business secrets and customer information.
[0100] The calculation of the keyword length influence factor in the sensitive information of the initial document data includes: obtaining the keyword length and a third preset value in the sensitive information of the initial document data; calculating the difference between the third preset value and the keyword length in the sensitive information of the initial document data to obtain a second difference value; and calculating the ratio of the keyword length in the sensitive information of the initial document data to the second difference value to obtain the keyword length influence factor in the sensitive information of the initial document data.
[0101] Specifically, the formula for calculating the influence factor of keyword length in sensitive information within the initial document data is as follows:
[0102] ,
[0103] in, This indicates the length of keywords in the sensitive information within the initial document data. This indicates the third preset value.
[0104] This application's embodiments consider the impact of keyword length on sensitivity and, by calculating a length influence factor, can balance the influence of keyword length on sensitivity, making the contribution calculation more reasonable and accurate. For example, both "financial" and "financial statements" are related to finance, but "financial statements" is more specific. Its length influence factor reflects this difference, thus giving it a more reasonable weight in the contribution calculation and avoiding length bias. Ignoring the impact of keyword length could lead to misjudgments of sensitivity. For instance, a shorter keyword may appear frequently in a document, but if its length is short, its actual sensitivity may not be high. By introducing a length influence factor, sensitivity assessment bias caused by differences in keyword length can be avoided, improving the reliability of the assessment results.
[0105] Because keyword length distributions can vary across different languages and document types, this method can adapt to multiple languages and document types by adjusting a third preset value, enhancing the flexibility and adaptability of the assessment. For example, keywords may be shorter in Chinese documents, while they may be longer in English documents. By appropriately setting the third preset value, it can be ensured that the length impact factor effectively reflects keyword sensitivity across different languages and document types. Furthermore, the formula for calculating the length impact factor quantifies the relationship between keyword length and the preset value, providing a scientific quantitative indicator for sensitivity assessment. This quantitative method makes the assessment process more objective and operable, avoiding the uncertainty caused by subjective judgment. By calculating the length impact factor, the contribution of keyword length to sensitivity can be accurately measured, thus providing a more accurate weight in the contribution calculation.
[0106] Suppose a company needs to conduct a sensitivity assessment of documents in its internal document management system to determine which keywords contribute most to the sensitivity of those documents. These documents may contain sensitive content such as trade secrets, customer information, and financial data. The company wants to use scientific methods to calculate the contribution of each keyword in order to better manage and protect these documents.
[0107] The implementation steps are as follows:
[0108] 1. Calculate the product of the squared frequency term of the keyword and the amount of data accessed (first product).
[0109] Keyword frequency statistics: Analyze each document to count the frequency of each sensitive keyword. For example, the keyword "financial statement" appears 5 times in a certain document.
[0110] Access data volume acquisition: Obtain the access data volume of this document from the access logs of the document management system, that is, the total number of times the document has been accessed. Assume that the document has been accessed 100 times in the past month.
[0111] Calculate the frequency square term: Square the frequency of the keyword. For example, the frequency square term for the keyword "financial statements" is 5^2 = 25.
[0112] Calculate the first product: Multiply the squared frequency term by the amount of data accessed. For example, the first product is 25 × 100 = 2500.
[0113] 2. Calculate the product of the keyword score and the length influence factor (second product).
[0114] Keyword Score Setting: A score is preset based on the sensitivity of the keywords. For example, the score for "financial statements" is 10.
[0115] Keyword Length Statistics: Calculates the length of keywords, for example, "financial statements" has 4 characters. Third Preset Value Setting: Sets a third preset value, let's say 10. Calculate Length Influence Factor: Calculates the difference between the third preset value and the keyword length (second difference): 10 - 4 = 6.
[0116] Calculate the ratio of keyword length to the second difference (length influence factor): 4 / 6 = 0.67. Calculate the second product: multiply the keyword score by the length influence factor. For example, the second product is 10 × 0.67 = 6.7.
[0117] 3. Calculate the logarithmic transformation of the amount of data accessed by keywords.
[0118] Logarithmic transformation: Perform a logarithmic transformation on the amount of data accessed for the keyword. Assuming the natural logarithm is used, taking the logarithm of 100 yields 4.6.
[0119] 4. Perform word segmentation on the document and count the total number of words.
[0120] Word segmentation: The document is segmented into individual words. For example, if the document content is "This is a detailed explanation of financial statements", the segmented result is "This is a detailed explanation of financial statements".
[0121] Total word count: Count the total number of words after word segmentation. For example, the total word count of this document is 8.
[0122] 5. Determine the contribution of keywords.
[0123] To calculate the first sum: add the first product to the second product. For example, the first sum is 2500 + 6.7 = 2506.7.
[0124] Calculate the first difference: Subtract the first sum from the logarithmic transformation of the accessed data amount. For example, the first difference is 2506.7 - 4.6 = 2502.1.
[0125] Preset value setting: Set the first preset value and the second preset value. For example, the first preset value is 100 and the second preset value is 5.
[0126] Calculate the first ratio: Calculate the ratio of the total number of words to the first preset value. For example, the first ratio is 8 / 100 = 0.08.
[0127] Calculate the second sum: Add the first ratio to the second preset value. For example, the second sum is 0.08 + 5 = 5.08.
[0128] Calculate the contribution: Calculate the ratio of the first difference to the second sum. For example, the contribution is 2502.1 / 5.08 = 492.54.
[0129] This application's embodiments, by comprehensively considering multiple factors such as keyword frequency, access data volume, and keyword length, can more accurately assess the contribution of each keyword to document sensitivity. Furthermore, the contribution of keywords can be dynamically adjusted based on document access frequency and content changes, ensuring that the assessment results always reflect the document's current state. Further, it helps enterprises identify the keywords that contribute most to document sensitivity, thereby concentrating resources on these key points, improving data protection efficiency, better protecting critical data, preventing sensitive information leakage, and enhancing overall data security.
[0130] The process involves generating key data corresponding to the initial document data based on keywords in sensitive information, their contribution, access permissions, and access frequency. This includes: assigning keyword scores to keywords in sensitive information based on their contribution and logarithmic transformation; assigning access permission scores based on access permission levels; obtaining frequency scores based on the proportion of access frequency of the initial document data to all document data; obtaining total word count scores based on the proportion of total word count of the initial document data to all document data; obtaining keyness scores for keywords in sensitive information based on keyword scores, access permission scores, frequency scores, and total word count scores; and sorting the keywords in sensitive information in the initial document data in descending order according to their keyness scores, outputting the key data corresponding to the initial document data.
[0131] Based on the above implementation steps, first configure the keyword score. Take a logarithmic transformation of the keyword contribution value of 492.54, assuming it to be 6.2. Then, configure the keyword score based on the logarithmic transformation value of 6.2. Assuming the logarithmic transformation value is directly used as the keyword score, the score for the keyword "financial statements" would be approximately 6.2.
[0132] Configure permission scores: The document's access permissions are set to "Senior Administrators," with a default permission level of "High." Configure permission scores based on the permission level; for example, assume a "High" permission score of 8.
[0133] Calculate the frequency score: The document's access frequency is 100 times, and the total access frequency of all documents is 1000 times, with a frequency percentage of 0.1. Configure the frequency score based on the frequency percentage. Assuming the frequency percentage is multiplied by 100 to obtain the frequency score, the frequency score is 10.
[0134] To obtain the total word count score: If the total word count of a document is 8 and the total word count of all documents is 1000, then the total word count percentage is 0.008. The total word count score is configured based on this percentage. For example, if the total word count percentage is multiplied by 100 to obtain the total word count score, then the total word count score is 0.8.
[0135] Keyword criticality score calculation: Keyword score = 6.2, Authority score = 8, Frequency score = 10, Total word count score = 0.8, Criticality score = 25.
[0136] Sort all keywords in descending order of their key scores: Suppose another keyword, "customer list," has a key score of 20, and another keyword, "meeting minutes," has a key score of 5.
[0137] The sorting results are: financial statements (25), customer list (20), meeting minutes (5).
[0138] Output key data: Output the sorted keywords and their key scores as key data for the document.
[0139] This application's embodiments, by comprehensively considering factors such as keyword contribution, access permissions, access frequency, and the total number of words in the document, can accurately identify truly critical sensitive information in a document. This method avoids misjudgments caused by relying on a single factor, improving the accuracy and reliability of the assessment. Furthermore, the calculation of keyword contribution takes into account the dynamic changes in access frequency and the total number of words in the document, enabling timely reflection of the document's actual usage and changes in sensitivity. This dynamic adaptability allows enterprises to continuously and accurately grasp the sensitivity status of documents and adjust their data protection strategies in a timely manner.
[0140] Furthermore, by calculating and ranking keyword severity scores, enterprises can clearly identify which keywords contribute the most to the sensitivity of documents, thus focusing limited security resources on these critical data. This method improves resource utilization efficiency and ensures that critical data receives focused protection. It effectively identifies and protects critical data, preventing the leakage of sensitive information. By identifying and protecting the keywords that contribute the most to the sensitivity of documents, enterprises can reduce the risk of data breaches and protect important assets such as trade secrets and customer information.
[0141] Step S104: Generate target document data based on the initial document data, the key data corresponding to the initial document data, and the operation data of sensitive information.
[0142] To generate target document data based on initial document data, key data corresponding to the initial document data, and sensitive information operation data, it is necessary to first obtain the access type and index information of the key data corresponding to the initial document data, then obtain the target operation items in the operation data of sensitive information, and finally generate the target document data based on the initial document data, the target operation items in the operation data of sensitive information, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data.
[0143] The target operation items include at least access control operations, anomaly detection operations, and backup operations.
[0144] There is an internal document management system for the company, which stores various sensitive and important document data, such as employee personal information, financial reports, trade secrets, etc.
[0145] The initial document data is taken as an example of a document containing employees' personal information, including key data such as employees' ID card numbers, bank card numbers, and home addresses.
[0146] Obtain the access type and index information for key data corresponding to the initial document data. Access type: Through system permission settings, determine that only specific personnel in the human resources department (such as HR specialists) and the employee themselves have access rights. For example, HR specialists can view and edit this key data, while employees can only view their own information. Index information: Create indexes for this key data, such as indexing by employee ID, name, etc., to facilitate quick location and retrieval.
[0147] The system tracks target operations within the data used to access sensitive information. Specifically, access control operations record every instance of access to critical data, including access time, visitor identity, and purpose. For example, when an HR specialist queries an employee's ID number, the system records this operation. Anomaly detection operations monitor access to critical data in real time. If anomalies are detected, such as frequent access to the same employee's critical data or a mismatch between the visitor's identity and access permissions, an alert is immediately issued. For example, if an employee outside of HR attempts to access another employee's bank card number, the system detects and records this anomaly. Backup operations regularly back up critical data to ensure data security and recoverability. For example, the system automatically backs up all employees' critical data every night and stores the backup files on a secure server.
[0148] Based on the initial document data, the target operation items in the sensitive information operation data, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data, target document data is generated. This target document data includes not only the original employee personal information but also the following:
[0149] Access logs: Record every access to key data, including access time, visitor identity, and purpose of access.
[0150] Anomaly detection log: Records abnormal behaviors detected by the system, including a detailed description of the abnormal behavior, the time of occurrence, and the key data involved.
[0151] Backup logs: Record information such as the time of the backup operation, the storage location of the backup file, and the integrity verification of the backup file.
[0152] This application's embodiments ensure that only authorized personnel can access critical data through access control operations, preventing data leakage. Anomaly detection operations can promptly identify and prevent potential security threats, protecting critical data from malicious attacks. Furthermore, indexing information makes critical data retrieval more efficient, enabling rapid location and retrieval of required information. Simultaneously, backup operations ensure data recoverability; even in the event of data loss or corruption, data can be recovered from backup files.
[0153] Furthermore, the access records, anomaly detection records, and backup records contained in the target document data provide strong support for data auditing and traceability. Enterprises can view the access and operation history of critical data at any time, ensuring compliant data use and quickly locating the cause and taking corrective action when problems occur.
[0154] The process of generating target document data involves several steps. First, based on the initial document data, the target operation items in the sensitive information operation data, the key data corresponding to the initial document data, and the access types and index information of the key data corresponding to the initial document data, the process includes: obtaining the position of the key data in the initial document data using the index position in the index information of the key data corresponding to the initial document data; configuring the target operation items in the sensitive information operation data and the sensitive information processing rules corresponding to the target operation items using a defined dictionary; performing key data detection based on the access types of the key data in the initial document data and the position of the key data in the initial document data to obtain sensitive information; obtaining the target operation items in the sensitive information operation data of the key data, and processing the sensitive information using the sensitive information processing rules corresponding to the target operation items to obtain processed key data; and updating the key data corresponding to the initial document data using the processed key data to generate the target document data.
[0155] A company has a customer information management system that stores detailed customer information, including name, ID number, bank card number, contact number, and home address. The company needs to manage and protect this sensitive information to ensure data security and compliance.
[0156] The initial document data is a table containing customer information, including: customer number, customer name, ID card number, bank card number, contact number, home address, etc.
[0157] The system has set up index information for each key data field to facilitate quick location and access. The index information is as follows:
[0158] ID card number: Column 3
[0159] Bank card number: Column 4
[0160] The system defines a dictionary of sensitive information operation data, which includes target operation items and corresponding sensitive information processing rules:
[0161] 1. Access Control:
[0162] Rule: Only authorized personnel are allowed to access sensitive information. Action: Record visitor identity and access time.
[0163] 2. Data anonymization:
[0164] Rule: De-identify sensitive information and hide some of it. Operation: Hide the middle 8 digits of the ID card number and the middle 10 digits of the bank card number.
[0165] 3. Anomaly Detection:
[0166] Rule: Detect abnormal access behavior, such as frequent access to the same sensitive information. Action: Log the abnormal behavior and issue an alert.
[0167] Based on the index information, the system determined the location of the key data within the initial document data:
[0168] The ID number is in column 3.
[0169] The bank card number is in column 4.
[0170] Based on the access type of the critical data (e.g., "access control"), the system performs detection operations at the location of the critical data. For example:
[0171] It was detected that user A accessed customer 001's ID card number and bank card number.
[0172] The system recorded the visitor's identity as user A, and the access time as 14:00 on March 26, 2025.
[0173] According to the defined rules for handling sensitive information, the system processes the detected sensitive information as follows:
[0174] The ID number "123456789012345678" is processed into "1234XXXXXXXX5678" (the middle 8 digits are hidden).
[0175] The bank card number "6228480000000000000" will be changed to "622848XXXXXXXXXX000" (the middle 10 digits will be hidden).
[0176] The system updates the initial document data with the processed key data to generate the final target document data.
[0177] This application embodiment ensures that only authorized personnel can access sensitive information by recording the visitor's identity and access time, preventing data leakage. Anomaly detection can promptly detect and block abnormal access behavior, protecting sensitive information from malicious attacks. Furthermore, sensitive information is anonymized to hide some information, so even if the data is leaked, the complete sensitive information cannot be directly obtained, thus protecting user privacy.
[0178] Furthermore, embodiments of this application also use index information to quickly locate the position of key data, thereby improving data retrieval efficiency. By defining sensitive information operation data and processing rules, automated processing is achieved, reducing manual intervention and improving processing efficiency.
[0179] Furthermore, through strict access control, data anonymization, and anomaly detection measures, data processing is ensured to comply with relevant laws and regulations, avoiding legal risks caused by data leaks and other issues. As can be seen from the above embodiments, this data processing method can effectively improve data security and management efficiency, protect user privacy, and meet regulatory requirements. It is applicable to fields such as enterprise internal customer information management systems, financial institution customer information management systems, and government department citizen information management systems.
[0180] According to the present invention, initial document data is first acquired, and then sensitive information within the initial document data is obtained based on fingerprint information. By acquiring sensitive information from the initial document data based on fingerprint information, key data in the document can be accurately located and identified, ensuring that only truly sensitive information is protected and avoiding omissions or misjudgments. Furthermore, based on the sensitive information in the initial document data and the access permissions of the initial document data, key data corresponding to the initial document data is generated. This generation of key data based on the access permissions of the initial document data ensures that only authorized users can access sensitive information. This access control mechanism effectively prevents unauthorized access and data leakage. Therefore, based on the initial document data, the corresponding key data, and the operation data of sensitive information, target document data is generated. Through the operation data of sensitive information, access and operation behaviors to sensitive information can be monitored in real time, promptly detecting and preventing potential leakage risks, effectively protecting data security and privacy.
[0181] Figure 2 An embodiment of the present invention provides a data leakage prevention device, such as... Figure 2 As shown, the device includes:
[0182] Data acquisition module 201 is used to acquire initial document data;
[0183] Information extraction module 202 is used to obtain sensitive information in the initial document data based on fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of the sensitive level;
[0184] The key data generation module 203 is used to generate key data corresponding to the initial document data based on sensitive information in the initial document data and the access permissions of the initial document data;
[0185] The target document generation module 204 is used to generate target document data based on the initial document data, the key data and sensitive information operation data corresponding to the initial document data.
[0186] Optionally, the key data generation module 203 is also used to obtain keywords from sensitive information in the initial document data;
[0187] Calculate the contribution of keywords in sensitive information within the initial document data;
[0188] The frequency of accessing initial document data;
[0189] Based on the keywords in the sensitive information of the initial document data, the contribution of the keywords in the sensitive information of the initial document data, the access permissions of the initial document data, and the access frequency of the initial document data, key data corresponding to the initial document data is generated.
[0190] Optionally, the key data generation module 203 is also used to calculate the product of the frequency squared term of the keywords in the sensitive information of the initial document data and the access data amount of the keywords in the sensitive information of the initial document data to obtain the first product;
[0191] The second product is obtained by multiplying the score of the keywords in the sensitive information of the initial document data by the length influence factor of the keywords in the sensitive information of the initial document data;
[0192] Calculate the logarithmic transformation of the amount of accessed data for keywords in the sensitive information of the initial document data;
[0193] Perform word segmentation on the initial document data and count the total number of words in the initial document data to obtain the total number of words in the initial document data;
[0194] Based on the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data, the first product, the second product, and the total number of words in the initial document data, the contribution of keywords in the sensitive information of the initial document data is determined.
[0195] Optionally, the key data generation module 203 is also used to add the first product and the second product to obtain a first sum;
[0196] The first difference is obtained by subtracting the first sum from the logarithmic transformation of the amount of accessed data of keywords in the sensitive information of the initial document data;
[0197] Obtain the first preset value and the second preset value;
[0198] Calculate the ratio of the total number of words in the initial document data to the first preset value to obtain the first ratio;
[0199] Calculate the sum of the first ratio and the second preset value to obtain the second sum value;
[0200] Calculate the ratio of the first difference to the second sum to obtain the contribution of keywords in the sensitive information of the initial document data.
[0201] Optionally, the key data generation module 203 is also used to configure corresponding keyword scores for the keywords in the sensitive information of the initial document data based on the contribution of the keywords in the sensitive information of the initial document data and the logarithmic transformation of the keywords in the sensitive information of the initial document data.
[0202] Configure the corresponding permission score based on the access permissions of the initial document data at the permission level;
[0203] A frequency score is obtained based on the proportion of the access frequency of the initial document data to all document data.
[0204] The total word count score is obtained based on the proportion of the total word count in the initial document data to all document data.
[0205] The keyness score of keywords in sensitive information in the initial document data is obtained by using keyword score, permission score, frequency score and total word count score;
[0206] Sort the keywords in the sensitive information of the initial document data in descending order according to their keyness scores, and output the key data corresponding to the initial document data.
[0207] Optionally, the apparatus further includes: calculating a length influence factor for keywords in sensitive information within the initial document data;
[0208] Calculate the length impact factor of keywords in sensitive information within the initial document data, including:
[0209] Extract the length and third preset value of keywords from sensitive information in the initial document data;
[0210] The second difference is obtained by calculating the difference between the length of the third preset value and the length of the keywords in the sensitive information of the initial document data;
[0211] The length of keywords in the sensitive information of the initial document data is calculated as the ratio of the length of the keywords to the second difference, thus obtaining the keyword length influence factor in the sensitive information of the initial document data.
[0212] Optionally, the target document generation module 204 is also used to obtain the access type and index information of the key data corresponding to the initial document data based on the key data corresponding to the initial document data;
[0213] The target operation item in the operation data for obtaining sensitive information;
[0214] Based on the initial document data, the target operation items in the sensitive information operation data, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data, the target document data is generated.
[0215] Optionally, the target document generation module 204 is also used to obtain the position of the key data in the initial document data through the index position in the index information of the key data corresponding to the initial document data;
[0216] The target operation items and the corresponding sensitive information processing rules in the operation data are configured by defining a dictionary.
[0217] By analyzing the access type of the key data corresponding to the initial document data and the location of the key data in the initial document data, the detection of the key data is performed to obtain sensitive information in the key data.
[0218] The target operation item in the operation data is obtained from the key data, and the sensitive information is processed using the sensitive information processing rules corresponding to the target operation item to obtain the processed key data.
[0219] The target document data is generated by updating the key data corresponding to the initial document data with the processed key data.
[0220] Optionally, the information extraction module 202 is also used to obtain the user ID through fingerprint information;
[0221] The initial document data is traversed by user ID, and a preset function is used to find sensitive information in the initial document data.
[0222] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used with the examples of this invention. The required structure for constructing such systems is apparent from the above description. Furthermore, this invention is not directed to any particular programming language. It should be understood that the contents of the invention described herein can be implemented using various programming languages, and the above description of specific languages is for the purpose of disclosing preferred embodiments of the invention.
[0223] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0224] Those skilled in the art will understand that modules, units, or components of the devices disclosed in the examples herein can be arranged in the devices described in this embodiment, or alternatively, can be located in one or more devices different from the devices in this example. The modules in the foregoing examples can be combined into a single module or, in addition, can be divided into multiple sub-modules.
[0225] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed herein and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed herein may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0226] Furthermore, some of the embodiments described herein are methods or combinations of method elements that can be implemented by a processor of a computer system or by other means of performing the functions. Therefore, a processor having the necessary instructions for implementing the methods or method elements forms means for implementing the methods or method elements. Furthermore, the elements described herein in the apparatus embodiments are examples of means for implementing the functions performed by elements for the purposes of carrying out the invention.
[0227] As used herein, unless otherwise specified, the use of ordinal numbers such as “first,” “second,” “third,” etc., to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects being described must have a given order in time, space, ordering, or any other manner.
[0228] Although the invention has been described with reference to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and edibility purposes, and not for the purpose of interpreting or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended specification. Regarding the scope of the invention, the disclosure made is illustrative rather than restrictive, and the scope of the invention is defined by the appended specification.
Claims
1. A method for preventing data leakage, characterized in that, include: Obtain initial document data; Sensitive information in the initial document data is obtained based on fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of the sensitive level; Based on the sensitive information in the initial document data and the access permissions of the initial document data, generate key data corresponding to the initial document data; Based on the initial document data, the key data corresponding to the initial document data, and the operation data of the sensitive information, target document data is generated; The process of generating key data corresponding to the initial document data based on sensitive information and access permissions of the initial document data includes: Extract keywords from the sensitive information in the initial document data; Calculate the contribution of keywords in the sensitive information of the initial document data; Obtain the access frequency of the initial document data; Based on the keywords in the sensitive information of the initial document data, the contribution of the keywords in the sensitive information of the initial document data, the access permissions of the initial document data, and the access frequency of the initial document data, key data corresponding to the initial document data is generated; The calculation of the contribution of keywords in the sensitive information of the initial document data includes: The first product is obtained by multiplying the squared frequency term of the keywords in the sensitive information of the initial document data with the access data amount of the keywords in the sensitive information of the initial document data; The second product is obtained by multiplying the score of the keywords in the sensitive information of the initial document data by the length influence factor of the keywords in the sensitive information of the initial document data; Calculate the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data; The initial document data is segmented into words, and the total number of words in the initial document data is counted to obtain the total number of words in the initial document data; Based on the first product, the second product, the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data, and the total number of words in the initial document data, the contribution of keywords in the sensitive information of the initial document data is determined.
2. The data leakage prevention method according to claim 1, characterized in that, The determination of the contribution of keywords in the sensitive information of the initial document data based on the logarithmic transformation of the access data volume of the first product, the second product, and the keywords in the sensitive information of the initial document data, and the total number of words in the initial document data, includes: Add the first product to the second product to obtain the first sum; The first difference is obtained by subtracting the first sum from the logarithmic transformation of the access data amount of keywords in the sensitive information of the initial document data; Obtain the first preset value and the second preset value; Calculate the ratio of the total number of words in the initial document data to the first preset value to obtain the first ratio; Calculate the sum of the first ratio and the second preset value to obtain the second sum value; The ratio of the first difference to the second sum is calculated to obtain the contribution of keywords in the sensitive information of the initial document data.
3. The data leakage prevention method according to claim 2, characterized in that, The process of generating key data corresponding to the initial document data based on keywords in sensitive information of the initial document data, the contribution of keywords in sensitive information of the initial document data, access permissions of the initial document data, and access frequency of the initial document data includes: Based on the contribution of keywords in the sensitive information of the initial document data and the logarithmic transformation of the keywords in the sensitive information of the initial document data, assign corresponding keyword scores to the keywords in the sensitive information of the initial document data; The permission score is configured according to the access permission level of the initial document data. A frequency score is obtained based on the proportion of the access frequency of the initial document data to all document data. The total word count score is obtained based on the proportion of the total word count of the initial document data to all document data. The key score of keywords in sensitive information in the initial document data is obtained by using the keyword score, the permission score, the frequency score, and the total word count score. The keywords in the sensitive information of the initial document data are sorted in descending order according to their keyness scores, and the key data corresponding to the initial document data is output.
4. The data leakage prevention method according to claim 1, characterized in that, The method further includes: Calculate the length impact factor of keywords in the sensitive information of the initial document data; The calculation of the length influence factor of keywords in sensitive information in the initial document data includes: Obtain the length and a third preset value of the keywords in the sensitive information of the initial document data; Calculate the difference between the length of the third preset value and the length of the keyword in the sensitive information of the initial document data to obtain the second difference value; The length of keywords in the sensitive information of the initial document data is calculated as the ratio of the length of the keywords to the second difference to obtain the length influence factor of keywords in the sensitive information of the initial document data.
5. The data leakage prevention method according to claim 1, characterized in that, The process of generating target document data based on the initial document data, the key data corresponding to the initial document data, and the operation data of the sensitive information includes: Based on the key data corresponding to the initial document data, obtain the access type and index information of the key data corresponding to the initial document data; The target operation item in the operation data for obtaining the sensitive information; The target document data is generated based on the initial document data, the target operation item in the operation data of the sensitive information, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data.
6. The data leakage prevention method according to claim 5, characterized in that, The process of generating the target document data based on the initial document data, the target operation item in the operation data of the sensitive information, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data includes: The position of the key data in the initial document data is obtained by using the index position in the index information of the key data corresponding to the initial document data; The target operation items in the operation data of the sensitive information are configured by defining a dictionary, as well as the sensitive information processing rules corresponding to the target operation items; By analyzing the access type of the key data corresponding to the initial document data and the position of the key data in the initial document data, the detection of the key data is performed to obtain the sensitive information in the key data. The target operation item in the operation data of the sensitive information in the key data is obtained, and the sensitive information is processed by the sensitive information processing rules corresponding to the target operation item to obtain the processed key data. The target document data is generated by updating the key data corresponding to the initial document data with the processed key data.
7. The data leakage prevention method according to claim 1, characterized in that, The process of obtaining sensitive information from the initial document data based on fingerprint information includes: The user ID is obtained using the fingerprint information; The initial document data is traversed using the user ID, and a preset function is used to find sensitive information in the initial document data.
8. A data leakage prevention device, characterized in that, include: The data acquisition module is used to acquire initial document data; An information extraction module is used to obtain sensitive information from the initial document data based on fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of the sensitive level; The key data generation module is used to generate key data corresponding to the initial document data based on the sensitive information in the initial document data and the access permissions of the initial document data; The target document generation module is used to generate target document data based on the initial document data, the key data corresponding to the initial document data, and the operation data of the sensitive information; The process of generating key data corresponding to the initial document data based on sensitive information and access permissions of the initial document data includes: Extract keywords from the sensitive information in the initial document data; Calculate the contribution of keywords in the sensitive information of the initial document data; Obtain the access frequency of the initial document data; Based on the keywords in the sensitive information of the initial document data, the contribution of the keywords in the sensitive information of the initial document data, the access permissions of the initial document data, and the access frequency of the initial document data, key data corresponding to the initial document data is generated; The calculation of the contribution of keywords in the sensitive information of the initial document data includes: The first product is obtained by multiplying the squared frequency term of the keywords in the sensitive information of the initial document data with the access data amount of the keywords in the sensitive information of the initial document data; The second product is obtained by multiplying the score of the keywords in the sensitive information of the initial document data by the length influence factor of the keywords in the sensitive information of the initial document data; Calculate the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data; The initial document data is segmented into words, and the total number of words in the initial document data is counted to obtain the total number of words in the initial document data; Based on the first product, the second product, the logarithmic transformation of the access data volume of keywords in the sensitive information of the initial document data, and the total number of words in the initial document data, the contribution of keywords in the sensitive information of the initial document data is determined.
Citation Information
Patent Citations
Remodification, identification and alarm system and method for sensitive archives
CN119046933A
Archive management system and method based on data analysis
CN119830308A