A method for tracing confidential document leakage using big data
Through file recognition technology and deep learning natural language processing technology, combined with fine-grained semantic comparison analysis of the two-way attention mechanism, the problem of difficulty in accurately traceing confidential file leakage in existing technology is solved, and the accuracy of identification of the association between leaked files and the original files is improved.
Patent Information
- Application Number
- CN202510141928.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-09
AI Technical Summary
The existing methods of leaking confidential files have limitations, and it is difficult to accurately identify the correlation between leaked files after format conversion, modification or tampering and the original confidential files.
File recognition technology is used to extract file contents, and context semantic analysis is performed through natural language processing technology based on deep learning to extract semantic features. Then, a fine-grained semantic comparison analysis is performed through a two-way attention mechanism to determine whether the suspected leaked file is a variant of the original confidential file.
Improve the accuracy and reliability of confidential file leakage traceability, and can more accurately capture the association between leaked files and original confidential files after various processing.
Smart Images

Figure CN119577843B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information security, and more specifically, to a method for tracing the leakage of confidential documents using big data. Background Art
[0002] In the field of information security, the protection of confidential documents is of vital importance. However, in today's era of high information circulation, the security of confidential documents has been challenged unprecedentedly. With the development of network technology and the increasing convenience of data sharing, the risk of confidential document leakage has also increased accordingly. Although traditional security measures such as firewalls, encryption technology and access control can prevent unauthorized access to a certain extent, they are powerless against file leakage caused by internal personnel's misoperation or malicious behavior.
[0003] When it comes to confidential document leaks, tracing the source of the leak and determining the path of the leak are crucial for taking timely remedial measures and holding accountable. To address this issue, most existing tracing methods rely on file attributes (such as file name, creation time, modification time, etc.) or direct comparison of file contents. However, these methods have obvious limitations. On the one hand, file attributes (such as name or size) are easily tampered with and cannot be used as a reliable basis for tracing to directly determine whether it is a leaked confidential file; on the other hand, when faced with leaked files that have been processed by format conversion, partial deletion, addition of interference information, or slight modification, it is often difficult to accurately identify their relevance to the original confidential file by directly comparing the file content.
[0004] Therefore, an optimized method of tracing the leakage of confidential documents using big data is desired. Summary of the invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a method for tracing the leakage of confidential files using big data, which uses file recognition technology to extract file contents from suspected leaked files and original confidential files respectively, and introduces natural language processing technology based on deep learning to perform contextual semantic analysis on the suspected leaked file contents and the original confidential file contents, so as to extract the semantic features of the suspected leaked file contents and the original confidential file contents, and then, by performing fine-grained semantic comparative analysis of the two based on a bidirectional attention mechanism, the deep-level semantic similarities and differences between the two are revealed to determine whether the suspected leaked file is a variant of the original confidential file. In this way, by performing fine-grained comparative analysis of the suspected leaked file and the original confidential file at the contextual semantic level, the association between the leaked file after various processing and the original confidential file can be more accurately captured, thereby improving the accuracy and reliability of tracing the leakage of confidential files.
[0006] According to one aspect of the present application, a method for tracing the leakage of confidential documents using big data is provided, which includes:
[0007] Obtaining suspected leaked documents;
[0008] Extracting file content from the suspected leaked file to obtain the suspected leaked file content;
[0009] Access to original confidential documents;
[0010] Extracting file content from the original confidential file to obtain original confidential file content;
[0011] Performing semantic coding on the suspected leaked file content and the original confidential file content to obtain a suspected leaked file content semantic coding vector and an original confidential file content semantic coding vector;
[0012] Performing a semantic fine-grained comparative analysis based on bidirectional attention on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain a bidirectional fine-grained comparative coding vector of the suspected leaked file-original confidential file content;
[0013] Based on the bidirectional fine-grained comparison encoding vector of the suspected leaked file and the original confidential file content, it is determined whether the suspected leaked file is a variant of the original confidential file.
[0014] Compared with the prior art, the method for tracing the leakage of confidential documents using big data provided by the present application uses file recognition technology to extract the file contents from the suspected leaked files and the original confidential files respectively, and introduces natural language processing technology based on deep learning to perform contextual semantic analysis on the suspected leaked file contents and the original confidential file contents, so as to extract the semantic features of the suspected leaked file contents and the original confidential file contents, and then, by performing fine-grained semantic comparative analysis of the two based on a bidirectional attention mechanism, the deep-level semantic similarities and differences between the two are revealed to determine whether the suspected leaked file is a variant of the original confidential file. In this way, by performing fine-grained comparative analysis of the suspected leaked files and the original confidential files at the contextual semantic level, the association between the leaked files after various processing and the original confidential files can be more accurately captured, thereby improving the accuracy and reliability of tracing the leakage of confidential files. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 The present invention is a flowchart of a method for tracing confidential document leakage using big data according to an embodiment of the present application.
[0017] Figure 2 This is a data flow diagram of a method for tracing confidential document leakage using big data according to an embodiment of the present application.
[0018] Figure 3 This is a flowchart of sub-step S6 of the method for tracing confidential document leakage using big data according to an embodiment of the present application.
[0019] Figure 4 This is a flowchart of sub-step S62 of the method for tracing confidential document leakage using big data according to an embodiment of the present application.
[0020] Figure 5 This is a flowchart of sub-step S63 of the method for tracing confidential document leakage using big data according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0022] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0023] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. Meanwhile, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.
[0024] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.
[0025] It is worth noting that in this application, all actions to obtain data are carried out in compliance with the relevant data protection laws and policies of the country where the data is located, and with the authorization given by the owner of the corresponding device.
[0026] In response to the technical problems described in the above background technology, this application proposes a method for tracing the leakage of confidential documents using big data, which uses file recognition technology to extract file contents from suspected leaked files and original confidential files respectively, and introduces natural language processing technology based on deep learning to perform contextual semantic analysis on the suspected leaked file contents and the original confidential file contents, so as to extract the semantic features of the suspected leaked file contents and the original confidential file contents, and then, by performing fine-grained semantic comparative analysis of the two based on a bidirectional attention mechanism, the deep-level semantic similarities and differences between the two are revealed to determine whether the suspected leaked file is a variant of the original confidential file. In this way, by performing fine-grained comparative analysis of the suspected leaked file and the original confidential file at the contextual semantic level, the association between the leaked file after various processing and the original confidential file can be more accurately captured, thereby improving the accuracy and reliability of tracing the leakage of confidential files.
[0027] Figure 1 The present invention is a flowchart of a method for tracing confidential document leakage using big data according to an embodiment of the present application. Figure 2 The following is a data flow diagram of a method for tracing confidential document leakage using big data according to an embodiment of the present application. Figure 1 and Figure 2 As shown, the method for tracing the leakage of confidential files using big data includes the following steps: S1, obtaining a suspected leaked file; S2, extracting file content from the suspected leaked file to obtain the suspected leaked file content; S3, obtaining the original confidential file; S4, extracting file content from the original confidential file to obtain the original confidential file content; S5, semantically encoding the suspected leaked file content and the original confidential file content to obtain a suspected leaked file content semantic encoding vector and an original confidential file content semantic encoding vector; S6, performing a semantic fine-grained comparative analysis based on bidirectional attention on the suspected leaked file content semantic encoding vector and the original confidential file content semantic encoding vector to obtain a suspected leaked file-original confidential file content bidirectional fine-grained comparative encoding vector; S7, determining whether the suspected leaked file is a variant of the original confidential file based on the suspected leaked file-original confidential file content bidirectional fine-grained comparative encoding vector.
[0028] In the above method for tracing confidential file leakage using big data, the step S1 is to obtain suspected leaked files. Specifically, in practical applications, the methods for obtaining suspected leaked files include but are not limited to network monitoring, internal system log auditing, employee reporting, etc.
[0029] In terms of network monitoring, modern enterprise network environments usually deploy a variety of network security devices, such as firewalls, intrusion detection systems (IDS), intrusion prevention systems (IPS), and data loss prevention (DLP) tools. These devices can not only monitor the data traffic in and out of the network in real time, but also identify abnormal behavior or potential threats according to preset security policies. For example, DLP tools can be configured to automatically scan outgoing emails, files uploaded to cloud storage services, or content sent through instant messaging tools to find out whether there is unauthorized outflow of confidential information. Once a data packet or file that may contain sensitive information is detected, the system will immediately issue an alarm and automatically isolate or block the operation according to a pre-defined procedure, while saving the relevant data for further analysis. Deep packet inspection (DPI) technology can parse the application layer protocol more deeply, helping security teams identify suspicious transmissions hidden in legitimate traffic, thereby improving the ability to capture leaks in complex attack modes. In addition, with the development of machine learning and artificial intelligence, more and more companies are beginning to adopt intelligent algorithms to assist network monitoring. For example, machine learning-based user entity behavior analysis (UEBA) can create a user behavior baseline and identify deviations from normal behavior, such as logging in at abnormal times, downloading large amounts of data, or frequently accessing specific resources, by comparing the user's daily activity patterns. This helps to detect improper behavior by insiders or lateral movement by external attackers early. Network traffic analysis (NTA) is also an effective supplementary means, which focuses on anomalies at the network communication level, such as traffic surges, use of non-standard ports, and the emergence of unknown protocols. NTA can provide a detailed view of events occurring on the network by collecting and analyzing network flow records, helping security analysts identify possible leakage paths. In particular, for advanced persistent threats (APTs) that attempt to bypass traditional security controls, NTA can provide additional visibility and contextual information.
[0030] Internal system log auditing is also an indispensable part. Almost all enterprise information systems generate a large number of log records, including logs of operating systems, databases, applications, and various middleware. By regularly reviewing and intelligently analyzing these logs, we can find changes in the behavior patterns of internal personnel accessing confidential resources, as well as whether there are any illegal operations or abnormal activities. For example, if a user suddenly starts to frequently access sensitive document libraries that are rarely accessed, or tries to download a large amount of information, this may be a signal of leakage risk. In addition to traditional text logs, modern systems also support structured log formats such as JSON and XML, which are easy for automated tools to parse and process. In addition, centralized log management systems (such as ELK Stack and Splunk) can unify the collection, indexing, and visualization of logs scattered on different servers, greatly improving audit efficiency. For large enterprises, you can also consider introducing a SIEM (Security Information and Event Management) platform, which can not only integrate log data from multiple sources, but also integrate with other security tools to achieve comprehensive threat detection and response.
[0031] In addition to relying on technology and automated tools, employee reporting channels are also an important source of suspected leaked files. To this end, companies can set up dedicated reporting hotlines, emails, or other convenient communication platforms.
[0032] In the above-mentioned method of tracing the leakage of confidential documents using big data, the step S2 extracts the file content from the suspected leaked file to obtain the suspected leaked file content. In a specific example of the present application, the step S2 includes: using OCR technology to identify the text content of the suspected leaked file to obtain the suspected leaked file content. It should be understood that, considering that in actual applications, suspected leaked files may exist in various forms, including but not limited to paper documents, electronic documents, PDF files, etc. Therefore, in order to uniformly process files of different forms and extract their text content, the present application uses OCR (optical character recognition) technology to identify the text content of the suspected leaked file, so as to convert non-pure text files such as PDF files, scans, and image files into a processable text format to obtain the suspected leaked file content, thereby providing basic data for subsequent file content analysis.
[0033] Specifically, choose the appropriate processing method according to the file type. For PDF files that already contain editable text, you can usually directly use the built-in text extraction tool or API interface to obtain the text content. For those PDF files composed of images or direct scans and picture files, you need to rely on OCR technology for text recognition. In order to improve the efficiency and accuracy of OCR processing, it is recommended to perform quality inspection and preprocessing of the files in advance. For example, if the file is blurred, tilted, or has shadows, you can use image processing algorithms to perform denoising, rotation correction, and smoothing to make the image clearer and easier to read.
[0034] Next, you need to choose the right OCR software or service. There are many mature OCR solutions on the market to choose from, including open source libraries such as Tesseract, and commercial products such as ABBYY FineReader, Google CloudVision API, etc. These tools can not only recognize text in multiple languages, but also support complex typesetting structures such as tables, charts, multi-column text, etc. In the selection process, factors such as recognition speed, accuracy, supported languages, and the ability to customize training models need to be considered. In addition, for specialized terms in specific industries or fields, customized training of OCR engines may be required to adapt to special needs and improve recognition effects.
[0035] Next, you need to configure parameters and optimize settings. Different OCR systems provide a series of adjustable options for adjusting recognition behavior. For example, you can specify the language to be recognized, set the page layout mode (automatic detection or fixed format), select the character set range (ASCII only or including special symbols), and even fine-tune some advanced parameters such as binarization threshold, connected domain filtering conditions, etc. according to the actual application scenario. Reasonable configuration of these parameters can help improve the quality of recognition results and reduce the occurrence of misrecognition.
[0036] After starting the OCR process, the system will scan the images in the input file page by page to find and parse the text information. During this process, the OCR engine will apply a series of intelligent algorithms to analyze image features, including but not limited to edge detection, morphological operations, template matching, etc. Through these algorithms, OCR can locate every area that may contain text and further subdivide it into individual characters. Then, combined with glyph similarity, contextual relevance and other linguistic rules, the most likely character combination is determined to finally generate a complete text string.
[0037] In order to ensure the accuracy of OCR output, a post-processing step is usually introduced. This step involves cleaning and correcting the initially recognized text. Common practices include removing extra spaces, correcting punctuation errors, and standardizing uppercase and lowercase letters. For some parts that are difficult to identify or uncertain, machine learning models can be used for secondary judgment, or a manual review interface can be provided to allow security experts to intervene in the review. In addition, for structured documents such as contracts and reports, natural language processing (NLP) technology can be used to further parse the text content, extract key fields, and build a semantic understanding model to more deeply explore the potential value of information.
[0038] Finally, store the OCR-converted text in a format that is easy to access and process, such as plain text files, XML, JSON, etc. This not only facilitates the subsequent automated analysis process, but also facilitates manual review. Considering confidentiality and security requirements, all intermediate results and final outputs should be encrypted and stored, and access rights should be strictly controlled. At the same time, a detailed logging mechanism should be established to track the timestamp, source file identifier, operator identity, and other information of each OCR operation to provide a basis for auditing and tracing.
[0039] In the above method for tracing confidential file leakage using big data, the step S3 is to obtain the original confidential file. Specifically, in practical applications, the original confidential file is usually stored in the company's confidential database or secure storage medium, and a legal authorized access mechanism, such as identity authentication, authority verification, etc., is required to ensure that the original file is accurate and has not been tampered with.
[0040] In order to ensure the acquisition of accurate and unaltered original confidential documents, first of all, before entering the acquisition process, the access requirements must be clearly defined and whether they comply with relevant laws and regulations and internal company policies. For each person who requests access to original confidential documents, a strict background check must be conducted to confirm the authenticity of the requester's identity and the legitimacy of the purpose of access. Once confirmed, the requester will be directed to a specially designed identity verification link. In this link, multi-factor authentication (MFA) can be used to enhance security, that is, not just relying on a single password, but combining physical tokens, biometric identification (such as fingerprint scanning, iris recognition), one-time verification codes (generated by SMS or dedicated applications) and other methods to verify the user's identity. Doing so can greatly reduce the risk of identity fraud and ensure that only authorized persons can access sensitive information.
[0041] After completing the identity verification, the next step is the permission verification step. According to the principle of least privilege, each employee or system component should only have the minimum permissions necessary to complete the work. Therefore, detailed permission settings are defined in the access control list (ACL) or role-based access control (RBAC) system. When a requester attempts to access a specific confidential file, the system checks the role or personal permission configuration to which the requester belongs to determine whether he has the right to view the target file. If the permissions are insufficient, the access request is automatically denied; if the permissions match, it is allowed to proceed to the next step.
[0042] After obtaining access permission, the actual acquisition phase begins. To ensure the integrity and confidentiality of the file during the entire transmission process, all communications must be conducted through encrypted channels. For example, the SSL / TLS protocol is used to protect network connections to prevent man-in-the-middle attacks or other forms of data theft. At the same time, a hash algorithm is used to calculate the digital fingerprint of the file, and the result is compared with a pre-stored checksum to verify whether the file content has changed during the transmission process.
[0043] Real-time monitoring and audit tracking are essential during file access. By deploying intrusion detection systems (IDS), log management systems, and behavioral analysis tools, users' activity patterns can be continuously monitored to capture abnormal behaviors in a timely manner. Once suspicious operations are detected, alarms are triggered immediately and actions are taken according to pre-set emergency plans. All access events will be recorded in detail, including timestamps, IP addresses, specific commands executed, and other information for future audits and investigations.
[0044] Finally, when access to the original confidential file is no longer required, the current session should be terminated immediately, cached data should be cleared, and all open file handles should be closed. For copies that are temporarily downloaded or copied to the local environment, strict cleanup procedures must be followed to ensure that these copies do not remain in non-secure areas.
[0045] In the above method for tracing the leakage of confidential documents using big data, the step S4 extracts the file content from the original confidential document to obtain the original confidential document content. Specifically, considering that the original confidential document may also exist in multiple forms, the present application also uses OCR technology to recognize the text content of the original confidential document to convert the original confidential document in a non-plain text format into a processable text format to obtain the original confidential document content.
[0046] In the above method of tracing confidential file leakage using big data, the step S5 semantically encodes the suspected leaked file content and the original confidential file content to obtain a semantic encoding vector of the suspected leaked file content and a semantic encoding vector of the original confidential file content. It should be understood that in an information leakage incident, confidential files may be modified and spread, and the main methods include: the file name is changed to cover the original attributes; the file content is partially modified, such as inserting a watermark, deleting or adding some sensitive information; the file format is converted (such as from PDF to image or from Word document to plain text); the internal metadata information is deleted (such as modifying the timestamp, author information, etc.). These modification methods make it difficult to directly compare file attributes or content. However, even if the file has been modified or formatted to a certain extent, its core semantic information is relatively difficult to be completely eliminated or tampered with. Therefore, the present application further performs semantic encoding on the suspected leaked file content and the original confidential file content to extract the deep semantic features of the file content, so as to more effectively identify the correlation between the two by performing semantic level comparative analysis on the original confidential file and the suspected leaked file, and confirm whether the leaked file is part of the original confidential data, or whether it is a variant version after modification or repackaging based on the original file. In a specific example of the present application, the step S5 includes: using a semantic encoder based on the Transformer architecture to semantically encode the suspected leaked file content and the original confidential file content to obtain the semantic encoding vector of the suspected leaked file content and the semantic encoding vector of the original confidential file content. It should be known to those skilled in the art that the Transformer architecture is significantly superior to the traditional recurrent neural network (RNN) and its variants because it can process the entire input sequence in parallel, greatly improving the training speed and efficiency. In addition, the Transformer architecture also introduces position encoding to retain the order information of the input sequence, and further enhances the feature expression capability through a feedforward neural network, making the subsequent comparative analysis more accurate and reliable. Moreover, the Transformer architecture allows the model to consider all positions in the text simultaneously during the encoding process, thereby capturing long-distance dependencies and contextual information. Specifically, when encoding the contents of suspected leaked files and the contents of original confidential files, the Transformer decomposes each input sequence (such as a sentence or paragraph) into multiple tokens, and then uses a multi-head self-attention layer to calculate the mutual influence between these tokens to generate a feature representation that reflects the global semantic structure.
[0047] In the above-mentioned method of tracing the leakage of confidential files using big data, the step S6 performs a semantic fine-grained comparative analysis based on bidirectional attention on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain a bidirectional fine-grained comparative coding vector of the suspected leaked file-original confidential file content. That is, the present application further reveals the semantic similarities and differences between the two by performing a semantic matching calculation on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content. In particular, considering that traditional semantic matching methods usually use metrics such as cosine similarity to measure the similarity between two vectors, but this method often only focuses on the overall similarity between vectors, while ignoring the local correlation and mutual influence between the features of each dimension within the vector, and is easily interfered by redundant information, resulting in inaccurate matching results. In this regard, the present application proposes a semantic fine-grained comparative analysis method based on bidirectional attention, which captures the local fine-grained semantic association between the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content by performing bidirectional correlation analysis on the two, and constructs a positive and negative bidirectional attention field. Through the interaction of positive and negative bidirectional attention, the model is guided to focus on the subtle differences and similarities between the semantic features of the two during the comparison process, thereby achieving more accurate and in-depth semantic comparative analysis.
[0048] Figure 3 FIG. 6 is a flowchart of sub-step S6 of the method for tracing confidential document leakage using big data according to an embodiment of the present application. Figure 3 As shown, the step S6 includes the steps of: S61, performing homography projection transformation on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain the homography projection coding vector of the suspected leaked file content semantics and the homography projection coding vector of the original confidential file content semantics; S62, constructing a positive and negative bidirectional attention balance field between the homography projection coding vector of the suspected leaked file content semantics and the homography projection coding vector of the original confidential file content semantics to obtain the positive and negative bidirectional attention balance field of the suspected leaked file-original confidential file content semantics; S63, based on the positive and negative bidirectional attention balance field of the suspected leaked file-original confidential file content semantics, performing semantic mapping comparison analysis on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content.
[0049] Specifically, the step S61 is expressed by the formula:
[0050]
[0051]
[0052] in, represents the semantic encoding vector of the suspected leaked file content, represents the semantic encoding vector of the original confidential file content, and are different homography projection matrices, and They respectively represent the semantic homography projection coding vector of the suspected leaked file content and the semantic homography projection coding vector of the original confidential file content.
[0053] That is, by performing homography projection transformation on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content, the two are mapped into a common feature space, so that the feature distance and semantic association of the two in the feature space can be more intuitive and easy to compare, thereby enhancing the effectiveness and accuracy of semantic comparative analysis. In this way, it is ensured that content from different sources can be evaluated under a unified standard, improving the ability to identify subtle differences between the two, and providing a solid foundation for in-depth analysis.
[0054] Figure 4 FIG. 6 is a flowchart of sub-step S62 of the method for tracing confidential document leakage using big data according to an embodiment of the present application. Figure 4 As shown, the step S62 includes the steps of: S621, calculating the forward attention score field of the semantic homography projection coding vector of the suspected leaked file content relative to the semantic homography projection coding vector of the original confidential file content to obtain the forward attention score field of the suspected leaked file-original confidential file content semantics; S622, calculating the reverse attention score field of the semantic homography projection coding vector of the original confidential file content relative to the semantic homography projection coding vector of the suspected leaked file content to obtain the reverse attention score field of the suspected leaked file-original confidential file content semantics; S623, constructing the suspected leaked file-original confidential file content semantics positive and negative bidirectional attention balance field based on the suspected leaked file-original confidential file content semantics forward attention score field and the suspected leaked file-original confidential file content semantics reverse attention score field.
[0055] More specifically, the step S621 is expressed by the formula:
[0056]
[0057] in, represents the transpose of a vector, represents the matrix multiplication operation, is the attention score scaling factor, Represents the semantic positive attention score field of the suspected leaked file-original confidential file content.
[0058] That is, the forward attention score field is calculated for the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content after the homography projection transformation, so as to highlight the importance of each part of the suspected leaked file content when it interacts semantically with the original confidential file content. In this way, it can be clarified which semantic information is the most critical and relevant, so as to focus on these important semantic fragments in the subsequent semantic comparative analysis, which helps to more accurately identify the parts of the suspected leaked file that are closely related to the original confidential file content, thereby improving the effectiveness and pertinence of the analysis.
[0059] More specifically, the step S622 is expressed by the formula:
[0060]
[0061] in, Represents the semantic reverse attention score field between suspected leaked files and original confidential files.
[0062] That is, similar to the calculation of the forward attention score field, the reverse attention score field is calculated from the perspective of the original confidential file content, so as to reversely identify the semantic feature parts of the original confidential file content that are closely related to the suspected leaked file content. By calculating the reverse attention score, it is possible to highlight which content in the original confidential file is most semantically related to the suspected leaked file, thereby providing a more comprehensive perspective for the two-way comparative analysis and ensuring that no key information is missed. This method helps to more accurately locate and understand the semantic connection between the two files.
[0063] More specifically, in a specific example of the present application, the step S623 includes: inputting the suspected leaked file-original confidential file content semantics positive attention score field and the suspected leaked file-original confidential file content semantics reverse attention score field into the attention balance modulation module based on the convolutional neural network to obtain the suspected leaked file-original confidential file content semantics positive and negative bidirectional attention balance field, which is expressed by the formula:
[0064]
[0065] in, Indicates cascade, Indicates the suspected leaked file-original confidential file content semantic positive and negative bidirectional attention balance field, Representation 3 3 convolution operations.
[0066] That is, in order to ensure that the semantic association information between the suspected leaked file and the original confidential file can be fully captured in the final semantic comparison analysis process, this application further constructs a suspected leaked file-original confidential file content semantic positive and negative bidirectional attention balance field, and by integrating positive and negative attention information, ensures that the final semantic comparison will neither favor any party nor ignore the key details of any party. In this way, the semantic interaction between the contents of the two files can be better balanced, thereby providing a more comprehensive and accurate comparison result, ensuring that all important semantic associations can be fully considered.
[0067] Figure 5 FIG. 6 is a flowchart of sub-step S63 of the method for tracing confidential document leakage using big data according to an embodiment of the present application. Figure 5 As shown, the step S63 includes the steps of: S631, mapping the semantic homography projection coding vector of the suspected leaked file content and the semantic homography projection coding vector of the original confidential file content to the semantic positive and negative bidirectional attention balance field of the suspected leaked file-original confidential file content respectively to obtain the semantic homography projection attention modulation coding vector of the suspected leaked file content and the semantic homography projection attention modulation coding vector of the original confidential file content; S632, calculating the position point-by-position point division between the semantic homography projection attention modulation coding vector of the suspected leaked file content and the semantic homography projection attention modulation coding vector of the original confidential file content to obtain the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content.
[0068] More specifically, the step S631 is expressed as follows:
[0069]
[0070]
[0071] in, and They respectively represent the semantic homography projection attention modulation coding vector of the suspected leaked file content and the semantic homography projection attention modulation coding vector of the original confidential file content.
[0072] That is, the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content after homography projection transformation are projected into the suspected leaked file-original confidential file content semantic positive and negative bidirectional attention balance field respectively, so as to adjust the feature representation of the two according to the attention weight distribution of the positive and negative bidirectional attention balance field, so that the features with high correlation are enhanced in the final feature representation, while the features with low correlation are weakened, thereby obtaining the homography projection attention modulation coding vector of the semantics of the suspected leaked file content and the homography projection attention modulation coding vector of the semantics of the original confidential file content. It should be understood that in the above-mentioned attention balance field projection process, the semantic expression of one party is essentially adjusted based on the semantic information of the other party, so that the generated homography projection attention modulation coding vector of the semantics of the suspected leaked file content and the homography projection attention modulation coding vector of the semantics of the original confidential file content can better reflect the semantic relevance and difference between the suspected leaked file and the original confidential file.
[0073] More specifically, the step S632 is expressed by the formula:
[0074]
[0075] in, A bidirectional fine-grained comparison encoding vector representing the suspected leaked file and the original confidential file content.
[0076] Here, by performing position-by-position point division operations, the semantic feature comparison and interactive coding of the semantic homography projection attention modulation coding vector of the suspected leaked file content and the semantic homography projection attention modulation coding vector of the original confidential file content are performed element by element to form a suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector. In this way, the similarities and differences between the suspected leaked file and the original confidential file in local fine-grained semantic features can be effectively quantified, thereby providing a more accurate and comprehensive basis for subsequent content verification and leakage identification.
[0077] In the above-mentioned method for tracing the leakage of confidential files using big data, the step S7 determines whether the suspected leaked file is a variant of the original confidential file based on the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content. In a specific example of the present application, the step S7 includes: inputting the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content into a content verification module based on a classifier to obtain a content verification result, and the content verification result is whether the suspected leaked file is a variant of the original confidential file. Here, the classifier is used to analyze the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content, identify the semantic similarities and differences between the two, and map such semantic similarities and differences to preset category labels, so as to determine whether the suspected leaked file is a variant of the original confidential file. Specifically, the classifier acquires the ability to understand specific types of data through training, and can detect content features even after format conversion, partial deletion, addition of interference information or slight modification. When the bidirectional fine-grained comparison encoding vector of the suspected leaked file-original confidential file content is input into the classifier, it uses the learned pattern recognition rules to evaluate the degree of consistency between the suspected leaked file and the original confidential file content at the semantic level, and outputs a content verification result based on this, which clearly indicates whether the suspected leaked file is a variant of the original confidential file.
[0078] In a specific example of the present application, the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content is input into a classifier-based content verification module to obtain a content verification result, including: using the fully connected layer of the content verification module to fully connect the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content to obtain a bidirectional fine-grained comparison fully connected coding vector of the suspected leaked file-original confidential file content; inputting the bidirectional fine-grained comparison fully connected coding vector of the suspected leaked file-original confidential file content into the Softmax classification function of the content verification module to obtain the probability value of the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content belonging to each classification label, wherein the classification labels include the suspected leaked file is a variant of the original confidential file and the suspected leaked file is not a variant of the original confidential file; the classification label corresponding to the largest of the probability values is determined as the content verification result.
[0079] In a preferred example of the present application, considering that the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content respectively represent the text semantic coding features of the suspected leaked file content and the text semantic coding features of the original confidential file content, when performing semantic fine-grained comparative analysis, the complexity of the positive and negative attention semantic interaction field caused by the semantic differences of the text source will lead to insufficient representation of the long-distance fine-grained semantic differences of the suspected leaked file-original confidential file content bidirectional fine-grained comparative coding vector, thereby reducing the expression effect of the suspected leaked file-original confidential file content bidirectional fine-grained comparative coding vector, and affecting the accuracy of the content verification result obtained by its input into the classifier-based content verification module.
[0080] Therefore, in one example, when the suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector is input into the classifier-based content verification module, the suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector is optimized, and the optimization includes the following steps:
[0081] First, the eigenvalues of the suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector are arranged in ascending order to form a suspected leaked file-original confidential file content bidirectional fine-grained comparison sequence coding vector;
[0082] Secondly, in response to the first of the two-way fine-grained comparison sequence encoding vectors of the suspected leaked file and the original confidential file content, The eigenvalue and The absolute value of the difference between the eigenvalues is less than or equal to the distance difference hyperparameter , calculate the The eigenvalues are related to the The weighted sum of the eigenvalues is the optimized The characteristic value is expressed as:
[0083]
[0084] in, and They represent the first order encoding vector of the bidirectional fine-grained comparison of the suspected leaked file and the original confidential file content. The eigenvalue and eigenvalues, and denote the first weight parameter and the second weight parameter respectively, The first vector representing the optimized bidirectional fine-grained comparison sequence encoding vector of the suspected leaked file and the original confidential file content eigenvalues;
[0085] Then, in response to the first of the two-way fine-grained comparison sequence encoding vectors of the suspected leaked file and the original confidential file content, The eigenvalue and The absolute value of the difference between the eigenvalues is greater than the distance difference hyperparameter , and calculate the square root of the sum of squares of all eigenvalues of the bidirectional fine-grained comparison encoding vector of the suspected leaked file-original confidential file content, which can be expressed as:
[0086]
[0087] in, The square root of the sum of squares of all eigenvalues of the bidirectional fine-grained comparison encoding vector of the suspected leaked file and the original confidential file content, , ... represents each eigenvalue of the bidirectional fine-grained comparison encoding vector of the content of the suspected leaked file and the original confidential file;
[0088] Next, the square root Multiply by 2 and then divide by the square of the length of the suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector to obtain the suspected leaked file-original confidential file content bidirectional fine-grained comparison space primitive value, which is expressed as:
[0089]
[0090] in, Indicates the length of the encoding vector for the bidirectional fine-grained comparison of the suspected leaked file and the original confidential file content. Indicates the spatial primitive value of the bidirectional fine-grained comparison between the suspected leaked file and the original confidential file content;
[0091] Then, the bidirectional fine-grained comparison space primitive value of the suspected leaked file-original confidential file content is multiplied by the first After the eigenvalues, calculate the product and the The weighted reduction between the eigenvalues is the optimized The characteristic value is expressed as:
[0092]
[0093] in, and denote the third weight parameter and the fourth weight parameter respectively, The first vector representing the optimized bidirectional fine-grained comparison sequence encoding vector of the suspected leaked file and the original confidential file content eigenvalues;
[0094] Finally, based on , the first Eigenvalues This is to obtain an optimized bidirectional fine-grained comparison encoding vector between the suspected leaked file and the original confidential file content.
[0095] In this way, in order to address the problem of insufficient representation of fine-grained semantic differences in the feature set of the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content under a predetermined feature value sequential distribution due to a long distance exceeding a predetermined local distribution interval threshold, a high-dimensional feature space primitive representation based on self-inner product fusion of the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content is used to capture the complex structure of the global network interaction of its feature values, thereby reconstructing the dynamic semantic search relationship between the feature values of the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content by simulating the scale-based high-dimensional feature space potential primitives, so as to achieve the coding reconstruction of the real sequence distribution behavior of the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content under a long distance, improve the coding expression effect of the bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content, and improve the accuracy of the content verification result obtained by inputting it into a content verification module based on a classifier.
[0096] In summary, the method of tracing the leakage of confidential documents using big data based on the embodiment of the present application is explained, which uses file recognition technology to extract file contents from suspected leaked files and original confidential files respectively, and introduces natural language processing technology based on deep learning to perform contextual semantic analysis on the suspected leaked file contents and the original confidential file contents, so as to extract the semantic features of the suspected leaked file contents and the original confidential file contents, and then, by performing fine-grained semantic comparative analysis of the two based on a bidirectional attention mechanism, the deep-level semantic similarities and differences between the two are revealed to determine whether the suspected leaked file is a variant of the original confidential file. In this way, by performing fine-grained comparative analysis of the suspected leaked file and the original confidential file at the contextual semantic level, the association between the leaked file after various processing and the original confidential file can be more accurately captured, thereby improving the accuracy and reliability of tracing the leakage of confidential files.
[0097] The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.
[0098] In the above embodiments, the description of each embodiment has its own emphasis. For the parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0099] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference to a figure in a claim should not be considered as limiting the claim to which it relates.
[0100] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0101] Finally, it should be noted that the above description has been given for the purpose of illustration and description. In addition, the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for tracing the leakage of confidential documents using big data, characterized in that: include: Obtaining suspected leaked documents; Extracting file content from the suspected leaked file to obtain the suspected leaked file content; Access to original confidential documents; Extracting file content from the original confidential file to obtain original confidential file content; Performing semantic coding on the suspected leaked file content and the original confidential file content to obtain a suspected leaked file content semantic coding vector and an original confidential file content semantic coding vector; A semantic fine-grained comparison analysis based on bidirectional attention is performed on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain a bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content, wherein the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content are subjected to homography projection transformation to map the two to a common feature space, and a semantic positive and negative bidirectional attention balance field of the suspected leaked file-original confidential file content is further constructed, and the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content after homography projection transformation are respectively projected into the semantic positive and negative bidirectional attention balance field of the suspected leaked file-original confidential file content to obtain a homography projection attention modulation coding vector of the suspected leaked file content semantics and a homography projection attention modulation coding vector of the original confidential file content semantics, and an element-by-element semantic feature comparison interactive coding is performed on the homography projection attention modulation coding vector of the suspected leaked file content semantics and the homography projection attention modulation coding vector of the original confidential file content semantics through position-by-position point division operation to form a bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content; Based on the bidirectional fine-grained comparison encoding vector of the suspected leaked file and the original confidential file content, determining whether the suspected leaked file is a variant of the original confidential file; When the bidirectional fine-grained comparison coding vector of the suspected leaked file and the original confidential file content is input into the classifier-based content verification module, the bidirectional fine-grained comparison coding vector of the suspected leaked file and the original confidential file content is optimized, and the optimization includes the following steps: Arrange the eigenvalues of the suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector in ascending order to form a suspected leaked file-original confidential file content bidirectional fine-grained comparison sequence coding vector; In response to the first of the two-way fine-grained comparison sequence encoding vectors of the suspected leaked file and the original confidential file content, The eigenvalue and The absolute value of the difference between the eigenvalues is less than or equal to the distance difference hyperparameter , calculate the The eigenvalues are related to the The weighted sum of the eigenvalues is the optimized eigenvalues; In response to the first of the two-way fine-grained comparison sequence encoding vectors of the suspected leaked file and the original confidential file content, The eigenvalue and The absolute value of the difference between the eigenvalues is greater than the distance difference hyperparameter , and calculating the square root of the sum of squares of all eigenvalues of the bidirectional fine-grained comparison encoding vector of the suspected leaked file-original confidential file content; Multiply the square root by 2 and then divide it by the square of the length of the suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector to obtain a suspected leaked file-original confidential file content bidirectional fine-grained comparison space primitive value; Multiply the bidirectional fine-grained comparison space primitive value of the suspected leaked file-original confidential file content by the first After the eigenvalues, calculate the product and the The weighted reduction between the eigenvalues is the optimized eigenvalues; Combinatorial Optimization feature values to obtain an optimized bidirectional fine-grained comparison encoding vector of the suspected leaked file and the original confidential file content.
2. The method for tracing confidential document leakage using big data according to claim 1, characterized in that: Extracting the file content from the suspected leaked file to obtain the suspected leaked file content includes: The OCR technology is used to identify the text content of the suspected leaked file to obtain the content of the suspected leaked file.
3. The method for tracing confidential document leakage using big data according to claim 2, characterized in that: The suspected leaked file content and the original confidential file content are semantically encoded to obtain a suspected leaked file content semantic encoding vector and an original confidential file content semantic encoding vector, including: The suspected leaked file content and the original confidential file content are semantically encoded using a semantic encoder based on a Transformer architecture to obtain a semantic encoding vector of the suspected leaked file content and a semantic encoding vector of the original confidential file content.
4. The method for tracing confidential document leakage using big data according to claim 3, characterized in that: Performing a semantic fine-grained comparative analysis based on bidirectional attention on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain a bidirectional fine-grained comparative coding vector of the suspected leaked file-original confidential file content, including: Performing homography projection transformation on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain a semantic homography projection coding vector of the suspected leaked file content and a semantic homography projection coding vector of the original confidential file content; Constructing a positive and negative bidirectional attention balance field between the semantic homography projection coding vector of the suspected leaked file content and the semantic homography projection coding vector of the original confidential file content to obtain a positive and negative bidirectional attention balance field of the semantic homography projection coding vector of the suspected leaked file content-original confidential file content; Based on the positive and negative bidirectional attention balance field of the semantics of the suspected leaked file-original confidential file content, a semantic mapping comparison analysis is performed on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain a bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content.
5. The method for tracing confidential document leakage using big data according to claim 4, characterized in that: Constructing a positive and negative bidirectional attention balance field between the semantic homography projection coding vector of the suspected leaked file content and the semantic homography projection coding vector of the original confidential file content to obtain a positive and negative bidirectional attention balance field of the semantic homography projection coding vector of the suspected leaked file content-original confidential file content, including: Calculating the forward attention score field of the semantic homography projection coding vector of the suspected leaked file content relative to the semantic homography projection coding vector of the original confidential file content to obtain the semantic forward attention score field of the suspected leaked file-original confidential file content; Calculating the reverse attention score field of the semantic homography projection coding vector of the original confidential file content relative to the semantic homography projection coding vector of the suspected leaked file content to obtain the suspected leaked file-original confidential file content semantic reverse attention score field; Based on the suspected leaked file-original confidential file content semantic positive attention score field and the suspected leaked file-original confidential file content semantic reverse attention score field, the suspected leaked file-original confidential file content semantic positive and negative bidirectional attention balance field is constructed.
6. The method for tracing confidential document leakage using big data according to claim 5, characterized in that: Based on the suspected leaked file-original confidential file content semantic positive attention score field and the suspected leaked file-original confidential file content semantic reverse attention score field, construct the suspected leaked file-original confidential file content semantic positive and negative bidirectional attention balance field, including: The semantic positive attention score field of the suspected leaked file-original confidential file content and the semantic negative attention score field of the suspected leaked file-original confidential file content are input into the attention balance modulation module based on the convolutional neural network to obtain the semantic positive and negative bidirectional attention balance field of the suspected leaked file-original confidential file content.
7. The method for tracing confidential document leakage using big data according to claim 6, characterized in that: Based on the semantic positive and negative bidirectional attention balance field of the suspected leaked file-original confidential file content, a semantic mapping comparison analysis is performed on the semantic coding vector of the suspected leaked file content and the semantic coding vector of the original confidential file content to obtain a bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content, including: Mapping the suspected leaked file content semantic homography projection coding vector and the original confidential file content semantic homography projection coding vector to the suspected leaked file-original confidential file content semantic positive and negative bidirectional attention balance field to obtain the suspected leaked file content semantic homography projection attention modulation coding vector and the original confidential file content semantic homography projection attention modulation coding vector; The suspected leaked file-original confidential file content bidirectional fine-grained comparison coding vector of the suspected leaked file-original confidential file content is obtained by calculating the position point-by-position difference between the semantic homography projection attention modulation coding vector of the suspected leaked file content and the semantic homography projection attention modulation coding vector of the original confidential file content.
8. The method for tracing confidential document leakage using big data according to claim 7, characterized in that: Based on the bidirectional fine-grained comparison encoding vector of the suspected leaked file and the original confidential file content, determining whether the suspected leaked file is a variant of the original confidential file includes: The bidirectional fine-grained comparison coding vector of the suspected leaked file and the original confidential file content is input into a classifier-based content verification module to obtain a content verification result, wherein the content verification result indicates whether the suspected leaked file is a variant of the original confidential file.
9. The method for tracing confidential document leakage using big data according to claim 8, characterized in that: Inputting the bidirectional fine-grained comparison encoding vector of the suspected leaked file and the original confidential file content into the classifier-based content verification module to obtain the content verification result, including: Use the fully connected layer of the content verification module to perform fully connected encoding on the suspected leaked file-original confidential file content bidirectional fine-grained comparison encoding vector to obtain a suspected leaked file-original confidential file content bidirectional fine-grained comparison fully connected encoding vector; Input the suspected leaked file-original confidential file content bidirectional fine-grained comparison fully connected encoding vector into the Softmax classification function of the content verification module to obtain the probability value of the suspected leaked file-original confidential file content bidirectional fine-grained comparison encoding vector belonging to each classification label, wherein the classification label includes the suspected leaked file is a variant of the original confidential file and the suspected leaked file is not a variant of the original confidential file; The classification label corresponding to the largest probability value among the probability values is determined as the content verification result.
Citation Information
Patent Citations
Shared data resource anomaly detection method, system, device, memory and product
CN116248412A
Method for preventing leakage of software backup data
CN119720279A