An implementation method for file confidentiality detection based on Tika

By building a logical binomial tree and a deep learning network using word embedding models based on Tika, the problem of parsing ability and sensitive information recognition in the existing technology when processing code files is solved, and effective identification and detection of sensitive information in the code is realized.

CN119577820BActive Publication Date: 2025-06-17YUHENG POWER STATION OF SHAANXI HUADIAN YUHENG COAL POWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411430159.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-06-17
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

The existing Tika-based file confidentiality detection methods have problems such as limited parsing capabilities, strong hidden sensitive information, multilingual and format complexity, and poor model applicability when processing code files, making it difficult to effectively identify sensitive information in the code.

Method used

Using the Tika-based file confidentiality detection implementation method, by obtaining sensitive information lexicon, detecting and locking sensitive vocabulary in the target text or code, building a logical binomial tree, using the word embedding model to calculate word embedding vectors, and detecting through a deep learning network to realize sensitive information recognition in the code.

Benefits of technology

This method can not only train natural language texts, but also effectively understand and analyze code language and identify sensitive information in the code, with good applicability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577820B_ABST
    Figure CN119577820B_ABST
Patent Text Reader

Abstract

A method for implementing file confidentiality detection based on Tika, the method comprising: obtaining a sensitive information thesaurus, and detecting and locking sensitive words in a target text or target code based on the sensitive information thesaurus. According to the context information of the sensitive words, a logical binary tree corresponding to the sensitive words is constructed through semantic analysis. The word embedding vectors corresponding to each node in the logical binary tree are calculated through a word embedding model, the similarity between any two nodes is calculated based on the word embedding vectors, and different nodes are classified. The word embedding vectors are used as the input of a deep learning network model, and the deep learning network is called to detect the word embedding vectors to obtain the confidentiality detection result of the target text or target code. The relationship between the parent node and the child node in the logical binary tree is used to represent the causal relationship of the context information of the sensitive words, and the relationship between the sibling nodes is used to represent the logical nesting relationship of the context information of the sensitive words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of text data security, and more specifically, relates to a method for implementing file confidentiality detection based on Tika. Background Art

[0002] With the acceleration of the digitalization process, data has become one of the most important assets of enterprises and organizations. However, data security risks are also increasing continuously, especially through the increasingly diverse channels of supply chain and industrial chain penetration. The state has successively introduced relevant laws and regulations on data security, clearly requiring strengthening data security protection to prevent the leakage of sensitive information. In this context, the confidentiality detection of files is particularly important. As an information carrier, files widely exist in various business activities. Ensuring the security of confidential information in files is not only a necessary measure to comply with laws and regulations, but also an important means to protect business secrets, maintain customer privacy, and prevent internal information leakage. Therefore, developing efficient and accurate file confidentiality detection methods has become a key requirement in the field of data security.

[0003] Apache Tika is an open-source framework developed by the Apache Software Foundation, aiming to detect and extract metadata and content of various document types. It provides a unified interface and supports the parsing of multiple file formats, including text, image, audio, video, documents, and compressed files, etc. Tika has extensive applications in content extraction and metadata parsing, especially performing well when processing common document formats such as Word and PDF.

[0004] Currently, the file detection method based on Tika mainly includes the following steps: Content extraction: Using Tika's parser to extract the text content and metadata information of the file; Content analysis: Performing keyword matching, regular expression detection on the extracted content, or combining technologies such as machine learning and deep learning (such as natural language processing) to identify sensitive information; Result output: Generating a detection report or warning information according to the analysis result. This method can effectively detect obvious sensitive information when processing ordinary text files and documents, and has a certain practicality.

[0005] However, the existing Tika-based file confidentiality detection methods have obvious deficiencies when dealing with code files (such as.ipynb,.py,.c,.cpp, etc.): Limited parsing ability: Tika's parsing of code files mainly stays at the plain text level and cannot deeply understand the syntax structure and logical relationships of the code. Code files have special programming language syntax, including complex structures such as variables, functions, classes, and comments. It is difficult to obtain meaningful content by directly extracting text. Strong concealment of sensitive information: Sensitive information in the code (such as hard-coded passwords, keys, API keys, personal identity information, etc.) is usually hidden in variable assignments, configuration files, or comments, and may be encoded or obfuscated. Traditional keyword matching and regular expression methods are difficult to cover comprehensively and are prone to false negatives or false positives. Multilingual and format complexity: Code files involve multiple programming languages, and the syntax and structures of different languages vary greatly. For example, Python code (such as.py or.ipynb files) and C / C++ code (such as.c,.cpp files) have significant differences in syntax and file structure. Tika lacks professional parsers for different programming languages and cannot effectively handle diverse code files. Poor applicability of existing models: Existing machine learning and deep learning models are mostly trained for natural language texts and have insufficient understanding and analysis capabilities for code languages, making it difficult to effectively identify sensitive information in the code. Therefore, the existing technologies have significant defects in the confidentiality detection of code files and cannot meet the actual application requirements. Summary of the Invention

[0006] To solve the deficiencies in the existing technology, the purpose of the present invention is to address the above-mentioned defects and further propose a method for implementing file confidentiality detection based on Tika.

[0007] The present invention adopts the following technical solutions.

[0008] The first aspect of the present invention discloses a method for implementing file confidentiality detection based on Tika, and the method includes:

[0009] Obtain a sensitive information vocabulary library, and detect and lock sensitive words in the target text or target code based on the sensitive information vocabulary library;

[0010] According to the context information of the sensitive words, construct a logical binary tree corresponding to the sensitive words through semantic analysis;

[0011] Calculate the word embedding vectors corresponding to each node in the logical binary tree through a word embedding model, and calculate the similarity between any two nodes based on the word embedding vectors, where the similarity is used to classify different nodes;

[0012] Use the word embedding vector as the input of a deep learning network model to call the deep learning network to detect the word embedding vector, and obtain the confidentiality detection result of the target text or target code;

[0013] Among them, the relationship between the parent node and the child node in the logical binary tree is used to represent the causal relationship of the context information of the sensitive vocabulary, and the relationship between sibling nodes is used to represent the logical nesting relationship of the context information of the sensitive vocabulary.

[0014] Further, the obtaining of the sensitive information vocabulary library and the detection and locking of sensitive vocabulary in the target text or target code based on the sensitive information vocabulary library include:

[0015] Parse all sub-files in the file directory to be confidentially detected through a Parser, and convert the parsing result into a string. The parsing result includes the sensitive vocabulary in the file to be confidentially detected;

[0016] Extract the sensitive vocabulary from the parsing result based on the pre-input sensitive information vocabulary library, and lock the sensitive vocabulary.

[0017] Further, the calculating of the word embedding vector corresponding to each node in the logical binary tree through a word embedding model, and the calculating of the similarity between any two nodes based on the word embedding vector, where the similarity is used to classify different nodes, includes:

[0018] Obtain the word embedding vector corresponding to each node in the logical binary tree through a Word2Vec word embedding model, and obtain the adjacency matrix V k associated with node v k ;

[0019] Among them, the expression of the adjacency matrix is:

[0020]

[0021] In the formula, if i = j, v ij represents the i-th child node of the vertex in the logical binary tree. If i ≠ j, v ij represents the distance between the j-th leaf node and node v k in the sub-binary tree with the i-th child node as the root node, i, j = 1, 2,..., N, and N represents the depth of the logical binary tree.

[0022] Further, after the calculating of the word embedding vector corresponding to each node in the logical binary tree through a word embedding model, and the calculating of the similarity between any two nodes based on the word embedding vector, where the similarity is used to classify different nodes, it also includes:

[0023] Data types obtained by classifying different nodes according to similarity, and based on the values corresponding to the nodes of the corresponding data type collected at any moment, combined with a preset weight learning matrix, calculate the mapping vector of each node in the logical binary tree; and

[0024] Based on the set of nodes of the same data type in the logical binary tree and the adjacency matrix associated with the nodes in the node set, calculate the neighborhood feature vector of any node in the logical binary tree with respect to any data type.

[0025] Further, calculating the word embedding vector corresponding to each node in the logical binary tree through a word embedding model, and calculating the similarity between any two nodes based on the word embedding vector, where the similarity is used to classify different nodes, and then further includes:

[0026] Call an activation function to calculate the attention scores of all data types according to the preset attention vector of any data type;

[0027] Calculate the influence factor between any two nodes according to the attention scores of all data types in the logical binary tree.

[0028] Further, calculating the word embedding vector corresponding to each node in the logical binary tree through a word embedding model, and calculating the similarity between any two nodes based on the word embedding vector, where the similarity is used to classify different nodes, and then further includes:

[0029] Based on a preset confidence threshold, merge the child nodes of the same data type, and call an activation function to calculate the feature vector of each data type according to the influence factor between any two nodes of the same data type; and

[0030] Establish a logical binary tree network model according to the heap embedding matrix of the logical binary tree, and perform an initialization process on the weight matrix in the logical binary tree network model, and convert the word embedding vector into a query vector, a key vector, and a value vector through a self-attention layer.

[0031] Further, using the word embedding vector as the input of a deep learning network model to call the deep learning network to detect the word embedding vector, and obtaining the confidentiality detection result of the target text or target code, including:

[0032] A prediction model based on a cross-entropy loss function, and update the weights and biases of the logical binary tree network model according to a feedforward network;

[0033] Predict the current file to be detected according to the trained logical binomial tree network model, and obtain the confidentiality detection result of the current file to be detected based on the word embedding vector of the current file to be detected;

[0034] Among them, the feedforward network is represented as two-layer linear transformation, namely the first-layer linear transformation and the second-layer linear transformation, and the feedforward network is obtained by processing the weights and biases of the first-layer linear transformation and the second-layer linear transformation by calling an activation function.

[0035] The second aspect of the present invention discloses a device for realizing file confidentiality detection based on Tika, and the device includes:

[0036] A sensitive word locking module, configured to obtain a sensitive information word library, and detect and lock sensitive words in the target text or target code based on the sensitive information word library;

[0037] A logical binomial tree construction module, configured to construct a logical binomial tree corresponding to the sensitive word through semantic analysis according to the context information of the sensitive word;

[0038] A node classification module, configured to calculate the word embedding vector corresponding to each node in the logical binomial tree through a word embedding model, and calculate the similarity between any two nodes based on the word embedding vector, and the similarity is used to classify different nodes;

[0039] A confidentiality detection module, configured to use the word embedding vector as the input of a deep learning network model to call the deep learning network to detect the word embedding vector, and obtain the confidentiality detection result of the target text or target code;

[0040] Among them, the relationship between the parent node and the child node in the logical binomial tree is used to represent the causal relationship of the context information of the sensitive word, and the relationship between the sibling nodes is used to represent the logical nesting relationship of the context information of the sensitive word.

[0041] The third aspect of the present invention discloses a terminal, including a processor and a storage medium; characterized in that:

[0042] The storage medium is used to store instructions;

[0043] The processor is configured to operate according to the instructions to execute the steps of the method described in the first aspect.

[0044] The fourth aspect of the present invention discloses a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the method described in the first aspect are realized.

[0045] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention has the following advantages:

[0046] By obtaining a sensitive information thesaurus and detecting and locking sensitive words in the target text or target code based on the sensitive information thesaurus. According to the context information of the sensitive words, a logical binary tree corresponding to the sensitive words is constructed through semantic analysis. The word embedding vectors corresponding to each node in the logical binary tree are calculated through a word embedding model, and the similarity between any two nodes is calculated based on the word embedding vectors to classify different nodes. The word embedding vectors are used as the input of a deep learning network model to call the deep learning network to detect the word embedding vectors, and the confidentiality detection result of the target text or target code is obtained. This method can not only be trained for natural language texts, but also understand and analyze code languages, and thus effectively identify sensitive information in the code, with good applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a schematic flowchart of a method for implementing file confidentiality detection based on Tika provided by the present invention;

[0048] Figure 2A is one of the schematic diagrams of the code confidentiality detection logical binary tree architecture of a method for implementing file confidentiality detection based on Tika provided by the present invention;

[0049] Figure 2B is another schematic diagram of the code confidentiality detection logical binary tree architecture of a method for implementing file confidentiality detection based on Tika provided by the present invention;

[0050] Figure 2C is yet another schematic diagram of the code confidentiality detection logical binary tree architecture of a method for implementing file confidentiality detection based on Tika provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following further describes the present application with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present application.

[0052] As Figure 1 shown, in one embodiment, a method for implementing file confidentiality detection based on Tika includes the following steps:

[0053] Step S110, obtain a sensitive information thesaurus and detect and lock sensitive words in the target text or target code based on the sensitive information thesaurus.

[0054] In some embodiments, a method for implementing file confidentiality detection based on Tika provided by the present invention obtains a sensitive information thesaurus, and detects and locks sensitive words in a target text or target code based on the sensitive information thesaurus, specifically including the following steps:

[0055] Step S111, parse all sub-files in the file directory to be detected for confidentiality through a Parser, and convert the parsing result into a string. The parsing result includes sensitive words in the file to be detected for confidentiality.

[0056] Step S112, extract sensitive words from the parsing result based on the pre-input sensitive information thesaurus, and lock the sensitive words.

[0057] Step S120, construct a logical binary tree corresponding to the sensitive word through semantic analysis according to the context information of the sensitive word.

[0058] Among them, the relationship between the parent node and the child node in the logical binary tree is used to represent the causal relationship of the context information of the sensitive word, and the relationship between the sibling nodes is used to represent the logical nesting relationship of the context information of the sensitive word.

[0059] Step S130, calculate the word embedding vector corresponding to each node in the logical binary tree through a word embedding model, and calculate the similarity between any two nodes based on the word embedding vector. The similarity is used to classify different nodes.

[0060] In some embodiments, a method for implementing file confidentiality detection based on Tika provided by the present invention calculates the word embedding vector corresponding to each node in the logical binary tree through a word embedding model, and calculates the similarity between any two nodes based on the word embedding vector. The similarity is used to classify different nodes, specifically including the following steps:

[0061] Step S131, obtain the word embedding vector corresponding to each node in the logical binary tree through the Word2Vec word embedding model, and obtain the adjacency matrix V k associated with the node v in the logical binary tree based on the word embedding vector k .

[0062] Among them, the expression of the adjacency matrix is:

[0063]

[0064] In the formula, if i = j, v ij represents the i-th child node of the vertex in the logical binary tree. If i ≠ j, v ij represents the j-th leaf node to the node v in the sub-binary tree with the i-th child node as the root node. kThe distance between them, i, j = 1, 2, …, N, where N represents the depth of the logical binomial tree.

[0065] In some embodiments, a method for implementing file confidentiality detection based on Tika provided by the present invention calculates the word embedding vectors corresponding to each node in the logical binomial tree through a word embedding model, and calculates the similarity between any two nodes based on the word embedding vectors. The similarity is used to classify different nodes. After that, the following steps are further included:

[0066] Step S210: According to the data types obtained by classifying different nodes according to the similarity, and based on the values corresponding to the nodes of the corresponding data types collected at any moment, combined with a preset weight learning matrix, calculate the mapping vectors of each node in the logical binomial tree.

[0067] Step S220: Based on the set of nodes of the same data type in the logical binomial tree and the adjacency matrix associated with the nodes in the set of nodes, calculate the neighborhood feature vectors of any node in the logical binomial tree with respect to any data type.

[0068] In some embodiments, a method for implementing file confidentiality detection based on Tika provided by the present invention calculates the word embedding vectors corresponding to each node in the logical binomial tree through a word embedding model, and calculates the similarity between any two nodes based on the word embedding vectors. The similarity is used to classify different nodes. After that, the following steps are further included:

[0069] Step S230: Call an activation function to calculate the attention scores of all data types according to the preset attention vectors of any data type.

[0070] Step S240: Calculate the influence factors between any two nodes according to the attention scores of all data types in the logical binomial tree.

[0071] In some embodiments, a method for implementing file confidentiality detection based on Tika provided by the present invention calculates the word embedding vectors corresponding to each node in the logical binomial tree through a word embedding model, and calculates the similarity between any two nodes based on the word embedding vectors. The similarity is used to classify different nodes. After that, the following steps are further included:

[0072] Step S250: Based on a preset confidence threshold, merge the child nodes of the same data type, and call an activation function to calculate the feature vectors of each data type according to the influence factors between any two nodes of the same data type.

[0073] Step S260, establish a logical binomial tree network model according to the heap embedding matrix of the logical binomial tree, initialize the weight matrix in the logical binomial tree network model, and convert the word embedding vector into a query vector, a key vector and a value vector through a self-attention layer.

[0074] Step S140, using the word embedding vector as the input of the deep learning network model to call the deep learning network to detect the word embedding vector and obtain the confidentiality detection result of the target text or target code.

[0075] In some embodiments, the present invention provides a method for implementing file confidentiality detection based on Tika, which uses a word embedding vector as the input of a deep learning network model to call the deep learning network to detect the word embedding vector and obtain a confidentiality detection result of the target text or target code, specifically comprising the following steps:

[0076] Step S141, based on the prediction model of the cross entropy loss function, and updating the weights and biases of the logistic binomial tree network model according to the feedforward network.

[0077] Step S142, predicting the current file to be detected according to the trained logical binomial tree network model, and obtaining the confidentiality detection result of the current file to be detected based on the word embedding vector of the current file to be detected.

[0078] Among them, the feedforward network is represented by two layers of linear transformation, namely the first layer linear transformation and the second layer linear transformation. The feedforward network is obtained by calling the activation function to process the weights and biases of the first layer linear transformation and the second layer linear transformation.

[0079] In a specific embodiment, the present invention provides a Tika-based file confidentiality detection implementation method. In order to specifically describe the technical defects mentioned in the background, a brief description is first given by combining code segment 1 and code segment 2 as an example.

[0080]

[0081]

[0082]

[0083] In this embodiment, code segment 1 is written in Java, and there are comments and variable names containing sensitive words such as "password" and "key", and the system detects hard-coded passwords. However, in fact, in file 11, the above API_KEY is a sample encrypted data used in the development environment. That is, in an actual engineering environment, real environment variables or configuration files will be used to load real encrypted data, as shown in file 12.

[0084] Code segment 2 is written in Python, which shows the key code for secondary development by calling the current artificial intelligence model through an external application. It is not difficult to see that the meaning expressed by its code is highly similar to that of code segment 1. However, although its name also involves obviously confidential content such as API_KEY, in essence, it is obtained through the key.env file. Therefore, whether there is a security risk is actually difficult to determine only from file 21. Only by judging whether file 22 is encrypted can it be determined whether it is confidential.

[0085] In this embodiment, both code segment 1 and code segment 2 are very likely to be misjudged as containing sensitive information, resulting in a detection result with potential security risks. In addition, the corresponding code segment 3 for code segment 1 and 2 has a risk of confidentiality.

[0086]

[0087] Similar to the above analysis, code segment 3 also contains sensitive information about the key. Although Tika will most likely still determine that code segment 3 has a risk of confidentiality, the logical process of its derivation that code segment 3 has a security risk is incorrect.

[0088] It should be noted that for the sake of convenience, only part of the above code is given. For example, taking code segment 2 as an illustration, it removes the irrelevant header files such as import os and import load_dotenv.

[0089] In this embodiment, first, the structural framework of Tika needs to be explained. Tika mainly includes:

[0090] A monitor for detecting the type of document and then selecting an appropriate parser.

[0091] A parser for parsing specific types of documents and extracting content and metadata.

[0092] Metadata, an object for storing and transmitting metadata information.

[0093] A content processor for processing the content output by the parser, which can write the content to a string, a file, or other output targets.

[0094] Among them, the monitor and metadata are usually based on fixed file formats and standard detection data. For the parser, in this embodiment, it mainly involves the following 2 interfaces.

[0095]

[0096]

[0097] Among them, the Parser interface is the core interface of Tika, which defines the methods that all parsers must implement. For example, specific file format parsers such as PDFParser and TXTParser need to inherit this interface. Its function is to parse the content in the input stream, extract text and metadata.

[0098] In this embodiment, the ParseContext class is used to pass additional information or services during the parsing process. It synthesizes the information of the context, and then realizes the complete control of the parsed content. For the implementation method of file confidentiality detection, especially for the security analysis of sensitive words, it is essential.

[0099] In this embodiment, an implementation method of file confidentiality detection based on Tika provided by the present invention includes steps 1 to 4:

[0100] Step 1, based on the sensitive information word library, lock the sensitive words in the text or code.

[0101] For the file to be detected for confidentiality, use Parser to parse all files in its directory, and convert all relevant information into strings, such as the above code segment 2. For the sensitive words to be detected, based on the pre-set sensitive word library, relevant sensitive words can be extracted, such as the characters "interface key", "key.env", and "api_key" in code segment 2.

[0102] Step 2, based on semantic analysis and combined with the context information of sensitive words, construct a logical binary tree corresponding to the sensitive words.

[0103] Figures 2A to 2C The logical binary trees corresponding to code segments 1 to 3 are respectively shown. The reason for choosing the logical binary tree is that the structural logic of the code is more complex. If a graph model is used, it is difficult to accurately describe the causal relationship between codes and the logical nesting relationship between codes. And if an ordinary binary tree model is used, it is difficult to describe the logical nesting relationship between codes.

[0104] In this embodiment, the causal relationship is represented in the form of parent and child nodes, and the logical nesting relationship is represented in the form of sibling nodes, as Figures 2A to 2C shown.

[0105] In Figure 2A , 01 to 10 can respectively represent sdk_key, API_KEY, "MyAPIKey12345", "API_KEY", API_KEY, properties, fileInput, keyPath, " / key.env", envPath, "Jupyter_notebook" in code segment 1.

[0106] In Figure 2B it, 01 - 05 can respectively represent load_dotenv in code segment 2,

[0107] "C: / Users / xxx / key.env", api_key, "API_KEY", API_KEY,'sk-xxxxxxxH'.

[0108] In Figure 2C it, 00 - 04 can respectively represent username, get_key, username + "_db_1234", username + "_api_key_888", "8888" in code segment 3.

[0109] It should be noted that Figures 2A to 2C nodes x1, x2 or x3 in it are blank nodes, which have no meaning by themselves, and are only used to determine the logical relationship between normal nodes and the distance value between normal nodes. In this embodiment, if not mentioned, "blank nodes" refer to non - blank normal nodes.

[0110] Taking Figure 2A as an example, the distance between nodes is regarded as the distance of their positions in the text. For the distance between a node and a blank node, it can be set arbitrarily, as long as the distance between nodes is accurate. For example, since the sum of the distance value between node 00 and node x2 and the distance value between node 01 and node x2 is a fixed value, but the distance value between any node 00 and node x2 or the distance value between node 01 and node x2 can be set arbitrarily.

[0111] It should be additionally noted here that although the technical term mentioned in this article is a logical binomial tree, actually from Figures 2A to 2C the perspective, it is more like a Fibonacci tree rather than a binomial tree. But this is only because we have not completed all the blank nodes! It should be noted that when calculating the j - th leaf node in the i - th sub - binomial tree, the specific values of i and j must be determined from the perspective of the binomial tree.

[0112] Step 3, based on the word embedding model, calculate the word embedding vector corresponding to each node, and based on the word embedding vector, calculate the similarity between any two nodes; and classify different nodes based on this similarity.

[0113] Specifically, the word embedding model can choose Word2Vec or BERT. In the present invention, Word2Vec is selected to obtain the word embedding vector x j .

[0114] Node v k Associated adjacency matrix V k , and its expression is:

[0115]

[0116] In the formula, if i = j, v ij represents the i-th child node of the vertex in the logical binomial tree. If i ≠ j, v ij represents the distance from the j-th leaf node to the node v in the sub-binomial tree with the i-th child node as the root node, where i, j = 1, 2, …, N. N represents the depth of the logical binomial tree. k

[0117] Take Figure 2A as an example: Since the depth of the sub-binomial tree with node x2 as the vertex is 2, therefore, it is actually the 3rd sub-binomial tree of the vertex in this logical binomial tree. Correspondingly, leaf nodes 00 and 02 respectively correspond to v 31 and v 32 . Since the depth of the sub-binomial tree with node 03 as the vertex is 5, therefore, it is actually the 6th sub-binomial tree of the vertex in this logical binomial tree. At this time, leaf nodes 10, 08, and 05 respectively correspond to v 65 , v 64 and v 62 .

[0118] Take Figure 2B as an example: Leaf nodes 00, 03, and 05 respectively correspond to v 43 , v 42 and v 21 . Note that at this time, the actual depth of the binomial tree with x2 as the root node is 3 instead of 2. In addition, node 01 is not a leaf node. A leaf node can be understood as the node determined by causal relationship rather than logical nesting relationship at the last time (the bottom layer).

[0119] Take Figure 2C as an example: Leaf nodes 00 and 02 respectively correspond to v 11 and v 41 . Note that the actual depth of the binomial tree with 01 as the root node is 3 instead of 1. That is to say, the actual depth of the binomial tree is greater than or equal to the number of child nodes of the root node (for example: Figure 2C 01 in it), and greater than or equal to the number of child nodes of the root node (for example: Figure 2B x2 in it) plus 1, and greater than or equal to the number of child nodes of the child node of the root node (for example: Figure 2B x3 in it) plus 2, and so on.

[0120] ​In this embodiment, only leaf nodes are concerned because only the final leaf nodes can truly determine the security leakage risk. This is the case for both the environment variables in Code Segment 1 and the actual stored keys in Code Segment 2. Further analysis of Code Segment 2 shows that if a logical binary tree is constructed only for File 21, its security risk is actually unknown. When File 22 is included, we are actually passing (transferring) the security risk to File 22.

[0121] This embodiment further includes steps A1 to A7:

[0122] Step A1: Calculate the mapping vector of each node in the logical binary tree according to the data type of the node.

[0123] Specifically, classify the nodes according to similarity to obtain different data types, and based on the values corresponding to the nodes of the corresponding data type collected at any moment, combined with a preset weight learning matrix, calculate the mapping vector of each node in the logical binary tree.

[0124] Step A2: Calculate the neighborhood feature vector of any node with respect to any type of data information.

[0125] Specifically, based on the set of nodes of the same data type in the logical binary tree and the adjacency matrix associated with the nodes in the set of nodes, calculate the neighborhood feature vector of any node in the logical binary tree with respect to any data type.

[0126] Step A3: Calculate the attention scores of all data types.

[0127] Specifically, call the activation function to calculate the attention scores of all data types according to the preset attention vector of any data type.

[0128] Step A4: Calculate the influence factor between any two nodes in the front and back times according to the attention scores.

[0129] Specifically, calculate the influence factor between any two nodes according to the attention scores of all data types in the logical binary tree.

[0130] Step A5: Merge the nodes with the same root based on a preset confidence threshold.

[0131] Specifically, when the influence factor between two nodes is greater than the preset confidence threshold (usually set to 0.95), then merge these two nodes to form a new node, and its value is the sum of the original two nodes.

[0132] It should be noted that the confidence threshold should be related to the total number of nodes, that is, when the number of nodes is larger, the confidence threshold is relatively smaller, and when the total number of nodes is smaller, the confidence threshold is relatively larger.

[0133] Step A6: Calculate the feature vectors of each data type according to the influence factors between two nodes in the front and back times under the same type.

[0134] Specifically, call the activation function to process the obtained influence factors and mapping vectors, and weight the processing results of the activation function to obtain the weighted feature vectors of the corresponding data types.

[0135] Step A7: Establish a heap embedding matrix to establish a logical binomial tree network model.

[0136] Specifically, first establish a logical binomial tree network model according to the heap embedding matrix of the logical binomial tree, and perform initialization processing on the weight matrix in the logical binomial tree network model. The word embedding vectors are transformed into query vectors, key vectors, and value vectors through the self-attention layer.

[0137] Step 4: Based on the deep learning network, input the word embedding vectors to obtain the detection results.

[0138] Specifically, it includes steps 4.1 to 4.2:

[0139] Step 4.1: A prediction model based on the cross-entropy loss function, and update the weights and biases of the logical binomial tree network model according to the feedforward network.

[0140] By performing multiple iterations on step 4.1, a more robust logical binomial tree network model is established.

[0141] Step 4.2: Predict the current file to be detected according to the trained logical binomial tree network model, and obtain the confidentiality detection result of the current file to be detected based on the word embedding vectors of the current file to be detected.

[0142] Among them, the feedforward network is represented as two-layer linear transformations, namely the first-layer linear transformation and the second-layer linear transformation. The feedforward network is obtained by calling the activation function to process the weights and biases of the first-layer linear transformation and the second-layer linear transformation.

[0143] Next, the device for implementing file confidentiality detection based on Tika provided by the present invention will be described. The device for implementing file confidentiality detection based on Tika described below can be correspondingly referred to the method for implementing file confidentiality detection based on Tika described above.

[0144] In one embodiment, a Tika-based file confidentiality detection implementation device includes a sensitive word locking module, a logical binary tree construction module, a node classification module, and a confidentiality detection module.

[0145] The sensitive word locking module is used to obtain a sensitive information word library and detect and lock sensitive words in the target text or target code based on the sensitive information word library.

[0146] The logical binary tree construction module is used to construct a logical binary tree corresponding to the sensitive word through semantic analysis according to the context information of the sensitive word.

[0147] The node classification module is used to calculate the word embedding vector corresponding to each node in the logical binary tree through a word embedding model, and calculate the similarity between any two nodes based on the word embedding vector, and the similarity is used to classify different nodes.

[0148] The confidentiality detection module is used to use the word embedding vector as the input of the deep learning network model to call the deep learning network to detect the word embedding vector and obtain the confidentiality detection result of the target text or target code.

[0149] Among them, the relationship between the parent node and the child node in the logical binary tree is used to represent the causal relationship of the context information of the sensitive word, and the relationship between the sibling nodes is used to represent the logical nesting relationship of the context information of the sensitive word.

[0150] In this embodiment, for the Tika-based file confidentiality detection implementation device provided by the present invention, the sensitive word locking module is specifically used for:

[0151] Parse all sub-files in the file directory to be confidentiality-detected through a Parser, and convert the parsing result into a string. The parsing result includes sensitive words in the file to be confidentiality-detected.

[0152] Extract sensitive words from the parsing result based on the pre-input sensitive information word library and lock the sensitive words.

[0153] In this embodiment, for the Tika-based file confidentiality detection implementation device provided by the present invention, the node classification module is specifically used for:

[0154] Obtain the word embedding vector corresponding to each node in the logical binary tree through the Word2Vec word embedding model, and obtain the adjacency matrix V k associated with the node v k .

[0155] Among them, the expression of the adjacency matrix is:

[0156]

[0157] In the formula, if i=j, v ij Represents the i-th child node of a vertex in a logical binomial tree. If i≠j, v ij Indicates that in the sub-binomial tree with the i-th child node as the root node, the j-th leaf node to the node v k The distance between them, i,j = 1, 2,…, N, where N represents the depth of the logical binomial tree.

[0158] In this embodiment, the present invention provides a Tika-based file confidentiality detection implementation device, which further includes a first computing module, which is used to:

[0159] The data types are obtained by classifying different nodes according to their similarities, and based on the values ​​corresponding to the nodes of the corresponding data types collected at any time, combined with the preset weight learning matrix, the mapping vector of each node in the logical binomial tree is calculated.

[0160] Based on a set of nodes of the same data type in the logical binomial tree and an adjacency matrix associated with the nodes in the node set, a neighborhood feature vector of any node in the logical binomial tree with respect to any data type is calculated.

[0161] In this embodiment, the present invention provides a Tika-based file confidentiality detection implementation device, which further includes a second computing module, which is used to:

[0162] Call the activation function to calculate the attention scores of all data types based on the preset attention vector of any data type.

[0163] Calculate the influence factor between any two nodes based on the attention scores of all data types in the logistic binomial tree.

[0164] In this embodiment, the present invention provides a Tika-based file confidentiality detection implementation device, which also includes a model building and initialization module for:

[0165] Based on the preset confidence threshold, the child nodes of the same data type are merged, and according to the influence factor between any two nodes of the same data type, the activation function is called to calculate the feature vector of each data type.

[0166] A logical binomial tree network model is established according to the heap embedding matrix of the logical binomial tree, and the weight matrix in the logical binomial tree network model is initialized. The word embedding vector is converted into a query vector, a key vector, and a value vector through the self-attention layer.

[0167] In this embodiment, the present invention provides a Tika-based file confidentiality detection implementation device, and the confidentiality detection module is specifically used for:

[0168] A prediction model based on the cross-entropy loss function, and update the weights and biases of the logical binomial tree network model according to the feedforward network.

[0169] Predict the current file to be detected according to the trained logical binomial tree network model, and obtain the confidentiality detection result of the current file to be detected based on the word embedding vector of the current file to be detected.

[0170] Among them, the feedforward network is represented as two-layer linear transformations, namely the first-layer linear transformation and the second-layer linear transformation, and the feedforward network is obtained by processing the weights and biases of the first-layer linear transformation and the second-layer linear transformation by calling an activation function.

[0171] This disclosure may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of this disclosure.

[0172] The computer-readable storage medium may be a tangible device that can retain and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0173] The computer-readable program instructions described herein may be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0174] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0175] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0176] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams. These computer-readable program instructions may also be stored in a computer-readable storage medium, which instructions cause a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams.

[0177] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0178] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A method for realizing file confidentiality detection based on Tika, characterized in that: The method comprises: Acquire a sensitive information word library, and detect and lock sensitive words in the target text or target code based on the sensitive information word library; According to the context information of the sensitive words, construct a logical binomial tree corresponding to the sensitive words through semantic analysis; Calculate the word embedding vector corresponding to each node in the logical binomial tree through the word embedding model, and calculate the similarity between any two nodes based on the word embedding vector, wherein the similarity is used to classify different nodes; Using the word embedding vector as an input of a deep learning network model to call the deep learning network to detect the word embedding vector and obtain a confidentiality detection result of the target text or target code; The relationship between the parent node and the child node in the logical binomial tree is used to represent the causal relationship of the context information of the sensitive words, and the relationship between the sibling nodes is used to represent the logical nested relationship of the context information of the sensitive words; The word embedding model is used to calculate the word embedding vector corresponding to each node in the logical binomial tree, and the similarity between any two nodes is calculated based on the word embedding vector, and the similarity is used to classify different nodes, and then it also includes: Based on a preset confidence threshold, child nodes of the same data type are merged, and according to the influence factor between any two nodes of the same data type, an activation function is called to calculate the feature vector of each data type; and A logical binomial tree network model is established according to the heap embedding matrix of the logical binomial tree, and a weight matrix in the logical binomial tree network model is initialized, and the word embedding vector is converted into a query vector, a key vector and a value vector through a self-attention layer; The using the word embedding vector as the input of the deep learning network model to call the deep learning network to detect the word embedding vector and obtain the confidentiality detection result of the target text or target code includes: A prediction model based on a cross entropy loss function, and updating the weights and biases of the logistic binomial tree network model according to a feedforward network; Predicting the current file to be detected according to the trained logical binomial tree network model, and obtaining the confidentiality detection result of the current file to be detected based on the word embedding vector of the current file to be detected; Among them, the feedforward network is represented by two layers of linear transformation, namely the first layer of linear transformation and the second layer of linear transformation. The feedforward network is obtained by calling the activation function to process the weights and biases of the first layer of linear transformation and the second layer of linear transformation.

2. The method for realizing file confidentiality detection based on Tika according to claim 1, characterized in that: The step of obtaining a sensitive information word library and detecting and locking sensitive words in a target text or target code based on the sensitive information word library includes: Parse all subfiles in the file directory to be tested for confidentiality by Parser, and convert the parsing results into strings, wherein the parsing results include sensitive words in the file to be tested for confidentiality; The sensitive words are extracted from the analysis results based on the pre-input sensitive information word library, and the sensitive words are locked.

3. The method for realizing file confidentiality detection based on Tika according to claim 1, characterized in that: The word embedding model is used to calculate the word embedding vector corresponding to each node in the logical binomial tree, and the similarity between any two nodes is calculated based on the word embedding vector, and the similarity is used to classify different nodes, including: The word embedding vector corresponding to each node in the logical binomial tree is obtained through the Word2Vec word embedding model, and the word embedding vector corresponding to the node v in the logical binomial tree is obtained based on the word embedding vector. k The associated adjacency matrix V k ; Wherein, the expression of the adjacency matrix is: In the formula, if i=j, v ij Represents the i-th child node of a vertex in a logical binomial tree. If i≠j, v ij In the sub-binomial tree whose root node is the i-th child node, the j-th leaf node to the node v k The distance between them, i,j = 1, 2,…, N, where N represents the depth of the logical binomial tree.

4. The method for realizing file confidentiality detection based on Tika according to claim 3 is characterized in that: The word embedding model is used to calculate the word embedding vector corresponding to each node in the logical binomial tree, and the similarity between any two nodes is calculated based on the word embedding vector, and the similarity is used to classify different nodes, and then it also includes: According to the data types obtained by classifying different nodes according to similarity, and based on the values ​​corresponding to the nodes of the corresponding data types collected at any time, combined with the preset weight learning matrix, the mapping vector of each node in the logical binomial tree is calculated; and Based on a node set of the same data type in the logical binomial tree and an adjacency matrix associated with nodes in the node set, a neighborhood feature vector of any node in the logical binomial tree with respect to any data type is calculated.

5. The method for realizing file confidentiality detection based on Tika according to claim 4 is characterized in that: The word embedding model is used to calculate the word embedding vector corresponding to each node in the logical binomial tree, and the similarity between any two nodes is calculated based on the word embedding vector, and the similarity is used to classify different nodes, and then it also includes: Call the activation function to calculate the attention scores of all data types based on the preset attention vector of any data type; According to the attention scores of all data types in the logical binomial tree, the influence factor between any two nodes is calculated.

6. A device for realizing file confidentiality detection based on Tika, applied to the method described in any one of claims 1 to 5, characterized in that: The device comprises: A sensitive word locking module, used to obtain a sensitive information word library, and detect and lock sensitive words in a target text or target code based on the sensitive information word library; A logical binomial tree construction module, used to construct a logical binomial tree corresponding to the sensitive word through semantic analysis according to the context information of the sensitive word; A node classification module, used to calculate the word embedding vector corresponding to each node in the logical binomial tree through a word embedding model, and calculate the similarity between any two nodes based on the word embedding vector, wherein the similarity is used to classify different nodes; A confidentiality detection module, used to use the word embedding vector as an input of a deep learning network model to call the deep learning network to detect the word embedding vector and obtain a confidentiality detection result of the target text or target code; The relationship between the parent node and the child node in the logical binomial tree is used to represent the causal relationship of the context information of the sensitive words, and the relationship between the sibling nodes is used to represent the logical nested relationship of the context information of the sensitive words.

7. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Code automatic abstracting method based on structure position awareness

    CN117407051A

  • Using machine learning to determine electronic document similarity

    US20200125648A1