A Sensitive Content Analysis Method and System Based on Continuous Neural Trees
By dividing and subgrouping each layer of the tree in the continuous neural tree, the problems of complexity and recognition accuracy in the analysis of sensitive content are solved, and higher recognition accuracy and feature retention are achieved.
Patent Information
- Application Number
- CN202510429001.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Continuous neural trees are too complex in the analysis of sensitive content, are difficult to train, and the recognition accuracy needs to be improved.
By dividing each layer of the tree in the continuous neural tree, the output of different layers is input into different subgroups in the same layer, and the correspondence between the subgroups and the first few layers is established, and then whether sensitive information is included.
It improves the accuracy of recognition, avoids interference from information at different levels, can retain shallow and deep features, and utilize the features of the intermediate layers.
Smart Images

Figure CN119938889B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security, and specifically to a sensitive content analysis method and system based on a continuous neural tree. Background Art
[0002] Sensitive content analysis is one of the hot issues in the field of information security today. In this era, information security and data privacy protection issues have attracted much attention. With the rapid development of big data technology, traditional information security technologies can no longer meet the needs of information security. Therefore, advanced technical means are needed to solve the problems of information leakage and sensitive content analysis. As an important artificial intelligence technology, neural networks have made many breakthroughs in the field of information security. Neural networks have good self-learning and adaptive capabilities and can effectively handle complex information security problems. As an extended form of neural networks, continuous neural trees have great potential for sensitive content analysis problems. However, in the process of use, continuous neural trees are too complex, difficult to train, and the recognition accuracy needs to be further improved. Summary of the Invention
[0003] To solve the above problems, the present invention provides a sensitive content analysis method based on a continuous neural tree. The method includes the following steps:
[0004] After preprocessing the text to be analyzed, it is used as the input of the continuous neural tree. The number of trees in the i-th layer of the continuous neural tree model is obtained, and the trees in the i-th layer are divided into i - 1 subgroups, and the corresponding relationship between the subgroups and the first i - 1 layers is established; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree.
[0005] Based on the corresponding relationship, the output of the layer corresponding to the subgroup is obtained, and the output of the layer corresponding to the subgroup and the input of the continuous neural tree model are used as the input of each tree in the subgroup.
[0006] The output of each subgroup in the continuous neural tree model is obtained, and it is determined whether sensitive information is included according to the output of each subgroup. If it is included, a prompt is issued.
[0007] Preferably, the establishment of the corresponding relationship between the subgroup and the first i - 1 layers is specifically:
[0008] The layers are arranged in the order of the layers in the continuous neural tree model, and the subgroups are sorted in the order of the subgroups in the layer where they are located;
[0009] The layers and subgroups with the same serial number are established with a corresponding relationship.
[0010] Preferably, the determination of whether sensitive information is included according to the output of each subgroup is specifically:
[0011] Obtain the results of the k-th subgroup in all layers, and merge the results of all the k-th subgroups; further obtain the merged results of N - 2 subgroups, and input the merged results of N - 2 subgroups into the first prediction unit to obtain the first prediction result; where k is a positive integer and k ≤ N - 2;
[0012] Merge the results of the last N - 2 layers of all layers, and input them into the second prediction unit to obtain the second prediction result;
[0013] Determine whether sensitive information is included according to the first prediction result and the second prediction result.
[0014] Preferably, the determining whether sensitive information is included according to the first prediction result and the second prediction result is specifically:
[0015] If the first prediction result and the second prediction result are the same, determine whether sensitive information is included according to the first prediction result or the second prediction result;
[0016] If the first prediction result and the second prediction result are different, issue a prompt for manual review.
[0017] Preferably, each node of the tree has a learnable vector;
[0018] After training the continuous neural tree model, judge the similarity between the learnable vectors of the nodes at the same position of each tree in the same subset and the similarity of the segmentation threshold θ; determine the similarity of the nodes at the same position according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ; furthermore, calculate the similarity between two trees according to the similarity of the nodes at the same position.
[0019] When the similarity between multiple trees in the same subset is greater than the threshold, merge the multiple trees; retrain or continue to train the merged continuous neural network model.
[0020] In another aspect, the present invention also provides a sensitive content analysis system based on a continuous neural tree, and the system includes the following modules:
[0021] A division module, configured to use the text to be analyzed after preprocessing as the input of the continuous neural tree, obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i - 1 subgroups, and establish a corresponding relationship between the subgroups and the previous i - 1 layers; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree;
[0022] An input determination module, configured to obtain the output of the layer corresponding to the subgroup based on the corresponding relationship, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup;
[0023] A result output module, which is used to obtain the output of each subgroup in the continuous neural tree model, determine whether sensitive information is included according to the output of each subgroup, and issue a prompt if it is included.
[0024] Preferably, the establishment of the correspondence between the subgroup and the previous i - 1 layers is specifically as follows:
[0025] Arrange the layers in the order of the continuous neural tree model, and sort the subgroups in the order of the subgroups in their respective layers;
[0026] Establish a correspondence between the layers and subgroups with the same serial numbers.
[0027] Preferably, the determination of whether sensitive information is included according to the output of each subgroup is specifically as follows:
[0028] Obtain the results of the k - th subgroup in all layers, and merge the results of all the k - th subgroups; further obtain the merged results of N - 2 subgroups, and input the merged results of the N - 2 subgroups into the first prediction unit to obtain the first prediction result; where k is a positive integer and k ≤ N - 2;
[0029] Merge the results of the last N - 2 layers of all layers and input them into the second prediction unit to obtain the second prediction result;
[0030] Determine whether sensitive information is included according to the first prediction result and the second prediction result.
[0031] Preferably, the determination of whether sensitive information is included according to the first prediction result and the second prediction result is specifically as follows:
[0032] If the first prediction result and the second prediction result are the same, determine whether sensitive information is included according to the first prediction result or the second prediction result;
[0033] If the first prediction result and the second prediction result are different, issue a prompt for manual review.
[0034] Preferably, each node of the tree has a learnable vector;
[0035] After training the continuous neural tree model, judge the similarity between the learnable vectors of the nodes at the same positions of each tree in the same subset and the similarity of the segmentation threshold θ; determine the similarity of the nodes at the same positions according to the similarity between the learnable vectors and the similarity of the segmentation threshold θ; further calculate the similarity of the two trees according to the similarity of the nodes at the same positions.
[0036] When the similarity of multiple trees in the same subset is greater than the threshold, merge the multiple trees; retrain or continue to train the merged continuous neural network model.
[0037] In addition, the present invention also provides a computer program product, which implements the method as described above when executed by a processor.
[0038] By dividing each layer of the continuous neural tree, the outputs of different layers are input into different subgroups in the same layer, preventing interference with information at different levels, being able to retain shallow features, deep features, and being able to utilize the features of the intermediate layer, improving the accuracy of recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of Embodiment 1;
[0040] Figure 2 is a structural diagram of the continuous neural tree;
[0041] Figure 3 is a structural diagram of the subgroup;
[0042] Figure 4 is a structural diagram of Embodiment 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element. In the present invention, if the collection of data related to personal privacy, such as faces, mobile phone usage information, etc., is involved, personal permission will be obtained in advance, including but not limited to oral reminders, posting posters, mobile phone reminders, etc.; if it involves conflicts with laws and regulations, it will be produced or used within the scope permitted by laws and regulations.
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0045] Embodiment 1, as Figure 1As shown in the figure, the present invention provides a sensitive content analysis method based on a continuous neural tree. The method includes the following steps:
[0046] S1. After preprocessing the text to be analyzed, use it as the input of the continuous neural tree. Obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i - 1 subgroups, and establish the corresponding relationship between the subgroups and the first i - 1 layers. Here, i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and non-output layers of the continuous neural tree.
[0047] When analyzing sensitive content, if it is to analyze sensitive information in the operating system, the sensitive information of the operating system in the text to be analyzed includes, but is not limited to, information related to unauthorized access, system log information, etc. Among them, the information related to unauthorized access includes, but is not limited to, source IP, timestamp, resource to be accessed, remote login or local login, operations on the resource, reasons for failure, reasons for success, etc. Of course, the text to be analyzed can also have other fields, which are only used as examples here. If it is to detect sensitive information in a text, such as a paragraph, divide the text into natural sentences, then perform word segmentation, and use the encoding of words in each sentence of the text to be analyzed as an input. Among them, preprocessing the text to be analyzed includes removing spaces, removing stop words, vectorization, etc.
[0048] The continuous neural tree model includes a multi-layer structure, as Figure 2 shown. Each layer contains multiple trees. In a specific embodiment, ODTSR trees are used. For the i-th layer in the continuous neural tree model, divide the trees in it into i - 1 subgroups, and then establish the corresponding relationship between the subgroups and the first i - 1 layers. For example, in the 3rd layer, there are 8 trees, which are divided into 2 subgroups, with 4 trees in each subgroup. The first subgroup corresponds to the 1st layer of the continuous neural tree model, and the second subgroup corresponds to the 2nd layer of the continuous neural tree model. Here, i is a positive integer, but 2 < i ≤ N, and N is the total number of non-input and non-output layers of the continuous neural tree. Assume that the continuous neural tree model has a total of 4 non-input and non-output layers, then N = 4, i = 3, 4. That is, divide the neural trees in the 3rd layer into 2 subgroups, and divide the neural trees in the 4th layer into 3 subgroups. The total number of non-input and non-output layers refers to the total number of layers of the continuous neural tree model including the input layer and the output layer minus 2, that is, the remaining number of layers after subtracting the input layer and the output layer.
[0049] In a more specific embodiment, the establishment of the corresponding relationship between the subgroups and the first i - 1 layers is specifically as follows:
[0050] Arrange the layers in the order of the sequence of layers in the continuous neural tree model, and sort the subgroups in the order of the subgroups in their respective layers;
[0051] For non-input / output layers, the layers are sorted in the order from front to back, that is, the non-input / output layers are sorted in the order of the first layer, the second layer, the third layer, etc., and then the subgroups are sorted according to the order of the layers where the subgroups are located, such as the first subgroup, the second subgroup, etc.
[0052] Establish a corresponding relationship between the layers and subgroups with the same serial numbers.
[0053] Establish the corresponding relationship between the layers and subgroups in the continuous neural tree model. Preferably, establish the corresponding relationship between the layers and subgroups with the same serial numbers.
[0054] In order to be able to optimize the continuous neural tree model subsequently, in a specific embodiment, the tree in the i-th layer is divided into i - 1 subgroups, specifically: when constructing the continuous neural tree model, set the structure of the tree according to the corresponding relationship between the layers and subgroups. For subgroups with smaller serial numbers, the depth of the tree is greater; for subgroups with larger serial numbers, the depth of the tree is smaller, so as to prevent the depth of the tree in the subgroup with a larger serial number from being greater, while its input itself contains the output of deeper layers in the continuous neural tree model, which extracts too much depth information.
[0055] In a preferred embodiment, the tree structures of each subgroup are the same, that is, they contain the same nodes and depths, as Figure 3 shown.
[0056] In another embodiment, each node of the tree has a learnable vector;
[0057] After training the continuous neural tree model, judge the similarity between the learnable vectors of the nodes at the same positions of each tree in the same subset and the similarity of the segmentation threshold θ; determine the similarity of the nodes at the same positions according to the similarity between the learnable vectors and the similarity of the segmentation threshold θ; and then calculate the similarity between the two trees according to the similarity of the nodes at the same positions.
[0058] In an alternative embodiment, the tree in the continuous neural network model adopts an ODT tree (Oblivious Decision Tree). After training the continuous neural tree model, or after completing several patch trainings, for the same subset, if the tree structures are the same, the learnable vectors of each node of the tree are the same or similar, and the segmentation thresholds are also the same or similar, then the final results of these two trees are the same. In order to simplify the structure of the continuous neural tree model, the present invention judges the learnable vectors of the nodes at the same positions for the trees with the same structure in the same subset and the similarity of the segmentation threshold θ. If the learnable vectors of all the nodes at the same positions of the two trees If the similarities with the splitting threshold θ are both large, then when the similarities of multiple trees in the same subset are greater than the threshold, the multiple trees are merged; the merged continuous neural network model is retrained or trained for the next patch. Among them, the learnable vector is also called the selection vector.
[0059] Among them, the similarity of nodes at the same position is determined according to the similarity of the learnable vector and the similarity of the splitting threshold θ. Specifically: it is judged whether the similarity of the learnable vector is greater than the first threshold, and whether the similarity of the splitting threshold θ is greater than the second threshold. If both are greater, the two nodes are considered similar.
[0060] Furthermore, if all the nodes at the same position of two trees are similar, the two trees are considered similar.
[0061] It should be noted that in one embodiment, the number of layers in the continuous neural tree model, the number, structure, and subgroups of trees in each layer are predetermined. During the training process, the number of trees in each layer can be continuously adjusted, for example, by merging trees according to the similarity of trees in the subgroup.
[0062] In a more specific embodiment, the process of merging trees is that the average value of the learnable parameters and the average value of the splitting threshold of the same nodes of the two trees to be merged are used as the learnable parameters and splitting threshold of the newly merged tree. This will involve changes in the number of outputs of all subtrees after merging subtrees, but through feature fusion including but not limited to averaging, convolution, etc., the dimensions before and after the change are made the same, which will not be elaborated here.
[0063] S2. Obtain the output of the layer corresponding to the subgroup based on the corresponding relationship, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup;
[0064] In one embodiment, the output of the layer corresponding to the subgroup and the input of the continuous neural tree model are used as the input of each tree in the subgroup. Among them, the layer corresponding to the subgroup refers to the layer corresponding to the subgroup determined through the corresponding relationship; the layer where the subgroup is located refers to the layer of the continuous neural tree model where the subgroup is located. The present invention can not only obtain shallow features and deep features, but also obtain intermediate layer features, avoiding only extracting shallow features and deep features, being able to extract features deeper than the shallow layer, and avoiding too many and too abstract deep features.
[0065] In an alternative embodiment of S2, after establishing the corresponding relationship between the subgroup and the layer, the output of the corresponding layer is used as all or part of the input of the subgroup. Specifically, the output of the layer corresponding to the subgroup and the output of the layer above the layer where the subgroup is located are used as the input of each tree in the subgroup.
[0066] S3. Obtain the output of each subgroup in the continuous neural tree model, and determine whether sensitive information is included according to the output of each subgroup. If it is included, issue a prompt.
[0067] In a specific embodiment, the determining whether sensitive information is included according to the output of each subgroup is specifically as follows:
[0068] Obtain the results of the k-th subgroup in all layers, and merge the results of all the k-th subgroups; further obtain the merged results of N - 2 subgroups, and input the merged results of N - 2 subgroups into the first prediction unit to obtain the first prediction result; where k is a positive integer and k ≤ N - 2;
[0069] For example, there are four layers. There are two subgroups A31 and A32 in the third layer, and three subgroups A41, A42, and A43 in the fourth layer. The first subgroups in all layers are A31 and A41 respectively, the second subgroups in all layers are distributed as A32 and A42 respectively, and the third subgroups in all layers are distributed as A43. Merge A31 and A41 to obtain the merged result 1, merge A32 and A42 to obtain the merged result 2. Since A43 has only one subgroup, the merged result is still itself, that is, the merged result 3. Input the merged results 1, 2, and 3 into the first prediction unit. Among them, all layers refer to the layers containing subgroups.
[0070] Merge the results of the last N - 2 layers of all layers and input them into the second prediction unit to obtain the second prediction result;
[0071] The results output by the last N - 2 layers of all layers are input into the second prediction unit to obtain the second prediction result.
[0072] Then, determine whether sensitive information is included according to the first prediction result and the second prediction result.
[0073] Preferably, the determining whether sensitive information is included according to the first prediction result and the second prediction result is specifically as follows:
[0074] If the first prediction result and the second prediction result are the same, then determine whether sensitive information is included according to the first prediction result or the second prediction result;
[0075] If the first prediction result and the second prediction result are different, then issue a prompt for manual review.
[0076] Embodiment 2. The present invention also provides a sensitive content analysis system based on a continuous neural tree, as Figure 4 shown. The system includes the following modules:
[0077] A partitioning module, which is used to take the text to be analyzed after preprocessing as the input of a continuous neural tree, obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i-1 subgroups, and establish a correspondence between the subgroups and the previous i-1 layers; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree.
[0078] An input determination module, which is used to obtain the output of the layer corresponding to the subgroup based on the correspondence, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup.
[0079] A result output module, which is used to obtain the output of each subgroup in the continuous neural tree model, and determine whether sensitive information is included according to the output of each subgroup. If it is included, a prompt is issued.
[0080] Preferably, the establishment of the correspondence between the subgroups and the previous i-1 layers is specifically as follows:
[0081] Arrange the layers in the order of the layers in the continuous neural tree model, and sort the subgroups in the order of the subgroups in the layer where they are located;
[0082] Establish a correspondence between the layers and subgroups with the same serial numbers.
[0083] Preferably, the determination of whether sensitive information is included according to the output of each subgroup is specifically as follows:
[0084] Obtain the results of the k-th subgroup in all layers, and merge the results of all the k-th subgroups; further obtain the merged results of N-2 subgroups, and input the merged results of the N-2 subgroups into the first prediction unit to obtain the first prediction result; where k is a positive integer and k ≤ N-2;
[0085] Merge the results of the last N-2 layers of all layers and input them into the second prediction unit to obtain the second prediction result;
[0086] Determine whether sensitive information is included according to the first prediction result and the second prediction result.
[0087] Preferably, the determination of whether sensitive information is included according to the first prediction result and the second prediction result is specifically as follows:
[0088] If the first prediction result and the second prediction result are the same, determine whether sensitive information is included according to the first prediction result or the second prediction result;
[0089] If the first prediction result and the second prediction result are different, issue a prompt for manual review.
[0090] Preferably, each node of the tree has a learnable vector.
[0091] After training the continuous neural tree model, judge the similarity between the learnable vectors of the nodes at the same position of each tree in the same subset and the similarity of the segmentation threshold θ; determine the similarity of the nodes at the same position according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ; furthermore, calculate the similarity between two trees according to the similarity of the nodes at the same position.
[0092] When the similarity between multiple trees in the same subset is greater than the threshold, merge the multiple trees; retrain or continue to train the merged continuous neural network model.
[0093] Embodiment 3. The present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program realizes the method described above when executed by a processor.
[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solution essentially or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Other embodiments can also be adopted; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A sensitive content analysis method based on a continuous neural tree, characterized in that: The method includes the following steps: The text to be analyzed is preprocessed and used as the input of the continuous neural tree. The sensitive content is the sensitive information in the operating system in the text to be analyzed. When analyzing the sensitive content, the sensitive information in the operating system includes information related to unauthorized access and system log information. The information related to unauthorized access includes source IP, timestamp, resource to be accessed, remote login, local login, operations on resources, reasons for failure, and reasons for success. Preprocessing the text to be analyzed includes removing spaces, removing stop words, and vectorizing. Obtain the number of trees in the i-th layer of the continuous neural tree model. Divide the trees in the i-th layer into i - 1 subgroups, and establish a correspondence between the subgroups and the previous i - 1 layers. Here, i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and non-output layers of the continuous neural tree. Based on the correspondence, obtain the output of the layer corresponding to the subgroup, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup. Obtain the output of each subgroup in the continuous neural tree model, and determine whether sensitive information is included according to the output of each subgroup. If it is included, a prompt is issued. The determination of whether sensitive information is included according to the output of each subgroup is specifically as follows: Obtain the results of the k-th subgroup in all layers, and merge the results of all the k-th subgroups. Obtain the merged results of N - 2 subgroups, and input the merged results of N - 2 subgroups into the first prediction unit to obtain the first prediction result. Here, k is a positive integer and k ≤ N - 2. Merge the results of the last N - 2 layers of all layers and input them into the second prediction unit to obtain the second prediction result. Determine whether sensitive information is included according to the first prediction result and the second prediction result. The determination of whether sensitive information is included according to the first prediction result and the second prediction result is specifically as follows: If the first prediction result and the second prediction result are the same, determine whether sensitive information is included according to the first prediction result or the second prediction result. If the first prediction result and the second prediction result are different, a prompt for manual review is issued.
2. The method according to claim 1, characterized in that The establishment of the correspondence between the subgroups and the previous i - 1 layers is specifically as follows: Arrange the layers in the order of the sequence of layers in the continuous neural tree model, and sort the subgroups in the order of the subgroups in their respective layers. Establish a correspondence between the layers and subgroups with the same serial numbers.
3. The method according to claim 1, characterized in that Each node of the tree has a learnable vector. After training the continuous neural tree model, judge the similarity between the learnable vectors of the nodes at the same positions of each tree in the same subset and the similarity of the segmentation threshold θ. Determine the similarity of the nodes at the same positions according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ. Furthermore, calculate the similarity of two trees according to the similarity of the nodes at the same positions. When the similarity of multiple trees in the same subset is greater than the threshold, merge the multiple trees. Retrain or continue to train the merged continuous neural network model.
4. A sensitive content analysis system based on a continuous neural tree, characterized in that: The system includes the following modules: A division module for preprocessing the text to be analyzed and using it as the input of the continuous neural tree. The sensitive content is the sensitive information in the operating system in the text to be analyzed. When analyzing sensitive content, the sensitive information in the operating system includes information related to unauthorized access and system log information. The information related to unauthorized access includes source IP, timestamp, resource to be accessed, remote login, local login, operations on resources, reasons for failure, and reasons for success. Preprocessing the text to be analyzed includes removing spaces, removing stop words, and vectorization. Obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i - 1 subgroups, and establish a correspondence between the subgroups and the first i - 1 layers. Here, i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree. An input determination module is used to obtain the output of the layer corresponding to the subgroup based on the correspondence, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup. A result output module is used to obtain the output of each subgroup in the continuous neural tree model, and determine whether sensitive information is included according to the output of each subgroup. If it is included, a prompt is issued. The determination of whether sensitive information is included according to the output of each subgroup is specifically as follows: Obtain the results of the k-th subgroup in all layers, and merge the results of all the k-th subgroups; obtain the merged results of N - 2 subgroups, and input the merged results of N - 2 subgroups into the first prediction unit to obtain the first prediction result. Here, k is a positive integer and k ≤ N - 2; merge the results of the last N - 2 layers of all layers and input them into the second prediction unit to obtain the second prediction result; determine whether sensitive information is included according to the first prediction result and the second prediction result. The determination of whether sensitive information is included according to the first prediction result and the second prediction result is specifically as follows: If the first prediction result and the second prediction result are the same, determine whether sensitive information is included according to the first prediction result or the second prediction result; if the first prediction result and the second prediction result are different, issue a prompt for manual review.
5. The system according to claim 4, characterized in that The establishment of the correspondence between the subgroups and the first i - 1 layers is specifically as follows: Arrange the layers in the order of the layers in the continuous neural tree model, and sort the subgroups in the order of the subgroups in their respective layers. Establish a correspondence between the layers and subgroups with the same serial numbers.
6. A computer storage device having a computer program stored thereon, characterized in that: The computer program, when executed by a processor, implements the method according to any one of claims 1 - 3.
Citation Information
Patent Citations
Production parameter optimization prediction method and device, equipment and storage medium
CN108647808A
Bamboo chip wormhole and mildew spot detection method based on convolutional flexible neural forest
CN111242895A