Sensitive content analysis method and system based on continuous neural tree

By dividing and subgrouping each layer of the tree in the continuous neural tree, the analysis of sensitive content is improved, and the complexity and recognition accuracy of continuous neural tree in sensitive content analysis is solved, and the accuracy of the analysis is improved.

CN119938889AActive Publication Date: 2025-05-06ZHONGKE LINGXUN (BEIJING) TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510429001.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Continuous neural trees are too complex in the analysis of sensitive content, are difficult to train, and the recognition accuracy needs to be improved.

Method used

By dividing each layer of the tree in the continuous neural tree, the output of different layers is input into different subgroups in the same layer, and the correspondence between the subgroups and the first few layers is established, and then whether sensitive information is included.

Benefits of technology

It improves the accuracy of recognition, avoids interference from information at different levels, can retain shallow and deep features, and utilize the features of the intermediate layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938889A_ABST
    Figure CN119938889A_ABST
Patent Text Reader

Abstract

The invention relates to the field of information security, and particularly discloses a sensitive content analysis method and system based on a continuous neural tree, a to-be-analyzed text is preprocessed to serve as the input of the continuous neural tree, the number of trees in the ith layer in a continuous neural tree model is obtained, the trees in the ith layer are divided into i-1 subgroups, and the subgroups are divided into i-1 subgroups; establishing a corresponding relationship between the subgroup and the previous i-1 layer; obtaining the output of the layer corresponding to the subgroup based on the corresponding relationship, and taking the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup; and obtaining the output of each subgroup in the continuous neural tree model, determining whether sensitive information is contained or not according to the output of each subgroup, and if yes, giving out a prompt. According to the method, mutual interference of features of different levels is overcome, the features of the middle level can be reserved, and the recognition accuracy of sensitive information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security, and specifically to a sensitive content analysis method and system based on a continuous neural tree. Background Art

[0002] Sensitive content analysis is one of the hot issues in the field of information security today. In this era, information security and data privacy protection issues have attracted much attention. With the rapid development of big data technology, traditional information security technologies can no longer meet the needs of information security. Therefore, advanced technical means are needed to solve the problems of information leakage and sensitive content analysis. As an important artificial intelligence technology, neural networks have made many breakthroughs in the field of information security. Neural networks have good self-learning and adaptive capabilities and can effectively handle complex information security problems. As an extended form of neural networks, continuous neural trees have great potential for sensitive content analysis problems. However, in the process of use, continuous neural trees are too complex, difficult to train, and the recognition accuracy needs to be further improved. Summary of the Invention

[0003] To solve the above problems, the present invention provides a sensitive content analysis method based on a continuous neural tree. The method includes the following steps: After preprocessing the text to be analyzed, it is used as the input of the continuous neural tree. The number of trees in the i-th layer of the continuous neural tree model is obtained, and the trees in the i-th layer are divided into i-1 subgroups, and the corresponding relationship between the subgroups and the first i-1 layers is established; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree. Based on the corresponding relationship, the output of the layer corresponding to the subgroup is obtained, and the output of the layer corresponding to the subgroup and the input of the continuous neural tree model are used as the input of each tree in the subgroup. The output of each subgroup in the continuous neural tree model is obtained, and it is determined whether sensitive information is included according to the output of each subgroup. If it is included, a prompt is issued.

[0004] Preferably, the establishment of the corresponding relationship between the subgroup and the first i-1 layers is specifically: The layers are arranged in the order of the layers in the continuous neural tree model, and the subgroups are sorted in the order of the subgroups in the layer where they are located; The layers and subgroups with the same serial number are established with a corresponding relationship.

[0005] Preferably, the determination of whether sensitive information is included according to the output of each subgroup is specifically: Obtain the results of the k-th subgroup in all layers, and merge the results of all the k-th subgroups; further obtain the merged results of N - 2 subgroups, and input the merged results of N - 2 subgroups into the first prediction unit to obtain the first prediction result; where k is a positive integer and k ≤ N - 2; Merge the results of the last N - 2 layers of all layers, and input them into the second prediction unit to obtain the second prediction result; Determine whether sensitive information is included according to the first prediction result and the second prediction result.

[0006] Preferably, the determining whether sensitive information is included according to the first prediction result and the second prediction result is specifically: If the first prediction result and the second prediction result are the same, then determine whether sensitive information is included according to the first prediction result or the second prediction result; If the first prediction result and the second prediction result are different, then issue a prompt for manual review.

[0007] Preferably, each node of the tree has a learnable vector; After training the continuous neural tree model, judge the similarity between the learnable vectors of the nodes at the same position of each tree in the same subset and the similarity of the segmentation threshold θ; determine the similarity of the nodes at the same position according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ; and then calculate the similarity of the two trees according to the similarity of the nodes at the same position. When the similarity of multiple trees in the same subset is greater than the threshold, merge the multiple trees; retrain or continue to train the merged continuous neural network model.

[0008] In another aspect, the present invention also provides a sensitive content analysis system based on a continuous neural tree, and the system includes the following modules: A division module, which is used to take the text to be analyzed after preprocessing as the input of the continuous neural tree, obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i - 1 subgroups, and establish the corresponding relationship between the subgroups and the first i - 1 layers; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree; An input determination module, which is used to obtain the output of the layer corresponding to the subgroup based on the corresponding relationship, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup; A result output module, which is used to obtain the output of each subgroup in the continuous neural tree model, determine whether sensitive information is included according to the output of each subgroup, and if it is included, issue a prompt.

[0009] Preferably, the establishing the corresponding relationship between the subgroups and the first i - 1 layers is specifically: Arrange the layers according to the order of the layers in the continuous neural tree model, and sort the subgroups according to the order of the subgroups in the layer; Establish a correspondence between layers and subgroups with the same sequence number.

[0010] Preferably, determining whether sensitive information is contained according to the output of each subgroup is specifically: Obtain the results of the k-th subgroups in all layers, and merge the results of all the k-th subgroups; further obtain the merged results of N-2 subgroups, and input the merged results of N-2 subgroups into the first prediction unit to obtain the first prediction result; wherein k is a positive integer, and k≤N-2; The results of all the last N-2 layers are combined and input into the second prediction unit to obtain the second prediction result; Determine whether sensitive information is included based on the first prediction result and the second prediction result.

[0011] Preferably, determining whether sensitive information is included according to the first prediction result and the second prediction result is specifically: If the first prediction result and the second prediction result are the same, determining whether sensitive information is contained according to the first prediction result or the second prediction result; If the first prediction result and the second prediction result are different, a prompt for manual review will be issued.

[0012] Preferably, each node of the tree has a learnable vector; After training the continuous neural tree model, the similarity of the learnable vectors of the nodes at the same position of each tree in the same subset and the similarity of the segmentation threshold θ are determined; the similarity of the nodes at the same position is determined according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ; and the similarity of the two trees is calculated according to the similarity of the nodes at the same position; When the similarity of multiple trees in the same subset is greater than a threshold, the multiple trees are merged; and the merged continuous neural network model is retrained or continued to be trained.

[0013] In addition, the present invention also provides a computer program product, which implements the above method when executed by a processor.

[0014] By dividing each layer of the continuous neural tree, the outputs of different layers are input into different subgroups in the same layer, which prevents interference with information at different levels, retains shallow and deep features, and utilizes the features of the intermediate layer, thereby improving recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of Embodiment 1; Figure 2It is a structural diagram of a continuous neural tree; Figure 3 It is a subgroup structure diagram; Figure 4 It is a structural diagram of the second embodiment. Specific implementation manners

[0016] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element. In the present invention, if the collection of personal privacy data is involved, such as face, mobile phone usage information, etc., personal permission will be obtained in advance, including but not limited to oral reminders, posting posters, mobile phone reminders, etc.; if it involves conflicts with laws and regulations, it will be produced or used within the scope permitted by laws and regulations.

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0018] Embodiment 1, as Figure 1 shown, the present invention provides a sensitive content analysis method based on a continuous neural tree, and the method includes the following steps: S1, after preprocessing the text to be analyzed, use it as the input of the continuous neural tree, obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i-1 subgroups, and establish the corresponding relationship between the subgroups and the first i-1 layers; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree; When analyzing sensitive content, if it is to analyze sensitive information in the operating system, the sensitive information of the operating system in the text to be analyzed includes, but is not limited to, information related to unauthorized access, system log information, etc. Among them, the information related to unauthorized access includes, but is not limited to, source IP, timestamp, resource to be accessed, remote login or local login, operations on resources, reasons for failure, reasons for success, etc. Of course, the text to be analyzed can also have other fields, which are only used as examples here. If it is to detect sensitive information in a text, such as a paragraph, the text is divided into natural sentences, and then word segmentation is performed. The encoding of words in each sentence of the text to be analyzed is used as an input. Among them, preprocessing the text to be analyzed includes removing spaces, removing stop words, vectorization, etc.

[0019] The continuous neural tree model includes a multi-layer structure, such as Figure 2 shown, each layer contains multiple trees. In a specific embodiment, ODTSR trees are used. For the i-th layer in the continuous neural tree model, the trees in it are divided into i - 1 subgroups, and then the corresponding relationship between the subgroups and the previous i - 1 layers is established. For example, in the 3rd layer, there are 8 trees, which are divided into 2 subgroups, with 4 trees in each subgroup. The first subgroup corresponds to the 1st layer of the continuous neural tree model, and the second subgroup corresponds to the 2nd layer of the continuous neural tree model. Among them, i is a positive integer, but 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree. Assume that the continuous neural tree model has a total of 4 non-input and output layers, then N = 4, i = 3, 4. That is, the neural trees in the 3rd layer are divided into 2 subgroups, and the neural trees in the 4th layer are divided into 3 subgroups. The total number of non-input and output layers refers to the total number of layers including the input layer and the output layer in the continuous neural tree model minus 2, that is, the number of layers remaining after subtracting the input layer and the output layer.

[0020] In a more specific embodiment, the establishment of the corresponding relationship between the subgroups and the previous i - 1 layers is specifically as follows: Arrange the layers in the order of the sequence of layers in the continuous neural tree model, and sort the subgroups in the order of their positions in the layer where they are located; For non-output layers, sort the layers in the order from front to back, that is, sort the non-input and output layers in the order of the first layer, the second layer, the third layer, etc., and then sort the subgroups in the order of their positions in the layer where they are located, such as the first subgroup, the second subgroup, etc.

[0021] Establish the corresponding relationship between the layers and subgroups with the same serial number.

[0022] Establish the corresponding relationship between the layers and subgroups in the continuous neural tree model. Preferably, establish the corresponding relationship between the layers and subgroups with the same serial number.

[0023] In order to optimize the continuous neural tree model later, in a specific embodiment, the tree in the i-th layer is divided into i-1 subgroups, specifically: when constructing the continuous neural tree model, the tree structure is set according to the correspondence between the layer and the subgroup. For the subgroup with a small sequence number, the depth of the tree is greater; for the subgroup with a large sequence number, the depth of the tree is smaller, so as to prevent the depth of the tree in the subgroup with a large sequence number from being greater, and its input itself contains the output of the deeper layer in the continuous neural tree model, so as to over-extract the depth information.

[0024] In a preferred embodiment, the structure of the tree of each subgroup is the same, that is, it contains the same nodes and depth, such as Figure 3 shown.

[0025] In another embodiment, each node of the tree has a learnable vector; After training the continuous neural tree model, the similarity of the learnable vectors of the nodes at the same position of each tree in the same subset and the similarity of the segmentation threshold θ are determined; the similarity of the nodes at the same position is determined according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ; and the similarity of the two trees is calculated according to the similarity of the nodes at the same position; In an alternative embodiment, the tree in the continuous neural network model adopts an ODT tree (Oblivious Decision Tree). After training the continuous neural tree model, or after completing several patch trainings, for the same subset, if the structure of the trees is the same, the learnable vectors of each node of the trees are the same or similar, and the segmentation thresholds are the same or similar, then the final results of the two trees are also the same. In order to simplify the structure of the continuous neural tree model, the present invention determines the learnable vectors of nodes at the same position for trees with the same structure in the same subset. and the similarity of the segmentation threshold θ, if the learnable vectors of all nodes at the same position in the two trees are and the segmentation threshold θ are both large, then when the similarity of multiple trees in the same subset is greater than the threshold, the multiple trees are merged; the merged continuous neural network model is retrained or the next patch is trained. Among them, the learnable vector is also called the selection vector.

[0026] The similarity of nodes at the same position is determined based on the similarity of the learnable vector and the similarity of the segmentation threshold θ, specifically: judging whether the similarity of the learnable vector is greater than a first threshold, and whether the similarity of the segmentation threshold θ is greater than a second threshold. If both are greater, the two nodes are considered similar.

[0027] Furthermore, if all nodes in the same position of two trees are similar, then the two trees are considered similar.

[0028] It should be noted that, in one embodiment, the number of layers in the continuous neural tree model and the number and structure of trees in each layer and subgroups are predetermined. During the training process, the number of trees in each layer can be continuously adjusted, for example, by merging trees according to the similarity of trees in subgroups.

[0029] In a more specific embodiment, the process of merging trees is to use the average values ​​of the learnable parameters and the average values ​​of the segmentation thresholds of the same nodes of the two trees to be merged as the learnable parameters and segmentation thresholds of the newly merged tree. This will involve changes in the number of outputs of all subtrees after the subtrees are merged. However, feature fusion, including but not limited to averaging and convolution, is used to make the dimensions before and after the change the same, which will not be repeated here.

[0030] S2, obtaining the output of the layer corresponding to the subgroup based on the corresponding relationship, and using the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup; In one embodiment, the output of the layer corresponding to the subgroup and the input of the continuous neural tree model are used as the input of each tree in the subgroup. The layer corresponding to the subgroup refers to the layer corresponding to the subgroup determined by the corresponding relationship; the layer where the subgroup is located refers to the layer of the continuous neural tree model where the subgroup is located. The present invention can not only obtain shallow features and deep features, but also obtain features of the middle layer, avoiding the extraction of only shallow features and deep features, being able to extract features deeper than the shallow layer, and avoiding too many deep features that are too abstract.

[0031] In an alternative embodiment of S2, after the correspondence between the subgroup and the layer is established, the output of the corresponding layer is used as all or part of the input of the subgroup. Specifically, the output of the layer corresponding to the subgroup and the output of the previous layer of the layer where the subgroup is located are used as the input of each tree in the subgroup.

[0032] S3, obtaining the output of each subgroup in the continuous neural tree model, and determining whether sensitive information is contained according to the output of each subgroup, and if so, issuing a prompt.

[0033] In a specific embodiment, determining whether sensitive information is contained according to the output of each subgroup is specifically as follows: Obtain the results of the k-th subgroups in all layers, and merge the results of all the k-th subgroups; further obtain the merged results of N-2 subgroups, and input the merged results of N-2 subgroups into the first prediction unit to obtain the first prediction result; wherein k is a positive integer, and k≤N-2; For example, there are four layers. The third layer has two subgroups A31 and A32, and the fourth layer has three subgroups A41, A42, and A43. The first subgroup of all layers is A31 and A41 respectively, the second subgroups of all layers are distributed as A32 and A42, and the third subgroups of all layers are distributed as A43. Combine A31 and A41 to get the combined result 1, combine A32 and A42 to get the combined result 2. Since A43 has only one subgroup, the combined result remains itself, that is, the combined result 3. Input the combined results 1, 2, and 3 into the first prediction unit. Among them, all layers refer to the layers containing subgroups.

[0034] Combine the results of the last N - 2 all layers and input them into the second prediction unit to obtain the second prediction result; The results output by the last N - 2 all layers are input into the second prediction unit to obtain the second prediction result.

[0035] Then, determine whether it contains sensitive information according to the first prediction result and the second prediction result.

[0036] Preferably, the determination of whether it contains sensitive information according to the first prediction result and the second prediction result is specifically as follows: If the first prediction result and the second prediction result are the same, then determine whether it contains sensitive information according to the first prediction result or the second prediction result; If the first prediction result and the second prediction result are different, then issue a prompt for manual review.

[0037] Embodiment 2, the present invention also provides a sensitive content analysis system based on a continuous neural tree, as Figure 4 shown, the system includes the following modules: A division module, which is used to take the text to be analyzed after preprocessing as the input of the continuous neural tree, obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i - 1 subgroups, and establish the corresponding relationship between the subgroups and the previous i - 1 layers; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree; An input determination module, which is used to obtain the output of the layer corresponding to the subgroup based on the corresponding relationship, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup; A result output module, which is used to obtain the output of each subgroup in the continuous neural tree model, determine whether it contains sensitive information according to the output of each subgroup, and if it contains, issue a prompt.

[0038] Preferably, the establishment of the corresponding relationship between the subgroups and the previous i - 1 layers is specifically as follows: Arrange the layers in the order of the layers in the continuous neural tree model, and sort the subgroups in the order of the subgroups in the layer where they are located; Establish a correspondence between layers and subgroups with the same sequence number.

[0039] Preferably, determining whether sensitive information is contained according to the output of each subgroup is specifically: Obtain the results of the k-th subgroups in all layers, and merge the results of all the k-th subgroups; further obtain the merged results of N-2 subgroups, and input the merged results of N-2 subgroups into the first prediction unit to obtain the first prediction result; wherein k is a positive integer, and k≤N-2; The results of all the last N-2 layers are combined and input into the second prediction unit to obtain the second prediction result; Determine whether sensitive information is included based on the first prediction result and the second prediction result.

[0040] Preferably, determining whether sensitive information is included according to the first prediction result and the second prediction result is specifically: If the first prediction result and the second prediction result are the same, determining whether sensitive information is contained according to the first prediction result or the second prediction result; If the first prediction result and the second prediction result are different, a prompt for manual review will be issued.

[0041] Preferably, each node of the tree has a learnable vector; After training the continuous neural tree model, the similarity of the learnable vectors of the nodes at the same position of each tree in the same subset and the similarity of the segmentation threshold θ are determined; the similarity of the nodes at the same position is determined according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ; and the similarity of the two trees is calculated according to the similarity of the nodes at the same position; When the similarity of multiple trees in the same subset is greater than a threshold, the multiple trees are merged; and the merged continuous neural network model is retrained or continued to be trained.

[0042] Embodiment 3: The present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described above is implemented.

[0043] Through the description of the above implementation methods, technicians in this field can clearly understand that each implementation method can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a computer product, and the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it, and other embodiments may also be used. Although the present invention has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A sensitive content analysis method based on a continuous neural tree, characterized in that: The method includes the following steps: The text to be analyzed is preprocessed and used as the input of the continuous neural tree. The number of trees in the i-th layer of the continuous neural tree model is obtained, the trees in the i-th layer are divided into i-1 subgroups, and the corresponding relationship between the subgroups and the previous i-1 layers is established; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree; Based on the corresponding relationship, the output of the layer corresponding to the subgroup is obtained, and the output of the layer corresponding to the subgroup and the input of the continuous neural tree model are used as the input of each tree in the subgroup; The output of each subgroup in the continuous neural tree model is obtained, and it is determined whether sensitive information is included according to the output of each subgroup. If it is included, a prompt is issued.

2. The method according to claim 1, characterized in that The establishment of the corresponding relationship between the subgroup and the previous i-1 layers is specifically as follows: The layers are arranged in the order of the layers in the continuous neural tree model, and the subgroups are sorted in the order of the subgroups in their respective layers; The layers and subgroups with the same serial number are established with a corresponding relationship.

3. The method according to claim 1, characterized in that The determination of whether sensitive information is included according to the output of each subgroup is specifically as follows: The results of the k-th subgroup in all layers are obtained, and the results of all the k-th subgroups are merged; the merged results of N-2 subgroups are obtained, and the merged results of N-2 subgroups are input into the first prediction unit to obtain the first prediction result; where k is a positive integer and k ≤ N-2; The results of the last N-2 layers of all layers are merged and input into the second prediction unit to obtain the second prediction result; It is determined whether sensitive information is included according to the first prediction result and the second prediction result.

4. The method according to claim 3, characterized in that The determination of whether sensitive information is included according to the first prediction result and the second prediction result is specifically as follows: If the first prediction result and the second prediction result are the same, it is determined whether sensitive information is included according to the first prediction result or the second prediction result; If the first prediction result and the second prediction result are different, a prompt for manual review is issued.

5. The method according to claim 1, characterized in that Each node of the tree has a learnable vector; After training the continuous neural tree model, the similarity between the learnable vectors of the nodes at the same position of each tree in the same subset and the similarity of the segmentation threshold θ are judged; the similarity of the nodes at the same position is determined according to the similarity of the learnable vectors and the similarity of the segmentation threshold θ; furthermore, the similarity of two trees is calculated according to the similarity of the nodes at the same position; When the similarity of multiple trees in the same subset is greater than the threshold, the multiple trees are merged; the merged continuous neural network model is retrained or continued to be trained.

6. A sensitive content analysis system based on a continuous neural tree, characterized in that: The system includes the following modules: A division module, which is used to preprocess the text to be analyzed and use it as the input of the continuous neural tree, obtain the number of trees in the i-th layer of the continuous neural tree model, divide the trees in the i-th layer into i-1 subgroups, and establish the corresponding relationship between the subgroups and the previous i-1 layers; where i is a positive integer, 2 < i ≤ N, and N is the total number of non-input and output layers of the continuous neural tree; An input determination module, which is used to obtain the output of the layer corresponding to the subgroup based on the corresponding relationship, and use the output of the layer corresponding to the subgroup and the input of the continuous neural tree model as the input of each tree in the subgroup; The result output module is used to obtain the output of each subgroup in the continuous neural tree model, and determine whether sensitive information is contained according to the output of each subgroup. If so, a prompt is issued.

7. The system according to claim 6, characterized in that The corresponding relationship between the subgroup and the first i-1 layers is established as follows: Arrange the layers according to the order of the layers in the continuous neural tree model, and sort the subgroups according to the order of the subgroups in the layer; Establish a correspondence between layers and subgroups with the same sequence number.

8. The system according to claim 6, characterized in that The method of determining whether sensitive information is contained according to the output of each subgroup is specifically as follows: Obtain the results of the k-th subgroups in all layers, merge the results of all the k-th subgroups; obtain the merged results of N-2 subgroups, and input the merged results of N-2 subgroups into the first prediction unit to obtain the first prediction result; wherein k is a positive integer, and k≤N-2; The results of all the last N-2 layers are combined and input into the second prediction unit to obtain the second prediction result; Determine whether sensitive information is included based on the first prediction result and the second prediction result.

9. The system according to claim 8, characterized in that The determining whether sensitive information is included according to the first prediction result and the second prediction result is specifically: If the first prediction result and the second prediction result are the same, determining whether sensitive information is contained according to the first prediction result or the second prediction result; If the first prediction result and the second prediction result are different, a prompt for manual review will be issued.

10. A computer storage device, wherein a computer program is stored on the storage device, characterized in that: When the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Production parameter optimization prediction method and device, equipment and storage medium

    CN108647808A

  • Bamboo chip wormhole and mildew spot detection method based on convolutional flexible neural forest

    CN111242895A

  • Long-tail distribution image recognition method based on hierarchical learning

    CN111738303A

  • Business information evaluation method and system for high-dimensional data

    CN114782078A

  • Malicious URL detection method based on hybrid binary neural tree

    CN116186251A