An unsupervised dense point labeling and auxiliary dense labeling method

By using unsupervised key point marking and assisted classification methods, the problem of inaccurate manual classification in existing technologies has been solved, realizing automated and scientific management of document classification levels and improving classification efficiency and accuracy.

CN115481429BActive Publication Date: 2026-02-10BEIJING JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210935013.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2026-02-10
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

In the current technology, the determination of document classification mainly relies on manual search and review, which has problems such as inaccurate classification, low classification efficiency, waste of resources and high risk of leakage, especially in classified documents where it is difficult to accurately mark the classified points.

Method used

We employ unsupervised dense point annotation and assisted dense point determination methods. By establishing a corpus statistical database, calculating the confidence of words and sentences, we construct a dense point vocabulary and sentence database. We then utilize a multi-feature fusion algorithm for dense point annotation, thereby improving annotation efficiency and accuracy.

Benefits of technology

It has automated and made the marking of classified information more scientific, reduced the randomness and subjectivity of classification, improved the accuracy and efficiency of classification, and met the needs of digital security work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481429B_ABST
    Figure CN115481429B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of unsupervised secret point marking and auxiliary secret marking method, comprising the following steps: 1) before training process, respectively establish corpus statistics library for different secret types;2) from corpus statistics library, calculate word confidence using algorithm, different types are sorted according to secret level confidence, and secret point word library is constructed;3) from corpus statistics library, extract secret point sentence of different secret level in the document of fixed secret using secret point sentence confidence evaluation method of multi-feature fusion, and construct secret point sentence library;4) secret point marking is carried out to the document of to-be-fixed secret using the secret point word library and secret point sentence library constructed;5) according to the secret marking result of to-be-marked document, it is included in corresponding category, and the record of relevant word in word statistics library is updated.The method improves the efficiency and accuracy of secret point marking, and effectively avoids the randomness and subjectivity of secret marking through auxiliary secret marking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic document security technology, specifically to an unsupervised method for marking secret points and assisting in determining confidentiality. Background Technology

[0002] The classification of state secrets is fundamental and fundamental to confidentiality management, and its importance is self-evident. How to achieve accurate classification is a pressing issue in current confidentiality work. With the continuous advancement of information technology and the ongoing improvement of e-government systems, various government agencies have built their own office automation systems, and some have even implemented paperless office practices. The forms in which state secrets are generated and stored have undergone significant changes. While digitalization has improved the efficiency of confidentiality work at all levels, it has also brought new challenges in all aspects.

[0003] Currently, the determination of document classification levels is mostly done manually, relying solely on the confidentiality knowledge and experience of the personnel involved. This inevitably leads to subjective judgments, resulting in inaccurate classifications, difficulties in applying classification standards, and the inability to pass on classification experience. On the one hand, due to the wide variety of classified matters, manual determination results in long cycles and low efficiency. On the other hand, the key to classifying a matter lies in its distinguishing features. However, the current method of classifying documents involves labeling the entire document. What is formally designated as "one" state secret by an organization may actually contain multiple secrets, or only a very small portion of the content may be classified as state secrets. In such cases, failing to clearly identify distinguishing features and simply managing documents as "one" state secret often leads to layers of derivation and classification, an excessive number of state secrets, wasted management resources, and increased risk of leaks. Determining the distinguishing features of state secrets involves identifying the truly essential attributes of state secrets—the critical and minimal information that, if leaked, would harm national security and interests—to provide a basis for classification. Therefore, based on manual classification, using computer technology to assist in classification to achieve accurate classification and marking of classified information, and to realize the standardization, scientification and digitalization of classification, is an urgent need for current confidentiality work. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the purpose of this invention is to provide an unsupervised dense point annotation and auxiliary dense point determination method, which improves the efficiency and accuracy of dense point annotation, and effectively avoids the randomness and subjectivity of dense point determination through auxiliary dense point determination.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] An unsupervised method for dense point annotation and auxiliary density determination, characterized by comprising the following steps:

[0007] Step 1: Establish separate corpus statistics databases for different secret types;

[0008] Step 2: Calculate word confidence based on the corpus statistics database obtained in Step 1, sort the confidence by security level according to different types, and construct a dense word database;

[0009] Step 3: Based on the corpus statistics database obtained in Step 1, use the multi-feature fusion dense sentence confidence evaluation algorithm to extract dense sentences of different security levels from the already classified documents and construct a dense sentence database.

[0010] Step 4: Use the key point vocabulary library built in Step 2 and the key point sentence library built in Step 3 to mark the key points in the document to be marked.

[0011] Step 5: Based on the confidentiality determination results of the documents to be annotated obtained in Step 4, include them in the corresponding categories and update the records of relevant words in the word statistics database.

[0012] Based on the above scheme, the algorithm for calculating word confidence based on the expected statistical database obtained in step 1 in step 2 is an improved SS3 algorithm. The confidence of word w belonging to category c is calculated through the function gv(w, c). The calculation of gv involves three functions, defined as follows (1):

[0013] gv(w, c) = lv σ (w,c)·sg λ (w,c)·sn ρ (w,c) (1);

[0014] In the above formula, gv(w, c) is the confidence level that word w belongs to category c;

[0015] lv σ (w, c) is a value assigned to a word based on the local probability of word w in category c. By defining the intra-class distribution coefficient and improving the local probability, the influence of intra-class distribution on word classification discrimination and the calculation bias caused by the differences between texts are considered. The specific definition is as shown in equation (2):

[0016]

[0017] Where, n c n represents the total number of texts in category c. w,C This represents the number of texts in category c that contain the word w, d c,j The j-th text in category c, where W is the set of all words, w i ∈W, and These represent the number of texts in category c containing the words with the most and fewest text counts, respectively. The word 'w' is in the text 'd'c,j Frequency of occurrence in It is text d c,j The number of words containing the word with the highest frequency.

[0018] sg λ (w, c) is used to represent the importance of word w to category c, when lv σ (w, c) is significantly larger than most other categories c. i lv σ (w, c) i When ), its output is close to 1; when for all categories c i lv σ (w, c) i When all the elements are close to each other, it outputs a value close to 0, as defined in equation (3):

[0019]

[0020] Among them, LV w ={lv σ (w, c) i )|c i ∈C}, that is, the set of all local values ​​of the word w; LV w , the median; That is, LV w The absolute median difference; For hyperparameters;

[0021] sn ρ (w, c) is a constraint function used to measure the uniqueness of the importance of word w to category c. It is related to the number of categories to which word w is important and will reduce the global value of word w proportionally. The specific definition is as shown in equation (4):

[0022]

[0023] in, That is, all categories of sg in C except c. λ (w, c) i The sum of ) This is a hyperparameter.

[0024] Based on the above scheme, the multi-feature fusion confidence evaluation algorithm for key sentences in step 3 is based on the multi-feature fusion confidence evaluation algorithm for key sentences. It evaluates the confidence of key sentences from three perspectives: key word features, position features, and summary word features. It is used to extract key sentences of different security levels in classified documents, accumulate experience with key sentences with complete semantic expression, and promptly incorporate newly encountered and newly generated state secrets of government agencies.

[0025] The calculation of the confidence score of the dense sentence mainly considers three features: dense word features, positional features, and summary word features. For each sentence in the document, the feature scores on the above three features are calculated, and the weighted sum is used as the confidence score of the dense sentence, as shown in the following formula (5):

[0026] CScore(s i )=γ1×classification(s i )+γ2×position(s i )+γ3×summary(s i (5);

[0027] Among them, s i Let represent the i-th sentence of text d, then d = {s1, s2, ..., s} |d|}, classification(s i ) represents sentence s i The dense word feature score, position(s) i ) represents sentence s i Location feature score, summary(s) i ) represents sentence s i The summary word feature scores are given, where γ1, γ2, and γ3 are real hyperparameters greater than 0, and γ1+γ2+γ3=1.

[0028] Building upon the aforementioned approach, the dense word feature is the most crucial characteristic for assessing whether a sentence is classified and, if so, at what level. The more words in a sentence possessing high dense word confidence, the greater the likelihood that the sentence is classified. This method utilizes the ISS3 algorithm to obtain the word... The score for the dense word feature of a sentence is defined as shown in equation (6):

[0029]

[0030] Where, n i,j Sentence s i In category c j The total number of words, m i For sentence s i The total number of words in the text, |s| represents the total number of sentences, w i,k Sentence s i The kth word, gv(w i,k c j ) indicates the word w i,k In category c j The gv value, The j-th component represents sentence s i In the corresponding category c j GV value, Representing vectors L1 norm, For vectors The maximum component value of GMAX is recorded, and the security level corresponding to GMAX is recorded as the sentence s. i The security classification label.

[0031] Considering positional characteristics, in classified documents, it is often impossible for the entire text to be classified, and it is even possible that a large portion of the text is not classified. The beginning and end of a document usually contain summary key content, making them more prone to containing classified sentences. Therefore, these parts should be assigned higher feature scores. This method defines the positional feature score of a sentence as shown in equation (7):

[0032]

[0033] Where i represents s i Let ||d| be the i-th sentence in text d, where ||d| represents the total number of sentences in text d. Initial position(s) i The value of position(s) decreases as the value of i increases. When i grows to half the total number of sentences, the position(s) value decreases. i The value of position(s) drops to its minimum value. As the value of i continues to increase, the position(s) value decreases. i The value increases, ensuring that the closer a sentence is to the beginning or end of the text, the higher its positional feature score.

[0034] Considering the characteristics of summary words, sentences introduced by summary words generally summarize important information. In classified documents, key points are very likely to be found in sentences that are summary-oriented and targeted. Therefore, this method incorporates the characteristics of summary words into the sentence confidence calculation, specifically defined as in equation (8):

[0035]

[0036] SList is a summary vocabulary for sentence s. i The words are traversed, and when the sentence contains a summary word from the summary word list, the feature score is 1, otherwise it is 0;

[0037] The above summary words include: so, therefore, in short, in general, and in conclusion.

[0038] Based on the above scheme, step 4, which involves marking the key points in the document to be classified as confidential, is as follows:

[0039] Step 4-1: Read in the document that needs to be annotated with dense dots;

[0040] Step 4-2: Use the jieba word segmentation tool to segment the above document content into individual words, remove stop words, and divide it into a series of word sets;

[0041] Step 4-3: Compare all the word sets generated in Step 4-2 with the dense word database, highlight all matching words in the original document, and display their gv values ​​as confidence scores;

[0042] Step 4-4: Determine the recommended security level of the document to be annotated based on the highest security level among the matching words marked in Step 4-3. Perform summation operations layer by layer to finally obtain the gv of the entire document. The category corresponding to the maximum component value is the classification result.

[0043] The unsupervised dense point annotation and auxiliary dense point determination method described in this invention has the following advantages:

[0044] Currently, document classification is mostly determined manually, relying on the confidentiality knowledge and experience of the personnel, which inevitably leads to subjective judgments, resulting in inaccurate classifications, difficulty in determining classification standards, and the inability to pass on classification experience. To address the new demands for confidentiality point annotation in the field of classification, and considering the lack of a public confidentiality point database and the fact that classified documents often lack confidentiality point annotations, this invention trains and constructs a confidentiality point lexicon and sentence database for already classified documents. It calculates the confidence scores of words and sentences and obtains the gv value of the entire document through aggregation operators to achieve the classification result. Alternatively, it can be combined with any auxiliary classification algorithm to provide a recommended classification level. This method improves the efficiency and accuracy of confidentiality point annotation and effectively avoids the randomness and subjectivity of classification through auxiliary classification. Attached Figure Description

[0045] The present invention includes the following figures:

[0046] Figure 1 A flowchart illustrating an unsupervised dense point annotation and auxiliary dense point determination method according to the present invention;

[0047] Figure 2 A schematic diagram illustrating the process of constructing a dense dot lexicon based on the TSS3 algorithm in this invention;

[0048] Figure 3 This invention provides a flowchart illustrating the process of constructing a dense sentence library through multi-feature fusion.

[0049] Figure 4 A schematic diagram illustrating the process of marking secret points in a document to be classified according to the present invention; Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] Example 1

[0052] Due to the sensitivity of classified documents, the results are presented here through an analogy experiment. Since news texts use more standardized language and closely resemble the requirements of real-world official document writing, the publicly available news dataset THUCNews, released by Tsinghua University, was selected for the analogy experiment. A simulated classification dataset for the education sector was also constructed. Three similar categories—finance, stocks, lottery, and real estate—were selected and classified as Top Secret, Confidential, Secret, and Internal, respectively. Education was classified as Public. A total of 24,000 articles were extracted.

[0053] like Figure 1 , Figure 2 , Figure 3 , Figure 4 As shown, an unsupervised method for dense point annotation and auxiliary density determination includes the following steps:

[0054] Step 1: Before the training process, establish a corpus statistics database for different types of secrets. Specifically, the secret levels include five categories: top secret, confidential, secret, internal, and public.

[0055] Step 2: Calculate word confidence scores from the corpus using an algorithm, sort the confidence scores by different types according to their density level, and construct a dense word database. Specifically:

[0056] The improved SS3 algorithm is used to calculate word confidence. The confidence of word w belonging to category c is calculated by the function gv(w,c), specifically referring to equations (1)-(4) mentioned above. Some calculation results are as follows:

[0057] Fund (0, 0.572, 0.161, 0, 0)

[0058] Futures (0, 0.236, 0.457, 0.183, 0)

[0059] Exam (0, 0, 0, 0.161, 0.362)

[0060] Assumption: Category c1 contains three texts d1, d2, and d3, containing three words w1, w2, and w3. σ = 1

[0061]

[0062] According to equation (2), we can calculate:

[0063]

[0064]

[0065]

[0066] Suppose that the lv values ​​of word w1 in the five classes are 0.1, 0.2, 0.3, 0.4, and 0.5, respectively, with λ = 1 and ρ = 1.

[0067] According to equation (3), we can calculate:

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075] According to equation (4), we can calculate:

[0076]

[0077]

[0078]

[0079]

[0080]

[0081] According to equation (1), we can calculate:

[0082] gv(w1, c1)=0.1×0×0.504=0

[0083] gv(w1, c2)=0.2×0×0.504=0

[0084] gv(w1, c3)=0.3×0.002×0.505=0

[0085] gv(w1, c4)=0.4×0.982×0.750=0.295

[0086] gv(w1,c5)=0.5×1×0.754=0.377

[0087] Then we can obtain

[0088] In the above calculation process, the set of classified documents is first preprocessed, which consists of three steps: word segmentation, stop word removal, and statistics. The existing set of classified documents is converted from text format to the data format input to the ISS3 algorithm and recorded in the word statistics table VSTable in the database. The ISS3 algorithm is then used to calculate the word count for each word. Each component represents the confidence level of the word belonging to the corresponding category. In this application scenario... This is a four-dimensional vector, where each dimension represents the confidence level of a word w belonging to one of the five security levels: Top Secret, Confidential, Secret, Internal, and Public. The confidence levels for all words are then obtained. Then, it is updated and stored in the GVTable. When marking secret words in documents to be classified, words in each GV component that are greater than the set threshold can be selected as marked secret words according to the pre-set threshold.

[0089] Step 3: Extract the dense sentences of different security levels from the corpus statistics database using the multi-feature fusion dense sentence confidence assessment method, and construct the dense sentence database;

[0090] In this embodiment, the classified document is segmented into sentences. Periods, exclamation marks, and question marks are used as sentence delimiters. A multi-feature fusion-based confidence assessment method for key sentences is used to calculate the sentence confidence for each segmented sentence. The multi-features refer to key word features, positional features, and summary word features, where the key word feature score is measured by each word stored in the key word library. Based on the sentence confidence calculated for each sentence and a pre-set threshold, a core layer and an outer layer of the key sentence library are constructed. The core layer is the most representative set of key sentences, and the outer layer is a set of substitute key sentences. After the classified document is segmented into sentences, a sentence set is formed. On the one hand, the multi-feature fusion-based confidence assessment method for key sentences is used to label key sentences; on the other hand, the semantic similarity is compared with the key sentence set in the key sentence library to label key sentences, and the labeled key sentences are placed in the key sentence library labeling layer. Specifically, referring to equations (5)-(8) above, some calculation results are as follows:

[0091] Two funds have been approved for issuance (0.583 Class 2).

[0092] Assessed by an appraisal agency qualified to conduct securities and futures related appraisal business (0.594 Category 3)

[0093] Registration for the April self-study exams is underway (0.634 Category 5)

[0094] Specifically, assuming the first sentence s of the document is "two funds were approved for issuance", after word segmentation, we can get: two / funds / approved / issuance.

[0095] For ease of calculation, assume that their gv vectors are as follows: two (0, 0, 0, 0, 0), fund (0, 0.5, 0.4, 0, 0), approved (0, 0, 0, 0, 0), issued (0, 0.2, 0.1, 0, 0), the dense word threshold is set to 0.05, the ratio of the largest dense word count to the total number of words in all sentences in each category is 2 / 3, and γ1, γ2, and γ3 are set to 0.8, 0.1, and 0.1, respectively.

[0096] According to equation (6), we can calculate:

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103] And record the security classification label as c2.

[0104] Since this is the first sentence and there are no concluding words, we can calculate the following based on equations (7) and (8):

[0105] position(s)=1,summary(s)=0

[0106] According to formula (5), the confidence level of the dense sentence can be calculated as follows:

[0107] CScoer(s)=0.8×0.583+0.1×1+0.1×0=0.5664

[0108] Step 4: Using the key point vocabulary database built in Step 2 and the key point sentence database built in Step 3, key point annotation is performed on the document to be identified as keyed. Specifically:

[0109] The process involves: reading in the document to be annotated with key points; using the jieba word segmentation tool to segment the document content into individual words and removing stop words to form a series of word sets; comparing all generated word sets with the key point vocabulary; highlighting all matching words in the original document and displaying their gv values ​​as confidence levels; determining the recommended key level of the document to be annotated based on the highest key level among the matching words; employing max pooling layer-by-layer operations to finally obtain the category corresponding to the maximum component value of the entire document, which is the classification result.

[0110] Step 5: Based on the classification results of the documents to be annotated, include them in the corresponding categories, update the records of relevant words in the word statistics database, and construct a sensitive word database for classified fields with incremental learning characteristics that can continuously learn and dynamically update. Specifically:

[0111] By combining the classification results of the auxiliary classification algorithm, the classification personnel determine the classification level of the document. After the number of documents accumulates to a certain amount, they are used as input to the word statistics database to update the word records. This eliminates the need to store all documents or to retrain from scratch every time a new training document is added, thus enabling incremental and continuous learning, and constantly learning new classification points from newly classified documents.

[0112] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

Claims

1. An unsupervised method for dense point labeling and auxiliary density determination, characterized in that, Includes the following steps: Step 1: Establish separate corpus statistics databases for different secret types; Step 2: Calculate word confidence based on the corpus statistics database obtained in Step 1, sort the confidence by security level according to different types, and construct a dense word database; Step 3: Based on the corpus statistics database obtained in Step 1, use the multi-feature fusion dense sentence confidence evaluation algorithm to extract dense sentences of different security levels from the documents with defined security levels, and construct a dense sentence database. Step 4: Use the key point vocabulary library built in Step 2 and the key point sentence library built in Step 3 to mark the key points in the document to be marked. Step 5: Based on the confidentiality assessment results of the documents to be annotated obtained in Step 4, include them in the corresponding categories and update the records of relevant words in the word statistics database. The confidence evaluation algorithm for dense sentence fusion based on multi-feature fusion described in step 3 is as follows (5): (5); in, Representing text The If there are 10 sentences, then Sentence The dense word feature score, Sentence Location feature score, Sentence Summary of word feature scores, For real hyperparameters greater than 0, and ; In the formula (5) The following equations are respectively: (6), (7), and (8): ; ; ; in, Sentence In category The total number of words, For sentences The total number of words in the Chinese dictionary. Indicates the total number of sentences. Sentence The One word, Words In category of value, The Each component represents a sentence. In the corresponding category of value, Representing vectors of Norm, For vectors The maximum component value is recorded simultaneously. The corresponding security level is used as a sentence Security classification label; ; in, express It is text The One sentence. Representing text The total number of sentences in the text; initial Value with The value increases and decreases, when When the number of sentences increases to half the total number of sentences, The value drops to its lowest value, as The value continues to increase, The value increases, ensuring that the closer a sentence is to the beginning or end of the text, the higher its positional feature score. ; ; in, To summarize the vocabulary, for sentences The words are traversed, and when the sentence contains a summary word from the summary word list, the feature score is 1, otherwise it is 0; The above summary words include: so, therefore, in short, in general, and in conclusion.

2. The unsupervised dense point labeling and auxiliary density determination method as described in claim 1, characterized in that: The algorithm for calculating word confidence based on the expected statistical database obtained in step 1, as described in step 2, is an improved SS3 algorithm, as shown in equation (1): ; In the above formula For words w Exclusive to Category c Confidence level; It assigns values ​​to words based on their local probabilities within a category. By defining the intra-class distribution coefficient and improving the local probability, it considers the impact of intra-class distribution on word classification discrimination, as well as the calculation bias caused by differences between texts. Used to indicate the importance of a word to a category; Used to measure words Category Important uniqueness; The following equations are respectively: (2), (3), and (4): ; in, Indicates category The total number of texts in the text. Indicates category Contains words The number of texts, category The first in This text, It is the collection of all words. and Representing categories The number of texts containing the words with the most and fewest texts. It is a word In the text Frequency of occurrence in It is text The number of words containing the word with the highest frequency. ; in, words The set of all local values; express the median; ,Right now The absolute median difference; For hyperparameters; ; in, ,Right now Except In addition, all categories The sum; This is a hyperparameter.

3. The unsupervised dense point labeling and auxiliary density determination method as described in claim 1, characterized in that: Step 4, which involves marking the key points in the document to be classified as confidential, includes the following steps: Step 4-1: Read in the document that needs to be annotated with dense dots; Step 4-2: Use the jieba word segmentation tool to segment the above document content into individual words, remove stop words, and divide it into a series of word sets; Step 4-3: Compare all the word sets generated in Step 4-2 with the dense word database, highlight all matching words in the original document, and display their... The value is used as the confidence level; Step 4-4: Determine the recommended security level of the document to be annotated based on the highest security level among the matched words marked in Step 4-3. Use the aggregation operator to perform layer-by-layer calculations to finally obtain the security level of the entire document. The category corresponding to the largest component value is the classification result.

Citation Information

Patent Citations

  • Unsupervised English sentence automatic simplification algorithm

    CN110096705A