A method, device and electronic device for detecting secret labels
By extracting the context in the text and using a pre-trained dense mark detection model for identification, the problem of inability to effectively detect dense marks in the prior art is solved in the case of document parsing exceptions or short context sentences, and the accuracy of dense level detection is improved.
Patent Information
- Application Number
- CN202111196100.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-14
AI Technical Summary
The prior art cannot effectively perform dense mark recognition when document parsing content is abnormal or context sentences are short, resulting in insufficient accuracy of dense level detection.
By extracting the context based on the preset secret code keyword and inputting it into the pre-trained secret code detection model, it is determined whether the context contains a secret code, thereby determining the secret level of the text to be detected.
Even when the document parsing content is abnormal or the context sentence is short, this method can accurately identify whether the context contains a dense mark, improving the accuracy of the dense level detection.
Smart Images

Figure CN113918973B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information security technology. Specifically, it relates to a method, device, and electronic device for detecting security labels. Background Art
[0002] A security label is a mark on an electronic document regarding the degree of secrecy, which can prevent the leakage of electronic documents, classify electronic documents with different degrees of secrecy, add marks to different categories of electronic documents through technical means, and enable the electronic documents to obtain corresponding levels of security protection. In order to monitor whether important classified documents appear in public spaces to prevent secrecy leakage, it is necessary to detect security labels for a large number of documents.
[0003] In the prior art, based on the detection of a standard security label template, the security label detection is regarded as a true word error detection problem, and an N-gram model is used to train the context sentences containing security label keywords. If the security label word is irrelevant to the context, it can be determined as a security label.
[0004] However, using the method of the prior art, effective security label recognition cannot be performed in the case of abnormal document parsing content or short context sentences. Summary of the Invention
[0005] An object of this application is to provide a method for detecting security labels to solve the problems that security labels cannot be effectively detected, the calculation and recognition ability is poor when the document parsing content is abnormal, and supervised learning is required to detect security labels well.
[0006] To achieve the above object, the technical solutions adopted in the embodiments of this application are as follows:
[0007] In a first aspect, this application provides a method for detecting security labels, including:
[0008] Extracting at least one context from the text to be detected based on a preset security label keyword;
[0009] Sequentially inputting the at least one context into a pre-trained security label detection model to obtain a security label detection result corresponding to each context, where the security label detection result includes a security label or not a security label;
[0010] If the security label detection result corresponding to at least one target context in each context is a security label, determining the security level of the text to be detected according to the at least one target context.
[0011] Optionally, the determining the security level of the text to be detected according to the at least one target context includes:
[0012] Respectively obtaining the security label keywords in each of the target contexts;
[0013] Determine the security label level corresponding to each of the target contexts according to the security label keywords in the respective target contexts;
[0014] Determine the security level of the text to be detected according to the security label level corresponding to each of the target contexts.
[0015] Optionally, determining the security level of the text to be detected according to the security label level corresponding to each of the target contexts includes:
[0016] Take the highest security label level among the security label levels corresponding to each of the target contexts as the security level of the text to be detected.
[0017] Optionally, before inputting the at least one context into the pre-trained security label detection model in sequence to obtain the security label detection results corresponding to each context, it further includes:
[0018] Obtain initial model parameters based on a preset task;
[0019] Construct an initial detection model based on the initial model parameters;
[0020] Train the initial detection model using pre-labeled training samples to obtain the security label detection model.
[0021] Optionally, the preset task is obtained based on a plurality of unlabeled training samples.
[0022] Optionally, the training the initial detection model using pre-labeled training samples to obtain the security label detection model includes:
[0023] Encode the training samples to obtain an encoded vector sequence;
[0024] Input the encoded vector sequence into the input layer of the initial detection model, and after being processed by multiple network layers of the initial detection model, perform encoding to obtain a predicted label encoding;
[0025] Based on a preset loss function and the predicted label encoding, correct the model parameters of the initial detection model to obtain the security label detection model.
[0026] Optionally, the extracting at least one context from the text to be detected based on the preset security label keywords includes:
[0027] Preprocess the text to be detected, and the preprocessing includes: removing preset characters from the text to be detected;
[0028] According to the characteristics of the classified keyword, the left and right context word characteristics, and the time sequence position characteristics, at least one context in the text to be detected is obtained. The length of each context is within a preset length range, includes the classified keyword, and the classified keyword is located at the center position of the context.
[0029] In a second aspect, the present application provides a classified label detection device, including:
[0030] An extraction module, configured to extract at least one context from the text to be detected based on a preset classified keyword;
[0031] A processing module, configured to sequentially input the at least one context into a pre-trained classified label detection model to obtain a classified label detection result corresponding to each context, where the classified label detection result includes a classified label or a non-classified label;
[0032] A determination module, configured to, if the classified label detection result corresponding to at least one target context in each context is a classified label, determine the classification level of the text to be detected according to the at least one target context.
[0033] Optionally, the determination module is specifically configured to:
[0034] Obtain the classified keywords in each of the target contexts respectively;
[0035] Determine the classified label level corresponding to each of the target contexts according to the classified keywords in each of the target contexts;
[0036] Determine the classification level of the text to be detected according to the classified label levels corresponding to each of the target contexts.
[0037] Optionally, the determination module is specifically configured to:
[0038] Use the highest classified label level among the classified label levels corresponding to each of the target contexts as the classification level of the text to be detected.
[0039] Optionally, the device further includes:
[0040] A first training module, configured to obtain initial model parameters based on a preset task;
[0041] A construction module, configured to construct an initial detection model based on the initial model parameters;
[0042] A second training module, configured to train the initial detection model using pre-labeled training samples to obtain the classified label detection model.
[0043] Optionally, the preset task is obtained based on a plurality of unlabeled training samples.
[0044] Optionally, the second training module is specifically configured to:
[0045] Encode the training samples to obtain an encoded vector sequence;
[0046] Input the encoded vector sequence into the input layer of the initial detection model, and after being processed by multiple network layers of the initial detection model, perform encoding to obtain a predicted label encoding;
[0047] Based on a preset loss function and the predicted label encoding, correct the model parameters of the initial detection model to obtain the secret label detection model.
[0048] Optionally, the extraction module is specifically configured to:
[0049] Preprocess the text to be detected, where the preprocessing includes: removing preset characters from the text to be detected;
[0050] Obtain at least one context in the text to be detected, where the length of each context is within a preset length range, and includes the secret label keyword, and the secret label keyword is located at the center position of the context.
[0051] In a third aspect, the present application provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the above-mentioned secret label detection method when executed.
[0052] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of the above-mentioned secret label detection method.
[0053] The beneficial effects of the present application are:
[0054] Extract at least one context from the text to be detected. The pre-trained secret label detection model can determine whether each context contains a secret label. If the secret label detection result of at least one target context in each context is a secret label, the secret level of the text to be detected can be further determined through the secret label detection result of the target context. Since the secret label detection model is pre-trained, even in the case of abnormal document parsing content or short context sentences, using this secret label detection model can still accurately identify whether the context contains a secret label, and thus can accurately detect the secret level of the document. Therefore, it can greatly improve the accuracy of secret level detection. Description of the Drawings
[0055] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0056] Figure 1 is a schematic flowchart of the secret label detection method provided by the embodiments of the present application;
[0057] Figure 2 is another schematic flowchart of the secret label detection method provided by the embodiments of the present application;
[0058] Figure 3 is yet another schematic flowchart of the secret label detection method provided by the embodiments of the present application;
[0059] Figure 4 is still another schematic flowchart of the secret label detection method provided by the embodiments of the present application;
[0060] Figure 5 is one of the schematic flowcharts of the secret label detection method provided by the embodiments of the present application;
[0061] Figure 6 is a module structure diagram of the secret label detection device provided by the embodiments of the present application;
[0062] Figure 7 is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the protection scope of the present application. Additionally, it should be understood that the schematic drawings are not drawn to the actual scale. The flowcharts used in the present application show the operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context relationships can be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0064] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application usually described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0065] Optionally, the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated thereafter, but does not exclude adding other features.
[0066] In the prior art, based on the detection of the standard secret label template, the secret label detection is regarded as a true word error detection problem, and the N-gram model is used to train the context sentences containing the secret label keywords. If the secret label word is irrelevant to the context, it can be judged as a secret label. However, the accuracy of the N-gram model may fluctuate with different texts to be detected. When the text to be detected is parsed and transferred, it is very likely that the positions of the characters are disordered, and the recognition and positions of the line break character "\n" and the tab character "\t" are disordered. In this case, the content parsed from the document is abnormal. If the method of the prior art is used, it will cause certain errors in the detection results. For example, when the content parsed from the document is "commercial\nsecret", there will be errors in the detection results of the document confidentiality level.
[0067] In addition, when the context sentences of the document are short, even if there are words similar to secret labels, but they are irrelevant in the short context, the N-gram model cannot effectively detect whether this word is a secret label; and for titles, book names, and film and television works names containing secret label keywords, etc., the secret labels cannot be accurately identified directly by the methods of the prior art.
[0068] To solve the above problems, the present application provides an inventive concept: regarding the recognition of secret labels as a text classification task for a piece of context containing secret label keywords, and detecting the secret labels in the context through a pre-trained secret label detection model, so that even in the case of abnormal content parsed from the document or short context sentences, the secret labels can still be effectively recognized, improving the security of the secret documents.
[0069] Figure 1 It is a schematic flowchart of a secret label detection method provided by the present application. The execution subject of this method can be an electronic device with computing and processing capabilities, such as: servers, desktop computers, laptop computers, etc. As Figure 1 shown, this method includes:
[0070] S101. Extract at least one context from the text to be detected based on the preset classified keyword.
[0071] Optionally, the character range value of the context on both sides of the classified keyword for a certain distance is a hyperparameter. Exemplarily, the context character range can be set to 10 characters, and the context character range can be preset. Exemplarily, the above 10 indicates that the context character range is 10 characters.
[0072] Exemplarily, first, the preset classified keywords include: the 3 keywords of "secret", "confidential", and "top secret". Those without the above 3 keywords are not classified. Secondly, after parsing documents such as official document tables, there may be misalignment of text positions, as well as misrecognition and misalignment of line break characters "\n" and tab characters "\t". After data analysis of the classified context, the line break symbol "\n" is an important classified feature. Here, the top 10 of the statistical results of the classified context word frequency are counted, as shown in Table 1. The top 10 of the statistical results of the non-classified context word frequency are shown in Table 2.
[0073] word context word frequency percentage disposal following text 3.76% message following text 3.71% garden following text 2.41% playing card preceding text 1.58% medical preceding text 1.58% a preceding text 1.54% base following text 0.66% task following text 0.45% residence following text 0.39% organization following text 0.36%
[0074] Table 1
[0075]
[0076]
[0077] Table 2
[0078] The text to be detected is a document such as an official document table that needs to be detected for classification. From the text to be detected, according to the position of the classified keyword, extract the context of the characters at a certain distance on both sides centered on the classified keyword, and at least one context can be obtained.
[0079] S102. Input at least one context into the pre-trained classified detection model in sequence, and obtain the classified detection results corresponding to each context. The classified detection results include classified or non-classified.
[0080] The above classified detection model can be a neural network model. The specific processes of training and using the classified detection model will be described in detail in the following embodiments.
[0081] Optionally, input each context into the above classified detection model in a certain order, for example, in the order of the context in the document to be detected. Correspondingly, the classified detection model can output the classified detection results corresponding to the context. After all contexts are input into the classified detection model, the classified detection results of all contexts can be obtained.
[0082] For each context, the obtained sensitive label detection result may be a sensitive label or a non-sensitive label. Optionally, when the sensitive label detection results of all contexts are non-sensitive labels, it indicates that the text to be detected does not contain sensitive labels, and the result that the text to be detected does not contain sensitive labels can be output and the process ends, without performing subsequent processing steps. When the sensitive label detection results of some or all of the contexts are sensitive labels, it indicates that the text to be detected is a document containing sensitive labels, and subsequent steps are continued to determine the classification level of the text to be detected.
[0083] S103. If the sensitive label detection result corresponding to at least one target context in each context is a sensitive label, determine the classification level of the text to be detected according to the at least one target context.
[0084] Using the above sensitive label detection model, the sensitive label detection result of each context can be obtained. The sensitive label detection results of each context may be the same, partially the same, or may be different from each other. The electronic device can finally determine the classification level of the text to be detected based on the sensitive label detection results of each context.
[0085] Exemplarily, the classification levels of the text to be detected may include: secret, confidential, top secret. And the priority of the classification levels can be: top secret > confidential > secret.
[0086] Exemplarily, the sensitive label styles are shown in Table 3.
[0087]
[0088] Table 3
[0089] In this embodiment, at least one context is extracted from the text to be detected, and the pre-trained sensitive label detection model can determine whether each context contains a sensitive label. If the sensitive label detection result of at least one target context in each context is a sensitive label, the classification level of the text to be detected can be further determined through the sensitive label detection result of the target context. Since the sensitive label detection model is pre-trained, even in the case of abnormal document parsing content or short context sentences, using this sensitive label detection model can still accurately identify whether the context contains a sensitive label, and thus can accurately detect the classification level of the document. Therefore, the accuracy of classification level detection can be greatly improved.
[0090] Figure 2 Another process schematic diagram of the sensitive label detection method provided by this application is as Figure 2 shown. An optional manner of the above step S103 includes:
[0091] S201. Obtain the sensitive label keywords in each target context respectively.
[0092] Optionally, the encrypted keyword in each target context can be obtained based on the processing method in step S501 described below, which will not be elaborated here.
[0093] Exemplarily, search for encrypted keywords in the text to be detected, such as "secret", "confidential", and "top secret".
[0094] S202. Determine the encrypted level corresponding to each target context according to the encrypted keywords in each target context.
[0095] Optionally, the encrypted level of the text to be detected can be judged according to the encrypted keywords in the text to be detected.
[0096] Exemplarily, when the encrypted keyword in the text to be detected is "secret", the encrypted level of this text to be detected is: secret. In other embodiments, the encrypted level can also be in other forms, which are not limited here.
[0097] Exemplarily, when the text to be detected is "a secret residence", it is assumed that the encrypted keyword of this text to be detected is "secret", and the encrypted level of this text to be detected is: non-encrypted.
[0098] S203. Determine the classification level of the text to be detected according to the encrypted level corresponding to each target context.
[0099] Optionally, the classification level of the text to be detected can be determined by the following method:
[0100] Take the highest encrypted level among the encrypted levels corresponding to each target context as the classification level of the text to be detected. Optionally, the encrypted level of the text to be detected is the highest level of the encrypted level that has appeared. Exemplarily, when "secret" and "confidential" appear in the text to be detected, the encrypted level of the text to be detected is "confidential". When "secret" and "top secret" appear in the text to be detected, the encrypted level of the text to be detected is "top secret". Preferably, in actual operation, the encrypted level can include but is not limited to "secret", "confidential", "top secret", and can also be in other forms, which are not limited here.
[0101] The training process of the aforementioned encrypted detection model will be described below.
[0102] Figure 3 Another process schematic diagram of the encrypted detection method provided by this application, as Figure 3 shown, before the above step S102, there is also an optional method, including:
[0103] S301. Obtain initial model parameters based on a preset task training.
[0104] Optionally, the parameters of the initial model are not randomly initialized. Instead, a set of model parameters are first obtained by training based on a preset task. The language model can predict the next word according to the context without the need for manually labeled corpus, and can learn rich semantic knowledge from an unlimited large-scale monolingual corpus, which is convenient to use and saves time costs.
[0105] S302. Construct an initial detection model based on the initial model parameters.
[0106] Optionally, a task can be constructed in advance through a large amount of unlabeled data to train and obtain a set of initial model parameters, and the initial detection model is constructed using the initial model parameters. For most natural language tasks, constructing a large-scale labeled dataset is a huge challenge because the annotation cost is very high, especially for tasks related to grammar and semantics. However, large-scale unlabeled corpora are relatively easy to construct. Therefore, in this step, a task is first constructed through a large amount of unlabeled data and the initial model parameters are obtained by training with this task. The initial model parameters are not random parameters like those in ordinary model initialization. Using the initial model parameters, pre-training can provide better model initialization, obtain better generalization performance, and accelerate the convergence speed of the subsequent initial detection model. S303. Use the pre-labeled training samples to train the initial detection model to obtain a classified label detection model.
[0107] Optionally, pre-training can be regarded as a kind of regularization to avoid overfitting to small data.
[0108] Optionally, the preset task is obtained based on multiple unlabeled training samples.
[0109] In this embodiment, a set of model parameters are first obtained by training based on a preset task, and the language model can predict the next word according to the context; then a task is constructed in advance through a large amount of unlabeled data to train and obtain a set of initial model parameters, and the initial detection model is constructed using the initial model parameters; finally, the pre-labeled training samples are used to train the initial detection model, and finally a classified label detection model is obtained. Therefore, the obtained classified label detection model has good generalization performance, can accelerate the convergence speed of the target task, has good fitting performance, and is convenient for subsequent accurate detection of classified labels.
[0110] Optionally, Figure 4 is another process schematic diagram of the classified label detection method provided by this application. As Figure 4 shown, an optional manner of the above step S303 includes:
[0111] S401. Encode the training samples to obtain an encoded vector sequence w = (w0, w1,..., w i ).
[0112] Exemplarily, the sequence sample is segmented to obtain sub - sequences of the segmentation. Then, the characters, words, and positions are encoded according to the dictionary to obtain the encoded sequence \(t=(c_0,c_1,\cdots,c i ,w_0,w_1,\cdots,w i ,p_0,p_1,\cdots,p i ). It can effectively count the word frequencies before and after the keyword of the secret label, as well as the position of the keyword of the secret label, which is convenient for establishing a secret label detection model later.
[0113] S402. Input the encoded vector sequence into the input layer of the initial detection model, and after being processed by multiple network layers of the initial detection model, it is encoded to obtain the predicted label encoding.
[0114] Optionally, input the encoded \(t\) vector obtained from the encoded sequence into the input layer of the initial detection model. After being processed by multiple layers of the Transformer network layer, the contextual embedding that captures context information is obtained. Then, the predicted label is encoded into a one - hot vector for subsequent encoding. Exemplarily, the secret label is \([1,0]\), and the non - secret label is \([0,1]\). Finally, the predicted label encoding is obtained.
[0115] S403. Based on the preset loss function and the predicted label encoding, correct the model parameters of the initial detection model to obtain the secret label detection model.
[0116] Optionally, based on the preset loss function and the predicted label encoding, use cross - entropy as the loss value for backpropagation to correct the parameters of the initial detection model, and finally obtain the secret label detection model.
[0117] Optionally, Figure 5 is a schematic flow diagram of one of the secret label detection methods provided by this application. As Figure 5 shown, an optional way of the above - mentioned step S101 includes:
[0118] S501. Pre - process the text to be detected, and the pre - processing includes: removing the preset characters in the text to be detected.
[0119] Exemplarily, because when parsing documents such as official document tables, the positions of characters may be disordered, generating a large number of preset characters, which will interfere with the subsequent secret label detection of the text. By pre - processing the text, the content of the detected text can be normalized, which is convenient for extracting the context of characters at a certain distance on the left and right of the secret label keyword later.
[0120] S502. Obtain at least one context in the text to be detected according to the secret label keyword features, the left and right context word features, and the temporal position features. The length of each context is within a preset length range, includes the secret label keyword, and the secret label keyword is located at the center position of the context.
[0121] Optionally, perform secret label determination according to the keyword features, the left and right context word features, and the temporal position features. The keyword features are: the three secret label keywords of "secret", "confidential", and "top secret". The left and right context word features are: the context sub-features of a certain distance to the left and right centered on the secret label keyword, and the range value of the left and right distances is a hyperparameter, which can be set according to requirements. The temporal position feature is to perform position encoding on the characters in the string containing the secret label keyword in the order from left to right as the position feature.
[0122] Exemplarily, the preset characters can be blank characters and special characters. Optionally, the special characters are real words or symbols that can be pasted into the text, including but not limited to: mathematical symbols, arrow symbols, unit symbols, tabulation characters, symbol patterns, Greek letters, Russian letters, Chinese pinyin, Chinese characters, weather symbols, chess symbols, currency symbols, etc.
[0123] Optionally, when the context length is an odd number n, the secret label keyword is located at (n±1) / 2; when the context length is an even number n, the secret label keyword is located at n / 2±1. Exemplarily, when the context length is 11, the secret label keyword is located at 5 and 6. When the context length is 10, the secret label keyword is located at 5 and 6.
[0124] Exemplarily, compared with the prior art, the detection results of the embodiments of the present application are shown in Table 4:
[0125]
[0126]
[0127] Table 4
[0128] Figure 6 For the module structure diagram of the secret label detection device provided by the embodiments of the present application, as Figure 6 shown, the device includes:
[0129] An extraction module 601, configured to extract at least one context from the text to be detected based on a preset secret label keyword;
[0130] A processing module 602, configured to sequentially input at least one context into a pre-trained secret label detection model to obtain the secret label detection results corresponding to each context, and the secret label detection results include secret label or non-secret label;
[0131] A determination module 603, configured to, if the detection result of the ciphertext label corresponding to at least one target context among each context is a ciphertext label, determine the classification level of the text to be detected according to the at least one target context.
[0132] Further, the determination module 603 is specifically configured to:
[0133] Obtain the ciphertext keywords in each target context respectively;
[0134] Determine the ciphertext label level corresponding to each target context according to the ciphertext keywords in each target context;
[0135] Determine the classification level of the text to be detected according to the ciphertext label levels corresponding to each target context.
[0136] Still further, the determination module 603 is specifically configured to:
[0137] Use the highest ciphertext label level among the ciphertext label levels corresponding to each target context as the classification level of the text to be detected.
[0138] In this embodiment, the apparatus further includes:
[0139] A first training module, configured to obtain initial model parameters based on a preset task;
[0140] A construction module, configured to construct an initial detection model based on the initial model parameters;
[0141] A second training module, configured to train the initial detection model using pre-labeled training samples to obtain a ciphertext label detection model.
[0142] Further, the preset task is obtained based on a plurality of unlabeled training samples.
[0143] Still further, the second training module is specifically configured to:
[0144] Encode the training samples to obtain an encoded vector sequence;
[0145] Input the encoded vector sequence into the input layer of the initial detection model, and after being processed by multiple network layers of the initial detection model, perform encoding to obtain a predicted label encoding;
[0146] Based on a preset loss function and the predicted label encoding, correct the model parameters of the initial detection model to obtain a ciphertext label detection model.
[0147] In this embodiment, the extraction module 601 is specifically configured to:
[0148] Perform preprocessing on the text to be detected, where the preprocessing includes: removing preset characters in the text to be detected;
[0149] Obtain at least one context in the text to be detected. The length of each context is within a preset length range, and it contains a sensitive keyword, and the sensitive keyword is located at the center position of the context.
[0150] An embodiment of the present application further provides an electronic device 70, as Figure 7 shown, which is a schematic structural diagram of the electronic device 70 provided by the present application, including: a processor 71, a storage medium 72, and a bus 73. The storage medium 72 stores machine-readable instructions executable by the processor 71. When the electronic device 70 runs, the processor 71 communicates with the storage medium 72 through the bus 73, and the processor 71 executes the machine-readable instructions to perform the steps of the above-mentioned sensitive label detection method when executed.
[0151] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it performs the steps of the above-mentioned sensitive label detection method.
[0152] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A secret label detection method, characterized in that, Including: Extracting at least one context from the text to be detected based on preset classified keyword; Sequentially inputting the at least one context into a pre-trained classified detection model to obtain a classified detection result corresponding to each context, where the classified detection result includes classified or non-classified; If the classified detection result corresponding to at least one target context in each context is classified, determining the classification level of the text to be detected according to the at least one target context; The extracting at least one context from the text to be detected based on preset classified keyword includes: Performing preprocessing on the text to be detected, where the preprocessing includes removing preset characters from the text to be detected; Obtaining the at least one context in the text to be detected according to classified keyword features, left and right context word features, and temporal position features, where the length of each context is within a preset length range, includes the classified keyword, and the classified keyword is located at the center position of the context.
2. The method according to claim 1, wherein The determining the classification level of the text to be detected according to the at least one target context includes: Respectively obtaining the classified keywords in each target context; Determining the classified level corresponding to each target context according to the classified keywords in each target context; Determining the classification level of the text to be detected according to the classified levels corresponding to each target context.
3. The method according to claim 2, wherein Determining the classification level of the text to be detected according to the classified levels corresponding to each target context includes: Taking the highest classified level among the classified levels corresponding to each target context as the classification level of the text to be detected.
4. The method according to any one of claims 1 to 3, characterized in that Before sequentially inputting the at least one context into a pre-trained classified detection model to obtain a classified detection result corresponding to each context, further including: Obtaining initial model parameters based on a preset task; Constructing an initial detection model based on the initial model parameters; Training the initial detection model using pre-labeled training samples to obtain the classified detection model.
5. The method according to claim 4, wherein The preset task is obtained based on multiple unlabeled training samples.
6. The method according to claim 4, wherein The training the initial detection model using pre-labeled training samples to obtain the classified detection model includes: Encoding the training samples to obtain an encoded vector sequence; Inputting the encoded vector sequence into the input layer of the initial detection model, and encoding it after being processed by multiple network layers of the initial detection model to obtain a predicted label encoding; Correcting the model parameters of the initial detection model based on a preset loss function and the predicted label encoding to obtain the classified detection model.
7. A secret label detection device, characterized in that, Including: An extraction module for extracting at least one context from the text to be detected based on preset classified keyword; A processing module for sequentially inputting the at least one context into a pre-trained classified detection model to obtain a classified detection result corresponding to each context, where the classified detection result includes classified or non-classified; A determination module for, if the classified detection result corresponding to at least one target context in each context is classified, determining the classification level of the text to be detected according to the at least one target context; The extraction module is specifically configured to: Preprocess the text to be detected, where the preprocessing includes: removing preset characters from the text to be detected; Obtain the at least one context in the text to be detected according to the ciphertext keyword feature, the left and right context word features, and the temporal position feature, where the length of each context is within a preset length range, and includes the ciphertext keyword, and the ciphertext keyword is located at the center position of the context.
8. An electronic device, characterized in that, Comprising: A processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the ciphertext detection method according to any one of claims 1 to 6 when executed.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it performs the steps of the ciphertext detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Official document judgment method and judgment system based on named entities
CN111626057A
Method and device for automatically determining secret grade of secret-related text
CN112347779A