Text geographic information security inspection method and system based on knowledge base

By establishing a knowledge base for geographic information inspection and a deep attention text matching model CW-DATM that integrates words, the problems of low efficiency and poor accuracy of geographic information confidentiality inspection in the prior art are solved, and efficient and accurate security detection of long texts is achieved.

CN115934871BActive Publication Date: 2025-08-15WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211498723.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-08-15
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The existing geographic information text confidentiality inspection methods have cumbersome inspection steps, wide coverage, and weak technical support. They cannot efficiently and accurately inspect confidential geographical information, lack professional geographic information keyword databases and confidential judgment rules. The deep text matching model has a great impact on the noise of long texts, and it is impossible to deeply explore fine-grained semantic information.

Method used

Establish a knowledge base for geographic information inspection, including a geographic information keyword database and a confidential judgment rule database, combine the deep attention text matching model CW-DATM with fusion words, obtain deep context semantic information through Bi-LSTM and multi-head self-attention mechanism, and use the multi-head interactive attention mechanism to extract the interactive semantic features between text pairs for security detection.

Benefits of technology

It realizes more comprehensive and professional text geographic information inspection, improves the accuracy of text semantic matching, and can efficiently detect confidential information in long texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115934871B_ABST
    Figure CN115934871B_ABST
Patent Text Reader

Abstract

This invention discloses a knowledge-base-based text geographic information security inspection method and system. First, using a geographic information inspection knowledge base, a set of files to be inspected for geographic information confidentiality is retrieved, along with the text containing the keywords within the files. Next, a text pair is generated using a deep word-integrated attention text matching model (CW-DATM), including the text to be inspected and a confidentiality determination rule. Finally, the CW-DATM is used to detect the security of the text. This invention introduces a geographic information inspection knowledge base, combines multi-granularity semantic features within the text, and interactive semantic features between texts, improving the accuracy of semantic matching performed by the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of geographic information confidentiality inspection and natural language processing, and relates to a text geographic information security inspection method and system, and specifically to a text geographic information security inspection method and system based on a knowledge base. Background Art

[0002] Geographic information is a critical strategic resource, of extraordinary significance to national defense, economic development, and social development. With the rapid development of information technology, the widespread use of geographic information and the highly specialized nature of geographic information technology have posed significant challenges to confidentiality management and prevention efforts. my country has always attached great importance to the security and confidentiality of geographic information, issuing legal norms and technical standards for the confidentiality management of surveying and mapping results, providing timely policy support for the security and confidentiality of geographic information and public services. Although geographic information confidentiality policies are continuously being refined, confidentiality inspections still face numerous shortcomings. The inspection procedures are cumbersome, the scope is wide, and the technical support is weak, making it difficult to efficiently and accurately identify confidential geographic information. Currently, confidentiality inspections for textual geographic information primarily rely on confidential keyword detection, semantic similarity calculation, regular expression matching, and machine learning methods to determine the security of text. However, these methods are incomplete and unavailable for many types of geographic information, resulting in a high rate of false positives and missed detections.

[0003] Text matching is a critical task in natural language processing. Traditional text matching methods require significant human effort, yet can only extract a small number of effective features. These models struggle to deeply explore the underlying semantics of the text and cannot accurately model text semantic similarity. In recent years, deep learning models have been widely used in text matching tasks, primarily encompassing three categories: deep text matching models based on single-semantic document representations, such as DSSM, CSSM, and LSTM-RNN, which only capture local information that is effective for text matching; deep text matching models based on multi-semantic document representations, such as uRAE, MV-LSTM, and MultiGranCNN, which comprehensively consider both local and global information in the text but struggle to capture structural information in the matching; and deep text matching models that directly model matching patterns, such as ARC-II, DeepMatch, and Match-SRNN, which directly capture both the degree of matching and the structure of the matching, but require a large amount of supervised text matching data to train the models, resulting in high resource consumption.

[0004] In summary, the existing methods for checking the confidentiality of geographic information texts have the following main problems: (1) Geographic information involves multiple industry sectors, with rich data types and different standards and specifications. There is a lack of a professional geographic information keyword library to check sensitive geographic information in texts; (2) The text confidentiality checking method is single, and the legal norms of surveying and mapping geographic information are not fully utilized. The corresponding confidentiality judgment rules are not designed according to the confidentiality characteristics of different types of geographic information, and there is a lack of the introduction of external knowledge; (3) Regular files involving geographic information usually contain large sections of text, while deep text matching models can only process texts of moderate length. Texts that are too long will bring a lot of irrelevant noise, affecting the mining of fine-grained semantic information in the text. Summary of the Invention

[0005] The present invention aims to solve the problems existing in the existing geographic information text confidentiality inspection technology and provide a text geographic information security inspection method and system based on a knowledge base.

[0006] The technical solution adopted by the method of the present invention is: a text geographic information security inspection method based on a knowledge base, comprising the following steps:

[0007] Step 1: For the set of files to be checked for geographic information confidentiality, use the geographic information check knowledge base to obtain files that match geographic information keywords and the text containing the keywords in the files;

[0008] The geographic information inspection knowledge base is composed of a geographic information keyword library and a confidentiality judgment rule library;

[0009] The geographic information keyword database includes a multi-industry geographic information vocabulary database, a Chinese place name and address database, and a sensitive geographic information vocabulary database;

[0010] The multi-industry geographic information vocabulary is constructed by collecting professional vocabulary related to geographic information based on the surveying and mapping standard system and product standards and specifications of surveying and mapping, remote sensing, navigation, planning, environment, public security, and transportation industries;

[0011] The Chinese place name and address database includes place name data, address data and point of interest data;

[0012] The sensitive geographic information vocabulary includes information on prisons, material storage depots, energy facilities, satellite observation stations, detention centers, industrial facilities, and radar stations;

[0013] The confidentiality judgment rule base is based on the three aspects of geographic information data confidentiality, map review and public application of basic geographic information, and studies the confidentiality characteristics of various geographic spatiotemporal data. According to the different characteristics of the data, corresponding confidentiality judgment rules are designed to form a rule base R = {R1, R2, ..., R N}, and each judgment rule R b There is a corresponding keyword list keyword_R b ;

[0014] Step 2: Generate text pairs for the word-integrated deep attention text matching model CW-DATM, including the text to be checked and the confidentiality judgment rules;

[0015] Step 3: Detect whether the text is safe by integrating the word-based deep attention text matching model CW-DATM.

[0016] The technical solution adopted by the system of the present invention is: a text geographic information security inspection system based on a knowledge base, including the following modules:

[0017] A file set preparation and knowledge base construction module is used to obtain files that match geographic information keywords and texts containing keywords in the files using the geographic information check knowledge base for the file set to be checked for geographic information confidentiality;

[0018] The geographic information inspection knowledge base is composed of a geographic information keyword library and a confidentiality judgment rule library;

[0019] The geographic information keyword database includes a multi-industry geographic information vocabulary database, a Chinese place name and address database, and a sensitive geographic information vocabulary database;

[0020] The multi-industry geographic information vocabulary is constructed by collecting professional vocabulary related to geographic information based on the surveying and mapping standard system and product standards and specifications of surveying and mapping, remote sensing, navigation, planning, environment, public security, and transportation industries;

[0021] The Chinese place name and address database includes place name data, address data and point of interest data;

[0022] The sensitive geographic information vocabulary includes information on prisons, material storage depots, energy facilities, satellite observation stations, detention centers, industrial facilities, and radar stations;

[0023] The confidentiality judgment rule base is based on the three aspects of geographic information data confidentiality, map review and public application of basic geographic information, and studies the confidentiality characteristics of various geographic spatiotemporal data. According to the different characteristics of the data, corresponding confidentiality judgment rules are designed to form a rule base R = {R1, R2, ..., R N}, and each judgment rule R b There is a corresponding keyword list keyword_R b ;

[0024] The confidentiality check text pair generation module is used to generate text pairs for the word-integrated deep attention text matching model CW-DATM, including the text to be checked and the confidentiality judgment rules;

[0025] The text matching and security detection module is used to detect whether the text is safe by integrating the deep attention text matching model CW-DATM that integrates words.

[0026] Compared with the existing method for checking the confidentiality of geographic information text, the beneficial effects of the present invention are: establishing a geographic information inspection knowledge base, including a geographic information keyword library and a confidentiality judgment rule library, which serves as the basis for keyword inspection and text confidentiality analysis, and can check text geographic information more comprehensively and professionally; designing a deep attention text matching model CW-DATM that integrates words, extracts the word vectors and word vectors of the text respectively and fuses them, obtains deep contextual semantic information through the Bi-LSTM model and multi-head self-attention mechanism, and then uses the multi-head interactive attention mechanism to extract the interactive semantic features between text pairs, and obtains the security judgment result of the text after passing through the fusion layer, pooling layer and prediction layer in sequence. The model has achieved good results in the text semantic matching task. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0028] Figure 2 This is a framework diagram of a geographic information inspection knowledge base according to an embodiment of the present invention;

[0029] Figure 3 It is a structural diagram of the word-fused deep attention text matching model CW-DATM according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to facilitate ordinary technicians in this field to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0031] Aiming at the problems existing in the existing geographic information text confidentiality inspection technology, the present invention introduces a geographic information inspection knowledge base and combines it with the deep attention text matching model CW-DATM that integrates words to deeply explore the fine-grained semantic features of the text and inspect the security of the text.

[0032] Please see Figure 1 The present invention provides a text geographic information security inspection method based on a knowledge base, comprising the following steps:

[0033] Step 1: For the set of files to be checked for geographic information confidentiality, use the geographic information check knowledge base to obtain files that match geographic information keywords and the text containing the keywords in the files;

[0034] The set of files to be checked for geographic information confidentiality in this embodiment includes Office files, PDF files, and text files containing geographic information. This invention aims to check whether conventional geographic information files contain confidential text. There are eight file extensions: doc, docx, xls, xlsx, ppt, pptx, pdf, and txt.

[0035] Please see Figure 2 ,The geographic information inspection knowledge base of this embodiment is composed of a geographic information keyword library and a confidentiality judgment rule library;

[0036] The geographic information keyword library of this embodiment consists of three parts: a multi-industry geographic information vocabulary, a Chinese place name and address database, and a sensitive geographic information vocabulary. The specific contents of the three vocabulary libraries are as follows:

[0037] (1) Multi-industry geographic information vocabulary: Geographic spatiotemporal data involves a wide range of industries, with different product standards and rich data types. The keywords involved in checking various types of geographic data are also different. We sorted out the "Surveying and Mapping Standard System" and the product standards and specifications of various industries such as surveying and mapping, remote sensing, navigation, planning, environment, public security, and transportation, and used the professional vocabulary (terms and abbreviations) related to geographic information to build a vocabulary.

[0038] (2) China Place Name and Address Database: To address the semantic inspection problem of fuzzy spatial scope of geographic spatiotemporal data, the national Tiandi Map and navigation software are used to collect place name data, address data and point of interest data to establish a national place name and address database.

[0039] (3) Sensitive geographic information lexicon: Sensitive facility geographic information is the focus of public security. The lexicon is composed of information such as prisons, material storage depots, energy facilities, satellite observation stations, detention centers, industrial facilities, and radar stations.

[0040] The confidentiality judgment rule base of this embodiment is based on the fact that different types of geographic spatiotemporal data have different specific confidentiality standards. Therefore, in comparison with four legal normative documents on surveying and mapping geographic information (see Table 1), starting from the three aspects of geographic information data confidentiality determination, map review and application of basic geographic information disclosure, the confidentiality characteristics (resolution, positioning accuracy, spatial range, geometric information, etc.) of various geographic spatiotemporal data are studied. According to the different characteristics of the data, corresponding confidentiality judgment rules are designed to form a rule base R = {R1, R2, ..., R N}, and each judgment rule R bThere is a corresponding keyword list keyword_R b ; In this embodiment, N takes the value of 63.

[0041] Table 1 Legal norms for surveying and mapping geographic information

[0042]

[0043]

[0044] This embodiment collects files that hit geographic information keywords and texts containing keywords in the files; the specific implementation steps are:

[0045] (1) Obtain files whose names contain keywords in the geographic information keyword library, and obtain a file set F = {F1, F2, ..., F f}.

[0046] (2) With the help of file parsing tools, the file contents are extracted, the text containing keywords and their context information are extracted, and the sentence set S = {S1, S2, ..., S s}, and each sentence S a There is a corresponding keyword list keyword_S a .

[0047] Step 2: Generate text pairs for the word-integrated deep attention text matching model CW-DATM, including the text to be checked and the confidentiality judgment rules;

[0048] In this embodiment, for each sentence set S in the file set F, each text to be checked S a (a=1,2,…,s) and r confidentiality judgment rules R b (b=1,2,…,r) forms a text pair (S a ,R b ) Input CW-DATM. The condition for forming a text pair is that S a With R b Keyword list keyword_S a with keyword_R b Table 2 gives an example of an input text pair.

[0049] Table 2 Input text pair examples

[0050]

[0051] Please see Figure 3 The deep attention text matching model CW-DATM of the word fusion in this embodiment consists of an embedding layer, an encoding layer, an interaction layer, a fusion layer, a pooling layer and a prediction layer;

[0052] The embedding layer in this embodiment obtains the word vectors and token vectors of two input texts and calculates the word-token fusion vector. The specific method is as follows:

[0053] (1) Input the text S to be inspected a and the classified information judgment rule R b into the Siamese network respectively, and use RoBERTa and Word2Vec to obtain the word vector and token vector representations of the text respectively.

[0054] (2) According to the different tokenization granularities, add all the word vectors contained in the word and take the average to obtain a new word vector, which is concatenated with the original word vector as the model input. Take a word vector w a in the text S i as an example. It consists of z word vectors c j and obtains the vector x i after word-token fusion. Then the input of the model is X = {x1, x2,..., x n}, 1 < i ≤ n, where n is the total number of word vectors in the text S a . The rule R b is also processed in the same way. The calculation formula for the word-token fusion vector x i is:

[0055]

[0056] The encoding layer in this embodiment obtains deep context semantic information through the Bi-LSTM model and the multi-head self-attention mechanism. The specific method is as follows:

[0057] (1) Use the bidirectional long short-term memory network Bi-LSTM to encode the model input X, and capture the long-distance dependence relationship of the text sequence from both positive and negative directions. The word-token fusion vector x i obtains the hidden layer representation containing context information through the forward LSTM and obtains the hidden layer representation containing context information through the reverse LSTM Concatenate the two vectors as the output of the Bi-LSTM model

[0058] (2) Introduce the multi-head self-attention mechanism to extract the deep features of the text, assign different weights to the output of the Bi-LSTM, and capture multiple features of semantically related information. For H = {h1, h2,..., h n}, combine Q e , K e and V e to obtain the calculation result H A of the multi-head self-attention weight:

[0059]

[0060]

[0061] in, and is the weight matrix, d k is the matrix dimension, A_head j (j=1,2,6…,h A ) is a single-head attention unit, h A is the number of splicing, E O is the weight matrix used for linear transformation.

[0062] The interaction layer of this embodiment uses the multi-head interactive attention mechanism to extract two input texts S a and R b The word interaction features between and Perform interaction calculations to capture the semantic dependencies between text pairs. The text interaction matrix H C The formula is:

[0063]

[0064]

[0065] in, and is the weight matrix, d k is the matrix dimension, C_head l (l=1,2,…,h C ) is a single-head attention unit, h C is the number of splicing, U O is the weight matrix used for linear transformation.

[0066] The fusion layer of this embodiment fuses the semantic information obtained by the encoding layer and the interaction layer to form a text S a For example, for H A1 、H C And the matrix obtained by multiplying the two is used for residual connection and layer normalization, G m is a weight matrix (m=1,2,3), then the fusion feature H O The calculation formula is:

[0067]

[0068] H m =GeLU(H m G m ) (7)

[0069]

[0070] The pooling layer of this embodiment combines the maximum pooling and average pooling methods to extract the key information of the text fusion representation. a and R b Fusion features and Perform pooling calculations to obtain P1 and P2 respectively. The formulas are as follows:

[0071]

[0072]

[0073] The prediction layer of this embodiment concatenates the vectors P1 and P2 obtained by the pooling operation and inputs them into a two-layer fully connected neural network. The text S is calculated through the sigmoid activation function. a and R b The semantic similarity of Thus predicting the text S a Is it safe? The calculation formula is:

[0074]

[0075] FC1(P)=ReLU(α1·P+β1) (12)

[0076] FC2(P)=ReLU(α2·P+β2) (13)

[0077] Among them, α1 and α2 are weight vectors, and β1 and β2 are biases.

[0078] Assume the text similarity threshold is 0.5, when When the text to be checked S a and confidentiality judgment rule R b The semantics of S a It can be considered safe, otherwise, S a It is unsafe.

[0079] Step 3: Detect whether the text is safe by integrating the word-based deep attention text matching model CW-DATM.

[0080] This embodiment uses the text matching results to determine the security of the file and outputs a security analysis report. The specific implementation steps are:

[0081] (1) Analyze the security of text by combining multiple tags. For the sentence set S of the file, each text S to be checked a There are r text pairs (Sa ,R b ), after the text matching model CW-DATM detects whether the text is safe, it outputs r predicted labels, where the value is 1 for safe text and 0 for unsafe text. a If there is a label of 1 in the corresponding r predicted labels, then S a Considered safe, S a _label=1;if S a The corresponding r predicted labels are all 0, then S a Can be considered unsafe, S a _label=0.

[0082] (2) Analyze the security of files by combining multiple labels. For each file in the file set F, the corresponding sentence set S contains s predicted labels S. a _label. If there is S a _label is 1, the file can be considered safe. If there are s predicted labels S a If all _labels are 0, the file can be considered unsafe.

[0083] (3) Output the security analysis results of the file, including the file name and the confidentiality of the file ("safe" or "unsafe"). If the file is unsafe, it is also necessary to output the text containing confidential information, as well as the corresponding keywords and confidentiality judgment rules.

[0084] The deep attention text matching model CW-DATM of the fusion word in this embodiment is a trained deep attention text matching model CW-DATM of the fusion word; the prediction value of the text semantic similarity in this embodiment is The binary cross entropy loss function is used as the objective function. To prevent overfitting, L2 regularization is added to the loss function, and the model is optimized using the Adam algorithm. The loss function is calculated as follows:

[0085]

[0086] Among them, N represents the total number of samples, y i Represents the true label of the sample, λ is the regularization parameter, and w represents the weight coefficient of the L2 regularization term.

[0087] The present invention introduces a geographic information inspection knowledge base and combines the multi-granularity semantic features within the text with the interactive semantic features between texts to improve the accuracy of the semantic matching performed by the model.

[0088] It should be understood that the above description of the preferred embodiment is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.

Claims

1. A text geographic information security inspection method based on a knowledge base, characterized in that: The following steps are involved: Step 1: For the set of files to be checked for geographic information confidentiality, use the geographic information check knowledge base to obtain files that match geographic information keywords and the text containing the keywords in the files; The geographic information inspection knowledge base is composed of a geographic information keyword library and a confidentiality judgment rule library; The geographic information keyword database includes a multi-industry geographic information vocabulary database, a Chinese place name and address database, and a sensitive geographic information vocabulary database; The multi-industry geographic information vocabulary is constructed by collecting professional vocabulary related to geographic information based on the surveying and mapping standard system and product standards and specifications of surveying and mapping, remote sensing, navigation, planning, environment, public security, and transportation industries; The Chinese place name and address database includes place name data, address data and point of interest data; The sensitive geographic information vocabulary includes information on prisons, material storage depots, energy facilities, satellite observation stations, detention centers, industrial facilities, and radar stations; The confidentiality judgment rule base is designed according to the different characteristics of the data to form a confidentiality judgment rule base. Rule base for judging texts , and each judgment rule There is a corresponding keyword list ; Step 2: Generate text pairs for the word-integrated deep attention text matching model CW-DATM, including the text to be checked and the confidentiality judgment rules; The deep attention text matching model CW-DATM for word fusion consists of an embedding layer, an encoding layer, an interaction layer, a fusion layer, a pooling layer, and a prediction layer; The embedding layer consists of the RoBERTa layer, the Word2Vec layer and the word fusion vector layer. and confidentiality judgment rules The RoBERTa model and the Word2Vec model are used to obtain the word vector and word vector representation of the text respectively; the word fusion vector layer is used to add and average all the word vectors of the characters contained in the word to obtain a new word vector, which is concatenated with the original word vector and used as the input of the encoding layer; The encoding layer is composed of a bidirectional long short-term memory network Bi-LSTM and a multi-head self-attention mechanism layer, which obtains deep contextual semantic information through the bidirectional long short-term memory network Bi-LSTM and the multi-head self-attention mechanism layer; The interaction layer uses a multi-head interactive attention mechanism layer to extract two input texts and The word interaction features between and Perform interactive computation to capture semantic dependencies between text pairs; The fusion layer, composed of a residual connection layer and a normalization layer, fuses the semantic information obtained from the encoding layer and the interaction layer; The pooling layer, combined with the maximum pooling layer and the average pooling layer, extracts the key information of the text fusion representation and and Fusion features and Perform pooling calculations and obtain and ; The prediction layer consists of two layers of fully connected neural networks and a Sigmoid layer; the vector obtained by the splicing pooling operation and , and input it into a two-layer fully connected neural network, and the text is calculated through the sigmoid activation function and The semantic similarity of , thereby predicting text Is it safe? When the text similarity threshold is greater than or equal to the text similarity threshold, the text to be checked and confidentiality judgment rules The semantics are similar, Can be considered safe, otherwise, for unsafe; Step 3: Detect whether the text is safe by integrating the word-based deep attention text matching model CW-DATM.

2. The text geographic information security inspection method based on a knowledge base according to claim 1 is characterized in that: Obtaining the file containing the geographic information keyword and the text containing the keyword in the file as described in step 1; The specific implementation includes the following sub-steps: Step 1.1: Get the files whose names contain keywords in the geographic information keyword library and get the file set ; Step 1.2: Extract the file content, extract the text containing keywords and their context information, and obtain the sentence set of each file , and each sentence There is a corresponding keyword list .

3. The knowledge base-based text geographic information security inspection method according to claim 2, characterized in that: In step 2, for the file set The set of sentences for each file in , each text to be checked and Rules for determining confidentiality Composition text pairs ,in, , ; The conditions for forming a text pair are, and Keyword list and Contains the same keywords.

4. The knowledge base-based text geographic information security inspection method according to claim 1, characterized in that: ; ; ; ; ; in, and is the weight vector, and For bias.

5. The knowledge base-based text geographic information security inspection method according to claim 4 is characterized in that: The word fusion vector layer, if the text A word vector in , which consists of word vectors After the word fusion layer, the word fusion vector is obtained ; , For text The total number of word vectors in .

6. The knowledge base-based text geographic information security inspection method according to claim 5, characterized in that: The encoding layer uses a bidirectional long short-term memory network Bi-LSTM to input Encode and capture long-distance dependencies of text sequences from both positive and negative directions; Word fusion vector The hidden layer representation containing contextual information is obtained through forward LSTM , after reverse LSTM, the hidden layer representation containing context information is obtained , concatenate the two vectors as the output of the bidirectional long short-term memory network Bi-LSTM ; The encoding layer uses a multi-head self-attention mechanism layer to extract deep features of the text, assigns different weights to the output of the bidirectional long short-term memory network Bi-LSTM, and captures multiple features of semantically relevant information; , combined with 、 and Get the multi-head self-attention weight calculation results : ; ; in, , , , 、 and is the weight matrix, is the matrix dimension, is a single-head attention unit, ; is the number of splicing, is the weight matrix used for linear transformation.

7. The knowledge base-based text geographic information security inspection method according to claim 6 is characterized in that: The interaction layer calculates the text interaction matrix : ; ; in, , , , 、 and is the weight matrix, is the matrix dimension, is a single-head attention unit, ; is the number of splicing, is the weight matrix used for linear transformation.

8. The knowledge base-based text geographic information security inspection method according to claim 7, characterized in that: The fusion layer targets text ,right 、 And the matrix obtained by multiplying the two is subjected to residual connection and layer normalization, then the fusion feature The calculation formula is: ; ; ; in, is the weight matrix, .

9. The knowledge base-based text geographic information security inspection method according to any one of claims 1 to 8, characterized in that: The deep attention text matching model CW-DATM for fused words is a trained deep attention text matching model CW-DATM for fused words; During the training process, the binary cross entropy loss function is used as the objective function. At the same time, to prevent overfitting, L2 regularization is added to the loss function, and the model is optimized using the Adam algorithm. The loss function is calculated as follows: ; in, represents the total number of samples, represents the true label of the sample, is the predicted value of text semantic similarity, is the regularization parameter, Represents the weight coefficient of the L2 regularization term.

10. A text geographic information security inspection system based on a knowledge base, used to implement the method described in any one of claims 1 to 9; characterized in that: Includes the following modules: A file set preparation and knowledge base construction module is used to obtain files that match geographic information keywords and texts containing keywords in the files using the geographic information check knowledge base for the file set to be checked for geographic information confidentiality; The geographic information inspection knowledge base is composed of a geographic information keyword library and a confidentiality judgment rule library; The geographic information keyword database includes a multi-industry geographic information vocabulary database, a Chinese place name and address database, and a sensitive geographic information vocabulary database; The multi-industry geographic information vocabulary is constructed by collecting professional vocabulary related to geographic information based on the surveying and mapping standard system and product standards and specifications of surveying and mapping, remote sensing, navigation, planning, environment, public security, and transportation industries; The Chinese place name and address database includes place name data, address data and point of interest data; The sensitive geographic information vocabulary includes information on prisons, material storage depots, energy facilities, satellite observation stations, detention centers, industrial facilities, and radar stations; The confidentiality judgment rule base is designed according to the different characteristics of the data to form a confidentiality judgment rule base. Rule base for judging texts , and each judgment rule There is a corresponding keyword list ; The confidentiality check text pair generation module is used to generate text pairs for the word-integrated deep attention text matching model CW-DATM, including the text to be checked and the confidentiality judgment rules; The text matching and security detection module is used to detect whether the text is safe by integrating the deep attention text matching model CW-DATM that integrates words.

Citation Information

Patent Citations

  • A place name information extraction and spatial positioning method in an internet text

    CN114091454A

  • Semantic sentiment analysis method fusing in-depth features and time sequence models

    US11194972B1