Small sample entity extraction method for military field

By combining the comprehensive method of BERT, IDCNN and CRF models, the problem of blurred boundaries and complex context in the entity extraction task in the military field is solved, efficient entity extraction effect is achieved, and training costs are reduced.

CN120163148APending Publication Date: 2025-06-17BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235072.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the military field, entity extraction tasks face unique challenges such as blurred boundaries, difficulty in obtaining data, high accuracy requirements, and complex context environments, and existing technologies are difficult to effectively solve these problems.

Method used

A small sample entity extraction method for military field is proposed. A comprehensive model of BERT pre-trained model, IDCNN model and CRF model is used to construct non-associated sentence pairs through random replacement sentence pairs, feature extraction and sequence annotation are performed to realize entity extraction.

Benefits of technology

While maintaining the advantages of the BERT model, this method fully captures text features, reduces training costs, shortens the development and deployment cycle of the model, and is suitable for small sample entity extraction tasks in the military field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163148A_ABST
    Figure CN120163148A_ABST
Patent Text Reader

Abstract

The invention discloses a military field-oriented small sample entity extraction method, which relates to the technical field of natural language processing, and comprises the following steps of: in a small sample military text, inputting a non-associated sentence pair into a BERT model for training to obtain vector representation; inputting the vector representation into an IDCNN model, adding an expansion width on classical convolution, and performing feature extraction on sentences by using a convolutional neural network; performing sequence labeling on the extracted features in combination with a CRF model; according to a given sequence and a corresponding label sequence, the score of each label is obtained through linear mapping, and the scores of all the labels are compared to obtain a global optimal label sequence. According to the method, the advantages of BERT, IDCNN and CRF models are fully utilized, while the advantages of the BERT model are kept, text features are fully captured, the model is trained based on a proper amount of training parameters, an effective entity extraction effect is achieved, meanwhile, the training cost is low, the service cycle is shortened, and the method is suitable for small sample entity extraction in the military field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a small-sample entity extraction method for the military field. Background Technique

[0002] Among many tasks in natural language processing, entity extraction occupies an important position, aiming to screen out information fragments with specific meanings from text. Currently, entity extraction technology is widely applied to various natural language processing tasks such as semantic analysis, question answering systems, machine translation, knowledge graphs, etc.

[0003] Currently, scholars at home and abroad have conducted a large number of studies on the entity extraction task. Early entity extraction technologies mostly adopted a combination of rules and vocabulary tables. This method identified entities in text by means of template rules written by experts and dictionary matching with a sufficient and rich vocabulary. However, the method based on rules and dictionaries not only has high requirements for expert knowledge, but also has poor portability, a cumbersome establishment process, poor adaptability, and high labor costs. Later, entity extraction technologies based on statistical methods emerged. This method mainly distinguishes features of different categories through manually designed statistical models to learn the characteristics and patterns of entities in text from a large amount of labeled training data, greatly reducing the dependence on manual rules. However, the entity extraction method based on statistics depends on a large amount of labeled data, complex feature engineering, and is difficult to adapt to new entities and language phenomena.

[0004] With the continuous progress of deep learning technology, the entity extraction method based on deep learning has become the dominant technology in the current NLP field due to its powerful automatic feature learning ability and excellent performance. However, in a specific field such as the military field, the entity extraction task faces unique challenges, including fuzzy boundaries, difficult data acquisition, high accuracy requirements, and complex context environments. These characteristics make the entity extraction work in this field more complex and meticulous than in other fields. With the emergence of the pre-trained model BERT, in order to address the unique challenges in the military field, researchers have proposed an entity extraction method that combines multiple neural networks such as BERT, BiLSTM, and CRF, which can more effectively handle fuzzy boundaries and complex context features, greatly improving the recall rate and F1 value. However, the model proposed by this method is relatively complex, and complex models usually require more computing resources and time to complete training, which not only increases the training cost but also may extend the cycle from model development to deployment. Therefore, a small-sample entity extraction method for the military field is proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide a small-sample entity extraction method for the military field to solve the problems raised in the above background technique.

[0006] To achieve the above object, the present invention provides the following technical solution: A small-sample entity extraction method for the military field, comprising the following steps:

[0007] Step 1, in small-sample military texts, select sentence pairs with context association, and randomly replace at least half of the sentences to construct non-associated sentence pairs for training;

[0008] Step 2, construct a BERT-IDCNN-CRF integrated model composed of a BERT pre-trained model, an IDCNN model, and a CRF model, input the non-associated sentence pairs obtained in Step 1 into the BERT model for training, and obtain vector representations that fuse context information;

[0009] Step 3, input the vector representations obtained in Step 2 into the IDCNN model, add a dilation width to the classical convolution, and use a convolutional neural network to extract features from the sentences;

[0010] Step 4, combine the CRF model to perform sequence labeling on the features extracted in Step 3;

[0011] Step 5, according to the given sequence and the corresponding label sequence, obtain the scores of each label through linear mapping, and obtain the globally optimal label sequence by comparing the scores of each label.

[0012] Preferably, in Step 1, it further includes preprocessing the constructed non-associated sentence pairs, including but not limited to word segmentation, removal of stop words and punctuation marks, and converting the preprocessed sentence pairs into a format recognizable by the BERT model.

[0013] Preferably, in Step 2, the process of obtaining the vector representations of non-associated sentence pairs through the BERT model is specifically as follows:

[0014] Use the framework to load the BERT model pre-trained with non-associated sentence pairs;

[0015] Pad or truncate the sentences to a fixed length and input them into the BERT model for training. During the training process, continuously adjust the parameters of the BERT model and monitor the performance of the BERT model;

[0016] Extract the vector representations of the sentence pairs from the preset layer of the BERT model.

[0017] Preferably, in the process of extracting the vector representations of non-associated sentence pairs in the BERT model, word embedding is used to represent the semantic information of words and positional embedding is used to capture positional information, strengthening the understanding of the word order or time step information in the input sequence. The model in the BERT model fills the time series sequence with 512 characters.

[0018] Preferably, in the third step, dilated convolution is applied to increase the receptive field of the convolutional layers of the IDCNN model, effectively capturing the long-distance dependencies of words in the sentence pair;

[0019] Among them, multiple dilated convolutional blocks are stacked together in the IDCNN model in sequence to obtain a feature map with context features of sequential data.

[0020] Preferably, in the third step, during the process of feature extraction of sentences using a convolutional neural network, a global pooling layer is used to integrate the information in the feature map output by the convolutional layers of the IDCNN model, converting the feature map into a global feature vector. Then, the feature vector is input into a fully connected layer to further process the global feature vector and convert it into a dimension matching the number of output classes, obtaining a feature vector that can be input into the CRF model.

[0021] Preferably, in the fifth step, the formula for calculating the score of each label in the CRF model is:

[0022] P i =W s h (t) +b

[0023] In the formula, P i represents the score of the label at the i-th position, and W s h (t) is the product of the hidden state h (t) at time step t and the weight matrix W, and b represents the bias term.

[0024] Preferably, in the fifth step, after obtaining the scores of each label through linear mapping, the relationship between different labels is modeled by the CRF model, and the Viterbi algorithm is used to decode on the CRF model to find the globally optimal label sequence.

[0025] Preferably, in the fifth step, obtaining the globally optimal label sequence according to the comparison of each label score is specifically:

[0026] Based on the label scores, the scoring function is set as:

[0027]

[0028] In the formula, s(x, y) represents the matching score from the input sequence x to the label sequence y, represents the transition score from one label y i-1 to the next label y i , represents the output score when the label corresponding to index i is y i ; among them, and It is gradually optimized through a learning process and represents the model's estimation of the probability of transitioning between different tags and selecting specific tags at different positions;

[0029] All possible tag sequences corresponding to a given input sequence are evaluated by calculating a scoring function, and the tag sequence with the highest score is selected as the prediction result.

[0030] Preferably, in the CRF model, W = (W i,j ) represents the maximum likelihood estimate calculated based on the training samples {x i , y j}, and its objective function for training the CRF model is:

[0031]

[0032] where λ and θ represent regularization parameters, and P represents the probability distribution from the source sequence to the predicted target sequence.

[0033] Compared with the prior art, the technical effects of the present invention:

[0034] A small-sample entity extraction method for the military field provided by the present invention makes full use of the advantages of the BERT, IDCNN, and CRF models. Through the BERT pre-trained language model, rich context information in a specific context is learned to generate word embedding vectors with deep semantic information. Through combining with the IDCNN model for multi-layer feature extraction, and finally using the CRF model to obtain the optimal predicted tag sequence. While maintaining the advantages of the BERT model, this method fully captures text features, trains the model based on an appropriate amount of training parameters, achieves effective entity extraction effects, and at the same time has low training costs, shortens the usage cycle, and is more suitable for small-sample entity extraction in the military field. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flowchart of the entity extraction method in an embodiment of the present invention.

[0036] Figure 2 is a graphical representation schematic diagram of the dilated convolution technique adopted in an embodiment of the present invention.

[0037] Figure 3 is a structural block diagram of the word and character encoding unit in an embodiment of the present invention.

[0038] Figure 4 is an overall architecture diagram of the BERT-IDCNN-CRF integrated model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] The present invention provides a small-sample entity extraction method for the military field as Figures 1 - 4 shown, including the following steps:

[0041] Step 1, in small-sample military texts, select sentence pairs with context associations (contextual connections) (these sentences are closely related logically, semantically, or thematically), and randomly replace half of the sentences to construct non-associated sentence pairs for training (not closely related to the previous sentences logically, semantically, or thematically), and perform preprocessing on the constructed non-associated sentence pairs including but not limited to word segmentation, removal of stop words, and punctuation marks, convert the preprocessed sentence pairs into a format recognizable by the BERT model. During word segmentation, the results need to be converted into tokens recognizable by the BERT model, which usually involves converting the text into indices in a series of predefined vocabularies, and special tokens required by the BERT model need to be added at the beginning and end of the text. These tokens help the model understand the boundaries and structure of the text;

[0042] Step 2, construct a BERT-IDCNN-CRF integrated model composed of a BERT pre-trained model, an IDCNN model, and a CRF model. Connect the IDCNN model and the CRF model through a convolutional neural network composed of a global pooling layer and a fully connected layer. Use the global pooling layer to reduce the dimension and integrate the feature map, which can integrate spatially distributed features into a global feature vector, achieving the functions of reducing the computational amount, reducing overfitting, and enhancing robustness. The fully connected layer is generally located after the global pooling layer, and its main function is to fuse and combine the information in the global feature vector to extract high-level features useful for classification or annotation tasks, and transform the dimension of the global feature vector to match the number of output categories for subsequent classification. Input the non-associated sentence pairs obtained in Step 1 into the BERT model for training to obtain vector representations that fuse context information. These vector representations capture long-distance dependencies in the language. In the BERT model, not only word embeddings are used to represent the semantic information of words, but also position embeddings are used to capture position information. And the model pads the time series sequence to 512 characters and trains and continuously adjusts the parameters to obtain the vector representations of the sentence pairs.

[0043] Among them, the main role of word embeddings is to map each word or sub-word in the text to a real-valued vector in a high-dimensional vector space. The purpose of this is to convert the words in the text into a form that can be understood and processed by the machine for subsequent calculations and analyses. Word embeddings can help the model understand the semantics and context information of each word, which is crucial for natural language processing tasks; in natural language processing, the order and position of words are also crucial for understanding the meaning of sentences. The semantic role of the same word in different positions may be different (for example, "I like to eat apples" and "Apples, I like to eat" use the same words, but due to the different word orders, the meanings of the sentences are completely different). Therefore, relying solely on word embeddings to capture semantic information is not enough. Position embeddings are also needed to represent the position information of each word in the sentence. Position embeddings are achieved by assigning a position vector to each word, enabling the model to capture the relative position relationship of words in the sentence, thereby better understanding the semantics of the sentence. This helps the model to more accurately distinguish and understand the role and meaning of each word in the sentence when processing non-related sentence pairs; by combining word embeddings and position embeddings, the BERT model can better understand the semantics and context relationship of the input text. Word embeddings capture the semantic information of each word, while position embeddings represent the position information of the words. The combination of the two can provide a more comprehensive and accurate semantic representation; in addition, since the BERT model can capture the semantic information and position information of words, this enables the BERT model to have higher performance when processing natural language processing tasks;

[0044] Furthermore, padding the time series sequence to 512 characters and training and continuously adjusting parameters are important steps in the process of using the BERT pre-trained model to extract the vector representation of non-related sentence pairs. These steps help ensure that the model meets its design limitations when processing the input text, improve the computational efficiency and memory usage efficiency, and optimize the performance of the model to adapt to different downstream tasks;

[0045] Step 3: Input the vector representation obtained in Step 2 into the IDCNN model. Add a dilation width to the classical convolution. The IDCNN model processes the word vector representation output by the BERT model through multiple dilated convolutional layers. Dilated convolution is achieved by inserting empty positions (i.e., holes) in the convolutional kernel. Stack multiple dilated convolutional blocks together in sequence in the IDCNN model to obtain a feature map with context features of sequential data. By applying dilated convolution, the scope of action of the convolutional layer of the IDCNN model is increased, effectively capturing the long-distance dependencies of words in the sentence pair. Then, use a convolutional neural network to extract features from the sentence. In the IDCNN model, the IDCNN convolutional neural network serves as a downstream processing module of BERT and utilizes the design of iterative dilated convolution to further refine the feature extraction process. The structure of the IDCNN model enables the model to better understand the long-distance dependencies in the text, which is particularly important for entity extraction tasks because entities usually have complex syntactic structures in the context information of the text.

[0046] Step 4: Combine the CRF model to perform sequence labeling on the features extracted in Step 3. The CRF model is a discriminative probability model that can directly model conditional probabilities, avoiding the complexity of calculating joint probabilities required by generative models. At the same time, the CRF model considers the global information of the entire sequence rather than just local information, thereby improving the accuracy of prediction. In addition, the CRF model is also flexible and can be combined with multiple features for modeling. After the IDCNN model extracts features from the sentence, a series of feature vectors are generated. These feature vectors are input into the CRF model. The CRF model can predict the label for each position in the input sequence based on these feature vectors and the dependencies between the labels. Finally, the CRF model outputs a complete labeled sequence, which represents the label information for each position in the input feature vectors. This method of combining the IDCNN model and the CRF model can make full use of the advantages of both. The IDCNN model is good at feature extraction, while the CRF model is good at handling label dependencies in sequence labeling, usually achieving better performance in the sequence labeling task of sentences.

[0047] Step 5: According to the given sequence and the corresponding label sequence, use the Viterbi decoding function of the CRF model to obtain the score of each label through linear mapping. Model the relationship between different labels through the CRF model, perform decoding on the CRF model using the Viterbi algorithm, and find the globally optimal label sequence. Obtain the globally optimal label sequence by comparing the scores of each label. The CRF model accurately models the constraints between labels by considering the mutual dependencies between labels, thereby achieving the global optimization of the entire output sequence.

[0048] Among them, in step two, the process of obtaining the vector representation of non-related sentence pairs through the BERT model is specifically as follows: Use the framework to load the pre-trained BERT model for non-related sentence pairs; Pad or truncate the sentences to a fixed length of 512 characters and input them into the BERT model for training. During the training process, continuously adjust the parameters of the BERT model and monitor the performance of the BERT model; Extract the vector representation of the sentence pair from the preset layer of the BERT model.

[0049] In step three, during the process of using the convolutional neural network to extract features from sentences, use the global pooling layer to integrate the information in the feature map output by the convolutional layer of the IDCNN model, convert the feature map into a global feature vector, and then input the feature vector into the fully connected layer to further process the global feature vector and convert it into a dimension that matches the number of output categories to obtain a feature vector that can be input into the CRF model.

[0050] It should be noted that the formula for calculating the score of each label in the CRF model is:

[0051] P i =W s h (t) +b

[0052] In the formula, P i represents the label score at the i-th position, W s h (t) is the product of the hidden state h (t) at time step t and the weight matrix W, and b represents the bias term.

[0053] It should be noted that obtaining the globally optimal label sequence by comparing the scores of each label is specifically as follows: Based on the label scores, set the scoring function as:

[0054]

[0055] In the formula, s(x,y) represents the matching score from the input sequence x to the label sequence y, represents the transition score from one label y i-1 to the next label y i , represents the output score when the label corresponding to index i is y i ; Among them, and are gradually optimized through the learning process, representing the model's estimation of the probability of transitioning between different labels and selecting specific labels at different positions;

[0056] W=(W i,j ) represents based on the training samples {x i ,y jThe calculated maximum likelihood estimate value, and the objective function for training the CRF model is:

[0057]

[0058] where λ and θ represent regularization parameters, and P represents the probability distribution from the source sequence to the predicted target sequence.

[0059] All possible label sequences corresponding to the given input sequence are evaluated by calculating the scoring function, and the label sequence with the highest score is selected as the prediction result.

[0060] It should be noted that dilated convolution refers to adding a dilation width where data will be skipped on top of a convolution with a constant kernel size by inserting blank positions (i.e., holes) in the convolution kernel, as Figure 2 shown.

[0061] In Figure 2 Figure (a) shows a conventional convolution operation with a 3×3 convolution kernel size and a dilation rate set to 1 at this time; the convolution in Figure (b) has a dilation width of 2 and a corresponding dilation rate of 2, which, based on the structure of Figure (a), expands the receptive field to 7×7; while in Figure (c), the dilation width is set to 4 and the dilation rate is 4, which expands the receptive field of the convolution on the basis of the convolution operation in Figure (b) to be equivalent to 15×15;

[0062] The IDCNN layer (IDCNN model) consists of 4 dilated convolution blocks with the same size and containing three convolutional layers with dilation widths of 1, 1, and 2. This design enables each block to gradually expand its receptive field when processing the input; since the parameters of each layer are independent and the receptive field expands exponentially with the increase in the number of layers, the entire network can observe the entire input sequence, thereby capturing long-range dependencies; this method significantly improves the computational efficiency because the number of parameters only increases linearly.

[0063] In this embodiment, the BERT model is based on the Transformer architecture; the Transformer encoding unit structure is as Figure 3As shown, the encoding unit realizes the effective processing of the input sequence through components such as the self-attention mechanism and the position feed-forward network, enabling the model to identify and process complex correlation relationships in the input data without explicitly considering the sequence order; the BERT model uses Transformer as the encoder, which can capture the context information before and after the text at the same time, greatly improving the accuracy of natural language processing (NLP) tasks. Traditional unidirectional language models can usually only capture unidirectional context information, while the bidirectional encoding ability of BERT enables it to more accurately understand the inherent meaning of language. The core of the Transformer architecture is the self-attention mechanism, which can process all words in the input sequence in parallel, unlike RNN that needs to be processed sequentially, which greatly improves the training speed, especially in the processing of long texts.

[0064] As Figure 4 As shown, the BERT-IDCNN-CRF integrated model architecture used in this embodiment is mainly divided into three layers: the BERT layer, the IDCNN layer, and the CRF layer; First, the BERT pre-trained model serves as the starting point of this architecture. By pre-training on a large-scale text, it provides a deep linguistic representation for the model; BERT captures rich semantic information of words in the context through bidirectional training. This method not only improves the quality of word vectors but also lays a solid foundation for the subsequent training of the model;

[0065] Next, the IDCNN convolutional neural network, as a downstream processing module of BERT, utilizes the design of iterative dilated convolution to further refine the feature extraction process; the structure of IDCNN enables the model to better understand the long-distance dependence relationships in the text, which is particularly important for entity extraction tasks because entities usually have complex syntactic structures in the context information of the text;

[0066] Finally, at the output layer of the model, the CRF layer accurately models the constraints between labels by considering the interdependence relationships between labels, thus achieving global optimization of the entire output sequence.

[0067] Overall, the performance improvement of the BERT-IDCNN-CRF integrated model benefits from its multiple attentions to language representation learning, context semantic information capture, efficient and accurate feature extraction, and label sequence optimization; during the training process, the model jointly optimizes the cross-entropy loss and the CRF conditional log-likelihood loss, enabling the model to better understand the relationship between the context information in the text and the label sequence. The advantage of this integrated model is that it better adapts to the complexity and diversity of entity recognition in entity extraction tasks through multi-level and global information processing.

[0068] The following is a further illustration of a small-sample entity extraction method for the military field provided in this embodiment through experimental examples:

[0069] The dataset used in this experimental example is sourced from unstructured text data on the Global Military Network. Eventually, 3,000 pieces of data are manually annotated. Data annotation reviewers check the data one by one to ensure the annotation quality, and the dataset is divided into training data, validation data, and test data according to the ratio of 7:2:1; this dataset covers eight entity categories, namely military events, weapons and equipment, personal names, time, military place names, military ranks and positions, military institutions, and military facilities. This disclosure uses the BIO annotation mode;

[0070] The experiment uses BERT-Base as the pre-trained language model, with a maximum sequence length of 128, a batch size of 16, and a learning rate of 5×10 -5 , and the dropout rate is 0.5; the convolutional kernels used in the IDCNN layer are 3×3, and the dilation widths are 1, 1, 2;

[0071] In this experiment, different settings are made for the convolutional kernel size of the IDCNN layer, and 10, 20, 50, and 100 are respectively used for testing. At the same time, the number of layers is fixed at 4 layers, and the number of convolutional kernels in the IDCNN layer and the number of convolutional layers are experimentally explored;

[0072] To more comprehensively evaluate the model performance, this experimental example adopts a set of evaluation index systems, which cover the mainstream entity granularity-based evaluation criteria in the current entity extraction field: precision P, recall R, and F1 value;

[0073] This disclosure's experiment selects the following 5 entity extraction models for comparative experiments with the BERT-IDCNN-CRF model:

[0074] (1) BiLSTM-CRF model, which uses pre-trained word vectors and combines the bidirectional long short-term memory network LSTM and the conditional random field CTF for training;

[0075] (2) IDCNN-CRF model, which combines the IDCNN model and the CRF model for training;

[0076] (3) Radical-BiLSTM-CRF model, which combines the root information of Chinese characters on the basis of the BiLSTM-CRF model;

[0077] (4) Lattice-LSTM-CRF model, which can better utilize the internal structure information of the text by introducing the Lattice structure;

[0078] (5) BERT-fine-tuning model, which is a fine-tuning mode of the BERT model for specific downstream tasks.

[0079] The experimental data is shown in the following table.

[0080] Table 1 Extraction Results of Different Types of Entities

[0081]

[0082] According to Table 1, it can be seen that the entity recognition accuracy of the military event and military facility categories is relatively good. This is mainly because the structure of military events is relatively complex, and at the same time, most military facilities have the situation of noun nesting.

[0083] Table 2 Extraction Results of Different Models

[0084]

[0085] The comparison between the BERT-IDCNN-CRF integrated model and other related works is shown in Table 3. From the experimental results, it can be seen that the BERT-IDCNN-CRF model proposed in this embodiment has higher accuracy in entity extraction than the models used in other comparative experiments, and at the same time, the time consumption is the shortest. It can be seen that this model combines the strong expression ability of the BERT model, the efficient feature extraction ability of the IDCNN model, and the prediction ability of the CRF model, thus improving the effect of entity recognition in a short time.

[0086] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A small sample entity extraction method for the military field, characterized in that: The following steps are involved: Step 1: In a small sample of military texts, select sentence pairs with context associations and randomly replace at least half of the sentences to construct non-associated sentence pairs for training; Step 2: Build a BERT-IDCNN-CRF comprehensive model consisting of a BERT pre-trained model, an IDCNN model, and a CRF model. Input the non-related sentence pairs obtained in step 1 into the BERT model for training to obtain a vector representation that incorporates context information. Step 3: Input the vector representation obtained in step 2 into the IDCNN model, add a dilation width to the classic convolution, and use the convolutional neural network to extract features of the sentence; Step 4: Combine the CRF model to perform sequence labeling on the features extracted in step 3; Step 5: According to the given sequence and the corresponding label sequence, the score of each label is obtained through linear mapping, and the global optimal label sequence is obtained by comparing the scores of each label.

2. According to the method for extracting small sample entities for the military field according to claim 1, it is characterized in that: In the step 1, it also includes preprocessing the constructed non-related sentence pairs, including but not limited to word segmentation, removal of stop words and punctuation marks, and converting the preprocessed sentence pairs into a format that can be recognized by the BERT model.

3. A method for extracting small sample entities for military fields according to claim 2, characterized in that: In step 2, the process of obtaining the vector representation of non-related sentence pairs through the BERT model is specifically as follows: Use the framework to load the BERT model of pre-trained non-related sentence pairs; Pad or truncate the sentences to a fixed length and input them into the BERT model for training. During the training process, continuously adjust the parameters of the BERT model and monitor the performance of the BERT model. Extract vector representations of sentence pairs from the preset layers of the BERT model.

4. According to the method for extracting small sample entities for the military field as claimed in claim 3, it is characterized in that: In the process of extracting vector representations of non-related sentence pairs in the BERT model, word embedding is used to represent the semantic information of words and position embedding is used to capture position information, thereby enhancing the understanding of word order or time step information in the input sequence. In the BERT model, the model fills the time series sequence with 512 characters.

5. A method for extracting small sample entities for the military field according to claim 4, characterized in that: In the step 3, dilated convolution is applied to increase the scope of the convolution layer of the IDCNN model, and effectively capture the long-distance dependency of words in the sentence pairs; Among them, the IDCNN model stacks multiple dilated convolutional blocks together in sequence to obtain a feature map with contextual features of sequence data.

6. A method for extracting small sample entities for the military field according to claim 5, characterized in that: In the step three, in the process of extracting features from sentences using a convolutional neural network, a global pooling layer is used to integrate the information in the feature map output by the convolutional layer of the IDCNN model, and the feature map is converted into a global feature vector. The feature vector is then input into a fully connected layer, and the global feature vector is further processed and converted into a dimension that matches the number of output categories, to obtain a feature vector that can be input into the CRF model.

7. A method for extracting small sample entities for the military field according to claim 6, characterized in that: In step 5, the score formula for calculating each tag in the CRF model is: P i =W s h (t) +b Where P i represents the label score of the i-th position, W s h (t) The hidden state h at time step t (t) The product of the weight matrix W, b represents the bias term.

8. The method for extracting small sample entities for the military field according to claim 7, characterized in that: In step five, after obtaining the score of each label through linear mapping, the relationship between different labels is modeled by the CRF model, and the Viterbi algorithm is used to decode on the CRF model to find the global optimal label sequence.

9. The method for extracting small sample entities for the military field according to claim 8, characterized in that: In step 5, the global optimal tag sequence is obtained by comparing the tags scores as follows: Based on the label score, the score function is set as: In the formula, s(x,y) represents the matching score from the input sequence x to the label sequence y. Represents a label y i-1 Move to next label y i The conversion score, Indicates that the label corresponding to index i is y i The output score of and It is gradually optimized through the learning process and represents the model's estimate of the probability of transferring between different labels and selecting specific labels at different positions; The scoring function is calculated to evaluate all possible label sequences corresponding to a given input sequence, and the label sequence with the highest score is selected as the prediction result.

10. The method for extracting small sample entities from military fields according to claim 9, characterized in that: In the CRF model, W = (W i,j ) represents the training sample {x i ,y j The maximum likelihood estimate calculated by}, the objective function used to train the CRF model is: Among them, λ and θ represent regularization parameters, and P represents the probability distribution from the source sequence to the predicted target sequence.