A Text Semantic Retrieval Method in a Military Scenario

By constructing a semantic representation model for military question-and-answer sentences and a text retrieval model, the problems of strong professionalism, user expressions and high accuracy in military scenarios are solved, and fast and accurate text data positioning is achieved, which improves the search efficiency and effect.

CN116150335BActive Publication Date: 2025-07-11THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211630251.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-07-11
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

The existing text search methods have problems such as strong professionalism, different user problem expressions, many professional vocabulary, difficulty in applying general pre-trained models, and difficulty in meeting the recall results, resulting in poor retrieval results.

Method used

Construct a semantic representation model for military question-answer sentences, use vector similarity search method to recall text collections related to question-answer sentences, and quickly and accurately locate the required text data through text retrieval fine-swer model, including the construction of military pre-trained model for offline parts, the offline construction of dual semantic retrieval model for text retrieval fine-swer model, and the real-time text retrieval task for online parts.

Benefits of technology

It significantly improves the complete and accurate text search rate and accuracy rate in military scenarios, improves the search efficiency and user experience, and can quickly and accurately locate the required text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150335B_ABST
    Figure CN116150335B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for text semantic retrieval in a military scenario. First, based on a military pre-trained model, a dual semantic retrieval model is constructed, fine-tuned on a military semantic retrieval dataset to form a question-answer pair language representation model, and a semantic vector library of military text data is obtained offline. A secondary inverted index is constructed by means of vector clustering. Second, based on the military pre-trained model, a text retrieval and refinement model is constructed and fine-tuned on a military semantic retrieval and refinement dataset. In the face of a real-time retrieval task, the question sentence language representation model is used to obtain the semantic vector representation of the question sentence, retrieve through vector similarity calculation, recall a text set that meets the user's needs, and use the text retrieval and refinement model to accurately locate specific text data and feedback it to the user. This method can accurately locate the data required by the user in real time from a large amount of military text data and can be used in scenarios such as massive text search and retrieval-based question answering in a military scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of semantic retrieval and intelligent question answering, and particularly relates to a method for text semantic retrieval in a military scenario. Background Art

[0002] With the increasing maturity of the digitalization and data engineering construction of our military, the military data assets have increased exponentially. Under the background of big data, the retrieval of military data resources has become an indispensable part for military users when performing various tasks. Facing the widely distributed and massive text data in the joint information service environment, how to help various military personnel quickly and accurately locate the required data in the ocean of data has become a key problem that needs to be solved urgently.

[0003] Traditional text semantic retrieval methods achieve the retrieval of massive text data through keyword retrieval. Specifically, first, the retrieval requirement is segmented to obtain the information of retrieval keywords / terms; then, by querying semantic dictionaries, synonym forests, and related dictionaries, etc., the information search scope is expanded; then, statistical algorithms such as BM25 (Best Match) with the term frequency statistics method such as TF-IDF (term frequency-inverse document frequency) as the core are used to recall the text set strongly related to the question sentence. This statistical method of "hard" matching with keywords solves the text retrieval problem at the word level, but ignores the information at the text semantic level, and only has a good effect in the usage scenarios where the search keywords are accurate and the search intention is clear.

[0004] With the development of natural language processing technology, researchers abstract text information retrieval as a text matching problem, and begin to try to construct a discriminant LTR (Learning to Rank) model, use a language representation model to obtain text feature representations, and recall the text set strongly related to the question sentence by calculating the similarity between texts. For example, Fu Jian et al. use a convolutional deep neural network model to extract the intra-sentence structure features of the question sentence and the answer sentence respectively and the interaction information between the two, and calculate the similarity, so as to improve the effect of the question answering task based on document retrieval. Shao Mingrui et al. apply a deep neural network model combining Transformer and attention mechanism to the question answering task based on FQA, so that the model can also complete the most similar matching of texts even when the data set is small. Zhu Zongkui et al. use the BERT pre-trained model to complete entity recognition and similarity calculation problems for the particularity of Chinese question sentences to optimize the performance of the Chinese knowledge graph question answering task. Compared with traditional methods, the text matching method based on deep learning can automatically learn the semantic information of question and answer sentences and is more excellent in text semantic matching.

[0005] The limitations of existing text retrieval methods in military scenarios are as follows: First, in military scenarios, various text retrieval requirements are highly specialized, and the clarity of various user questions varies. The traditional text semantic retrieval method based on word frequency statistics has poor performance. Second, there are many professional terms in military text corpora, and existing general pre-trained language models are difficult to apply. A large amount of military corpus data needs to be collected to retrain and generate a military pre-trained model. Third, the method based on the recall text similarity ranking is difficult to meet the high text precision requirements in military scenarios, and the recall results need to be further processed to obtain the best text answers. Summary of the Invention

[0006] Object of the Invention: The technical problem to be solved by the present invention is to provide a text semantic retrieval method in a military scenario for a large amount of military text data, aiming at the deficiencies of the prior art. By constructing a military question-answer sentence semantic representation model, deeply representing the semantic features of questions with strong professionalism and inconsistent expression quality, and using the vector similarity retrieval method to recall a text set strongly related to the question. On this basis, using the text retrieval re-ranking model, quickly and accurately locate the required text data, improve the recall rate and precision rate of text retrieval from the semantic perspective, and further improve the user experience.

[0007] To solve the above technical problems, the present invention discloses a text semantic retrieval method in a military scenario, which is mainly divided into two parts: offline and online. The specific content is as follows:

[0008] Offline part: Step 1, construction of a military pre-trained model: Construct a military text corpus dataset, and on this basis, construct a pre-trained model suitable for military scenarios; Step 2, offline construction of a dual semantic retrieval model: Generate a question-answer pair language representation model and a military text data semantic vector library. Based on the military pre-trained model, construct a dual semantic retrieval model, train and fine-tune it on the military semantic retrieval dataset to generate a question-answer pair language representation model, and use the answer language representation model to obtain the semantic vector representation of the military text to be retrieved, and construct a military text data semantic vector library; Construction of the secondary index of the military data text library to be retrieved: Based on the military text data semantic vector library, use the K-means clustering algorithm to construct a secondary inverted index of semantic vectors; Step 3, construction of a text retrieval re-ranking model: Traverse all questions in the military semantic retrieval dataset, obtain a text set strongly related to each question, and construct a military semantic retrieval re-ranking dataset. Based on the military pre-trained model, construct a multi-class re-ranking model, and train and fine-tune it on the military semantic retrieval re-ranking dataset to generate a text retrieval re-ranking model.

[0009] Online part: Step 4. For the real-time text retrieval task, ① input the user's data requirements, use the question language representation model to obtain the semantic representation vector of the question; ② through vector similarity calculation and retrieval, obtain the text set strongly relevant to the user's requirements; ③ use the text retrieval re-ranking model to obtain the text answer that best meets the user's requirements and feedback it to the user.

[0010] Furthermore, the military pre-trained model in Step 1 is constructed offline, including the following steps:

[0011] Step 1-1. Collect military raw corpus data;

[0012] Step 1-2. Clean and transform the redundant characters, stop words, and simplified and traditional Chinese characters in the military raw corpus data for data preprocessing, perform word segmentation on the preprocessed data, and collect the word list data in the semantic word library, synonym forest, related word library, and extended word library in the existing military information retrieval and intelligent question-answering systems to form a military word list;

[0013] Step 1-3. Select a pre-trained model in the field of natural language processing; for each piece of military raw corpus, use the military word list for mapping and transformation, that is, find the corresponding ordinal position of each word and digitize it to construct a military text corpus dataset;

[0014] Step 1-4. Based on the military text corpus dataset, set the model training parameters to train the pre-trained model to form a military pre-trained model.

[0015] Furthermore, the dual semantic retrieval model in Step 2 is constructed offline, including the following steps:

[0016] Step 2-1. Collect military retrieval question-answering corpus data, and through the mapping and transformation of the military word list, digitize the text data to construct a military semantic retrieval dataset;

[0017] Step 2-2. Based on the military pre-trained model, construct a dual semantic retrieval model. After training the dual semantic retrieval model, obtain a question-answer pair language representation model. The two language representation models are the left and right branch network models of the dual semantic retrieval model respectively;

[0018] Step 2-3. For the military data text set to be retrieved, use the answer language representation model to generate a military text data semantic vector library;

[0019] Step 2-4. For the military text data semantic vector library, use a clustering algorithm to construct a secondary inverted index.

[0020] Furthermore, in Step 2-1, for the text retrieval task, mainly using the retrieval question-answering corpus data in the raw corpus data to construct a military semantic retrieval dataset, including the following steps:

[0021] Step 2-1-1: Set the Q&A pair data in the original corpus as positive samples and set the label to 1;

[0022] Step 2-1-2: For each question in the positive samples, construct negative samples in two ways: ① Relevant but non-optimal answer samples: Use the BM25 algorithm to retrieve the top three texts that are most relevant to the question except the answer from the original corpus data, and set the similarity degrees to 0.8, 0.5, and 0.3 respectively; ② Irrelevant samples: Use the random sampling method to randomly select two text data from the military text corpus except the positive samples and those in method ① or other Q&A pairs, and set the similarity degree to 0;

[0023] Step 2-1-3: Use the military word list to map and transform the text data in the positive and negative sample sets to construct a military semantic retrieval data set.

[0024] Furthermore, Step 2-2 constructs a dual semantic retrieval model based on the military pre-trained model and obtains a Q&A pair language representation model after model training, including the following steps:

[0025] Step 2-2-1: For the military semantic retrieval task, construct two input branch networks, namely the dual semantic retrieval model; use the military pre-trained model generated in Step 1 to encode the military Q&A pair data respectively to obtain the feature representations of the questions and answers, and then use vector similarity calculation to obtain the similarity between the Q&A sentences;

[0026] Step 2-2-2: Perform training and fine-tuning on the dual semantic retrieval model on the military semantic retrieval data set. After the dual semantic retrieval model is trained, obtain the Q&A pair language representation model.

[0027] Furthermore, Step 2-4 constructs a secondary inverted index for the military text data semantic vector library using a clustering algorithm, including the following steps:

[0028] Step 2-4-1: Use the K-means algorithm with the Euclidean distance as the distance formula to perform preliminary clustering on the military text data semantic vector library, and set the initial number of clustering categories to C; use the corresponding category serial number and the class center point vector of each semantic vector as the first-level index;

[0029] Step 2-4-2: For each set of semantic vectors of each category, use the K-means algorithm with the Euclidean distance as the distance formula to perform re-clustering and division on the military text data semantic vector library, and set the number of sub-categories to K, where K is not less than 10; use the corresponding sub-category serial number and the sub-class center point vector of each semantic vector as the second-level index.

[0030] Furthermore, the offline construction of the text retrieval and refinement model in step 3 includes step 3-1: constructing a military semantic retrieval and refinement dataset, and step 3-1 includes the following steps:

[0031] Step 3-1-1: Traverse each question in the military semantic retrieval dataset, and use the question language representation model generated in step 2-2 to obtain the semantic vector representation of the question;

[0032] Step 3-1-2: For each semantic vector representation of the question, adopt a two-level inverted index method to quickly locate the Top-N text set strongly related to the question semantics from the military text data semantic vector library; N represents the number of retrieved texts; in a military scenario, if the user hopes to retrieve M texts strongly related to the requirement, and the value range of M is 1 to 5, then the number of retrieved texts N = 10 * M;

[0033] Step 3-1-3: For the N retrieval results of each question, if the answer text in the original question-answer pair exists in the retrieval results, then this answer text is a positive sample, and the label is set to 1, and the other retrieval results are defined as negative samples, and the label corresponding to each result is set to 0; if the correct answer text does not exist in the retrieval results, then randomly delete one retrieval result, and define the other retrieval results as negative samples, and set the correct answer text as a positive sample;

[0034] Step 3-1-4: Use the military word list to map and transform the positive and negative sample text data to construct a military semantic retrieval and refinement dataset. The data of this military semantic retrieval and refinement dataset includes the question and N retrieved texts, and the label is an N-dimensional 01 vector, where 1 represents the text answer that best matches the question.

[0035] Furthermore, the offline construction of the text retrieval and refinement model in step 3 also includes step 3-2: constructing a text retrieval and refinement model based on the military pre-trained model, and step 3-2 includes the following steps:

[0036] Step 3-2-1: Using the military pre-trained model in step 1 as the backbone, construct a text retrieval and refinement model. This model takes the question and N retrieved texts as inputs and an N-dimensional classification vector as the output to discriminate the text in the potential answers that best matches the question;

[0037] Step 3-2-2: Fine-tune and train the text retrieval and refinement model on the military semantic retrieval and refinement dataset obtained in step 3-1.

[0038] Furthermore, step 4 for text semantic retrieval for real-time tasks includes the following steps:

[0039] Step 4-1: When facing a real-time search task, after cleaning the data of the question, use the question language representation model to obtain the semantic vector representation of the question;

[0040] Step 4-2: Use vector similarity to retrieve and recall text sets that are strongly related to user needs;

[0041] Step 4-3: Use the military vocabulary list to map and transform the question and the recall results of step 4-2, and use this as input to use the text retrieval and ranking model to accurately locate the answer text that meets the user's needs and feedback it to the user.

[0042] Furthermore, step 4-2 uses vector similarity to retrieve and recall a text set that is strongly related to the user's needs, including the following steps:

[0043] Step 4-2-1: Traverse the C class center point vectors in step 2-4-1, use the vector dot product to calculate the similarity between each class center point and the question semantic vector, sort them by similarity, and select the class with the top (2*M) similarity ranking as the next retrieval target;

[0044] Step 4-2-2: For the 2*M classes in step 4-2-1, traverse the K subclasses respectively, and use the vector dot product to calculate the similarity between the center point of each subclass and the question semantic vector; sort the similarities of the center points of the 2*M*K subclasses from large to small, and select the top (4*10*M) subclasses as the candidate set of retrieval texts;

[0045] Step 4-2-3: Traverse step 4-2-2 to retrieve each text semantic vector in the text candidate set, use vector dot product to calculate the similarity between each vector and the question semantic vector, and sort them from large to small by similarity, and select the Top (N) texts as the full set of recalled texts.

[0046] Beneficial effects:

[0047] Compared with the prior art, the present invention has the following significant advantages: First, the recall rate is significantly improved. Traditional text retrieval methods are prone to losing some query results due to the problem of "ambiguous demand description" during the "hard" matching of keywords. The present invention utilizes a military question-answer pair language representation model to obtain the semantic features of the question-answer pair, and obtains the full set of potential answers through semantic feature similarity matching, thereby greatly improving the recall rate of the retrieval.

[0048] Second, the precision rate is significantly improved. In order to meet the needs of users in military scenarios to accurately obtain the required data, the present invention constructs a text retrieval ranking model based on a military pre-training model to obtain the text that best meets user needs from potential answers, greatly improving the precision rate of text retrieval.

[0049] Thirdly, the real-time performance is significantly improved. The text semantic retrieval method provided by the present invention can be divided into two parts: an offline system and an online system. Among them, the vector representation of military text data involving a large amount of calculations is completed by the offline system. The online system only needs to complete three parts: vectorizing text requirements, vector similarity retrieval, and refined ranking of search results. Among them, in the vector similarity retrieval part, by constructing a secondary inverted index retrieval method, the time loss is relatively small, and it can respond to user needs in a relatively timely manner, and the retrieval efficiency is improved. Brief Description of the Drawings

[0050] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0051] Figure 1 is the process of the text semantic retrieval method in a military scenario;

[0052] Figure 2 is the network structure of the Bert pre-trained model;

[0053] Figure 3 is the network structure of the dual semantic retrieval model;

[0054] Figure 4 is the network structure of the multi-class refined ranking model. Specific Embodiments

[0055] The following will describe the embodiments of the present invention in conjunction with the drawings.

[0056] The present invention proposes a text semantic retrieval method in a military scenario. By leveraging the powerful semantic feature representation ability of pre-trained models in the field of natural language processing, a dual semantic retrieval model and a text retrieval refined ranking model are respectively constructed to improve the text retrieval ability in a military scenario. The specific steps are as follows:

[0057] Step 1: Offline construction of a military pre-trained model: Construct a military text corpus dataset, select an open-source pre-trained model, and train it in the military text corpus dataset to form a military pre-trained model, including the following steps:

[0058] Step 1-1: Collect military raw corpus data: The military raw corpus data mainly includes three parts. ① Collect text data such as military movement intelligence data, military documents, and historical archive materials in the military business information system; ② Obtain relevant data in the military field from the Internet; ③ Collect question-and-answer pair data in military information retrieval, intelligent question-and-answer, and dialogue systems, mainly user search log information, specifically including user requirements and answer information.

[0059] Step 1-2: Preprocessing of the original corpus data: Clean the redundant characters, stop words, simplified and traditional Chinese characters, etc. in the military original corpus data to obtain the cleaned data. Perform word segmentation on the cleaned data, and collect the vocabulary data in the semantic vocabulary, synonym forest, related vocabulary, and extended vocabulary in the existing military information retrieval and intelligent question-answering systems, and integrate them to form a military vocabulary list.

[0060] Step 1-3: Construction of the military pre-trained model and the military text corpus dataset: As Figure 2 shown, in this embodiment, the BERT model is selected as the pre-trained model; the military original corpus is preprocessed using the military vocabulary list to form a military text corpus dataset. Since the BERT pre-trained model includes two training tasks: the masked language model and the next sentence prediction model. Therefore, based on the military text corpus dataset, two training datasets for the BERT model are constructed. The two training datasets for the BERT model include the next sentence prediction model dataset and the masked language model dataset.

[0061] Step 1-3-1: Construction of the next sentence prediction model dataset: Randomly intercept 20-30 training corpora from each military text corpus. Each corpus contains two consecutive sentences (sentence A and sentence B), that is, it forms a "next sentence" relationship, which is used as the positive sample; replace sentence B in the training corpus with any other sentence in the original corpus to form a "not the next sentence" relationship, forming a negative sample, and thus construct the next sentence prediction dataset;

[0062] Step 1-3-2: Construction of the masked language model dataset: Randomly mask some words in each military text corpus in a random masking manner. The random masking ratio of words is 15%. Use the sentences before and after the masking as the model input and output, and thus construct the masked language model dataset. To solve the problem of mismatch between pre-training and subsequent tasks, the random masking of words should be operated according to the following strategy: 80% of the randomly masked words are marked with [Mask], 10% of the randomly masked words are randomly replaced with other words, and 10% of the randomly masked words are randomly kept unchanged.

[0063] Step 1-4: Training of the military pre-trained model: Obtain the pre-trained model in the field of natural language processing from the Internet. In the present invention, an open-source Chinese BERT model is obtained from the Internet as the initial pre-trained model. Since the open-source pre-trained model is trained on a public dataset and is not applicable in the military scenario, it needs to be retrained on the next sentence prediction dataset and the masked language model dataset prepared in Step 1-3. The pseudo-code for training the military pre-trained model is shown in Table 1, and the network structure of the pre-trained model is as Figure 2 shown;

[0064] Table 1 Pseudo-code for the training process of the pre-trained model

[0065]

[0066]

[0067] Model Input Representation: The text is vectorized and represented at the input layer of the BERT model. For each training corpus in steps 1 - 3, the following input processing is performed. Add the "[CLS]" flag at the head of each training corpus and the "[SEP]" flag at the tail, and then input the sentence into the model for vectorized representation. Set the maximum corpus length to 512. The input vectorized representation is composed of the sum of word vectors, text vectors, and position vectors, with a dimension of 768. Among them, the word vectors are obtained by looking up the word id (sequence number) in the word segmentation list in steps 1 - 2, and then querying the word vector table (random vector representation) to convert each word into a vector representation; the text vector specifically represents whether the input is the same sentence text. If it is the same, the label is set to 0. If there are two sentences in the text, the label value corresponding to each word in the second sentence is 1; the position vectors are obtained in the form of random vector representation according to the sequence number of each word in the text.

[0068] Model Structure Selection: The present invention selects the BERT Lager network model as the initial pre - trained model. The number of layers of this model's Trm (Transformer blocks) is 24, the dimension of the hidden layer is 1024, the number of self - attention mechanism modules is 16, and the total number of parameters is approximately 340M.

[0069] Model Output Representation: In the Bert pre - trained model, the model input and output positions correspond one by one. The position corresponding to [CLS] represents the sentence - level semantic relationship feature, and the position corresponding to each other word represents the word - level semantic feature. For the two tasks of predicting the following sentence and masked language model in the Bert pre - training, the following two output formats are designed.

[0070] Output for the Task of Predicting the Following Sentence: The task of predicting the following sentence is a simple binary classification problem, which considers the sentence - level semantic relationship feature. Take the hidden layer vector corresponding to the [CLS] position, with a dimension of 1024. After passing through the [1024, 2] linear layer and using the softmax function, the prediction of the relationship between the two sentences can be obtained.

[0071] Output for the Masked Language Model Task: The purpose of the masked language model task is to predict the word corresponding to each position in the sentence, which can be understood as a multi - classification problem. That is, take the hidden layer vector corresponding to each word position, multiply it by the transposed matrix of the word vector matrix after passing through the [1024, 768] fully - connected layer, and use the softmax function to obtain a vector with a dimension equal to the number of word segmentations, representing the probability of each word.

[0072] Model training: Select the Adam optimizer to train the Bert model. Set the maximum number of training times to 100, the model learning rate to 10e-5, and the dropout random loss probability to 0.2. Observe the convergence of the loss value during the training process. You can appropriately adjust the learning rate according to the situation until the loss value converges, and then save the model.

[0073] So far, the military pre-trained model training is completed.

[0074] Step 2: Offline construction of the dual semantic retrieval model: Construct a military semantic retrieval dataset. Based on the military pre-trained model, construct a dual semantic retrieval model, and perform training and fine-tuning on the military semantic retrieval dataset to generate a question-answer pair language representation model, including a question language representation model and an answer language representation model; collect a set of military data texts to be retrieved, and the set of military data texts to be retrieved is military data texts collected by users according to business requirements; for the set of military data texts to be retrieved, use the answer language representation model to offline generate a semantic vector library of military text data, and use a clustering algorithm to construct a secondary inverted index to improve the retrieval efficiency.

[0075] Step 2-1: Construction of the military semantic retrieval dataset: Step 2-1-1, select question-answer pair data from the original corpus data and use it as positive samples, with the similarity set to 1. Step 2-1-2, in order to distinguish negative samples with relatively high similarity to the answer from positive samples as much as possible and improve the robustness of the model. The present invention uses the following two methods to construct negative samples. ① Relevant but non-optimal answer samples: Use the BM25 algorithm to retrieve the top three texts most relevant to the question except the answer from the original corpus, with the similarities set to 0.8, 0.5, and 0.3 respectively; ② Irrelevant samples: Use the random sampling method to randomly select two text data from the original corpus data (except positive samples and method ①) or other question-answer pairs, with the similarity set to 0. Step 2-1-3, for the set of positive and negative samples of the question-answer pair, use the list of word tables in the military context obtained in step 1-2 to map and transform the question-answer pair text data, so as to construct a military semantic retrieval dataset.

[0076] Step 2-2: Construction of the dual semantic retrieval model: Step 2-2-1, for the military semantic retrieval task, construct two input branch networks, that is, the dual semantic retrieval model. Specifically, use the military pre-trained model in step 1 to encode the question-answer pairs in the military semantic retrieval dataset to obtain the feature representations of the question and answer sentences, and use vector similarity calculation to obtain the similarity between the question and answer sentences. You can choose vector dot product, Euclidean distance, cosine distance, etc. as the similarity calculation method. In this embodiment, after the inner product of the semantic vectors of the question-answer pair, use the Sigmoid function to obtain the similarity between the question and answer sentences. The specific network model is as Figure 3 shown, specifically divided into three parts.

[0077] (1) Question language representation model: Use the main network of the military pre-trained model (the non-input / output part of the model) to represent the features of the question. The input preprocessing operations are the same as those in steps 1-4. Select the word vector of the question as the input of the question language representation model, and its dimension is 768.

[0078] (2) Answer language representation model: Use the main network of the military pre-trained model (the non-input / output part of the model) to represent the features of the candidate answers. The input preprocessing operations are the same as those in steps 1-4. Select the word vector of the answer as the input of the answer language representation model, and its dimension is 768.

[0079] (3) Similar pair calculation: Select the sentence-level feature representation of the question and answer (the output vector at the [CLS] position of the question and answer language representation model), and use the vector dot product to obtain the similarity of the question and answer pair.

[0080] Step 2-2-2, Construction of the question and answer pair language representation model: Since the training tasks are different, the pre-trained model cannot be directly used as the question and answer pair language representation model. To improve the accuracy of semantic retrieval, it is necessary to perform fine-tuning training on the dual semantic retrieval model on the military semantic retrieval dataset. Select the Adam optimizer to perform fine-tuning training on the dual semantic retrieval model, set the number of training times to 5, the model learning rate to 10e-5, and the dropout random loss probability to 0.5. After the model training is completed, the question and answer pair language representation model can be obtained.

[0081] Step 2-3: Vector representation of the military data text collection to be retrieved: For each piece of corpus in the military data text collection to be retrieved, perform preprocessing using the same input preprocessing operations as in steps 1-4, and then use the answer language representation model in step 2-2 to perform vector characterization on each piece of corpus to obtain the military text data semantic vector library.

[0082] Step 2-4: Construction of the secondary inverted index of the text to be retrieved: Use the K-means clustering algorithm to perform clustering analysis on the military text data semantic vector library and construct a secondary inverted index to improve the vector retrieval efficiency. The K-means algorithm is as follows in the table:

[0083] Table 2 Pseudo-code of the K-Means algorithm

[0084]

[0085]

[0086] Step 2-4-1: Use the K-means algorithm with the Euclidean distance as the distance formula to perform preliminary clustering on the military text data semantic vector library. For the military corpus data to be retrieved in the millions, set the number of clusters C to 1000. Use the corresponding category serial number and the class center point vector of each semantic vector as the first-level index;

[0087] Step 2-4-2: For each semantic vector set, use the K-means algorithm with Euclidean distance as the distance formula to cluster the military text data semantic vector library again, and set the number of subdivision categories to K, which is not less than 10. The subdivision category number and subclass center point vector corresponding to each semantic vector are used as secondary indexes.

[0088] Step 3: Offline construction of text retrieval and ranking model: Build a military semantic retrieval and ranking dataset, build a multi-classification ranking model based on the military pre-trained model, and train and fine-tune it on the military semantic retrieval and ranking dataset to generate a text retrieval and ranking model.

[0089] Step 3-1: Construct a military semantic retrieval ranking dataset.

[0090] Step 3-1-1: Traverse each question in the military semantic retrieval dataset and use the question language representation model generated in step 2-2 to obtain the question semantic vector representation;

[0091] Step 3-1-2: For each question semantic vector representation, use the secondary inverted index method to quickly locate the text set (Top-N) that is strongly related to the question semantics from the military text data semantic vector library. N represents the number of recalled texts, which depends on the specific application scenario. In the military scenario, if the user wants to retrieve M texts that are strongly related to the demand, and M ranges from 1 to 5, then the number of recalled texts N = 10*M. The specific steps are as follows:

[0092] Step 3-1-2-1: Traverse the C class center vectors in step 2-5-1, use the vector dot product to calculate the similarity between each class center point and the question vector, and sort them by similarity. Select the class with the top (2*M) similarity ranking as the next retrieval target. M is the number of texts the user wants to obtain. The top (2*M) basically covers the texts that the user wants to retrieve and are strongly related to the needs.

[0093] Step 3-1-2-2: For the M classes in step 3-1-2-1, traverse the K subclasses corresponding to each class in step 2-5-2, and use the vector dot product to calculate the similarity between the center point of each character class and the question vector. Sort the similarities of the 2*M*K subclass center points from large to small, and select the top (4*10*M) subclasses as the candidate set of retrieval texts.

[0094] Step 3-1-2-3: Traverse the 4*10*M subcategories in step 3-1-2-2, use vector dot product to calculate the similarity between each vector and the question vector, and sort them from large to small by similarity. Select the top (N) texts as the full set of recalled texts, where N is 10*M.

[0095] Step 3-1-3: For the N recall results of each question sentence, if the answer text in the original Q&A pair exists in the recall results, then this answer text is a positive sample, and the label is set to 1. Define the other recall results as negative samples, and set the corresponding label of each result to 0. If the correct answer text does not exist in the recall results, randomly delete one recall result, define the other recall results as negative samples, and set the correct answer text as the positive sample.

[0096] Step 3-1-4: Use the military vocabulary list to map and convert the positive and negative sample text data to construct a refined ranking dataset for military semantic retrieval. The data of this refined ranking dataset for military semantic retrieval includes the question sentence and N recall texts, and the label is an N-dimensional "01" vector, where "1" represents the text answer that best matches the question sentence.

[0097] Step 3-2: Construction of the text retrieval refined ranking model: Based on the military pre-trained model in Step 1, construct a multi-class text refined ranking model. The specific network structure is as Figure 4 shown, specifically including:

[0098] (1) Model input: This model takes the question sentence and the entire set of potential answers (N recall texts) as input. For each question sentence and answer set, perform the following data preprocessing: (1) Concatenate the 1+N text levels to form a data corpus. When concatenating, add the "[CLS]" flag at the head of the question sentence and the "[SEP]" flag at the end of each corpus; (2) Use the military vocabulary list generated in Step 1-2 for mapping and conversion to digitalize the input corpus. That is, by looking up the word segmentation list in Step 1-2, obtain the id (ordinal number) of the word, and then query the word vector table (random vector representation) to convert each word into a vector representation, which is used as the model input.

[0099] (2) Main network structure:

[0100] Use the main network structure of the military pre-trained model (the non-input and output parts of the model) as the main network structure of the text retrieval refined ranking model.

[0101] (3) Model output layer:

[0102] Construct a fully connected classification network at the model output position corresponding to [CLS]: Specifically, use two fully connected layers of [1024, 1024] and [1024, N] for feature extraction and dimensionality reduction, and use the softmax function to discriminate the best matching answer.

[0103] Due to different training tasks, it is necessary to perform model training and fine-tuning on the military semantic retrieval and sorting dataset. The present invention selects the Adam optimizer to perform fine-tuning training on the multi-classification text sorting model, sets the number of training times to 5, the model learning rate to 10e-5, and the dropout random loss probability to 0.2. After the multi-classification text sorting model training is completed, a text retrieval and sorting model is generated.

[0104] Step 4: Text semantic retrieval for real-time tasks: Input user data requirements, first use the question language representation model generated in step 2 to obtain the question semantic representation vector; then obtain the text collection that is strongly related to the user's needs through vector similarity calculation and retrieval; finally, use the text retrieval and sorting model in step 3 to obtain the text answer that best meets the requirements and feedback it to the user. The specific steps are as follows:

[0105] Step 4-1: According to the user data requirements, after data cleaning of the question, use the question language representation model generated in step 2-2 to obtain the question semantic vector representation;

[0106] Step 4-2: Use vector similarity to retrieve and recall text sets that are strongly related to user needs;

[0107] Step 4-2-1: Traverse the C class center point vectors in step 2-4-1, use the vector dot product to calculate the similarity between each class center point and the question vector, and sort them by similarity. Select the class with the top (2*M) similarity ranking as the next retrieval target. M is the number of texts the user wants to obtain. The top (2*M) basically covers the texts that the user wants to retrieve and are strongly related to the needs.

[0108] Step 4-2-2: For the 2*M classes in step 4-2-1, traverse the K subclasses respectively, and use the vector dot product to calculate the similarity between the center point of each word class and the question vector. Sort the similarities of the center points of the 2*M*K subclasses from large to small, and select the top (4*10*M) subclasses as the candidate set of retrieval texts.

[0109] Step 4-2-3: Traverse step 4-2-2 to retrieve each text semantic vector in the text candidate set, use the vector dot product to calculate the similarity between each vector and the question vector, and sort them from large to small by similarity, and select the Top (N) texts as the full set of recalled texts, where N is 10*M.

[0110] Step 4-3: Use the military vocabulary list to map and transform the question and the recall results of step 4-2-4, and use this as input to use the text retrieval and ranking model to accurately locate the answer text and feedback it to the user.

[0111] like Figure 1As shown in the figure, for the real-time data requirements of users, first, the model is represented by question language to obtain the semantic vector representation of the data requirements; second, vector similarity calculation and retrieval are used to recall a text set strongly relevant to the user's needs from the military text semantic vector library to be searched; then, the text retrieval re-ranking model is used to obtain the best text answer; finally, the corresponding original text data is fed back to the user.

[0112] Principle of this embodiment:

[0113] This embodiment fully draws on the powerful language learning pre-training models in the field of natural language processing and designs a text semantic retrieval method in a military scenario. First, on the basis of using the pre-training model to represent the semantic features of questions, vector similarity calculation and retrieval are used to improve the recall rate of text retrieval in a military scenario from the semantic level; second, the text retrieval re-ranking model is used to accurately locate the answers required by users and improve the precision rate of text retrieval in a military scenario, further improving the user experience.

[0114] Specifically: Offline part Step 1: Construction of military pre-training model part: Select the BERT model as the initial pre-training model, and obtain the original corpus by collecting text data in various information systems, question-and-answer systems, dialogue systems, historical file libraries, etc. in the military field. And for the BERT model, the language model and the next sentence prediction model are masked to construct a military text corpus dataset, and a military pre-training model is generated through model training. This model can be used as a general model for text processing in the military field; Step 2: Generation of question-and-answer pair language representation model and text corpus text data semantic vector library: Based on the military pre-training model, a dual semantic retrieval model is constructed, and a question-and-answer pair language representation model is generated through training in the military semantic retrieval dataset, which can fully represent the semantic features of question-and-answer pairs and improve the recall rate of text retrieval in a military scenario from the semantic level; Construction of the secondary index of the text library to be retrieved: Select the K-means algorithm to perform secondary clustering on the text semantic vector library to be retrieved, and generate a secondary inverted index of the text library to improve the speed of vector retrieval text recall; Step 3: Construction of text retrieval re-ranking model: Based on the military pre-training model, a text retrieval re-ranking model is constructed to select the text that best meets the user's needs from the recalled texts and improve the precision rate of text retrieval in a military scenario.

[0115] In summary, the present invention provides a text semantic retrieval method in a military scenario. To fully consider the text semantic features and avoid the problem of poor text retrieval effect when the user's demand description is unclear, the pre-training model is fully utilized to obtain the semantic-level features between question-and-answer pairs, and through vector similarity calculation and search technology, it helps various military users to accurately locate the required data from the massive text data and improve the user experience.

[0116] Its main contributions are as follows: First, it proposes a method for constructing a military pre-trained model. By collecting various types of text data in military information systems and the military domain on the Internet, an original corpus is obtained. Based on open-source pre-trained language models, a military pre-trained model is trained. As a general model, this model can be used for natural language processing tasks in subsequent military scenarios such as event classification, event extraction, sentiment extraction, and intelligent question answering.

[0117] Second, regarding the problem of text retrieval in military scenarios, it proposes a method for constructing a dual semantic retrieval model. Based on the military pre-trained model, a question-answer pair language model is constructed respectively to obtain the sentence-level features of the question-answer pair. Using the method of calculating the similarity of feature vectors, a text set (Top-K) strongly related to the question is obtained. Through training and fine-tuning on the military semantic retrieval data set, this method finally obtains a question-answer pair language representation model. Using the answer text language representation model, a vector representation library of the entire text set can be obtained offline.

[0118] Third, to improve the retrieval efficiency of massive military text data, a method for constructing a two-level inverted index is proposed. The K-means algorithm is used to perform secondary clustering on the semantic vector library of the text to be retrieved. Based on the results of the secondary clustering, a two-level inverted index of the text library is constructed, greatly improving the text retrieval efficiency.

[0119] Fourth, to meet the need for accurately positioning text answers in military scenarios, a method for constructing a multi-class fine-rank model is proposed. This model takes the pre-trained model as the main body, takes the question and the potential answer set as the input, and takes the probability of whether the potential answer set is the best answer as the output. Through fine-tuning training on the military semantic retrieval fine-rank data set, this method finally obtains a fine-rank model. For real-time search tasks, the question feature vector representation is obtained using the question language representation model. Through similarity calculation and retrieval, a potential answer set is generated, and then through the text retrieval fine-rank model, the information required by the user is accurately positioned in real time, improving the efficiency of the user's information retrieval.

[0120] In specific implementation, this application provides a computer storage medium and a corresponding data processing unit. Among them, this computer storage medium can store a computer program. When the computer program is executed by the data processing unit, it can run the content of the invention of a method for text semantic retrieval in military scenarios and some or all of the steps in each embodiment provided by the present invention. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0121] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a computer program, that is, a software product. This computer program software product can be stored in a storage medium, including several instructions to enable a device including a data processing unit (which can be a personal computer, server, single-chip microcomputer, MUU or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0122] The present invention provides a method for text semantic retrieval in a military scenario. There are many methods and ways to specifically implement this technical solution. The above is only the specific implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A text semantic retrieval method in a military scenario, characterized in that, It includes the following steps: Step 1. Offline construction of military pre-trained model: Construct a military text corpus dataset; select an open-source pre-trained model and train it in the military text corpus dataset to form a military pre-trained model; Step 2. Offline construction of dual semantic retrieval model: Construct a military semantic retrieval dataset; based on the military pre-trained model, construct a dual semantic retrieval model, train and fine-tune it on the military semantic retrieval dataset to generate a question-answer pair language representation model, including a question language representation model and an answer language representation model; collect the set of military data texts to be retrieved, and for the set of military data texts to be retrieved, use the answer language representation model to offline generate a semantic vector library of military text data, and use a clustering algorithm to construct a secondary inverted index; Step 3. Offline construction of text retrieval fine-ranking model: Construct a military semantic retrieval fine-ranking dataset, based on the military pre-trained model, construct a multi-classification fine-ranking model, and train and fine-tune it on the military semantic retrieval fine-ranking dataset to generate a text retrieval fine-ranking model; Step 4. Text semantic retrieval for real-time tasks: Input the user's data requirements. First, use the question language representation model generated in Step 2 to obtain the question semantic representation vector; then, through vector similarity calculation and retrieval, obtain a text set strongly relevant to the user's requirements; finally, use the text retrieval fine-ranking model in Step 3 to obtain the text answer that meets the requirements and feedback it to the user.

2. The text semantic retrieval method in a military scenario according to claim 1, wherein The offline construction of the military pre-trained model in Step 1 includes the following steps: Step 1-1. Collect military raw corpus data; Step 1-2. Perform data preprocessing such as cleaning and converting redundant characters, stop words, and simplified and traditional Chinese characters in the military raw corpus data, perform word segmentation on the preprocessed data, and collect the word list data in the semantic word library, synonym forest, related word library, and extended word library in the existing military information retrieval and intelligent question-answer systems to form a military word list; Step 1-3. Select a pre-trained model in the field of natural language processing; for each piece of military raw corpus, use the military word list for mapping and conversion, that is, find the corresponding ordinal position of each word and digitize it to construct a military text corpus dataset; Step 1-4. Based on the military text corpus dataset, set the model training parameters to train the pre-trained model to form a military pre-trained model.

3. The method for text semantic retrieval in a military scenario according to claim 2, wherein, The offline construction of the dual semantic retrieval model in Step 2 includes the following steps: Step 2-1. Collect military retrieval question-answer corpus data, perform digital conversion on the text data through the military word list mapping, and construct a military semantic retrieval dataset; Step 2-2. Based on the military pre-trained model, construct a dual semantic retrieval model. After training the dual semantic retrieval model, obtain a question-answer pair language representation model, and the two language representation models are the left and right branch network models of the dual semantic retrieval model respectively; Step 2-3. For the set of military data texts to be retrieved, use the answer language representation model to generate a semantic vector library of military text data; Step 2-4. For the semantic vector library of military text data, use a clustering algorithm to construct a secondary inverted index.

4. A method for text semantic retrieval in a military scenario according to claim 3, characterized in that Step 2-1 For the text retrieval task, mainly retrieve the Q&A corpus data from the original corpus data, and construct a military semantic retrieval dataset, including the following steps: Step 2-1-1: Set the Q&A pair data in the original corpus as positive samples, and set the label to 1; Step 2-1-2: For each question in the positive samples, construct negative samples in two ways: ① Related but non-optimal answer samples: Use the BM25 algorithm to retrieve the top three texts most relevant to the question except the answer from the original corpus data, and set the similarity degrees to 0.8, 0.5, and 0.3 respectively; ② Unrelated samples: Use the random sampling method to randomly select two text data from the military text corpus except the positive samples and those in method ① or other Q&A pairs, and set the similarity degree to 0; Step 2-1-3: Use the military word list to map and transform the text data in the positive and negative sample sets to construct a military semantic retrieval dataset.

5. A method for text semantic retrieval in a military scenario according to claim 4, characterized in that Step 2-2 Based on the military pre-trained model, construct a dual semantic retrieval model, and obtain a Q&A pair language representation model after model training, including the following steps: Step 2-2-1: For the military semantic retrieval task, construct two input branch networks, namely the dual semantic retrieval model; use the military pre-trained model generated in Step 1 to encode the military Q&A pair data respectively to obtain the feature representations of the questions and answers, and then use vector similarity calculation to obtain the similarity between the Q&A sentences; Step 2-2-2: Perform training and fine-tuning on the dual semantic retrieval model on the military semantic retrieval dataset. After the dual semantic retrieval model is trained, obtain the Q&A pair language representation model.

6. The text semantic retrieval method in a military scenario according to claim 5, characterized in that Step 2-4 For the military text data semantic vector library, use the clustering algorithm to construct a secondary inverted index, including the following steps: Step 2-4-1: Use the K-means algorithm with the Euclidean distance as the distance formula to perform preliminary clustering on the military text data semantic vector library, and set the initial number of clustering categories to C; Take the category serial number corresponding to each semantic vector and the category center point vector as the first-level index; Step 2-4-2: For each set of semantic vectors of each category, use the K-means algorithm with the Euclidean distance as the distance formula to perform re-clustering and division on the military text data semantic vector library, and set the number of detailed categories to K, where K is not less than 10; Take the detailed category serial number corresponding to each semantic vector and the sub-category center point vector as the second-level index.

7. A method for text semantic retrieval in a military scenario according to claim 6, characterized in that Step 3 Offline construction of the text retrieval re-ranking model includes Step 3-1: Construct a military semantic retrieval re-ranking dataset. Step 3-1 includes the following steps: Step 3-1-1: Traverse each question in the military semantic retrieval dataset, and use the question language representation model generated in Step 2-2 to obtain the semantic vector representation of the question; Step 3-1-2: For each semantic vector representation of the question, adopt the secondary inverted index method to quickly locate the Top-N text set strongly relevant to the question semantics from the military text data semantic vector library; N represents the number of retrieved texts; in the military scenario, if the user hopes to retrieve M texts strongly relevant to the demand, and the value range of M is 1-5, then the number of retrieved texts N = 10*M; Step 3-1-3: For each question sentence, among the N recall results, if the answer text in the original question-answer pair exists in the recall result, the answer text is a positive sample, and the label is set to 1. The other recall results are defined as negative samples, and the corresponding labels of each result are set to 0. If the correct answer text does not exist in the recall result, a recall result is randomly deleted, and the other recall results are defined as negative samples, and the correct answer text is set as a positive sample. Step 3-1-4: Use the military vocabulary list to map and transform the positive and negative sample text data to construct a military semantic retrieval and ranking dataset. The military semantic retrieval and ranking dataset includes questions and N recalled texts. The labels are N-dimensional 01 vectors, where 1 represents the text answer that best matches the question.

8. A method for text semantic retrieval in a military scenario according to claim 7, characterized in that, Step 3: Offline construction of the text retrieval and ranking model includes step 3-2: Based on the military pre-training model, a text retrieval and ranking model is constructed. Step 3-2 includes the following steps: Step 3-2-1: Using the military pre-trained model in step 1 as the backbone, build a text retrieval and ranking model; this model takes the question and N recalled texts as input, and outputs an N-dimensional classification vector to identify the text that best matches the question among the potential answers; Step 3-2-2: Fine-tune the text retrieval ranking model on the military semantic retrieval ranking dataset obtained in step 3-1.

9. The method for text semantic retrieval in a military scenario according to claim 8, wherein Step 4: Text semantic retrieval for real-time tasks, including the following steps: Step 4-1: When facing real-time search tasks, after data cleaning of the question, use the question language representation model to obtain the question semantic vector representation; Step 4-2: Use vector similarity to retrieve and recall text sets that are strongly related to user needs; Step 4-3: Use the military vocabulary list to map and transform the question and the recall results of step 4-2, and use this as input to use the text retrieval and ranking model to accurately locate the answer text that meets the user's needs and feedback it to the user.

10. A method for text semantic retrieval in a military scenario according to claim 9, characterized in that Step 4-2 uses vector similarity to retrieve and recall text sets that are strongly related to user needs, including the following steps: Step 4-2-1: Traverse the C class center point vectors in step 2-4-1, use the vector dot product to calculate the similarity between each class center point and the question semantic vector, sort them by similarity, and select the class with the top (2*M) similarity ranking as the next retrieval target; Step 4-2-2: For the 2*M classes in step 4-2-1, traverse the K subclasses respectively, and use the vector dot product to calculate the similarity between the center point of each subclass and the question semantic vector; sort the similarities of the center points of the 2*M*K subclasses from large to small, and select the top (4*10*M) subclasses as the candidate set of retrieval texts; Step 4-2-3: Traverse step 4-2-2 to retrieve each text semantic vector in the text candidate set, use vector dot product to calculate the similarity between each vector and the question semantic vector, and sort them from large to small by similarity, and select the Top (N) texts as the full set of recalled texts.

Citation Information

Patent Citations

  • Public opinion clustering method, device and equipment

    CN113032566A

  • Intelligent question and answer method, device and equipment and storage medium

    CN114416927A