Text semantic matching method and device for medical intelligent question answering
By constructing an embedding layer, a semantic coding layer, a multi-level fine-grained feature extraction layer and a feature fusion layer, a two-way long and short-term memory network and a variety of attention mechanisms, the problem of insufficient semantic information capture in Chinese text semantic matching is solved, and the matching accuracy is improved.
Patent Information
- Application Number
- CN202210996504.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-19
AI Technical Summary
The prior art is difficult to effectively capture the semantic information of Chinese text at word granularity, word granularity and text level, resulting in insufficient semantic matching accuracy of Chinese text.
By constructing an embedding layer, a semantic coding layer, a multi-level fine-grained feature extraction layer and a feature fusion layer, a bidirectional long and short-term memory network and a variety of attention mechanisms, the fine-grained semantic features of the text and the semantic interaction features between text are extracted to generate the final matching feature vector.
It improves the accuracy of semantic matching in Chinese text, can better capture semantic dependencies within and between text, and enhances the prediction accuracy of the model.
Smart Images

Figure CN115269808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and natural language processing, and particularly relates to a text semantic matching method and device for medical intelligent question answering. Background Art
[0002] Medical intelligent question answering can automatically find questions with similar semantics in a question-answering knowledge base for the questions raised by patients and push the answers to users, greatly reducing the burden of manual replies from doctors; for the thousands of different questions raised by patients, how to find standard questions with similar semantics to them is the core of a medical intelligent question-answering system; the essence of this technology is to measure the matching degree between the questions raised by patients and the standard questions in the question-answering knowledge base, and its essence is a text semantic matching task.
[0003] The text semantic matching task aims to measure whether the semantics contained in two texts are consistent, which is consistent with the core objectives of many natural language processing tasks. Calculating the semantic matching degree of natural language texts is a very challenging task, and existing methods have not yet been able to perfectly solve this problem.
[0004] Existing methods mainly focus on the similarity discrimination of English texts, model the semantic information within the same text at the word granularity level, and model the semantic interaction information between texts at the text level. However, Chinese texts are more complex than English texts. Chinese has rich semantic information at both the character granularity level and the word granularity level. How to better capture the semantic information at the character granularity, word granularity, and text level to better determine the semantic similarity between texts is a challenging task; aiming at the deficiencies of existing text semantic matching methods, the present invention proposes a text semantic matching method and device for medical intelligent question answering; this method captures the fine-grained semantic features of the same text and the semantic interaction features between texts at multiple levels; its core idea is to extract the character and word granularity features of the text through a multi-layer encoding structure combined with multiple attention mechanisms, capture the fine-grained semantic features of the same text and the semantic interaction features between texts, finally merge various relevant features, and then perform various matching operations to generate a final matching feature vector and determine the similarity of the texts. Summary of the Invention
[0005] The technical task of the present invention is to provide a text semantic matching method and device for medical intelligent question answering, which extracts the character and word granularity features of the text, captures the fine-grained semantic features of the same text and the semantic interaction features between texts, finally merges various relevant features, and then performs various matching operations to generate a final matching feature vector and determine the similarity of the texts.
[0006] The technical task of the present invention is achieved in the following way. A text semantic matching method for medical intelligent question answering. This method constructs and trains a text semantic matching model composed of an embedding layer, a semantic encoding layer, a multi-level fine-grained feature extraction layer, a feature fusion layer, and a prediction layer to extract text character and word granularity features, capture the fine-grained semantic features of the same text and the semantic interaction features between texts, and finally merge various relevant features, and then perform various matching operations to generate a final matching feature vector and determine the similarity of the texts. Specifically as follows:
[0007] The embedding layer performs embedding operations on the input text at the character granularity and the word granularity respectively to obtain the text character embedding representation and the word embedding representation.
[0008] The semantic encoding layer receives the text character embedding representation and the word embedding representation and uses a bidirectional long short-term memory network (BiLSTM) for encoding, and outputs the text character and word granularity features.
[0009] The multi-level fine-grained feature extraction layer performs same-text and inter-text encoding operations on the text character and word granularity features output by the semantic encoding layer to obtain the fine-grained semantic features of the same text and the semantic interaction features between texts.
[0010] The feature fusion layer merges the relevant features, and then performs various matching operations to generate a final matching feature vector.
[0011] The prediction layer inputs the final matching feature vector into a multi-layer perceptron to obtain a floating-point numerical value, compares it with a preset threshold, and determines whether the semantics of the text match according to the comparison result.
[0012] Preferably, the embedding layer includes a character-word mapping conversion table, an input layer, and a character-word vector mapping layer, and outputs the text character embedding representation and the word embedding representation.
[0013] Among them, the character-word mapping conversion table: the mapping rule starts with the number 1, and then increases sequentially according to the order in which each character or word is entered into the character-word table, so as to form a character-word mapping conversion table. Then, use Word2Vec to train the character-word vector model to obtain the character-word vector matrix of each character and word.
[0014] Input layer: The input layer includes four inputs. For each text in the training dataset or the text to be predicted, word breaking and word segmentation preprocessing are performed to obtain txt P_char, txt Q_char, txt P_word, and txt Q_word respectively, where the suffixes char and word indicate that the corresponding text has undergone word breaking or word segmentation processing. They are formalized as: (txt P_char, txt Q_char, txt P_word, txt Q_word); each character and word in the input text is converted into a corresponding digital identifier according to the word and character mapping conversion table;
[0015] Word and character vector mapping layer: Load the word and character vector matrix obtained during the training of the word and character mapping conversion table construction step to initialize the weight parameters of the current layer; for the input texts txt P_char, txt_Q_char, txt P_word, and txt Q_word, obtain their corresponding text character embeddings and word embeddings txt P_char_embed, txt Q_char_embed, txt P_word_embed, and txt Q_word_embed
[0016] Preferably, the implementation details of the semantic encoding layer are as follows:
[0017] Taking text P as an example, this module receives the text P character and word embeddings and uses a bidirectional long short-term memory network BiLSTM for encoding to obtain the text P character and word granularity features, denoted as The specific formula is as follows:
[0018]
[0019]
[0020] Among them, N represents the length of the character granularity feature and the word granularity feature. Formula (1) represents encoding the text P character embedding using the bidirectional long short-term memory network BiLSTM, where, represents the text P character granularity feature at the i-th position obtained by encoding through the bidirectional long short-term memory network BiLSTM, represents the text P character granularity feature at the i-th position obtained by encoding through the forward long short-term memory network LSTM, represents the text P character granularity feature at the i-th position obtained by encoding through the backward LSTM; the meanings of the symbols in Formula (2) are basically the same as those in Formula (1), represents the text P word granularity feature at the j-th position obtained by encoding through the bidirectional long short-term memory network BiLSTM, It represents the word-grained feature at the j-th position of text P obtained by forward LSTM encoding. It represents the word-grained feature at the j-th position of text P obtained by backward LSTM encoding.
[0021] Similarly, similar operations to those on text P are performed on text Q to obtain the character- and word-grained features of text Q, denoted as Q c and Q w .
[0022] More preferably, the implementation details of the multi-level fine-grained feature extraction layer are as follows:
[0023] For the character- and word-grained features of the text output by the semantic encoding layer, intra-text and inter-text encoding operations are performed to obtain intra-text fine-grained semantic features and inter-text semantic interaction features; it includes two sub-modules. The first sub-module is responsible for extracting intra-text fine-grained semantic features, mainly using multiple attention modules to encode different grain sizes of the same text to obtain intra-text fine-grained semantic features; the second sub-module is responsible for extracting inter-text semantic interaction features, mainly using several layers of encoding structures between texts to obtain inter-text semantic interaction features.
[0024] Extraction of intra-text fine-grained semantic features by the first sub-module:
[0025] First, for the convenience of subsequent description, in the first part, text P is taken as an example to define the following attention module:
[0026] Define the soft alignment attention module, denoted as SOA, and the formula is as follows:
[0027]
[0028] Where represents the character-grained feature at the i-th position of text P, from formula (1), represents the word-grained feature at the j-th position of text P, from formula (2), represents the soft alignment attention weight between the character-grained feature at the i-th position and the word-grained feature at the j-th position of text P, represents mapping the soft alignment attention weight by softmax operation to a value between 0 and 1, represents using the soft alignment attention to re-represent the character-grained feature at the i-th position of text P by weighted summation of all word-grained features of text P, represents using the soft alignment attention to re-represent the word-grained feature at the j-th position of text P by weighted summation of all character-grained features of text P;
[0029] Define the multiplicative alignment attention module, denoted as MUA, and the formula is as follows:
[0030]
[0031] Among them, TimeDistributed(Dense()) means that the same Dense() layer operation is performed on the tensors of each time step, ⊙ represents element-wise multiplication operation, tanh represents the activation function, and P c represents the character-level feature of text P, represents the character-level feature of text P after being processed by the Dense() layer, and P w represents the word-level feature of text P, represents the multiplicative alignment attention weight, represents mapping the multiplicative alignment attention weight to a value between 0 and 1 through the softmax operation, represents using multiplicative alignment attention to re-represent the character-level feature at the i-th position of text P as a weighted sum of all word-level features of text P, represents using multiplicative alignment attention to re-represent the word-level feature at the j-th position of text P as a weighted sum of all character-level features of text P;
[0032] Define the subtraction alignment attention module as SUA, and the formula is as follows:
[0033]
[0034] Among them, TimeDistributed(Dense()) means that the same Dense() layer operation is performed on the tensors of each time step, - represents element-wise subtraction operation, tanh represents the activation function, and P c represents the character-level feature of text P, represents the character-level feature of text P after being processed by the Dense() layer, and P w represents the word-level feature of text P, represents the subtraction alignment attention weight, represents mapping the subtraction alignment attention weight to a value between 0 and 1 through the softmax operation, represents using subtraction alignment attention to re-represent the character-level feature at the i-th position of text P as a weighted sum of all word-level features of text P, represents using subtraction alignment attention to re-represent the word-level feature at the j-th position of text P as a weighted sum of all character-level features of text P;
[0035] Define the self-alignment attention module as SEA, and the formula is as follows:
[0036]
[0037] Among them, It represents the character-level feature of the i-th position in text P. It represents the character-level feature of the j-th position in text P. It represents the self-alignment attention weight between the character-level feature of the i-th position in text P and the character-level feature of the j-th position in text P. It represents the value mapped to 0-1 by performing the softmax operation on the self-alignment attention weight. It represents using self-alignment attention to re-represent the character-level feature of the i-th position in text P by weighted summation of all character-level features in text P.
[0038] In subsequent descriptions, the operation of formula (3) is represented by the SOA symbol, the operation of formula (4) is represented by the MUA symbol, the operation of formula (5) is represented by the SUA symbol, and the operation of formula (6) is represented by the SEA symbol.
[0039] The first-layer encoding structure uses multiple attention modules to extract the fine-grained initial semantic features of the same text:
[0040] First, using soft alignment attention, the character-level feature P of text P c is soft-aligned with the word-level feature P of text P w to obtain the soft-aligned feature of text P at the character level and the soft-aligned feature of text P at the word level as shown in formula (7):
[0041]
[0042] Second, using multiplicative alignment attention, the character-level feature P of text P c is multiplicatively aligned with the word-level feature P of text P w to obtain the multiplicative alignment feature of text P at the character level and the multiplicative alignment feature of text P at the word level as shown in formula (8):
[0043]
[0044] After that, using subtractive alignment attention, the character-level feature P of text P c is subtractively aligned with the word-level feature P of text P w to obtain the subtractive alignment feature of text P at the character level and the subtractive alignment feature of text P at the word level as shown in formula (9):
[0045]
[0046] Similarly, performing a similar process on text Q as on text P yields the soft alignment features of text Q at the character granularity level. The soft alignment features of text Q at the word granularity level The multiplicative alignment features of text Q at the character granularity level The multiplicative alignment features of text Q at the word granularity level The subtractive alignment features of text Q at the character granularity level The subtractive alignment features of text Q at the word granularity level That is, the extraction of the fine-grained initial semantic features of the same text is completed;
[0047] The second-layer encoding structure enhances the fine-grained semantic features of the same text and completes the extraction of the fine-grained semantic features of the same text:
[0048] First, add the soft alignment features of text P at the character granularity level in formula (7) to the character granularity feature P of text P in formula (1) c to obtain the deep soft alignment features of text P at the character granularity level as shown in formula (10):
[0049]
[0050] Subsequently, add the multiplicative alignment features of text P at the character granularity level in formula (8) to the character granularity feature Pc of text P in formula (1) to obtain the deep multiplicative alignment features of text P at the character granularity level as shown in formula (11):
[0051]
[0052] Then, add the subtractive alignment features of text P at the character granularity level in formula (9) to the character granularity feature P of text P in formula (1) c to obtain the deep subtractive alignment features of text P at the character granularity level as shown in formula (12):
[0053]
[0054] Next, connect the deep soft alignment features of text P at the character granularity level in formula (10) the deep multiplicative alignment features of text P at the character granularity level in formula (11) the deep subtractive alignment features of text P at the character granularity level in formula (12) to obtain the high-level feature P' of text P at the character granularity level c , as shown in formula (13):
[0055]
[0056] Similar to the character level, first, the soft alignment feature of text P at the word level in formula (7) is added to the word-level feature P of text P in formula (2) w to obtain the deep soft alignment feature of text P at the word level As shown in formula (14):
[0057]
[0058] Subsequently, the multiplicative alignment feature of text P at the word level in formula (8) is added to the word-level feature P of text P in formula (2) w to obtain the deep multiplicative alignment feature of text P at the word level As shown in formula (15):
[0059]
[0060] Then, the subtractive alignment feature of text P at the word level in formula (9) is added to the word-level feature P of text P in formula (2) w to obtain the deep subtractive alignment feature of text P at the word level As shown in formula (16):
[0061]
[0062] Next, the deep soft alignment feature of text P at the word level in formula (14) the deep multiplicative alignment feature of text P at the word level in formula (15) the deep subtractive alignment feature of text P at the word level in formula (16) are concatenated to obtain the high-level feature P' of text P at the word level w , as shown in formula (17):
[0063]
[0064] The deep soft alignment feature of text P at the character level in formula (10) is concatenated with the deep soft alignment feature of text P at the word level in formula (14) to obtain the deep semantic feature P' of text P deep , as shown in formula (18):
[0065]
[0066] Similarly, perform similar processing on text Q as on text P to obtain the deep soft alignment features of text Q at the character granularity level. Deep multiplicative alignment features of text Q at the character granularity level Deep subtractive alignment features of text Q at the character granularity level High-level features Q' of text Q at the character granularity level c And the deep soft alignment features of text Q at the word granularity level Deep multiplicative alignment features of text Q at the word granularity level Deep subtractive alignment features of text Q at the word granularity level High-level features Q' of text Q at the word granularity level w And the deep semantic features Q' of text Q deep , which completes the extraction of fine-grained semantic features of the same text.
[0067] The second sub-module: extraction of inter-text semantic interaction features:
[0068] The first-layer encoding structure simultaneously uses several layers of encoding structures to extract the initial semantic interaction features between texts:
[0069] At the character granularity, first, take the character granularity feature P of text P in formula (1) c and the character granularity feature Q of text Q c to perform soft alignment attention to obtain the soft alignment interaction features of text P at the character granularity level and the soft alignment interaction features of text Q at the character granularity level As shown in formula (19):
[0070]
[0071] Secondly, take the character granularity feature P of text P in formula (1) c and the character granularity feature Q of text Q c to perform subtractive alignment attention to obtain the subtractive alignment interaction features of text P at the character granularity level and the subtractive alignment interaction features of text Q at the character granularity level As shown in formula (20):
[0072]
[0073] At the word granularity, it is similar to that at the character granularity. First, take the word granularity feature P of text P in formula (2) w and the word granularity feature Q of text Q w to perform soft alignment attention to obtain the soft alignment interaction features of text P at the word granularity level and the soft alignment interaction features of text Q at the word granularity level As shown in formula (21):
[0074]
[0075] Then, the word-grained feature P of text P in formula (2) w and the word-grained feature Q of text Q w are subjected to subtraction alignment attention to obtain the word-grained text P subtraction alignment interaction feature and the word-grained text Q subtraction alignment interaction feature As shown in formula (22):
[0076]
[0077] The second-layer encoding structure enhances the initial semantic interaction features between texts and completes the extraction of the semantic interaction features between texts:
[0078] At the character granularity, first, the character-grained soft alignment interaction feature of text P in formula (19) is added to the character-grained feature P of text P in formula (1) c to obtain the character-grained deep soft alignment interaction feature of text P As shown in formula (23):
[0079]
[0080] After that, the character-grained subtraction alignment interaction feature of text P in formula (20) is added to the character-grained feature P of text P in formula (1) c to obtain the character-grained deep subtraction alignment interaction feature of text P As shown in formula (24):
[0081]
[0082] Finally, the character-grained deep soft alignment interaction feature of text P in formula (23) is concatenated with the character-grained deep subtraction alignment interaction feature of text P in formula (24) to obtain the character-grained high-level interaction feature P′ of text P c as shown in formula (25):
[0083]
[0084] At the word granularity, first, the word-grained soft alignment interaction feature of text P in formula (21) is added to the word-grained feature of text P in formula (2) w to obtain the word-grained deep soft alignment interaction feature of text P As shown in formula (26):
[0085]
[0086] After that, subtract and align the interaction features of the word-grained text P in formula (22) from the word-grained feature P of the text P in formula (2) w and add them to obtain the deep subtraction alignment interaction feature of the word-grained text P as shown in formula (27):
[0087]
[0088] Finally, concatenate the deep soft alignment interaction feature of the word-grained text P in formula (26) with the deep subtraction alignment interaction feature of the word-grained text P in formula (27) to obtain the high-level interaction feature P″ of the word-grained text P w , as shown in formula (28):
[0089]
[0090] Concatenate the deep subtraction alignment interaction feature of the character-grained text P in formula (24) with the deep subtraction alignment interaction feature of the word-grained text P in formula (27) to obtain the deep semantic interaction feature P″ of the text P deep , as shown in formula (29):
[0091]
[0092] Similarly, perform a similar process on text Q as on text P to obtain the deep soft alignment interaction feature of the character-grained text Q the deep subtraction alignment interaction feature of the character-grained text Q the high-level interaction feature Q″ of the character-grained text Q c , the deep soft alignment interaction feature of the word-grained text Q the deep subtraction alignment interaction feature of the word-grained text Q the high-level interaction feature Q″ of the word-grained text Q w , and the deep semantic interaction feature Q″ of the text Q deep , i.e., the extraction of the semantic interaction features between texts is completed.
[0093] More preferably, the implementation details of the feature fusion layer are as follows:
[0094] First, for the convenience of subsequent description, the following definitions are made:
[0095] Define the operation of subtracting vectors and then taking the absolute value bit by bit as AB, as shown in formula (30):
[0096] AB(P, Q) = |P - Q| (30)
[0097] where P and Q are two different vectors, representing the operation of taking the absolute value bit by bit after subtracting the two vectors P and Q;
[0098] Define the bitwise multiplication operation of vectors as MU, as shown in formula (31):
[0099] MU(P, Q) = P ⊙ Q (31)
[0100] where P and Q are two different vectors, representing the operation of bitwise multiplication of vectors P and Q;
[0101] In the following description, the AB symbol is used to represent the operation of formula (30), the MU symbol is used to represent the operation of formula (31), the SOA symbol is used to represent the operation of formula (3), the MUA symbol is used to represent the operation of formula (4), the SUA symbol is used to represent the operation of formula (5), and the SEA symbol is used to represent the operation of formula (6);
[0102] The feature fusion layer is divided into two sub-modules. The first sub-module merges multiple relevant features, and the second sub-module performs multiple matching operations to obtain the final matching feature vector;
[0103] The first sub-module merges multiple relevant features:
[0104] Connect the high-level feature P of the word-grained text P in formula (13) c ' with the high-level interaction feature P of the word-grained text P in formula (25) c ' to obtain the aggregated feature of the word-grained text P And perform self-attention on the aggregated feature of the word-grained text P to obtain the deep aggregated feature of the word-grained text P As shown in formula (32):
[0105]
[0106] At the word level, similar to the word-grained level, connect the high-level feature P' of the word-level text P in formula (17) w with the high-level interaction feature P'' of the word-level text P in formula (28) w to obtain the aggregated feature of the word-level text P And perform self-attention on the aggregated feature of the word-level text P to obtain the deep aggregated feature of the word-level text P As shown in formula (33):
[0107]
[0108] After that, connect the deep aggregation feature of text P at the character granularity level in formula (32) with the deep aggregation feature of text P at the word granularity level in formula (33) Then perform a max pooling operation to obtain the pooled semantic feature P' of text P, as shown in formula (34):
[0109]
[0110] Next, connect the deep semantic feature P' of text P in formula (18) deep with the deep semantic interaction feature P'' of text P in formula (29) deep to obtain the deep aggregation feature of text P As shown in formula (35):
[0111]
[0112] Similarly, perform the same operations on text Q as on text P to obtain the aggregation feature of text Q at the character granularity level the deep aggregation feature of text Q at the character granularity level the aggregation feature of text Q at the word granularity level the deep aggregation feature of text Q at the word granularity level the pooled semantic feature Q' of text Q, the deep aggregation feature of text Q
[0113] Then, connect the deep aggregation feature of text P in formula (35) with the deep aggregation feature of text Q to perform soft alignment attention to obtain the soft-aligned deep aggregation feature of text P and the soft-aligned deep aggregation feature of text Q As shown in formula (36):
[0114]
[0115] Subsequently, perform a max pooling operation on the soft-aligned deep aggregation feature of text P in formula (36) to obtain the pooled deep aggregation feature P'' of text P, and perform a max pooling operation on the soft-aligned deep aggregation feature of text Q to obtain the pooled deep aggregation feature Q'' of text Q, as shown in formula (37):
[0116]
[0117] The second sub-module performs multiple matching operations to obtain the final matching feature vector:
[0118] First, the absolute value subtraction is performed on the semantic feature P' of the pooled text P and the semantic feature Q' of the pooled text Q in formula (34) to obtain the subtraction matching feature PQ ab , as shown in formula (38):
[0119] PQ ab = AB(P' - Q') (38)
[0120] Second, the dot product is performed on the semantic feature P' of the pooled text P and the semantic feature Q' of the pooled text Q in formula (34) to obtain the dot product matching feature PQ mu , as shown in formula (39):
[0121] PQ mu = MU(P', Q') (39)
[0122] Third, the absolute value subtraction is performed on the deep aggregation feature P″ of the pooled text P and the deep aggregation feature Q″ of the pooled text Q in formula (37) to obtain the deep subtraction matching feature PQ' ab , as shown in formula (40):
[0123] PQ' ab = AB(P″, Q″) (40)
[0124] After that, the dot product is performed on the deep aggregation feature P″ of the pooled text P and the deep aggregation feature Q″ of the pooled text Q in formula (37) to obtain the deep dot product matching feature PQ' mu , as shown in formula (41):
[0125] PQ' mu = MU(P″, Q″) (41)
[0126] Finally, the semantic feature P' of the pooled text P, the semantic feature Q' of the pooled text Q, the subtraction matching feature PQ in formula (38) ab , the dot product matching feature PQ in formula (39) mu , the deep subtraction matching feature PQ' in formula (40) ab , the deep dot product matching feature PQ' in formula (41) mu in formula (42) are concatenated to obtain the final matching feature vector F, as shown in formula (42):
[0127] F = [P'; Q'; PQ ab ; PQ mu ; PQ' ab ; PQ' mu (42)
[0128] Preferably, the implementation details of the prediction layer are as follows:
[0129] Taking the final matching feature vector F as the input, using a three-layer fully connected layer and applying the ReLU activation function after the first and second fully connected layers for activation, and applying the sigmoid function after the third fully connected layer for activation, so as to obtain a matching degree value between [0,1], denoted as y pred ; finally, by comparing with the set threshold of 0.5, it is judged whether the text semantics match; that is, y pred ≥0.5, it is predicted that the semantics of the text match, otherwise, they do not match; when the text semantics matching model has not been trained, it needs to be trained on the training dataset constructed according to the semantics matching knowledge base to optimize the model parameters; when the model is trained, the prediction layer can predict whether the semantics of the target text match.
[0130] Preferably, the text semantics matching knowledge base includes obtaining the original data by downloading the dataset on the network, preprocessing the original data, and summarizing the sub-knowledge bases;
[0131] Obtaining the original data by downloading the dataset on the network: Downloading the publicly available text semantics matching dataset or the manually constructed dataset on the network and using it as the original data for constructing the text semantics matching knowledge base;
[0132] Preprocessing the original data: Preprocessing the original data for constructing the text semantics matching knowledge base, performing word breaking and word segmentation operations on each text therein to obtain the text semantics matching word breaking processing knowledge base and the word segmentation processing knowledge base;
[0133] Summarizing the sub-knowledge bases: Summarizing the text semantics matching word breaking processing knowledge base and the text semantics matching word segmentation processing knowledge base to construct the text semantics matching knowledge base;
[0134] The text semantics matching model is obtained by training using the training dataset, and the construction process of the training dataset includes constructing training positive examples, constructing training negative examples, and constructing the training dataset;
[0135] Constructing training positive examples: For each text in the text semantics matching knowledge base, if its semantics are consistent, the text can be used to construct training positive examples;
[0136] Constructing training negative examples: Selecting a text txt P, and then randomly selecting a text txt Q that does not match the text txt P from the text semantics matching knowledge base, and combining txt P and txt Q to construct negative examples;
[0137] Construct the training dataset: Combine all the positive example data and negative example data obtained after the operations of constructing training positive examples and constructing training negative examples, and shuffle their order to construct the final training dataset;
[0138] After the text semantic matching model is constructed, it is trained and optimized through the training dataset, as follows:
[0139] Construct the loss function: As can be seen from the implementation process of the prediction layer, y pred is the calculated value of the matching degree obtained after being processed by the text semantic matching model; while y true is the true label of whether two text semantics match, and its value is limited to 0 or 1. Cross-entropy is used as the loss function;
[0140] Construct the optimization function: Use the Adam optimization function; on the training dataset, optimize and train the text semantic matching model.
[0141] A text semantic matching device for medical intelligent question answering, which includes a text semantic matching knowledge base construction unit, a training dataset generation unit, a text semantic matching model construction unit, and a text semantic matching model training unit;
[0142] The specific functions of each unit of the aggregated text knowledge base are as follows:
[0143] The text semantic matching knowledge base construction unit is used to obtain a large amount of text data, and then preprocess the text data to obtain a text semantic matching knowledge base that meets the training requirements;
[0144] The training dataset generation unit, for the data in the text semantic matching knowledge base, if its semantics are consistent, the text is used to construct training positive examples, otherwise, it is used to construct training negative examples. Mix all the positive example data and negative example data to obtain the training dataset;
[0145] The text semantic matching model construction unit is used to construct a word mapping conversion table, construct an input layer, construct a word vector mapping layer, construct a semantic encoding layer, construct a multi-level fine-grained feature extraction layer, construct a feature fusion layer, and construct a prediction layer;
[0146] The text semantic matching model training unit is used to construct a training loss function and an optimization function to complete the training of the model.
[0147] A storage medium in which multiple instructions are stored, and the instructions are loaded and executed by a processor to perform the steps of the above-mentioned text semantic matching method for medical intelligent question answering.
[0148] An electronic device, the electronic device includes:
[0149] The above storage medium; and
[0150] A processor for executing instructions in the storage medium.
[0151] The text semantic matching method and device for medical intelligent question answering of the present invention have the following advantages:
[0152] (1) By performing embedding operations at the word and character granularities of the text, semantic information contained in different granularities of the text is extracted, making the extracted semantic features more detailed and rich;
[0153] (2) By using a bidirectional long short-term memory network to perform semantic encoding on the text, bidirectional semantic dependencies of the text can be better captured;
[0154] (3) By constructing a fine-grained feature extraction layer, semantic features of different granularities and different levels can be captured, and as many semantic features of different granularities and deeper levels as possible can be extracted;
[0155] (4) By using an attention mechanism to perform semantic encoding on the text, dependency relationships between texts and between different granularities within the text can be effectively captured, making the generated text matching tensor have rich interaction features, thereby improving the prediction accuracy of the model;
[0156] (5) By performing a max pooling operation, invalid information in the matching tensor can be effectively filtered and valid information can be strengthened, making the matching process more accurate, thereby improving the accuracy of text semantic matching. Description of the Drawings
[0157] The present invention will be further described below with reference to the drawings.
[0158] Figure 1 It is a flowchart of a text semantic matching method for medical intelligent question answering;
[0159] Figure 2 It is a flowchart of constructing a text semantic matching knowledge base;
[0160] Figure 3 It is a flowchart of constructing a training data set;
[0161] Figure 4 It is a flowchart of constructing a text semantic matching model;
[0162] Figure 5 It is a flowchart of training a text semantic matching model;
[0163] Figure 6 It is a schematic diagram of a semantic encoding layer model (taking text P as an example);
[0164] Figure 7 It is a schematic diagram of the structure for extracting fine-grained semantic features of the same text (taking text P as an example);
[0165] Figure 8 Schematic diagram for extracting semantic interaction feature structure among texts
[0166] Figure 9 Schematic diagram of the feature fusion layer
[0167] Figure 10 Schematic diagram of the structure of a text semantic matching device for medical intelligent question answering Specific implementation manner
[0168] The text semantic matching method and device for medical intelligent question answering of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments
[0169] Embodiment 1
[0170] The main framework structure of the present invention includes an embedding layer, a semantic encoding layer, a multi-level fine-grained feature extraction layer, a feature fusion layer, and a prediction layer. Among them, the embedding layer performs embedding operations on the input text at the character granularity and word granularity respectively, and outputs the text character embedding representation and word embedding representation. The structure of the semantic encoding layer is as Figure 6 shown. Taking text P as an example, it receives the text P character embedding representation and word embedding representation and uses a bidirectional long short-term memory network BiLSTM for encoding to output the text P character and word granularity features, and transmits the text P character and word granularity features to the multi-level fine-grained feature extraction layer. The multi-level fine-grained feature extraction layer includes two sub-modules. The first sub-module, as Figure 7 shown, is responsible for extracting the fine-grained semantic features of the same text, and mainly uses multiple attention modules to encode different granularities of the same text to obtain the fine-grained semantic features of the same text; the second sub-module, as Figure 8 shown, is responsible for extracting the semantic interaction features among texts, and mainly uses several layers of encoding structures among texts to obtain the semantic interaction features among texts; in the first sub-module, as Figure 7As shown, taking text P as an example, the first-layer encoding structure uses multiple attention modules to extract the initial semantic features of the same text at a fine-grained level. Specifically: First, using soft-alignment attention, the character-level features of text P are soft-aligned with the word-level features of text P to obtain the soft-aligned features of text P at the character level and the soft-aligned features of text P at the word level. Second, using multiplicative-alignment attention, the character-level features of text P are multiplicatively aligned with the word-level features of text P to obtain the multiplicative-aligned features of text P at the character level and the multiplicative-aligned features of text P at the word level. After that, using subtractive-alignment attention, the character-level features of text P are subtractively aligned with the word-level features of text P to obtain the subtractive-aligned features of text P at the character level and the subtractive-aligned features of text P at the word level, thus completing the extraction of the initial semantic features of the same text at a fine-grained level. The second-layer encoding structure enhances the initial semantic features of the same text at a fine-grained level and completes the extraction of the semantic features of the same text at a fine-grained level. Specifically: First, at the character level, the soft-aligned features of text P at the character level are added to the character-level features of text P to obtain the deep soft-aligned features of text P at the character level. Subsequently, the multiplicative-aligned features of text P at the character level are added to the character-level features of text P to obtain the deep multiplicative-aligned features of text P at the character level. Then, the subtractive-aligned features of text P at the character level are added to the character-level features of text P to obtain the deep subtractive-aligned features of text P at the character level. Next, the deep soft-aligned features of text P at the character level, the deep multiplicative-aligned features of text P at the character level, and the deep subtractive-aligned features of text P at the character level are concatenated to obtain the high-level features of text P at the character level. At the word level, it is similar to the character level. First, the soft-aligned features of text P at the word level are added to the word-level features of text P to obtain the deep soft-aligned features of text P at the word level. Subsequently, the multiplicative-aligned features of text P at the word level are added to the word-level features of text P to obtain the deep multiplicative-aligned features of text P at the word level. Then, the subtractive-aligned features of text P at the word level are added to the word-level features of text P to obtain the deep subtractive-aligned features of text P at the word level. Next, the deep soft-aligned features of text P at the word level, the deep multiplicative-aligned features of text P at the word level, and the deep subtractive-aligned features of text P at the word level are concatenated to obtain the high-level features of text P at the word level, and the deep soft-aligned features of text P at the character level and the deep soft-aligned features of text P at the word level are concatenated to obtain the deep semantic features of text P, thus completing the extraction of the semantic features of the same text at a fine-grained level. In the second sub-module, such as Figure 8As shown, the first-layer encoding structure uses several encoding structures to extract the initial semantic interaction features between texts. Specifically: at the character granularity, first, the character granularity features of text P and the character granularity features of text Q are subjected to soft alignment attention to obtain the soft alignment interaction features of text P at the character granularity level and the soft alignment interaction features of text Q at the character granularity level. Secondly, the character granularity features of text P and the character granularity features of text Q are subjected to subtraction alignment attention to obtain the subtraction alignment interaction features of text P at the character granularity level and the subtraction alignment interaction features of text Q at the character granularity level; at the word granularity, it is similar to that at the character granularity. The word granularity features of text P and the word granularity features of text Q are subjected to soft alignment attention to obtain the soft alignment interaction features of text P at the word granularity level and the soft alignment interaction features of text Q at the word granularity level. Then, the word granularity features of text P and the word granularity features of text Q are subjected to subtraction alignment attention to obtain the subtraction alignment interaction features of text P at the word granularity level and the subtraction alignment interaction features of text Q at the word granularity level; the second-layer encoding structure enhances the initial semantic interaction features between texts and completes the extraction of semantic interaction features between texts. Specifically: at the character granularity, first, the soft alignment interaction features of text P at the character granularity level are added to the character granularity features of text P to obtain the deep soft alignment interaction features of text P at the character granularity level. After that, the subtraction alignment interaction features of text P at the character granularity level are added to the character granularity features of text P to obtain the deep subtraction alignment interaction features of text P at the character granularity level. Finally, the deep soft alignment interaction features of text P at the character granularity level and the deep subtraction alignment interaction features of text P at the character granularity level are concatenated to obtain the high-level interaction features of text P at the character granularity level; at the word granularity, first, the soft alignment interaction features of text P at the word granularity level are added to the word granularity features of text P to obtain the deep soft alignment interaction features of text P at the word granularity level. After that, the subtraction alignment interaction features of text P at the word granularity level are added to the word granularity features of text P to obtain the deep subtraction alignment interaction features of text P at the word granularity level. Finally, the deep soft alignment interaction features of text P at the word granularity level and the deep subtraction alignment interaction features of text P at the word granularity level are concatenated to obtain the high-level interaction features of text P at the word granularity level, and the deep subtraction alignment interaction features of text P at the character granularity level and the deep subtraction alignment interaction features of text P at the word granularity level are concatenated to obtain the deep semantic interaction features of text P; for text Q, similar operations are performed to obtain the high-level interaction features of text Q at the character granularity level, the deep semantic interaction features of text Q, and the high-level interaction features of text Q at the word granularity level, that is, the extraction of semantic interaction features between texts is completed. Feature fusion layer, such as Figure 9As shown, the first sub-module merges multiple relevant features, and the second sub-module performs multiple matching operations to obtain the final matching feature vector. Among them, the first sub-module merges multiple relevant features. Taking text P as an example, specifically: at the character granularity, the high-level features of text P at the character granularity level are concatenated with the high-level interaction features of text P at the character granularity level to obtain the aggregated features of text P at the character granularity level, and the self-attention is applied to the aggregated features of text P at the character granularity level to obtain the deep aggregated features of text P at the character granularity level; at the word granularity, similar to the character granularity, the high-level features of text P at the word granularity level are concatenated with the high-level interaction features of text P at the word granularity level to obtain the aggregated features of text P at the word granularity level, and the self-attention is applied to the aggregated features of text P at the word granularity level to obtain the deep aggregated features of text P at the word granularity level. Then, the deep aggregated features of text P at the character granularity level and the deep aggregated features of text P at the word granularity level are concatenated and the max-pooling operation is performed to obtain the pooled semantic features of text P. Similarly, the same operation as text P is performed on text Q to obtain the pooled semantic features of text Q; then, the deep semantic interaction features of text P are concatenated with the deep semantic features of text P to obtain the deep aggregated features of text P. Similarly, the same operation as text P is performed on text Q to obtain the deep aggregated features of text Q. Then, the soft alignment attention is applied to the deep aggregated features of text P and the deep aggregated features of text Q to obtain the soft-aligned deep aggregated features of text P and the soft-aligned deep aggregated features of text Q. Subsequently, the max-pooling operation is performed on the soft-aligned deep aggregated features of text P to obtain the pooled deep aggregated features of text P, and the max-pooling operation is performed on the soft-aligned deep aggregated features of text Q to obtain the pooled deep aggregated features of text Q; the second sub-module performs multiple matching operations to obtain the final matching feature vector. Specifically: first, the absolute value subtraction is performed on the pooled semantic features of text P and the pooled semantic features of text Q to obtain the subtraction matching feature. Second, the dot product is performed on the pooled semantic features of text P and the pooled semantic features of text Q to obtain the dot product matching feature. Third, the absolute value subtraction is performed on the pooled deep aggregated features of text P and the pooled deep aggregated features of text Q to obtain the deep subtraction matching feature. Fourth, the dot product is performed on the pooled deep aggregated features of text P and the pooled deep aggregated features of text Q to obtain the deep dot product matching feature. Finally, the pooled semantic features of text P, the pooled semantic features of text Q, the subtraction matching feature, the dot product matching feature, the deep subtraction matching feature, and the deep dot product matching feature are concatenated to obtain the final matching feature vector. The prediction layer inputs the final matching feature vector into a multi-layer perceptron to obtain a floating-point numerical value, which is used as the matching degree and compared with a preset threshold. According to the comparison result, it is determined whether the semantics of the text match. Specifically as follows:
[0171] (1) The embedding layer performs embedding operations on the input text at the character granularity and the word granularity respectively, and outputs the text character embedding representation and the word embedding representation;
[0172] (2) The semantic encoding layer receives the word embedding representation and the character embedding representation and encodes them using a bidirectional long short-term memory network (BiLSTM), outputting text character and word granularity features;
[0173] (3) The multi-level fine-grained feature extraction layer performs intra-text and inter-text encoding operations on the text character and word granularity features output by the semantic encoding layer to obtain intra-text fine-grained semantic features and inter-text semantic interaction features;
[0174] (4) The feature fusion layer merges various relevant features and then performs various matching operations to obtain the final matching feature vector;
[0175] (4) The prediction layer inputs the final matching feature vector into a multi-layer perceptron to obtain a floating-point numerical value, compares it with a preset threshold, and determines whether the semantics of the text match according to the comparison result.
[0176] Embodiment 2:
[0177] As shown in the appendix Figure 1 The text semantic matching method for medical intelligent question answering of the present invention is specifically as follows:
[0178] S1. Construct a text semantic matching knowledge base. As shown in the appendix Figure 2 The specific steps are as follows:
[0179] S101. Download the dataset on the network to obtain the original data: Download the publicly available text semantic matching dataset or the artificially constructed dataset on the network and use it as the original data for constructing the text semantic matching knowledge base;
[0180] For example: There are many publicly available text semantic matching datasets for medical intelligent question answering on the network, and there are also a large number of question-answer data pairs in many medical community forums. The present invention collects these data and downloads them to obtain the original data for constructing the text semantic matching knowledge base;
[0181] Examples of text pairs are as follows:
[0182] txt P What are the symptoms of a cold? txt Q Which symptoms can indicate a cold?
[0183] S102. Preprocess the original data: Preprocess the original data for constructing the text semantic matching knowledge base, perform word breaking and word segmentation operations on each text therein to obtain a text semantic matching word breaking processing knowledge base and a word segmentation processing knowledge base;
[0184] Taking the txt P shown in S101 as an example, after performing the word breaking processing operation on it, "What are the symptoms of a cold?" is obtained; after using the Jieba word segmentation tool to perform the word segmentation operation on it, "What are the symptoms of a cold?" is obtained.
[0185] S103, Aggregate sub-knowledge bases: Aggregate the text semantic matching word-breaking processing knowledge base and the text semantic matching word-segmentation processing knowledge base to construct a text semantic matching knowledge base;
[0186] Aggregate the text semantic matching word-breaking processing knowledge base and the text semantic matching word-segmentation processing knowledge base obtained in S102 into the same folder to obtain a text semantic matching knowledge base. The process is as shown in the appendix. Figure 2 It should be noted here that the data processed by the word-breaking operation and the data processed by the word-segmentation operation will not be merged into the same file. That is, the text semantic matching knowledge base actually contains two independent sub-knowledge bases.
[0187] S2. Construct a training data set for the text semantic matching model: For each text in the text semantic matching knowledge base, if its semantics are the same, this text can be used to construct a training positive example; if its semantics are different, this text can be used to construct a training negative example; Mix a certain number of positive example data and negative example data to construct the training data set required by the model. As shown in the appendix, the specific steps are as follows: Figure 3 As shown below, the specific steps are as follows:
[0188] S201. Construct training positive example data: Construct two texts with the same text semantics as positive example data, formalized as: (txt P_char, txt Q_char, txt P_word, txt Q_word, 1);
[0189] Example: For txt P and txt Q shown in step S101, after the word-breaking operation and word-segmentation operation in step S102, the formalized positive example data is:
[0190] ("What are the symptoms of a cold?", "What symptoms can be used to judge a cold?", "What are the symptoms of a cold?", "What symptoms can be used to judge a cold?", 1).
[0191] S202. Construct training negative example data: Select a certain text contained in it and randomly select a text that does not match it for combination; Construct these two texts with different semantics as negative example data. Using an operation similar to step S201, the negative example data can be formalized as: (txt P_char, txt Q_char, txt P_word, txt Q_word, 0). The meanings of each symbol are the same as those in step S201, and 0 indicates that the semantics of these two texts do not match and it is a negative example;
[0192] Example: The example is very similar to constructing training positive examples and will not be elaborated here.
[0193] S203. Construct a training data set: Merge all the positive example data and negative example data obtained after the operations in steps S201 and S202, shuffle their order, and construct the final training data set. Whether it is positive example data or negative example data, they all contain 5 dimensions, namely txt P_char, txt Q_char, txt P_word, txt Q_word, 0 or 1.
[0194] S3. Construct a text semantic matching model: The main operations are to construct a word mapping conversion table, an input layer, a word vector mapping layer, a semantic encoding layer, a multi-level fine-grained feature extraction layer, a feature fusion layer, and a prediction layer; as shown in the appendix Figure 4 as follows:
[0195] S301. Construct a word mapping conversion table: The word table is constructed by the text semantic matching word breaking processing knowledge base and word segmentation processing knowledge base obtained after step S102. After the word table is constructed, each word or character in the table is mapped to a unique digital identifier, and its mapping rule is: starting from the number 1, and then increasing sequentially according to the order in which each word or character is entered into the word table, so as to form the word mapping conversion table required by the present invention;
[0196] Example: Taking the content after step S102, "What are the symptoms of a cold?", "What are the symptoms of a cold?", the word table and word mapping conversion table are constructed as follows:
[0197] Words Cold Flu Manifestation Appearance Of Symptom Symptoms All Have Which Those Mapping 1 2 3 4 5 6 7 8 9 10 11 Words ? Cold Manifestation Symptom All have Which Mapping 12 13 14 15 16 17
[0198] Subsequently, the present invention uses Word2Vec to train the word vector model to obtain the word vector matrix char_embedding_matrix of each word or character;
[0199] Example: In Keras, the code implementation for the above description is as follows:
[0200] w2v_model = models.Word2Vec(w2v_corpus, size = EMB_DIM, window = 5, min_count = 1, sg = 1, workers = 4, seed = 1234, iter = 25)
[0201] embedding_matrix = np.zeros([len(tokenizer.word_index)+1, EMB_DIM])
[0202] tokenizer=Tokenizer(num_words=len(word_set))
[0203] for word, idx in tokenizer.word_index.items():
[0204] embedding_matrix[idx,:]=w2v_model.wv[word]
[0205] Among them, w2v_corpus is all the data in the text semantic matching knowledge base; EMB_DIM is the vector dimension, and this model sets EMB_DIM to 300, and word_set is the word list.
[0206] S302, constructing the input layer: the input layer includes four inputs, and txt P_char, txt Q_char, txt P_word and txt Q_word are obtained from the training data set samples of the input layer respectively, and are formalized as: (txt P_char, txt Q_char, txt P_word, txt Q_word);
[0207] For each character and word in the input text, the present invention converts it into a corresponding digital identifier according to the character-word mapping conversion table constructed in step S301;
[0208] For example: Using the text shown in step S201 as an example to form an input data, the result is as follows:
[0209] ("What are the symptoms of a cold?", "What symptoms can be diagnosed as a cold?", "What are the symptoms of a cold?", "What symptoms can be diagnosed as a cold?")
[0210] Each piece of input data contains 4 subtexts; according to the word mapping conversion table in step S301, it is converted into numerical representation (assuming that "可", "以", "判", "断", "为", "可以", "判" that appear in txt Q but not in txt P are mapped to 18, 19, 20, 21, 22, 23, 24 respectively). The 4 subtexts of the input data are combined and represented as follows:
[0211] (“1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12”, “3, 4, 10, 11, 6, 7, 18, 19, 20, 21, 22, 1, 2, 12”, “13, 14, 5, 15, 16, 17, 12”, “14, 17, 15, 23, 24, 22, 13, 12”).
[0212] S303. Construct a word-character vector mapping layer: Initialize the weight parameters of the current layer by loading the word-character vector matrix obtained from training in the step of constructing the word mapping conversion table; for the input texts txt P_char, txt Q_char, txt P_word, and txt Q_word, obtain their corresponding text character embeddings and word embeddings txt P_char_embed, txt Q_char_embed, txt P_word_embed, and txt Q_word_embed; each text in the text semantic matching knowledge base can convert the text information into vector form through the word-character vector mapping method.
[0213] For example: In Keras, the code implementation for the above description is as follows:
[0214] embedding_layer = Embedding(embedding_matrix.shape[0], emb_dim, weights=[embedding_matrix], input_length=input_dim, trainable=False)
[0215] Among them, embedding_matrix is the word-character vector matrix obtained from training in the word mapping conversion table step, embedding_matrix.shape[0] is the size of the word-character table of the word-character vector matrix, emb_dim is the dimension of the output text character embeddings and word embeddings, and input_length is the length of the input sequence.
[0216] The corresponding texts txt P_char, txt Q_char, txt P_word, and txt Q_word, after being processed by the Embedding layer of Keras, obtain the corresponding text character embeddings and word embeddings txt P_char_embed, txt Q_char_embed, txt P_word_embed, and txt Q_word_embed.
[0217] S304. Construct a semantic encoding layer:
[0218] Taking text P as an example, this module receives the text P character and word embeddings and uses a bidirectional long short-term memory network BiLSTM for encoding to obtain the text P character and word granularity features, denoted as The specific formula is as follows:
[0219]
[0220]
[0221] Among them, N represents the lengths of the character-level features and word-level features. Formula (1) represents encoding the character embedding representation of text P using a bidirectional long short-term memory network BiLSTM. Among them, represents the character-level feature of the i-th position in text P encoded by the bidirectional long short-term memory network BiLSTM, represents the character-level feature of the i-th position in text P encoded by the forward long short-term memory network LSTM, represents the character-level feature of the i-th position in text P encoded by the backward LSTM; the symbols in Formula (2) have basically the same meanings as those in Formula (1). represents the word-level feature of the j-th position in text P encoded by the bidirectional long short-term memory network BiLSTM, represents the word-level feature of the j-th position in text P encoded by the forward LSTM, represents the word-level feature of the j-th position in text P encoded by the backward LSTM.
[0222] Similarly, perform operations on text Q similar to those on text P to obtain the character and word-level features of text Q, denoted as Q c 、Q w .
[0223] S305. Construct a multi-level fine-grained feature extraction layer:
[0224] The multi-level fine-grained feature extraction layer takes the character and word-level features of the text output by the semantic encoding layer as input; performs intra-text and inter-text encoding operations to obtain the intra-text fine-grained semantic features and inter-text semantic interaction features; includes two sub-modules. The first sub-module is responsible for extracting the intra-text fine-grained semantic features, mainly using multiple attention modules to encode different granularities of the same text to obtain the intra-text fine-grained semantic features, as shown in Figure 7 ; The second sub-module is responsible for extracting the inter-text semantic interaction features, mainly using several layers of encoding structures between texts to obtain the semantic interaction features between texts, as shown in Figure 8 .
[0225] S30501. Extraction of intra-text fine-grained semantic features by the first sub-module:
[0226] First, for the convenience of subsequent description, take text P as an example in the first part and define the following attention module:
[0227] Define a soft alignment attention module, denoted as SOA, and the formula is as follows:
[0228]
[0229] where represents the character-level feature at the \(i\)-th position of text \(P\), from formula (1), represents the word-level feature at the \(j\)-th position of text \(P\), from formula (2), represents the soft alignment attention weight between the character-level feature at the \(i\)-th position and the word-level feature at the \(j\)-th position of text \(P\), represents mapping the soft alignment attention weight to a value between 0 and 1 through a softmax operation, represents using soft alignment attention to re - represent the character-level feature at the \(i\)-th position of text \(P\) as a weighted sum of all word-level features of text \(P\), represents using soft alignment attention to re - represent the word-level feature at the \(j\)-th position of text \(P\) as a weighted sum of all character-level features of text \(P\);
[0230] Define the multiplicative alignment attention module, denoted as MUA, with the formula as follows:
[0231]
[0232] where, TimeDistributed(Dense()) represents performing the same Dense() layer operation on the tensor of each time step, ⊙ represents element-wise multiplication operation, tanh represents the activation function, \(P\) c represents the character-level feature of text \(P\), represents the character-level feature of text \(P\) after being processed by the Dense() layer, \(P\) w represents the word-level feature of text \(P\), represents the multiplicative alignment attention weight, represents mapping the multiplicative alignment attention weight to a value between 0 and 1 through a softmax operation, represents using multiplicative alignment attention to re - represent the character-level feature at the \(i\)-th position of text \(P\) as a weighted sum of all word-level features of text \(P\), represents using multiplicative alignment attention to re - represent the word-level feature at the \(j\)-th position of text \(P\) as a weighted sum of all character-level features of text \(P\);
[0233] Define the subtractive alignment attention module denoted as SUA, with the formula as follows:
[0234]
[0235] where, TimeDistributed(Dense()) represents performing the same Dense() layer operation on the tensor of each time step, - represents element-wise subtraction operation, tanh represents the activation function, \(P\) c represents the character-level feature of text \(P\), Represents the character-level features of text P after being processed by the Dense() layer, P w Represents the word-level features of text P Represents the subtraction alignment attention weights Represents mapping the subtraction alignment attention weights to values between 0 and 1 through a softmax operation Represents using subtraction alignment attention to re - represent the character-level feature at the i-th position of text P as a weighted sum of all word-level features of text P Represents using subtraction alignment attention to re - represent the word-level feature at the j-th position of text P as a weighted sum of all character-level features of text P
[0236] Define the self - alignment attention module as SEA, and the formula is as follows
[0237]
[0238] Among them, Represents the character-level feature at the i-th position of text P Represents the character-level feature at the j-th position of text P Represents the self - alignment attention weights between the character-level feature at the i-th position of text P and the character-level feature at the j-th position of text P Represents mapping the self - alignment attention weights to values between 0 and 1 through a softmax operation Represents using self - alignment attention to re - represent the character-level feature at the i-th position of text P as a weighted sum of all character-level features of text P
[0239] In subsequent descriptions, use the SOA symbol to represent the operation of formula (3), the MUA symbol to represent the operation of formula (4), the SUA symbol to represent the operation of formula (5), and the SEA symbol to represent the operation of formula (6);
[0240] S3050101. The first - layer encoding structure uses multiple attention modules to extract the fine - grained initial - level semantic features of the same text:
[0241] First, use soft alignment attention to align the character-level features P of text P c with the word-level features P of text P w to obtain the soft - alignment feature of text P at the character level and the soft - alignment feature of text P at the word level As shown in formula (7):
[0242]
[0243] Secondly, using multiplicative alignment attention, the character-level feature P of text P c and the word-level feature P of text P w are subjected to multiplicative alignment attention to obtain the character-level multiplicative alignment feature of text P and the word-level multiplicative alignment feature of text P As shown in formula (8):
[0244]
[0245] After that, using subtractive alignment attention, the character-level feature P of text P c and the word-level feature P of text P w are subjected to subtractive alignment attention to obtain the character-level subtractive alignment feature of text P and the word-level subtractive alignment feature of text P As shown in formula (9):
[0246]
[0247] Similarly, performing a similar process on text Q as on text P, the character-level soft alignment feature of text Q the word-level soft alignment feature of text Q the character-level multiplicative alignment feature of text Q the word-level multiplicative alignment feature of text Q the character-level subtractive alignment feature of text Q the word-level subtractive alignment feature of text Q That is, the extraction of the fine-grained initial semantic features of the same text is completed;
[0248] S3050102. The second-layer encoding structure enhances the fine-grained initial semantic features of the same text to complete the extraction of the fine-grained semantic features of the same text:
[0249] First, the character-level soft alignment feature of text P in formula (7) is added to the character-level feature P of text P in formula (1) c to obtain the character-level deep soft alignment feature of text P As shown in formula (10):
[0250]
[0251] Subsequently, the character-level multiplicative alignment feature of text P in formula (8) is added to the character-level feature P of text P in formula (1) c to obtain the character-level deep multiplicative alignment feature of text P As shown in formula (11):
[0252]
[0253] Then, subtract and align the character-level text P features in formula (9) with the character-level text P features P in formula (1) c and add them to obtain the deep subtraction alignment features of text P at the character level As shown in formula (12):
[0254]
[0255] Next, connect the deep soft alignment features of text P at the character level in formula (10) the deep multiplication alignment features of text P at the character level in formula (11) the deep subtraction alignment features of text P at the character level in formula (12) to obtain the high-level features P' of text P at the character level c , as shown in formula (13):
[0256]
[0257] Similar to the character level, at the word level, first, add the soft alignment features of text P at the word level in formula (7) to the word-level features P of text P in formula (2) w to obtain the deep soft alignment features of text P at the word level As shown in formula (14):
[0258]
[0259] Subsequently, add the multiplication alignment features of text P at the word level in formula (8) to the word-level features P of text P in formula (2) w to obtain the deep multiplication alignment features of text P at the word level As shown in formula (15):
[0260]
[0261] Then, add the subtraction alignment features of text P at the word level in formula (9) to the word-level features P of text P in formula (2) w to obtain the deep subtraction alignment features of text P at the word level As shown in formula (16):
[0262]
[0263] Next, connect the deep soft alignment features of text P at the word granularity level in formula (14) The deep multiplication alignment features of text P at the word granularity level in formula (15) The deep subtraction alignment features of text P at the word granularity level in formula (16) Obtain the high-level feature P' of text P at the word granularity level w , as shown in formula (17):
[0264]
[0265] Connect the deep soft alignment features of text P at the character granularity level in formula (10) With the deep soft alignment features of text P at the word granularity level in formula (14) Obtain the deep semantic feature P' of text P deep , as shown in formula (18):
[0266]
[0267] Similarly, perform similar processing on text Q as on text P to obtain the deep soft alignment features of text Q at the character granularity level The deep multiplication alignment features of text Q at the character granularity level The deep subtraction alignment features of text Q at the character granularity level The high-level feature Q' of text Q at the character granularity level c And the deep soft alignment features of text Q at the word granularity level The deep multiplication alignment features of text Q at the word granularity level The deep subtraction alignment features of text Q at the word granularity level The high-level feature Q' of text Q at the word granularity level w 、The deep semantic feature Q' of text Q deep , that is, the extraction of fine-grained semantic features of the same text is completed;
[0268] S30502. Extraction of semantic interaction features between texts:
[0269] S3050201. The first-layer coding structure uses several layers of coding structures to extract the initial semantic interaction features between texts:
[0270] At the character granularity, first, the character granularity feature P of text P in formula (1) c And the character granularity feature Q of text Q c Perform soft alignment attention to obtain the soft alignment interaction features of text P at the character granularity level And the soft alignment interaction features of text Q at the character granularity level As shown in formula (19):
[0271]
[0272] Secondly, the character granularity feature P of text P in formula (1) c and the character granularity feature Q of text Q c are subjected to subtraction alignment attention to obtain the text P subtraction alignment interaction feature at the character granularity level and the text Q subtraction alignment interaction feature at the character granularity level As shown in formula (20):
[0273]
[0274] Similar to the word granularity, first, the word granularity feature P of text P in formula (2) w and the word granularity feature Q of text Q w are subjected to soft alignment attention to obtain the text P soft alignment interaction feature at the word granularity level and the text Q soft alignment interaction feature at the word granularity level As shown in formula (21):
[0275]
[0276] Then, the word granularity feature P of text P in formula (2) w and the word granularity feature Q of text Q w are subjected to subtraction alignment attention to obtain the text P subtraction alignment interaction feature at the word granularity level and the text Q subtraction alignment interaction feature at the word granularity level As shown in formula (22):
[0277]
[0278] S3050202. The second-layer coding structure enhances the initial semantic interaction features between texts to complete the extraction of semantic interaction features between texts:
[0279] At the character granularity, first, the text P soft alignment interaction feature at the character granularity level in formula (19) is added to the character granularity feature P of text P in formula (1) c to obtain the text P deep soft alignment interaction feature at the character granularity level As shown in formula (23):
[0280]
[0281] After that, the text P subtraction alignment interaction feature at the character granularity level in formula (20) is combined with the character granularity feature P of text P in formula (1) cAdd them to obtain the deep subtraction alignment interaction feature of text P at the character granularity level As shown in formula (24):
[0282]
[0283] Finally, concatenate the deep soft alignment interaction feature of text P at the character granularity level in formula (23) with the deep subtraction alignment interaction feature of text P at the character granularity level in formula (24) to obtain the high-level interaction feature P of text P at the character granularity level c ″, as shown in formula (25):
[0284]
[0285] At the word granularity, first, add the soft alignment interaction feature of text P at the word granularity level in formula (21) to the word granularity feature of text P in formula (2) w to obtain the deep soft alignment interaction feature of text P at the word granularity level As shown in formula (26):
[0286]
[0287] After that, add the subtraction alignment interaction feature of text P at the word granularity level in formula (22) to the word granularity feature of text P in formula (2) w to obtain the deep subtraction alignment interaction feature of text P at the word granularity level As shown in formula (27):
[0288]
[0289] Finally, concatenate the deep soft alignment interaction feature of text P at the word granularity level in formula (26) with the deep subtraction alignment interaction feature of text P at the word granularity level in formula (27) to obtain the high-level interaction feature P of text P at the word granularity level″ w , as shown in formula (28):
[0290]
[0291] Concatenate the deep subtraction alignment interaction feature of text P at the character granularity level in formula (24) with the deep subtraction alignment interaction feature of text P at the word granularity level in formula (27) to obtain the deep semantic interaction feature P of text P″ deep , as shown in formula (29):
[0292]
[0293] Similarly, by performing similar processing on text Q as on text P, the deep soft alignment interaction features of text Q at the character granularity level can be obtained. The deep subtraction alignment interaction features of text Q at the character granularity level The high-level interaction features Q″ of text Q at the character granularity level c and the deep soft alignment interaction features of text Q at the word granularity level The deep subtraction alignment interaction features of text Q at the word granularity level The high-level interaction features Q″ of text Q at the word granularity level w and the deep semantic interaction features Q″ of text Q deep , which completes the extraction of the semantic interaction features between texts.
[0294] S306. Construct a feature fusion layer:
[0295] First, for the convenience of subsequent description, the following operations are defined:
[0296] Define the operation of subtracting vectors and then taking the absolute value bit by bit as AB, as shown in formula (30):
[0297] AB(P, Q) = |P - Q| (30)
[0298] where P and Q are two different vectors, representing the operation of taking the absolute value bit by bit after subtracting the two vectors P and Q;
[0299] Define the bitwise multiplication operation of vectors as MU, as shown in formula (31):
[0300] MU(P, Q) = P ⊙ Q (31)
[0301] where P and Q are two different vectors, representing the operation of multiplying the vectors P and Q bit by bit;
[0302] In subsequent descriptions, the AB symbol is used to represent the operation of formula (30), the MU symbol is used to represent the operation of formula (31), the SOA symbol is used to represent the operation of formula (3), the MUA symbol is used to represent the operation of formula (4), the SUA symbol is used to represent the operation of formula (5), and the SEA symbol is used to represent the operation of formula (6);
[0303] The feature fusion layer is divided into two sub-modules. The first sub-module merges various relevant features, and the second sub-module performs various matching operations to obtain the final matching feature vector, as Figure 9 shown.
[0304] S30601. The first sub-module merges various relevant features:
[0305] Connect the high-level features \(P'\) of the character-level text \(P\) in formula (13) c with the high-level interaction features \(P'\) of the character-level text \(P\) in formula (25) c to obtain the aggregated features of the character-level text \(P\) and use the aggregated features of the character-level text \(P\) to perform self-attention to obtain the deep aggregated features of the character-level text \(P\) As shown in formula (32):
[0306]
[0307] At the word level, similar to the character level, connect the high-level features \(P'\) of the word-level text \(P\) in formula (17) w with the high-level interaction features \(P''\) of the word-level text \(P\) in formula (28) w to obtain the aggregated features of the word-level text \(P\) and use the aggregated features of the word-level text \(P\) to perform self-attention to obtain the deep aggregated features of the word-level text \(P\) As shown in formula (33):
[0308]
[0309] After that, connect the deep aggregated features of the character-level text \(P\) in formula (32) with the deep aggregated features of the word-level text \(P\) in formula (33) Then perform max pooling operation to obtain the pooled semantic features \(P'\) of text \(P\), as shown in formula (34):
[0310]
[0311] Next, connect the deep semantic features \(P'\) of text \(P\) in formula (18) deep with the deep semantic interaction features \(P''\) of text \(P\) in formula (29) deep to obtain the deep aggregated features of text \(P\) As shown in formula (35):
[0312]
[0313] Similarly, perform the same operations on text \(Q\) as on text \(P\) to obtain the aggregated features of the character-level text \(Q\) the deep aggregated features of the character-level text \(Q\) the aggregated features of the word-level text \(Q\) the deep aggregated features of the word-level text \(Q\) The semantic feature Q' of the text Q after pooling and the deep aggregation feature of the text Q
[0314] Then, the deep aggregation feature of the text P in formula (35) and the deep aggregation feature of the text Q are subjected to soft-alignment attention to obtain the soft-aligned deep aggregation feature of the text P and the soft-aligned deep aggregation feature of the text Q As shown in formula (36):
[0315]
[0316] Subsequently, the soft-aligned deep aggregation feature of the text P in formula (36) is subjected to max pooling operation to obtain the deep aggregation feature P″ of the text P after pooling, and the soft-aligned deep aggregation feature of the text Q is subjected to max pooling operation to obtain the deep aggregation feature Q″ of the text Q after pooling. As shown in formula (37):
[0317]
[0318] S30602. Perform multiple matching operations to obtain the final matching feature vector:
[0319] First, the absolute value subtraction is performed on the semantic feature P' of the text P after pooling and the semantic feature Q' of the text Q after pooling in formula (34) to obtain the subtraction matching feature PQ ab , as shown in formula (38):
[0320] PQ ab = AB(P' - Q') (38)
[0321] Second, the dot product is performed on the semantic feature P' of the text P after pooling and the semantic feature Q' of the text Q after pooling in formula (34) to obtain the dot product matching feature PQ mu , as shown in formula (39):
[0322] PQ mu = MU(P', Q') (39)
[0323] Third, the absolute value subtraction is performed on the deep aggregation feature P″ of the text P after pooling in formula (37) and the deep aggregation feature Q″ of the text Q after pooling in formula (37) to obtain the deep subtraction matching feature PQ' ab , as shown in formula (40):
[0324] PQ' ab = AB(P″, Q″) (40)
[0325] After that, the deep aggregation feature P″ of the pooled text P in formula (37) and the deep aggregation feature Q″ of the pooled text Q in formula (37) are dot - multiplied to obtain the deep dot - product matching feature PQ′ mu , as shown in formula (41):
[0326] PQ′ mu = MU(P″, Q″) (41)
[0327] Finally, connect the semantic feature P′ of the pooled text P in formula (34), the semantic feature Q′ of the pooled text Q, the subtraction matching feature PQ in formula (38) ab , the dot - product matching feature PQ in formula (39) mu , the deep subtraction matching feature PQ′ in formula (40) ab , and the deep dot - product matching feature PQ′ in formula (41) mu to obtain the final matching feature vector F, as shown in formula (42):
[0328] F = [P′; Q′; PQ ab ; PQ mu ; PQ′ ab ; PQ′ mu (42)
[0329] S307. Construct a prediction layer:
[0330] Take the final matching feature vector as the input, use a three - layer fully - connected layer, and use the ReLU activation function after the first and second fully - connected layers for activation, and use the sigmoid function after the third fully - connected layer for activation, so as to obtain a matching degree value between [0, 1], denoted as y pred ; Finally, by comparing with the set threshold 0.5, it is judged whether the text semantics match; that is, when y pred ≥0.5, it is predicted that the semantics of the text match, otherwise, they do not match;
[0331] When the text semantics matching model has not been trained, it needs to be trained on the training data set constructed according to the semantics matching knowledge base to optimize the model parameters; when the model is trained, the prediction layer can predict whether the semantics of the target text match.
[0332] S4. Train the text semantics matching model: Train the text semantics matching model constructed in step S3 on the training data set obtained in step S2, as shown in the appendix Figure 5 as follows:
[0333] S401. Construct a loss function: As known from step S307, y pred is the matching degree value obtained after being processed by the text semantics matching model; and ytrue is the true label indicating whether the semantics of two texts match. Its value is limited to 0 or 1. Cross-entropy is used as the loss function, and the formula is as follows:
[0334]
[0335] S402. Construct an optimization function:
[0336] After testing various optimization functions, the model finally selects the Adam optimization function as the optimization function for this model, and its hyperparameters are all set to the default values in Keras;
[0337] For example, the optimization function and its settings described above are represented by the following code in Keras:
[0338] optim = keras.optimizers.Adam()
[0339] The proposed model can achieve excellent results on the medical intelligent Q&A dataset.
[0340] Example 3:
[0341] As shown in the appendix Figure 10 The text semantic matching device for medical intelligent Q&A based on Example 2 includes:
[0342] A text semantic matching knowledge base construction unit, which is used to obtain a large amount of text data and then perform preprocessing operations on it to obtain a text semantic matching knowledge base that meets the training requirements;
[0343] A training dataset generation unit. For the data in the text semantic matching knowledge base, if their semantics are consistent, the text is used to construct training positive examples; otherwise, it is used to construct training negative examples. All positive example data and negative example data are mixed to obtain a training dataset;
[0344] A text semantic matching model construction unit: used to construct a word mapping conversion table, construct an input layer, construct a word vector mapping layer, construct a semantic encoding layer, construct a multi-level fine-grained feature extraction layer, construct a feature fusion layer, and construct a prediction layer;
[0345] A text semantic matching model training unit: used to construct a training loss function and an optimization function to complete the training of the model.
[0346] Example 4:
[0347] Based on the storage medium of Example 2, which stores multiple instructions that are loaded and executed by a processor to perform the steps of the text semantic method for medical intelligent Q&A in Example 2.
[0348] Example 5:
[0349] An electronic device based on Embodiment 4, the electronic device comprising: the storage medium of Embodiment 4; and
[0350] a processor configured to execute instructions in the storage medium of Embodiment 4.
[0351] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text semantic matching method for medical intelligent question answering, characterized in that This method constructs and trains a text semantic matching model composed of an embedding layer, a semantic encoding layer, a multi-level fine-grained feature extraction layer, a feature fusion layer, and a prediction layer to extract text character and word-level features, capture the fine-grained semantic features of the same text and the semantic interaction features between texts, and finally merge various relevant features, and then perform various matching operations to generate the final matching feature vector and determine the similarity of the texts. Specifically as follows: The embedding layer performs embedding operations on the input text at the character level and the word level respectively, and outputs the text character embedding representation and the word embedding representation. The semantic encoding layer receives the text character embedding representation and the word embedding representation and encodes them using a bidirectional long short-term memory network (BiLSTM), and outputs the text character and word-level features. The multi-level fine-grained feature extraction layer performs same-text and inter-text encoding operations on the text character and word-level features output by the semantic encoding layer to obtain the fine-grained semantic features of the same text and the semantic interaction features between texts. The feature fusion layer merges various relevant features, and then performs various matching operations to generate the final matching feature vector. The prediction layer inputs the final matching feature vector into a multi-layer perceptron to obtain a floating-point numerical value, uses it as the matching degree to compare with a preset threshold, and determines whether the semantics of the texts match according to the comparison result. Among them, the implementation details of the semantic encoding layer are as follows: Taking text P as an example, this module receives the word and character embedding representations of text P and encodes them using a bidirectional long short-term memory network (BiLSTM) to obtain the character and word granularity features of text P, denoted as The specific formula is as follows: Among them, N represents the lengths of the character-level features and the word-level features. Formula (1) indicates encoding the character embeddings of text P using a bidirectional long short-term memory network (BiLSTM). Among them, represents the character-level feature of the i-th position in text P encoded by the bidirectional long short-term memory network (BiLSTM), represents the character-level feature of the i-th position in text P encoded by the forward long short-term memory network (LSTM), represents the character-level feature of the i-th position in text P encoded by the backward LSTM; the symbols in Formula (2) have basically the same meanings as those in Formula (1). represents the word-level feature of the j-th position in text P encoded by the bidirectional long short-term memory network (BiLSTM), represents the word-level feature of the j-th position in text P encoded by the forward LSTM, represents the word-level feature of the j-th position in text P encoded by the backward LSTM; Similarly, perform similar operations on text Q as on text P to obtain the character and word granularity features of text Q, denoted as Q c 、Q w 。 2. The text semantic matching method for medical intelligent question answering according to claim 1, wherein The embedding layer includes a character-word mapping conversion table, an input layer, and a character-word vector mapping layer, and outputs the text character embedding representation and the word embedding representation. Among them, the character-word mapping conversion table: The mapping rule starts with the number 1, and then increases sequentially according to the order in which each character or word is entered into the character-word table, so as to form a character-word mapping conversion table. Then, use Word2Vec to train the character-word vector model to obtain the character-word vector matrix of each character or word. Input layer: The input layer includes four inputs. For each text in the training dataset or the text to be predicted, perform word segmentation and tokenization preprocessing to obtain txt P_char, txt Q_char, txt P_word, and txt Q_word respectively, where the suffixes char and word indicate that the corresponding text has been word-segmented or tokenized, and formalize them as: (txt P_char, txt Q_char, txt P_word, txt Q_word); for each character and word in the input text, convert them into corresponding digital identifiers according to the character-word mapping conversion table. Character-word vector mapping layer: Load the character-word vector matrix obtained in the step of constructing the character-word mapping conversion table to initialize the weight parameters of the current layer; obtain the corresponding text character embedding representation and word embedding representation txt P_char_embed, txt Q_char_embed, txt P_word_embed, and txt Q_word_embed for the input texts txt P_char, txt Q_char, txt P_word, and txt Q_word.
3. The text semantic matching method for medical intelligent question answering according to claim 1, characterized in that, The implementation details of the multi-level fine-grained feature extraction layer are as follows: For the text word and word granularity features output by the semantic encoding layer, perform intra-text and inter-text encoding operations to obtain intra-text fine-grained semantic features and inter-text semantic interaction features; it includes two sub-modules. The first sub-module is responsible for extracting intra-text fine-grained semantic features, mainly using multiple attention modules to encode different granularities of the same text to obtain intra-text fine-grained semantic features; the second sub-module is responsible for extracting inter-text semantic interaction features, mainly using several layers of encoding structures between texts to obtain inter-text semantic interaction features; The first sub-module: Extracting intra-text fine-grained semantic features First, for the convenience of subsequent description, in the first part, take text P as an example and define the following attention modules: Define the soft alignment attention module, denoted as SOA, and the formula is as follows: Among them represents the character-level feature at the $i$-th position of text $P$, which comes from formula (1). represents the word-level feature at the $j$-th position of text $P$, which comes from formula (2). represents the soft alignment attention weight between the character-level feature at the $i$-th position and the word-level feature at the $j$-th position of text $P$. represents the result of applying the softmax operation to the soft alignment attention weight to map it to a value between 0 and 1. represents that using the soft alignment attention, the character-level feature at the $i$-th position of text $P$ can be re-represented as a weighted sum of all word-level features of text $P$. represents that using the soft alignment attention, the word-level feature at the $j$-th position of text $P$ can be re-represented as a weighted sum of all character-level features of text $P$. Define the multiplication alignment attention module, denoted as MUA, and the formula is as follows: Among them, TimeDistributed(Dense()) means that the Dense() layer operation is performed on the tensors of each time step, ⊙ represents element-wise multiplication operation, tanh represents the activation function, and P c represents the character-level features of text P, represents the character-level features of text P after being processed by the Dense() layer, and P w represents the word-level features of text P, represents the multiplicative alignment attention weights, represents mapping the multiplicative alignment attention weights to values between 0 and 1 through the softmax operation, represents that the character-level feature at the i-th position of text P can be re-represented by the weighted sum of all word-level features of text P using multiplicative alignment attention, represents that the word-level feature at the j-th position of text P can be re-represented by the weighted sum of all character-level features of text P using multiplicative alignment attention; Define the subtraction alignment attention module, denoted as SUA, and the formula is as follows: Among them, TimeDistributed(Dense()) means that the same Dense() layer operation is performed on the tensors of each time step, - represents the element-wise subtraction operation, tanh represents the activation function, and P c represents the character-level features of text P, represents the character-level features of text P after being processed by the Dense() layer, P w represents the word-level features of text P, represents the subtraction alignment attention weights, represents mapping the subtraction alignment attention weights to values between 0 and 1 by performing a softmax operation, represents that the character-level feature at the i-th position of text P can be re-represented by the weighted sum of all word-level features of text P using subtraction alignment attention, represents that the word-level feature at the j-th position of text P can be re-represented by the weighted sum of all character-level features of text P using subtraction alignment attention; Define the self-alignment attention module, denoted as SEA, and the formula is as follows: Among them, represents the character-level feature at the $i$-th position of text $P$, represents the character-level feature at the $j$-th position of text $P$, represents the self-alignment attention weight between the character-level feature at the $i$-th position of text $P$ and the character-level feature at the $j$-th position of text $P$, represents the value obtained by mapping the self-alignment attention weight through the softmax operation to a value between 0 and 1, represents the re-representation of the character-level feature at the $i$-th position of text $P$ by weighted summation of all character-level features of text $P$ using self-alignment attention; In subsequent descriptions, use the SOA symbol to represent the operation of formula (3), use the MUA symbol to represent the operation of formula (4), use the SUA symbol to represent the operation of formula (5), and use the SEA symbol to represent the operation of formula (6); The first layer of the encoding structure uses multiple attention modules to extract the intra-text fine-grained initial semantic features: First, using soft alignment attention, the character-level features \(P\) of text \(P\) c and the word-level features \(P\) of text \(P\) w are subjected to soft alignment attention to obtain the character-level soft alignment features of text \(P\) and the word-level soft alignment features of text \(P\) As shown in formula (7): Secondly, using multiplicative alignment attention, the character-level features P of text P c and the word-level features P of text P w are subjected to multiplicative alignment attention to obtain the character-level multiplicative alignment features of text P and the word-level multiplicative alignment features of text P as shown in Formula (8): After that, subtractively aligned attention is used to obtain the character-level features of text P c from the word-level features of text P w to obtain the subtractively aligned features of text P at the character level and the subtractively aligned features of text P at the word level as shown in Equation (9): Similarly, performing a similar process on text Q as on text P yields the soft alignment features of text Q at the character granularity level. Soft alignment features of text Q at the word granularity level Multiplication alignment features of text Q at the character granularity level Multiplication alignment features of text Q at the word granularity level Subtraction alignment features of text Q at the character granularity level ⊙ Subtraction alignment features of text Q at the word granularity level That is, the extraction of the fine-grained initial semantic features of the same text is completed. The second layer of the encoding structure enhances the intra-text fine-grained initial semantic features to complete the extraction of intra-text fine-grained semantic features: First, add the character-level text P soft alignment feature in formula (7) to the character-level feature P of text P in formula (1) c to obtain the deep soft alignment feature of text P at the character level as shown in formula (10): Subsequently, the word-grained text P multiplication alignment feature in formula (8) is added to the word-grained feature P of text P in formula (1) c to obtain the deep multiplication alignment feature of word-grained text P as shown in formula (11): Then, subtract and align the text P subtraction alignment feature at the character granularity level in formula (9) from the text P character granularity feature P in formula (1) c and add them to obtain the deep subtraction alignment feature of text P at the character granularity level as shown in formula (12): Next, the deep soft alignment features of the character-level text P in formula (10) The deep multiplication alignment features of the character-level text P in formula (11) The deep subtraction alignment features of the character-level text P in formula (12) are concatenated to obtain the high-level features P of the character-level text P c ', as shown in formula (13): At the word granularity level, it is similar to that at the character granularity level. First, the soft alignment feature of text P at the word granularity level in formula (7) is added to the word granularity feature P of text P in formula (2) w to obtain the deep soft alignment feature of text P at the word granularity level as shown in formula (14): Subsequently, the word-grained text P multiplicative alignment feature in formula (8) is added to the word-grained feature P of text P in formula (2) w to obtain the deep multiplicative alignment feature of word-grained text P as shown in formula (15): Then, subtract and align the text P feature at the word granularity level in formula (9) from the text P word granularity feature P in formula (2) w and add them to obtain the deep subtraction alignment feature of text P at the word granularity level as shown in formula (16): Next, connect the deep soft alignment features of text P at the word granularity level in formula (14) The deep multiplication alignment features of text P at the word granularity level in formula (15) The deep subtraction alignment features of text P at the word granularity level in formula (16) Obtain the high-level feature P of text P at the word granularity level w ', as shown in formula (17): The deep soft alignment feature of text P at the word granularity level in connection formula (10) and the deep soft alignment feature of text P at the word granularity level in formula (14) to obtain the deep semantic feature P' of text P deep , as shown in formula (18): Similarly, performing a similar process on text Q as on text P yields the deep soft alignment features of text Q at the character granularity level. The deep multiplication alignment features of text Q at the character granularity level The deep subtraction alignment features of text Q at the character granularity level The high-level features Q' of text Q at the character granularity level c And the deep soft alignment features of text Q at the word granularity level The deep multiplication alignment features of text Q at the word granularity level The deep subtraction alignment features of text Q at the word granularity level The high-level features Q' of text Q at the word granularity level w The deep semantic features Q' of text Q deep , which completes the extraction of fine-grained semantic features of the same text. The second sub-module: Extracting inter-text semantic interaction features The first layer of the encoding structure uses several layers of encoding structures at the same time to extract the initial inter-text semantic interaction features: At the character granularity level, first, the character granularity feature P of text P in formula (1) c and the character granularity feature Q of text Q c are subjected to soft alignment attention to obtain the soft alignment interaction feature of text P at the character granularity level and the soft alignment interaction feature of text Q at the character granularity level as shown in formula (19): Secondly, the character-grained feature P of text P in formula (1) c and the character-grained feature Q of text Q c are subjected to subtraction alignment attention to obtain the character-grained text P subtraction alignment interaction feature and the character-grained text Q subtraction alignment interaction feature as shown in formula (20): Similar to the character level, first, the word-level features of text P in formula (2) w and the word-level features of text Q w are subjected to soft alignment attention to obtain the word-level soft alignment interaction features of text P and the word-level soft alignment interaction features of text Q As shown in formula (21): Then, the word granularity feature P of text P in formula (2) w and the word granularity feature Q of text Q w are subjected to subtraction alignment attention to obtain the word granularity-level text P subtraction alignment interaction feature and the word granularity-level text Q subtraction alignment interaction feature as shown in formula (22): The second layer of the encoding structure enhances the initial inter-text semantic interaction features to complete the extraction of inter-text semantic interaction features: At the word granularity level, first, the text P soft-alignment interaction feature at the word granularity level in formula (19) is added to the word granularity feature P of text P in formula (1) c to obtain the deep soft-alignment interaction feature of text P at the word granularity level as shown in formula (23): After that, subtract and align the interaction features of the character-level text P in formula (20) from the character-level feature P of the text P in formula (1) c and add them to obtain the deep subtraction alignment interaction features of the character-level text P as shown in formula (24): Finally, connect the deep soft alignment interaction features of the word-level text P in formula (23) with the deep subtraction alignment interaction features of the word-level text P in formula (24) to obtain the high-level interaction feature P of the word-level text P c ”, as shown in formula (25): At the word granularity level, first, add the word granularity-level text P soft alignment interaction feature in formula (21) to the word granularity feature of text P in formula (2) to obtain P w the deep soft alignment interaction feature of text P at the word granularity level as shown in formula (26): After that, subtract and align the interaction features of the word-grained text P in formula (22) from the word-grained feature P of the text P in formula (2) w and add them to obtain the deep subtraction-aligned interaction features of the word-grained text P as shown in formula (27): Finally, connect the deep soft alignment interaction features of text P at the word granularity level in formula (26) with the deep subtraction alignment interaction features of text P at the word granularity level in formula (27) to obtain the high-level interaction feature P″ of text P at the word granularity level w , as shown in formula (28): The deep subtraction alignment interaction feature of the character-grained text P in the connection formula (24) and the deep subtraction alignment interaction feature of the word-grained text P in the formula (27) to obtain the deep semantic interaction feature P″ of the text P deep , as shown in the formula (29): Similarly, performing a similar process on text Q as on text P yields the deep soft alignment interaction features of text Q at the character granularity level. The deep subtraction alignment interaction features of text Q at the character granularity level The high-level interaction features Q″ of text Q at the character granularity level c and the deep soft alignment interaction features of text Q at the word granularity level The deep subtraction alignment interaction features of text Q at the word granularity level The high-level interaction features Q″ of text Q at the word granularity level w and the deep semantic interaction features Q″ of text Q deep , thus completing the extraction of semantic interaction features between texts.
4. The text semantic matching method for medical intelligent question answering according to claim 1, characterized in that The implementation details of the feature fusion layer are as follows: First, for the convenience of subsequent description, make the following definitions: Define the operation of subtracting vectors and then taking the absolute value bit by bit, denoted as AB, as shown in formula (30): AB(P, Q) = |P - Q| (30) Where P and Q are two different vectors, indicating the operation of taking the absolute value bit by bit after subtracting the two vectors P and Q; Define the vector bitwise multiplication operation, denoted as MU, as shown in formula (31): MU(P, Q) = P ⊙ Q (31) Where P and Q are two different vectors, indicating the operation of multiplying the vectors P and Q bit by bit; In subsequent descriptions, use the AB symbol to represent the operation of formula (30), use the MU symbol to represent the operation of formula (31), use the SOA symbol to represent the operation of formula (3), use the MUA symbol to represent the operation of formula (4), use the SUA symbol to represent the operation of formula (5), and use the SEA symbol to represent the operation of formula (6); The feature fusion layer is divided into two sub-modules. The first sub-module merges multiple relevant features, and the second sub-module performs multiple matching operations to obtain the final matching feature vector; The first sub-module merges multiple relevant features: Connect the high-level features P of the word-grained text P in formula (13) c ' with the high-level interaction features P of the word-grained text P in formula (25) c ” to obtain the aggregated features of the word-grained text P And take the aggregated features of the word-grained text P Perform self-attention to obtain the deep aggregated features of the word-grained text P As shown in formula (32): At the word granularity level, similar to the character granularity level, the high-level features P' of the text P at the word granularity level in the connection formula (17) w and the high-level interaction features P'' of the text P at the word granularity level in the formula (28) w are used to obtain the aggregated features of the text P at the word granularity level and the aggregated features of the text P at the word granularity level are subjected to self-attention to obtain the deep aggregated features of the text P at the word granularity level as shown in the formula (33): Subsequently, connect the deep aggregation features of text P at the word granularity level in formula (32) with the deep aggregation features of text P at the word granularity level in formula (33) Then perform a max-pooling operation to obtain the pooled semantic feature P' of text P, as shown in formula (34): Next, connect the deep semantic feature P' of the text P in formula (18) deep with the deep semantic interaction feature P″ of the text P in formula (29) deep to obtain the deep aggregation feature of the text P as shown in formula (35): Similarly, perform the same operations on text Q as on text P to obtain the word-granularity-level text Q aggregation features Word-granularity-level deep aggregation features of text Q Phrase-granularity-level aggregation features of text Q Phrase-granularity-level deep aggregation features of text Q Semantic features Q' of text Q and deep aggregation features of text Q after pooling Then, the deep aggregation feature of text P in formula (35) and the deep aggregation feature of text Q are subjected to soft-alignment attention to obtain the deep aggregation feature of text P after soft alignment and the deep aggregation feature of text Q after soft alignment as shown in formula (36): Subsequently, the deep aggregated features of the soft-aligned text P in formula (36) are subjected to a max pooling operation to obtain the deep aggregated features P” of the pooled text P, and the deep aggregated features of the soft-aligned text Q are subjected to a max pooling operation to obtain the deep aggregated features Q” of the pooled text Q, as shown in formula (37): The second sub-module performs multiple matching operations to obtain the final matching feature vector: First, take the absolute value subtraction of the semantic feature P' of the pooled text P and the semantic feature Q' of the pooled text Q in formula (34) to obtain the subtraction matching feature PQ ab , as shown in formula (38): PQ ab = AB(P' - Q') (38) Secondly, the semantic feature P' of the pooled text P and the semantic feature Q' of the pooled text Q in formula (34) are multiplied to obtain the dot product matching feature PQ mu , as shown in formula (39): PQ mu = MU(P', Q') (39) Again, take the absolute value subtraction of the deep aggregated feature P” of the pooled text P in formula (37) and the deep aggregated feature Q” of the pooled text Q in formula (37) to obtain the deep subtraction matching feature PQ'. ab , as shown in formula (40): PQ' ab = AB(P”,Q”) (40) After that, the deep aggregated feature \(P''\) of the pooled text \(P\) in formula (37) is dot-multiplied with the deep aggregated feature \(Q''\) of the pooled text \(Q\) in formula (37) to obtain the deep dot-multiplication matching feature \(PQ'\). mu , as shown in formula (41): PQ' mu = MU(P”,Q”) (41) Finally, connect the semantic feature P' of the pooled text P, the semantic feature Q' of the pooled text Q in formula (34), the subtraction matching feature PQ in formula (38) ab , the dot product matching feature PQ in formula (39) mu , the deep subtraction matching feature PQ' in formula (40) ab , the deep dot product matching feature PQ' in formula (41) mu to obtain the final matching feature vector F, as shown in formula (42): F = [P'; Q'; PQ ab ; PQ mu ; PQ' ab ; PQ' mu (42).
5. The text semantic matching method for medical intelligent Q&A according to claim 1, characterized in that The implementation details of the prediction layer are as follows: Taking the final matching feature vector F as the input, a three-layer fully connected layer is used, and the ReLU activation function is used for activation after the first and second fully connected layers, and the sigmoid function is used for activation after the third fully connected layer, so as to obtain a matching degree value between [0,1], denoted as y pred ; Finally, by comparing with the set threshold of 0.5, it is judged whether the text semantics match; that is, y pred ≥0.5, it is predicted that the semantics of the text match, otherwise, they do not match; When the text semantic matching model has not been trained, it needs to be trained on the training dataset constructed according to the semantic matching knowledge base to optimize the model parameters; when the model training is completed, the prediction layer can predict whether the semantics of the target text match.
6. The text semantic matching method for medical intelligent question answering according to claim 1, characterized in that The text semantic matching knowledge base includes obtaining raw data by downloading datasets on the network, preprocessing the raw data, and aggregating sub-knowledge bases; Obtaining raw data by downloading datasets on the network: Download publicly available text semantic matching datasets or artificially constructed datasets on the network and use them as the raw data for constructing the text semantic matching knowledge base; Preprocessing the raw data: Preprocess the raw data used to construct the text semantic matching knowledge base, perform word breaking and word segmentation operations on each text therein to obtain the text semantic matching word breaking processing knowledge base and the word segmentation processing knowledge base; Aggregating sub-knowledge bases: Aggregate the text semantic matching word breaking processing knowledge base and the text semantic matching word segmentation processing knowledge base to construct the text semantic matching knowledge base; The text semantic matching model is obtained by training using the training dataset, and the construction process of the training dataset includes constructing training positive examples, constructing training negative examples, and constructing the training dataset; Constructing training positive examples: For each text in the text semantic matching knowledge base, if its semantics are the same, the text can be used to construct training positive examples; Constructing training negative examples: Select a text txt P, and then randomly select a text txt Q that does not match the text txt P from the text semantic matching knowledge base, and combine txt P and txt Q to construct negative examples; Constructing the training dataset: Combine all the positive example data and negative example data obtained after the operations of constructing training positive examples and constructing training negative examples, and shuffle their order to construct the final training dataset; After the text semantic matching model is constructed, it is trained and optimized through the training dataset as follows: Construct the loss function: As can be seen from the implementation process of the prediction layer, y pred is the calculated value of the matching degree obtained after being processed by the text semantic matching model; while y true is the true label of whether two text semantics match, and its value is limited to 0 or 1. The cross-entropy is used as the loss function; Constructing an optimization function: Use the Adam optimization function; perform optimization training on the text semantic matching model on the training dataset.
7. A text semantic matching device for medical intelligent question answering, characterized in that, The device includes a text semantic matching knowledge base construction unit, a training dataset generation unit, a text semantic matching model construction unit, and a text semantic matching model training unit, which respectively implement the steps of the text semantic matching method for medical intelligent question answering described in claims 1-6.
8. A storage medium storing multiple instructions, characterized in that, The instructions are loaded by the processor and execute the steps of the text semantic matching method for medical intelligent question answering described in claims 1-6.
9. An electronic device, characterized in that, The electronic device includes: The storage medium and the processor described in claim 8, which are used to execute the instructions in the storage medium.
Citation Information
Patent Citations
Text pair semantic matching method oriented to bank consultation service and based on multi-granularity alignment
CN113065358A
Natural language processing method, training method, and data processing device
WO2020147369A1