Text matching method and device based on pinyin modeling for medical question answering
Through pinyin modeling and text semantic matching technology, a text semantic matching knowledge base is constructed, and BiLSTM and hierarchical expansion convolutional networks are used to process text, which solves the problem of unprofessional and repetitive patient question descriptions in online medical Q&A and improves text matching performance and answer speed.
Patent Information
- Application Number
- CN202310819853.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-06
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2043-07-06
AI Technical Summary
In online medical Q&A communities, patients' question descriptions are unprofessional and contain a lot of repetition, resulting in heavy workload for doctors and slow response times. Pinyin input errors also affect text matching performance.
Build a text semantic matching knowledge base, use pinyin modeling, use BiLSTM and hierarchical dilated convolutional networks to process text, and combine multi-correlation feature interaction and three-dimensional convolutional feature encoding to improve text matching performance.
Effectively alleviate the impact of typos, improve text matching performance, reduce doctors' workload, and increase the speed of answering questions.
Smart Images

Figure CN116991991B_ABST
Abstract
Description
[0001] Technical issues
[0002] The present invention relates to the technical field of natural language processing, and in particular to a text matching method and device based on pinyin modeling for medical question answering.
[0003] With the rapid development of the internet, the traditional medical industry has begun to embrace the internet, resulting in a large number of patient-friendly applications. Among these, online medical Q&A communities have experienced rapid growth, significantly reducing the number of steps patients have to seek medical help. Patients can post questions in online communities, and doctors with relevant expertise can answer their questions, eliminating the traditional medical consultation process of visiting a hospital, registering, and waiting in line. However, as the number of users asking questions online has increased, the number of questions has far outstripped the number of answers. This problem stems from two factors: first, the number of patients far exceeds the number of professional doctors, and the number of doctors available to answer questions in the community is even smaller. Second, patients lack professional medical knowledge and are unable to express their questions professionally. Consequently, they often provide different descriptions of the same question, resulting in a large number of duplicate questions in the Q&A community. Text matching technology is used to match patient questions with question-answer pairs stored in a database (i.e., duplicate question-answer pairs), and then select the corresponding answers to provide to users. This reduces doctors' workload and speeds up answering.
[0004] Pinyin input method is a mainstream method for people to type, however, typos are often made when using this method to input text. Wrong text will cause the deep learning model to learn wrong semantics, resulting in a decrease in matching performance, so alleviating the impact of typos is crucial to improving text matching technology. An effective method is to integrate Pinyin into the model building process. Although the text with typos is different, it has the same Pinyin. Using Pinyin can retain some correct semantics to alleviate the impact of typos. Therefore, the present invention constructs a text matching model based on Pinyin, and applies it to medical Q&A, so as to select the most relevant answer from the existing question and answer pairs based on the questions raised by the user. Summary of the invention:
[0005] The technical task of this invention is to provide a text matching method and device, storage medium, and electronic device for medical question-and-answer (Q&A) based on pinyin modeling. This method solves the problem of using natural language processing technology to match patients' questions with existing questions and provide feedback to patients on the answers to the matched preset questions, thereby reducing the doctor's workload and speeding up the answering process. By integrating pinyin into the modeling process, this invention effectively alleviates the impact of typos in patients' input questions on the question's semantics, effectively improving the performance of text matching and further promoting the development of medical Q&A.
[0006] The technical task of the present invention is achieved in the following manner: a text matching method based on pinyin modeling for medical question answering, the method comprising the following steps:
[0007] S1. Build a text semantic matching knowledge base: Collect frequently asked medical question and answer texts from the Internet and pre-process the texts to build a text semantic matching knowledge base;
[0008] S2. Constructing a training dataset for a text semantic matching model: For the text pairs obtained in step S1, if they are semantically consistent with the text pairs in the knowledge base, the text pairs are used to construct positive training examples; otherwise, they are used to construct negative training examples. A large amount of positive and negative data are mixed to obtain a training dataset.
[0009] S3. Build a text semantic matching model: Build a text semantic matching model based on pinyin;
[0010] S4. Training the text semantic matching model: The text semantic matching model constructed in step S3 is trained on the text semantic matching model training dataset obtained in step S2.
[0011] A text matching device for medical question answering based on pinyin modeling, the device comprising:
[0012] The text semantic matching knowledge base construction unit is used to obtain a large amount of text pair data and then pre-process the data to obtain a text semantic matching knowledge base that meets the training requirements;
[0013] The text semantic matching model training dataset generation unit generates a text pair in the semantic matching knowledge base. If the text pairs have the same semantics, the text pairs are used to construct positive training examples; otherwise, they are used to construct negative training examples. A large amount of positive and negative data are mixed to obtain a training dataset.
[0014] A text semantic matching model construction unit is used to construct a word mapping conversion table, a pinyin mapping conversion table, an input encoding module, a word vector mapping layer, a pinyin vector mapping layer, a text semantic encoding layer, a multi-correlation feature interaction layer, a three-dimensional convolutional feature encoding layer, and a prediction layer;
[0015] The text semantic matching model training unit is used to construct the loss function and optimization function required in the model training process and complete the model training.
[0016] A storage medium stores a plurality of instructions, which are loaded by a processor to execute the steps of the above-mentioned text matching method based on pinyin modeling for medical question answering.
[0017] An electronic device comprises: the above-mentioned storage medium; and a processor for executing instructions in the storage medium.
[0018] Technical effects:
[0019] (1) The present invention processes the text semantic coding layer to extract the character representation of the text, the word representation of the text and the pinyin representation of the text, which can more comprehensively learn the semantic information of the text. In addition, the pinyin information can effectively retain the correct semantics in the text;
[0020] (2) The present invention uses multiple correlation feature interaction layers and multiple calculation methods, namely, dot multiplication, absolute value subtraction, and multiplication, to construct a correlation matrix between texts. Compared with a single calculation method, more correlation matrices are generated to improve the performance of text matching;
[0021] (3) The present invention constructs a correlation matrix based on multiple granularity information between texts, namely characters, words, and pinyin, through a multi-correlation feature interaction layer. Compared with single-granularity information, it generates more correlation matrices and further constructs a correlation feature graph, which can effectively improve the performance of the model.
[0022] (4) The present invention constructs a correlation cube based on the correlation feature map through a three-dimensional convolutional feature encoding layer, and further uses a three-dimensional convolutional neural network model to generate a matching representation of the input text pair. This method models the matching representation of the input text pair from a higher dimension, thereby achieving better performance;
[0023] (5) The present invention can effectively improve the effect of semantic matching by comprehensively using the text semantic coding layer, the multi-correlation feature interaction layer, and the three-dimensional convolution feature coding layer.
[0024] The present invention will be further described below with reference to the accompanying drawings.
[0025] Figure 1 Flowchart of a text matching method based on pinyin modeling for medical question answering
[0026] Figure 2 Flowchart for building a text semantic matching knowledge base
[0027] Figure 3 Flowchart for building a training dataset for a text semantic matching model
[0028] Figure 4 Flowchart for building a text semantic matching model
[0029] Figure 5 Flowchart for training text semantic matching model
[0030] Figure 6Schematic diagram of the structure of the hierarchical expansion convolutional network
[0031] Figure 7 Schematic diagram for constructing the structure of the dot product correlation matrix
[0032] Figure 8 Schematic diagram of the framework of the text matching model based on pinyin modeling for medical question answering Specific implementation method:
[0033] The text matching method and device for medical question answering based on pinyin modeling of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments of the specification.
[0034] Example 1:
[0035] The overall framework structure of the present invention is as follows Figure 8 As shown. Figure 8 It can be seen that the main framework structure of the present invention includes a text semantic encoding layer, a multi-correlation feature interaction layer, a three-dimensional convolution feature encoding layer and a prediction layer. Among them, the text semantic encoding layer first embeds the character granularity information, word granularity information and pinyin granularity information of the input text to obtain character embedding representation, word embedding representation and pinyin embedding representation; then uses a bidirectional LSTM, namely BiLSTM, to encode them respectively to generate character context representation, word context representation and pinyin context representation, and pass them into the hierarchical expansion convolution network. The structural diagram of the hierarchical expansion convolution network is shown in the figure. Figure 6As shown, it takes the contextual representation as input, i.e., the character / word / pinyin contextual representation, and encodes it through a dilated convolutional neural network with a dilation rate of 1 to obtain the first-layer convolutional representation, i.e., the first-layer convolutional representation of the character / word / pinyin. It then encodes it through a dilated convolutional neural network with a dilation rate of 2 to obtain the second-layer convolutional representation. It further encodes it through a dilated convolutional neural network with a dilation rate of 3 to obtain the third-layer convolutional representation. Finally, the first, second, and third-layer convolutional representations are concatenated to obtain the dilated convolutional text representation. The character contextual representation, word contextual representation, and pinyin contextual representation are encoded through hierarchical dilated convolutional networks, respectively, to obtain the dilated convolutional character text representation, dilated convolutional word text representation, and dilated convolutional pinyin text representation. The text semantic encoding layer further concatenates the character contextual representation, word contextual representation, and pinyin contextual representation with the dilated convolutional character text representation, dilated convolutional word text representation, and dilated convolutional pinyin text representation, respectively, to generate the text character representation, text word representation, and text pinyin representation, which are then passed to the multi-correlation feature interaction layer. The multi-correlation feature interaction layer receives the representations of two texts obtained through the text semantic encoding layer, namely, the character representation of text P, the character representation of text Q, the word representation of text P, the word representation of text Q, the pinyin representation of text P, and the pinyin representation of text Q; first, based on different calculation methods, a dot product correlation matrix, an absolute value subtraction correlation matrix, and a multiplication correlation matrix are constructed respectively; then they are stacked to construct a correlation feature map based on dot product, a correlation feature map based on multiplication, and a correlation feature map based on absolute value subtraction. Among them, the structural diagram of constructing the dot product correlation matrix is as follows: Figure 7As shown, it takes the character representation of text P, the word representation of text P, the pinyin representation of text P, the character representation of text Q, the word representation of text Q, and the pinyin representation of text Q as input; performs dot multiplication of the character representation of text P with the character representation of text Q, the pinyin representation of text Q, and the word representation of text Q, respectively, to obtain the dot product correlation matrix between the characters of text P and the characters / pinyin / words of text Q; performs dot multiplication of the pinyin representation of text P with the character representation of text Q, the pinyin representation of text Q, and the word representation of text Q, respectively, to obtain the dot product correlation matrix between the pinyin of text P and the characters / pinyin / words of text Q; performs dot multiplication of the word representation of text P with the character representation of text Q and the pinyin representation of text Q, respectively, to obtain the dot product correlation matrix between the words of text P and the characters / pinyin of text Q. The multi-correlation feature interaction layer passes the dot product-based correlation feature map, the multiplication-based correlation feature map, and the absolute value subtraction-based correlation feature map into the three-dimensional convolutional feature encoding layer. The three-dimensional convolutional feature encoding layer stacks these to construct a relevance cube. This is then encoded using a three-dimensional convolutional neural network to obtain a matching representation of the input text pair, which is then passed to the prediction layer. The prediction layer encodes the matching representation of the received input text pair using a fully connected layer to obtain a matching value.
[0036] Example 2:
[0037] As attached Figure 1 As shown, the text matching method for medical question answering based on pinyin modeling of the present invention comprises the following steps:
[0038] S1. Build a text semantic matching knowledge base: Collect frequently asked medical question and answer texts from the Internet and pre-process the texts to build a text semantic matching knowledge base;
[0039] S2. Constructing a training dataset for a text semantic matching model: For the text pairs obtained in step S1, if they are semantically consistent with the text pairs in the knowledge base, the text pairs are used to construct positive training examples; otherwise, they are used to construct negative training examples. A large amount of positive and negative data are mixed to obtain a training dataset.
[0040] S3. Build a text semantic matching model: Build a text semantic matching model based on pinyin;
[0041] S4. Training the text semantic matching model: The text semantic matching model constructed in step S3 is trained on the text semantic matching model training dataset obtained in step S2.
[0042] S1. Build a text semantic matching knowledge base: Collect frequently asked medical question and answer texts from the Internet and pre-process the texts to build a text semantic matching knowledge base;
[0043] S101. Data collection: Collect frequently asked medical question texts as raw data for the text semantic matching knowledge base;
[0044] For example, we collect question text pairs from medical Q&A and use them as raw data. The text pair examples are as follows:
[0045] Txt P What causes hemoptysis after strenuous exercise? TxtQ What causes hemoptysis after strenuous exercise?
[0046] S102, pre-processing the original data for constructing a text semantic matching knowledge base, performing a hyphenation operation, a word segmentation operation, and a pinyin conversion operation on each text therein, to obtain a text semantic matching hyphenation processing knowledge base, a word segmentation processing knowledge base, and a pinyin processing knowledge base;
[0047] Each text in the original data obtained in step S101 for constructing the text semantic matching knowledge base is pre-processed with hyphenation, word segmentation, and pinyin conversion, thereby obtaining a text semantic matching hyphenation processing knowledge base, a word segmentation processing knowledge base, and a pinyin processing knowledge base; the specific steps of the hyphenation operation are: taking each character in the Chinese text as a unit, each text is segmented using a space as a delimiter; the specific steps of the word segmentation operation are: using the Jieba word segmentation tool and selecting the default precise mode to segment each text; the specific steps of the pinyin conversion operation are: using the PyPinyin toolkit to convert each character in the text into pinyin and save it; in this operation, in order to avoid the loss of semantic information, all contents in the text, including punctuation marks, special characters, and stop words, are retained;
[0048] For example, taking the txt P shown in S101 as an example, after the word segmentation operation is performed on it, the result is "After strenuous exercise, coughing up blood, what's wrong?"; after the word segmentation operation is performed on it using the Jieba word segmentation tool, the result is "After strenuous exercise, coughing up blood, what's wrong?"; after the word segmentation operation is performed on it using the PyPinyin toolkit, the result is "ju lie yun dong houke xue, shi zen mo le?".
[0049] S103, summarizing the text semantic matching hyphenation processing knowledge base, the text semantic matching word segmentation processing knowledge base, and the text semantic matching pinyin processing knowledge base to construct a text semantic matching knowledge base;
[0050] The text semantic matching word segmentation processing knowledge base, text semantic matching word segmentation processing knowledge base, and text semantic matching pinyin processing knowledge base obtained in step S102 are aggregated into the same folder to obtain a text semantic matching knowledge base. The process is as follows: Figure 2As shown; it should be noted here that the data processed by the hyphenation operation, the data processed by the word segmentation operation, and the data processed by the pinyin conversion operation will not be merged into the same file, that is, the text semantic matching knowledge base actually contains three independent sub-knowledge bases; each preprocessed text retains the ID information of its original text.
[0051] S2. Constructing a training dataset for a text semantic matching model: For the text pairs obtained in step S1, if they are semantically consistent with the text pairs in the knowledge base, the text pairs are used to construct positive training examples; otherwise, they are used to construct negative training examples. A large amount of positive and negative data are mixed to obtain a training dataset.
[0052] S201. Construct training positive data: construct two texts with consistent text semantics as positive data; since the text semantic matching knowledge base actually contains three independent sub-knowledge bases, the same text will have three initial representations, namely (txt P_char,txt Q_char,1), (txt P_word,txt Q_word,1), and (txt P_pinyin,txt Q_pinyin,1); merge them together, and the final constructed positive data can be formalized as: (txt P_char,txt Q_char,txt P_word,txt Q_word,txt P_pinyin,txt Q_pinyin,1).
[0053] Among them, txt P_char and txt Q_char refer to the character sequence of text P and the character sequence of text Q in the text semantic matching word segmentation processing knowledge base, respectively; txt P_word and txt Q_word refer to the word sequence of text P and the word sequence of text Q in the text semantic matching word segmentation processing knowledge base, respectively; txt P_pinyin and txt Q_pinyin refer to the pinyin sequence of text P and the pinyin sequence of text Q in the text semantic matching pinyin processing knowledge base, respectively; and the 1 here indicates that the semantics of the two texts match, which is a positive example;
[0054] For example, for txt P and txt Q shown in S101, after the hyphenation operation, word segmentation operation, and pinyin conversion operation, the constructed positive example data format is:
[0055] (“I coughed up blood after strenuous exercise, what happened?”,“What is the reason for coughing up blood after strenuous exercise?”,“I coughed up blood after strenuous exercise, what happened?”,“What is the reason for coughing up blood after strenuous exercise?”,“ju lie yun dong hou ke xue,shi zenmo le?”,“ju lie yun dong hou ke xue shi shen mo yuan yin?”,1).
[0056] S202, constructing training negative examples: For each positive example text obtained in step S201, select a text contained in it, and randomly select a text that does not match it to combine; these two semantically inconsistent texts are constructed as negative example data; using operations similar to step S201, the negative example data can be formalized as: (txt P_char, txt Q_char, txt P_word, txt Q_word, txt P_pinyin, txt Q_pinyin, 0); the meaning of each symbol is the same as in step S201, and 0 indicates that the semantics of the two texts do not match, which is a negative example;
[0057] For example: Since the construction method of negative examples is very similar to that of positive examples, I will not go into details here;
[0058] S203. Construct a training data set: merge all the positive example text data and negative example text data obtained after the operations in step S201 and step S202, and shuffle their order to construct the final training data set; both positive example data and negative example data contain seven dimensions, namely, txt P_char, txt Q_char, txt P_pinyin, txt Q_pinyin, txt P_word, txt Q_word, 0 or 1.
[0059] S3. Build a text semantic matching model: The process of building a text semantic matching model is as follows Figure 4 As shown, the main operations are to build a word mapping conversion table, build a pinyin mapping conversion table, build an input encoding module, build a word vector mapping layer, build a pinyin vector mapping layer, build a text semantic encoding layer, build a multi-correlation feature interaction layer, build a three-dimensional convolution feature encoding layer, and build a prediction layer; the specific steps are as follows:
[0060] S301, constructing a word mapping conversion table: The word table is constructed by using the text semantic matching word segmentation processing knowledge base and the word segmentation processing knowledge base obtained after the processing in step S102; after the word table is constructed, each word in the table is mapped to a unique numerical identifier, and the mapping rule is: starting with the number 1, and then sorting each word in the order in which it was entered into the word table in ascending order, thereby forming the required word mapping conversion table;
[0061] For example, the content processed in step S102 is “coughing up blood after strenuous exercise, what’s wrong?” and “coughing up blood after strenuous exercise, what’s wrong?”. The word table and word mapping conversion table are constructed as follows:
[0062]
[0063] Afterwards, the present invention uses Word2Vec to train a word vector model to obtain a word vector matrix char_word_embedding_matrix for each word;
[0064] For example, in Keras, the code described above is implemented as follows:
[0065] w2v_model_char_word=genism.models.Word2Vec(w2v_corpus_char_word, size=EMBDIM, window=5, min_count=1, sg=1, workers=4, seed=1234, iter=25)
[0066] tokenizer=Tokenizer(num_words=len(char_word_set))
[0067] tokenizer.fit_on_texts(corpus)
[0068] L=len(tokenizer.char_word_index)
[0069] char_word_embedding_matrix=np.zeros([len(tokenizer.word_index)+1,EMBDIM])
[0070] for char_word,idx in tokenizer.word_index.items():
[0071] if char_word in w2v_model.wv:
[0072] char_word_embedding_matrix[idx,:]=w2v_model.wv[char_word]
[0073] Among them, w2v_corpus_char_word is the training corpus for word segmentation and word segmentation, that is, all the data in the knowledge base for word segmentation and text semantic matching; EMBDIM is the vector dimension of the word. This model sets EMBDIM to 300, and char_word_set is the word list.
[0074] S302, constructing a pinyin mapping conversion table: The pinyin table is constructed by matching the text semantics obtained after step S102 with the pinyin processing knowledge base; after the pinyin table is constructed, each pinyin in the table is mapped to a unique digital identifier, and the mapping rule is: starting with the number 1, and then sorting each pinyin in the order in which it is entered into the pinyin table in ascending order, thereby forming the required pinyin mapping conversion table;
[0075] For example, the content processed in step S102, "ju lie yun dong hou ke xue, shi zenmo le?", is constructed as follows:
[0076] Pinyin ju lie yun dong hou ke xue , shi Zen mo le ? Mapping 1 2 3 4 5 6 7 8 9 10 11 12 13
[0077] Afterwards, the present invention uses Word2Vec to train the pinyin vector model to obtain the pinyin vector matrix pinyin_embedding_matrix of each word;
[0078] For example, in Keras, the code described above is implemented in the same way as in S301, except that the parameters are changed from char_word to pinyin. Due to space limitations, I will not go into details here.
[0079] Among them, the w2v_corpus_char_word in the example in S301 is replaced with w2v_corpus_piniyin, which is the pinyin processing training corpus, that is, the text semantics matches all the data in the pinyin processing knowledge base; the pinyin vector dimension is PY_EMBDIM, and this model sets PY_EMBDIM to 70; char_word_set is replaced with pinyin_set, which is the pinyin table.
[0080] S303. Construct the input module: The input module includes six inputs; for each text in the training dataset or the text to be predicted, use the corresponding modules in S1 and S2 to preprocess it, and respectively obtain the character sequence txt P_char of text P, the character sequence txt Q_char of text Q, the word sequence txt P_word of text P, the word sequence txt Q_word of text Q, the pinyin sequence txt P_pinyin of text P, and the pinyin sequence txt Q_pinyin of text Q, and formalize them as: (txt P_char, txt Q_char, txt P_word, txt Q_word, txt P_pinyin, txt Q_pinyin);
[0081] For each character, word, and pinyin in the input text, the present invention converts them into corresponding digital identifiers according to the word mapping conversion table and pinyin mapping conversion table constructed in step S301 and step S302;
[0082] ("Why do I cough up blood after strenuous exercise?", "What causes hemoptysis after strenuous exercise?", "Why do I cough up blood after strenuous exercise?", "What causes hemoptysis after strenuous exercise?", "ju lie yun dong hou ke xue,shi zenmo le?", "ju lie yun dong hou ke xue shi shen mo yuan yin?")
[0083] Each input data contains 6 texts; for the first four texts, according to the word mapping conversion table in step S301, convert them into numerical representations; for the next two texts, according to the pinyin mapping conversion table in step S302, convert them into numerical representations; the combined representation results of the 6 texts of the input data are as follows (assuming that the mapping relationships that appear in text 2 but not in text 1 are: "什": 17, "原": 18, "因": 19, "什么": 20, "原因": 21, "shen": 14, "yuan": 15, "yin": 16):
[0084] ("1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13", "1, 2, 3, 4, 5, 6, 7, 9, 17, 11, 18, 19, 13", "14, 5, 15, 8, 9, 16, 12, 13", "14, 5, 16, 9, 20, 21, 13", "1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13", "1, 2, 3, 4, 5, 6, 7, 9, 14, 11, 15, 16, 13").
[0085] S304, constructing a word vector mapping layer: initializing the weight parameters of the current layer by loading the word vector matrix trained in the step of constructing a word mapping conversion table; for the input texts txt P_char, txt Q_char and txt P_word, txt Q_word, obtaining their corresponding vectors txt P_char_embed, txt Q_char_embed, txt P_word_embed, txt Q_word_embed; text semantic matching, word segmentation processing, each text in the knowledge base is converted into a vector form through word vector mapping, i.e., character embedding representation and word embedding representation;
[0086] For example, in Keras, the code described above is implemented as follows:
[0087] char_word_embedding_layer=Embedding(char_word_embedding_matrix.shape[0],char_word_emb_dim,weights=[char_word_embedding_matrix],
[0088] input_length=input_dim, trainable=False)
[0089] Among them, char_word_embedding_matrix is the word vector matrix trained in the step of building the word mapping conversion table, char_word_embedding_matrix.shape[0] is the size of the word table of the word vector matrix, char_word_emb_dim is the dimension of the output word vector, and input_length is the length of the input sequence.
[0090] The corresponding texts txt P_char, txt Q_char, txt P_word, and txt Q_word are processed by the Keras Embedding layer to obtain the character embedding representations txt P_char_embed and txt Q_char_embed, and the word embedding representations txt P_word_embed and txt Q_word_embed of the corresponding texts.
[0091] S305, constructing a pinyin vector mapping layer: initializing the weight parameters of the current layer by loading the pinyin vector matrix trained in the step of constructing the pinyin mapping conversion table; for the input texts txt P_pinyin and txt Q_pinyin, obtaining their corresponding vectors txt P_pinyin_embed and txt Q_pinyin_embed; for each text in the text matching pinyin processing knowledge base, the text pinyin information is converted into a vector form, i.e., a pinyin embedding representation, by means of pinyin vector mapping;
[0092] For example: In Keras, the code implementation described above is basically the same as in S304, except that the parameters are changed from char to pinyin related; due to space limitations, I will not go into details here.
[0093] The corresponding texts txt P_pinyin and txt Q_pinyin are processed by the Embedding layer of Keras to obtain the corresponding text pinyin embedding representations txt P_pinyin_embed and txt Q_pinyin_embed.
[0094] S306, constructing a text semantic encoding layer: This layer receives the character embedding representation, word embedding representation output by the word vector mapping layer, and the pinyin embedding representation output by the pinyin vector mapping layer as input; first, using BiLSTM to encode the character embedding representation and word embedding representation respectively to generate a character context representation and a word context representation; at the same time, the pinyin embedding representation is encoded using a fully connected neural network and then passed to the BiLSTM to generate a pinyin context representation; further using a hierarchical dilated convolutional network encoding to obtain a character text representation after dilation convolution, a word text representation after dilation convolution, and a pinyin text representation after dilation convolution; and connecting them with the corresponding context representation to generate a character representation of the text, a word representation of the text, and a pinyin representation of the text;
[0095] S30601, BiLSTM feature extraction: The text semantic encoding layer first uses BiLSTM to encode the character embedding representation and word embedding representation of the two texts respectively to generate character context representation and word context representation. For the pinyin embedding representation, it is first encoded using a fully connected neural network and then passed to BiLSTM to generate the pinyin context representation. For text P, the specific operation is shown in the following formula:
[0096]
[0097] Where N is the length of the text; represents the character embedding representation of the text P at the i-th position under the character granularity c; Represents the word embedding representation of the text P at the jth position under the word granularity w; represents the pinyin embedding representation of text P at the mth position under pinyin granularity p; Dense represents the fully connected neural network encoding; and Represent the character context representation of text P, the word context representation of text P and the pinyin context representation of text P respectively; similarly, for text Q, we get Q respectively. c' , Q w' , Q p' , its symbolic meaning is similar to that of text P, except that Q represents text Q, and the rest represent similarities with this text, which can make their meaning clear, so we will not elaborate on them one by one here;
[0098] For example, in Keras, the code described above is implemented as follows:
[0099] w_embedding=embedding_layer(w_input)
[0100] c_embedding=embedding_layer(c_input)
[0101] py_embedding=py_embedding_layer(py_input)
[0102] w_r=Bidirectional(LSTM(300, return_sequences='True', dropout=0.2), merge_mode='sum')(w_embedding)
[0103] c_r=Bidirectional(LSTM(300, return_sequences='True', dropout=0.2), merge_mode='sum')(c_embedding)
[0104] py_embedding2=Dense(300)(py_embedding)
[0105] py_r=Bidirectional(LSTM(300, return_sequences='True', dropout=0.2), merge_mode='sum')(py_embedding2)
[0106] Among them, w_input, c_input and py_input represent text input at word granularity, text input at character granularity and text input at pinyin granularity respectively.
[0107] S30602, Hierarchical Dilated Convolutional Network Coding: Encode the character context representation, word context representation, and pinyin context representation of two texts using a hierarchical dilated convolutional network to generate a dilated convolutional character text representation, a dilated convolutional word text representation, and a dilated convolutional pinyin text representation of the corresponding texts;
[0108] Specifically, the hierarchical dilated convolutional network consists of three layers of dilated convolutional networks with dilation rates of 1, 2, and 3 respectively. The character / word / pinyin context representation is passed through the three layers of dilated convolutional networks to obtain the character / word / pinyin representation after the first layer of convolution, the character / word / pinyin representation after the second layer of convolution, and the character / word / pinyin representation after the third layer of convolution. Then, they are connected according to the characters, words, and pinyin to generate the character / word / pinyin text representation after dilated convolution. For the convenience of explanation, the character context representation P of the text P is used here. c' For example, the calculation formula is as follows:
[0109]
[0110] Among them, Dilated_CNN represents a one-dimensional dilated convolutional neural network; dilation_rate is the dilation rate; Represents the character representation of the text P after the first layer of convolution; Represents the character representation of the text P after the second convolution layer; Represents the character representation of the text P after the third convolution layer; Represents the character text representation of the text P after dilation and convolution; [;] represents the connection operation; similarly, for P w' 、P p' , Q c' , Q w' , Q p' Process, get Its symbolic meaning is Similarly, the only difference is that Q represents text Q, superscript w represents word, and superscript p represents pinyin. By analogy, the meaning can be made clear, so I will not go into details here.
[0111] S30603, Feature Connection: For text P, the text semantic encoding layer uses the connection operation to respectively connect P c' and P w' and P p' and Perform the join to generate the character representation of text P, the word representation of text P, and the pinyin representation of text P. The specific formula is as follows:
[0112]
[0113] in, Represents the character representation of the text P; The word representation representing the text P; Represents the pinyin representation of text P; similarly, for text Q, we get and Its meaning is similar to the corresponding symbol of text P, the only difference is that Q represents text Q, which will not be repeated here;
[0114] For example, in Keras, the code described above is implemented as follows:
[0115] w_l=Conv1D(300,3,padding='same',dilation_rate=1,activation='relu')(w_r)
[0116] w_l2=Conv1D(300,3,padding='same',dilation_rate=2,activation='relu')(w_l)
[0117] w_l3=Conv1D(300,3,padding='same',dilation_rate=3,activation='relu')(w_l2)
[0118] w_r=concatenate([w_r,w_l,w_l2,w_l3])
[0119] c_l=Conv1D(300,3,padding='same',dilation_rate=1,activation='relu')(c_r)
[0120] c_l2=Conv1D(300,3,padding='same',dilation_rate=2,activation='relu')(c_l)
[0121] c_l3=Conv1D(300,3,padding='same',dilation_rate=3,activation='relu')(c_l2)
[0122] c_r=concatenate([c_r,c_l,c_l2,c_l3])
[0123] py_l=Conv1D(300,3,padding='same',dilation_rate=1,activation='relu')(py_r)
[0124] py_l2=Conv1D(300,3,padding='same',dilation_rate=2,activation='relu')(py_l)
[0125] py_l3=Conv1D(300,3,padding='same',dilation_rate=3,activation='relu')(py_l2)
[0126] py_r=concatenate([py_r,py_l,py_l2,py_l3])
[0127] S307, constructing a multi-correlation feature interaction layer: This layer receives the character representation of the text P in step S306 Word representation of text P Pinyin representation of text P Character representation of text Q Word representation of text Q and the pinyin representation of the text Q As input; construct the point product correlation matrix, absolute value subtraction correlation matrix, and multiplication correlation matrix respectively by point multiplication, absolute value subtraction, and multiplication operations; then construct the correlation feature map according to the correlation matrix;
[0128] S30701. Constructing a correlation matrix: The structural diagram for constructing a dot product correlation matrix is as follows: Figure 7As shown; it takes the character representation of text P, the word representation of text P, the pinyin representation of text P, the character representation of text Q, the word representation of text Q and the pinyin representation of text Q as input, and performs dot product of the character representation of text P with the character representation of text Q, the pinyin representation of text Q and the word representation of text Q, to obtain the dot product correlation matrix of text P characters and text Q characters / pinyin / words; performs dot product of the pinyin representation of text P with the character representation of text Q, the pinyin representation of text Q and the word representation of text Q, to obtain the dot product correlation matrix of text P pinyin and text Q characters / pinyin / words; performs dot product of the word representation of text P with the character representation of text Q and the pinyin representation of text Q, to obtain the dot product correlation matrix of text P words and text Q characters / pinyin; taking the dot product of the character representation of text P with the character representation of text Q, the pinyin representation of text Q and the word representation of text Q as an example, its specific formula is as follows:
[0129]
[0130] in, The subscript a in the matrix indicates that the matrix is calculated by dot product, and the subscript ij indicates the dot product correlation matrix of text P characters and text Q characters. The value at row i and column j in ; The superscripts c and c represent that the matrix is generated according to the character representation of text P and the character representation of text Q respectively; similarly Respectively represent The value at the i-th row and j-th column in the matrix; the superscripts c and p represent that the matrix is generated according to the character representation of text P and the pinyin representation of text Q, and c and w represent that the matrix is generated according to the character representation of text P and the word representation of text Q; Represent the dot product correlation matrix of text P characters and text Q pinyin / words respectively; Character representation of text P The character representation at position i in ; Character representation of the text Q The character representation at the j-th position in ; Represents the pinyin representation of the text Q The pinyin representation at the j-th position in ; Word representation representing text Q The word representation at the jth position in ; The generation method of other point product correlation matrices is similar to it, and we get, Its symbolic meaning is Similarly, by analogy with this, the meaning can be made clear, so I will not elaborate on them here;
[0131] For example, in Keras, the code described above is implemented as follows:
[0132] pc_qc=Dot(axes=-1)([pc,qc])
[0133] pc_qpy=Dot(axes=-1)([pc,qpy])
[0134] pc_qw=Dot(axes=-1)([pc,qw])
[0135] ppy_qpy=Dot(axes=-1)([ppy,qpy])
[0136] ppy_qw=Dot(axes=-1)([ppy,qw])
[0137] ppy_qc=Dot(axes=-1)([ppy,qc])
[0138] pw_qc=Dot(axes=-1)([pw,qc])
[0139] pw_qpy=Dot(axes=-1)([pw,qpy])
[0140] Among them, pc represents the character representation of text P; ppy represents the pinyin representation of text P; pw represents the word representation of text P; qc represents the character representation of text Q; qpy represents the pinyin representation of text Q; qw represents the word representation of text Q.
[0141] The construction of the multiplication correlation matrix and the construction of the absolute value subtraction correlation matrix are similar to the construction of the dot product correlation matrix; the only difference is the generation method of the correlation matrix. Taking the character representation of text P and the character representation of text Q as an example, the specific formula is as follows:
[0142]
[0143] Among them, ⊙ represents element-wise dot product; |-| represents absolute value subtraction; TimeDistributed represents the use of fully connected neural network encoding (Dense) at each time step; tanh represents the tanh function; and The superscript meaning is the same as The subscripts mul and sub respectively represent that the correlation matrix is obtained by multiplication and absolute value subtraction; similar to constructing the dot product correlation matrix; using the same method to encode different inputs, we can get and The meaning of its symbols and construction formula can be obtained by analogy, so I will not go into details here;
[0144] For example, in Keras, the code described above is implemented as follows:
[0145] pc_qc_multi=multi_attn(pc,qc)
[0146] pc_qpy_multi=multi_attn(pc,qpy)
[0147] pc_qw_multi=multi_attn(pc,qw)
[0148] ppy_qpy_multi=multi_attn(ppy,qpy)
[0149] ppy_qc_multi=multi_attn(ppy,qc)
[0150] ppy_qw_multi=multi_attn(ppy,qw)
[0151] pw_qc_multi=multi_attn(pw,qc)
[0152] pw_qpy_multi=multi_attn(pw,qpy)
[0153] pc_qc_abs=abs_attn(pc,qc)
[0154] pc_qpy_abs=abs_attn(pc,qpy)
[0155] pc_qw_abs=abs_attn(pc,qw)
[0156] ppy_qpy_abs=abs_attn(ppy,qpy)
[0157] ppy_qc_abs=abs_attn(ppy,qc)
[0158] ppy_qw_abs=abs_attn(ppy,qw)
[0159] pw_qc_abs=abs_attn(pw,qc)
[0160] pw_qpy_abs=abs_attn(pw,qpy)
[0161] Among them, multi_attn and abs_attn are the defined multiplication and absolute value subtraction functions, and their codes are as follows:
[0162] def multi_attn(input_1,input_2):
[0163] attention=multiply([input_1,input_2])
[0164] attention=TimeDistributed(Dense(30,activation='tanh'))(attention)
[0165] return attention
[0166] def abs_attn(input_1,input_2):
[0167] attention=Lambda(lambda x:K.abs(x[0]-x[1]))([input_1,input_2])
[0168] attention=TimeDistributed(Dense(30,activation='tanh'))(attention)
[0169] return attention
[0170] S30702. Constructing a correlation feature map: The multi-correlation feature interaction layer stacks the correlation matrices obtained in step S30701 to construct a correlation feature map. The specific formula is as follows:
[0171]
[0172] Taking the dot product correlation matrix as an example, the dot product correlation matrix of text P characters and text Q characters / pinyin, the dot product correlation matrix of text P pinyin and text Q characters / pinyin, the dot product correlation matrix of text P words and text Q characters / pinyin, the dot product correlation matrix of text P characters and text Q words, and the dot product correlation matrix of text P pinyin and text Q words are stacked to construct four dot product-based correlation feature maps, namely A'1, A'2, A'3, and A'4; similarly, for the multiplication correlation matrix and the absolute value subtraction correlation matrix, four multiplication-based correlation feature maps, namely M'1, M'2, M'3, and M'4, and four absolute value subtraction-based correlation feature maps, namely U'1, U'2, U'3, and U'4, are formed respectively.
[0173] For example, in Keras, the code described above is implemented as follows:
[0174] pc = Lambda(stack_dim1)([pc_qc, pc_qpy])
[0175] ppy = Lambda(stack_dim1)([ppy_qc, ppy_qpy])
[0176] pw = Lambda(stack_dim1)([pw_qc, pw_qpy])
[0177] qw = Lambda(stack_dim1)([pc_qw, ppy_qw])
[0178] pc_multi = Lambda(stack_dim1)([pc_qc_multi, pc_qpy_multi])
[0179] ppy_multi = Lambda(stack_dim1)([ppy_qc_multi, ppy_qpy_multi])
[0180] pw_multi = Lambda(stack_dim1)([pw_qc_multi, pw_qpy_multi])
[0181] qw_multi = Lambda(stack_dim1)([pc_qw_multi, ppy_qw_multi])
[0182] pc_abs = Lambda(stack_dim1)([pc_qc_abs, pc_qpy_abs])
[0183] ppy_abs = Lambda(stack_dim1)([ppy_qc_abs, ppy_qpy_abs])
[0184] pw_abs = Lambda(stack_dim1)([pw_qc_abs, pw_qpy_abs])
[0185] qw_abs = Lambda(stack_dim1)([pc_qw_abs, ppy_qw_abs])
[0186] Among them, the code for stack_dim1 is as follows:
[0187] def stack_dim1(a):
[0188] x = K.stack(a, axis = 1)
[0189] return x
[0190] S308: Construct a three-dimensional convolutional feature encoding layer: further stack the correlation feature map based on dot multiplication, the correlation feature map based on multiplication, and the correlation feature map based on absolute value subtraction obtained in step S307 to construct a correlation cube; then encode the correlation cube through a three-dimensional convolutional neural network to generate a matching representation of the input text pair;
[0191] S30801. Constructing a correlation cube: The three-dimensional convolutional feature encoding layer first stacks the correlation feature map based on dot multiplication, the correlation feature map based on multiplication, and the correlation feature map based on absolute value subtraction obtained in step S307 to construct a correlation cube. The specific formula is as follows:
[0192]
[0193] Where F represents the correlation cube, stack represents the stacking operation, and the meanings of other symbols are the same as those in formula (6);
[0194] S30802, 3D Convolutional Neural Network Encoding: Use a 3D convolutional neural network to encode the relevance cube to generate a matching representation of the input text pair. The specific formula is as follows:
[0195] F′=Conv3d(F) (19)
[0196] Among them, F′ represents the matching representation of the input text pair, and Conv3d represents the three-dimensional convolutional neural network.
[0197] For example, in Keras, the code described above is implemented as follows:
[0198] stack2=Lambda(stack_dim1)([pc,ppy,pw,qw,pc_multi,ppy_multi,pw_multi,qw_multi,pc_abs,ppy_abs,pw_abs,qw_abs])
[0199] conv_sim=Conv3D(32, kernel_size=[3,3,3], strides=[2,2,2], activation='relu', padding='same')(stack2)
[0200] S309, construct prediction layer: The prediction layer takes the matching representation of the input text pair in step S308 as input, flattens it using the Flatten method, and then encodes it using a three-layer fully connected neural network to predict the matching degree y of the input text pair. pred and compare it with the preset threshold. If y pred If the value is greater than or equal to the threshold, the text is considered to be semantically matched, otherwise it is not matched.
[0201] S4, training the text semantic matching model: using the text semantic matching model training data set obtained in step S2 to train the text semantic matching model constructed in step S3;
[0202] S401, constructing a loss function: using the modified binary cross entropy as the loss function based on the matching degree of the input text pair obtained in step S309, the formula is as follows:
[0203]
[0204]
[0205]
[0206] Where θ(x) is the unit step function, thre is the threshold value, which is set to 0.6 in the present invention, and L represents the modified cross entropy formula; y pred Represents the matching degree of the predicted input text pair; y true The true label representing whether the input text pair matches;
[0207] For example, in Keras, the code described above is implemented as follows:
[0208] m=0.6
[0209] theta=lambda t:(K.sign(t)+1.) / 2.
[0210] Loss=-(1-theta(y_true-m)*theta(y_pred-m)-theta(1-m-y_true)*theta(1-m-y_pred))*(y_true*K.log(y_pred+1e-8)+(1y_true)*K.log(1-y_pred+1e-8))
[0211] S402. Construct an optimization function: Use the Adam algorithm as the optimization function of the model; all hyperparameters are set to the default values in Keras;
[0212] For example, the optimization function and its settings described above are expressed in Keras using the following code:
[0213] optim=Keras.optimizers.Aadm()
[0214] When the method model has not been fully trained, it needs to be trained on the training dataset to optimize the model parameters; when the model training is completed, the model predicts the matching degree of the input text pair.
[0215] The proposed model can achieve excellent results on the medical question answering dataset.
[0216] Example 3:
[0217] The text matching device based on pinyin modeling for medical question answering based on Example 2 includes a text semantic matching knowledge base construction unit, a text semantic matching model training data set generation unit, a text semantic matching model construction unit, and a text semantic matching model training unit, which respectively implement the functions of steps S1, S2, S3, and S4 in the text matching method based on pinyin modeling for medical question answering. The specific functions of each unit are as follows:
[0218] The text semantic matching knowledge base construction unit is used to obtain a large amount of text pair data and then pre-process the data to obtain a text semantic matching knowledge base that meets the training requirements;
[0219] The text semantic matching model training dataset generation unit generates a text pair in the semantic matching knowledge base. If the text pairs have the same semantics, the text pairs are used to construct positive training examples; otherwise, they are used to construct negative training examples. A large amount of positive and negative data are mixed to obtain a training dataset.
[0220] A text semantic matching model construction unit is used to construct a word mapping conversion table, a pinyin mapping conversion table, an input encoding module, a word vector mapping layer, a pinyin vector mapping layer, a text semantic encoding layer, a multi-correlation feature interaction layer, a three-dimensional convolutional feature encoding layer, and a prediction layer;
[0221] The text semantic matching model training unit is used to construct the loss function and optimization function required in the model training process and complete the model training.
[0222] Example 4:
[0223] The storage medium based on Example 2 stores multiple instructions, which are loaded by a processor to execute the steps of the text matching method based on pinyin modeling for medical questions and answers in Example 2.
[0224] Example 5:
[0225] Based on the electronic device of embodiment 4, the electronic device includes: the storage medium of embodiment 4; and a processor for executing instructions in the storage medium of embodiment 4.
[0226] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features therein may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text matching method based on pinyin modeling for medical question answering, characterized by: The method comprises the following steps: S1. Build a text semantic matching knowledge base: Collect frequently asked medical question and answer texts from the Internet and pre-process the texts to build a text semantic matching knowledge base; S2. Constructing a training dataset for a text semantic matching model: For the text pairs obtained in step S1, if they are semantically consistent, the text pairs are used to construct training positive examples. On the contrary, it is used to construct training negative examples; a large amount of positive and negative data are mixed to obtain a training data set; S3. Build a text semantic matching model: Build a text semantic matching model based on pinyin information; S4, training the text semantic matching model: training the text semantic matching model constructed in step S3 on the text semantic matching model training dataset obtained in step S2; The specific steps for building a text semantic matching model in S3 are as follows: S301, constructing a word mapping conversion table; S302, constructing a pinyin mapping conversion table; S303, constructing an input module; S304, constructing a word vector mapping layer; S305, constructing a pinyin vector mapping layer; S306, constructing a text semantic encoding layer: This layer encodes the character embedding representation and word embedding representation output by the word vector mapping layer and the pinyin embedding representation output by the pinyin vector mapping layer to obtain a character representation of the text, a word representation of the text, and a pinyin representation of the text; S307. Construct a multi-correlation feature interaction layer: This layer takes the character representation of the text, the word representation of the text, and the pinyin representation of the text as input; constructs a dot product correlation matrix, an absolute value subtraction correlation matrix, and a multiplication correlation matrix using dot product, absolute value subtraction, and multiplication operations respectively; and then constructs a correlation feature map based on the correlation matrix; S308, constructing a three-dimensional convolutional feature encoding layer: stacking the correlation feature maps to construct a correlation cube, and further encoding the correlation cube using a three-dimensional convolutional neural network to generate a matching representation of the input text pair; S309, build prediction layer: use Flatten method to flatten the matching representation of input text pairs, and then encode it through a three-layer fully connected neural network to predict the matching degree y of the input text pairs pred , and compare it with the preset threshold. If it is greater than or equal to the threshold, the text is considered to be semantically matched, otherwise it is not matched.
2. The text matching method based on pinyin modeling for medical question answering according to claim 1 is characterized in that: The specific steps of constructing the text semantic matching model in step S3 are as follows: S301, constructing a word mapping conversion table: the mapping rule is to start with the number 1, and then sort each word in the order in which it is entered into the word table, thereby forming the required word mapping conversion table; Afterwards, Word2Vec is used to train the word vector model to obtain the word vector matrix of each word; S302, constructing a pinyin mapping conversion table: starting with the number 1, and then sorting each pinyin in the order in which it is entered into the pinyin table, thereby forming the required pinyin mapping conversion table; Afterwards, Word2Vec is used to train the pinyin vector model to obtain the pinyin vector matrix of each pinyin; S303, constructing an input module: The input module includes six inputs; for each text in the training data set or the text to be predicted, preprocess it using the corresponding modules in S1 and S2, and obtain the character sequence txtP_char of text P, the character sequence txtQ_char of text Q, the word sequence txtP_word of text P, the word sequence txtQ_word of text Q, the pinyin sequence txtP_pinyin of text P, and the pinyin sequence txtQ_pinyin of text Q respectively; for each character, word, and pinyin in the input text, convert it into a corresponding digital identifier according to the word mapping conversion table and the pinyin mapping conversion table; S304, constructing a word vector mapping layer: initializing the weight parameters of the current layer by loading the word vector matrix trained in the step of constructing a word mapping conversion table; for the input texts txt P_char, txt Q_char and txt P_word, txt Q_word, obtaining their corresponding vectors txt P_char_embed, txt Q_char_embed, txt P_word_embed, txt Q_word_embed; for each text in the text semantic matching knowledge base, the text word information is converted into a vector form, i.e., a character embedding representation and a word embedding representation, by means of word vector mapping; S305. Construct a pinyin vector mapping layer: initialize the weight parameters of the current layer by loading the pinyin vector matrix trained in the step of constructing the pinyin mapping conversion table; for the input texts txt P_pinyin and txt Q_pinyin, obtain their corresponding vectors txt P_pinyin_embed and txt Q_pinyin_embed; each text in the text semantic matching knowledge base is converted into a vector form, i.e., a pinyin embedding representation, by means of pinyin vector mapping.
3. The text matching method based on pinyin modeling for medical question answering according to claim 2 is characterized in that: The text semantic coding layer construction process in step S306 is as follows: S306, constructing a text semantic encoding layer: This layer receives the character embedding representation, word embedding representation output by the word vector mapping layer, and the pinyin embedding representation output by the pinyin vector mapping layer as input; first, using a bidirectional long short-term memory network BiLSTM to encode the character embedding representation and word embedding representation respectively to generate a character context representation and a word context representation; at the same time, the pinyin embedding representation is encoded using a fully connected neural network and then passed to the BiLSTM to generate a pinyin context representation; further using a hierarchical dilated convolutional network encoding to obtain a character text representation after dilated convolution, a word text representation after dilated convolution, and a pinyin text representation after dilated convolution; and connecting them with the corresponding context representation to generate a character representation of the text, a word representation of the text, and a pinyin representation of the text; S30601, BiLSTM feature extraction: The text semantic encoding layer first uses BiLSTM to encode the character embedding representation and word embedding representation of the two texts to generate character context representation and word context representation; For the pinyin embedding representation, it is first encoded using a fully connected neural network and then passed into BiLSTM to generate the pinyin context representation; For text P, the specific operation is described by the following formula: Where N is the length of the text; represents the character embedding representation of the text P at the i-th position under the character granularity c; Represents the word embedding representation of the text P at the jth position under the word granularity w; represents the pinyin embedding representation of text P at the mth position under pinyin granularity p; Dense represents the fully connected neural network encoding; and Represent the character context representation of text P, the word context representation of text P and the pinyin context representation of text P respectively; for text Q, the character context representation Q of text Q is obtained respectively c' , the word context representation Q of text Q w' , the pinyin context representation of text Q p' ; S30602, Hierarchical Dilated Convolutional Network Coding: Encode the character context representation, word context representation, and pinyin context representation of two texts using a hierarchical dilated convolutional network to generate a dilated convolutional character text representation, a dilated convolutional word text representation, and a dilated convolutional pinyin text representation of the corresponding texts; The hierarchical dilated convolutional network consists of three layers of dilated convolutional networks with dilation rates of 1, 2, and 3, respectively. The character / word / pinyin context representation is passed through the three layers of dilated convolutional networks to obtain the character / word / pinyin representation after the first convolution layer, the character / word / pinyin representation after the second convolution layer, and the character / word / pinyin representation after the third convolution layer. These representations are then concatenated according to the character, word, and pinyin to generate the dilated convolution text representation of the character / word / pinyin. The calculation formula is as follows: Among them, Dilated_CNN represents a one-dimensional dilated convolutional neural network; dilation_rate is the dilation rate; Represents the character representation of the text P after the first layer of convolution; Represents the character representation of the text P after the second convolution layer; Represents the character representation of the text P after the third convolution layer; Represents the character text representation of the text P after dilation convolution; [;] represents the connection operation; w' 、P p' , Q c' , Q w' , Q p' Processing, get the word text representation after the text P expansion convolution Pinyin text representation of text P after dilated convolution Character text representation after text Q dilation convolution Text Q word text representation after dilated convolution Pinyin text representation after text Q dilation convolution S30603, Feature Connection: For text P, the text semantic encoding layer uses the connection operation to respectively connect P c' and P w' and P p' and Perform the join to generate the character representation of text P, the word representation of text P, and the pinyin representation of text P. The specific formula is as follows: in, Represents the character representation of the text P; The word representation representing the text P; Represents the pinyin representation of text P; for text Q, obtain the character representation of text Q Word representation of text Q and the pinyin representation of the text Q 4. The text matching method based on pinyin modeling for medical question answering according to claim 2 is characterized in that: The process of constructing the multi-correlation feature interaction layer in step S307 is as follows: S307, constructing a multi-correlation feature interaction layer: This layer receives the character representation of the text P in step S306 Word representation of text P Pinyin representation of text P Character representation of text Q Word representation of text Q and the pinyin representation of the text Q As input; construct the point product correlation matrix, absolute value subtraction correlation matrix, and multiplication correlation matrix by point multiplication, absolute value subtraction, and multiplication operations respectively; then construct the correlation feature map according to the correlation matrix; S30701. Construct a correlation matrix: Take the character representation of text P, the word representation of text P, the pinyin representation of text P, the character representation of text Q, the word representation of text Q, and the pinyin representation of text Q as input, perform dot product of the character representation of text P with the character representation of text Q, the pinyin representation of text Q, and the word representation of text Q, respectively, to obtain a dot product correlation matrix of the characters of text P and the characters / pinyin / words of text Q; perform dot product of the pinyin representation of text P with the character representation of text Q, the pinyin representation of text Q, and the word representation of text Q, respectively, to obtain a dot product correlation matrix of the pinyin of text P and the characters / pinyin / words of text Q; Perform dot product of the word representation of text P with the character representation of text Q and the pinyin representation of text Q to obtain the dot product correlation matrix of the words of text P and the characters / pinyin of text Q. The specific formula is as follows: in, The subscript a in the matrix indicates that the matrix is calculated by dot product, and the subscript ij indicates the dot product correlation matrix of text P characters and text Q characters. The value at row i and column j in ; The superscripts c and c represent that the matrix is generated according to the character representation of text P and the character representation of text Q respectively; Respectively represent The value at the i-th row and j-th column in the matrix; the superscripts c and p represent that the matrix is generated according to the character representation of text P and the pinyin representation of text Q, and c and w represent that the matrix is generated according to the character representation of text P and the word representation of text Q; Represent the dot product correlation matrix of text P characters and text Q pinyin / words respectively; Character representation of text P The character representation at position i in ; Character representation of the text Q The character representation at the j-th position in ; Represents the pinyin representation of the text Q The pinyin representation at the j-th position in ; Word representation representing text Q The word representation at the jth position in the text is obtained in sequence, and the dot product correlation matrix of the text P pinyin and the text Q characters is obtained in sequence The dot product correlation matrix of text P pinyin and text Q pinyin The dot product correlation matrix of text P pinyin and text Q words The dot product correlation matrix of text P words and text Q characters The dot product correlation matrix of words in text P and pinyin in text Q The specific formulas for constructing the multiplication correlation matrix and the absolute value subtraction correlation matrix are as follows: Among them, ⊙ represents element-wise dot product; |-| represents absolute value subtraction; Dense represents fully connected neural network encoding; TimeDistributed represents the use of fully connected neural network encoding at each time step; tanh represents the tanh function; and The superscript meaning is the same as The subscripts mul and sub respectively represent that the correlation matrix is obtained by multiplication and absolute value subtraction; different inputs are encoded to obtain the multiplication correlation matrix of text P characters and text Q pinyin in turn Multiplication correlation matrix of text P characters and text Q words Multiplication correlation matrix of text P pinyin and text Q characters Multiplication correlation matrix of text P pinyin and text Q pinyin Multiplication correlation matrix of text P pinyin and text Q words Multiplication correlation matrix of text P words and text Q characters Multiplication correlation matrix of words in text P and pinyin in text Q Absolute value subtraction correlation matrix of text P characters and text Q pinyin Absolute value subtraction correlation matrix of text P characters and text Q words Correlation matrix of text P pinyin and text Q character absolute value subtraction Absolute value subtraction correlation matrix of text P pinyin and text Q pinyin Correlation matrix of text P pinyin and text Q word absolute value subtraction Correlation matrix of absolute value subtraction between words in text P and characters in text Q Absolute value subtraction correlation matrix of words in text P and pinyin in text Q S30702. Constructing a correlation feature map: The multi-correlation feature interaction layer stacks the correlation matrices obtained in step S30701 to construct a correlation feature map. The specific formula is as follows: The dot product correlation matrix of text P characters and text Q characters / pinyin, the dot product correlation matrix of text P pinyin and text Q characters / pinyin, the dot product correlation matrix of text P words and text Q characters / pinyin, the dot product correlation matrix of text P characters and text Q words, and the dot product correlation matrix of text P pinyin and text Q words are stacked respectively to construct four dot product-based correlation feature maps, namely A'1, A'2, A'3, and A'4; for the multiplication correlation matrix and the absolute value subtraction correlation matrix, four multiplication-based correlation feature maps, namely M'1, M'2, M'3, and M'4, and four absolute value subtraction-based correlation feature maps, namely U'1, U'2, U'3, and U'4, are formed respectively.
5. The text matching method based on pinyin modeling for medical question answering according to claim 4 is characterized in that: The process of constructing the three-dimensional convolutional feature coding layer in step S308 is as follows: S308: Construct a three-dimensional convolutional feature encoding layer: further stack the correlation feature map based on dot multiplication, the correlation feature map based on multiplication, and the correlation feature map based on absolute value subtraction obtained in step S307 to construct a correlation cube; then encode the correlation cube through a three-dimensional convolutional neural network to generate a matching representation of the input text pair; S30801. Constructing a correlation cube: The three-dimensional convolutional feature encoding layer first stacks the correlation feature map based on dot multiplication, the correlation feature map based on multiplication, and the correlation feature map based on absolute value subtraction obtained in step S307 to construct a correlation cube. The specific formula is as follows: Where F represents the correlation cube, stack represents the stacking operation, and the meanings of other symbols are the same as those in formula (6); S30802, 3D Convolutional Neural Network Encoding: Use a 3D convolutional neural network to encode the relevance cube to generate a matching representation of the input text pair. The specific formula is as follows: F′=Conv3d(F) (8) Among them, F′ represents the matching representation of the input text pair, and Conv3d represents the three-dimensional convolutional neural network.
6. The text matching method for medical question and answer based on pinyin modeling according to claim 1 or 2, characterized in that: The specific steps of step S4: training the text semantic matching model are as follows: S401, constructing a loss function: using the modified binary cross entropy as the loss function based on the matching degree of the input text pair obtained in step S309, the formula is as follows: Where θ(x) is the unit step function, thre is the threshold, which is set to 0.6, and L represents the modified cross entropy formula; y pred Represents the matching degree of the predicted input text pair; y true The true label representing whether the input text pair matches; S402. Construct an optimization function: Use the Adam algorithm as the optimization function of the model; all hyperparameters are set to the default values in Keras; When the method model has not been fully trained, it needs to be trained on the training dataset to optimize the model parameters; when the model training is completed, the model predicts the matching degree of the input text pair.
7. A text matching device based on phonetic modeling for medical question answering, which implements the text matching method based on phonetic modeling for medical question answering according to any one of claims 1 to 6, characterized in that: The device comprises, The text semantic matching knowledge base construction unit is used to obtain a large amount of text pair data and then pre-process the data to obtain a text semantic matching knowledge base that meets the training requirements; The text semantic matching model training data set generation unit is used to construct training positive examples for text pairs in the text pair semantic matching knowledge base if their semantics are consistent. On the contrary, it is used to construct training negative examples; a large amount of positive and negative data are mixed to obtain a training data set; A text semantic matching model construction unit is used to construct a word mapping conversion table, a pinyin mapping conversion table, an input encoding module, a word vector mapping layer, a pinyin vector mapping layer, a text semantic encoding layer, a multi-correlation feature interaction layer, a three-dimensional convolutional feature encoding layer, and a prediction layer; The text semantic matching model training unit is used to construct the loss function and optimization function required in the model training process and complete the model training.
8. A storage medium storing a plurality of instructions, characterized in that: The instructions are loaded by a processor to execute the text matching method based on pinyin modeling for medical question answering according to any one of claims 1-6.
9. An electronic device, characterized in that: The electronic device comprises: the storage medium according to claim 8; and a processor for executing instructions in the storage medium.
Citation Information
Patent Citations
Text semantic matching method and device for intelligent questions and answers of fire safety knowledge
CN114547256A
Method and device of matching speech input to text
US20140379335A1