A method and device for retrieving similar sentence pairs
By segmenting and translating words in the encoder-decoder framework of the NMT model, the problem of contextual relationship utilization and mapping relationship generation in the existing deep structured semantic model when calculating sentence similarity is solved, and a more accurate search of similar sentence pairs is achieved.
Patent Information
- Application Number
- CN201910176655.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-03-08
AI Technical Summary
The existing deep structured semantic models are difficult to utilize the contextual relationship of the language model when calculating sentence pair similarity, and cannot generate the mapping relationship between words and words between sentence pairs.
Using the encoder-decoder framework based on the NMT model, the query sentence pairs are segmented through the encoder encoder to generate eigenvectors, and the decoder decoder is used to translate the eigenvectors based on the word segmentation feature vectors to generate eigenvectors of similar sentence pairs.
Effectively combining the context semantics of the query sentence pairs, similar sentence pairs that are more similar to the query sentence pairs are generated, overcoming the limitations of the existing discriminant model.
Smart Images

Figure CN111666299B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for retrieving similar sentence pairs. Background Art
[0002] In the prior art, in the process of sentence retrieval, the problem of calculating the similarity between sentence pairs is usually involved. Currently, there are many ways to calculate the similarity between similar sentence pairs, such as using deep structured semantic models (DSSM) to calculate the dot product between text features to obtain the similarity score of sentence pairs.
[0003] However, existing deep structured semantic models are essentially discriminative models, which cannot make good use of the contextual relationships of the language model and cannot generate word-to-word mapping relationships between sentence pairs. Summary of the invention
[0004] In view of the above problems, the present invention is proposed to provide a method and device for retrieving similar sentence pairs that overcome the above problems or at least partially solve the above problems.
[0005] According to one aspect of the present invention, a method for retrieving similar sentence pairs is provided, comprising:
[0006] Obtaining a query sentence pair to be retrieved, and inputting the query sentence pair into an encoder of a preset NMT model;
[0007] Using an encoder to segment the query sentence pair, and generating a feature vector corresponding to each segmentation;
[0008] Inputting the feature vector of each word segmentation into the decoder of the NMT model in sequence according to a preset time interval;
[0009] A decoder is used to translate the query sentence pair according to the feature vectors of each word segment received in sequence, to obtain feature vectors of similar sentence pairs of the query sentence pair, and to decode the feature vectors of the similar sentence pairs into corresponding similar sentence pairs.
[0010] Optionally, a decoder is used to translate the query sentence pair according to the feature vectors of each word segment received in sequence, including:
[0011] After the feature vector of the word segmentation of the query sentence pair is first input into the decoder, the decoder is used to translate the received feature vector of the word segmentation into similar sentence pairs to obtain a translation result;
[0012] Selecting a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results according to a preset algorithm;
[0013] The feature vector of the similar sentence pair selected this time and the feature vector of the next word segment input after the preset time are used as the next input of the decoder decoder, the decoder decoder is used to translate the received input content to obtain a translation result, and a specified number of feature vectors of similar sentence pairs that meet the preset conditions are selected from the translation result according to a preset algorithm, and this cycle is repeated until the translation of the feature vector of the last word segment of the query text is completed.
[0014] Optionally, the feature vectors of similar sentence pairs that meet preset conditions include:
[0015] The feature vectors of similar sentence pairs whose correlation scores with the feature vector input to the decoder are greater than the specified score.
[0016] Optionally, the preset algorithm includes a beam search algorithm.
[0017] Optionally, before the query sentence pair is input into an encoder of a preset NMT model, the method further includes:
[0018] Obtain a candidate set containing multiple candidate sentence pairs;
[0019] A dictionary tree is constructed for the multiple candidate sentence pairs in the candidate set.
[0020] Optionally, selecting a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results according to a preset algorithm includes:
[0021] Based on the constructed dictionary tree and according to the beam search algorithm, a specified number of feature vectors of similar sentence pairs that meet preset conditions are selected from the translation results.
[0022] Optionally, before the query sentence pair is input into an encoder of a preset NMT model, the method further includes:
[0023] Collect multiple candidate sentence pairs;
[0024] Select multiple groups of sentence pairs with similar semantics from the collected multiple candidate sentence pairs, and mark the corresponding similar sentence pairs for each group of sentence pairs with similar semantics;
[0025] The NMT model is trained based on annotated sentence pairs with similar semantics.
[0026] According to another aspect of the present invention, there is also provided a similar sentence pair retrieval device, comprising:
[0027] A first acquisition module is adapted to acquire a query sentence pair to be retrieved, and input the query sentence pair into an encoder of a preset NMT model;
[0028] A word segmentation module, adapted to segment the query sentence pair using an encoder to generate a feature vector corresponding to each word segmentation;
[0029] An input module, adapted to sequentially input the feature vectors of each word segment into a decoder of the NMT model at preset time intervals;
[0030] The translation module is adapted to use a decoder to translate the query sentence pair according to the feature vectors of each word segment received in sequence, obtain the feature vectors of similar sentence pairs of the query sentence pair, and decode the feature vectors of the similar sentence pairs into corresponding similar sentence pairs.
[0031] Optionally, the translation module is further adapted to:
[0032] After the feature vector of the word segmentation of the query sentence pair is first input into the decoder, the decoder is used to translate the received feature vector of the word segmentation into similar sentence pairs to obtain a translation result;
[0033] Selecting a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results according to a preset algorithm;
[0034] The feature vector of the similar sentence pair selected this time and the feature vector of the next word segment input after the preset time are used as the next input of the decoder decoder, the decoder decoder is used to translate the received input content to obtain a translation result, and a specified number of feature vectors of similar sentence pairs that meet the preset conditions are selected from the translation result according to a preset algorithm, and this cycle is repeated until the translation of the feature vector of the last word segment of the query text is completed.
[0035] Optionally, the feature vectors of similar sentence pairs that meet preset conditions include:
[0036] The feature vectors of similar sentence pairs whose correlation scores with the feature vector input to the decoder are greater than the specified score.
[0037] Optionally, the preset algorithm includes a beam search algorithm.
[0038] Optionally, it also includes:
[0039] A second acquisition module, adapted to acquire a candidate set including a plurality of candidate sentence pairs before the first acquisition module inputs the query sentence pair into an encoder of a preset NMT model;
[0040] The construction module is adapted to construct a dictionary tree for the plurality of candidate sentence pairs in the candidate set.
[0041] Optionally, the translation module is further adapted to:
[0042] Based on the constructed dictionary tree and according to the beam search algorithm, a specified number of feature vectors of similar sentence pairs that meet preset conditions are selected from the translation results.
[0043] Optionally, it also includes:
[0044] A collecting module, adapted for the first acquisition module to collect a plurality of candidate sentence pairs before inputting the query sentence pair into an encoder of a preset NMT model;
[0045] The tagging module is adapted to select a plurality of sentence pairs having similar semantics from the collected plurality of candidate sentence pairs, and tag the corresponding similar sentence pairs for each group of sentence pairs having similar semantics;
[0046] The training module is adapted to train the NMT model based on annotated sentence pairs with similar semantics.
[0047] According to another aspect of the present invention, a computer storage medium is provided, wherein the computer storage medium stores computer program code, and when the computer program code is executed on a computing device, the computing device is caused to execute the method for retrieving similar sentence pairs described in any of the above embodiments.
[0048] According to another aspect of the present invention, a computing device is provided, comprising: a processor; a memory storing computer program code; when the computer program code is executed by the processor, the computing device executes the similar sentence pair retrieval method described in any of the above embodiments.
[0049] In an embodiment of the present invention, after the query sentence pair to be retrieved is obtained, the query sentence pair is input into the encoder of a preset NMT (Neural Machine Translation) model, and the encoder is used to segment the query sentence to generate a feature vector corresponding to each segmentation. Then, the feature vector of each segmentation is sequentially input into the decoder of the NMT model at preset time intervals, and the decoder is used to translate the query sentence pair according to the feature vectors of each segmentation received sequentially to obtain feature vectors of similar sentence pairs of the query sentence pair. Finally, the feature vectors of the similar sentence pairs are decoded into corresponding similar sentence pairs. Therefore, the embodiment of the present invention uses the encoder of the NMT model to segment the query sentence, and then uses the decoder to translate the query sentence pair according to the respectively input segmented words. Since each segmented word is respectively input to the decoder at a preset time interval, the segmented word feature vector input this time and the previous translation result can be combined for translation during the translation process. Different from the existing discriminant model that directly calculates whether two feature vectors are similar, the present invention uses the NMT model to translate similar sentence pairs in combination with the contextual semantics of the query sentence pair, thereby effectively obtaining similar sentence pairs that are more semantically similar to the query sentence pair.
[0050] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below.
[0051] Based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0053] Figure 1 A schematic diagram showing a flow chart of a method for retrieving similar sentence pairs according to an embodiment of the present invention;
[0054] Figure 2 A schematic diagram showing a similar sentence pair retrieval process according to an embodiment of the present invention;
[0055] Figure 3A schematic diagram showing the structure of a similar sentence pair retrieval device according to an embodiment of the present invention;
[0056] Figure 4 A schematic structural diagram of a similar sentence pair retrieval device according to another embodiment of the present invention is shown. DETAILED DESCRIPTION
[0057] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0058] In order to solve the above technical problem, an embodiment of the present invention provides a method for retrieving similar sentence pairs. Figure 1 FIG. 1 is a flow chart showing a method for retrieving similar sentence pairs according to an embodiment of the present invention. Figure 1 The method at least includes steps S102 to S108.
[0059] Step S102: obtaining a query pair to be retrieved, and inputting the query pair into an encoder of a preset NMT model.
[0060] In this embodiment, the NMT model is based on an encoder-decoder framework.
[0061] Step S104: Use an encoder to segment the query sentence pair to generate a feature vector corresponding to each segmentation.
[0062] Step S106: Input the feature vector of each word segment into the decoder of the NMT model in sequence according to a preset time interval.
[0063] In this step, the feature vectors of each word segment are sequentially input into the decoder, which means that the feature vectors of each word segment are sequentially input into the decoder according to the order of the positions of each word segment in the query sentence pair.
[0064] Step S108, using a decoder to translate the query sentence pair according to the feature vectors of each word segment received in sequence, obtain feature vectors of similar sentence pairs of the query sentence pair, and decode the feature vectors of the similar sentence pairs into corresponding similar sentence pairs.
[0065] The embodiment of the present invention uses an encoder of an NMT model to segment the query sentence, and then uses a decoder to translate the query sentence pair according to the respectively input segmented words. Since each segmented word is respectively input to the decoder at a preset time interval, the segmented word feature vector input this time and the previous translation result can be combined for translation during the translation process. Compared with the existing discriminant model that directly calculates whether two feature vectors are similar, the present invention uses the NMT model to combine the contextual semantics of the query sentence pair to translate similar sentence pairs, thereby effectively obtaining similar sentence pairs that are more semantically similar to the query sentence pair.
[0066] Referring to step S108 above, in one embodiment of the present invention, the specific execution steps of using a decoder to translate the query sentence pair according to the feature vectors of each word segment received in sequence are as follows:
[0067] Step 1: After the feature vector of the word segment of the query sentence pair is first input into the decoder, the decoder is used to translate the feature vector of the received word segment into similar sentence pairs to obtain the translation result. Since the feature vector of each word segment of the query sentence pair is input into the decoder in sequence, the word segment first input into the decoder is actually the first word segment of the query sentence pair.
[0068] Step 2: Select a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results according to a preset algorithm.
[0069] In this step, the feature vectors of similar sentence pairs that meet the preset conditions can be feature vectors of similar sentence pairs whose correlation scores with the feature vectors input to the decoder are greater than the specified score. The purpose is to select feature vectors of similar sentence pairs with greater correlation with the feature vectors input to the decoder from the translation results. This embodiment does not make any specific limitation on the specified score.
[0070] Step 3: Use the feature vector of the similar sentence pair selected this time and the feature vector of the next word segmentation input after a preset time as the next input of the decoder, and use the decoder to translate the received input content to obtain a translation result.
[0071] Steps 2 to 3 are repeatedly executed in a loop until the feature vector translation of the last word segment of the query text is completed.
[0072] In this embodiment, in order to facilitate the translation of the feature vectors of each word segment, a recurrent neural network, such as a recurrent neural network RNN (Recurrent Neural Network), can be defined in the NMT model, so that each step in this embodiment is executed through the recurrent neural network.
[0073] In one embodiment of the present invention, the preset algorithm in the above embodiment can adopt a preset beam search algorithm. By selecting a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results through the beam search algorithm, the recall space of similar sentence pairs can be continuously expanded. Of course, other algorithms can also be used, which will not be specifically introduced in the embodiment of the present invention.
[0074] In order to more clearly illustrate the above embodiment, a specific example is now used to introduce the retrieval process of similar sentence pairs.
[0075] See also Figure 2 For example, the query sentence pair is "3-day tour in Beijing". After the decoder decodes the word segmentation of "3-day tour in Beijing", "Beijing", "3 days", and "tour" are obtained. The feature vectors corresponding to "Beijing", "3 days", and "tour" are input into the decoder in sequence according to the preset time interval.
[0076] First, the feature vector of "Beijing" is input into the decoder, and the decoder translates the feature vector of "Beijing" into similar sentence pairs to obtain multiple translation results, such as "Tianjin", "Shanghai", "Beijing", etc. In this step, if the feature vectors of similar sentence pairs that meet the preset conditions in the translation results do not reach the specified number, then the feature vectors of similar sentence pairs that meet the preset conditions can be directly selected. For example, if the beam search algorithm calculates that only "Beijing" meets the preset conditions, and the set specified number of selections is 3, the feature vector corresponding to "Beijing" can be directly selected.
[0077] Then, the feature vectors of "Beijing" and "3 days" are input into the decoder, and the decoder translates similar sentence pairs of the feature vectors of "Beijing" and "3 days" to obtain multiple translation results, such as "Beijing 1 day", "Beijing 3 days", "Beijing travel", etc. The beam search algorithm is used to select three feature vectors that meet the preset conditions from the translation results, which are the feature vectors corresponding to "Beijing 1 day", "Beijing 3 days", and "Beijing travel".
[0078] Finally, the feature vectors of "Beijing 1 day", "Beijing 3 days", "Beijing travel" and the word "travel" are input into the decoder. The decoder translates similar sentence pairs based on the input content to obtain multiple translation results, such as "Beijing 1 day tour", "Beijing 3 day tour", "Beijing travel guide", etc. The beam search algorithm is used to select three feature vectors that meet the preset conditions from the translation results, which are the feature vectors corresponding to "Beijing 1 day tour", "Beijing 3 day tour", and "Beijing travel guide". The final translation result is the final search result, that is, the similar sentence pairs retrieved that are similar to the query sentence pair.
[0079] In one embodiment of the present invention, before inputting the query sentence pair into the encoder of the preset NMT model, a candidate set including multiple candidate sentence pairs may be obtained first, and a dictionary tree may be constructed for the multiple candidate sentence pairs in the candidate set.
[0080] In the process of selecting a specified number of feature vectors of similar sentence pairs that meet the preset conditions from the translation results according to the preset algorithm, a specified number of feature vectors of similar sentence pairs that meet the preset conditions can also be selected from the translation results based on the constructed dictionary tree and according to the beam search algorithm. That is, the dictionary tree can be used to continuously prune the beam search translation results that do not meet the preset conditions, and finally obtain the most similar specified number of similar sentence pairs as the search results.
[0081] In one embodiment of the present invention, in order to enable the NMT model to retrieve similar sentence pairs for a query sentence pair more accurately and quickly, the NMT model based on the encoder-decoder framework is trained before the query sentence pair is input into the encoder of the preset NMT model. The training method adopted in this embodiment is to train the NMT model through real annotated similar semantic sentence pairs.
[0082] Specifically, first, multiple candidate sentence pairs are collected. Then, multiple groups of sentence pairs with similar semantics are selected from the collected multiple candidate sentence pairs, and corresponding similar sentence pairs are annotated for each group of sentence pairs with similar semantics. Finally, the NMT model is trained based on the annotated sentence pairs with similar semantics. The trained NMT model can more efficiently retrieve similar sentence pairs for query sentence pairs.
[0083] Based on the same inventive concept, an embodiment of the present invention also provides a similar sentence pair retrieval device. Figure 3 FIG. 4 is a schematic diagram showing a structure of a similar sentence pair retrieval device according to an embodiment of the present invention. Figure 3The similar sentence pair retrieval device 300 includes a first acquisition module 310 , a word segmentation module 320 , an input module 330 , and a translation module 340 .
[0084] The functions of the components or devices of the similar sentence pair retrieval device 300 according to the embodiment of the present invention and the connection relationship between the components are now introduced:
[0085] A first acquisition module 310 is adapted to acquire a query sentence pair to be retrieved, and input the query sentence pair into an encoder of a preset NMT model;
[0086] The word segmentation module 320 is coupled to the first acquisition module 310 and is adapted to segment the query sentence pair using an encoder to generate a feature vector corresponding to each word segmentation;
[0087] An input module 330, coupled to the word segmentation module 320, is adapted to sequentially input the feature vectors of each word segmentation into a decoder of the NMT model at a preset time interval;
[0088] The translation module 340 is coupled to the input module 330 and is adapted to use a decoder to translate the query sentence pair according to the feature vectors of each word segment received in sequence, obtain feature vectors of similar sentence pairs of the query sentence pair, and decode the feature vectors of the similar sentence pairs into corresponding similar sentence pairs.
[0089] In one embodiment of the present invention, the translation module 340 is also suitable for first inputting the feature vector of the word segment of the query sentence pair into the decoder decoder for the first time, and then using the decoder decoder to translate the feature vector of the received word segment into similar sentence pairs to obtain a translation result. Then, according to the preset algorithm, a specified number of feature vectors of similar sentence pairs that meet the preset conditions are selected from the translation results. Furthermore, the feature vector of the similar sentence pair selected this time and the feature vector of the next word segment input after the preset time are used as the next input of the decoder decoder, and the decoder decoder is used to translate the received input content to obtain a translation result, and according to the preset algorithm, a specified number of feature vectors of similar sentence pairs that meet the preset conditions are selected from the translation results, and this cycle is repeated until the translation of the feature vector of the last word segment of the query text is completed.
[0090] In one embodiment of the present invention, the feature vectors of similar sentence pairs that meet the preset conditions include feature vectors of similar sentence pairs whose correlation scores with the feature vectors input to the decoder are greater than a specified score. The specified score is not specifically limited here.
[0091] In an embodiment of the present invention, the preset algorithm introduced above may adopt a beam search algorithm, and the embodiment of the present invention does not specifically limit the preset algorithm.
[0092] The embodiment of the present invention also provides a similar sentence pair retrieval device. Figure 4 FIG. 4 is a schematic diagram showing a structure of a similar sentence pair retrieval device according to an embodiment of the present invention. Figure 4 In addition to the first acquisition module 310, the word segmentation module 320, the input module 330, and the translation module 340, the similar sentence pair retrieval device 300 also includes a second acquisition module 350, a construction module 360, a collection module 370, a labeling module 380, and a training module 390. For the introduction of the first acquisition module 310, the word segmentation module 320, the input module 330, and the translation module 340, please refer to the above embodiment, which will not be repeated here.
[0093] The second acquisition module 350 is coupled to the first acquisition module 310 and is adapted to acquire a candidate set including a plurality of candidate sentence pairs before the first acquisition module 310 inputs the query sentence pair into an encoder of a preset NMT model.
[0094] The construction module 360 is coupled to the second acquisition module 350 and is adapted to construct a dictionary tree for multiple candidate sentence pairs in the candidate set.
[0095] The collecting module 370 is adapted for the first acquiring module 310 to collect a plurality of candidate sentence pairs before inputting the query sentence pairs into the encoder of the preset NMT model.
[0096] The marking module 380 is coupled with the collecting module 370 and is adapted to select a plurality of groups of sentence pairs having similar semantics from the collected plurality of candidate sentence pairs, and mark the corresponding similar sentence pairs for each group of sentence pairs having similar semantics.
[0097] The training module 390 is coupled to the annotation module 380 and the first acquisition module 310 respectively, and is suitable for training the NMT model based on the annotated sentence pairs with similar semantics.
[0098] In one embodiment of the present invention, the translation module 340 is further adapted to select a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results based on the constructed dictionary tree and according to the beam search algorithm.
[0099] According to another aspect of the present invention, a computer storage medium is provided, which stores computer program code. When the computer program code runs on a computing device, the computing device executes the method for retrieving similar sentence pairs in any of the above embodiments.
[0100] According to another aspect of the present invention, there is provided a computing device, comprising: a processor; a memory storing computer program code; when the computer program code is executed by the processor, the computing device executes the method for retrieving similar sentence pairs of any of the above embodiments.
[0101] According to any one of the above preferred embodiments or a combination of multiple preferred embodiments, the embodiments of the present invention can achieve the following beneficial effects:
[0102] In an embodiment of the present invention, after the query sentence pair to be retrieved is obtained, the query sentence pair is input into the encoder of the preset NMT model, and the encoder is used to segment the query sentence to generate a feature vector corresponding to each segmentation. Then, the feature vector of each segmentation is sequentially input into the decoder of the NMT model at preset time intervals, and the decoder is used to translate the query sentence pair according to the feature vectors of each segmentation received in sequence to obtain feature vectors of similar sentence pairs of the query sentence pair. Finally, the feature vectors of the similar sentence pairs are decoded into corresponding similar sentence pairs. Therefore, the embodiment of the present invention uses the encoder of the NMT model to segment the query sentence, and then uses the decoder to translate the query sentence pair according to the respectively input segmented words. Since each segmented word is respectively input to the decoder at a preset time interval, the segmented word feature vector input this time and the previous translation result can be combined for translation during the translation process. Different from the existing discriminant model that directly calculates whether two feature vectors are similar, the present invention uses the NMT model to translate similar sentence pairs in combination with the contextual semantics of the query sentence pair, thereby effectively obtaining similar sentence pairs that are more semantically similar to the query sentence pair.
[0103] Those skilled in the art can clearly understand that the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments, and for the sake of brevity, they are not further described here.
[0104] In addition, the functional units in various embodiments of the present invention may be physically independent of each other, or two or more functional units may be integrated together, or all functional units may be integrated into one processing unit. The above integrated functional units may be implemented in the form of hardware, or in the form of software or firmware.
[0105] Those skilled in the art can understand that if the integrated functional unit is implemented in the form of software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can essentially or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, which includes a number of instructions to enable a computing device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention when running the instructions. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store program codes.
[0106] Alternatively, all or part of the steps of implementing the aforementioned method embodiments may be accomplished by hardware associated with program instructions (such as a computing device such as a personal computer, a server, or a network device), and the program instructions may be stored in a computer-readable storage medium. When the program instructions are executed by a processor of a computing device, the computing device executes all or part of the steps of the methods described in the embodiments of the present invention.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that within the spirit and principles of the present invention, the technical solutions described in the aforementioned embodiments can still be modified, or some or all of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate from the protection scope of the present invention.
Claims
1. A method for retrieving similar sentence pairs, comprising: Obtaining a query sentence pair to be retrieved, and inputting the query sentence pair into an encoder of a preset NMT model; Using an encoder to segment the query sentence pair, and generating a feature vector corresponding to each segmentation; Inputting the feature vector of each word segmentation into the decoder of the NMT model in sequence according to a preset time interval; A decoder is used to translate the query sentence pair according to the feature vectors of each word segment received in sequence, to obtain feature vectors of similar sentence pairs of the query sentence pair, and to decode the feature vectors of the similar sentence pairs into corresponding similar sentence pairs; The decoder is used to translate the query sentence pair according to the feature vectors of each word segment received in sequence, including: After the feature vector of the word segmentation of the query sentence pair is first input into the decoder, the decoder is used to translate the received feature vector of the word segmentation into similar sentence pairs to obtain a translation result; Selecting a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results according to a preset algorithm; The feature vector of the similar sentence pair selected this time and the feature vector of the next word segment input after the preset time are used as the next input of the decoder decoder, the decoder decoder is used to translate the received input content to obtain a translation result, and a specified number of feature vectors of similar sentence pairs that meet the preset conditions are selected from the translation result according to a preset algorithm, and this cycle is repeated until the translation of the feature vector of the last word segment of the query sentence is completed.
2. The method according to claim 1, wherein: The feature vectors of similar sentence pairs that meet the preset conditions include: The feature vectors of similar sentence pairs whose correlation scores with the feature vector input to the decoder are greater than the specified score.
3. The method according to claim 2, wherein: The preset algorithm includes a beam search algorithm.
4. The method according to claim 3, wherein: Before inputting the query sentence pair into the encoder of the preset NMT model, the method further includes: Obtain a candidate set containing multiple candidate sentence pairs; A dictionary tree is constructed for the multiple candidate sentence pairs in the candidate set.
5. The method according to claim 4, wherein: Selecting a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results according to a preset algorithm includes: Based on the constructed dictionary tree and according to the beam search algorithm, a specified number of feature vectors of similar sentence pairs that meet preset conditions are selected from the translation results.
6. The method according to claim 1 or 2, wherein: Before inputting the query sentence pair into the encoder of the preset NMT model, the method further includes: Collect multiple candidate sentence pairs; Select multiple groups of sentence pairs with similar semantics from the collected multiple candidate sentence pairs, and mark the corresponding similar sentence pairs for each group of sentence pairs with similar semantics; The NMT model is trained based on annotated sentence pairs with similar semantics.
7. A similar sentence pair retrieval device, comprising: A first acquisition module is adapted to acquire a query sentence pair to be retrieved, and input the query sentence pair into an encoder of a preset NMT model; A word segmentation module, adapted to segment the query sentence pair using an encoder to generate a feature vector corresponding to each word segmentation; An input module, adapted to sequentially input the feature vectors of each word segment into a decoder of the NMT model at preset time intervals; A translation module, adapted to use a decoder to translate the query sentence pair according to the feature vectors of each word segment received in sequence, obtain feature vectors of similar sentence pairs of the query sentence pair, and decode the feature vectors of the similar sentence pairs into corresponding similar sentence pairs; Wherein, the translation module is also suitable for: After the feature vector of the word segmentation of the query sentence pair is first input into the decoder, the decoder is used to translate the received feature vector of the word segmentation into similar sentence pairs to obtain a translation result; Selecting a specified number of feature vectors of similar sentence pairs that meet preset conditions from the translation results according to a preset algorithm; The feature vector of the similar sentence pair selected this time and the feature vector of the next word segment input after the preset time are used as the next input of the decoder decoder, the decoder decoder is used to translate the received input content to obtain a translation result, and a specified number of feature vectors of similar sentence pairs that meet the preset conditions are selected from the translation result according to a preset algorithm, and this cycle is repeated until the translation of the feature vector of the last word segment of the query sentence is completed.
8. The device according to claim 7, wherein: The feature vectors of similar sentence pairs that meet the preset conditions include: The feature vectors of similar sentence pairs whose correlation scores with the feature vector input to the decoder are greater than the specified score.
9. The device according to claim 8, wherein: The preset algorithm includes a beam search algorithm.
10. The device according to claim 9, wherein: Also includes: A second acquisition module, adapted to acquire a candidate set including a plurality of candidate sentence pairs before the first acquisition module inputs the query sentence pair into an encoder of a preset NMT model; The construction module is adapted to construct a dictionary tree for the plurality of candidate sentence pairs in the candidate set.
11. The device according to claim 10, wherein: The translation module is also adapted to: Based on the constructed dictionary tree and according to the beam search algorithm, a specified number of feature vectors of similar sentence pairs that meet preset conditions are selected from the translation results.
12. The device according to claim 7 or 8, wherein: Also includes: A collecting module, adapted for the first acquisition module to collect a plurality of candidate sentence pairs before inputting the query sentence pair into an encoder of a preset NMT model; The tagging module is adapted to select a plurality of sentence pairs having similar semantics from the collected plurality of candidate sentence pairs, and tag the corresponding similar sentence pairs for each group of sentence pairs having similar semantics; The training module is adapted to train the NMT model based on annotated sentence pairs with similar semantics.
13. A computer storage medium storing a computer program code, which, when executed on a computing device, causes the computing device to execute the method for retrieving similar sentence pairs according to any one of claims 1 to 6.
14. A computing device comprising: processor; a memory storing computer program code; When the computer program code is executed by the processor, it causes the computing device to execute the similar sentence pair retrieval method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text translation method, device, storage medium and computer device
CN109145315A
Context-aware peer-to-peer communication
TW201409978A