An English synonym recommendation method integrating dictionary and context information

By combining dictionary and context information, combined with mask language model, filtering and matching synonyms, the problem of inaccurate synonyms in the prior art is solved, and a more accurate and diverse synonym recommendation is achieved.

CN114417825BActive Publication Date: 2025-05-13SHANGHAI YIZHE INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210060733.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-05-13
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The existing methods of English synonyms recommendation lack context information when using dictionaries, resulting in inaccurate synonyms recommendation; while when using mask language model, there are many candidate words but not necessarily synonyms, especially when the context information is insufficient.

Method used

By integrating dictionary and context information, enter the words to be searched and their sentences, calculate the similarity between the dictionary definition and the text context, and sort the dictionary definition; then use the mask language model to obtain candidate words, calculate the semantic similarity between the candidate words and the words to be searched, filter out qualified candidate words, and match them with the dictionary definition with the greatest similarity, and merge synonyms.

Benefits of technology

The definition and synonyms of the words to be queryed that are most in line with the current sentence context are realized, which enriches the diversity of synonyms under each definition and improves the accuracy of synonyms recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114417825B_ABST
    Figure CN114417825B_ABST
Patent Text Reader

Abstract

A method for recommending English synonyms by integrating dictionary and context information comprises the following steps: inputting a query word, obtaining the word's part of speech, English definition and synonyms; using the context information provided by the dictionary definition, calculating the context similarity between the word and the text in which the query word is located, and arranging the dictionary definition in descending order; masking the word to be queried, constructing a masked text, inputting it into a masked language model, and obtaining candidate words based on the context; calculating the semantic similarity between the candidate word and the word to be queried, and screening out qualified candidate words; matching the qualified candidate words with the dictionary definition based on the reordered dictionary definition and dictionary synonyms, merging synonyms, and generating a new recommended synonym list. The present invention overcomes the shortcomings of the prior art, and after inputting the word to be queried and the sentence in which it is located, it can output the definition and synonyms of the word to be queried that best fit the current sentence context, and enrich the diversity of synonyms under each definition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to an English synonym recommendation method integrating dictionary and context information. Background Art

[0002] In English writing scenarios, the synonym recommendation function can help authors write more authentic articles. The existing synonym recommendation methods are divided into the following two types:

[0003] First, query various English dictionaries and return the word's part of speech, definition, and synonyms. The advantage of this method is that the word's part of speech, definition, and synonyms are corresponding, and users can find the corresponding word definition and synonyms based on the specific text. The disadvantage is that the dictionary has limited and fixed synonyms, and does not use context information, multiple parts of speech, and polysemous words, which increases the user's query time. For example, "good" has 3 parts of speech and 28 definitions in the open source dictionary WordNet.

[0004] Second, using the masked language model, by masking the words to be queried, using context information, predicting the masked words, and returning candidate words as synonyms. The advantage of this method is that the candidate words are diverse and conform to the context of the text to be queried. The disadvantage is that these candidate words are actually co-occurring words, and are not necessarily synonyms with the words to be queried, and may even be antonyms, especially when the text is short and the context information is insufficient. For example, if the text to be queried is "I am fine." and the word to be queried is "fine", masking "fine", and inputting the masked language model, the first five candidate words returned are "here", "sorry", "ready", "alive", and "not". These words are not synonyms with "fine" and are not qualified recommended replacements. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides an English synonym recommendation method that integrates dictionary and context information, overcomes the shortcomings of the prior art, inputs the word to be queried and the sentence in which it is located, and can output the definition and synonyms of the word to be queried that best fit the current sentence context, and enrich the diversity of synonyms under each definition.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0007] An English synonym recommendation method integrating dictionary and context information comprises the following steps:

[0008] Step S1: Input a query word and obtain the word's part of speech, English definition and synonyms;

[0009] Step S2: using the context information provided by the dictionary definition to calculate the degree of similarity between the context and the text containing the queried word; and arranging the dictionary definitions in descending order according to the degree of similarity;

[0010] Step S3: Mask the word to be queried, construct a masked text, input it into the masked language model, and obtain candidate words based on the context;

[0011] Step S4: Calculate the semantic similarity between the candidate words and the query word, and set a threshold to filter out qualified candidate words;

[0012] Step S5: Based on the re-ordered dictionary definitions and dictionary synonyms, the qualified candidate words are matched with the dictionary definitions, the synonyms are merged, and a new recommended synonym list is generated.

[0013] Preferably, the step S2 specifically includes the following steps:

[0014] Step S21: input the dictionary definition of the query word and the text in which the word is located into the text representation model Sentence Bert respectively, and construct the dictionary definition vector and the text context vector;

[0015] Step S22: Calculate the cosine similarity between the definition and the text. The specific formula is as follows:

[0016]

[0017]

[0018]

[0019] Among them, context_vector is the context vector of the query text, m is the number of words in the query text, and v i is the i-th word vector of the query text; definition_vector j is the jth definition vector of the query word, n is the number of words in the definition text, v k is the k-th word vector of the definition text; def_con_sim j is the cosine similarity between the jth definition vector and the context text vector;

[0020] Step S23: Arrange the dictionary definitions in descending order according to the similarity, and the definitions that are closer to the text context are ranked at the top.

[0021] Preferably, the masked language model in step S3 is a BERT model.

[0022] Preferably, the step S4 specifically includes the following steps:

[0023] Step S41: input the candidate word and the query word into the text representation model Sentence Bert respectively, and output the candidate word vector and the query vector;

[0024] Step S42: Calculate the cosine similarity between the candidate word vector and the query vector. The formula is as follows:

[0025]

[0026] Among them, query_vector is the query vector, alter_vector i is the i-th candidate word vector, query_alter_sim i is the cosine similarity between the i-th candidate word vector and the query vector;

[0027] Step S43: According to the similarity between the candidate word vector and the query vector, candidate words with a similarity greater than 0.8 are regarded as qualified candidate words.

[0028] Preferably, the step S5 specifically includes the following steps:

[0029] Step S51: Calculate the cosine similarity between the qualified candidate word vector and the dictionary definition vector, the formula is as follows:

[0030]

[0031] Among them, alter_def_sim ij is the i-th qualified candidate word vector good_alter_vector i and the jth definition vector definition_vector j The cosine similarity of

[0032] Step S52: Match the qualified candidate words with the dictionary definition with the greatest similarity, and merge the dictionary synonyms and the qualified candidate words as recommended synonyms.

[0033] The present invention provides an English synonym recommendation method integrating dictionary and context information. The method has the following beneficial effects: by calculating the contextual similarity between the contextual information provided by the dictionary definition and the text in which the queried word is located, and sorting them, and obtaining candidate words based on the context, screening out qualified candidate words by setting a threshold, and then matching the qualified candidate words with the dictionary definition with the greatest similarity, merging the dictionary synonyms and the qualified candidate words as recommended synonyms; thereby being able to output the definition and synonyms of the queried word that best fit the current sentence context, and enriching the diversity of synonyms under each definition. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the present invention or the technical solutions in the prior art, the drawings required for describing the prior art are briefly introduced below.

[0035] Figure 1 Flow chart of the present invention;

[0036] Figure 2 The present invention recommends a full flow chart of word definitions and synonym generation;

[0037] Figure 3 Flowchart for obtaining dictionary definitions and synonyms in the present invention;

[0038] Figure 4 A flowchart of defining the recommended sorting in the present invention;

[0039] Figure 5 A flowchart of context-based candidate word generation in the present invention;

[0040] Figure 6 Flow chart of screening qualified candidate words in the present invention;

[0041] Figure 7 Flowchart of matching qualified candidate words and definitions in the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention.

[0043] like Figure 1-7 As shown, a method for recommending English synonyms by integrating dictionary and context information includes the following steps:

[0044] Step S1: Enter the word to be searched, and obtain the word's part of speech, English definition, and synonyms; the open source English dictionary used in this application is WordNet. For example, if you enter "good", the definition and corresponding synonyms are returned, such as Figure 3 As shown;

[0045] Step S2: Since the order of dictionary definitions is fixed, the context information provided by the dictionary definitions is used to calculate the degree of similarity between the context and the text containing the queried word; and then the dictionary definitions are arranged in descending order according to the degree of similarity. If the similarity is higher and the context is more matched, the dictionary definition is positioned higher.

[0046] This application constructs a dictionary definition vector and a text context vector by inputting the dictionary definition of the query word and the text in which the word is located into the text representation model Sentence Bert respectively; and then calculates the cosine similarity between the definition and the text. The specific formula is as follows:

[0047]

[0048]

[0049]

[0050] Among them, context_vector is the context vector of the query text, m is the number of words in the query text, and v i is the i-th word vector of the query text; definition_vector j is the jth definition vector of the query word, n is the number of words in the definition text, v k is the k-th word vector of the definition text; def_con_sim j It is the cosine similarity between the j-th definition vector and the context text vector; then the dictionary definitions are sorted in descending order according to the similarity, and the definitions that are closer to the text context are ranked at the front.

[0051] For example, the query text is "It's a good idea.", the query item is "good", and the three dictionary definitions of "good" are selected as "of moral excellence", "resulting favorably", and "in excellent physical condition". The similarities with the query text are 0.40, 0.77, and 0.70 respectively. Therefore, the definitions are reordered as "resulting favorably", "in excellent physical condition", and "of moral excellence". Figure 4 shown.

[0052] Step S3: Mask the word to be queried, construct a masked text, and input it into the masked language model. The masked language model used in this application is BERT. Then, based on the context information, output the candidate words of the top 20 masked words.

[0053] For example, if the query text is "It's a good idea" and the query item is "good", "good" is masked as "It's a[MASK]idea", and candidate words such as "great", "wonderful", and "bad" are output. Figure 5 As shown in the figure, “bad” is obviously not a synonym of “good”, so further screening is required.

[0054] Step S4: In order to exclude antonymous candidate words, the semantic similarity between the candidate words and the query word is calculated, and a threshold is set to screen out qualified candidate words;

[0055] This application inputs the candidate word and the query word into the text representation model Sentence Bert respectively, outputs the candidate word vector and the query vector; then calculates the cosine similarity of the candidate word vector and the query vector, the formula is as follows:

[0056]

[0057] Among them, query_vector is the query vector, alter_vector i is the i-th candidate word vector, query_alter_sim i is the cosine similarity between the i-th candidate word vector and the query vector;

[0058] Then, according to the similarity between the candidate word vector and the query vector, the candidate words with a similarity greater than 0.8 are regarded as qualified candidate words.

[0059] like Figure 6 As shown, for example, the similarities between "good" and "great", "wonderful" and "bad" are 0.92, 0.85 and 0.43 respectively, so "bad" is excluded and "great" and "wonderful" are retained.

[0060] Step S5: Based on the re-ordered dictionary definitions and dictionary synonyms, the qualified candidate words are matched with the dictionary definitions, the synonyms are merged, and a new recommended synonym list is generated.

[0061] This application calculates the cosine similarity between the qualified candidate word vector and the dictionary definition vector, and the formula is as follows:

[0062]

[0063] Among them, alter_def_sim ij is the i-th qualified candidate word vector good_alter_vector i and the jth definition vector definition_vector j The cosine similarity of ; then the qualified candidate word is matched with the dictionary definition with the greatest similarity, and the dictionary synonyms and the qualified candidate word are merged as the recommended synonyms.

[0064] like Figure 7For example, the qualified candidate words "great" and "wonderful" for "good" have similarities with the three dictionary definitions of "good" as shown in Table 1: "great" matches the definition "resulting favorably", and "wonderful" matches the definition "in excellent physical condition".

[0065] Table 1 Similarity between the dictionary definitions of the words “great”, “wonderful” and the word “good”

[0066] of moral excellence resulting favorably in excellent physical condition great 0.48 0.80 0.78 wonderful 0.43 0.72 0.73

[0067] Based on the above process, when the query text is "It's a good idea" and the query item is "good", dictionary definitions and synonyms are obtained from the open source English dictionary WordNet, qualified candidate words are obtained from the language model, and finally the top 3 recommended definitions and synonyms are output.

[0068] The present application calculates the similarity between the context information provided by the dictionary definition and the context of the text in which the queried word is located, sorts them, obtains candidate words based on the context, screens out qualified candidate words by setting a threshold, and then matches the qualified candidate words with the dictionary definition with the greatest similarity, merges the dictionary synonyms and the qualified candidate words as recommended synonyms; thereby, the definition and synonyms of the queried word that best fit the current sentence context can be output, and the diversity of synonyms under each definition can be enriched.

[0069] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for recommending English synonyms by integrating dictionary and context information, characterized in that: The following steps are involved: Step S1: Input a query word and obtain the word's part of speech, English definition and synonyms; Step S2: using the context information provided by the dictionary definition to calculate the degree of similarity between the context and the text containing the queried word; and arranging the dictionary definitions in descending order according to the degree of similarity; Step S3: Mask the word to be queried, construct a masked text, input it into the masked language model, and obtain candidate words based on the context; Step S4: Calculate the semantic similarity between the candidate words and the query word, and set a threshold to filter out qualified candidate words; Step S5: Based on the re-ordered dictionary definitions and dictionary synonyms, matching qualified candidate words with dictionary definitions, merging synonyms, and generating a new recommended synonym list; Wherein, the step S2 specifically includes the following steps: Step S21: input the dictionary definition of the query word and the text in which the word is located into the text representation model SentenceBert respectively, and construct the dictionary definition vector and the text context vector; Step S22: Calculate the cosine similarity between the definition and the text; Step S23: Arrange the dictionary definitions in descending order according to the similarity, and the definitions that are closer to the text context are ranked at the top; In step S3, the masked language model is a BERT model; The step S4 specifically comprises the following steps: Step S41: input the candidate word and the query word into the text representation model Sentence Bert respectively, and output the candidate word vector and the query vector; Step S42: Calculate the cosine similarity between the candidate word vector and the query vector; Step S43: According to the similarity between the candidate word vector and the query vector, candidate words with a similarity greater than 0.8 are regarded as qualified candidate words.

2. The English synonym recommendation method integrating dictionary and context information according to claim 1, characterized in that: The step S22: calculating the cosine similarity between the definition and the text, the specific formula is as follows: Among them, context_vector is the context vector of the query text, m is the number of words in the query text, and v i is the i-th word vector of the query text; definition_vector j is the jth definition vector of the query word, n is the number of words in the definition text, v k is the k-th word vector of the definition text; def_con_sim j is the cosine similarity between the jth definition vector and the context text vector.

3. The English synonym recommendation method integrating dictionary and context information according to claim 1, characterized in that: Step S42: Calculate the cosine similarity between the candidate word vector and the query vector. The formula is as follows: Among them, query_vector is the query vector, alter_vector i is the i-th candidate word vector, query_alter_sim i is the cosine similarity between the i-th candidate word vector and the query vector.

4. The English synonym recommendation method integrating dictionary and context information according to claim 1, characterized in that: The step S5 specifically comprises the following steps: Step S51: Calculate the cosine similarity between the qualified candidate word vector and the dictionary definition vector, the formula is as follows: Among them, alter_def_sim ij is the i-th qualified candidate word vector good_alter_vector i and the jth definition vector definition_vector j The cosine similarity of Step S52: Match the qualified candidate words with the dictionary definition with the greatest similarity, and merge the dictionary synonyms and the qualified candidate words as recommended synonyms.

Citation Information

Patent Citations

  • Reader-centered personalized English text simplification method

    CN113705223A