Method, apparatus, storage medium and device for extracting attention phrases in text

By constructing inverted index and multimodal matching models, the text is detected and collected, segmented and screened candidate fragments, the problem of the incomplete identification of similar variants of phrases in the text in the prior art is solved, and the accuracy of recognition is improved.

CN119646115BActive Publication Date: 2025-06-17BEIJING UCAP INTERNET TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411682398.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-06-17
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In the prior art, when synonyms or synonyms dictionaries are used to identify similar variants of phrases in text, the recognition cannot be fully recognized, resulting in poor recognition effect.

Method used

By constructing inverted index and multimode matching models, using multimode matching models to detect the input text, multiple fragments are obtained, and fragments are collected using inverted indexes, segmenting and filtering candidate fragments according to preset editing distances, and finally combining them into similar variants of the phrase of concern.

Benefits of technology

Improve the accuracy of recognition of similar variants of attention phrases in the text, allowing more comprehensive recognition of attention phrases in the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646115B_ABST
    Figure CN119646115B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, storage medium and device for extracting a concerned phrase in a text, belonging to the field of computer technology. The method includes: obtaining an inverted index and a multi-mode matching model constructed for the concerned phrase; using the multi-mode matching model to detect the input text to obtain a plurality of fragments; using the inverted index to collect the plurality of fragments to obtain a first fragment set; segmenting the first fragment set according to a preset edit distance to obtain multiple groups of second fragment sets; for each group of second fragment sets, screening candidate fragments according to the edit distance between the fragments in the second fragment set and the concerned phrase; combining the candidate fragments in each group of second fragment sets into a similar variant of the concerned phrase, and identifying the similar variant with an edit distance greater than a predetermined threshold from the concerned phrase as the concerned phrase. The present application can use the improved multi-mode detection model and the inverted index to identify the similar variant of the concerned phrase in the text as the concerned phrase, improving the accuracy of extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, storage medium, and device for extracting phrases of interest in text. Background Art

[0002] In the broad field of Natural Language Processing (NLP), extracting and accurately identifying similar variants of phrases of interest in text has become a crucial and challenging task. Especially in the field of government text monitoring, even expressions with similar meanings must maintain the original words and appearances, which makes quickly and precisely capturing various similar variants of phrases of interest in text a key link to improving the system's efficiency. For example, "Zhonghua Ren43 Min Gongheguo" in the text needs to be recognized as "People's Republic of China".

[0003] In related technologies, we can construct a thesaurus or synonym dictionary for phrases of interest, and match similar variants in the text with the thesaurus or synonym dictionary, so as to recognize the similar variants as the correct phrases of interest.

[0004] However, the thesaurus or synonym dictionary is compiled manually, with a limited coverage, and its update speed is difficult to keep up with the dynamic development of the language, resulting in the inability to comprehensively recognize phrases of interest in the text. Summary of the Invention

[0005] This application provides a method, apparatus, storage medium, and device for extracting phrases of interest in text, which is used to solve the problem that when using a thesaurus or synonym dictionary, similar variants in the text cannot be comprehensively recognized as the correct phrases of interest. The technical solutions are as follows:

[0006] According to the first aspect of this application, a method for extracting phrases of interest in text is provided. The method includes:

[0007] Obtain an inverted index and a multi-pattern matching model constructed for the phrases of interest, where the multi-pattern matching model is generated according to the phrases of interest, the similar character set of the phrases of interest, and the ignored character set;

[0008] Use the multi-pattern matching model to detect the input text to obtain multiple fragments;

[0009] Use the inverted index to collect the multiple fragments to obtain a first fragment set;

[0010] According to a preset edit distance, split the first fragment set to obtain multiple groups of second fragment sets;

[0011] For each set of second fragments, candidate fragments are filtered according to the edit distance between the fragments in the second fragment set and the concerned phrase;

[0012] The candidate fragments in each set of second fragments are combined into similar variants of the concerned phrase, and the similar variants with an edit distance greater than a predetermined threshold from the concerned phrase are identified as the concerned phrase.

[0013] In a possible implementation, the use of the multi-mode matching model to detect the input text to obtain multiple fragments includes:

[0014] Obtain the i-th character in the input text, where i is a positive integer;

[0015] If the i-th character is an ignored character in the ignored character set, update i to i + 1 and continue to execute the step of obtaining the i-th character in the input text;

[0016] If the i-th character is not an ignored character, find all similar characters of the i-th character in the similar character set, match the i-th character and all similar characters with the trie in the multi-mode matching model, and generate fragments according to the matching results;

[0017] Wherein, the fragment includes a phrase extracted from the text, the start position and the end position of the phrase in the text, and the phrase matched in the trie according to the extracted phrase.

[0018] In a possible implementation, the filtering of candidate fragments according to the edit distance between the fragments in the second fragment set and the concerned phrase includes:

[0019] Select m consecutive fragments from the n fragments in the second fragment set, where 1 ≤ m ≤ n;

[0020] Determine a candidate phrase according to the start position of the first fragment and the end position of the m-th fragment among the m fragments, and retain the candidate phrases with an edit distance less than a predetermined threshold from the concerned phrase;

[0021] Among all the obtained candidate phrases, filter out the candidate phrases with the longest length and non-overlapping start positions and end positions with each other, and determine the fragments corresponding to the candidate phrases as candidate fragments.

[0022] In a possible implementation, the combining of the candidate fragments in each set of second fragments into similar variants of the concerned phrase includes:

[0023] For each set of second fragments, each candidate fragment in the second fragment set is combined according to the start position and the end position to obtain a similar variant of the concerned phrase.

[0024] In a possible implementation manner, the collecting the multiple fragments by using the inverted index to obtain a first fragment set includes:

[0025] Using the inverted index to determine the concerned phrase to which each fragment belongs;

[0026] Extracting the first character and the last character from the concerned phrase, and supplementing the fragments of the first character and the fragments of the last character;

[0027] Among all the obtained fragments, screening the fragments with the longest length and non-overlapping start positions and end positions with each other to obtain a first fragment set.

[0028] In a possible implementation manner, the identifying a similar variant with an edit distance greater than a predetermined threshold from the concerned phrase as the concerned phrase includes:

[0029] Calculating the edit distance between the similar variant and the concerned phrase;

[0030] If the edit distance is less than the predetermined threshold, then identifying the similar variant as the concerned phrase.

[0031] In a possible implementation manner, the obtaining an inverted index and a multi-pattern matching model constructed for a concerned phrase includes:

[0032] Splitting the concerned phrase, and constructing an inverted index according to the splitting result;

[0033] Obtaining a set of similar characters for each character in the concerned phrase, where the set of similar characters includes at least one of a set of homophonic characters, a set of phonetically similar characters, and a set of visually similar characters;

[0034] Obtaining a set of ignored characters, where the set of ignored characters includes at least one of a set of numbers and a tag set, and the tag set is a tag to be ignored set according to the application field;

[0035] Constructing a multi-pattern matching model according to the set of similar characters and the set of ignored characters.

[0036] According to the second aspect of the present application, there is provided an apparatus for extracting a concerned phrase in a text, the apparatus includes:

[0037] An obtaining module, configured to obtain an inverted index and a multi-pattern matching model constructed for a concerned phrase, where the multi-pattern matching model is generated according to the concerned phrase, the set of similar characters of the concerned phrase, and the set of ignored characters;

[0038] A detection module, configured to detect the input text by using the multi-pattern matching model to obtain multiple fragments;

[0039] An aggregation module, configured to aggregate the multiple fragments by using the inverted index to obtain a first fragment set;

[0040] A segmentation module, configured to segment the first fragment set according to a preset edit distance to obtain multiple groups of second fragment sets;

[0041] A screening module, configured to, for each group of second fragment sets, screen candidate fragments according to the edit distance between the fragments in the second fragment sets and the concerned phrase;

[0042] An identification module, configured to combine the candidate fragments in each group of second fragment sets into a similar variant of the concerned phrase, and identify the similar variant whose edit distance from the concerned phrase is greater than a predetermined threshold as the concerned phrase.

[0043] According to a third aspect of the present application, there is provided a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the method for extracting the concerned phrase in the text as described above.

[0044] According to a fourth aspect of the present application, there is provided a computer device, which includes the device for extracting the concerned phrase in the above text.

[0045] The beneficial effects of the technical solution provided by the present application at least include:

[0046] By generating an improved multi-pattern detection model according to the concerned phrase, the similar character set of the concerned phrase, and the ignored character set, detecting the input text by using the multi-pattern matching model to obtain multiple fragments; then aggregating the multiple fragments by using the inverted index of the concerned phrase to obtain a first fragment set; then, segmenting the first fragment set according to a preset edit distance to obtain multiple groups of second fragment sets; for each group of second fragment sets, screening candidate fragments according to the edit distance between the fragments in the second fragment sets and the concerned phrase; combining the candidate fragments in each group of second fragment sets into a similar variant of the concerned phrase, and identifying the similar variant whose edit distance from the concerned phrase is greater than a predetermined threshold as the concerned phrase, the extraction accuracy is improved. Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0048] Figure 1 It is a flowchart of a method for extracting concerned phrases in a text provided by an embodiment of the present application;

[0049] Figure 2 It is a flowchart of a method for extracting concerned phrases in a text provided by an embodiment of the present application;

[0050] Figure 3 It is a schematic diagram for extracting instances of concerned phrases provided by an embodiment of the present application;

[0051] Figure 4 It is a structural block diagram of an apparatus for extracting concerned phrases in a text provided by an embodiment of the present application. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0053] As Figure 1 shown, it shows a method flowchart of a method for extracting concerned phrases in a text provided by an embodiment of the present application. The method for extracting concerned phrases in the text can be applied to a computer device. The method for extracting concerned phrases in the text may include:

[0054] Step 101: Obtain an inverted index and a multi-pattern matching model constructed for concerned phrases. The multi-pattern matching model is generated according to concerned phrases, a similar character set of the concerned phrases, and an ignored character set.

[0055] A concerned phrase is a proper noun that needs to be accurately recognized without errors. For example, the concerned phrase may be "the People's Republic of China".

[0056] When constructing the inverted index, the concerned phrases can be segmented, and the inverted index can be constructed according to the attribution information of the segmentation.

[0057] When constructing the multi-pattern matching model, it is necessary to obtain a similar character set and an ignored character set of the concerned phrases, and use the multi-pattern matching algorithm to generate a tree-structured multi-pattern matching model for the concerned phrases, the similar character set, and the ignored character set.

[0058] Step 102: Use the multi-pattern matching model to detect the input text to obtain multiple fragments.

[0059] If the input text includes similar variants of the concerned phrases, the similar variants need to be recognized as the concerned phrases. Among them, the similar variants are obtained by replacing some characters in the concerned phrases with similar characters, inserting ignored characters in the concerned phrases, deleting some characters in the concerned phrases, etc. For example, if the input text is "The Constitution of the Republic of China 2 is the fundamental law of the Republic of the Chinese People 43", and the concerned phrase is "the People's Republic of China", the similar variants include "the Republic of China 2" and "the Republic of the Chinese People 43".

[0060] The multi-mode matching model includes a trie composed of multiple concerned phrases. When using the multi-mode matching model, the input text will be preprocessed according to the set of similar characters and the set of ignored characters, so as to find the corresponding phrases in the trie. Then, the phrases in the text can be fragmented based on the phrases found in the trie for the text phrase, as well as the starting position and ending position of the text phrase.

[0061] For example, if the input text is "The Constitution of the Republic of China 2 is the fundamental law of the Republic of the Chinese People 43", the phrase in the text is "the Republic of China 2", and the phrase found in the multi-mode matching model is "the People's Republic of China", then the generated fragment can be recorded as ("the Republic of China 2", 0, 3, "the People's Republic of China").

[0062] Step 103: Use the inverted index to collect multiple fragments to obtain the first fragment set.

[0063] Specifically, the phrases found in the multi-mode matching model can be used to find the corresponding concerned phrases in the inverted index, and then all the fragments corresponding to one concerned phrase are combined into the first fragment set.

[0064] For example, the "the People's Republic of China" corresponding to the "the People's Republic of China" in the fragment ("the Republic of China 2", 0, 3, "the People's Republic of China"), and the "the People's Republic of China" corresponding to the "the People's Republic of China" in the fragment ("the Republic", 3, 6, "the Republic"), then the first fragment set corresponding to "the People's Republic of China" is {("the Republic of China 2", 0, 3, "the People's Republic of China"), ("the Republic", 3, 6, "the Republic")}.

[0065] Step 104: Split the first fragment set according to the preset edit distance to obtain multiple groups of second fragment sets.

[0066] The preset edit distance is set for the concerned phrases. The preset edit distances for different concerned phrases can be the same or different.

[0067] Specifically, the fragments in the first fragment set can be divided according to the number of the preset edit distance to obtain multiple groups of second fragment sets.

[0068] For example, if the preset edit distance is 4 and the first fragment set includes fragments 1-6, one way is to divide fragments 1-4 into a second fragment set and fragments 5-6 into another second fragment set; another way is to divide fragments 1-2 into a second fragment set and fragments 3-6 into another second fragment set.

[0069] Step 105: For each group of second fragment sets, screen candidate fragments according to the edit distance between the fragments in the second fragment set and the focus phrase.

[0070] Specifically, at least one fragment in the second fragment set can be used to form a candidate phrase according to the start position and end position, calculate the edit distance between the candidate phrase and the focus phrase. If the edit distance is less than the preset edit distance, retain the fragment corresponding to the candidate phrase to obtain candidate fragments; if the edit distance is greater than the preset edit distance, delete the combination method of the fragments corresponding to the candidate phrase.

[0071] Step 106: Combine the candidate fragments in each group of second fragment sets into similar variants of the focus phrase, and identify the similar variants with an edit distance greater than a predetermined threshold from the focus phrase as the focus phrase.

[0072] Specifically, form similar variants from the candidate fragments according to the start position and end position, calculate the edit distance between the similar variant and the focus phrase. If the edit distance is less than the predetermined threshold, identify the similar variant as the focus phrase; if the edit distance is greater than the predetermined threshold, do not identify the similar variant as the focus phrase.

[0073] In summary, for the method for extracting a focus phrase in a text provided in an embodiment of the present application, an improved multi-mode detection model is generated according to the focus phrase, the similar character set of the focus phrase, and the ignored character set, and the input text is detected using the multi-mode matching model to obtain multiple fragments; then, the multiple fragments are grouped using the inverted index of the focus phrase to obtain a first fragment set; then, according to the preset edit distance, the first fragment set is split to obtain multiple groups of second fragment sets; for each group of second fragment sets, candidate fragments are screened according to the edit distance between the fragments in the second fragment set and the focus phrase; the candidate fragments in each group of second fragment sets are combined into similar variants of the focus phrase, and the similar variants with an edit distance greater than a predetermined threshold from the focus phrase are identified as the focus phrase, improving the extraction accuracy.

[0074] As Figure 2 shown, it shows a flowchart of the method for extracting a focus phrase in a text provided in an embodiment of the present application. The method for extracting a focus phrase in this text can be applied to a computer device. The method for extracting a focus phrase in this text may include:

[0075] Step 201: Split the concerned phrase and construct an inverted index according to the splitting result.

[0076] When constructing the inverted index, the concerned phrase can be segmented, and an inverted index can be constructed according to the attribution information of the segmentation.

[0077] For example, if the concerned phrase is "People's Republic of China", the segmented words after splitting include: "Zhonghua", "Huaren", "Gonghe", "Renmin", "Mingong", "Guohe", "Zhonghuaren", "Renmingong", "Mingonghe", "Gongheguo", "Huarenmin", "Zhonghuarenmin", "Mingongheguo", "Huarenmingong", "Renmingonghe", "Huarenmingonghe", "Zhonghuarenmingong", "Renmingongheguo", "Huarenmingongheguo", "Zhonghuarenmingonghe", "People's Republic of China".

[0078] Step 202: Obtain the similar character set for each character in the concerned phrase. The similar character set includes at least one of the homophone set, the phonetically similar character set, and the visually similar character set.

[0079] The homophone set refers to the set of characters with the same pronunciation as the characters in the concerned phrase. The phonetically similar character set refers to the set of characters with a pronunciation similar to the characters in the concerned phrase. The visually similar character set refers to the set of characters with a similar shape to the characters in the concerned phrase.

[0080] Step 203: Obtain the ignored character set. The ignored character set includes at least one of the number set and the tag set. The tag set is the set of tags that need to be ignored according to the application field.

[0081] The number set includes 0, 1, 2, 3, 4, 5, 6, 7, 8, 9.

[0082] The tag set is the set of tags set according to the application field. For example, when applied in the web page field, the tag set can include HyperText Markup Language (HTML) tags.

[0083] Step 204: Construct a multi-pattern matching model according to the similar character set and the ignored character set.

[0084] Among them, the multi-pattern matching model is generated according to the segmentation of the concerned phrase, the similar character set of the concerned phrase, and the ignored character set.

[0085] Step 205: Use the multi-pattern matching model to detect the input text and obtain multiple fragments.

[0086] If the input text includes similar variants of the concerned phrases, the similar variants need to be recognized as the concerned phrases. Among them, the similar variants are obtained by replacing some characters in the concerned phrases with similar characters, inserting ignored characters in the concerned phrases, deleting some characters in the concerned phrases, etc. For example, if the input text is "The Constitution of the Zhong2hua Republic is the fundamental law of the Zhonghua People's Republic", and the concerned phrase is "People's Republic of China", the similar variants include "Zhong2hua Republic" and "Zhonghua People's Republic".

[0087] The multi-mode matching model includes a trie tree composed of multiple concerned phrases. When using the multi-mode matching model, the input text will be preprocessed according to the similar character set and the ignored character set to find the corresponding phrases in the trie tree.

[0088] Specifically, using the multi-mode matching model to detect the input text, multiple fragments can be obtained, including: obtaining the i-th character in the input text, where i is a positive integer; if the i-th character is an ignored character in the ignored character set, then update i to i + 1 and continue to execute the step of obtaining the i-th character in the input text; if the i-th character is not an ignored character, find all similar characters of the i-th character in the similar character set, match the i-th character and all similar characters with the trie tree in the multi-mode matching model, and generate fragments according to the matching results; among them, the fragments include the phrases extracted from the text, the start position and end position of the phrases in the text, and the phrases matched by the extracted phrases in the trie tree.

[0089] For example, if the input text is "The Constitution of the Zhong2hua Republic is the fundamental law of the Zhonghua People's Republic", for the first character "Zhong", start from the root node of the trie tree to find the "Zhong" node; for the second character "2", since "2" is a number in the ignored character set, delete "2"; for the third character "hua", start from the "Zhong" node and look down to find the "hua" node; for the fourth character "gong", start from the "hua" node and look down, and no corresponding node is found. Start from the root node of the trie tree to find the "gong" node, and so on.

[0090] Then, the phrases in the text can be used to generate fragments based on the phrases found in the trie tree and the start position and end position of the phrases in the text. For example, if the phrase in the text is "Zhong2hua" and the phrase found in the multi-mode matching model is "Zhonghua", the generated fragment can be recorded as ("Zhong2hua", 0, 3, "Zhonghua").

[0091] If the input text is "The Constitution of the Zhong2hua Republic is the fundamental law of the Zhonghua People's Republic", the fragments output include: ("Zhong2hua", 0, 3, "Zhonghua"), ("Republic", 3, 6, "Republic"), ("Repub", 3, 5, "Repub"), ("Zhonghua People's Rep", 9, 16, "Zhonghua People's Rep"), ("Zhonghua People", 9, 15, "Zhonghua People"), ("Zhonghua Person", 9, 12, "Zhonghua Person"), ("Zhonghua", 9, 11, "Zhonghua").

[0092] Step 206, use the inverted index to collect multiple fragments to obtain the first fragment set.

[0093] In this embodiment, the phrases found in the multi-modal matching model can be used to find the corresponding concerned phrases in the inverted index, and then all the fragments corresponding to one concerned phrase are combined into the first fragment set.

[0094] Specifically, using the inverted index to collect multiple fragments to obtain the first fragment set may include:

[0095] (1) Use the inverted index to determine the concerned phrase to which each fragment belongs.

[0096] For example, if the concerned phrase is "People's Republic of China", the first fragment set is [("Zhong2hua", 0, 3, "Zhonghua"), ("Republic", 3, 6, "Republic"), ("Repub", 3, 5, "Repub"), ("Zhonghua People's Rep", 9, 16, "Zhonghua People's Rep"), ("Zhonghua People", 9, 15, "Zhonghua People"), ("Zhonghua Person", 9, 12, "Zhonghua Person"), ("Zhonghua", 9, 11, "Zhonghua")].

[0097] (2) Extract the first character and the last character from the concerned phrase, and supplement the fragments of the first character and the fragments of the last character.

[0098] Optionally, when there are no fragments of the first character and / or no fragments of the last character of the concerned phrase in the first fragment set, the fragments of the first character and the fragments of the last character also need to be supplemented.

[0099] Taking the concerned phrase "People's Republic of China" and the above first fragment set as an example, the fragments of the first character ("Zhong", 0, 1, "Zhong") and the fragments of the last character ("Guo", 16, 17, "Guo") need to be supplemented.

[0100] (3) Among all the obtained fragments, filter out the fragments with the longest length and non-overlapping start positions and end positions with each other to obtain the first fragment set.

[0101] Since there is overlap between ("中2华", 0, 3, "中华") and ("中", 0, 1, "中"), keep the longest one ("中2华", 0, 3, "中华") and delete ("中", 0, 1, "中").

[0102] Since there is overlap between ("Republic", 3, 6 "Republic") and ("Republic", 3, 5 "Republic"), keep the longest one ("Republic", 3, 6 "Republic") and delete ("Republic", 3, 5 "Republic").

[0103] Since there are overlaps between ("中化人43民共", 9, 16, "中国人民共"), ("中化人43民", 9, 15, "中国人民"), ("中化人", 9, 12, "中国人"), and ("中化", 9, 11, "中国"), the longest one ("中化人43民共", 9, 16, "中国人民共") is retained, and ("中化人43民", 9, 15, "中国人民"), ("中化人", 9, 12, "中国人"), and ("中化", 9, 11, "中国") are deleted.

[0104] After the above screening, the first fragment set finally obtained is [(“中2华”, 0, 3, “中国”), (“人民”, 3, 6 “人民”), (“中化人43人民共”, 9, 16, “中国人民”), (“国”, 16, 17, “国”)].

[0105] Step 207 , dividing the first fragment set according to a preset edit distance to obtain multiple groups of second fragment sets.

[0106] The preset edit distance is set for the focus phrase, and the preset edit distances of different focus phrases may be the same or different.

[0107] Specifically, the fragments in the first fragment set may be divided according to the number of preset edit distances to obtain multiple groups of second fragment sets.

[0108] For example, the preset edit distance is 4, and the first fragment set includes fragments 1-6. Then, one method is to divide fragments 1-4 into a group of second fragment sets, and divide fragments 5-6 into another group of second fragment sets; another method is to divide fragments 1-2 into a group of second fragment sets, and divide fragments 3-6 into another group of second fragment sets.

[0109] Step 208: For each set of second fragments, candidate fragments are screened according to the edit distance between the fragments in the second fragment set and the focus phrase.

[0110] Specifically, screening candidate fragments according to the edit distance between the fragments in the second fragment set and the focus phrase may include:

[0111] (1) Select m consecutive fragments from the n fragments of the second fragment set, where 1 ≤ m ≤ n.

[0112] Among them, the value of m ranges from 1 to n. Suppose n = 4. When m = 4, fragments 1 - 4 form a candidate phrase; when m = 3, fragments 1 - 3 form a candidate phrase, and fragments 2 - 4 form a candidate phrase; when m = 2, fragments 1 - 2 form a candidate phrase, fragments 2 - 3 form a candidate phrase, and fragments 3 - 4 form a candidate phrase; when m = 1, fragment 1 forms a candidate phrase, fragment 2 forms a candidate phrase, fragment 3 forms a candidate phrase, and fragment 4 forms a candidate phrase.

[0113] (2) Determine the candidate phrases according to the start position of the first fragment and the end position of the mth fragment among the m fragments, and retain the candidate phrases whose edit distance from the concerned phrase is less than a predetermined threshold.

[0114] Specifically, at least one fragment in the second fragment set can be combined into candidate phrases according to the start position and the end position. Calculate the edit distance between the candidate phrase and the concerned phrase. If the edit distance is less than the preset edit distance, retain the fragments corresponding to the candidate phrase to obtain candidate fragments; if the edit distance is greater than the preset edit distance, delete the combination method of the fragments corresponding to the candidate phrase.

[0115] (3) Among all the obtained candidate phrases, screen out the candidate phrases with the longest length and non - overlapping start positions and end positions with each other, and determine the fragments corresponding to the candidate phrases as candidate fragments.

[0116] For example, the first concerned phrase is "People's Republic of China", and the corresponding candidate fragments are: ("Zhong2Hua", 0, 3, "ZhongHua"), ("Republic of China", 3, 6, "Republic of China"); the second concerned phrase is "People's Republic of China", and the corresponding candidate fragments are: ("ZhongHuaRen43MinGong", 9, 16, "ZhongHuaRenMinGong"), ("Guo", 16, 17, "Guo").

[0117] Step 209: Combine the candidate fragments in each group of the second fragment set into similar variants of the concerned phrase, and identify the similar variants whose edit distance from the concerned phrase is greater than the predetermined threshold as the concerned phrase.

[0118] Specifically, combining the candidate fragments in each group of the second fragment set into similar variants of the concerned phrase includes: for each group of the second fragment set, combine the respective candidate fragments in the second fragment set according to the start position and the end position to obtain similar variants of the concerned phrase.

[0119] Such as Figure 3As shown, a similar variant of the first concerned phrase "People's Republic of China" is "中2华人民公国", and a similar variant of the second concerned phrase "People's Republic of China" is "中化人43民公国".

[0120] Specifically, identifying a similar variant whose edit distance with the focus phrase is greater than a predetermined threshold as the focus phrase may include: calculating the edit distance between the similar variant and the focus phrase; if the edit distance is less than the predetermined threshold, identifying the similar variant as the focus phrase.

[0121] In this embodiment, the edit distance between the similar variant and the focus phrase may be calculated. If the edit distance is less than a predetermined threshold, the similar variant is identified as the focus phrase; if the edit distance is greater than the predetermined threshold, the similar variant is not identified as the focus phrase.

[0122] In summary, the method for extracting a focus phrase in a text provided by an embodiment of the present application generates an improved multimodal detection model based on the focus phrase, a similar character set of the focus phrase, and an ignored character set, and uses the multimodal matching model to detect the input text to obtain multiple fragments; then, the multiple fragments are grouped using the inverted index of the focus phrase to obtain a first fragment set; then, the first fragment set is segmented according to a preset edit distance to obtain multiple groups of second fragment sets; for each group of second fragment sets, candidate fragments are screened according to the edit distance between the fragments in the second fragment set and the focus phrase; the candidate fragments in each group of second fragment sets are combined into similar variants of the focus phrase, and similar variants whose edit distance with the focus phrase is greater than a predetermined threshold are identified as the focus phrase, thereby improving the accuracy of extraction.

[0123] like Figure 4 As shown, it shows a structural block diagram of a device for extracting a focused phrase in a text provided by an embodiment of the present application, and the device for extracting a focused phrase in a text can be applied to a computer device. The device for extracting a focused phrase in a text can include:

[0124] An acquisition module 410 is used to acquire an inverted index and a multi-mode matching model constructed for a focused phrase, where the multi-mode matching model is generated based on the focused phrase, a similar character set of the focused phrase, and an ignored character set;

[0125] A detection module 420 is used to detect the input text using a multi-mode matching model to obtain multiple fragments;

[0126] A collection module 430 is used to collect multiple fragments by using the inverted index to obtain a first fragment set;

[0127] A segmentation module 440, configured to segment the first fragment set according to a preset edit distance to obtain multiple sets of second fragment sets;

[0128] A screening module 450, configured to screen candidate fragments for each group of second fragment sets according to the edit distance between the fragments in the second fragment sets and a concerned phrase;

[0129] An identification module 460, configured to combine the candidate fragments in each group of second fragment sets into similar variants of the concerned phrase, and identify the similar variants with an edit distance greater than a predetermined threshold from the concerned phrase as the concerned phrase.

[0130] In an optional embodiment, the detection module 420 is further configured to:

[0131] Obtain the i-th character in the input text, where i is a positive integer;

[0132] If the i-th character is an ignored character in the ignored character set, update i to i + 1, and continue to execute the step of obtaining the i-th character in the input text;

[0133] If the i-th character is not an ignored character, find all similar characters of the i-th character in the similar character set, match the i-th character and all similar characters with the trie in the multi-pattern matching model, and generate fragments according to the matching results;

[0134] Wherein, the fragments include phrases extracted from the text, the starting position and the ending position of the phrases in the text, and the phrases matched by the extracted phrases in the trie.

[0135] In an optional embodiment, the screening module 450 is further configured to:

[0136] Select m consecutive fragments from the n fragments of the second fragment set, where 1 ≤ m ≤ n;

[0137] Determine a candidate phrase according to the starting position of the first fragment and the ending position of the m-th fragment among the m fragments, and retain the candidate phrases with an edit distance less than a predetermined threshold from the concerned phrase;

[0138] Among all the obtained candidate phrases, screen the candidate phrases with the longest length and non-overlapping starting positions and ending positions with each other, and determine the fragments corresponding to the candidate phrases as candidate fragments.

[0139] In an optional embodiment, the identification module 460 is further configured to:

[0140] For each group of second fragment sets, combine the respective candidate fragments in the second fragment set according to the starting position and the ending position to obtain similar variants of the concerned phrase.

[0141] In an optional embodiment, the collection module 430 is further configured to:

[0142] Use the inverted index to determine the focus phrases to which each fragment belongs;

[0143] Extracting the first word and the last word from the focus phrase, and supplementing the first word fragment and the last word fragment;

[0144] Among all the obtained fragments, the fragments with the longest length and non-overlapping starting and ending positions are selected to obtain the first fragment set.

[0145] In an optional embodiment, the identification module 460 is further configured to:

[0146] Calculate the edit distance between similar variant domain focus phrases;

[0147] If the edit distance is less than a predetermined threshold, the similar variant is identified as a phrase of interest.

[0148] In an optional embodiment, the acquisition module 410 is further configured to:

[0149] Split the focus phrases and build an inverted index based on the split results;

[0150] Obtaining a similar character set for each character in the focus phrase, the similar character set comprising at least one of a homophone character set, a phonetically similar character set, and a similar-looking character set;

[0151] Obtain an ignored character set, where the ignored character set includes at least one of a number set and a label set, and the label set is a label that needs to be ignored and is set according to an application field;

[0152] A multi-modal matching model is constructed based on similar character sets and ignored character sets.

[0153] In summary, the device for extracting a focus phrase in a text provided by an embodiment of the present application generates an improved multimodal detection model according to the focus phrase, a similar character set of the focus phrase, and an ignored character set, and uses the multimodal matching model to detect the input text to obtain multiple fragments; then uses the inverted index of the focus phrase to group the multiple fragments to obtain a first fragment set; then, the first fragment set is segmented according to a preset edit distance to obtain multiple groups of second fragment sets; for each group of second fragment sets, candidate fragments are screened according to the edit distance between the fragments in the second fragment set and the focus phrase; the candidate fragments in each group of second fragment sets are combined into similar variants of the focus phrase, and similar variants whose edit distance with the focus phrase is greater than a predetermined threshold are identified as the focus phrase, thereby improving the accuracy of extraction.

[0154] An embodiment of the present application provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned method for extracting focus phrases in a text.

[0155] An embodiment of the present application provides a computer device, and the computer device includes a device for extracting a concerned phrase in any of the above texts.

[0156] It should be noted that when the device for extracting a concerned phrase in the text provided in the above embodiment extracts the concerned phrase in the text, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device for extracting a concerned phrase in the text is divided into different functional modules to complete all or part of the functions described above. In addition, the device for extracting a concerned phrase in the text provided in the above embodiment and the embodiment of the method for extracting a concerned phrase in the text belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0157] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.

[0158] The above does not intend to limit the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. A method for extracting focus phrases in a text, characterized in that: The method comprises: Obtaining an inverted index and a multi-mode matching model constructed for a focus phrase, wherein the multi-mode matching model is generated according to the focus phrase, a similar character set of the focus phrase, and an ignored character set; Using the multi-mode matching model to detect the input text, and obtain multiple fragments; The plurality of fragments are grouped by using the inverted index to obtain a first fragment set; Splitting the first fragment set according to a preset edit distance to obtain multiple groups of second fragment sets; For each set of second fragments, screening candidate fragments according to the edit distance between the fragments in the second fragment set and the focus phrase; Combining candidate fragments in each group of the second fragment set into similar variants of the focus phrase, and identifying similar variants whose edit distance with the focus phrase is less than a predetermined threshold as the focus phrase; The method of using the inverted index to group the multiple fragments to obtain a first fragment set includes: using the inverted index to determine the focus phrase to which each fragment belongs; extracting the first character and the last character from the focus phrase, and supplementing the fragment of the first character and the fragment of the last character; among all the obtained fragments, screening the fragments with the longest length and whose starting positions and ending positions do not overlap with each other to obtain the first fragment set.

2. The method for extracting focus phrases in a text according to claim 1, characterized in that: The multi-mode matching model is used to detect the input text to obtain multiple fragments, including: Get the i-th character in the input text, where i is a positive integer; If the i-th character is an ignored character in the ignored character set, then i is updated to i+1, and the step of obtaining the i-th character in the input text is continued; If the i-th character is not the ignored character, searching for all similar characters of the i-th character in the similar character set, matching the i-th character and all similar characters with the dictionary tree in the multi-mode matching model, and generating fragments according to the matching results; The fragments include phrases extracted from the text, the starting position and the ending position of the phrases in the text, and phrases matched in the dictionary tree according to the extracted phrases.

3. The method for extracting focus phrases in a text according to claim 2, characterized in that: The screening of candidate fragments according to the edit distance between the fragments in the second fragment set and the focus phrase comprises: Selecting m consecutive fragments from the n fragments of the second fragment set, 1≤m≤n; Determine candidate phrases according to the starting position of the first fragment and the ending position of the mth fragment among the m fragments, and retain candidate phrases whose edit distance with the focus phrase is less than a predetermined threshold; Among all the obtained candidate phrases, the candidate phrases with the longest length and whose starting positions and ending positions do not overlap with each other are selected, and the fragments corresponding to the candidate phrases are determined as candidate fragments.

4. The method for extracting focus phrases in a text according to claim 2, characterized in that: The step of combining the candidate fragments in each group of the second fragment set into similar variants of the concerned phrase comprises: For each group of second fragment sets, each candidate fragment in the second fragment set is combined according to the starting position and the ending position to obtain a similar variant of the focus phrase.

5. The method for extracting focus phrases in a text according to claim 1, characterized in that: The step of identifying a similar variant having an edit distance from the focus phrase less than a predetermined threshold as the focus phrase comprises: Calculating the edit distance between the similar variants and the phrase of interest; If the edit distance is less than a predetermined threshold, the similar variant is identified as the focus phrase.

6. The method for extracting focus phrases in a text according to any one of claims 1 to 5, characterized in that: The step of obtaining an inverted index and a multi-mode matching model constructed for a focus phrase includes: Splitting the focus phrases, and constructing an inverted index according to the splitting results; Obtaining a similar character set for each character in the focus phrase, wherein the similar character set includes at least one of a homophone character set, a phonetically similar character set, and a similar-looking character set; Obtain an ignored character set, wherein the ignored character set includes at least one of a number set and a label set, and the label set is a label that needs to be ignored and is set according to an application field; A multi-modal matching model is constructed according to the similar character set and the ignored character set.

7. A device for extracting a focus phrase in a text, characterized in that: The device comprises: An acquisition module, used for acquiring an inverted index and a multi-mode matching model constructed for a focus phrase, wherein the multi-mode matching model is generated according to the focus phrase, a similar character set of the focus phrase, and an ignored character set; A detection module, used to detect the input text using the multi-mode matching model to obtain multiple fragments; A collection module, configured to collect the plurality of fragments by using the inverted index to obtain a first fragment set; A segmentation module, configured to segment the first fragment set according to a preset edit distance to obtain multiple groups of second fragment sets; A screening module, configured to screen candidate fragments for each set of second fragments according to the edit distance between the fragments in the second fragment set and the focus phrase; an identification module, configured to combine candidate fragments in each group of the second fragment set into similar variants of the focus phrase, and identify the similar variants whose edit distance with the focus phrase is less than a predetermined threshold as the focus phrase; The collection module is also used to: determine the focus phrase to which each fragment belongs by using the inverted index; extract the first character and the last character from the focus phrase, and supplement the fragment of the first character and the fragment of the last character; among all the obtained fragments, select the fragments with the longest length and whose starting positions and ending positions do not overlap with each other, to obtain a first fragment set.

8. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the method for extracting focus phrases in a text as described in any one of claims 1 to 6.

9. A computer device, characterized in that: The computer device comprises: the device for extracting focus phrases in the text as described in claim 7.

Citation Information

Patent Citations

  • Repeated material entity recognition method based on mutually different feature vectors

    CN112861918A

  • Characterization processing method and device for medical data, equipment, medium and product

    CN116994689A