Method and apparatus for matching text content

Through the method of marker word segmentation and similarity calculation, the problem of low name consistency verification efficiency in the prior art is solved, and efficient and accurate text content matching is achieved.

CN114510933BActive Publication Date: 2025-07-22ALL CHINA MARKETING RES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210036857.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-13
Publication Date
2025-07-22
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

In the prior art, the name consistency verification method that splits a single word for comparison is inefficient, resulting in redundant matching results and requires a lot of manual screening, which reduces the matching efficiency of text content.

Method used

The target text content is segmented using marker word types, natural language processing technology is used to split and analyze words, build index words and generate synonyms, determine the matching text content through similarity comparison, and calculate the similarity value using the weight model to improve matching accuracy.

Benefits of technology

It improves the matching accuracy of text content, reduces the redundancy of matching results, reduces the error of manual screening, and improves matching efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114510933B_ABST
    Figure CN114510933B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for matching text content, relating to the technical field of natural language processing, and mainly aiming to solve the problem of low matching efficiency of existing text content. The method includes: obtaining target text content to be matched; segmenting the target text content according to the type of marked words to obtain a segmentation result, where the type of marked words is used to represent the type of index words to be indexed and matched; if the segmentation result matches the index words, comparing the similarity value between the segmentation result and the index words with a screening similarity threshold, and determining the text content that matches the target text content based on the similarity comparison result, where the index words are constructed based on the comparison text content corresponding to the target text content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method and device for matching text content. Background Art

[0002] With the rapid development of big data technology, more and more application fields need to perform big data management on enterprise data. Especially for objects without a unique identity code, it is usually necessary to use a unique name to identify the identity for relevant data management, such as enterprise name management, paper name management, test question management, etc. For example, in the process of using an enterprise name as the unique identity for relevant business management, it is necessary to perform consistency verification of the name, that is, compare one or more enterprise names with multiple enterprise names in the existing business data to determine the consistency of the enterprise entity, so as to ensure the authenticity of the enterprise identity.

[0003] Currently, the existing consistency verification of names usually splits the name as text content into single characters for one-by-one comparison to determine the consistency of the name's main body. However, splitting the single characters for one-by-one comparison greatly reduces the matching efficiency, resulting in redundant matching results. Moreover, due to the characteristics of word composition, splitting the single characters for comparison also requires a large amount of manual screening, increasing the burden of matching, thus reducing the matching efficiency of text content. Summary of the Invention

[0004] In view of this, the present invention provides a method and device for matching text content, mainly aiming to solve the problem of low matching efficiency of existing text content.

[0005] According to one aspect of the present invention, a method for matching text content is provided, including:

[0006] Obtain the target text content to be matched;

[0007] Segment the target text content according to the marker word type to obtain a segmentation result, where the marker word type is used to represent the type of index word to be indexed and matched, and the index word is constructed based on the comparison text content corresponding to the target text content;

[0008] If the segmentation result matches the index word, then compare the similarity value between the segmentation result and the index word with a screening similarity threshold, and determine the text content matching the target text content based on the similarity comparison result.

[0009] Further, before obtaining the target text content to be matched, the method further includes:

[0010] Obtain the comparison text content and split it according to the types of marked words, where the types of marked words include regional word types, feature range word types, and business form word types;

[0011] Construct an index relationship according to the split words to determine the index words, where the index relationship is used to represent the matching order during index matching;

[0012] Generate a text matching thesaurus that matches the index relationship and the index words, where the text matching thesaurus contains a thesaurus of synonyms corresponding to the index words, so as to perform index matching based on the synonymous words in the thesaurus of synonyms.

[0013] Further, the tokenization of the target text content according to the types of marked words results in:

[0014] Using natural language processing technology, split and parse the words in the target text content according to the types of marked words to determine the types of marked words corresponding to the words;

[0015] Mark the words according to the types of marked words to obtain a tokenization result containing word content that matches the types of marked words.

[0016] Further, after the tokenization of the target text content according to the types of marked words to obtain the tokenization result, it also includes:

[0017] Determine the index words for index matching according to the types of marked words of the word content in the tokenization result, and the thesaurus of synonyms corresponding to the index words;

[0018] Compare the word content with the synonymous words in the thesaurus of synonyms according to the index relationship of the index words;

[0019] If the synonymous word matches the word content, it is determined that the tokenization result matches the index word.

[0020] Further, comparing the similarity value between the tokenization result and the index word with a screening similarity threshold, and determining the text content that matches the target text content based on the similarity comparison result includes:

[0021] Determine the marking parameters of the tokenization result and the weight values of the types of marked words;

[0022] Calculate the similarity value of the tokenization result based on the ratio of the synonymous words corresponding to the index words in the thesaurus of synonyms to the tokenization result, as well as the marking parameters and the weight values;

[0023] If the similarity value is greater than or equal to the screening similarity threshold, the comparison text content for constructing the index term is determined as the text content matching the target text content.

[0024] Further, the determining the weight value of the marker word type includes:

[0025] Obtain the uniqueness parameter, interference parameter, and experience parameter of the comparison text content;

[0026] Based on the weight model function, calculate the weight values of different marker word types corresponding to the uniqueness parameter, the interference parameter, and the experience parameter.

[0027] Further, after obtaining the target text content to be matched, the method further includes:

[0028] Parse at least one target word in the target text content, and obtain at least one comparison word in the comparison text content matching the target text content;

[0029] If the target word is exactly matched with the comparison word, it is determined that the target text content is exactly matched with the comparison text content, and the comparison file content is output as the text content matching the target text content;

[0030] If the target word is not exactly matched with at least one of the comparison words, it is determined that the target text content is not exactly matched with the comparison text content, and the target text content is segmented according to the marker word type.

[0031] According to another aspect of the present invention, there is provided a text content matching device, including:

[0032] An acquisition module for acquiring the target text content to be matched;

[0033] A segmentation module for segmenting the target text content according to the marker word type to obtain a segmentation result, where the marker word type is used to characterize the type of index term to be indexed and matched, and the index term is constructed based on the comparison text content corresponding to the target text content;

[0034] A determination module for, if the segmentation result is matched with the index term, comparing the similarity value between the segmentation result and the index term with the screening similarity threshold, and determining the text content matching the target text content based on the similarity comparison result.

[0035] Further, the device further includes:

[0036] A splitting module, configured to obtain comparison text content and split it according to the types of tagging words, where the types of tagging words include regional word types, feature range word types, and business form word types;

[0037] A building module, configured to build an index relationship according to the split words to determine index words, where the index relationship is used to represent the matching order during index matching;

[0038] A generating module, configured to generate a text matching thesaurus that matches the index relationship and the index words, where the text matching thesaurus contains a thesaurus of synonyms corresponding to the index words, so as to perform index matching based on the synonymous words in the thesaurus of synonyms.

[0039] Further, the word segmentation module includes:

[0040] A splitting unit, configured to use natural language processing technology to split and analyze the words in the target text content according to the types of tagging words to determine the types of tagging words corresponding to the words;

[0041] A tagging unit, configured to tag the words according to the types of tagging words to obtain a word segmentation result containing word content that matches the types of tagging words.

[0042] Further, the device further includes: a comparison module,

[0043] The determining module is further configured to determine the index words for index matching and the thesaurus of synonyms corresponding to the index words according to the types of tagging words of the word content in the word segmentation result;

[0044] The comparison module is configured to compare the word content with the synonymous words in the thesaurus of synonyms according to the index relationship of the index words;

[0045] The determining module is further configured to determine that the word segmentation result matches the index word if the synonymous word matches the word content.

[0046] Further, the determining module includes:

[0047] A first determining unit, configured to determine the marking parameters of the word segmentation result and the weight values of the types of tagging words;

[0048] A calculating unit, configured to calculate the similarity value of the word segmentation result based on the ratio of the synonymous words corresponding to the index words in the thesaurus of synonyms to the word segmentation result, as well as the marking parameters and the weight values;

[0049] A second determination unit, configured to determine the comparison text content for constructing the index word as the text content matching the target text content if the similarity value is greater than or equal to the screening similarity threshold.

[0050] Further, the first determination unit is specifically configured to obtain the uniqueness parameter, interference parameter, and experience parameter of the comparison text content; calculate the weight values corresponding to the uniqueness parameter, the interference parameter, and the experience parameter for different types of marked words based on the weight model function.

[0051] Further, the apparatus further includes: a parsing module, an output module

[0052] The parsing module is configured to parse at least one target word in the target text content and obtain at least one comparison word in the comparison text content matching the target text content;

[0053] The output module is configured to determine that the target text content is completely matched with the comparison text content and output the comparison file content as the text content matching the target text content if the target word is completely matched with the comparison word;

[0054] The determination module is configured to determine that the target text content is not completely matched with the comparison text content to perform word segmentation on the target text content according to the type of marked word if the target word is not completely matched with at least one of the comparison words.

[0055] According to another aspect of the present invention, there is provided a storage medium storing at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the above text content matching method.

[0056] According to still another aspect of the present invention, there is provided a terminal including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0057] The memory is configured to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above text content matching method.

[0058] By means of the above technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:

[0059] The present invention provides a method and apparatus for matching text content. Compared with the prior art, in the embodiments of the present invention, a target text content to be matched is obtained; the target text content is segmented according to the type of marked words, and a segmentation result is obtained. The type of marked words is used to represent the type of index words to be indexed and matched, and the index words are constructed based on the comparison text content corresponding to the target text content; if the segmentation result matches the index words, then the similarity value between the segmentation result and the index words is compared with a screening similarity threshold, and the text content matching the target text content is determined based on the similarity comparison result, greatly improving the matching accuracy of the text content, reducing the redundancy of the text content matching result, avoiding the error of manual screening, reducing the burden of text content matching, and thus improving the matching efficiency of the text content.

[0060] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0062] Figure 1 shows a flowchart of a method for matching text content provided by an embodiment of the present invention;

[0063] Figure 2 shows a flowchart of another method for matching text content provided by an embodiment of the present invention;

[0064] Figure 3 shows a flowchart of yet another method for matching text content provided by an embodiment of the present invention;

[0065] Figure 4 shows a schematic diagram of the implementation based on a matching engine provided by an embodiment of the present invention;

[0066] Figure 5 shows a flowchart of still another method for matching text content provided by an embodiment of the present invention;

[0067] Figure 6 shows a block diagram of a device for matching text content provided by an embodiment of the present invention;

[0068] Figure 7The figure shows a schematic structural diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners

[0069] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0070] For the verification of enterprise name consistency, it is usually to split the name into single characters as text content and compare them one by one to determine the main body consistency of the name. However, splitting the single characters and comparing them one by one greatly reduces the matching efficiency, making the matching results redundant. And due to the characteristics of word composition, splitting the single characters for comparison also requires a large amount of manual screening, increasing the burden of matching, thus reducing the matching efficiency of the text content. An embodiment of the present invention provides a method for matching text content, as Figure 1 shown, the method includes:

[0071] 101. Obtain the target text content to be matched.

[0072] In an embodiment of the present invention, the target text content is used to represent a text object that needs to match whether there is the same text content within a preset text range. The business application scope of the target text content includes, but is not limited to, text contents such as enterprise names, thesis titles, and test questions. For example, if it is in the field of enterprise names, the target text content to be matched can be "Beijing Some Tongren Information Co., Ltd.", and then the method in step 102 is used for word segmentation. Among them, the target text content can be formed in different languages such as Chinese or English, and is a text data containing multiple words, so that the target text content can be split into words. The embodiment of the present invention does not make specific limitations.

[0073] It should be noted that the current execution entity can be a processor, a server, a terminal, a component, etc. for performing text content matching, so as to provide text content matching for users. Therefore, the target text content to be matched in the current execution end can be obtained by receiving the text content input by the user or from options (outputting options through a pre-configured option interface for the user to select).

[0074] 102. Segment the target text content according to the marker word type to obtain a word segmentation result.

[0075] In the embodiments of the present invention, in order to compare the text content with the index terms to determine whether they match, at this time, the target text content is segmented according to the markup terms. Among them, the markup term type is used to represent the type of index terms to be indexed and matched, so as to match with the index terms according to the segmented results. Among them, the markup term type for different business application scopes at least includes regional terms, feature scope terms, and business form terms. For example, for the business application scope of business names, the regional term can be the region where the enterprise name belongs, such as "Beijing", "Liaoning", etc., the feature scope term is the business scope term, such as "technology", "Internet", etc., and the business form can be the organizational form, such as "limited liability company", "firm". The markup term type can also include enterprise names and non-business markup terms. For example, the non-business markup term is used to represent all word contents that cannot be classified, such as a certain person, so as to perform index matching as a type of words. Another example is that for the application scope of thesis titles, the regional term can be the research field scope of the thesis, such as "biology", "information control", the feature scope term is the scientific research innovation scope term, such as "based on neural network...", "based on wind energy...", and the business form can be the thesis system form, such as "research", "analysis", "design". The embodiments of the present invention do not make specific limitations.

[0076] It should be noted that when tokenizing the target text content, it is split by marking the word types. At this time, specifically, based on natural language processing techniques, the target text content can be split according to word features, such as name phrases. At the same time, the words obtained by splitting are integrated according to the marked word types. For example, when splitting according to word features, the words "liability" and "company" are obtained. However, according to the marked word type "limited liability company", "liability" and "company" are integrated to obtain the final result of splitting the target text content, that is, "limited liability company". The embodiments of the present invention do not make specific limitations. At the same time, since the index words are constructed based on the comparison text content that matches the target text content, at this time, the index words are not only represented as a single word, but a thesaurus containing multiple synonymous words constructed based on the comparison text content. That is, by finding the index words, the corresponding thesaurus can be found. This thesaurus contains thesauruses corresponding to different index words, so as to perform one-to-one matching based on the synonymous words in the thesaurus. It is also possible to construct the synonymous relationship between the index words and the words corresponding to different marked word types based on the comparison text content, so as to match between the tokenization results obtained by splitting and the index words. The embodiments of the present invention do not make specific limitations. The comparison text content in the embodiments of the present invention is at least one text object selected from the preset text range to be compared with the target text content to determine whether they match. For example, the target text content to be matched can be "Beijing Some Communication People Information Co., Ltd.", and a comparison text content can be "Some Communication People (Beijing) Research Co., Ltd.". At this time, index words are constructed based on multiple comparison text contents, and then one-to-one matching is performed between the index words and the tokenization results of the target text content. In addition, a text database containing all the comparison text contents is pre-established in the current execution end, so as to serve as the data basis for obtaining the comparison text content. When the target text content to be matched is obtained, at least one comparison text content can be obtained from the text database to construct index words.

[0077] 103. If the tokenization result matches the index word, then compare the similarity value between the tokenization result and the index word with the screening similarity threshold, and determine the text content that matches the target text content based on the similarity comparison result.

[0078] In the embodiments of the present invention, since the word segmentation result contains word contents of different tag word types, therefore, the matching between the word segmentation result and the index word is specifically the matching with the synonyms in the thesaurus corresponding to this index word. At this time, it can also be determined whether there is a match between the word segmentation result and the index word based on the synonym relationship established in advance between the word segmentation result and the index word. The embodiments of the present invention do not make specific limitations. In addition, since the word segmentation result can contain word contents of multiple tag word types, therefore, when performing index matching, the word contents can be matched with the index word in sequence. If each word content matches the index word, it is determined that the word segmentation result matches the index word. In order to obtain an accurate matching result, calculate the similarity value between the word segmentation result and the index word, and then compare it with the screening similarity threshold to determine the final matching result as the target text content. For example, if the similarity value between the word segmentation result and the index word is greater than the screening similarity threshold, the text content that matches the target text content is determined, which is the comparison text content for constructing the index word.

[0079] In another embodiment of the present invention, for further limitation and illustration, as Figure 2 shown, before step 101 of obtaining the target text content to be matched, the method further includes:

[0080] 201. Obtain comparison text content, and split it according to the tag word type based on the comparison text content;

[0081] 202. Build an index relationship according to the split words, and determine the index word;

[0082] 203. Generate a text matching thesaurus that matches the index relationship and the index word.

[0083] In the embodiments of the present invention, in order to implement the matching based on the index word and the word segmentation result obtained by splitting, the index word is pre-constructed based on the comparison text content, so as to perform the matching based on the word segmentation result and the index word. Specifically, obtain at least one comparison text content stored in the current execution entity, and split the comparison text content according to the tag word type. At this time, each word in the comparison text content can be integrated and split according to natural language processing technology, and the implementation method is the same as that for the target comparison text content, and will not be elaborated here. At the same time, establish an index relationship for the words of each tag word type obtained after splitting, and each word split according to the tag word type is used as an index word. Thus, the constructed index relationship is used to represent the matching order during index matching. For example, match in sequence according to the index words corresponding to the region word type, the feature range word type, and the business form word type. The embodiments of the present invention do not make specific limitations.

[0084] It should be noted that, in order to implement the applicable scope of matching based on index words, a text matching thesaurus is generated based on the index relationship and index words. At this time, the text matching thesaurus contains synonym thesauruses corresponding to different index words, so as to perform index matching based on the synonyms in the synonym thesaurus. That is, a text matching thesaurus contains multiple index words determined by splitting the comparison text content, and the synonym thesaurus corresponding to each index word. At this time, the synonym thesaurus includes synonyms matched by different index words, so that when matching the word segmentation result with the index word, if the word segmentation result is the same as the synonym in the synonym thesaurus of the index word, the word segmentation result matches the index word. For example, if the index word is "Beijing", and the synonym thesaurus includes synonyms "Beijing Municipality" and "Capital", then when the word segmentation result is "Beijing Municipality", it matches "Beijing Municipality" in the synonym thesaurus of the index word "Beijing", so the word segmentation result "Beijing Municipality" matches the index word "Beijing" to calculate the similarity value between the word segmentation result "Beijing Municipality" and the index word "Beijing".

[0085] In another embodiment of the present invention, for further limitation and explanation, step 102 performs word segmentation on the target text content according to the marker word type, and the obtained word segmentation result includes:

[0086] Using natural language processing technology, split and analyze the words in the target text content according to the marker word type to determine the marker word type corresponding to the words;

[0087] Mark the words according to the marker word type to obtain a word segmentation result containing word content matching the marker word type.

[0088] In order to implement word segmentation of the target text content according to the marker word type, specifically, based on natural language processing technology, each word in the target text content is split and analyzed according to the marker word type, that is, the words in the target text content can be split and analyzed according to the word characteristics of the determined Chinese phrases. For example, if the target text content is "Beijing Some Tongren Information Co., Ltd.", then according to the regional word type, feature range word type, and business form word type, the words "Beijing", "Some Tongren", "Information", "Limited", and "Company" in "Beijing Some Tongren Information Co., Ltd." are split and analyzed to determine the marker word type corresponding to each word, that is, "Beijing" corresponds to the regional word type, "Information" corresponds to the feature range word type, and the combination of "Limited" and "Company" corresponds to the business form word type. Among them, "Some Tongren" can be used as a non-business marker word, so as to mark the marker word type of each of the above words to obtain the word segmentation result.

[0089] In another embodiment of the present invention, for further limitation and explanation, such as Figure 3As shown, after step 102 tokenizes the target text content according to the tag word type and obtains the tokenization result, it further includes:

[0090] 301. Determine the index words that match the index according to the tag word type of the word content in the tokenization result, and the thesaurus corresponding to the index words;

[0091] 302. Compare the word content with the synonymous words in the thesaurus according to the index relationship of the index words;

[0092] 303. If the synonymous word matches the word content, determine that the tokenization result matches the index word.

[0093] Since the tokenization result includes all the word contents with tag word types obtained after splitting the target text content, in order to achieve the matching between the tokenization result and the index word, specifically, by corresponding the tag word type of each index word and the tag word type in the process of splitting each word content, determine each word content in the tokenization result and the index word for corresponding matching, so as to compare the word content with the thesaurus in the thesaurus of the matching index word according to the index relationship of the index word. For example, match the words "Beijing City" and "Limited Company" included in the tokenization result in sequence according to the index relationship, that is, compare with the synonymous words "Beijing City", "Capital" in the thesaurus of index word 1 "Beijing", and the synonymous words "Limited Company" in the thesaurus of index word 2 "Limited Liability Company" respectively in sequence, that is, "Beijing City" in the tokenization result matches index word 1 "Beijing", and "Limited Company" matches index word 2 "Limited Company", as Figure 4 shown in the matching process.

[0094] It should be noted that since the tag word type can include a non-business tag word method to cover the tag types of words other than the region word type, feature range word type, and business form word type, therefore, when constructing the index word, it can be constructed based on the pre-configured tag word type, so as to further limit the words used to identify and distinguish identities when matching the tokenization result with the index word, so as to improve the matching accuracy of the text content.

[0095] In another embodiment of the present invention, for further limitation and explanation, step 103 compares the similarity value between the tokenization result and the index word with the screening similarity threshold, and determines the text content that matches the target text content based on the similarity comparison result, including:

[0096] Determine the marking parameters of the tokenization result and the weight value of the tag word type;

[0097] Calculate the similarity value of the word segmentation result based on the ratio of the synonymous words corresponding to the index words in the thesaurus to the word segmentation result, as well as the marking parameter and the weight value.

[0098] If the similarity value is greater than or equal to the screening similarity threshold, determine the comparison text content for constructing the index word as the text content matching the target text content.

[0099] Since there can be multiple word segmentation results and multiple index words, when the word segmentation result matches the index word, it means that each word segmentation result matches the corresponding index word. At this time, in order to further determine the text content matching the target text content, calculate the similarity between the word segmentation result and the index word, and then compare it with the screening similarity threshold to determine the text content corresponding to the target text content based on the comparison result. Specifically, determine the marking parameter corresponding to the word segmentation result and the weight value corresponding to the marking type to calculate the similarity value. Among them, the marking parameter K is whether it matches the marker in the marker word type after word segmentation. If it matches, K = 1; if it does not match, K = 0; the weight value W is the calculation weight value corresponding to the marker word type, and the similarity value rank is calculated based on the similarity calculation function.

[0100] Among them, the similarity calculation function is , where n is the number of word contents and index words in the word segmentation result, K is the marking parameter, W is the weight value, and M is the ratio of the synonymous words corresponding to the index words in the thesaurus to the word segmentation result. For example, the percentage of the enterprise name hitting the synonymous words in the thesaurus can be calculated according to the ratio of the number of matches with the synonyms. The embodiments of the present invention do not make specific limitations.

[0101] It should be noted that since a similarity value can be calculated between each word content in the word segmentation result and the corresponding index word, at this time, the screening similarity threshold is a preset similarity threshold for screening word contents, which can be multiple different similarity thresholds for matching index words or a similarity threshold. The embodiments of the present invention do not make specific limitations. When the similarity value is greater than or equal to the screening similarity threshold, determine the comparison text content for constructing the index word as the text content matching this target text content. Preferably, the screening similarity threshold is 60% - 90%.

[0102] In another embodiment of the present invention, for further limitation and explanation, the step of determining the weight value of the marker word type includes: obtaining the uniqueness parameter, interference parameter, and experience parameter of the comparison text content; calculating the weight values of different marker word types corresponding to the uniqueness parameter, the interference parameter, and the experience parameter based on the weight model function.

[0103] Specifically, when calculating the similarity value, the weight value of the marker word type can be determined based on the uniqueness parameter, the interference parameter, and the empirical parameter to perform accurate matching of the text content. Among them, the uniqueness parameter f(X0) is the weighted average value of the proportion of words with the same marker word type but different word contents, and the interference parameter f(X1) is the weighted average value of the proportion of words with the same marker word type and the same word contents. At this time, the weighted average value is used to represent the distribution weighted value of the words that only appear once in the target text content to be matched. The empirical parameter Z is a preset artificial empirical adjustment parameter value to calculate the weight values of different marker word types corresponding to the uniqueness parameter, the interference parameter, and the empirical parameter based on the weight model function. Among them, the weight model function Weight is expressed as Weight = f(X0) - f(X1) + Z, and the embodiments of the present invention do not make specific limitations.

[0104] In another embodiment of the present invention, for further limitation and illustration, as Figure 5 shown, after step 101 obtains the target text content to be matched, the method further includes:

[0105] 401. Analyze at least one target word in the target text content and obtain at least one comparison word in the comparison text content that matches the target text content;

[0106] 402. If the target word completely matches the comparison word, determine that the target text content completely matches the comparison text content, and output the comparison file content as the text content matched by the target text content;

[0107] 403. If the target word does not completely match at least one of the comparison words, determine that the target text content does not completely match the comparison text content, and perform word segmentation on the target text content according to the marker word type.

[0108] To improve the matching efficiency and accuracy of text content, before tokenizing the target text content according to the token type, each word in the target text content can be first matched one by one with each word in the comparison text content to determine whether there is a perfect match. If there is an imperfect match, the step of tokenizing according to the token type is executed. Among them, the target text content is parsed into words through natural language processing technology to obtain at least one target word, and at the same time at least one comparison word in the comparison text content is obtained. At this time, the comparison word can be obtained by pre-parsing the comparison text content into words, and then in the way of one-by-one comparison and matching, it is judged whether the target word is exactly the same as the comparison word. If the target word and the comparison word are perfectly matched, it means that the target text content and the comparison text content are matched, and the content of the comparison file can be output as the text content matched with the target text content.

[0109] An embodiment of the present invention provides a method for matching text content. Compared with the prior art, in the embodiment of the present invention, the target text content to be matched is obtained; the target text content is tokenized according to the token type to obtain a tokenization result, and the token type is used to represent the type of index words to be indexed and matched, and the index words are constructed based on the comparison text content corresponding to the target text content; if the tokenization result matches the index words, then the similarity value between the tokenization result and the index words is compared with a screening similarity threshold, and based on the similarity comparison result, the text content matched with the target text content is determined, which greatly improves the matching accuracy of the text content, reduces the redundancy of the matching results of the text content, and avoids the error of manual screening, reduces the burden of text content matching, and thus improves the matching efficiency of the text content.

[0110] Further, as an implementation of the above Figure 1 shown method, an embodiment of the present invention provides a text content matching device, as Figure 6 shown, the device includes:

[0111] An acquisition module 51, configured to acquire the target text content to be matched;

[0112] A tokenization module 52, configured to tokenize the target text content according to the token type to obtain a tokenization result, and the token type is used to represent the type of index words to be indexed and matched;

[0113] A determination module 53, configured to, if the tokenization result matches the index words, compare the similarity value between the tokenization result and the index words with a screening similarity threshold, and determine the text content matched with the target text content based on the similarity comparison result, where the index words are constructed based on the comparison text content corresponding to the target text content.

[0114] Further, the device further includes:

[0115] A splitting module, configured to obtain the comparison text content and split it according to the marker word types, where the marker word types include regional word types, feature range word types, and business form word types;

[0116] A construction module, configured to construct an index relationship according to the split words and determine index words, where the index relationship is used to represent the matching order during index matching;

[0117] A generation module, configured to generate a text matching thesaurus that matches the index relationship and the index words, where the text matching thesaurus includes a synonym thesaurus corresponding to the index words, so as to perform index matching based on the synonymous words in the synonym thesaurus.

[0118] Further, the word segmentation module includes:

[0119] A splitting unit, configured to use natural language processing technology to split and analyze the words in the target text content according to the marker word types and determine the marker word types corresponding to the words;

[0120] A marking unit, configured to mark the words according to the marker word types to obtain a word segmentation result including the word content matching the marker word types.

[0121] Further, the device further includes: a comparison module,

[0122] The determination module is further configured to determine the index words for index matching and the corresponding synonym thesaurus according to the marker word types of the word content in the word segmentation result;

[0123] The comparison module is configured to compare the word content with the synonymous words in the synonym thesaurus according to the index relationship of the index words;

[0124] The determination module is further configured to determine that the word segmentation result matches the index words if the synonymous words match the word content.

[0125] Further, the determination module includes:

[0126] A first determination unit, configured to determine the marking parameters of the word segmentation result and the weight values of the marker word types;

[0127] A calculation unit, configured to calculate the similarity value of the word segmentation result based on the ratio of the synonymous words corresponding to the index words in the synonym thesaurus to the word segmentation result, as well as the marking parameters and the weight values;

[0128] A second determination unit, configured to determine the comparison text content for constructing the index term as the text content matching the target text content if the similarity value is greater than or equal to the screening similarity threshold.

[0129] Further, the first determination unit is specifically configured to obtain the uniqueness parameter, interference parameter, and experience parameter of the comparison text content; calculate the weight values corresponding to the uniqueness parameter, the interference parameter, and the experience parameter for different types of marked words based on the weight model function.

[0130] Further, the apparatus further includes: a parsing module, an output module.

[0131] The parsing module is configured to parse at least one target word in the target text content and obtain at least one comparison word in the comparison text content matching the target text content.

[0132] The output module is configured to determine that the target text content and the comparison text content are completely matched and output the comparison file content as the text content matching the target text content if the target word and the comparison word are completely matched.

[0133] The determination module is configured to determine that the target text content and the comparison text content are not completely matched to perform word segmentation on the target text content according to the type of marked word if the target word and at least one of the comparison words are not completely matched.

[0134] An embodiment of the present invention provides a text content matching apparatus. Compared with the prior art, in the embodiment of the present invention, by obtaining the target text content to be matched; performing word segmentation on the target text content according to the type of marked word to obtain a word segmentation result, where the type of marked word is used to represent the type of index word to be indexed and matched, and the index word is constructed based on the comparison text content corresponding to the target text content; if the word segmentation result matches the index word, comparing the similarity value between the word segmentation result and the index word with the screening similarity threshold, and determining the text content matching the target text content based on the similarity comparison result, the matching accuracy of the text content is greatly improved, the redundancy of the text content matching result is reduced, the error of manual screening is avoided, the burden of text content matching is reduced, and thus the matching efficiency of the text content is improved.

[0135] According to an embodiment of the present invention, there is provided a storage medium storing at least one executable instruction, and the computer executable instruction can execute the text content matching method in any of the above method embodiments.

[0136] Figure 7The structural schematic diagram of a terminal provided according to an embodiment of the present invention is shown. The specific embodiments of the present invention do not limit the specific implementation of the terminal.

[0137] As Figure 7 shown, the terminal may include: a processor 602, a communications interface 604, a memory 606, and a communication bus 608.

[0138] Among them: the processor 602, the communications interface 604, and the memory 606 communicate with each other through the communication bus 608.

[0139] The communications interface 604 is used to communicate with network elements of other devices such as clients or other servers.

[0140] The processor 602 is used to execute the program 610, and specifically can execute the relevant steps in the matching method embodiment of the above text content.

[0141] Specifically, the program 610 may include program code, and the program code includes computer operation instructions.

[0142] The processor 602 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the terminal may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0143] The memory 606 is used to store the program 610. The memory 606 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0144] The program 610 is specifically used to cause the processor 602 to perform the following operations:

[0145] Obtain the target text content to be matched;

[0146] Segment the target text content according to the marker word type to obtain a segmentation result, where the marker word type is used to characterize the type of index word to be indexed and matched;

[0147] If the word segmentation result matches the index word, compare the similarity value between the word segmentation result and the index word with a screening similarity threshold, and determine the text content that matches the target text content based on the similarity comparison result. The index word is constructed based on the comparison text content corresponding to the target text content.

[0148] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0149] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for matching text content, characterized in that, Including: Obtain the target text content to be matched. Before obtaining the target text content to be matched, the method further includes: obtaining comparison text content, and splitting it according to the types of marker words, where the types of marker words include regional word types, feature range word types, and business form word types; constructing an index relationship according to the split words to determine index words, where the index relationship is used to represent the matching order during index matching; generating a text matching thesaurus that matches the index relationship and the index words, and the text matching thesaurus contains a synonym thesaurus corresponding to the index words, so as to perform index matching based on the synonymous words in the synonym thesaurus. Segment the target text content according to the types of marker words to obtain a segmentation result, including: splitting the target text content according to word features based on natural language processing technology, and then integrating the split words according to the types of marker words. The types of marker words are used to represent the types of index words to be subjected to index matching, and the index words are constructed based on the comparison text content corresponding to the target text content. If the segmentation result matches the index word, then compare the similarity value between the segmentation result and the index word with a screening similarity threshold, and determine the text content that matches the target text content based on the similarity comparison result, including: determining the marking parameters of the segmentation result and the weight values of the types of marker words. The weight values of the types of marker words are determined based on uniqueness parameters, interference parameters, and empirical parameters. Among them, the uniqueness parameter is the weighted average of the proportion of words with the same type of marker word but different word contents, and the interference parameter is the weighted average of the proportion of words with the same type of marker word and the same word content; calculate the similarity value of the segmentation result based on the ratio of the synonymous words corresponding to the index words in the synonym thesaurus to the segmentation result, the marking parameters, and the weight values; if the similarity value is greater than or equal to the screening similarity threshold, then determine the comparison text content for constructing the index word as the text content that matches the target text content.

2. The method according to claim 1, wherein After segmenting the target text content according to the types of marker words to obtain a segmentation result, the method further includes: Determine the index words for index matching and the corresponding synonym thesaurus according to the types of marker words of the word contents in the segmentation result. Compare the word contents with the synonymous words in the synonym thesaurus according to the index relationship of the index words. If the synonymous words match the word contents, then determine that the segmentation result matches the index word.

3. The method according to claim 1, wherein The determination of the weight values of the types of marker words includes: Obtain the uniqueness parameter, interference parameter, and empirical parameter of the comparison text content. Calculate the weight values of different types of marker words corresponding to the uniqueness parameter, the interference parameter, and the empirical parameter based on a weight model function.

4. The method according to any one of claims 1-3, characterized in that After obtaining the target text content to be matched, the method further includes: Parse at least one target word in the target text content, and obtain at least one comparison word in the comparison text content that matches the target text content; If the target word exactly matches the comparison word, determine that the target text content exactly matches the comparison text content, and output the comparison text content as the text content matching the target text content; If the target word does not exactly match at least one of the comparison words, determine that the target text content does not exactly match the comparison text content, and perform word segmentation on the target text content according to the marker word type.

5. A matching device for text content, characterized in that, Includes: An acquisition module for acquiring the target text content to be matched. Before acquiring the target text content to be matched, the device further includes: acquiring the comparison text content, and splitting it according to the marker word type. The marker word type includes the regional word type, the feature range word type, and the business form word type; constructing an index relationship according to the split words, determining the index word, and the index relationship is used to represent the matching order during index matching; generating a text matching thesaurus that matches the index relationship and the index word, and the text matching thesaurus contains a synonym thesaurus corresponding to the index word, so as to perform index matching based on the synonymous words in the synonym thesaurus; A word segmentation module for performing word segmentation on the target text content according to the marker word type to obtain a word segmentation result, including: splitting the target text content according to the word features based on natural language processing technology, and then integrating the split words according to the marker word type. The marker word type is used to represent the type of the index word to be subjected to index matching, and the index word is constructed based on the comparison text content corresponding to the target text content; A determination module for, if the word segmentation result matches the index word, comparing the similarity value between the word segmentation result and the index word with a screening similarity threshold, and determining the text content matching the target text content based on the similarity comparison result, including: determining the marker parameter of the word segmentation result, and the weight value of the marker word type, and the weight value of the marker word type is determined based on the uniqueness parameter, the interference parameter, and the experience parameter. The uniqueness parameter is the weighted average of the proportion of words with the same marker word type but different word contents, and the interference parameter is the weighted average of the proportion of words with the same marker word type and the same word contents; calculating the similarity value of the word segmentation result based on the ratio of the synonymous words corresponding to the index word in the synonym thesaurus to the word segmentation result, and the marker parameter and the weight value; if the similarity value is greater than or equal to the screening similarity threshold, determine the comparison text content constructing the index word as the text content matching the target text content.

6. A storage medium storing at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the text content matching method according to any one of claims 1-4.

7. A terminal, comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the text content matching method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Standard text matching method and device, storage medium and electronic equipment

    CN112541051A